Source-linked AI summary
Beyond similarity: A network approach for identifying and delimiting biogeographical regions
Daril A. Vilhena, Alexandre Antonelli
TL;DR
Bioregion delimitation has been limited by similarity methods and by biases in real species-occurrence data. The paper applies a community-detection approach to associational networks that captures higher-order presence–absence patterns, and reports more detailed, qualitatively different regions that often align with recognised biogeographical boundaries. The authors conclude that network methods provide useful tools for objective classification and delimitation, although environmental comparisons must remain post-hoc because SDM variables are already used in delineation.
Problem
Similarity approaches cannot objectively identify the optimal number of bioregions or realistic, reproducible boundaries, while occurrence data remain geographically and taxonomically biased.
Method
The paper abstracts species distributions as bipartite occurrence networks and applies community detection to higher-order presence–absence relationships among species and localities.
Results
Network analyses produced quantitatively and qualitatively different regionalisations, recovering more expert-recognised bioregions and boundaries than similarity-based comparisons across empirical and hypothetical datasets.
Takeaways & Limitations
Network methods offer tools to classify, delimit, and better understand biodiversity and hold potential to improve identification of the world’s bioregions.
Takeaways & Limitations
Comparisons between inferred bioregions and environmental gradients must be post-hoc because those environmental variables are already major components of SDMs.
Abstract
from arXiv · showhide
Biogeographical regions (geographically distinct assemblages of species and communities) constitute a cornerstone for ecology, biogeography, evolution and conservation biology. Species turnover measures are often used to quantify biodiversity patterns, but algorithms based on similarity and clustering are highly sensitive to common biases and intricacies of species distribution data. Here we apply a community detection approach from network theory that incorporates complex, higher order presence-absence patterns. We demonstrate the performance of the method by applying it to all amphibian species in the world (c. 6,100 species), all vascular plant species of the USA (c. 17,600), and a hypothetical dataset containing a zone of biotic transition. In comparison with current methods, our approach tackles the challenges posed by transition zones and succeeds in identifying a larger number of commonly recognised biogeographical regions. This method constitutes an important advance towards objective, data derived identification and delimitation of the world's biogeographical regions.
I. INTRODUCTION
Bioregions are important units for macroecology, evolution, climate-change research, and conservation, but their classification and delimitation remain controversial. Similarity-based approaches rely on species turnover measures that can be distorted by geographic, environmental, taxonomic, and sampling biases, motivating a network-based alternative.
- Motivation: Bioregions support analyses of niche conservatism, climate-change vulnerability, resilience, and conservation priorities across taxa.They can help assess lineage responses to ecological barriers and target protection toward threatened bioregions rather than individual species.
- Conceptual challenges: Terminology and classification remain unsettled, with multiple overlapping names and internally hierarchical systems used for biogeographical regions.The WWF-adopted terrestrial system recognises 8 realms, 14 biomes, and 867 ecoregions.
- Existing approaches: Similarity measures such as Jaccard, Sørenson, and β-similarity quantify relationships between regions through shared-species proportions.These measures are based theoretically on beta diversity and compare shared species with the total species in the regions.
- Existing approaches: Turnover measures can be biased by geographic distance, environmental gradients, competitive exclusion, spatial clustering, taxonomic sampling, and incomplete distributional information.Consequently, the number of shared species cannot always be trusted as a gauge for identifying bioregions.
- Study aim: The study introduces a data-driven associational-network method and tests it on hypothetical data, global amphibians, and U.S. vascular plants.The datasets differ in geographic scale, grain size, and sampling methodology, while the reported delimitations are congruent with expert-generated regions.
A. Amphibians
For global amphibians, the network represents localities and species together, allowing higher-order occurrence relationships to reveal structure beyond direct similarity. It identifies more detailed quantitative and qualitative regional patterns, including major biogeographic boundaries and distinct subregions.
- Network approach: The occurrence network links localities through the species they contain, while broad cluster separations represent realms and differently coloured groups represent bioregions.Widespread species create cross-realm links between otherwise separated clusters.
- Amphibian results: 10 zoogeographic realms and 55 bioregion-level regions were the optimal representation of the full amphibian dataset.A similarity analysis based on pβsim identified 19 bioregions as optimal.
- Amphibian results: The network and similarity approaches differed both in the number of regions identified and in major biogeographic boundaries.The comparison therefore concerns regional resolution as well as the qualitative placement of boundaries.
- Amphibian results: The network method detected both Wallace’s and Weber’s lines, with Weber’s line emerging as the major boundary between Oceanian and Oriental faunas.It also revealed Sulawesi and intervening islands as distinct subregions of the Oriental realm.
B. Vascular plants
In U.S. vascular plants, similarity clustering produced partitions dominated by state boundaries and weak biogeographic structure, whereas the network approach produced more numerous, broadly coherent clusters across plant datasets. Dataset differences and coarse sampling still constrained fine-scale delimitation.
- Similarity approach: Similarity clustering selected 11 clusters for all plants, 22 for trees, and 14 for non-trees, with limited biogeographic structure.The partitions were dominated by rigid state boundaries and failed to distinguish several recognised regions, including the Everglades, Pacific Coast, and Rocky Mountains.
- Similarity approach: With 40 non-optimal similarity clusters, some deeper structure appeared, but unrealistic boundaries and state-level biases remained.The Great Plains and Southwest desert became more visible in some datasets, while state-boundary artefacts persisted.
- Network approach: The network analysis used county–species occurrence links and a two-level map equation because the pilot analysis revealed little hierarchical structure.The authors associate the limited hierarchy with coarse county grain and low spatial scale.
- Network results: 25 clusters for all plants, 19 for trees, and 16 for non-trees were selected as optimal by the map equation.The stochastic partitioning algorithm was run 1000 times, retaining the partition with the minimum map-equation score.
- Network results: Network results broadly agreed across plant datasets but differed in the visibility of the Everglades, West Coast, Pacific Northwest, and Great Plains.These differences may reflect biological traits as well as sampling issues.
- Limitations: Large counties obscure finer boundaries in the southern Arizona deserts, while state-level biases are evident in Louisiana tree data.The Louisiana bias does not appear in the other two plant datasets.
C. Hypothetical dataset
The hypothetical transition-zone dataset exposes how clustering choices can absorb or isolate transitional areas. The network method instead assigns the transition cells their own clusters while retaining the northern and southern faunal zones.
- Similarity approach: Simpson similarity with UPGMA engulfed the transition zone in the Northern realm with two clusters, but isolated it as a distinct cluster with three.Because the data are symmetric, swapping matrix rows causes the transition zone to be engulfed by the Southern realm.
- Network approach: The network method selected four clusters: Southern fauna with grid cells 1–14, Northern fauna with grid cells 17–30, and separate clusters for cells 15 and 16.This partition was slightly preferred over a two-cluster solution that evenly divided the data into two biogeographic zones.
III. DISCUSSION
Network analyses differed from similarity methods both in the number of regions identified and in their boundaries. Across amphibian, plant, and hypothetical datasets, the approach recovered finer biogeographical structure while exposing biases in similarity-based delineations.
- The network method produced quantitative and qualitative differences from species-similarity approaches, affecting regional resolution, areas, and boundaries.
- Global amphibian bioregions: Many expert-based regions were recovered for amphibians, including the Brazilian Cerrado, Atlantic forest, and Guianan highlands.
- Global amphibian bioregions: The inferred Amazonian–Andean boundary better matched commonly recognised topographic, climatic, and evolutionary distinctions than the similarity-based delimitation.
- United States plant bioregions: Similarity clustering of United States plants identified few biogeographical regions at its optimum and was strongly biased by political state boundaries.
- United States plant bioregions: County-level occurrence data introduced richness bias because county sizes differ substantially, while similarity outputs also reflected state-level aggregation artifacts.
- Hypothetical transition zone: For the hypothetical transition zone, network clustering isolated the transition cells as their own cluster without species, whereas similarity clustering either absorbed or isolated them depending on cluster number.
- The authors argue that biogeographic methods should accommodate geographically and taxonomically biased occurrence data rather than rely on potentially circular SDM-based delineations.
IV. METHODS
The method models species distributions as a network linking species and grid cells, then hierarchically partitions the network into bioregions and realms.
- Species and grid cells are hierarchically classified together into clusters representing bioregions and realms.
- The workflow first constructs a network meaningful for biogeographic analysis, then applies network clustering algorithms to produce bioregions.
A. Delimiting bioregions with networks
The paper represents species occurrences as a bipartite association network, allowing shared occurrences and higher-order presence-absence patterns to inform bioregion delineation.
- A bipartite network links species nodes to locality nodes, with no links between nodes within the same set.
- 2-paths quantify species co-occurrences and shared species between pairs of localities, supporting comparisons beyond direct similarity.
- Higher-order paths capture joint occurrences and complex presence-absence patterns, including partially overlapping species ranges.
- The adjacency matrix encodes species occurrences, while its square yields taxon co-occurrences and locality-level shared-species counts.
- The binary presence-absence matrix B records taxa by localities, forming the network’s species–locality incidence structure.
B. Clustering the bipartite species network
The method clusters the bipartite occurrence network with the map equation, using a random-walk communication objective to identify bioregions and hierarchical realms.
- The map equation is extended to bipartite networks to cluster species and grid-cell associations.
- Its random walk repeatedly moves from a grid cell to a species and then to another cell within that species’ geographic range.
- Strong biogeographic structure keeps the walk within bioregions for long intervals, with crossings mainly through cross-bioregion species.
- The map equation balances detail against communication length, favoring shorter visit descriptions when biogeographic structure is strong.
- Hierarchical map-equation partitions can reveal realms and nested bioregions.
C. Method validation and performance
The study validates its network method on global amphibians, U.S. vascular plants, and a hypothetical transition-zone dataset, while benchmarking against βsim-based similarity clustering.
- c. 6,100 amphibian species from the IUCN database provide a globally distributed empirical test using native range shapefiles.
- 22,918 native vascular plant taxa, corresponding to 17,600 species across 50 states and 3143 U.S. counties, provide a challenging benchmark.
- The hypothetical dataset contains a transition zone where widespread Northern and Southern biota species co-occur, testing cluster-selection pitfalls.
- The benchmark computes βsim between county assemblages because it is considered less sensitive to species-richness differences than conventional measures such as Jaccard.
- UPGMA converts county βsim distances into a hierarchical dendrogram, with the optimum cluster number selected at the evaluation curve’s knee.
V. CONCLUSIONS
The study argues that species distribution data and network methods can improve objective bioregion identification, while emphasizing unresolved data, methodological, theoretical, and phylogenetic challenges.
- Robust bioregionalisation requires objective identification, delimitation, and nomenclature of global bioregions.
- Species distribution data alone hold substantial unrealised potential for these bioregionalisation goals.
- Phylogenetic turnover measures require robust, well-sampled species-level phylogenies, which are lacking for many organismal groups.
- Major remaining challenges include the quantity and quality of occurrence data, methodological integration, and ecological theory linking bioregion patterns to their causes.
- Network methods offer tools to classify, delimit, and better understand biodiversity as increasingly available data support finer bioregion delineation.