Source-linked AI summary

Seeing biodiversity: perspectives in machine learning for wildlife conservation

Devis Tuia, Benjamin Kellenberger, Sara Beery, Blair R. Costelloe, Silvia Zuffi, Benjamin Risse, Alexander Mathis, Mackenzie W. Mathis, Frank van Langevelde, Tilo Burghardt, Roland Kays, Holger Klinck, Martin Wikelski, Iain D. Couzin, Grant van Horn, Margaret C. Crofoot, Charles V. Stewart, Tanya Berger-Wolf

arXiv:2110.12951v1cs.LGcs.CV

TL;DR

Animal ecology must analyze rapidly growing sensor data despite limited traditional monitoring and processing capacity. The paper synthesizes how machine learning can scale wildlife monitoring and conservation, while emphasizing interdisciplinary collaboration and boundaries from dataset bias, ethics, computational cost, and sensor deployment.

  • Problem

    Animal ecology needs accurate population and behavioral information from increasingly large sensor datasets, but conventional observation and analysis remain limited in scale and efficiency.

  • Method

    The paper presents success stories and opportunities at the interface of machine learning, new-generation sensors, and ecological workflows.

  • Results

    The paper reports performance improvements and existing successful applications of machine learning and sensors for animal ecology, including MegaDetector, AIDE, DeepLabCut, and Wildlife Insights.

  • Takeaways & Limitations

    Effective wildlife conservation applications require ecological expertise, machine-learning expertise, strict quality control, and cross-disciplinary training.

  • Takeaways & Limitations

    Applications remain constrained by geographic, sensor, species, and sampling biases, ethical risks in data sharing, computational costs, and sensor-specific deployment limitations.

Abstract

from arXiv · show

Data acquisition in animal ecology is rapidly accelerating due to inexpensive and accessible sensors such as smartphones, drones, satellites, audio recorders and bio-logging devices. These new technologies and the data they generate hold great potential for large-scale environmental monitoring and understanding, but are limited by current data processing approaches which are inefficient in how they ingest, digest, and distill data into relevant information. We argue that machine learning, and especially deep learning approaches, can meet this analytic challenge to enhance our understanding, monitoring capacity, and conservation of wildlife species. Incorporating machine learning into ecological workflows could improve inputs for population and behavior models and eventually lead to integrated hybrid modeling tools, with ecological models acting as constraints for machine learning models and the latter providing data-supported insights. In essence, by combining new machine learning approaches with ecological domain knowledge, animal ecologists can capitalize on the abundance of data generated by modern sensor technologies in order to reliably estimate population abundances, study animal behavior and mitigate human/wildlife conflicts. To succeed, this approach will require close collaboration and cross-disciplinary education between the computer science and animal ecology communities in order to ensure the quality of machine learning approaches and train a new generation of data scientists in ecology and conservation.

New sensors expand available data types for animal ecology

New sensors broaden wildlife observation across space, time, taxa, and data types, while machine learning helps process the resulting large and heterogeneous datasets. Different sensor platforms offer complementary strengths but also introduce constraints involving coverage, disturbance, generalization, labeling, and computational scale.

  • Sensor categories: Sensor data span imagery, soundscapes, positional data, and physiological or behavioral measurements collected by stationary, mobile, and on-animal devices.These platforms support monitoring individuals, species, habitats, activity, and behavior across spatial and temporal scales.
  • Stationary sensors: Stationary sensors provide close-range, continuous monitoring for presence, identification, behavior, and predator-prey analysis, but their observations are strongly correlated in space and time.Placement and animal habits limit which species are captured and reduce independence among observations from the same sensor.
  • Stationary sensors: Camera traps are inexpensive and widely deployed, while machine learning can filter blank images and support annotation and species prediction at scale.Production-ready accuracy remains limited by poor generalization across geographies, acquisition periods, and sensor types.
  • Stationary sensors: Bioacoustic monitoring is cost-effective and less affected by light and weather than camera traps, but long-term datasets can exceed terabyte scale and require noise-robust, generalizable models.Large, diverse labeled datasets remain unavailable for many animal groups and confounding signals.
  • Remote sensing: Remote sensing from animals, drones, aircraft, and satellites expands monitoring coverage and enables movement, habitat, detection, counting, posture, and individual-identification analyses.UAVs offer agile, relatively inexpensive acquisition, whereas satellites provide broad coverage but often insufficient resolution for direct wildlife observation.
  • Remote sensing: Drones remain constrained by battery capacity, legislation, cost, and possible wildlife disturbance, whose severity depends on species, equipment, and flight characteristics.Lower-altitude flights can affect group and individual mammal behavior, while fixed-wing platforms trade greater footprint for higher cost and less close access.
  • Community science: Community science increases image availability and geographic, pose, and viewing-angle diversity, but volunteer annotations can contain errors and require substantial verification.Open platforms such as Wildbook combine community or social-media imagery with computer vision for species classification and individual identification.

Machine learning to scale-up and automate animal ecology and conservation research

Machine learning tools translate increasingly abundant sensor data into ecological information, supporting automated detection, identification, reconstruction, and biodiversity analysis. These applications can scale monitoring and provide non-invasive insights, while remaining constrained by data and reconstruction challenges.

  • From sensors to ecological information: Sensor imagery can be converted into abundance maps, individual re-identification, herd tracking, and three-dimensional environmental or phenotypical reconstructions.These outputs are described as animal biometrics and form usable information for ecological research.
  • From sensors to ecological information: Computer vision maps aerial imagery to animal localization, movement tracking, landscape photogrammetry, posture estimation, behavior inference, and individual re-identification.The figure establishes a shared vocabulary between ecology tasks and corresponding computer-vision tasks.
  • Automated detection and identification: Deep-learning and computer-vision methods support automatic individual identification, including facial recognition of chimpanzees from a million-image dataset.The chimpanzee study reported greater than 90% accuracy and substantially faster identification than manual human re-identification.
  • Behavior and reconstruction: Three-dimensional shape recovery and pose estimation provide non-invasive information about animal health, age, reproductive status, posture, kinematics, and behavior.Marker-less deep-learning methods and user-friendly toolboxes have improved extraction of animal postures from video.
  • Behavior and reconstruction: Environmental reconstruction from wildlife imagery remains difficult because large video datasets, appearance ambiguity, drift, critical configurations, and camera motion can reduce reconstruction quality.External sensor cues are usually integrated to achieve satisfactory reconstructions from video data.
  • Biodiversity modeling: Machine learning improves biodiversity analyses by reducing errors in species-richness prediction and outperforming generalized linear models for plant-pollinator interactions.Tree-based methods rank features, select covariates, and model nonlinear relationships in species-distribution tasks.

Attention points and opportunities

The paper identifies opportunities for integrating machine learning with ecology, alongside requirements for responsible deployment, domain knowledge, quality control, and interdisciplinary collaboration.

  • Opportunities: Machine and deep learning are presented as accelerators for wildlife research and conservation, with opportunities for genuinely integrated approaches across disciplines.
  • Responsible use: Responsible use requires attention to geographic, sensor, species, and sampling biases, because models may be accurate only within specific conditions.The paper calls for transparency about training data and intended model usage to avoid potentially catastrophic accuracy drop-offs.
  • Responsible use: Open ecological datasets require careful curation, annotation, evaluation metrics, code maintenance, data attribution, and protection against harmful disclosure.The paper specifically warns that shared animal locations can expose wildlife to poachers and drone paths can reveal ranger locations.
  • Responsible use: Model errors and unquantified uncertainty can produce erroneous or catastrophic conservation decisions, making quality control and accountability essential.The paper also identifies laboratories as development spaces for tools later adopted in field studies of free-moving animals.
  • Responsible use: Machine learning carries substantial financial, hardware, cloud-computing, and energy costs for conservation organizations.The paper notes that cloud costs can reach several thousand dollars per month for a single virtual machine.
  • Involving domain knowledge from the start: Hybrid models can incorporate ecological knowledge and fundamental laws into machine-learning systems rather than relying solely on correlations from generic black-box models.
  • Involving domain knowledge from the start: 17.9% higher mean species identification precision was reported for Context R-CNN on Snapshot Serengeti after integrating image features over time.The model uses prior knowledge about slowly varying camera-trap backgrounds and low-frequency sampling with dropouts.
  • Towards a new generation of biodiversity models: GeoLifeCLEF remains an open biodiversity-modeling challenge, with 1.9 million observations covering over 31,000 species and only approximately 26% top-30 accuracy by the 2021 winners.The task provides positive observations without absence data and predicts likely species for geospatial grid cells.

Conclusions

The paper argues that machine learning and deep learning can scale ecological studies toward global understanding of the animal world. Its success stories also show that broader adoption requires close cooperation, quality control, and interdisciplinary development of hybrid and large-scale habitat models.

  • Animal ecology and wildlife conservation need scalable analysis of growing data streams to estimate populations, understand behavior, and address poaching and biodiversity loss.
  • Machine learning and deep learning are proposed as tools for scaling local ecological studies toward global understanding of the animal world.
  • The Perspective synthesizes success stories, performance improvements, existing tools, attention points, and future research directions at the interface of machine learning and animal ecology.
  • Further progress depends on cooperation between ecologists and machine-learning specialists, strict quality control, detailed design knowledge, hybrid models, and habitat-distribution models at scale.

Glossary

The glossary defines core machine-learning and data-science terminology used throughout the paper, including artificial intelligence, big data, classification, computer vision, convolutional neural networks, and data science.

  • Artificial intelligence is defined as the concept of a machine performing higher-level semantic reasoning, while big data refers to analysis volumes too large for conventional hardware.
  • Classification assigns an entire image or video to a single category, and computer vision performs image manipulation and understanding tasks, often using machine learning.
  • A convolutional neural network is a deep-learning model containing at least one convolution layer whose shared filter weights reduce required neurons and provide limited translation invariance.
  • Data science is presented as an interdisciplinary or multidisciplinary research field associated with big-data analysis.
Loading 2110.12951v1…