Source-linked AI summary

Markerless tracking of user-defined features with deep learning

Alexander Mathis, Pranav Mamidanna, Taiga Abe, Kevin M. Cury, Venkatesh N. Murthy, Mackenzie W. Mathis, Matthias Bethge

arXiv:1804.03142v1cs.CVq-bio.NCq-bio.QMstat.ML

TL;DR

Researchers need efficient, nonintrusive ways to quantify specific animal movements without manually labeling videos or using predetermined reflective markers. The paper introduces transfer-learning-based markerless tracking and shows that a few hundred labeled frames can achieve performance comparable to human labeling across laboratory behaviors.

  • Problem

    Quantifying specific aspects of animal behavior is time-consuming manually and can require intrusive, predetermined reflective markers.

  • Method

    The framework adapts pretrained deep neural network feature detectors through transfer learning to track user-defined body parts from a few hundred labeled video frames.

  • Results

    Across mouse and Drosophila behaviors, the framework achieved body-part tracking comparable to human labeling after training on only a few hundred images.

  • Takeaways & Limitations

    The method converts videos into semantically meaningful, low-dimensional pose sequences that support behavioral clustering and analysis and may enable real-time feedback.

  • Takeaways & Limitations

    Performance can degrade on test images that differ substantially from the training set, and random selection of very small training sets introduces generalization variability.

Abstract

from arXiv · show

Quantifying behavior is crucial for many applications in neuroscience. Videography provides easy methods for the observation and recording of animal behavior in diverse settings, yet extracting particular aspects of a behavior for further analysis can be highly time consuming. In motor control studies, humans or other animals are often marked with reflective markers to assist with computer-based tracking, yet markers are intrusive (especially for smaller animals), and the number and location of the markers must be determined a priori. Here, we present a highly efficient method for markerless tracking based on transfer learning with deep neural networks that achieves excellent results with minimal training data. We demonstrate the versatility of this framework by tracking various body parts in a broad collection of experimental settings: mice odor trail-tracking, egg-laying behavior in drosophila, and mouse hand articulation in a skilled forelimb task. For example, during the skilled reaching behavior, individual joints can be automatically tracked (and a confidence score is reported). Remarkably, even when a small number of frames are labeled ($\approx 200$), the algorithm achieves excellent tracking performance on test frames that is comparable to human accuracy.

I. INTRODUCTION

The introduction motivates markerless animal pose estimation as a less invasive and more flexible alternative to reflective markers and model-based methods. It presents transfer learning with deep neural networks as achieving human-level accuracy with minimal training data.

  • Motivation: Accurate behavioral quantification is essential for understanding the brain, and videography can reveal previously unknown movement features.The introduction situates behavioral measurement within a broader tradition of using technology to study movement.
  • Limitations of existing methods: Reflective markers enable accurate animal pose tracking but are expensive, potentially distracting, and require feature locations to be predefined before recording.Markers reduce video’s low invasiveness and constrain which body parts can be tracked.
  • Limitations of existing methods: Skeleton and active-contour models can be effective and fast, but they require sophisticated models that are difficult to develop and fit, limiting flexibility.
  • Approach: The study applies state-of-the-art human-limb detection methods to animal pose estimation in laboratory settings.
  • Approach: ≈200 training images can train the DeeperCut feature-detector architecture to achieve human-level labeling accuracy through transfer learning.The architecture is described as one of the best-performing pose-estimation algorithms, and transfer learning enables minimal-data training.

II. RESULTS

The results focus on feature detectors derived from Deep Residual Neural Networks rather than the full DeeperCut system. Trail-tracking evaluation used human-labeled frames and compared predicted labels with human labels during training and testing.

  • Feature detectors: The study evaluates ResNet-based feature detectors with readout layers that predict body-part locations, distinguishing them from the full DeeperCut system.Full DeeperCut performance requires training on thousands of labeled images; the results instead focus on its feature-detector subset.
  • Trail-tracking evaluation: 1,080 frames were used for trail-tracking evaluation, with 80% used for training across three splits.The evaluation included snout, ears, and tail-base labels applied by a human.
  • Trail-tracking evaluation: Trail-tracking performance was assessed using cross-entropy loss and root mean square error between human and predicted labels on training and test images.RMSE was evaluated every 50,000 training iterations for three data splits.

A. Benchmarking

DeepLabCut accurately tracked multiple mouse body parts during odor-guided navigation despite challenging video conditions, achieving human-level test error with limited training data. Performance remained strong across reduced training sets, showed few systematic outliers, and improved slightly with deeper networks.

  • Dataset and labeling: 1,080 frames from videos of 7 mice and 2 cameras were manually labeled for the snout, ears, and tail base.The dataset included frames from multiple videos and all four tracked body parts.
  • Human annotation variability: 2.69 ± 0.1 pixels was the average within-labeler variability across 4,320 body-part image pairs.This variability was smaller than the low-resolution mouse snout width.
  • Tracking accuracy: 80%/20% training/test splits produced test RMSE matching average human variability across tracked body parts.DeepLabCut was evaluated on all body parts and on snout/tail subsets using held-out test images.
  • Training-set efficiency: Less than 5 pixels of average error was achieved with only 10% of images used for training, despite increasing error with fewer training images.Test RMSE declined only slowly from 80% to 10% training fractions.
  • Error robustness: Human and algorithm predictions produced only a few outliers, with no systematic error detected across images.The comparison used one split with a 50% training-set size.
  • Network architecture: ResNet-101 with intermediate supervision achieved 2.88 ± 0.06 pixels versus 3.09 ± 0.04 for ResNet-50 on identical 50% splits.ResNet-101 without intermediate supervision achieved 2.90 ± 0.09 pixels.

B. Generalization & transfer learning

DeepLabCut generalized across novel mice and transferred from single-mouse training to detecting multiple body parts across multiple mice. This transfer supports extensions such as multi-mice tracking for social-behavior studies.

  • Generalization to novel mice: DeepLabCut generalized to novel mice during odor-trail tracking.
  • Transfer learning: A network trained only on single-mouse images detected multiple body parts across multiple mice within the same frame.
  • Transfer to multi-mice tracking: Single-mouse-trained feature detectors readily transferred to multi-mice tracking, enabling extensions useful for social-behavior studies.

C. The power of end-to-end training

The architecture shares a deep ResNet preprocessing network while using body-part-specific deconvolution layers, potentially enabling localization of one body part from other labeled parts. This hypothesis was tested using networks trained only on snout/tail data, alongside examples of Drosophila predictions on held-out frames and varied postures.

  • Architecture: DeeperCut shares the ResNet preprocessing network across body parts but uses body-part-specific deconvolution layers.This architecture was hypothesized to facilitate localization of one body part based on other labeled body parts.
  • Drosophila tracking: DeepLabCut body-part predictions closely match human annotator labels on held-out Drosophila frames and generalize across postures and orientations.The figure describes examples from frames excluded from training and from a fly not included in the training data.

D. Drosophila in 3D behavioral chamber

DeepLabCut tracked 12 user-defined features on freely behaving fruit flies in a 3D chamber using 589 labeled frames from six animals. Training on 95% of the data yielded low test error and robust generalization across flies, orientations, and backgrounds.

  • Experimental setup: The tracked features captured freely behaving flies exploring a cubical environment with an agar-based egg-laying substrate.Flies frequently changed orientation and moved along walls and the ceiling, altering the appearance of body features.
  • Experimental setup: Researchers labeled 12 body points across 589 frames from six flies, including head, thorax, leg, abdominal, and ovipositor features.Only features visible in each frame were labeled, covering diverse orientations and postures.
  • Tracking performance: 4.17 ± 0.32 pixels test error was achieved after training on 95% of the data, compared with 1.39 ± 0.01 pixels average training error.Values are reported as mean ± s.e.m. across n=3 splits.
  • Tracking performance: Generalization to flies excluded from training was excellent, with feature detectors remaining robust to changes in orientation and background.The paper illustrates these results with example test frames and network-applied labels.

E. Digit Tracking During Reaching

DeepLabCut enabled markerless tracking of individual mouse-hand digits during skilled reaching, despite complex articulations and intrusive-marker limitations. With 141 training frames, it achieved low test error and supported confidence-aware analysis of digit trajectories and hand postures.

  • Motivation: Markerless tracking addresses the difficulty and intrusiveness of placing markers on the mouse’s small, complex hand during reaching.Tracking is challenging because of complex hand articulations, the other hand in the background, and varying image statistics.
  • Digit labeling: The network labeled 13 points per frame across four visible digits and the wrist, including digit tips, joints, and the hand base.Each visible digit contributed three points: the tip, mid-digit joint, and base joint, approximately corresponding to the PIP and MCP joints.
  • Tracking performance: 141 training frames yielded an average test error of 5.21 ± 0.28 pixels and an average training error of 1.16 ± 0.03 pixels.The errors are reported as mean ± s.e.m.
  • Confidence estimation: DeeperCut produces point locations from probability fields representing the likelihood that a body part occupies each image region.This enables each point estimate to be accompanied by a confidence measure of label location.
  • Downstream analyses: Extracted body-part locations support digit trajectories across reaches, comparisons of movement patterns, dimensionality reduction of hand postures, and semantically defined skeletons.The digit trajectory example uses frame-by-frame prediction without temporal filtering.

III. DISCUSSION

The discussion presents DeepLabCut as an efficient, transferable framework for markerless posture tracking that can be tailored with limited labeling and applied across experimental settings. It also emphasizes careful dataset curation, confidence-guided refinement, and strategies for occlusions and temporal information.

  • Contributions: Transfer learning enables deep learning models to be adapted efficiently to laboratory tracking tasks with reduced training-data requirements.The authors first estimated human labeling accuracy and then demonstrated the deep architecture’s performance for odor-guided navigation.
  • Dataset labeling and fine-tuning: A training set should be error-free, visually diverse, consistently labeled, and carefully curated to prevent large errors on unusual test images.Recommended diversity includes poses, individuals, luminance conditions, and cameras.
  • Dataset labeling and fine-tuning: Confidence estimates support iterative dataset expansion and post-hoc fine-tuning by identifying informative or problematic frames.Sequences can be sampled around high-probability points, while low probabilities may also signal occlusions rather than errors.
  • Speed and accuracy of DeepLabCut: DeepLabCut converts large videos into computationally tractable, low-dimensional behavioral time series and performs pose extraction rapidly on modern hardware.Users pre-select body parts expected to provide the most information about the behavior.
  • Speed and accuracy of DeepLabCut: Temporal filtering may improve performance in some contexts, but the study focuses on high-precision frame-by-frame prediction from within-frame visual information.The authors note that end-to-end video architectures face hardware and dimensionality challenges.
  • Contributions: The open-source toolbox supports frame extraction, training-data generation, network training, and feature localization, with application beyond mice and Drosophila.A typical workflow involves a few hours of labeling and a day of network training before applying the model to novel videos.

IV. METHODS

The methods apply transfer-learning-based deep feature detectors to diverse animal-behavior videos, then evaluate predictions with body-part RMSE and human variability. An open-source toolbox supports frame selection, annotation checks, network training, and automatic labeling of novel videos.

  • Behavioral datasets: The study analyzes mouse odor trail-tracking, mouse reach-and-pull joystick behavior, and Drosophila egg-laying behavior using labeled video frames.Odor-tracking data included 1 080 frames from 7 mice, while the joystick task included 159 labeled frames across four digits.
  • Deep feature detector architecture: The feature detectors use a 50-layer ResNet initialized with ImageNet weights and optimize score-map cross-entropy with stochastic gradient descent.The architecture builds on DeeperCut-style body-part detectors and state-of-the-art residual-network features.
  • Deep feature detector architecture: Around 500 000 training steps were sufficient for convergence, with training taking 24-36h on an NVIDIA GTX 1080 Ti GPU.Training used batch size 1, a decreasing learning rate, and image-rescaling augmentation.
  • Evaluation and error measures: Predicted body-part locations are obtained from score-map peaks and refined using learned correspondences between score-map grids and ground-truth joint positions.For multiple mice, local score-map maxima provide body-part predictions.
  • Evaluation and error measures: Performance is quantified with pairwise body-part Euclidean distance, reported as RMSE, while human variability is defined as the RMSE between first and second annotations.The analysis also reports 95% confidence intervals using mean ±1.96 times the standard error of the mean.
  • Implementation: The accompanying open-source Python toolbox supports training tailored networks from labeled images and automatically labeling novel video data.It includes tools for selecting training frames, checking annotations, generating training data, and evaluating test-frame performance.

V. SUPPLEMENTARY MATERIALS

Supplementary materials document the labeled dataset and evaluate deeper network architectures. The materials also direct readers to video demonstrations of the tracking system.

  • Video demonstrations: Supplementary materials related to the main figures include a link to video demonstrations of the tracking.The demonstrations are hosted at the Mouse Motor Lab website.
  • Labeled dataset: The labeled dataset comprised 1 080 random frames with human-applied labels illustrating typical image variability.Figure S1 also reports average RMSE in pixels across all body parts above the images.
  • Network architectures: Deeper ResNet-101 and ResNet-101ws architectures strongly reduced training error and modestly improved test performance without overfitting.The comparison used three splits with 50% training-set size; added part-loss layers characterize ResNet-101ws.
Loading 1804.03142v1…