Source-linked AI summary

Decoding Children's Gait Behavior

Yifan Shen, Boyi Li, Meihuan Huang, Yuanzhe Liu, Xu Cao, Jinyang Jin, Zhengyuan Li, Anglin Liu, Junho Kim, Jingyuan Zhu, Lan Fangzhou, Jianguo Cao, Jintai Chen, Ismini Lourentzou, James Matthew Rehg

arXiv:2608.00371v1cs.CV

TL;DR

Fine-grained pediatric gait assessment matters clinically, but existing visual and sensor-based approaches struggle with subtle, phase-sensitive abnormalities. The paper introduces the CGV dataset and ChildGait-Video framework, achieving 70%–93% agreement with expert annotations across all scoring items and a 0.83 average F1 score.

  • Problem

    Pediatric gait assessment is clinically important, yet existing evaluation approaches face practical constraints and subjective, variable expert interpretation.

  • Method

    The paper introduces the multi-view CGV dataset and ChildGait-Video, combining RGB video with pose, segmentation, and spatiotemporal modeling for EVGS inference.

  • Results

    70%–93% agreement with expert annotations across all scoring items and a 0.83 average F1 score are achieved for typical-versus-atypical distinction.

  • Takeaways & Limitations

    CGV establishes a benchmark for pediatric gait research, while ChildGait-Video demonstrates automated scoring from standard RGB camera setups.

  • Takeaways & Limitations

    The study provides model-level validation, with future work needed on open-scene data from homes, rehabilitation centers, and community environments.

Abstract

from arXiv · show

We introduce a new problem domain for human action recognition: the fine-grained analysis of children's gait behaviors from standard RGB video. We specifically target the ambulatory patterns of children aged 3-17 years. Such behaviors arise naturally in the diagnosis and treatment of several critical developmental and neuromuscular disorders, such as cerebral palsy and hemiplegia. Despite their clinical value, current 3D sensor-based gait analysis systems are expensive, intrusive, and often impractical for young subjects. To address this, we introduce a new dataset comprising over 1,100 high-frame-rate (60 FPS) video sequences from 110 subjects, accompanied by synchronized, anonymized pose sequences. In each session, the child performs a 5-second "walk-around" task, capturing the gait cycle from multiple viewpoints. Crucially, we demonstrate that current state-of-the-art approaches, including gait foundation models and Multimodal Large Language Models (MLLMs), fail to effectively resolve these clinical nuances. We identify the key technical challenges in analyzing these erratic and subtle motor patterns and describe a unified end-to-end framework for decoding fundamental components of pediatric gait. Through comprehensive experimental results, we demonstrate the potential of this dataset to drive novel research questions and establish a rigorous baseline for automated child gait assessment.

1 Introduction

This work introduces pediatric gait analysis from standard RGB video as a clinically motivated computer-vision challenge and presents the Children Gait Video (CGV) dataset to support it. CGV combines multi-view child gait videos, clinical EVGS assessments, and frame-level annotations, while evaluations show that current VLMs are ineffective at decoding subtle gait abnormalities.

  • Motivation: Pediatric gait analysis is clinically important, but existing systems can be expensive, space-intensive, intrusive, and difficult to deploy broadly.Physical-space requirements can bias adoption toward well-resourced clinical sites, while body-attached sensors may be poorly tolerated or induce reactivity.
  • Motivation: Computer vision could enable reliable, rich, and non-intrusive gait measurement using only one or two camera views, supporting scalable early screening.The approach could bring gait assessment into everyday clinical settings without encumbering children.
  • Problem: Adult-centric gait foundation models fail to capture the high entropy, high intra-class variance, and inconsistent motion patterns of developing motor systems.Existing computer-vision gait methods largely focus on healthy adults and assume mature, stable, highly periodic gait.
  • Dataset: Over 1,100 video sessions from 110 child subjects form the Children Gait Video (CGV) dataset, covering 3–17-year-olds across anterior, posterior, and lateral walking views.The standardized protocol uses brief 3–5 second sequences.
  • Dataset: CGV pairs expert-clinician Edinburgh Visual Gait Score annotations with frame-level instance masks, bounding boxes, and anatomical keypoints tailored to developing bodies.The dataset supports both clinical assessment and detailed visual analysis of pediatric gait.
  • Evaluation: A comprehensive evaluation finds that out-of-the-box open and closed VLMs are ineffective at decoding subtle signs of gait abnormalities from video.The evaluation specifically examines inferring EVGS from children’s videos.

2 Related Work

Related work spans the evolution of gait datasets, advances in visual gait modeling, and clinical approaches for screening and assessment. These efforts range from identity-focused benchmarks and deep spatiotemporal representations to quantitative clinical gait metrics.

  • Human Gait Datasets: Gait datasets have evolved from general biometric recognition toward specialized health and clinical applications.Foundational CASIA and OU-ISIR databases advanced appearance-invariant and multi-view identity recognition, while SUSTech1K and CMU MoBo expanded visual gait analysis.
  • Gait Vision Modeling: Visual gait modeling combines silhouette, pose, skeleton, graph, video, Transformer, and foundation-model approaches to address appearance, viewpoint, and temporal challenges.Pose-based methods pair 2D keypoint detectors with 3D pose lifting or triangulation, while spatiotemporal modeling uses graph convolutions, video architectures, diffusion, and large vision pipelines.
  • Children’s Gait Screening and Assessment: Clinical gait screening extracts actionable developmental and diagnostic information from motion signals using spatiotemporal parameters and kinetic ground reaction forces.Reported parameters include velocity, cadence, step length, and double support time, alongside generalized scales such as FGA and UPDRS.

3 Challenges in Pediatric Gait Analysis

Pediatric gait analysis from 2D video presents challenges that differ from adult-centric datasets, including developmental skeletal variation, subtle kinematic patterns, and occlusion from complex interactions. These challenges also motivate collaboration between computer vision and clinical researchers on early motor-development questions.

  • Technical challenges: Children’s gait modeling faces a substantial domain shift because pediatric limb proportions and centers of mass do not scale linearly from adult data.The passage identifies anthropometric differences between pediatric and adult populations as a core technical challenge.
  • Technical challenges: Subtle gait features, such as maximum knee flexion during swing, create demanding fine-grained analysis requirements.The supplied passage specifically cites maximum knee flexion during swing as an example of a gait feature requiring analysis.
  • Technical challenges: Physical guidance or encouragement from parents and clinicians produces visual clutter and frequent inter-person occlusions that challenge tracking, segmentation, and gait-modeling pipelines.Young children may require assistance during walking, creating multi-person interactions that complicate automated video analysis.
  • Clinical motivation: Computer vision and clinical collaboration can investigate whether subtle gait deviations indicate autism spectrum disorder and whether BCI technologies improve gait kinematics.The passage frames pediatric gait analysis as an intersection for studying early motor development and related clinical questions.

4 Children Gait Video (CGV) Dataset

The CGV dataset is an open-sourced resource for children’s gait modeling, combining high-resolution videos with clinically grounded, fine-grained EVGS annotations and abnormality diagnoses. Its collection protocol, expert validation, privacy safeguards, and patient-level splitting support clinically valid and leakage-resistant evaluation.

  • Dataset composition: CGV is the first open-sourced children’s gait video dataset, containing 339,236 frames from 1,185 2K videos with gait-abnormality diagnoses.Diagnoses include cerebral palsy, traumatic brain injury, developmental dysplasia of the hip, toe-in, and idiopathic toe walking.
  • Data collection: The two-camera setup records coronal and sagittal views along an 8-meter walkway, with the lateral camera covering the middle four meters.The calibrated viewpoint is designed to capture 2–3 complete gait cycles per subject.
  • Clinical annotations: Each evaluation covers 2 × 17 EVGS scoring items across the left and right limbs, with annotations adhering to the EVGS reference guide.The dataset captures 2–3 complete gait cycles per subject and uses a three-point ordinal scale for each parameter.
  • Expert validation: ICC = 0.93 and 93.8% average scoring accuracy demonstrate strong agreement between pediatric experts and provide a human annotation baseline.An additional pediatrician completed a separate annotation round and was evaluated as a human baseline.
  • Ethical considerations: The dataset received retrospective IRB approval under approval number 202106202 and follows a multi-layered protocol for participant protection and data security.The protocol addresses privacy for pediatric facial, pose, and behavioral data.
  • Experimental setup: Experiments use a strict Patient ID split, assign each patient exclusively to training or test data, balance positive and negative samples, and set the train-to-test patient ratio to 6:1.The test set is randomly selected under this object-level split.

5 Benchmarking Children’s Gait Analysis

The benchmark evaluates zero-shot multimodal large language models on uniformly sampled video frames for children’s gait analysis. All tested models fail to reliably resolve the clinical nuances of subtle pediatric gait patterns.

  • Baselines: The benchmark evaluates Gemini 3 Pro, GPT-5.2, Qwen3-VL-235B, GLM4.6V, Qwen3.5-9B, and InternVL3-8B as zero-shot multimodal LLM baselines.Each model receives T = 16 uniformly sampled frames from the center window of downsampled 30 FPS video and the prompt P.
  • Experimental Results: 50% to 60% average accuracy per limb leaves all listed MLLMs only marginally above random guessing on binary scoring tasks.The results show that these models do not effectively resolve the required clinical nuances.
  • Experimental Results: The performance gap reflects difficulty capturing subtle kinematic deviations and fine-grained temporal dynamics critical for clinical gait analysis.This limitation persists despite the models’ strong general visual–language reasoning abilities.

6 Decoding Children’s Gait via Vision-Language Models

This section introduces Qwen3-VL-ChildGait, a fine-tuned multimodal framework for mapping pediatric gait videos to fine-grained CGV scores. Despite multimodal inputs and task-specific fine-tuning, the model shows limited overall improvement, indicating that VLMs struggle with subtle child-gait recognition.

  • Framework: Qwen3-VL-ChildGait fine-tunes Qwen3-VL-4B to map pediatric gait patterns to all fine-grained CGV items.The framework formulates gait analysis as a sequence-to-label prediction task.
  • Multimodal Representation: Each frame combines RGB imagery, skeletal keypoints, and instance segmentation masks to expose joint angles and global body morphology.The composite representation is Ft = {It, Kt, Mt}, enabling attention to fine-grained and global visual evidence while filtering clinical background noise.
  • Clinical Goal-Oriented Fine-tuning: Training downsamples videos to 30 FPS, samples four temporal windows with 16 frames each, and minimizes negative log-likelihood of categorical gait-score tokens.Testing uses 16 frames from a single central temporal window, with prompts directing the model to synthesize visual trajectories into gait-quality ratings.
  • Experimental Results: 2% and 1%: fine-tuning increases average accuracy for the left and right limbs by these amounts, respectively, versus zero-shot Qwen3-VL-4B.The reported averages show no significant overall improvement over the base model.
  • Experimental Results: +11.5% for Initial Contact in Stance (IC), but -1.5% for Varus/Valgus in Stance (HVV), producing mixed item-level effects after fine-tuning.The counterintuitive pattern supports the conclusion that VLMs do not perform well on subtle child-gait recognition.

7 Decoding Children’s Gait via Video-based Models

ChildGait-Video is an end-to-end video model that maps visual features directly to EVGS items, using kinematic prompts and foreground token pruning to capture pediatric gait patterns. It outperforms gait-analysis and VideoMAE v2 baselines, reaching 70%–93% accuracy across scoring items and an average F1-score of 0.83.

  • Model Design: ChildGait-Video directly maps video features to EVGS items, avoiding pose-angle pipelines that discard appearance information and accumulate multi-stage errors.It uses a VideoMAE v2 vision encoder followed by an MLP prediction head to model global motion patterns and interdependencies.
  • Model Design: Skeletal keypoints provide token-level kinematic prompts that encode local anatomical topology alongside original appearance features.The prompts are rendered onto RGB frames before ViT patch embedding, bridging the gap between generic motion representations and pediatric anatomical priors.
  • Model Design: Mask-guided token dropping removes background patches, focuses self-attention on the child’s gait, and reduces fine-tuning computation.The retained foreground patch count satisfies N_foreground ≪ N_total.
  • Experimental Results: 69% and 72% average accuracies: fine-tuned VideoMAE v2 surpasses prior gait baselines by 15% and 19% over BiggerGait for the left and right limbs.Prior gait models achieve average accuracies ranging from 45% to 56%, with SwinGait reaching 54% for the left limb and GaitBase reaching 56% for the right limb.
  • Experimental Results: 70% ∼93% accuracy: ChildGait-Video achieves SoTA performance across all 34 scoring items, improving average accuracy by 15% and 12% over fine-tuned VideoMAE v2.Its average F1-score is 0.83, and the improvement is statistically significant with p = 2.3 × 10−4 < 0.001.
  • Temporal Analysis: 90.7% performance: increasing the temporal input from 8 to 32 frames improves results, while the gain beyond 16 frames is only 3.7%.Sixteen frames capture critical gait-cycle sub-phases, whereas longer windows increase computational complexity quadratically.

8 Discussion

The CGV dataset and ChildGait-Video framework aim to automate pediatric gait evaluation and establish a comprehensive visual benchmark. The framework achieves 70%–93% agreement with expert annotations, while future work targets unconstrained, multi-site deployment for broader generalization and longitudinal tracking.

  • Motivation: CGV aims to democratize pediatric healthcare by automating clinical gait evaluation beyond expert-dependent visual inspection and toward in-the-wild environments.Traditional 2D tools such as the Edinburgh Visual Gait Score rely heavily on expert visual inspection.
  • Framework: ChildGait-Video establishes a specialized multi-stage pipeline combining instance segmentation, anatomical pose estimation, and spatiotemporal modeling.The pipeline is introduced as part of the first comprehensive visual benchmark for children’s gait analysis.
  • Results: 70%–93% agreement with expert annotations is achieved across all items for typical and atypical distinctions.This result is reported for the proposed framework’s empirical evaluation.
  • Future work: Future work will extend the framework to open-scene data from homes, rehabilitation centers, and community environments.The proposed expansion is intended to address broader pathological gait patterns and unconstrained acquisition scenarios.
  • Future work: Wild deployment could enable reproducible longitudinal tracking and cross-site standardization for large-scale outcome evaluation and data-driven rehabilitation research.These capabilities are described as essential for scaling evaluation and rehabilitation research.

9 Conclusion · Appendix

The paper establishes fine-grained pediatric gait understanding from standard RGB video as a new computer vision problem and introduces CGV, a multi-view dataset with synchronized anonymized pose, segmentation, and clinically grounded EVGS-derived annotations. Benchmarking shows that contemporary zero-shot MLLMs and fine-tuned VLMs remain unreliable for phase-sensitive clinical gait scoring.

  • 9 Conclusion: The work introduces fine-grained pediatric gait understanding from standard RGB videos as a new computer vision problem.
  • 9 Conclusion: CGV is presented as a large-scale multi-view dataset for pediatric gait understanding.
  • 9 Conclusion: CGV includes synchronized anonymized pose data.
  • 9 Conclusion: CGV includes segmentation annotations.
  • 9 Conclusion: CGV includes clinically grounded EVGS-derived annotations.
  • 9 Conclusion: Systematic benchmarking finds contemporary zero-shot MLLMs unreliable for phase-sensitive clinical gait scoring.
  • 9 Conclusion: Systematic benchmarking finds fine-tuned VLMs unreliable for phase-sensitive clinical gait scoring.
  • 9 Conclusion: The paper proposes ChildGait-Video to address the challenges identified in pediatric gait understanding.

A CGV Details · B EVGS Scoring Criteria · C Prompt Design

The paper details the CGV dataset demographics, defines EVGS as a 17-item three-point ordinal assessment for each limb, and introduces structured prompts for MLLM-based gait analysis. The prompts combine expert role specification, explicit evaluation instructions, embedded EVGS criteria, and strict JSON output formatting.

  • A CGV Details: The CGV dataset comprises 110 pediatric patients representing the target clinical population.The dataset is intended to provide demographic coverage for visual gait analysis.
  • A CGV Details: Participants span 2.6–16.6 years, with mean age µ = 8.40 and σ = 3.14 years.The age distribution is visualized with a fitted normal curve and covers a clinically relevant developmental window.
  • A CGV Details: The dataset contains 58.2% male and 41.8% female subjects, yielding a nearly balanced gender distribution.This composition is described as supporting diversity and fairness in training and evaluation.
  • B EVGS Scoring Criteria: EVGS evaluates 17 observational gait items for each limb using coronal and sagittal video recordings on a three-point ordinal scale.Scores are categorized as Normal, Moderate deviation, or Marked deviation.
  • C Prompt Design: The structured prompt architecture contains System Persona, Task Instructions, and Output Formatting.This organization instructs MLLMs to analyze children’s gait using the defined clinical scoring workflow.
  • C Prompt Design: The system persona assigns the model an expert pediatrician role specializing in observational gait analysis.This initialization is intended to activate knowledge of anatomical priors, gait-cycle phases, and kinematic deviations.
  • C Prompt Design: Task instructions require systematic assessment of both coronal and sagittal recordings for 17 items per limb, while output formatting requires strict JSON.The prompt embeds the exact EVGS criteria and requests structured results for evaluation and inspection.

D Evaluation of Skeleton Baselines

The evaluation tests SkeletonGait++ and ScoNet as skeleton-based baselines for children's gait analysis. Despite their strength in standard pose and gait recognition, these methods struggle with consistent clinical scoring across joints, while ScoNet pretrained on Scoliosis1K performs best.

  • The evaluation compares two strong skeleton-based baselines, SkeletonGait++ and ScoNet, for children's gait analysis.
  • 58%–65% average accuracy is achieved for the left limb and 60%–64% for the right limb by general skeleton-based baselines.These results reflect sub-optimal consistency in precise clinical scoring across all joints.
  • 68% and 70% average accuracies are the peaks achieved by ScoNet pretrained on Scoliosis1K.

E Module Ablation Study

The module ablation study shows that TKP and MPP improve pediatric gait recognition, while random masking degrades performance. TKP supplies anatomical priors, and MPP more effectively removes background noise while preserving structural integrity than a bounding box mask.

  • Token-Level Kinematic Prompting (TKP): TKP raises average accuracy (L-AVG/R-AVG) from 69%/72% to 72%/74%.The module provides anatomical priors from expert annotations, directing the network toward clinically relevant joint dynamics rather than generic spatial features.
  • Mask-Guided Patch Pruning (MPP): The bounding box mask achieves 75%/77% for L/R-AVG by removing some background.This strategy provides a moderate gain but is outperformed by MPP.
  • Mask-Guided Patch Pruning (MPP): MPP outperforms the bounding box mask by 3% and 2% on the left and right limbs, respectively.Its advantage comes from adaptively filtering background noise while preserving structural integrity.
  • Random Mask Strategy: Random masking drops performance by 1% to 2% compared with baseline, reaching 68%/70%.Random masking can obscure crucial kinematic joints or inadequately eliminate background noise.
Loading 2608.00371v1…