Source-linked AI summary
SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation
Soroush Mehraban, Xin Lei Lin, Vida Adeli, Majid Mirmehdi, Amirhossein Dadashzadeh, Clint Hansen, Andrea Iaboni, Babak Taati
TL;DR
Monocular gait estimation is limited by small, capture-specific datasets and restricted visual variation. The paper introduces SynthGait-19K and Gait2Vid, which unify real MoCap through SMPL and generate controllable RGB videos for cross-representation benchmarking. Synthetic supervision transfers effectively to real video, while spatial parameters are more domain-sensitive and better HMR reconstruction does not necessarily improve gait estimation.
Problem
Existing monocular gait datasets are limited in scale, viewpoints, visual diversity, and shareability, restricting controlled evaluation of generalization and domain shift.
Method
The paper builds SynthGait-19K by unifying heterogeneous MoCap recordings through SMPL, generating depth-conditioned RGB videos with controllable cameras and scenes, and benchmarking multiple gait-estimation representations.
Results
Synthetic supervision transfers effectively to real video across direct RGB and pose-based models, while spatial gait parameters are more sensitive to visual domain shift and improved HMR reconstruction does not necessarily improve gait estimation.
Takeaways & Limitations
SynthGait-19K supports controlled study of viewpoint, training-data scale, synthetic-to-real shift, and gait estimation across substantially different intermediate representations.
Takeaways & Limitations
The dataset incompletely covers severe occlusion, assistive devices, broader clothing and clinical diversity, longer-horizon phenomena, and cohorts beyond GPJATK with PD4T validation.
Abstract
from arXiv · showhide
Accurate estimation of clinically meaningful gait parameters from monocular video is important for scalable mobility assessment, yet progress is limited by the small scale, restricted viewpoints, and limited visual diversity of existing datasets. We introduce SynthGait-19k, a physically grounded synthetic video dataset containing 19,272 walking videos derived from 6,427 MoCap sequences across 437 subjects, with paired SMPL motion and annotations for six gait parameters. To construct the dataset, we develop Gait2Vid, which unifies heterogeneous MoCap recordings through SMPL and synthesizes diverse RGB walking videos under controllable viewpoints and scene appearances. We assess the generated videos for consistency with their conditioning gait kinematics and validate extracted gait events against force-platform measurements. Using SynthGait-19K, we benchmark direct RGB, pose-based, biomechanical, and human-mesh-recovery approaches and analyze viewpoint, training-data scale, and synthetic-to-real domain shift. We also introduce GaitXFormer as a direct RGB reference model for estimating gait parameters. Synthetic supervision transfers effectively to real videos across both GaitXFormer and a pose-based architecture, demonstrating utility across different representations. We further find that spatial gait parameters are more sensitive to visual domain shift and that improved HMR reconstruction alone does not necessarily translate to improved downstream gait estimation.
1. Introduction
Gait estimation from monocular video could support scalable clinical mobility assessment, but existing datasets limit generalization analysis through small scale, restricted viewpoints, and limited visual diversity. SynthGait-19K addresses these constraints with physically grounded synthetic videos, unified motion annotations, and a benchmark spanning multiple representations.
- Motivation: Gait features support diagnosis, fall-risk prediction, and rehabilitation tracking, but conventional IMU and MoCap assessment requires specialized laboratory setups.Controlled measurements can also diverge from everyday gait patterns.
- Motivation: Existing video datasets are expensive to collect and share, often restricting environments, viewpoints, subject appearance, and controlled distribution-shift analysis.Force-plate or MoCap supervision and biometric privacy concerns constrain large-scale collection.
- Dataset: SynthGait-19K contains 19,273 RGB walking videos from 6,427 MoCap sequences across 437 subjects, paired with SMPL motion and six gait-parameter annotations.The source data span five cohorts, including healthy or asymptomatic and clinical populations, with multiple camera configurations and visual appearances.
- Approach: Gait2Vid unifies heterogeneous MoCap recordings through SMPL, controllable virtual cameras, and depth-conditioned video diffusion to generate diverse indoor and outdoor RGB walking videos.The generated videos are evaluated for conditioning-kinematics consistency, while gait events are independently validated against force-platform measurements.
- Benchmark: The benchmark covers HMR, biomechanical, pose-based, and direct RGB methods, with GaitXFormer as a direct RGB reference model.Synthetic supervision transfers effectively to real video for both GaitXFormer and a pose-based STT model, while improved HMR reconstruction does not necessarily improve gait estimation.
2. Related Work
Prior gait-estimation datasets provide valuable synchronized measurements but remain small and capture-specific, while general synthetic human datasets emphasize reconstruction rather than quantitative gait analysis. SynthGait-19K complements this landscape with diverse RGB videos derived from real MoCap while preserving motion and gait-label correspondence.
- Human-motion reconstruction: HMR methods reconstruct 3D pose and shape before deriving gait events and spatiotemporal parameters, but they are optimized primarily for reconstruction rather than downstream gait accuracy.UnderPressure is one example of deriving gait quantities from reconstructed motion.
- Gait datasets: Existing synchronized gait datasets remain relatively small and capture-specific, limiting independent variation of viewpoint, appearance, and scene while preserving underlying motion.GPJATK provides synchronized MoCap and calibrated multi-view RGB, alongside dedicated clinical and instrumented-laboratory cohorts.
- Synthetic human data: SURREAL, AGORA, and BEDLAM increase synthetic diversity in motion, appearance, clothing, scenes, or cameras, but primarily target general human reconstruction rather than quantitative gait analysis.
- Positioning: Unlike projected-pose synthetic musculoskeletal gait data, Gait2Vid starts from real-world MoCap across multiple cohorts and generates diverse RGB videos retaining fitted 3D motion and gait annotations.This correspondence supports controlled analysis of viewpoint and synthetic-to-real visual variation.
3. SynthGait-19K Dataset
SynthGait-19K combines heterogeneous public MoCap sources into a large paired RGB-video and gait-annotation resource. Its videos preserve fitted SMPL motion and six parameter labels across diverse source cohorts and viewpoints.
- Composition: 19,272 RGB walking videos derive from 6,427 unique MoCap sequences totaling 671 minutes across 437 subjects and five public datasets.The source cohorts include healthy or asymptomatic participants and multiple clinical populations.
- Annotations: Each video is paired with fitted SMPL motion and annotations for cadence, walking speed, step length, step width, stooped posture, and arm swing.
4. Gait2Vid: Dataset Construction Pipeline
Gait2Vid constructs SynthGait-19K by converting heterogeneous MoCap into a common SMPL representation, rendering controllable depth-conditioned views, and generating diverse RGB walking videos. The resulting dataset preserves motion-derived gait annotations while varying scenes and viewpoints.
- Unified motion representation: Gait2Vid converts source motions with different marker layouts and joint conventions into a common SMPL representation for video generation and gait annotation.
- Video synthesis: The pipeline renders fitted walking motions from controllable virtual cameras and uses depth-conditioned video diffusion to synthesize RGB videos.Figure 2 describes depth maps formed from synthetic cameras and scene geometry as conditioning for video generation.
- Dataset organization: Table 1 organizes SynthGait-19K and the independent GPJATK real-video benchmark by subjects, sequences, videos, and viewpoint.
- Scene variation: Camera and scene sampling produces variation in subject appearance, background, lighting, and recording context across indoor and outdoor templates.
- Gait annotation: Heel strikes from fitted SMPL motion support cadence calculation, while pelvis and foot displacements, neck–pelvis displacement, and wrist range define the remaining gait parameters.Each generated video remains paired with labels derived from its conditioning motion.
5. Benchmark Protocol
The benchmark evaluates gait estimation on real synchronized RGB–MoCap data and compares direct RGB, pose-based, HMR, and biomechanical approaches using standardized gait annotations and metrics.
- Evaluation Dataset and Metrics: GPJATK provides 152 walking sequences from 32 subjects and 608 RGB videos across four camera configurations, with synchronized MoCap-derived gait references.
- Evaluation Dataset and Metrics: The preferred-view protocol assigns side views to walking speed, step length, stooped posture, and arm swing; front and back views to step width; and all views to cadence.
- Evaluation Dataset and Metrics: Pearson correlation measures inter-sequence variation, while the overall score averages six parameter-wise correlations using Fisher transformation.
- GaitXFormer: GaitXFormer encodes walking videos with a Video ViT, uses six gait-query tokens, and predicts six parameters through parameter-specific linear heads.
- Comparison Methods: The benchmark compares HMR, biomechanical, pose-based, and direct video-based methods under a common gait-estimation protocol.
6. Results and Analysis
Experiments show that synthetic supervision transfers to real gait estimation, but transfer depends on gait quantity, representation, viewpoint, and domain shift; better motion reconstruction alone is insufficient.
- Validation: Generated videos show lower pose and knee-angle errors than corresponding real videos, while velocity error is nearly unchanged, indicating no observed degradation in conditioning-motion adherence.
- Validation: 2.31 frames (≈77 ms) is the mean heel-strike timing error across 6,092 force-platform-observed contacts, with 91.7% localized within 5 frames.
- Real-World Gait Estimation Benchmark: Across HMR and biomechanical methods, performance varies by gait parameter: OpenCap Monocular performs strongly for walking speed, WHAM for step width, and FastHMR for arm swing.
- Real-World Gait Estimation Benchmark: 0.82 is STT’s Fisher-averaged correlation after SynthGait-19K training on real GPJATK, compared with 0.84 for GaitXFormer, while STT achieves the strongest step-length result.
- Efficiency: 0.27 s is GaitXFormer’s runtime for a 5-second clip on one RTX3090, compared with 1.21–139.10 s for evaluated baselines including preprocessing.
- Synthetic-to-Real Transfer Analysis: GaitXFormer’s Fisher-averaged correlation decreases from 0.87 on GPJATK-VACE to 0.84 on real GPJATK, with larger reductions for step length and step width.
- Synthetic-to-Real Transfer Analysis: WHAM’s overall correlation remains similar across generated and real GPJATK, but step width improves from 0.56 to 0.69 while stooped posture decreases from 0.79 to 0.64.
- Effect of SynthGait-19K Supervision on HMR: WHAM reconstruction improves by 7.43% in body-pose error, 2.78% in pelvis MPJPE, and 2.53% in PA-MPJPE, while gait correlation changes only from 0.7045 to 0.7080.
7. Conclusion
SynthGait-19K provides paired synthetic RGB, SMPL motion, and gait annotations for controlled studies of transfer and viewpoint, with useful but bounded real-video generalization.
- Conclusion: SynthGait-19K combines heterogeneous MoCap recordings with unified SMPL motion and six gait-parameter annotations, enabling controlled evaluation of viewpoint and visual-domain shift.
- Conclusion: Synthetic supervision transfers effectively to real video for several gait parameters, while spatial quantities remain more sensitive to domain shift.
- Conclusion: Improved HMR reconstruction does not necessarily improve downstream gait estimation, although predicted gait quantities show consistent associations with clinical severity on PD4T.
- Limitations: The dataset incompletely covers severe occlusion, assistive devices, broader clothing and clinical diversity, and longer-horizon phenomena such as freezing of gait.
- Limitations: Evaluation centers on GPJATK, with PD4T validation; broader cohorts are needed to better characterize generalization.
Supplementary Material
The supplementary material details how SynthGait-19K is rendered, annotated, and validated, emphasizing controllable viewpoints, stable video synthesis, physically defined gait features, and force-platform agreement. It also shows that the synthetic dataset covers gait distributions relevant to real evaluation data while extending beyond them in several parameters.
- Motion and viewpoint construction: Gait2Vid unifies heterogeneous MoCap recordings into SMPL and renders depth videos from frontal, back, side, and random viewpoints with varied camera pitch.Camera pitch is sampled from −5° to 45°, while random viewpoints uniformly vary yaw from 0° to 360°.
- Depth rendering and video synthesis: A synthetic planar ground provides stable scene geometry, preventing camera drift artifacts and improving temporal consistency during RGB video generation.Static ground depth implicitly enforces a static camera configuration.
- Depth rendering and video synthesis: RGB videos are synthesized from depth sequences and text prompts describing subject and environment attributes using a depth-conditioned video-generation model.The generation campaign used Wan2.1-VACE-14B with TeaCache, but reported compute includes discarded candidate samples and is configuration-specific.
- Gait parameter extraction: Gait annotations derive from heel strikes and SMPL joint trajectories, covering walking speed, cadence, step length, step width, stooped posture, and arm swing.Walking speed uses pelvis displacement; cadence uses heel-strike timing; spatial features use foot, neck, and pelvis trajectories.
- Force-platform validation: 2.305 frames (≈76.8 ms) is the mean absolute timing error for 6,092 force-platform-observed heel strikes, with 82.3% within 3 frames and 91.7% within 5 frames.The evaluation measures timing accuracy for contacts observed by force plates, not recall over complete walking sequences.
- Gait parameter distributions: SynthGait-19K provides broader coverage across most gait parameters than GPJATK, whose distributions are smaller and more concentrated.GPJATK largely falls within the synthetic training distribution, while SynthGait-19K extends beyond it for cadence, step width, stooped posture, and arm swing.
C.1. Architecture and Training
GaitXFormer encodes fixed-length cropped walking clips with a Video Vision Transformer, then uses six parameter-specific gait queries and independent prediction heads. It is trained end-to-end with normalized-parameter mean-squared error.
- Input processing: GaitXFormer processes 32-frame, person-centered clips and z-score-normalizes gait parameters using SynthGait-19K training statistics.Shorter clips are zero-padded; frames are resized to 256 × 192.
- Video encoder: The Video Vision Transformer uses a V-JEPA2 initialization, 24 transformer layers, 1024-dimensional embeddings, temporal stride 2, and 16 × 16 spatial patches.All encoder layers are fine-tuned during SynthGait-19K training.
- Gait-query decoder: Six learnable gait-query tokens represent cadence, walking speed, step length, step width, stooped posture, and arm swing.The queries provide parameter-specific representations from the encoded video tokens.
- Gait-query decoder: Each gait query cross-attends to encoded video tokens through a three-layer decoder without self-attention among queries.Removing query self-attention keeps each query focused on evidence for its corresponding gait parameter.
- Prediction and training: Independent linear heads map decoded representations to six scalar estimates, and training minimizes mean-squared error over normalized gait parameters.Optimization uses AdamW with a 1 × 10^-4 learning rate for 100 epochs and global batch size 16.
D. Benchmark Method Details
The benchmark compares pose, biomechanical, HMR, and direct-video approaches under shared real-world evaluation settings. It also examines clinical classification and the transferability of synthetic supervision across architectures.
- Pose-based evaluation: STT† is trained from scratch on SynthGait-19K with six regressors operating on 81-frame, 16-FPS sequences of 25 BODY_25 joints.The released STT model was trained on side-view cerebral-palsy videos and transfers poorly to GPJATK for cadence.
- HMR evaluation: OpenCap Monocular is evaluated as a released WHAM-based monocular 3D motion pipeline without additional adaptation.The benchmark tests whether stronger mesh reconstruction translates into downstream gait estimation.
- HMR evaluation: WHAM adaptation freezes most of the HMR pipeline and trains a lightweight temporal adapter to refine SMPL pose, shape, geometry, and traveled distance.Evaluation covers both SMPL reconstruction on GPJATK-VACE and downstream gait estimation on GPJATK.
- Clinical classification: PD4T contributes 418 annotated trials from 30 participants and 1,666 straight-walking clips for UPDRS severity classification.The experiment freezes each video encoder and trains a lightweight MLP classifier, comparing GaitXFormer with V-JEPA2.
- Clinical classification: UPDRS 2 is substantially underrepresented, while GaitXFormer produces a cleaner confusion-matrix diagonal and reduces adjacent-class confusion relative to V-JEPA2.Both models perform better on the more frequent classes, with weaker recognition of UPDRS 2.
E.2. Patient-Aware Clinical-Anchor Evaluation
The clinical-anchor analysis aggregates videos within participant-score groups and uses participant-level uncertainty procedures. It evaluates feature changes between adjacent clinical states using direction-specific expectations and standardized response means.
- Patient-aware aggregation: Videos sharing a participant and clinical score are grouped before gait predictions are averaged and correlated with clinical score.This prevents multiple videos from the same participant from being treated as independent samples.
- Uncertainty estimation: Confidence intervals use cluster bootstrap resampling that retains all rows belonging to sampled patients, preserving within-participant dependencies.The reported 95% interval uses the 2.5th and 97.5th bootstrap percentiles.
- Directional analysis: Within-patient directional analysis compares adjacent clinical-score states and computes feature changes for each gait feature.The analysis is restricted to comparisons within the same participant.
- Directional analysis: Expected change signs are negative for walking speed, step length, and arm swing, but positive for stooped posture as severity worsens.A pair is counted as expected-direction change using the feature-specific sign convention.
- Effect sizes and error analysis: Effect size is computed as the standardized response mean, while preferred-view MAE results show complementary error patterns across methods and parameters.STT† has particularly low cadence and arm-swing error; PBL and OpenCap-M have the lowest walking-speed and step-length errors, respectively.
F.2. Viewpoint Sensitivity Analysis
The analyses show that gait-estimation performance depends on viewpoint, input design, occlusion pattern, and model representation. GaitXFormer has strong view-averaged performance, but parameter-specific visual requirements and deployment limitations remain important.
- View-averaged comparison: GaitXFormer achieves the strongest view-averaged performance, with average correlation 0.79 versus 0.65 for WHAM and 0.60 for PromptHMR.Its gains are particularly clear for walking speed and step length, while it remains competitive for step width.
- Viewpoint-wise analysis: WHAM viewpoint effects are parameter-specific: walking speed is stable, step width benefits from frontal and back views, and stooped posture is strongest from the side.Arm swing performs better from back and oblique views, whereas step length remains challenging across viewpoints.
- Decoder ablation: The cross-attention decoder obtains the best Fisher-averaged correlation among decoder designs and modestly improves cadence, step width, and stooped posture.It is adopted as the default because it combines a modest gain with a simple task-specific design.
- Training and input ablations: Full encoder fine-tuning outperforms frozen and LoRA strategies on Fisher-averaged correlation, while 32 input frames achieve the highest overall correlation of 0.84.Longer input at 64 frames does not provide the preferred setting used by the final model.
- Occlusion robustness: Persistent upper-body occlusion reduces stooped-posture correlation from 0.75 to 0.20, while lower-body occlusion reduces step-width correlation from 0.67 to 0.27.Random persistent occlusion causes the largest overall drop, reducing Fisher z-averaged correlation from 0.84 to 0.75.
- Paired comparison: GaitXFormer is comparable to STT† under matched synthetic supervision but exceeds WHAM by Δ = 0.133 in average correlation, with advantages in cadence, walking speed, step length, and arm swing.The STT† difference is Δ = 0.013 and its 95% confidence interval includes zero.
- Deployment and scope: A smartphone prototype detects and crops the walking subject before running the full RGB-to-gait pipeline, while deployment remains bounded by unrepresented conditions and nonclinical scope.The paper cautions that performance may vary with severe occlusion, assistive devices, clothing, camera placement, and population characteristics.