Source-linked AI summary
SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking
Muhammad Saif Ullah Khan, Didier Stricker
TL;DR
Spinal motion estimation lacks large-scale, biomechanically verified 3D annotations for unconstrained movement. This work introduces biomechanics-aware simulation and SIMSPINE, achieving improved 2D spine metrics while providing baselines for 2D detection, multiview triangulation, and monocular lifting.
Problem
Existing approaches lack scalable, biomechanically accurate, clinically relevant, and unconstrained 3D estimation of healthy and pathological spine motion.
Method
The framework augments an existing 3D pose dataset with vertebra-level keypoints generated from musculoskeletal models and releases SIMSPINE with pretrained 2D, multiview, and monocular baselines.
Results
Across 2D detection, multiview triangulation, and monocular lifting, the baselines establish reference performance, including indoor AUC improvement from 0.61 to 0.80 and SpineTrack APS improvement from 0.91 to 0.93.
Takeaways & Limitations
SIMSPINE provides a large-scale open resource and unified benchmark for reproducible, anatomically grounded spine pose estimation under natural conditions.
Takeaways & Limitations
The resource models only selected spinal joints, omits intervertebral translations and dynamics, and is validated for geometric plausibility rather than absolute in vivo accuracy.
Abstract
from arXiv · showhide
Modeling spinal motion is fundamental to understanding human biomechanics, yet remains underexplored in computer vision due to the spine's complex multi-joint kinematics and the lack of large-scale 3D annotations. We present a biomechanics-aware keypoint simulation framework that augments existing human pose datasets with anatomically consistent 3D spinal keypoints derived from musculoskeletal modeling. Using this framework, we create the first open dataset, named SIMSPINE, which provides sparse vertebra-level 3D spinal annotations for natural full-body motions in indoor multi-camera capture without external restraints. With 2.14 million frames, this enables data-driven learning of vertebral kinematics from subtle posture variations and bridges the gap between musculoskeletal simulation and computer vision. In addition, we release pretrained baselines covering fine-tuned 2D detectors, monocular 3D pose lifting models, and multi-view reconstruction pipelines, establishing a unified benchmark for biomechanically valid spine motion estimation. Specifically, our 2D spine baselines improve the state-of-the-art from 0.63 to 0.80 AUC in controlled environments, and from 0.91 to 0.93 AP for in-the-wild spine tracking. Together, the simulation framework and SIMSPINE dataset advance research in vision-based biomechanics, motion analysis, and digital human modeling by enabling reproducible, anatomically grounded 3D spine estimation under natural conditions.
1. Introduction
Spinal motion is biomechanically important but difficult to estimate because vertebral kinematics are complex, subtle, and insufficiently annotated in 3D. SIMSPINE addresses this gap with biomechanics-aware simulation, open vertebral annotations, and pretrained estimation baselines for unconstrained motion.
- The spine bears axial loads, enables locomotion, protects the spinal cord, and exhibits nonlinear, interdependent motion across many articulating vertebrae.
- Traditional motion capture misses vertebral rotations, postural sway, and compensatory pelvic tilts relevant to spinal stability, load distribution, and injury risk.
- Existing RGB approaches remain constrained by controlled capture requirements or limited to 2D annotations with insufficient biomechanical verification.
- SIMSPINE augments a large-scale 3D pose dataset with anatomically valid vertebral keypoints generated through biomechanics-aware musculoskeletal simulation.
- The framework provides an open benchmark with unconstrained motions, pretrained 2D and 3D baselines, and 15 simulated landmarks spanning cervical, thoracic, and lumbar regions.
2. Related Work
Prior spine analysis relies on imaging, clinical reconstruction, or RGB tracking, while scalable 3D pose methods generally overlook anatomical validity and intervertebral coherence. SIMSPINE connects musculoskeletal simulation with vision-based spine estimation.
- The human spine has 33 vertebrae, 24 mobile vertebrae, and motion segments permitting three rotations with small constrained translations.
- Radiographs and EOS support static alignment and upright 3D reconstruction, while CT and MRI provide structural or soft-tissue analysis mainly in static postures.
- Dual-fluoroscopy with model-based tracking provides highly accurate in vivo kinematics but is costly and dose-intensive.
- RGB tracking methods require controlled views or remain affected by projection ambiguity and limited biomechanical validity.
- OpenSim and related toolchains provide musculoskeletal models and simulation components that can be linked with vision for large-scale image annotation.
- Modern monocular 3D lifting methods scale to full-body reconstruction but overlook anatomical validity and intervertebral coherence.
3. Methodology
The methodology augments Human3.6M with sparse vertebral keypoints by combining multi-view detection, OpenSim inverse kinematics, and forward-kinematic marker generation. Quality control and biomechanical analyses assess whether the resulting motions are anatomically plausible and action-sensitive.
- Simulation framework: The framework augments Human3.6M with sparse 3D vertebral positions, per-vertebra rotations, and subject-scaled biomechanical models.It preserves the standard Human3.6M train/test splits and timestamps.
- Simulation framework: Synchronized multi-view RGB spinal detections are robustly triangulated with calibrated cameras to produce pseudo-3D spinal keypoints.The method uses nine spinal keypoints per frame and prunes inconsistent reconstructions using reprojection and view-consistency thresholds.
- Simulation framework: Pseudo-3D spinal points are synchronized and merged with Human3.6M markers as targets for subject-scaled OpenSim inverse kinematics.The merged marker set excludes overlapping pelvis, spine, neck, head-top, and nose labels where appropriate.
- Simulation framework: The model represents lumbar intervertebral joints with three rotational DOFs, while cervical motion uses one aggregate 3-DOF joint and rigid upper segments.Virtual vertebral markers are then generated from the IK solution through forward kinematics, alongside anatomical-axis rotations.
- Curation and validation: Quality control rejects implausible curvature, clamps gimbal-wrap discontinuities, and applies temporal smoothing and interpolation to ensure motion continuity.The curated output contains 37 markers and 62 kinematic axes, including 56 Euler angles.
- Curation and validation: 32–41° LLA and 28–39° TKA averages fall within reported biomechanical ranges, while lumbar ROM curves reproduce established segmental trends.The simulations also show action-sensitive curvature distributions and physiologically credible aggregate neck motion.
4. Baselines and Experiments
SIMSPINE supports three spine-aware pose-estimation tasks—2D detection, multiview reconstruction, and monocular lifting—while ablations identify effective data mixing and dataset usage. Results establish baseline performance, geometric consistency, and important simulation scope boundaries.
- Experiments: SIMSPINE benchmarks 2D RGB pose estimation, multiview 3D reconstruction, and monocular 3D lifting with deployment-oriented baselines.The experiments define training and evaluation protocols for these three tasks.
- Ablation Studies: Using 2% of SIMSPINE with per-batch mixing provides balanced indoor and outdoor performance while avoiding larger synthetic-data fractions.The selected configuration uses about 31k indoor images, roughly matching 33k outdoor SpineTrack samples.
- Multiview 3D Reconstruction: 31.8 mm MPJPE and 29.5 mm P-MPJPE are achieved by the multiview triangulation baseline, while GT 2D produces sub-millimeter P-MPJPE.Detector-based reconstructions remain in the 20–40 mm range because similarity alignment does not remove inter-view noise.
- Monocular 3D Lifting: Full-body monocular lifting improves detected-2D P-MPJPE from 18.6 mm to 16.3 mm and GT-2D P-MPJPE from 17.5 mm to 13.5 mm.The comparison indicates better vertebral localization when global body context is included.
- Limitations: The simulation is kinematics-only, models limited articulated spine regions, and represents nominal healthy motion from indoor Human3.6M captures.It is intended as a scalable benchmarking and pretraining proxy rather than clinical measurement, requiring later fine-tuning on biomechanically validated data.
5. Conclusion
The framework produces anatomically constrained 3D spine motion and releases SIMSPINE with pretrained baselines across key estimation tasks. Its evidence supports a practical bridge between computer vision and musculoskeletal modeling, despite simplifications and domain limits.
- SIMSPINE provides 15 spine-driving keypoints with per-segment rotations across 2.14M frames.
- The released baselines span 2D detection, multiview triangulation, and monocular lifting.
- Code, models, and SIMSPINE annotations will be released for research use, while full-body keypoints require licensed Human3.6M data for reproduction.