Source-linked AI summary

Evaluation of Pose Tracking Accuracy in the First and Second Generations of Microsoft Kinect

Qifei Wang, Gregorij Kurillo, Ferda Ofli, Ruzena Bajcsy

arXiv:1512.04134v1cs.CVcs.AI

TL;DR

The paper addresses limited comparative evidence on Kinect skeletal-tracking accuracy by evaluating Kinect 1 and Kinect 2 against optical motion capture. Using exercises across subjects and viewpoints, it measures joint and bone-length accuracy and models localization errors. Overall, Kinect 2 provides more accurate and robust pose tracking, with notable lower-leg and foot limitations.

  • Problem

    The study addresses limited evidence on Kinect 2 accuracy and the absence of a concurrent comparison of Kinect 1 and Kinect 2 against motion capture.

  • Method

    The authors compare both Kinect systems with marker-based motion capture across 12 activities, three viewpoints, joint and limb-length measures, and Gaussian-uniform error mixtures.

  • Results

    Overall, Kinect 2 has better joint-estimation accuracy and more robust tracking, including greater robustness to occlusions and body rotation.

  • Takeaways & Limitations

    Kinect 1 can be exchanged with Kinect 2 for the majority of motions, while lower legs and feet require particular caution with Kinect 2.

Abstract

from arXiv · show

Microsoft Kinect camera and its skeletal tracking capabilities have been embraced by many researchers and commercial developers in various applications of real-time human movement analysis. In this paper, we evaluate the accuracy of the human kinematic motion data in the first and second generation of the Kinect system, and compare the results with an optical motion capture system. We collected motion data in 12 exercises for 10 different subjects and from three different viewpoints. We report on the accuracy of the joint localization and bone length estimation of Kinect skeletons in comparison to the motion capture. We also analyze the distribution of the joint localization offsets by fitting a mixture of Gaussian and uniform distribution models to determine the outliers in the Kinect motion data. Our analysis shows that overall Kinect 2 has more robust and more accurate tracking of human pose as compared to Kinect 1.

I. INTRODUCTION

The paper evaluates Kinect 1 and Kinect 2 skeletal tracking against marker-based motion capture, addressing limited evidence for Kinect 2 and the lack of a concurrent comparison. It examines dynamic activities, viewpoints, joint and limb-length errors, and localization outliers.

  • Markerless motion capture is increasingly used for human-computer interaction, healthcare, surveillance, and other movement-analysis applications.
  • Kinect 2 accuracy had been reported only to a limited extent, and no concurrent Kinect 1–Kinect 2 comparison had been published to the authors’ knowledge.
  • Prior work had extensively evaluated Kinect 1, while Kinect 2 had comparatively little skeletal-tracking evidence alongside optical motion capture.
  • The study compares both Kinect generations with marker-based motion capture across 12 activities, including standing, sitting, slow, and fast movements.
  • The evaluation also varies three horizontal orientation angles and analyzes joint-position errors, extracted limb lengths, and localization-error distributions.

III. ACQUISITION SYSTEMS

Kinect 1 uses structured-light depth sensing and SDK-based skeletal tracking, with depth and color captured at up to 30 Hz. Its tracking estimates body parts from depth images using a random-decision-forest approach.

  • Kinect 1 captures color data at 640 × 480 pixels and depth data at 320 × 240 pixels, with acquisition rates up to 30 Hz.
  • Structured light projects a pseudorandom infrared dot pattern, and stereo triangulation estimates 3D point positions from their projections.
  • Depth accuracy typically ranges from about 1–4 cm over 1–4 m, decreasing with the square of distance.
  • The Kinect SDK estimates candidate body parts from depth images using a random decision forest trained on synthetically generated human poses and shapes.
  • Kinect 1’s SDK tracks up to two users and provides 3D locations for 20 joints per skeleton.

C. Kinect 2

Kinect 2 provides higher-resolution color and depth data and uses time-of-flight depth acquisition. Its SDK tracks more users and joints while retaining a methodology similar to Kinect 1’s.

  • Kinect 2 provides 1920 × 1080 color images and 512 × 424 depth images, both higher resolution than Kinect 1.
  • Time-of-flight sensing estimates surface-point distance from the phase shift of modulated infrared light.
  • The Kinect 2 SDK appears to use methodology similar to Kinect 1 while using GPU computation to reduce latency and improve performance.
  • Kinect 2 tracks up to six users and provides 3D locations for 25 joints per skeleton.
  • Compared with Kinect 1, Kinect 2 adds hand-tip, thumb-tip, and neck joints.

D. Calibration and Data Acquisition

The two Kinects and the optical motion-capture system were captured concurrently after hardware setup, clock synchronization, geometric calibration, and temporal alignment. Figure 1 shows the resulting skeletons in the motion-capture coordinate space.

  • The Kinect systems captured skeletal data at 30 Hz using separate SDK versions and a PC configuration supporting full frame rates.
  • Temporal synchronization used Network Time Protocol, with timestamps supplied by the motion-capture server over a local network.
  • Geometric calibration used a checkerboard with three motion-capture markers placed in three positions, followed by rigid coordinate transformation.
  • Figure 1 displays Kinect 1, Kinect 2, and motion-capture skeletons after geometric and temporal alignment.

E. Data Processing

The data-processing pipeline calibrates and synchronizes Kinect skeletons with motion capture, retains 20 common joints, models tracking offsets to remove outliers, and evaluates bone-length variability.

  • Kinect skeleton sequences were mapped into motion-capture coordinates using a calibration-derived rigid transformation, then temporally aligned and resampled.
  • The analysis retained 20 joints common to Kinect 1, Kinect 2, and motion capture, ignoring remaining joints.
  • Joint offsets were modeled as a mixture of Gaussian and uniform distributions, with maximum likelihood estimating the model parameters.
  • After fitting the mixture model, data were classified as on-track or off-track, and off-track observations were excluded from accuracy evaluation.
  • Bone-length variability was evaluated from distances between bone endpoint joints for Kinect and segment-length parameters from motion-capture BVH files.

IV. EXPERIMENTS

The experiments compare Kinect 1 and Kinect 2 with motion capture across 12 exercises, 10 subjects, and three orientations, analyzing common joint positions and limb lengths.

  • Experimental protocol: The protocol included 12 exercises: six sitting or sit-to-stand exercises involving a chair and six standing exercises without props.
  • Experimental protocol: Data were collected from 10 subjects, with five repetitions per exercise except Jogging, which used ten steps.
  • Experimental protocol: Each exercise was recorded at 0°, 30°, and 60° subject orientations relative to the Kinect cameras.
  • Data alignment: Kinect joint coordinates were transformed into the motion-capture global coordinate system and temporally synchronized to Kinect 2 timestamps.
  • Measured quantities: Accuracy analysis used 20 joints common to all three systems and assessed bone lengths for upper and lower extremities.

V. RESULTS AND DISCUSSION

Results are reported as averages across subjects, with sitting and standing values representing means across exercises within each respective set.

  • Reported results are averaged across all subjects, while sitting and standing values average the exercises within each corresponding set.

A. Joint Position Accuracy

Joint tracking errors are generally smaller and less variable in Kinect 2 than Kinect 1, although active joints become less accurate as viewpoint angle increases. Mixture-model outlier removal further reduces measured offsets and variability.

  • Most joint offsets range between 50 mm and 100 mm in sitting exercises for both Kinect systems.
  • Kinect 1 consistently has its largest offsets at ROOT, HIP L, and HIP R, whereas Kinect 2 shows smaller offsets there and larger offsets in lower-extremity joints.
  • For most joints, SDs range from 10 mm to 50 mm, while highly active joints typically exceed 50 mm variability.
  • Mean offsets and SDs of active joints increase with viewpoint angle, especially on the skeleton side turned away from the camera because occlusion increases pose uncertainty.
  • Overall SDs are larger in Kinect 1 than Kinect 2, with feet and hands showing especially large viewpoint-dependent variability during Jogging and Punching.
  • After outlier removal, the mean and SD of most joint offsets decrease, demonstrating that mixture-model fitting can significantly improve data accuracy.

B. Bone Length Accuracy

Bone-length stability provides a structural measure of Kinect pose-tracking robustness. Kinect 2 generally produces smaller bone-length differences and variability than Kinect 1, particularly for upper legs and across larger viewpoints.

  • Bone-length variance and SD over time measure the robustness of the extracted kinematic model because human skeleton segments are expected to remain relatively constant.
  • Kinect bone length is defined as the l2 distance between two successive joint positions, while motion-capture bone length remains constant after calibration.
  • Kinect 1 usually has larger bone-length offsets and SDs than Kinect 2, especially for upper legs because of large vertical hip-joint offsets.
  • Mean bone-length differences and SDs are smaller in Kinect 2 across sitting and standing viewpoints, suggesting a more robust kinematic skeleton.

C. Summary of Findings

Kinect 2 generally provides more accurate, less variable, and more robust skeletal tracking than Kinect 1, while retaining notable foot and ankle offsets. Its lower latency and reduced outlier component further distinguish its tracking performance.

  • Joint-position accuracy: Kinect 1 hip joints are offset upward by about 200 mm, whereas Kinect 2 has generally smaller, more anthropometric offsets.The Kinect 1 offset is especially relevant when calculating knee and hip angles in sitting positions.
  • Joint-position accuracy: Kinect 2 foot and ankle joints are offset from the ground plane by about 100 mm or more, making foot orientation unreliable.Tracking becomes more accurate once the foot is lifted; the offset may originate from time-of-flight artifacts near large planar surfaces.
  • Robustness: Kinect 2 has smaller uniform-distribution components, indicating fewer outliers and more robust tracking, including under partial body occlusions.The mixture analysis also supports more reliable tracking by Kinect 2 during occlusions.
  • Performance: Kinect 2 tracks with much smaller latency than Kinect 1, particularly during fast exercises such as Jogging and Punching.The comparison also found smaller differences and variance in actual limb lengths for Kinect 2.

VI. CONCLUSION

The paper compares pose estimation from the first and second generations of Microsoft Kinect with standard motion capture technology. It concludes that Kinect 2 is generally more accurate and robust, while identifying lower-leg offsets as a remaining issue.

  • VI. CONCLUSION: Compared with standard motion capture, Kinect 2 provides more accurate joint estimation and more robust tracking to occlusions and body rotation than Kinect 1.The study concludes that Kinect 1 can be exchanged with Kinect 2 for the majority of motions.
  • VI. CONCLUSION: Kinect 2’s lower legs retain large offsets, possibly caused by time-of-flight artifacts, unlike Kinect 1’s structured-light system.For Kinect 1, the largest offsets occur in the pelvic area.
  • VI. CONCLUSION: Classifying joint-position offsets with a Gaussian-uniform mixture model can reduce standard deviations by 30% to 40% by excluding outliers and compensating for offsets.The authors suggest this can support more accurate human motion analysis.
Loading 1512.04134v1…