Source-linked AI summary
MMGait: Benchmarking and Unifying Gait Recognition across Heterogeneous Modalities
Saihui Hou, Chenye Wang, Qingyuan Cai, Aoqi Li, Yongzhen Huang
TL;DR
RGB-centered gait benchmarks do not systematically compare the heterogeneous photometric, geometric, and motion cues available from human walking. The paper introduces MMGait and OmniGait++ to evaluate and unify recognition across modalities, finding that one jointly trained model remains competitive with task-specific experts while supporting varying modality subsets.
Problem
RGB-centered benchmarks provide limited evidence about how heterogeneous gait modalities differ, align, and complement one another.
Method
MMGait provides sequence-level multi-sensor correspondence and shared evaluation, while OmniGait++ combines modality-specific encoders, shared identity learning, and adaptive fusion.
Results
OmniGait++ supports single-modal, cross-modal, and multi-modal recognition in one jointly trained model while remaining competitive with separately trained task-specific models.
Takeaways & Limitations
MMGait establishes a common testbed for heterogeneous gait sensing, and OmniGait++ demonstrates unified recognition under changing modality availability.
Abstract
from arXiv · showhide
Gait recognition is commonly studied using RGB videos or their derived silhouettes and poses. Yet human walking produces heterogeneous photometric, geometric, and motion cues that cannot be systematically examined with RGB-centered benchmarks. We present MMGait, a large-scale multi-sensor benchmark that brings visible, infrared, depth, LiDAR, and radar observations into sequence-level correspondence. It provides diverse modalities spanning appearance, contours, geometry, motion, and body structure. Under a shared impostor-augmented protocol, we evaluate single-modal recognition, cross-modal recognition via directed retrieval, and multi-modal recognition using task-specific experts. Across settings, modality rankings vary with probe conditions, cross-modal alignment remains difficult, and fusion often provides complementary gains. This analysis exposes a scalability problem: individual modalities, modality pairs, and fusion configurations are typically handled by separately trained experts. We formulate Omni-Modal Gait Recognition, which unifies single-modal, cross-modal, and multi-modal recognition within a shared identity space. OmniGait++ uses modality-specific front ends followed by a shared identity encoder to preserve modality-dependent cues while learning comparable identity descriptors. An anchor-guided fusion module aggregates modality subsets of varying size without frame-level synchronization. A jointly trained checkpoint covers all three recognition settings and accommodates modality subsets of different compositions and cardinalities. Experiments show OmniGait++ remains competitive with task-specific experts in many shared settings and extends to higher-cardinality fusion unavailable to fixed-pair models. The results establish MMGait as a common testbed for heterogeneous gait sensing and demonstrate the feasibility of unified recognition under varying modality availability.
I. INTRODUCTION
The paper introduces MMGait, a broad multi-sensor benchmark for comparing heterogeneous gait cues, and OmniGait++, a unified model for recognition across changing modality availability.
- Benchmark and evaluation: MMGait combines five sensing streams into twelve gait modalities with sequence-level correspondence across 1,015 identities, ten views, and three walking conditions.The benchmark covers appearance, contours, depth, point clouds, projected geometry, event motion, and pose observations.
- Benchmark and evaluation: The benchmark uses an impostor-augmented protocol spanning single-modal recognition, directed cross-modal retrieval, and multi-modal fusion.These settings assess standalone discriminability, identity compatibility across modalities, and complementary value.
- Benchmark and evaluation: Results show that modality rankings vary with probe conditions, while standalone strength, cross-modal compatibility, and fusion utility are distinct properties.The findings motivate evaluating heterogeneous cues together rather than relying on RGB-centered benchmarks.
- Omni-Modal Gait Recognition: OmniGait++ uses modality-specific encoders and shared identity learning to support single-modal, cross-modal, and multi-modal recognition within one model.Its adaptive fusion handles different modality subsets without requiring a separately trained expert for every configuration.
- Paper scope: The expanded study adds a 1,015-identity benchmark, a common evaluation protocol, and variable-cardinality fusion beyond the preliminary conference work.The extension includes a 19-configuration fusion registry and re-runs the recognition experiments on the expanded benchmark.
B. Cross-Modal, Multi-Modal, and Unified Gait Recognition
Prior gait studies typically optimize cross-modal or multi-modal recognition for predefined modality pairs or combinations. MMGait broadens the sensing basis with sequence-level correspondence across five streams and twelve modalities, while retaining native timing differences.
- Prior recognition settings: Existing cross-modal methods improve alignment but remain optimized for predefined modality pairs, while multi-modal methods generally target fixed complementary combinations.Related work covers camera–LiDAR alignment, silhouette–pose fusion, and camera–LiDAR integration, but does not provide a general variable-subset framework.
- MMGait benchmark: MMGait combines five sensing streams into twelve modalities, including appearance, contours, pose, depth, point clouds, projected geometry, and event-based motion.The benchmark records 1,015 identities across ten directions and three walking conditions, with 482,327 released sequences.
- Data acquisition: The acquisition route yields ten walking directions at approximately 36° intervals, with four full-route recordings spanning normal walking, backpack carrying, and clothing change.The five-pointed-star route traverses five segments in opposite directions, and each full route is divided into direction-specific sequences.
- Data alignment: MMGait associates independently acquired streams at the sequence level using identity, sequence type, and view labels rather than imposing frame-level synchronization.IR, LiDAR, and radar use separate interfaces without a common hardware trigger or time base, while native frame counts are retained.
- Data processing: Stream-specific processing preserves modality characteristics while producing RGB and IR appearance and silhouettes, depth crops, LiDAR and radar point clouds, and projected-depth maps.The resulting sequence-length distributions differ across streams because acquisition rates and temporal boundaries are independently determined.
C. Unified Evaluation Protocols
The benchmark uses identity-disjoint training and testing with two gallery protocols that share probes and positive templates but differ in impostor composition. The same protocol framework supports single-modal, directed cross-modal, and multi-modal recognition.
- Dataset split: The identity-disjoint split assigns 300 identities to training and 715 identities to testing.
- Gallery protocols: Protocol V1 uses a compact view-specific normal-walking template gallery, whereas Protocol V2 adds non-target instances across available views and sequence types.Probes and positive templates remain unchanged, so the protocol difference isolates gallery enlargement and impostor diversity.
- Recognition settings: Single-modal recognition, directed cross-modal retrieval, and multi-modal recognition all use the same two protocols.The configuration may represent one modality, an ordered probe–gallery pair, or a modality subset.
- Metrics: Evaluation computes Rank-1 accuracy and average precision for probe-view-by-template-view matrices, then averages across cells to obtain final Rank-1 and mAP.Queries use BG-01 and CL-01 instances, while positive templates are normal-walking instances at the designated template view.
- Responsible use: MMGait access requires an application, institutional affiliation, intended research purpose, and agreement to a research-use license prohibiting redistribution.The conditions apply to both raw sensor observations and derived modality data, while pseudonymization does not eliminate biometric privacy risks.
IV. SYSTEMATIC EVALUATION OF GAIT MODALITIES
The systematic evaluation shows that gait-modality performance depends on probe condition, gallery protocol, representation, and method rather than following a universal modality ranking. Protocol enlargement particularly exposes modality-specific robustness differences.
- Protocol robustness: Under Protocol V2, RGB and IR appearance remain comparatively stable under BG, while depth and event modalities are more sensitive to the enlarged gallery, especially for CL probes.The enlarged impostor gallery causes substantial degradation across nearly all modalities under CL, with modality-dependent magnitude.
- Method comparisons: DeepGaitV2-P3D consistently achieves the best Rank-1 and mAP for RGB silhouettes, while SkeletonGait leads 2D pose under BG and GPGait++ leads under CL.With GPGait++ fixed, estimated 3D pose remains substantially below 2D pose.
- Modality rankings: Across protocols, RGB and IR appearance lead under BG, whereas LiDAR projected depth and IR appearance lead under CL; depth also outperforms RGB appearance under CL.Event and radar modalities remain comparatively weak under Protocol V2, particularly for CL probes.
- Within-stream representations: Appearance is more discriminative than the RGB silhouette under BG but less discriminative under CL, whereas IR appearance consistently outperforms the IR silhouette.The IR gap may reflect removed near-infrared cues and errors from silhouette segmentation.
- Geometric encodings: LiDAR projected depth slightly outperforms native LiDAR point clouds, while radar point clouds outperform radar projections.The opposite pattern may reflect greater information loss when sparse radar returns are projected onto a regular image plane.
- Overall findings: No evaluated modality–model pair is uniformly superior across probe conditions and gallery protocols.Single-modal discriminability depends jointly on probe condition, input representation, and recognition method.
1) Cross-Modal Baseline:
The baselines show that cross-modal alignment is difficult even when modalities are individually discriminative, while fusion gains depend on probe conditions and sensor complementarity.
- Pairwise Retrieval: 51.2% Cross-Modal Avg. Rank-1 under BG falls to 19.4% under CL, showing that clothing change further increases cross-modal matching difficulty.The corresponding mAP values are 52.3% under BG and 21.9% under CL.
- Pairwise Retrieval: RGB–IR silhouette retrieval achieves the strongest cross-modal performance, consistent with their shared contour abstraction.
- Pairwise Retrieval: Depth and LiDAR projected depth remain difficult to align despite strong standalone recognition, averaging only 36.2% Rank-1 under BG and 17.6% under CL.Standalone discriminability therefore does not guarantee cross-modal descriptor compatibility.
- Intra- and Inter-Sensor Fusion: 96.1% Two-Stream Avg. Rank-1 under BG decreases to 58.2% under CL, while fusion complementarity varies across intra-sensor configurations.The corresponding mAP values are 95.8% under BG and 60.6% under CL; pose and event contribute more under CL, whereas event slightly degrades the anchor under BG.
- Intra- and Inter-Sensor Fusion: All five inter-sensor configurations outperform both the RGB-silhouette anchor and their auxiliary single-modal baselines, with IR appearance performing best among them.Radar projected depth is useful despite weak standalone discriminability, whereas IR silhouettes provide limited gains.
- Discussion and Key Findings: Because modality rankings, cross-modal compatibility, and fusion utility diverge, exhaustive task-specific experts become increasingly impractical as modality combinations expand.The discussion motivates adaptive fusion that responds to the available modality combination rather than treating every auxiliary modality as uniformly beneficial.
V. OMNI-MODAL GAIT RECOGNITION
Omni-Modal Gait Recognition replaces separate experts with one jointly trained identity space for single-, cross-, and multi-modal recognition. OmniGait++ preserves modality-specific processing and uses anchor-guided fusion to support variable modality subsets without frame-level synchronization.
- V. OMNI-MODAL GAIT RECOGNITION: OmniGait++ replaces separate task-specific experts with one jointly trained model supporting single-, cross-, and multi-modal recognition.The formulation operates on different modality inputs and combines modality-aware identity encoding with anchor-guided variable-cardinality temporal fusion.
- Problem Formulation: The model focuses on nine image-like modalities spanning all five sensing streams while retaining appearance, contours, geometry, motion, and body structure.Native LiDAR and radar point clouds and 3D pose are left to extensions with dedicated encoders.
- Problem Formulation: OmniGait++ associates modalities by identity, sequence type, and view without requiring one-to-one frame correspondence across streams with different acquisition rates and temporal boundaries.
- Overview of OmniGait++: Its private-to-shared architecture preserves modality-dependent cues, aligns heterogeneous descriptors, and adapts each modality’s contribution to the available subset.The three requirements are preserving physical cues, making identity descriptors compatible, and avoiding uniformly beneficial-modality assumptions.
- Overview of OmniGait++: Anchor-guided fusion aggregates auxiliary private features into the anchor timeline and spatial layout before applying the shared identity encoder and descriptor head.The shared pathway handles subsets with different compositions and cardinalities without configuration-specific networks.
- Overview of OmniGait++: Unlike preliminary OmniGait’s fixed pairwise fusion, OmniGait++ represents auxiliaries as a variable-size set and uses shared temporal attention for higher-cardinality subsets.
C. Modality-Aware Shared Identity Encoder
The modality-aware shared identity encoder separates low-level modality adaptation from high-level identity learning. Independent private prefixes preserve heterogeneous patterns, while shared encoding and projections produce comparable descriptors across modalities and fusion subsets.
- C. Modality-Aware Shared Identity Encoder: OmniGait++ decomposes encoding into modality-specific front ends and a shared identity pathway to reconcile heterogeneous low-level statistics with comparable descriptors.The private component absorbs input-dependent appearance and geometry, while shared parameters promote a common high-level representation.
- C. Modality-Aware Shared Identity Encoder: Independent tokenizers and private encoders handle incompatible channel structures and physical meanings before features enter shared processing or fusion.Silhouettes encode contours, RGB and IR retain appearance, projected depth encodes geometry, events emphasize motion, and pose heatmaps encode body structure.
- C. Modality-Aware Shared Identity Encoder: After private Stage s, both single-modal and fused features pass through the shared encoder, whose convolutional parameters are shared across modalities and fusion subsets.
- C. Modality-Aware Shared Identity Encoder: Temporal pooling, horizontal-part pooling, and shared part-specific projections produce final descriptors for both individual modalities and fused subsets.Average- and max-pooled responses are combined within each horizontal part to retain distributed and salient information.
3) Modality-Aware Normalization:
The method preserves modality-specific information while aggregating heterogeneous sequences through anchor-guided, region-wise attention that supports variable modality sets without frame-level synchronization.
- 2) Part-Aware Temporal Aggregation:: Anchor-guided temporal fusion uses RGB silhouettes as a stable reference and produces a fused representation with the anchor’s shape and timeline.The anchor suppresses appearance variation while retaining body contours; auxiliary modalities are aggregated into its output reference.
- 2) Part-Aware Temporal Aggregation:: Horizontal regions are processed independently so local gait cues can receive modality-dependent evidence during temporal aggregation.The fusion regions are distinct from the horizontal parts used by the output head.
- 2) Part-Aware Temporal Aggregation:: Modality embeddings identify token sources after auxiliary modalities are gathered into a common sequence.This source information is added before attention-based aggregation.
- 2) Part-Aware Temporal Aggregation:: Anchor tokens form queries, while auxiliary tokens provide keys and values, allowing attention to select full local feature maps across modalities and times.The anchor is excluded from the key–value sets, and cosine-similarity attention accommodates unequal sequence lengths without frame-level correspondence.
- 2) Part-Aware Temporal Aggregation:: A residual fusion output retains the anchor representation while a learnable scalar controls the auxiliary correction when added modalities are noisy or weakly complementary.The fusion equation adds the transformed attended context to the anchor feature.
E. Unified Identity Learning and Inference
OmniGait++ jointly learns modality-specific identity descriptors, directed cross-modal alignment, and fused-descriptor geometry so one representation supports the paper’s recognition settings.
- E. Unified Identity Learning and Inference:: The training objectives first learn discriminative features within each modality, then enforce directed cross-modal alignment, and finally supervise fused descriptors.The resulting descriptors are reused directly by the corresponding inference modes.
- E. Unified Identity Learning and Inference:: Shared classification and independently computed intra-modal triplet losses give all nine modalities common identity supervision while preserving modality-specific retrieval geometry.The intra-modal triplet losses are averaged across modalities and do not introduce cross-modal distances.
- E. Unified Identity Learning and Inference:: Explicit cross-modal metric supervision is applied to ordered retrieval pairs among RGB silhouettes, IR silhouettes, depth, and LiDAR projected depth.The first modality in each ordered pair supplies anchors, with same- and different-identity descriptors from the second modality.
- E. Unified Identity Learning and Inference:: Fused descriptors use a dedicated classifier and configuration-wise triplet losses so subsets with different compositions and cardinalities remain identity-discriminative.Fusion triplets are formed separately within each training configuration before averaging.
4) Overall Objective:
The overall objective balances classification and metric-learning terms across modalities, retrieval directions, and fusion subsets, while one checkpoint serves the supported inference modes under the evaluation protocols.
- 4) Overall Objective:: Five equally weighted loss terms combine single-modal and fusion classification with intra-modal, cross-modal, and fusion triplet objectives.A common margin is used for the three triplet objectives, and group-wise averaging balances the evaluated task dimensions.
- 4) Overall Objective:: At inference, classifiers are discarded and Euclidean retrieval uses pre-BN descriptors for single-modal, directed cross-modal, and same-subset fusion-to-fusion recognition.Cross-modal retrieval independently encodes probe and gallery modalities without invoking fusion.
- 4) Overall Objective:: The same jointly trained checkpoint supports single-modal, cross-modal, and multi-modal recognition without task-specific fine-tuning, score-level fusion, or separate networks.The main experiments use one checkpoint across the three operating modes.
- 4) Overall Objective:: Protocol V2 reports Rank-1 accuracy and mAP separately for BG and CL using an impostor-augmented gallery and averaged off-diagonal view pairs.The protocol reuses the same probe, positive-template, and impostor definitions for omni-modal and expert studies.
- 4) Overall Objective:: The evaluation instantiates nine single-modal tasks, twelve directed cross-modal tasks, and fusion configurations organized around an RGB-silhouette anchor and cue families.The fusion registry covers appearance/contour, geometry, and motion/structure auxiliaries, including 19 variable-cardinality configurations.
B. Performance Evaluation
OmniGait++ retains modality-dependent single-modal strengths, enables direct cross-modal retrieval, and gains from complementary fusion, although performance depends on modality composition and cardinality.
- B. Performance Evaluation: 45.3% Rank-1 accuracy under CL lets the unified RGB branch outperform the RGB-specific DeepGaitV2-P3D baseline by 6.8 percentage points.Under BG, RGB reaches 99.0% Rank-1 accuracy; under CL, IR is strongest at 55.7%.
- 2) Cross-Modal Recognition:: OmniGait++ performs better than pair-specific CL-Gait on five of six modality pairs under both probe conditions, with RGB Sil.↔IR Sil. the exception.RGB and IR silhouettes achieve 78.0% bidirectional-average Rank-1 accuracy under BG.
- 2) Cross-Modal Recognition:: Standalone discriminability and shared physical content do not guarantee cross-modal compatibility, as silhouettes and geometric modalities remain difficult to align.The jointly trained identity space nevertheless enables direct heterogeneous-modality retrieval without pair-specific fine-tuning.
- 3) Multi-Modal Recognition:: Every Two-Stream fusion surpasses both its RGB-silhouette anchor and auxiliary modality under BG and CL, while comparisons with MultiGait++ favor RGB or IR appearance auxiliaries.Under CL, RGB Sil.+IR reaches 75.4% Rank-1 accuracy, and RGB Sil.+Radar rises from 37.8% to 54.7%.
- 3) Multi-Modal Recognition:: Adding RGB and IR appearance to Omni-5 improves CL Rank-1 by 7.1 percentage points, whereas adding event and pose to Omni-9 contributes only 0.5 percentage points.The results indicate diminishing increments from auxiliary cues after appearance, contour, and geometry are combined.
4) Comparison with OmniGait:
OmniGait++ improves over OmniGait across shared operating modes and extends unified recognition beyond fixed RGB-silhouette pairwise fusion. Ablations identify Stage 3 fusion and modality-aware normalization as important design choices, while MMGait’s controlled scope motivates broader future data collection.
- Comparison with OmniGait: OmniGait++ extends one checkpoint from fixed RGB-silhouette pairwise fusion to within-group, across-group, and omni-fusion inputs of varying composition and cardinality.This broadens coverage beyond OmniGait’s fixed auxiliary-modality design.
- Ablation Study: Stage 3 fusion provides the strongest overall balance across six task groups, whereas Stage 4 reduces Omni-Fusion Rank-1 under clothing change by 8.6 percentage points.Stage 2 helps BG cross-modal retrieval slightly but is less effective for most fusion settings.
- Ablation Study on Modality-Aware Normalization: Removing modality-aware normalization lowers cross-modal BG Rank-1 from 57.0% to 43.2% and degrades all fusion groups.The shared encoder and BNNeck normalization improve consistency across unified operating modes.
- Scope and Future Directions: MMGait’s matched recordings control identities, sensors, views, and walking conditions, while future collections should add environmental and device diversity for transfer evaluation.The benchmark’s controlled setting reduces confounding but does not itself cover broader capture domains.
VIII. CONCLUSION
MMGait provides a controlled, identity-disjoint basis for comparing heterogeneous gait sensing across recognition settings. OmniGait++ unifies these settings in one checkpoint, remains competitive with task-specific models, and offers scalable coverage of variable modality subsets.
- Conclusion: MMGait combines five sensing streams and twelve modalities with sequence correspondence, identity-disjoint evaluation, and an impostor-augmented protocol.The framework supports consistent comparison of single-modal, cross-modal, and multi-modal recognition.
- Conclusion: OmniGait++ uses private modality front ends, shared identity encoding, and anchor-guided temporal fusion to support modality subsets with different sizes and compositions.One jointly trained checkpoint covers the unified operating modes without requiring frame-level synchronization.
- Conclusion: The full OmniGait++ checkpoint requires 24.12M parameters versus approximately 222.19M for the covered collection of task-specific experts.The comparison aggregates nine single-modal, six cross-modal, and eight Two-Stream experts.
- Conclusion: Expanding from Omni-5 to Omni-7 and Omni-9 adds 3.42M and 3.41M Active Params, with corresponding increases of 97.21G and 97.14G FLOPs.The reported growth is predictable as modality pairs are added.