Source-linked AI summary

RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts

Diwas Lamsal, Juha Carlon, Reinhard Claeys, Maxim Yudayev, Louis Flynn, Tom Verstraten, David Beckwée, Eva Swinnen, Mihai Bâce, Bart Vanrumste, Benjamin Filtjens

arXiv:2609.08090v1cs.AIcs.CV

TL;DR

Assistive locomotion systems need benchmarks that reflect clinical populations, daily mobility, and precisely timed mode transitions, but existing resources are often narrow or healthy-participant focused. RevalExo constructs a clinically feasible inertial–visual benchmark and evaluates recognition, cross-population generalization, and vision-guided IMU-only transfer. Multimodal fusion performs best overall, while transition recognition and transfer across populations remain difficult.

  • Problem

    Existing locomotion benchmarks often use healthy adults, limited tasks, or labels too coarse to evaluate mode transitions, restricting evidence under realistic clinical mobility demands.

  • Method

    RevalExo records 27 participants across three cohorts using lower-body IMUs and synchronized egocentric video for a clinically feasible subset, then benchmarks recognition, cross-population generalization, and vision-guided IMU-only transfer.

  • Results

    Multimodal fusion provides the strongest overall performance, while visual features retain higher performance than inertial features under population shift in this setup.

  • Takeaways & Limitations

    RevalExo exposes transition recognition and cross-population and cross-modal transfer as open challenges for clinical assistive locomotion research.

  • Takeaways & Limitations

    Synchronized egocentric video was omitted for 14 participants because the added 1.4 kg equipment load raised safety and fatigue concerns.

Abstract

from arXiv · show

Assistive devices for people with mobility impairments, such as powered exoskeletons, rely on accurate locomotion mode recognition to adapt control strategies and provide appropriate assistance during daily activities. However, public benchmarks are typically collected from healthy adults, lack temporally precise labels necessary for detecting mode transitions, or focus on a limited set of tasks. To support development and evaluation under realistic clinical constraints and daily mobility demands, we introduce RevalExo, a functional daily-activity benchmark for inertial and visual locomotion mode recognition. RevalExo is built around a standardized, clinically and ecologically validated daily-activity protocol reflecting the cumulative everyday mobility demands in ageing and clinical populations. The benchmark includes 27 participants across three cohorts: older adults without mobility impairments, stroke survivors, and older adults with probable sarcopenia. The full cohort was recorded with lower-body IMUs, while synchronized egocentric video was collected for a clinically feasible subset of 13 participants. RevalExo provides 10.1 hours of frame-level annotations across 11 locomotion modes, including 5.1 hours of paired inertial--visual recordings. We benchmark three challenges: unimodal and multimodal locomotion mode recognition across multiple horizons, cross-population generalization from older adults without mobility impairments to clinical cohorts, and vision-guided knowledge transfer to IMU-only models. Results confirm consistent gains from fusing inertial and visual inputs but reveal a substantial gap between general recognition ($\sim$93\% F1) and recognition during transitions ($\sim$68\% F1), alongside persistent challenges in cross-population generalization and cross-modal transfer. We release RevalExo to stimulate further research on these open challenges.

1 Introduction

Accurate locomotion mode recognition is needed for assistive devices to adapt support, but existing benchmarks poorly represent clinical populations, daily mobility, and mode transitions. RevalExo addresses these gaps with a clinically feasible, multimodal daily-activity benchmark.

  • Motivation: Stroke and sarcopenia create substantial mobility-related disability, motivating wearable assistance to restore independent mobility.Stroke affects nearly 12 million people annually, while sarcopenia affects an estimated 10–27% of older adults and increases fall and injury risk.
  • Motivation: Locomotion mode recognition must identify current movement and ideally predict upcoming changes so assistive devices can adapt support in time.IMUs capture how a person moves but not the upcoming environment, whereas egocentric vision provides terrain context but has its own limitations.
  • Benchmark gap: Existing public datasets usually cover narrow tasks and healthy adults, while clinical datasets often have small cohorts, IMU-only sensing, or simplified protocols.These limitations reduce their fit for evaluating systems under realistic daily mobility demands and clinical constraints.
  • Benchmark gap: Transition phases are difficult to recognize because they are rare, variable, and sensor readings may not clearly match either adjacent locomotion mode.Transitions are also the moments when assistive control must react, motivating dedicated transition evaluation and cross-population testing.
  • RevalExo: RevalExo uses the FATIG’AGE daily-activity protocol to benchmark inertial, visual, and multimodal recognition under realistic clinical constraints.It includes 27 participants across three cohorts, 10.1 hours of frame-level annotations across 11 modes, and 5.1 hours of paired inertial–visual recordings; the benchmark also evaluates multiple horizons and cross-population generalization.

2 Related Work

Prior locomotion resources vary in cohort coverage, sensing, duration, and label precision, with limited clinical and paired multimodal data. RevalExo combines frame-level labels, clinical cohorts, synchronized sensing, and evaluation across recognition, transfer, and deployment-oriented settings.

  • Dataset landscape: Existing locomotion datasets differ in cohort size, clinical coverage, sensor modalities, recording duration, and label granularity.These dimensions frame the comparison between RevalExo and representative resources.
  • Dataset landscape: Earlier datasets include small transition or multimodal collections, but several lack synchronized inertial–visual data or frame-level mode boundaries.For example, some resources use session-level scenarios, while others provide only a limited number of modes or participants.
  • Dataset landscape: Large multimodal corpora can support transferable pretraining, but they generally lack the frame-level locomotion labels needed for direct mode-recognition evaluation.They complement rather than replace task-specific benchmarks.
  • Clinical coverage: Most public locomotion datasets contain healthy participants, whereas clinical resources often use simplified protocols, small cohorts, or lack standardized temporally precise labels.RevalExo addresses this gap with clinical-cohort coverage, paired modalities where feasible, and frame-level annotations across 11 modes.
  • Methods and deployment: The benchmark compares lightweight and higher-capacity architectures for inertial, visual, and fused recognition, including models designed to capture temporal dynamics.It also evaluates transfer to unseen clinical cohorts and training-time vision support for IMU-only deployment.
  • Methods and deployment: Knowledge distillation and related cross-modal methods use video or multimodal teachers to transfer information to IMU-only students.This addresses privacy, power, compute, and wearability constraints associated with deploying egocentric cameras.

3 Dataset

RevalExo combines synchronized inertial and egocentric video sensing with a clinically validated daily-activity protocol spanning three cohorts. It provides frame-level labels, reliability measurements, and cohort-specific recordings across 11 locomotion modes.

  • Acquisition setup: RevalExo records lower-body IMUs for all participants and synchronized egocentric video for a clinically feasible subset.The setup uses seven Xsens IMUs, Pupil Core smart glasses, external cameras, and on-body logging hardware.
  • Protocol: The dataset follows FATIG’AGE, a standardized protocol capturing sequenced and cumulative everyday mobility demands in ageing and clinical populations.Trials include repeated locomotion tasks connected by level-ground walking and may end early because of fatigue or clinical safety judgment.
  • Annotation: Frame-level annotation assigns exactly one locomotion mode per frame, while transition predictions are defined within ±0.25 s of mode boundaries during evaluation.Three trained annotators used synchronized video streams, reference guidance, and live demonstrations; ambiguous cases were resolved with supervising researchers.
  • Annotation reliability: Fleiss’ κ = 0.944 and mean pairwise Cohen’s κ = 0.943 indicate high frame-level agreement across the 11 classes.Boundary F1 is 0.866 at ±250 ms and 0.957 at ±500 ms; agreement is lower for sit-to-stand and stand-to-sit.
  • Participants: 27 participants span non-impaired older adults, stroke survivors, and older adults with probable sarcopenia, with cohort-specific eligibility criteria.Probable sarcopenia is operationalized using reduced sit-to-stand performance and handgrip strength thresholds.
  • Dataset contents: 10.1 hours of annotated data cover 11 locomotion modes, including approximately 5.1 hours of synchronized inertial–visual recordings.The remaining 14 participants contribute inertial-only recordings.

4 Experimental Setup

The experiments evaluate current and predictive locomotion recognition across multiple horizons, modalities, populations, and transfer strategies. The setup standardizes temporal inputs, subject-wise validation, and model comparisons for inertial, visual, fused, and vision-guided systems.

  • Task formulation: Locomotion recognition is formulated as multi-horizon classification at τ ∈{0,0.1,0.2,0.3,0.5,1.0} s.τ=0 measures current-state recognition, whereas τ>0 evaluates prediction of the future locomotion mode.
  • Inputs: Inertial and video models use 2 s temporal windows, while image models use a single frame.The same temporal duration is applied to both temporal modalities.
  • Evaluation: Performance is reported using subject-averaged macro F1 at τ=0 and τ=0.5 s, including separate transition-window results within ±0.25 s of mode changes.Full results across all six horizons are provided in the appendix.
  • Benchmark tasks: The benchmark evaluates recognition with LOSO-CV on 13 multimodal subjects, cross-population generalization to stroke and sarcopenic cohorts, and vision-guided transfer to IMU-only models.Cross-population training uses seven non-impaired older adults and tests paired stroke survivors plus broader clinical cohorts for inertial-only evaluation.
  • Compared methods: Unimodal baselines include DCL for inertial data and ResNet, MobileNet-v3, X3D, and MViT for visual data.Multimodal methods combine encoder features by concatenation or averaging and include additional fusion architectures.
  • Vision-guided transfer: Vision-guided transfer compares contrastive pretraining with knowledge distillation from a frozen visual teacher to an inertial-only student.Distillation variants include vanilla KD, FitNet-style feature distillation, CRD, and NKD.

5 Results

Results show that multimodal fusion improves locomotion recognition, but transition recognition, cross-population generalization, and deployment efficiency remain difficult. Vision-guided transfer and frozen visual backbones provide useful gains, though practical accuracy–throughput trade-offs persist.

  • Locomotion Mode Recognition: 66.2% transition macro F1 at τ=0 s is achieved by the best inertial+image model, versus 53.4% inertial-only and 54.2% image-only.Fusion combines complementary kinematic and scene information, with especially pronounced gains during transitions.
  • Locomotion Mode Recognition: 92.8% overall F1 at τ=0 s is reached by MViT-DCL, dropping to 89.6% at τ=0.5 s, while transition F1 remains 68.2%.Inertial-only models degrade most steeply as prediction horizon increases, whereas fusion models retain higher scores and degrade more gradually.
  • Cross-Population Generalization: 57.7% overall F1 for stroke survivors and 39.9% for the sarcopenic cohort are obtained by inertial-only models at τ=0 s.The results indicate substantial cross-population degradation and high inter-subject variability relative to within-subset evaluation.
  • Cross-Population Generalization: 88.1% overall F1 at τ=0 s is achieved by the best fused model in the multimodal clinical subset, declining to 85.4% at τ=0.5 s.Vision transfers more consistently across subjects, although both overall and transition scores are lower than in the LOSO setting.
  • Cross-Population Generalization: 79.8% overall F1 at τ=0 s is obtained by frozen R18-DCL, exceeding the 65.6% best inertial baseline.Frozen ResNet-18 and MViT also outperform the inertial baseline, suggesting generic pretrained visual features provide terrain and context cues without environment-specific tuning.
  • Vision-Guided Knowledge Transfer: 7.4 pp on SR and 6.9 pp on ST are the τ=0 s gains of CP+FitNets over the IMU-only baseline.CP+FitNets achieves the best overall macro F1 at both horizons in both clinical cohorts, while transition recognition remains difficult and some improvements are not significant at N=10.
  • Inference Performance: Fusion models provide a practical accuracy–throughput balance on a Jetson Orin Nano, whereas the most accurate models substantially reduce throughput.The trade-off highlights the need for efficient architectures that maintain transition-robust recognition within embedded compute, power, and latency budgets.

6 Limitations and Future Work

RevalExo prioritizes clinically meaningful evaluation through standardized cross-cohort testing, but fixed-scene effects, limited multimodal coverage, and transition-focused evaluation constrain interpretation and motivate targeted extensions.

  • Cross-population evaluation: Cross-population evaluation uses a standardized rehabilitation environment, strengthening cohort comparisons while potentially allowing visual models to exploit the fixed scene layout.Multi-site evaluation is needed to separate generalizable visual cues from fixed-scene effects and establish cross-environment robustness.
  • Multimodal coverage: Clinical safety constrained the multimodal subset because chest-mounted hardware could introduce risks or reduce usable session time under fatigue-limited protocols.Lightweight wireless hardware could extend multimodal coverage and strengthen future cross-population comparisons.
  • Transition evaluation: Transition performance is reported within ±0.25 s windows around mode boundaries, while future work should add lead-time evaluation and balance detection speed against stability.Event-level metrics complement the transition-window analysis with detection delay, missed-transition rate, and false-transition rate.
  • Open challenges: 68.2% versus 92.8% macro F1 shows that transition-window recognition remains substantially below overall recognition for the best multimodal model.Inertial-only performance also degrades substantially under cross-population transfer, leaving both issues central to clinical assistive deployment.

7 Conclusion

RevalExo is a functional daily-activity benchmark covering older adults and clinical cohorts with inertial data and clinically feasible synchronized video. Its tasks show benefits from multimodal and vision-guided approaches while exposing transition and population-shift gaps.

  • Benchmark scope: RevalExo contains 10.1 hours of frame-level annotations from 27 participants across older-adult, stroke-survivor, and probable-sarcopenia cohorts.It includes 5.1 hours of synchronized egocentric video alongside lower-body IMU recordings where clinically feasible.
  • Benchmark tasks: Three benchmark tasks cover locomotion mode recognition, cross-population generalization, and vision-guided knowledge transfer.The benchmark is designed for inertial and visual recognition under realistic clinical constraints.
  • Main findings: Multimodal fusion provides the strongest overall performance, while visual features retain higher performance than inertial features under population shift in this setup.Vision-guided training can improve IMU-only inference for cohorts without paired video.
  • Remaining gaps: Transition-window recognition remains far below overall recognition, and inertial-only performance degrades substantially under cross-population transfer.These gaps remain central for clinical assistive deployment.

Appendix B Effect of Input Window Size

The appendix evaluates DCL across input-window durations and reports the associated recognition gains and inference-speed trade-off on the Jetson Orin Nano.

  • Evaluation setup: Table 8 reports LOSO mean ± SD macro F1 (%) across 13 subjects while varying DCL (a/g) input windows from 0.5 s to 3.0 s.Inference speed is measured on a Jetson Orin Nano.
  • Window-length gains: +7.7 pp at τ=0 s and +8.3 pp at τ=0.5 s result from increasing the input window from 0.5 s to 2.0 s.The gains are reported for DCL (a/g) recognition performance.
  • Window-length trade-off: +0.7 pp at τ=0 s and +1.9 pp at τ=0.5 s result from extending the window from 2.0 s to 3.0 s, with nearly half the inference speed.The longer window therefore provides diminishing returns relative to the 2.0 s setting.

Appendix C Transition Event Evaluation

Dense 30 Hz predictions are evaluated with event-level transition metrics, showing that fusion detects transitions quickly and reliably while inertial-only recognition is slower and misses more events.

  • Metrics: Detection delay, missed transition rate within 250 ms, false transitions per minute, and detection rate define the event-level evaluation.The protocol uses dense 30 Hz frame-level predictions on the 13-subject multimodal subset.
  • Fusion performance: 70 ms average detection delay, 13.9% missed transitions within 250 ms, and 95.5% detection rate make unfiltered fusion best on the reported transition metrics.These results are accompanied by the highest detection rate and are reported without filtering.
  • Inertial-only trade-off: 33.6 versus 55.0 false transitions per minute favors IMU-only DCL on raw false-transition rate, but its 203 ms mean delay and 24.5% miss rate are worse.Without visual access to upcoming terrain, inertial-only DCL waits for gait-pattern changes before committing to terrain-driven modes.
  • Visual-only trade-off: 0 ms median detection delay for ResNet-18 does not imply the best overall transition performance because it has a 122.1 false-transition rate and 85.1% detection rate.The delay is computed only over detected transitions, and the image branch can react when new terrain enters view.

Appendix D Additional Per-Class Results

Appendix D reports additional per-class results for representative models, focusing on how performance changes across horizons and transition windows. Longer-horizon degradation is especially pronounced for short and boundary-sensitive transition classes.

  • Table 10 provides the remaining per-class breakdowns for the representative models reported in the main paper.
  • Per-class degradation at longer horizons is most pronounced during transition windows, especially for short and boundary-sensitive classes.
  • The appendix reports transition-window scores at τ=0.0 s and τ=0.5 s, alongside overall scores at τ=0.5 s.

Appendix E Full Per-Horizon Results

Appendix E extends the main locomotion mode recognition results across all six prediction horizons. It reports subject-averaged macro F1 results for the 13-subject multimodal subset using LOSO-CV, with model and input abbreviations defined in the table.

  • Table 11 extends the main locomotion mode recognition table across all six prediction horizons.
  • Training logs for the extended results are available from Hugging Face.
  • The results report mean ± SD macro F1 (%) over the 13-subject multimodal subset using LOSO-CV.
  • Table 11 identifies the best result per column and defines abbreviations for DCL, MNV3-S, R18/R50, a/g, and A/C.
Loading 2609.08090v1…