Source-linked AI summary

DeMMO: Longitudinal and Cross-Disease Modelling of Digital Mobility Outcomes via Multi-Task Learning

Menghui Zhou, Zhipeng Yuan, Vitaveska Lanfranchi, Po Yang

arXiv:2608.25073v1cs.LGcs.AI

TL;DR

Existing DMO research and temporal multi-task models do not jointly model evolving multivariate relationships across diseases and outcomes, especially with non-overlapping cohorts. DeMMO represents each objective with longitudinal DMO mappings and learns signed relations for selective information sharing. It achieves the strongest reported prediction performance across outcomes and identifies outcome-specific longitudinal DMO patterns for later validation.

  • Problem

    Existing studies and temporal multi-task methods do not jointly model multivariate longitudinal DMO–outcome relationships across diseases, particularly when cohorts do not share participants.

  • Method

    DeMMO uses visit-specific longitudinal DMO coefficient matrices with temporal regularisation, feature selection, and automatic signed relation learning across disease–outcome objectives without shared participants.

  • Results

    0.722 overall nMSE and 0.515 weighted correlation across four outcomes and five visits, significantly outperforming nine strong baselines.

  • Takeaways & Limitations

    Stability selection identifies longitudinally consistent but outcome-differing DMO patterns as candidates for clinical validation and disease monitoring.

  • Takeaways & Limitations

    Selected DMO candidates may substitute for correlated measures and require clinical validation; generalisability to independent cohorts, additional diseases, and longer follow-up remains future work.

Abstract

from arXiv · show

Digital mobility outcomes (DMOs) derived from wearable sensors characterise mobility in daily life and offer a promising means of monitoring disease progression. Yet most DMO studies examine one disease at one visit; they do not model how multivariate DMO relationships with multiple clinical outcomes evolve jointly across diseases. Technically, existing temporal multi-task frameworks can model progression within an individual disease, but they do not jointly model multiple prediction outcomes across diseases, particularly when disease cohorts do not share participants. To address these gaps, we propose DeMMO, an interpretable framework for longitudinal, multi-disease, and multi-outcome learning. DeMMO represents each disease-outcome objective by a longitudinal DMO coefficient matrix and combines temporal regularisation with stable and visit-specific feature selection. Its central technical contribution is an automatic cross-disease and cross-outcome relation-learning mechanism that learns signed relations directly from these longitudinal mappings, enabling selective information sharing without paired participants. We evaluate DeMMO on the recently released, large-scale, multicentre Mobilise-D dataset, which provides a new opportunity to study 24 harmonised real-world DMOs over five visits across multiple mobility-limiting conditions. Against nine strong linear, longitudinal, and deep-regression baselines, DeMMO achieves the best overall and outcome-specific prediction performance, with significant improvements over the strongest baselines. Stability selection further identifies reliable longitudinal DMO patterns for subsequent clinical validation and disease monitoring. The implementation code and experimental results are available at https://github.com/menghui-zhou/DeMMO.

1 INTRODUCTION

Wearable-derived DMOs offer denser real-world mobility monitoring, but existing studies and models do not jointly capture multivariate longitudinal relationships across diseases and outcomes. DeMMO addresses this gap with selective relation learning and identifies outcome-specific patterns for prediction and validation.

  • Motivation: Periodic clinical assessments provide intermittent, resource-intensive mobility information that may miss gradual, fluctuating, or transient changes.Wearable DMOs capture daily-life walking across days and contexts, providing denser longitudinal observations.
  • Research gap: Most DMO studies are cross-sectional, while longitudinal evidence remains disease-specific, limited in sample size, or restricted to prespecified measures.Consequently, relationships between comprehensive DMO profiles and multiple clinical outcomes remain unclear.
  • Research gap: Analysing longitudinal DMOs requires separating stable from visit-specific effects and learning shared structure across outcomes with different, non-overlapping cohorts.These requirements combine evolving DMO–outcome relationships with cross-disease information sharing without paired participants.
  • Methodological contribution: DeMMO learns a signed cross-disease and cross-outcome relation graph directly from longitudinal DMO mappings.The framework retains outcome-specific temporal evolution and feature selection while enabling selective knowledge transfer across unpaired cohorts.
  • Evaluation: Across five visits and several mobility-limiting conditions, DeMMO jointly studies multiple outcomes and outperforms nine strong linear, longitudinal, and deep-regression baselines.The reported improvements over the strongest baselines are statistically significant.
  • Practical contribution: Longitudinal stability selection identifies robust, outcome-specific DMO patterns as candidates for clinical validation and disease monitoring.This extends the analysis beyond prediction toward reliable candidate DMO identification.

2 METHOD

DeMMO models four clinical objectives across three disease cohorts and five visits using longitudinal DMO mappings, temporal regularisation, sparse selection, and learned cross-objective relations. Its formulation supports selective information sharing without requiring shared participants across diseases.

  • Dataset and objectives: The Mobilise-D analysis uses 24 weekly DMOs across five visits and four objectives spanning PD, MS, and PFF; COPD is excluded because outcomes are unavailable at T2 and T4.PD contributes H&Y stage and MDS–UPDRS Part III; MS contributes EDSS; PFF contributes SPPB impairment.
  • Longitudinal formulation: Each visit-specific clinical-status task regresses outcomes on 24 DMOs, with columns representing DMO–outcome relationships and rows representing longitudinal DMO trajectories.The coefficient matrices preserve how each DMO’s association evolves across the five visits.
  • Longitudinal formulation: Fused temporal regularisation encourages adjacent visits to share coefficients while retaining localised changes, and sparse group Lasso selects both visit-specific effects and DMOs active across visits.The element-wise penalty controls visit-specific sparsity, whereas the row-wise penalty controls DMO-level sparsity.
  • Multi-disease and multi-outcome extension: The multi-disease extension treats each disease–outcome pair as a separate longitudinal prediction objective, accommodating different outcome counts and outcome-specific data availability.PD contributes two objectives, while MS and PFF contribute one each.
  • Automatic relation learning: DeMMO learns a symmetric, zero-diagonal relation matrix whose blocks represent within-disease and cross-disease relationships among longitudinal DMO mappings.The relation penalty controls cross-objective information sharing and stabilises estimation; an ℓ2 penalty is used because weak relationships may remain useful.
  • Unified objective: The unified objective combines longitudinal regression, fused temporal regularisation, sparse group DMO selection, and automatic relation learning, optimised by alternating coefficient and relation subproblems.DeMMO retains objective-specific coefficient matrices and does not require shared participants across disease cohorts.

3 EXPERIMENTS

DeMMO is evaluated on four clinical prediction objectives across five Mobilise-D visits against nine linear, longitudinal, and deep-regression baselines. It achieves the strongest overall performance while learning cross-objective relations and stable, outcome-specific DMO patterns.

  • Compared methods and training settings: The evaluation uses 24 weekly DMOs, four clinical objectives, five visits, and participant-level 70%/10%/20% train-validation-test splits.COPD is excluded because required clinical outcomes are unavailable at T2 and T4.
  • Compared methods and training settings: DeMMO is compared with nine baselines spanning conventional linear regression, interpretable longitudinal models, and deep regression methods.Non-temporal methods are fitted independently to each of the 20 outcome–visit tasks.
  • Hyperparameter sensitivity: The sensitivity analysis finds that sparsity penalties most affect validation performance, temporal smoothing is moderately beneficial, and λR supports selective sharing.Performance is comparatively insensitive to λA over the tested range.
  • Prediction performance: DeMMO ranks first or second in 18 of 20 tasks and is uniquely best in 12, with particularly strong performance for PFF at later visits.It achieves the lowest RMSE for PD H&Y at T2–T4 and for PD MDS–UPDRS at T2–T4.
  • Prediction performance: DeMMO reduces overall nMSE from 0.735 to 0.722 and increases wR from 0.500 to 0.515 versus the strongest baseline, with significant gains.The corresponding p-values are 0.0015 and 0.0148.
  • Learned cross-objective relations: The learned relation matrix contains positive and negative cross-objective edges, including strong within-PD and cross-disease relations derived from longitudinal DMO mappings.The strongest positive relations are PD outcome pairs at 0.49 and MS EDSS–PFF SPPB impairment at 0.47; negative relations include PD MDS–UPDRS–MS EDSS at −0.32.
  • Longitudinal stability selection: Stability selection identifies outcome-specific DMO profiles that are generally stable across visits, while no DMO exceeds a mean selection probability of 0.4 across all outcomes.Correlated DMOs may substitute for one another, so selected candidates require subsequent clinical validation.

4 CONCLUSION

The paper concludes that DeMMO models longitudinal DMO relationships across multiple diseases and outcomes while sharing information across cohorts without common participants. Its results support outcome-specific modelling, but the identified DMO patterns still require validation in broader independent settings.

  • Conclusion: DeMMO learns signed cross-outcome structure from DMO mappings and shares information across cohorts without requiring common participants.The framework is designed for longitudinal modelling across multiple outcomes and diseases.
  • Conclusion: Across four outcomes and five visits, DeMMO achieves overall nMSE of 0.722 ± 0.032 and weighted correlation of 0.515 ± 0.030, significantly outperforming nine baselines.These results are reported on the Mobilise-D analysis.
  • Conclusion: Stability selection finds DMO relevance generally consistent across visits but different across outcomes, supporting outcome-specific modelling rather than a universal DMO panel.The resulting patterns are presented as candidates for clinical validation.
  • Conclusion: Future work will assess generalisability in independent cohorts, additional diseases, and longer follow-up studies.This defines the principal scope boundary for the reported DMO patterns.

A.3 MULTI-TASK LEARNING FOR LONGITUDINAL MODELLING

Conventional longitudinal multi-task learning models related future time-point outcomes within a disease, whereas DeMMO defines each disease–outcome pair as an objective and learns signed relations across longitudinal DMO mappings.

  • Existing longitudinal multi-task learning: Conventional longitudinal multi-task learning treats future time points as related tasks predicted from shared baseline features.Structured sparsity, temporal smoothness, and adaptive regularisation identify shared and time-specific biomarkers.
  • DeMMO task definition: DeMMO instead represents each disease–outcome pair as one prediction objective with a longitudinal coefficient matrix.This task definition distinguishes clinical outcomes from temporal tasks.
  • Cross-outcome relation learning: DeMMO learns a symmetric signed relation graph from coefficient mappings, allowing aligned or oppositely oriented DMO patterns to share information selectively.The graph has a zero diagonal and operates at the parameter level rather than through raw outcome covariance.
  • Cross-disease sharing: The relation mechanism supports selective sharing across non-overlapping disease cohorts without pooling participant records or assuming identical effects.MS and PFF outcomes come from different participants and clinical scales, while the two PD outcomes share a cohort.

A.4 DEEP LEARNING FOR DISEASE PROGRESSION MODELLING

Deep models can represent complex clinical time series, but DeMMO uses structured linear modelling because its inputs are a small fixed DMO panel and its goal includes visit-specific feature identification. The optimisation alternates coefficient and relation-matrix updates under structured regularisation and exact constraints.

  • Deep learning context: Deep clinical models learn patient representations or temporal patterns for predicting diseases, diagnoses, medications, and future conditions.They are valuable when signal lies in dense, heterogeneous event sequences.
  • Model choice: DeMMO’s structured multi-task model provides more interpretable effects and more appropriate inductive biases than a high-capacity deep sequence model for this setting.The inputs are a small clinically defined DMO panel observed on five fixed visits, with visit-specific DMO identification as a scientific objective.
  • Empirical comparison: Neural regression and continuous-label representation-learning baselines test whether nonlinear capacity or label-aware embeddings improve prediction in the available sample regime.The comparison does not claim structured linear models dominate deep learning generally.
  • Optimisation: DeMMO alternates updates to longitudinal coefficient matrices with updates to the symmetric, zero-diagonal relation matrix.The coefficient block uses accelerated proximal gradient with backtracking, while the relation block is parameterised through upper-triangular entries.
  • Structured regularisation: The coefficient penalty combines fused Lasso, element-wise Lasso, and group Lasso through an exact row-wise decomposition.Fused-Lasso signal approximation identifies temporally fused and visit-specific coefficients, followed by group shrinkage that can remove a complete DMO trajectory.
  • Relation update: With fixed coefficient matrices, the relation subproblem is strongly convex when λA > 0 and has a unique minimiser.The constrained update uses K = M(M −1)/2 free upper-triangular edge variables; here M = 4 and K = 6.
  • Algorithm: Algorithm 1 initialises the relation matrix at zero, warm-starts coefficient updates, and stops when the complete objective’s relative change meets a convergence criterion or an iteration limit.Backtracking and Nesterov acceleration are used within coefficient updates.

B.3.1 CONVERGENCE AND COMPUTATIONAL COMPLEXITY

Each coefficient block is solved by accelerated proximal gradient, while the relation block has a unique minimiser for λA > 0. Alternating optimisation converges to a block-coordinate stationary point under standard conditions, but not necessarily to a global minimiser.

  • Convergence: The coefficient subproblem is convex and non-smooth, and its exact composite proximal operator lets accelerated proximal gradient converge to its global minimiser.The relation-matrix subproblem is strongly convex when λA > 0 and therefore has a unique minimiser.
  • Convergence: The unified objective is biconvex rather than jointly convex, so alternating optimisation is not guaranteed to reach a global minimiser and may depend on initialisation.With exact or sufficiently accurate block updates, each outer iteration does not increase the objective; under standard regularity conditions, accumulation points are block-coordinate stationary.
  • Computational complexity: The composite proximal operator costs O(MpT) with a linear-time one-dimensional fused-Lasso solver.The coefficient-update cost also includes accelerated proximal-gradient iterations and regression-gradient evaluations.
  • Computational complexity: The relation update remains inexpensive when the number of prediction objectives is small because its direct linear-system solve costs O(K^3), where K = M(M −1)/2.Its system construction depends on the pT-dimensional objective representations.
  • Computational complexity: In this study, p = 24, T = 5, M = 4, and K = 6, so overall cost is dominated by repeated regression-gradient evaluations across participants, visits, and objectives.The composite proximal operations and relation-matrix update are small for these dimensions.

C NON-CONVEX EXTENSIONS OF DEMMO

DeMMO-var1 and DeMMO-var2 replace convex sparsity penalties with concave square-root penalties to reduce shrinkage of strong effects while retaining sparsity. They differ in whether DMO selection and temporal smoothness are independently controlled or coupled.

  • Motivation: Convex fused-Lasso, Lasso, and group-Lasso penalties may over-shrink weak, correlated, or visit-specific informative DMO effects.Such shrinkage may reduce predictive performance and recovery of clinically meaningful longitudinal patterns.
  • Non-convex extensions: DeMMO-var1 and DeMMO-var2 use a concave square-root penalty to reduce shrinkage of sufficiently strong DMO effects while retaining sparsity.Both variants retain the longitudinal regression loss and relation-learning regulariser.
  • DeMMO-var1: DeMMO-var1 treats DMO selection and temporal smoothness as separate components, allowing their strengths to be controlled independently.This separates the two regularisation functions within the longitudinal model.
  • DeMMO-var2: DeMMO-var2 couples DMO selection and temporal smoothness in a single DMO-level penalty.Retention depends jointly on coefficient magnitude and temporal variation.
  • Variant comparison: Comparing the variants tests whether longitudinal DMO modelling benefits more from independent or coupled regularisation.The variants therefore encode different relationships between feature retention and temporal variation.

C.1 DEMMO-VAR1: SEPARATE SELECTION AND TEMPORAL SMOOTHING

DeMMO-var1 replaces the original sparsity penalties with a composite non-convex penalty while retaining separate fused temporal regularisation. Its selection criterion combines DMO-level and visit-specific coefficient information.

  • DeMMO-var1 replaces Lasso and group-Lasso penalties with a composite ℓ(0.5,1) penalty while retaining separate fused temporal regularisation.
  • The outer square root promotes DMO-level sparsity, while the inner ℓ1 norm can remove visit-specific coefficients.
  • Temporal smoothness is controlled independently by λFL.
  • Unlike DeMMO-var2, DeMMO-var1 separates coefficient selection from temporal variation when determining whether to retain a DMO.

C.3 OPTIMISATION

The non-convex variants are optimised by iteratively reweighting convex surrogates, alternating coefficient and relation-matrix updates. This procedure converges to a stationary point under standard assumptions but increases computational cost relative to convex DeMMO.

  • Non-convex variants require additional difference-of-convex iterations, making model selection and training more expensive than convex DeMMO.Their added predictive flexibility is accompanied by longer computation.
  • Non-convex penalties are replaced at each reweighting step by weighted convex ℓ1 and fused-Lasso penalties.
  • For fixed A, accelerated proximal-gradient updates solve the coefficient block; for fixed W, a closed-form quadratic subproblem updates the symmetric relation matrix.
  • Both variants are initialised from convex DeMMO, providing a stable coefficient pattern and avoiding the square-root penalty’s singular point.
  • The procedure converges to a stationary point under boundedness and accurate-subproblem assumptions, but global optimality is not guaranteed.

D EXPERIMENTAL EVALUATION OF THE NON-CONVEX VARIANTS

The evaluation compares convex DeMMO with two non-convex variants using overall, objective-specific, and visit-level prediction metrics. Convex DeMMO achieves the strongest aggregate results across the comparison.

  • 0.722±0.032 overall nMSE is achieved by DeMMO, compared with 0.728±0.043 for DeMMO-var1 and 0.731 ± 0.043 for DeMMO-var2.
  • 0.515 ± 0.030 overall wR is achieved by DeMMO, exceeding 0.509 ± 0.040 for DeMMO-var1 and 0.505 ± 0.038 for DeMMO-var2.
  • DeMMO obtains the lowest nMSE for every clinical outcome and the highest wR for three of four outcomes.DeMMO-var1 is marginally higher only for PD MDS–UPDRS wR.
  • DeMMO yields the lowest RMSE in 17 of 20 visit-level tasks, while each non-convex variant leads in two specified tasks.
  • The consistently weaker aggregate performance of the variants suggests shrinkage-induced bias is not a major limitation in this setting.

E ADDITIONAL EXPERIMENTAL RESULTS

Additional results show that DeMMO is usually competitive across outcomes and visits, but its performance is not uniformly dominant. Later-visit behaviour supports selective information sharing when longitudinal data become sparse.

  • Outcome-specific observations: DeMMO is best or tied for best across all five PD H&Y visits, with clearest advantages at T2–T4.
  • Outcome-specific observations: DeMMO becomes best for PFF from later visits, when the sample size decreases sharply across follow-up.This pattern suggests temporal and cross-objective information sharing may provide useful regularisation when outcome-specific data are sparse.
  • Patterns across visits and methods: DeMMO ranks first or second in 18 of 20 outcome–visit comparisons and is uniquely best in 12.
  • Patterns across visits and methods: DeMMO is not uniformly dominant, with simpler linear models performing better for PD MDS–UPDRS T1 and PFF T1.
  • Patterns across visits and methods: Overall results support selective rather than uniform information sharing, preserving outcome-specific behaviour while helping when longitudinal observations are limited.
Loading 2608.25073v1…