Source-linked AI summary
Evidence-Grounded Mapping of Multimodal Human Sensing Psychological Transdiagnostic Dimensions
Xiyun Hu, Xiangyuan Xue, Yuting Lyu, Hanya Shao, Jingping Nie
TL;DR
The paper addresses how to translate heterogeneous passive-sensing and self-report data into meaningful transdiagnostic psychopathology profiles when observed B-HiTOP responses are unavailable. It builds a clinician-in-the-loop GLOBEM benchmark and compares direct scoring with two-stage semantic abstraction across evidence modalities. Two-stage prediction improves compatibility for EMA and questionnaire evidence but reduces it for passive sensing and combined evidence, while producing more conservative scores.
Problem
Existing sensing studies commonly target single disorders, questionnaire scores, or clinical outcomes, while GLOBEM lacks observed B-HiTOP responses for direct diagnostic-accuracy evaluation.
Method
The study constructs 14,592 GLOBEM participant-day instances aligned to 29 B-HiTOP items across five spectra and evaluates direct versus two-stage predictions using AI assessment and clinician verification.
Results
Two-stage prediction raises mean C from 7.6% to 53.4% for EMA and questionnaire evidence, but lowers it from 80.6% to 21.8% for passive sensing and from 77.8% to 58.2% for combined evidence.
Takeaways & Limitations
Semantic abstraction can organize heterogeneous self-report evidence but can become an information bottleneck for indirect behavioral sensing signals.
Takeaways & Limitations
The benchmark evaluates evidence compatibility rather than agreement with observed B-HiTOP responses, and only 29 of 45 items were retained with uneven directness of support.
Abstract
from arXiv · showhide
Mobile and wearable sensing enables longitudinal observation of behavior, yet translating these signals into meaningful mental health constructs remains difficult. We introduce a clinician-in-the-loop benchmark for evaluating whether large language models (LLMs) can generate evidence-grounded Brief Hierarchical Taxonomy of Psychopathology (B-HiTOP) item profiles from passive sensing, ecological momentary assessment (EMA), and questionnaire evidence. Using the Generalization of Longitudinal Behavior Modeling (GLOBEM) dataset, we construct 14,592 participant-day instances and align multimodal evidence to 29 B-HiTOP items across five spectra. Since GLOBEM lacks B-HiTOP responses, we evaluate evidence compatibility (C) rather than diagnostic accuracy, separating substantive predictions from abstentions when evidence is insufficient for item-level scoring. Two-stage prediction improves C for EMA and questionnaire evidence, but reduces C under passive sensing and combined evidence and produces more conservative score distributions across models, spectra, and evidence settings. Overall, semantic abstraction helps organize heterogeneous self-report evidence while becoming an information bottleneck for indirect behavioral sensing signals.
1 Introduction
The paper frames multimodal sensing as an opportunity for longitudinal mental-health observation but argues that existing disorder- or score-specific formulations inadequately represent transdiagnostic psychopathology. It therefore introduces a clinician-verified benchmark using GLOBEM evidence and B-HiTOP profiles.
- Mobile and wearable devices enable longitudinal observation of sleep, mobility, activity, social interaction, affect, and momentary self-reports.
- Existing digital-phenotyping studies often target one disorder, questionnaire score, or clinical outcome, limiting representation of overlapping psychopathology.
- HiTOP organizes psychopathology into dimensional spectra, aligning with behavioral observations that are rarely specific to one clinical condition.
- GLOBEM combines repeated-wave passive sensing, EMA, and questionnaire measures across multiple behavioral domains for transdiagnostic comparison.
- The benchmark constructs 14,592 participant-day instances and aligns evidence with 29 B-HiTOP items across five spectra, comparing direct and two-stage prediction with AI and clinician evaluation.
2 Benchmark Construction
The benchmark treats each participant-day as an evidence-grounded transdiagnostic profiling instance, using multimodal observations to score clinically defensible B-HiTOP items. Because GLOBEM lacks observed B-HiTOP responses, retained items are selected through clinician-verified evidence alignment.
- Task and participant-day construction: Each participant-day receives scores from 1 (“not at all”) to 4 (“a lot”) for retained B-HiTOP items using passive sensing, EMA, questionnaire, or combined evidence.
- Task and participant-day construction: The benchmark uses 14,592 instances from four GLOBEM waves, with identical instances across evidence settings so comparisons vary model input rather than sample.
- Evidence representation: GLOBEM evidence includes RAPIDS features covering sleep, screen use, physical activity, mobility, proximity, communication, and Wi-Fi context.
- Clinician-verified evidence alignment: Clinician-verified alignment retains items only when sensing, EMA, or questionnaire measures provide a defensible basis for reasoning.
- Clinician-verified evidence alignment: The benchmark retains 29 of 45 B-HiTOP items: 8 directly supported and 21 partially supported across five spectra.
3 Model Inference and Evaluation Design
The evaluation compares direct scoring with a two-stage abstraction-and-scoring pipeline across three evidence settings while holding task inputs and output requirements constant. Compatibility is judged against evidence and separates substantive predictions from abstentions when evidence is insufficient.
- Evidence settings: The benchmark tests passive sensing, EMA and questionnaire, and combined evidence conditions while keeping instructions, item wording, scale, reference, and output schema fixed.
- Prediction pipelines: The direct pipeline scores all 29 items from original evidence, whereas the two-stage pipeline first generates a psychological abstraction and then scores items from it.
- Output validation: Outputs are invalidated when unparsable, incomplete, duplicated, or outside the 1–4 score range; spectrum scores are averages, but item-level evaluation remains primary.
- Evidence-grounded evaluation: The evaluator classifies predictions as reasonable when evidence supports their direction and severity, and unreasonable when they contradict evidence or assign scores 2–4 without evidence.
- Abstention handling: A score of 1 with no relevant evidence is treated as abstention rather than symptom absence because insufficient evidence cannot establish that a symptom is absent.
- Evidence compatibility: Evidence compatibility C measures reasonable predictions among substantive scores after abstentions are excluded, distinguishing compatibility from prediction coverage and severity differentiation.
4 Results
The two-stage pipeline substantially improved evidence compatibility for EMA and questionnaire evidence but reduced it for passive sensing and combined evidence. Effects varied by spectrum and were consistent across waves, while score distributions shifted toward more conservative predictions.
- Pipeline effects: 53.4% versus 7.6% mean C under EMA and questionnaire evidence favored the two-stage pipeline.Evidence compatibility C excluded abstentions and therefore reflected substantive predictions.
- Pipeline effects: 80.6% versus 21.8% C under passive sensing, and 77.8% versus 58.2% under combined evidence, favored direct prediction.The two-stage pipeline reduced compatibility in both settings.
- Spectrum effects: The two-stage pipeline improved compatibility for Internalizing, Somatoform, Detachment, and Disinhibition under EMA and questionnaire evidence, but not Thought Disorder.Under passive sensing and combined evidence, compatibility decreased across all retained spectra, with especially pronounced declines for Internalizing and Thought Disorder.
- Score distributions: The two-stage pipeline shifted scores toward 1 across all spectra, most strongly for Thought Disorder.Because abstentions were excluded from C, the shift indicates reduced severity differentiation and lower prediction coverage rather than mechanically higher compatibility.
- Wave effects: The modality-dependent pipeline effects were consistent across Waves 1–4, with direct prediction stronger for passive sensing and combined evidence and two-stage prediction stronger for self-report evidence.The wave-level analysis was descriptive, and between-wave variation was limited.
5 Discussion and Conclusion
Semantic abstraction improved item-level scoring for heterogeneous self-report evidence but reduced compatibility when behavioral sensing details were central. The benchmark therefore supports modality-dependent representation choices while highlighting limits in coverage and evaluation.
- Discussion: Semantic abstraction reorganized distributed self-report signals for item-level scoring but reduced compatibility for passive sensing and combined evidence.The authors characterize abstraction as a modality-dependent representation choice rather than a uniformly superior strategy.
- Limitations: The benchmark measures evidence compatibility rather than agreement with observed B-HiTOP responses, which GLOBEM does not provide.Abstention-excluded C does not capture prediction coverage, and score-1 concentration can accompany compatible predictions.
- Limitations: Only 29 of 45 B-HiTOP items were retained, and several relied on indirect evidence.The authors call for testing generalization across datasets, populations, and longitudinal splits.