Source-linked AI summary
Reading Cognition as Decisions Unfold in Words: A Factorized Inverse Decision Model
Jiawen Kang, Dongrui Han, Xixin Wu, Helen Meng
TL;DR
Action-only inverse decision models miss response dynamics in verbalized cognitive tasks. FIDM separates action selection from action execution, recovering interpretable, non-redundant factors and complementary information in cognitive-screening data.
Problem
Existing inverse decision formulations primarily use action trajectories, despite verbalized cognitive tasks also containing response dynamics such as verbal production, interaction, and hesitation.
Method
FIDM factorizes task-execution likelihood into action and effort factors with separate individual-specific parameters, using language-model-derived traces from verbal transcripts.
Results
FIDM selectively recovers non-redundant action–effort distinctions and provides interpretable, task-localized evidence complementary to clinical, behavioral, and language representations.
Takeaways & Limitations
Extending inverse decision models beyond action trajectories yields an interpretable way to characterize individual differences in sequential task execution.
Takeaways & Limitations
The evaluation centers on one cognitive-assessment task, so generality across other sequential behaviors remains empirically unestablished.
Abstract
from arXiv · showhide
Inverse decision modeling infers latent properties of decision processes from observed behavior, but existing formulations rely primarily on action trajectories. In verbalized cognitive tasks, task execution also produces response dynamics that action-only formulations leave unmodeled, such as verbal production, interaction, and hesitation. We propose a factorized inverse decision model (FIDM) that decomposes each individual's task-execution likelihood into an action factor and an effort factor, governed by separate individual-specific parameters. From raw verbal transcripts, a language model produces structured task-execution traces for factorized inference. On data from 400 older adults performing a grocery-shopping dialog task for cognitive screening, controlled recovery shows selective estimation of the intended factors, while matched semi-synthetic conditions show that FIDM preserves action-execution distinctions even when aggregate behavioral summaries are matched. Action evidence further localizes task-defined deviations across participants. In cognitive-status classification, FIDM provides information complementary to clinical scores, trajectory summaries, and frozen language representations, with consistent gains across all evaluated baselines in the binary setting.
1 Introduction
Existing inverse decision models primarily infer cognition from action trajectories, leaving verbal response dynamics such as fluency and hesitation unmodeled. FIDM addresses this gap by separating action selection from action execution in verbalized cognitive tasks and applying the resulting interpretable profiles to cognitive screening.
- Motivation: Existing inverse decision modeling infers latent decision-process properties from observed behavior but relies primarily on action trajectories.This action-centered formulation leaves response dynamics unmodeled, even though participants with similar actions may differ in fluency, hesitation, and need for guidance.
- Model contribution: FIDM factorizes task-execution likelihood into separate action and effort factors, preserving individual differences in action choices and response dynamics.Participant-specific parameters govern each factor, yielding an interpretable action–effort profile from verbalized task-execution traces.
- Study setting: The study tests FIDM on verbal task-execution data from over 400 older adults performing a grocery-shopping dialog task designed for cognitive screening.Participants describe how they would navigate a store to collect ingredients for a dish.
- Results: FIDM’s action–effort decomposition is selectively recoverable, is not reducible to aggregate behavioral statistics, and provides task-localized evidence complementary to conventional clinical and data-driven representations.These evaluations use cognitive-screening data from over 400 older adults.
2 Related Work
Related work frames inverse decision modeling as inferring latent decision-process properties from behavior, distinguishing task- or objective-focused approaches from this work’s individual-specific action and effort parameters. It also situates the approach among models using response times and structured cognitive assessments of speech, planning, and navigation.
- Inverse decision modeling: IDM infers latent decision-process properties from behavior; this work fixes the task model and estimates individual-specific action and effort parameters from choices and response dynamics.The central distinction is whether inference targets the task or differences across individuals.
- Inverse reinforcement learning: Inverse reinforcement learning infers reward functions from state–action trajectories, with extensions modeling trajectory probabilities, unknown constraints, time-varying rewards, or dynamics uncertainty [15].These approaches primarily infer properties of the task or objective, unlike the present individual-focused formulation.
- Bayesian inverse planning: Bayesian inverse planning infers latent mental states from goal-directed actions, including goals, beliefs, and planning limitations such as bounded search, partial plans, replanning, and failed actions [17].The reviewed models invert action-generation models to characterize latent states and planning behavior.
- Evidence accumulation models: Evidence-accumulation models combine response time and choice, with drift-diffusion models representing evidence accumulation, decision boundaries, and non-decision time.Reinforcement-learning diffusion models connect learned values to accumulation.
- Cognitive assessment: Cognitive assessment uses structured tasks to probe observable abilities through speech, planning, and navigation, including picture description, discourse tasks, Multiple Errands, and Naturalistic Action Test.Virtual multiple-errands and supermarket tasks provide comparable controlled settings for multistep assessment.
3 Factorized Inverse Decision Model
The Factorized Inverse Decision Model represents verbalized task execution with a shared sequential task model and separates individual-specific action choices from context-dependent execution signals. Its factorized likelihood estimates these complementary parameter blocks independently, enabling action selection and execution effort to be interpreted relative to task context.
- 3 Factorized Inverse Decision Model: The shared task model fixes the sequential structure and task-value reference, while η_j controls action sensitivity and θ_j governs the conditional execution-signal distribution.Lower η_j values yield more even action probabilities, whereas larger values favor higher-Q actions.
- 3 Factorized Inverse Decision Model: FIDM models each participant’s trace with an action factor for choices and an effort factor for aligned execution signals, both defined over the shared task model M.Execution signals can include response latency, verbal production, and revision behavior.
- 3 Factorized Inverse Decision Model: The factorized objective separately estimates η̂_j from the realized state–action sequence and θ̂_j from execution signals conditional on that sequence.Clinical outcomes and other external measurements are excluded from estimation and used only in subsequent analyses.
- 3 Factorized Inverse Decision Model: FIDM interprets choices by their task-value gaps and execution signals by their expected distribution in context, rather than by route summaries or raw magnitudes alone.Thus, participants with similar action paths can still receive different inferred execution-effort profiles.
4 Model Instantiation for the Grocery-Shopping Task
FIDM is instantiated on the navigation-and-purchase component of HK-GSDT by converting timestamped dialog into grounded movement traces and aligned execution observations. The model separates movement-level action likelihoods from segment-level speech and hesitation measures across the task’s phase-dependent objectives.
- Task instantiation: FIDM is instantiated on HK-GSDT’s navigation-and-purchase component, using the shared supermarket environment and task objectives for individual-level inference.Timestamped dialog supplies action and execution observations.
- Task instantiation: The task is represented as a discrete grid world whose states track location, task progress, and phase across shopping, a bakery request, and checkout.Each phase activates its corresponding objective and phase-dependent action values.
- Trace construction: A language model grounds route descriptions into grid movements, while deterministic replay reconstructs states and grounded purchase and payment events update task progress.Grounding uses preceding dialog context and current location.
- Factorized observations: Action likelihoods are evaluated per grounded movement, whereas execution is measured per dialog segment using speech duration, character count, turn count, and positive pause duration.Execution measurements are conditioned on the number of aligned movements because route segments may encode several movements.
5 Experiments
Across controlled and real-trace experiments, FIDM selectively recovers action and effort dimensions, preserves distinctions obscured by aggregate summaries, and localizes task-defined deviations. FIDM also provides complementary information for cognitive-status classification, improving every binary baseline and achieving strong three-class results.
- Controlled recovery: Controlled semi-synthetic recovery tests whether FIDM separately recovers action sensitivity, speech duration, character count, dialog turns, and positive pause duration from generated traces.Action sequences are sampled from the task-value policy, while effort observations are generated at the segment level and all parameters are re-estimated using the same inference procedure.
- Matched observable behavior: On matched semi-synthetic behavior, aggregate summaries remain at chance, whereas FIDM action, effort, and full representations achieve AUCs of 0.929, 0.989, and 0.996.This shows that factorized representations retain action–execution distinctions obscured by aggregate statistics.
- Action and effort evidence: Localized examples show that action unexpectedness can mark erroneous turns, backtracking, or missed purchases, while effort residuals can remain elevated in only selected segments, demonstrating factor decoupling.The contrasts include typical routes with elevated effort and irregular routes whose effort evidence is confined to a small subset of segments.
- Binary classification: FIDM improves every binary-classification baseline across all four metrics, with BERT+FIDM reaching 75.63% held-out AUC and MacBERT+FIDM reaching 72.23% held-out balanced accuracy.FIDM alone also has the strongest standalone cross-validation performance, with 70.39% AUC and 65.12% balanced accuracy.
- Three-class classification: For three-class classification, GSDT score+FIDM achieves the strongest held-out results: 78.13% macro AUC and 71.72% balanced accuracy.FIDM consistently improves held-out GSDT-score representations, while gains for trajectory and language representations depend on the metric.
6 Discussion and Limitations
Validation supports FIDM’s selective, non-redundant recovery of action and effort distinctions that aggregate summaries can obscure, while localization depends on model-relative unexpectedness rather than visual deviation alone. The study is limited by evaluation on one cognitive-assessment task and a cohort of 400 participants, leaving broader generality and larger-scale validation unresolved.
- Validation: FIDM selectively recovers non-redundant action–effort distinctions by conditioning execution behavior on the realized action sequence.This preserves differences between what participants do and how they carry out those actions, which aggregate behavioral summaries may obscure.
- Localization: Participant-level localization can identify high-evidence regions that lack visually apparent deviations because unexpectedness is defined relative to the fitted task model and available alternatives.Thus, model-relative evidence need not coincide with visual deviation alone.
- Limitations: The empirical evaluation centers on one cognitive-assessment task, so generality across other forms of sequential behavior remains empirically unestablished.The factorization is not tied to cognitive assessment tasks, but broader applicability requires testing.
- Limitations: The study includes 400 participants—comparable to established clinically collected datasets but modest by conventional machine-learning standards—motivating larger-cohort validation.The passage identifies cohort size as a principal limitation.
7 Conclusion
FIDM separates action selection from action execution within a shared sequential-task model. Across controlled validation and cognitive-screening data, its factors provide interpretable, non-redundant, and complementary information about task execution.
- 7 Conclusion: FIDM separates how individuals select actions from how those actions are executed under a shared sequential-task model.
- 7 Conclusion: Across controlled validation and real-world cognitive-screening data, FIDM yields an interpretable and non-redundant characterization of task execution.
- 7 Conclusion: The resulting factors carry information complementary to conventional behavioral, clinical, and language representations.
A Task and Model Details … A.3 Effort Likelihoods
The appendix specifies the HK-GSDT task, transcript-to-trace grounding and alignment, and FIDM’s effort-likelihood parameterization. It also defines population-conditioned residuals for participant-level localization and visualization.
- A.1 HK-GSDT Task Model: The HK-GSDT asks participants to verbally guide an assessor-controlled figure through fixed supermarket goals from entrance to exit.The sequence includes item collection, an assessor-introduced bakery request, checkout, and exit.
- A.1 HK-GSDT Task Model: FIDM models the task as a deterministic grid world, with phase changes representing task-defined active objectives rather than behavior-inferred latent states.Therefore, identical movements can receive different action values under different goals.
- A.2 Transcript Grounding and Effort Alignment: Timestamped participant–assessor dialog is converted into ordered task-execution traces by grounding route descriptions into grid movements and replaying states.Grounded purchase and payment actions update task progress.
- A.2 Transcript Grounding and Effort Alignment: Transcript grounding used a locally deployed DeepSeek-v4 Flash model because IRB constraints required protected-health-information processing to remain on local infrastructure.This restricted grounding to locally deployable models.
- A.2 Transcript Grounding and Effort Alignment: Action likelihoods are evaluated per grounded movement, whereas execution effort is measured per speech segment and conditioned on its aligned movement count.A single route-description segment may encode several consecutive grid movements.
- A.2 Transcript Grounding and Effort Alignment: All 400 task-execution traces passed automated structural checks, while manual auditing found 19 of 20 transcript–trajectory pairs semantically consistent.One audited pair contained a route completion insufficiently supported by the transcript.
- A.3 Effort Likelihoods: Effort localization computes residuals against population-level, movement-conditioned expectations; these residuals support localization and visualization, while participant parameters come from the likelihoods.For segment u of participant j, movement count enters through x_ju = log(1 + n_ju).
- A.3 Effort Likelihoods: FIDM parameterizes four execution-channel likelihoods using population-level parameters and participant-specific effort parameters, with separate models for pause presence and positive pause duration.The Gamma likelihood uses mean and shared shape, while negative-binomial channels provide additional count-channel parameterizations.
B Experimental Details … C.2 Matched Observable Behavior
The experiments compare trajectory and frozen-language representations, combine standardized feature blocks for classification, and validate factor recovery under controlled and matched-observable conditions. Matched conditions closely align raw and per-movement summaries, while the comparison remains limited to the handcrafted summaries evaluated in the main text.
- B.1 Representations: The trajectory baseline represents each participant with 14 action-summary statistics, including counts, movement directions, revisits, backtracks, and state–movement consistency.
- B.1 Representations: Frozen Chinese BERT and MacBERT utterance encoders produce one 768-dimensional participant representation by chunk pooling and token-length-weighted aggregation.
- B.2 Feature Fusion and Classification: Feature fusion median-imputes missing low-dimensional values, standardizes blocks separately, scales them by inverse square root of dimensionality, and concatenates them.Preprocessing is fitted only on training data and separately within each cross-validation training fold.
- B.2 Feature Fusion and Classification: Downstream experiments use L2-regularized logistic regression, with the full pipeline refitted on 332 training participants for evaluation on 68 held-out participants.
- C.1 Controlled Factor Recovery: Semi-synthetic controlled recovery independently manipulates action sensitivity, speech duration, character count, dialog-turn count, and positive pause duration, reporting standardized partial slopes across all five manipulations.The missing-observation condition removes 30% of effort observations, and a separate condition tests recovery under greater observation variability.
- C.2 Matched Observable Behavior: Matched A/B pairs share state–value contexts and total movement exposure, while differing in action sensitivity and effort segmentation with execution offsets chosen to match aggregate totals in expectation.Both pair members use the same cross-validation fold.
- C.2 Matched Observable Behavior: 0.072 was the maximum absolute standardized mean difference across five repetitions for matched raw and per-movement summaries.Table 6 reports factor recovery in this matched-observable experiment, with mean differences defined as condition B minus condition A.
- C.2 Matched Observable Behavior: The matched-observable comparison covers only raw and length-normalized summaries evaluated in the main text and does not rule out approximation by other handcrafted representations.
C.3 Dataset-Level Action Localization
Dataset-level action localization identifies task-defined deviations using route divergence and target-opportunity bypasses, then evaluates their concentration among participants’ most unexpected actions.
- C.3 Dataset-Level Action Localization: Route-divergence onset marks a movement after which the shortest-path distance to the next observed task completion increases.The next observed task completion serves as a retrospective task anchor, not a participant’s latent intention.
- C.3 Dataset-Level Action Localization: Target-opportunity bypass occurs when a participant leaves an unfinished target shelf despite an available purchase action.
- C.3 Dataset-Level Action Localization: Localization computes within-participant AUC, macro-averages across eligible participants, uses participant-bootstrap confidence intervals, and measures top-decile enrichment against baseline event rates.Top-decile enrichment compares event rates among each participant’s most unexpected 10% of actions with corresponding baseline rates.