Source-linked AI summary
Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences
Natasha Ureyang, Sebastian Porsdam Mann, Yuxin Liu, Zuriel Hassirim, Melanie Almonte, Wenhao Chen, Joyce Ng, Thant Nay Lin, Aung Thiha, Gerald CH Koh, Brian David Earp, Pin Sym Foong
TL;DR
Human surrogates often predict patient preferences inaccurately, and prior P4 systems largely treat values as static ratings rather than context-dependent reasoning. This paper introduces P4-DT, which trains on dilemma decisions, explanations, and feedback, and reports 81.7% accuracy across 12 dyads, exceeding chance and unassisted surrogates. The findings are preliminary and limited by the small, single-panel sample and remote scenario-based testing.
Problem
Human surrogates often predict patient preferences inaccurately, while prior P4 approaches largely overlook the context-dependent nature of medical values.
Method
P4-DT uses a values survey, five varied medical dilemmas, free-text explanations, and prediction feedback to construct an individual preference policy without population-level data.
Results
81.7% accuracy was achieved by P4-DT across 12 patient–surrogate dyads, compared with 55.0% for unassisted surrogates and 61.7% for assisted surrogates.
Takeaways & Limitations
The findings provide preliminary proof-of-concept for P4 decision support based on richer, decision-grounded preference reasoning without population-level training data.
Takeaways & Limitations
The small sample from a single recruitment panel and remote scenario-based testing limit evaluation across diverse backgrounds, family dynamics, and real-time clinical nuance.
Abstract
from arXiv · showhide
In serious illness, human surrogates often struggle to accurately predict patient preferences (68% accuracy), causing decision conflict. Personalized Patient Preference Predictor (P4) agents offer a potential solution, but prior prototypes treat values as static ratings, ignoring the contextual, situation-dependent nature of medical choices. Grounded in the 'logic of care', we present P4-DT (Dilemma Training), a P4 agent that constructs a patient decision policy by engaging users with varied medical dilemmas, eliciting individual preference reasoning through bi-directional training. In a study with 12 patient-surrogate dyads, P4-DT predicted patient treatment choices with 81.7% accuracy, significantly exceeding chance (OR = 5.61 [2.03, 15.51], p < .001) and outperforming both unassisted surrogates (55.0%; OR = 3.67 [1.59, 8.47], p = .002) and surrogates assisted by P4-DT (61.7%). Comparative prompt analyses showed that incorporating contextual scenario decisions and open-ended text improved accuracy by 15.0 percentage points over initial values ratings alone. We discuss implications for further testing and designing of context-aware AI agents that embody richer human experience to partner in complex decision-making.
1 Introduction
Human surrogates often predict patient preferences inaccurately, while prior P4 agents largely model values as static ratings. P4-DT instead elicits preference reasoning through varied medical dilemmas and uses those decisions, explanations, and feedback to predict treatment preferences.
- Motivation: 68% average surrogate accuracy and variability across studies contribute to discordance in serious-illness decisions.Personal experiences, surrogate preferences, and insufficient communication about patients’ wishes can contribute to conflict.
- Research gap: Prior P4 approaches use person-specific data, but recent prototypes largely represent patient values as quantitative ratings.An LLM-enhanced system using quantitative value ratings achieved 72.6% prediction accuracy.
- Approach: P4-DT constructs predictions from a values survey, five varied medical-dilemma decisions, free-text explanations, and feedback, without population-level data.The design treats preference reasoning as something elicited through engagement with concrete situations and possible actions.
- Headline result: 81.7% accuracy was achieved by P4-DT across 12 patient–surrogate dyads, versus 55.0% for surrogates and 61.7% for assisted surrogates.The study compares the agent, unassisted surrogates, and surrogates assisted by P4-DT.
2 Methods
The study evaluated P4-DT with patient–surrogate dyads using remote sessions, dilemma-based training, independent surrogate predictions, and prompt-variation analyses. Accuracy was defined as directional concordance between patient choices and predictions, with mixed-effects logistic regression used for hypothesis tests.
- Participants: 12 dyads completed synchronous remote sessions, with each dyad comprising a prospective patient and a trusted surrogate.Trusted surrogates were nominated by prospective patients; 10 were spouses, one a sibling, and one a parent.
- Scenarios: Scenario repositories and clinician refinement provided systematic coverage of impairment, recovery, pain, illness trajectory, prognosis, and decision-making capacity.Scenario presentations were adapted from international patient decision-aid guidelines.
- Patient training: Patients completed a values dashboard, five training dilemmas, treatment preferences, confidence ratings, prediction feedback, and optional dashboard refinements.The dashboard included quality-of-life preferences, goals of care, pain–financial-cost trade-offs, and an open-ended field.
- Surrogate evaluation: Surrogates independently predicted patient preferences on five testing scenarios and then re-evaluated those predictions with P4-DT assistance.Initial surrogate predictions were made without access to model outputs or patient responses.
- Analysis: Accuracy measured directional concordance, while mixed-effects logistic regression tested performance against chance and unassisted surrogates.The models included participant- and scenario-level random intercepts.
- Prompt variations: Prompt variations compared the original prompt with values-only and scenario-plus-narrative prompts excluding values ratings.The analysis tested whether richer scenario and narrative data improved accuracy beyond values data alone.
3 Results
P4-DT predicted patient treatment preferences more accurately than unassisted human surrogates and exceeded chance. Prompt analyses further showed that scenario decisions and open-ended inputs were important to performance, whereas removing the values survey did not change accuracy.
- Prediction accuracy: 81.7% directional concordance (49/60) was achieved by P4-DT, compared with 55.0% (33/60) for unassisted surrogates.P4-DT’s concordance had κw = 0.69, while surrogate concordance had κw = 0.23.
- Prediction accuracy: OR = 3.67 [1.59, 8.47], p = .002, for P4-DT correctly predicting preferences versus unassisted surrogates.The estimated odds of correct prediction were 3.67 times higher for P4-DT than for TO.
- Prediction accuracy: 81.7% accuracy for P4-DT exceeded 55.0% for TO and 61.7% for TO + P4-DT in the figure comparison.The figure compares the AI agent, unassisted human surrogate, and human surrogate assisted by P4-DT.
- Prompt variations: 81.7% accuracy was retained without the values survey, whereas values ratings alone produced 66.7% accuracy.The lower values-only result was driven largely by increased false negatives.
4 Discussion and Limitations
The findings provide preliminary proof-of-concept for context-aware P4 decision support without population-level training data, while emphasizing that the evidence remains constrained by study scale, setting, and model sensitivity. Surrogates assisted by P4-DT performed below P4-DT alone in this study.
- Discussion: P4-DT achieved 81.7% accuracy alone, compared with 61.7% for assisted surrogates and 55.0% for unassisted surrogates.The authors suggest that conflicting surrogate views or insufficiently persuasive reasoning may explain the assisted-surrogate deficit, requiring future testing.
- Discussion: The study offers preliminary proof-of-concept for LLM decision support without population-level training data.The authors frame P4 agents as potential supplements to surrogate decision-makers rather than substitutes.
- Limitations: n=12 dyads from one recruitment panel limits evaluation across socio-cultural backgrounds, health literacy levels, and complex family dynamics.The authors identify this as a limitation requiring larger-scale study for deeper robustness and generalizability analyses.
- Limitations: Remote scenario-based testing cannot fully capture emotional volatility, real-time clinical nuance, and evolving acute-care trajectories.The limitation concerns the controlled remote testing setting rather than the reported accuracy calculation.
- Limitations: P4-DT performance remains sensitive to prompt architecture and the underlying base-model architecture.Future work should optimize and validate these components.
B Accuracy by Session
Across 12 participant dyads, P4-DT generally predicted patient preferences more accurately than human surrogates.
- 81.7% accuracy for P4-DT versus 55.0% for TO across 12 participant dyads.P4-DT outperformed TO in 7 dyads, matched it in 4, and was outperformed once.