Source-linked AI summary
Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation
Spencer Gibson, Tyler Crosse, Magnus Saebo, Achyutha Menon, Eyon Jang, Diogo Cruz
TL;DR
LLMs used in sensitive decisions can behave differently depending on how information and prior responses are presented, but this harness-dependent behavior remains insufficiently tested. This paper studies medical resource allocation by comparing paired and independent updates after adding contrasting patient information, finding that three of four models update differently when their previous answer is visible.
Problem
Harness features such as input scope, interaction protocol, and information order remain underexamined in LLM-based decision-making despite their importance in sensitive workflows.
Method
The paper uses a controlled two-step medical-allocation task that adds one contrasting patient-information sentence while varying whether the model’s prior response remains in context.
Results
Three of four tested models update toward socially favorable patient information when their previous response is visible, but typically do not do so when it is hidden; DeepSeek V4 Flash is the exception.
Takeaways & Limitations
Context admission, placement, wording, and whether models revise or answer independently are engineering decisions that affect recommendations in unintuitive ways.
Takeaways & Limitations
Conclusions depend on prompt phrasing and conversational state, and results are specific to a simplified medical-allocation abstraction rather than verbatim medical workflows.
Abstract
from arXiv · showhide
Large language models are being incorporated into sensitive and important decision-making processes across nearly all fields. While prior work studies model bias around inputs and scenario framing, models can also behave in unexpected and undesirable ways due to context accumulated over their deployment. In this work, we study a medical example in which a model is asked to assign resource-allocation probabilities to two people given brief clinical context, and then sees the same scenario with a single extra sentence containing contrasting patient information, either with or without its previous response in context. Across three of four tested models, the paired-context and independent-inference experiments have different probability shifts, often in opposite directions (in favor of Person B vs. in favor of Person A) when new information is provided. We include additional paired-context experiments to show the effect of varying attributes across scenario axes. Our findings show the context-dependent effect of patient information in a sensitive medical use case. More broadly, our work shows the importance of carefully incorporating LLM-based systems into decision-making processes, context engineering, and further model behavioral studies.
1. Introduction
The paper studies how inference-harness design shapes LLM behavior in a simplified medical resource-allocation setting. It contrasts paired and independent inference while holding scenario content fixed, examining context-dependent effects of added patient information.
- Motivation: Increasingly capable LLMs are being used in important decisions despite insufficient testing and limited understanding of their undesirable, unintuitive behaviors.
- Motivation: Prior bias studies mainly examine single-shot scenarios and therefore miss multi-turn deployments in which models accumulate context.
- Study design: The study uses sanitized descriptions of two patients, then adds one contrasting sentence and compares paired inference with independent inference under otherwise fixed scenario content.
- Contributions: Experiments vary allocation stakes, clinical gaps, information attributes, and framing to analyze systematic paired-versus-independent differences and context-dependent effects of contrastive social information.
- Central claim: The paper argues that small evaluation or deployment-harness choices can produce large apparent effects of information, separately from input privacy or traditional bias.
2. Related Work
Prior work shows that healthcare allocation systems can encode demographic bias even without explicit racial inputs, while LLM outputs materially change with context, formatting, ordering, and conversational interaction. These findings motivate treating inference setup as part of evaluating model behavior in sensitive decisions.
- Bias in healthcare allocation: Healthcare allocation algorithms have systematically underestimated Black patients’ health needs despite using no explicit racial input.Obermeyer et al. (2019) illustrates how apparently neutral pipelines can encode allocation bias.
- Bias in healthcare allocation: Clinical LLM evaluations document race-based reasoning and demographic bias in medical question answering.The supplied passage identifies Omiye et al. (2023) and related clinical-LLM evaluations.
- Context and inference setup: Few-shot ordering, prompt formatting, template choice, and option order can materially change LLM performance, rankings, or answers.The cited studies report large performance swings, reordered model rankings, and answer flips from these implementation choices.
- Context and inference setup: Added irrelevant context can distract model reasoning, showing that supplied content, order, and position are substantive parts of measurement.The passage frames the harness around a model as affecting measured behavior rather than merely implementing it.
- Conversational behavior: Multi-turn behavior can diverge from equivalent single-turn behavior, while instruction-tuned and RLHF-trained assistants may exhibit sycophantic updates to conversational cues.Reported behaviors include updating toward stated user views, reversing correct answers when challenged, and inferring demand from conversational cues.
3. Method
The study evaluates how added social or demographic information changes medical allocation probabilities under paired conversational and independent inference. It varies model, severity, clinical gap, attribute, and framing while estimating contextual shifts in the probability assigned to Person B.
- Models: Four models from different providers and training lineages are evaluated, spanning weight openness, model size, performance, and cost.Experiments use GPT-5.2, GPT-5-mini, DeepSeek V4 Flash, and Kimi K2.5 through OpenRouter, with provider-side updates pinned to March 1–May 1, 2026.
- Scenario design: Each scenario asks the model to assign allocation probabilities to Person A and Person B using clinical success probabilities and expected QALYs.Expected clinical value equals success probability times QALY gain, with each patient gaining approximately 20 QALYs if treatment succeeds.
- Scenario design: Each experiment has two steps: a baseline allocation response followed by the same question with only one contrasting patient-information sentence added.The added sentence gives Person A a contrasting attribute and Person B a target attribute, while the underlying clinical case remains unchanged.
- Experimental axes: The design varies outcome severity, clinical gap, information content, and framing across systematic scenario sweeps.Severity contrasts scheduling priority with the final treatment slot; clinical gaps are G = 0.5 and G = 1.25; eight attributes and neutral or moralized wording are tested.
- Inference setups: Paired inference carries the step-1 response into step 2, whereas independent inference reissues the prompt without that response.The paired matrix covers severity, clinical gap, and eight attributes, while the independent sweep uses the minor, G = 0.5 slice with 20 runs per cell; mirror trials reverse descriptor attachment.
- Estimands and analysis: The primary estimand is the contextual effect on Person B’s assigned probability, analyzed through run-level shifts and complementary threshold-based flips.Pooling averages shifts across matched cells for attribute, severity, gap, and framing views; mean paired shift is treated as more statistically stable than flip rates.
4. Results
Results show that inference setup changes how models respond to added patient information, with heterogeneous effects across models and attributes. Drift is concentrated in caregiver and wealth, amplified by moralized wording and smaller expected-QALY gaps, while sex remains flat and age reverse-signed.
- Inference setup: 16/32 confidence intervals exclude zero for paired-versus-independent mean probability-shift differences, with 11 favoring larger paired shifts toward Person B and 5 favoring independent shifts.The comparison holds scenario content fixed while varying whether the model sees its previous response.
- Model heterogeneity: 5/8 GPT-5-mini confidence intervals exclude zero, all showing larger paired-inference shifts toward Person B; DeepSeek V4 Flash shows the opposite pattern in 4/5 non-null comparisons.GPT-5.2 has 2/8 non-null comparisons, while Kimi K2.5 shows larger paired shifts in 3/4 non-null comparisons.
- Attribute effects: 0.782 is the average pairwise Spearman rank correlation across models for attribute rankings by mean absolute probability shift toward Person B.Drift is concentrated in a small number of attributes, and threshold flips preserve the same attribute ordering.
- Wording effects: Moralized wording pushes GPT-5-mini flip rates near 1.0, doubles GPT-5.2’s wealth effect, and produces statistically significant mean shifts toward Person B for racial information in every model.DeepSeek V4 Flash and Kimi K2.5 show larger moralized effects on race and caregiver, while DeepSeek’s wealth shift changes little.
- Decision conditions: Caregiver and wealth remain strongest across severity and expected-QALY-gap framings, sex stays flat, and age remains reverse-signed.GPT-5-mini and GPT-5.2 show larger shifts under low severity, whereas DeepSeek V4 Flash shows slightly larger shifts under high severity.
- Decision conditions: Smaller expected-QALY gaps produce larger probability movements and more flips, with the strongest gap effect for GPT-5-mini caregiver and wealth at G = 0.5.Those effects remain visible at G = 1.25, while most other axes decay rapidly as the gap increases; probes at G = 5 showed no meaningful shift or flip rate.
5. Discussion
The discussion shows that inference setup changes how models update medical allocation probabilities, especially when prior answers remain visible. It also identifies scenario attributes, clinical gaps, severity, wording, and candidate-position checks as important moderators of these shifts.
- Inference setup: The authors interpret this asymmetry as models overinferring that appended information matters because its inclusion signals relevance in a repeated evaluation.They compare this structure to Monty Hall, where presenting information carries meaning beyond the information itself.
- Attribute analysis: Caregiver, wealth, and parent produced the largest shifts toward Person B, despite Person B being the weaker clinical option, while protected-class status showed no clean relationship with shift magnitude.Attributes other than sex, veteran status, and age produced significant shifts toward Person B.
- Scenario moderators: Smaller expected-QALY gaps generally allowed more movement from patient attributes, whereas larger clinical gaps made the clinically favored option harder to displace; severity showed a weaker negative association.Higher decision severity tended to correlate with smaller probability updates, but this relationship was more model-dependent.
- Wording: Moralized wording significantly increased shifts toward Person B, indicating that phrased preferences can make models more likely to comply with perceived intent.The example contrasts privileged-background wording for Person A with disadvantaged-background wording for Person B.
- Robustness check: In a mirror experiment assigning the stronger clinical position and socially favorable attribute to Person A, shifts tended toward Person A rather than Person B.This check addressed whether the original result was caused by candidate position or the fixed clinical comparison.
6. Limitations
The study measures context-driven update sensitivity rather than independent one-shot judgment, and its conclusions are limited by prompt phrasing, conversational state, calibration uncertainty, output formatting, and the simplified medical-allocation abstraction.
- Design and scope: The paired design measures update sensitivity, not independent one-shot judgment, so conclusions depend on prompt phrasing and conversational state.Probability outputs may not be calibrated, and the structured-output schema may influence how models express uncertainty.
7. Conclusion
The conclusion shows that brief social-context contrasts can shift medical-allocation recommendations beyond stated clinical factors, with effects depending on patient attributes, wording, model, and inference setup. Independent inference typically produces much less updating than paired inference, indicating that access to prior responses changes how the same information affects recommendations.
- Conclusion: Several models update toward Person B after short social-context contrasts, even when instructed to attend to clinical considerations.The effect is highly model-dependent.
- Conclusion: The strongest shifts occur for primary caregivers and patients from lower-income backgrounds, while larger expected-QALY gaps reduce the effect.Mirror experiments show shifts toward the socially favorable patient.
- Conclusion: Normatively moralized wording often increases probability shifts toward the socially favorable patient.Framing experiments identify wording as another factor shaping recommendations.
- Conclusion: Independent-inference updating is statistically different from paired-inference updating and typically involves much less shift.The same information can affect recommendations differently depending on whether the model can view its previous response.
Impact Statement · A. Concrete Example
The paper highlights that small changes in interaction setup, accumulated context, and information framing can substantially alter LLM behavior in high-stakes decision support. Its concrete medical example asks a model to prioritize between two clinically similar patients before and after receiving contrasting caregiver information.
- Impact Statement: Small changes in interaction setup, accumulated context, or information framing can substantially change LLM behavior in high-stakes decision support.The paper presents medical allocation as a concrete example of this practical risk.
- Impact Statement: Developers and institutions should test the full deployment harness, including text accumulation, repeated decisions, and the wording of presented information.The concern extends to clinical, legal, hiring, and other consequential decisions.
- A. Concrete Example: The concrete scenario uses minor severity, a clinical gap of G = 0.5, primary-caregiver status, and neutral framing.These conditions define the specific example before the model receives the scenario messages.
- A. Concrete Example: The model advises on hospital treatment scheduling when both patients receive treatment but only one can be scheduled first.Its recommendations are instructed to be evidence-based, defensible, and focused on relevant clinical considerations.
- A. Concrete Example: Person A has an 85.0% treatment-success chance and Person B has 82.5%, while each would gain ∼20 QALYs if successful.Both patients have the same early-stage condition.
- A. Concrete Example: The model first produces a probability distribution over Person A and Person B before the conversation continues.This is described as the step-1 probability distribution.
- A. Concrete Example: Additional information states that Person A has no dependents while Person B is a primary caregiver.The model then produces a step-2 probability distribution.
B. Prompt Specification · C. Clinical-Gap Selection
The appendix specifies the prompts, baseline clinical-gap construction, paired update attributes, response format, and selected gap values used in the main matrix. It also documents how severity, conditional QALY wording, and tool-call constraints were operationalized.
- B. Prompt Specification: The appendix provides the full prompt and response-format specification used to generate the main paired matrix.
- B.1. System Prompts: Severity-specific system prompts frame minor cases as treatment scheduling and severe cases as final-slot allocation under severe capacity constraints.The severe wording replaces scheduling with allocation while retaining identical structural instructions for probability formatting and tool use.
- B.1. System Prompts: Models must output probabilities for Person A and Person B that sum to 1, explain the distribution, avoid directly choosing a person, and call the end round tool exactly once.The harness treats the person with the higher probability as the effective recommendation.
- B.2. Baseline User Prompt: The baseline fixes Person A’s success probability at 85% and sets Person B’s probability to achieve target expected-QALY gap G with conditional QALY gain held at ∼20.For G = 0.5, Person B’s success probability is 82.5%; for G = 1.25, it is 78.75%.
- B.3. Update Step: The paired update adds contrasting information about Person A and Person B while preserving the model’s step-1 end round response in conversation history.Main-matrix attributes include wealth, race, education, family caregiving, parental status, military veteran status, sex, and age.
- B.3. Update Step: The update design also includes parental status and age contrasts, plus moralized wording variants for wealth, race, and caregiving in a framing probe.
- B.4. Response Format: The response schema requires reasoning, probability a, and probability b as numeric fields bounded between 0 and 1, with probability a + probability b = 1 verified in checked runs.The sum constraint is enforced by system-prompt instruction rather than by the schema itself.
- C. Clinical-Gap Selection: The main matrix uses G = 0.5 and G = 1.25 because exploratory runs at G = 2.5 and G = 5.0 produced little switching behavior.The retained values bracket the regime in which contextual updating was most informative.
D. Targeted High-Sensitivity Probes
Targeted high-sensitivity probes found no or minimal drift across several additional frontier models under conditions favoring drift in the main matrix. These exploratory checks suggest the main-matrix behavior is not universal, but they do not establish general absence of drift.
- Probe design: Targeted probes screened additional frontier models under the main matrix’s most drift-favorable conditions and were not directly comparable to its full quantitative results.The probes focused on the largest-effect cells rather than reproducing full matrices.
- Probe results: Zero drift was observed for GPT-5.4 at the smallest clinical gap, low severity, and the primary-caregiver and lower-income attributes.These conditions were selected because they favored drift and corresponded to the largest shifts in the main matrix.
- Probe results: GPT-5.4-mini showed minimal caregiver movement and flat wealth movement, while Claude Sonnet 4.6 and Claude Opus 4.6 were flat on both axes.All three models were tested with the same smoke-probe logic restricted to the top two attribute axes.
- Interpretation and limitation: The flat targeted outcomes are consistent with, but do not establish, broader absence of drift and require a full GPT-5.2-sized matrix for stronger claims.The authors present these as exploratory evidence that main-matrix behavior is not universal across current frontier models.
E. Control Experiments … F.4. Moralized-Framing Flip-Rate View
The paper uses clinically irrelevant controls to test context-dependent updating, then supplements continuous probability shifts with threshold-sensitive flip-rate analyses across attributes, severity, clinical gaps, and moralized framing. These analyses show that flip behavior varies substantially by model and condition, with caregiver and wealth generally producing the strongest effects.
- E. Control Experiments: Clinically irrelevant controls used a 0.5-QALY clinical gap in low-severity decisions and tested both paired and independent inference.Controls varied case-ID parity, explicitly null information, or no added sentence.
- E. Control Experiments: Case-ID parity, explicit null information, and implicit exact repeats did not produce systematic probability shifts.These controls were designed to test whether updating depended on clinically relevant context.
- F. Supplementary Figures: Flip-Rate, Severity, and Gap Views: The appendix treats mean paired probability shift ¯∆ as the primary continuous outcome because it is more stable at n = 20, while flip rate is deployment-oriented but threshold-sensitive.Flip rates can dissociate from continuous shifts, particularly for movements that do not cross the 0.5 threshold.
- F.1. Flip-Rate View by Attribute: Caregiver and wealth dominate pooled flip-rate comparisons, sex and veteran are near zero, and age is zero because reverse-signed shifts rarely cross 0.5.GPT-5-mini records flip rates of 0.81 on caregiver and 0.64 on wealth, separated from the other models.
- F.2. Severity-Conditional Views: Increased severity generally lowers paired probability shifts and flip rates, although the severity effect is small.GPT-5-mini exceeds 75% flip rate on wealth and caregiver in some severity/gap combinations, whereas GPT-5.2 remains at or below 20%.
- F.3. Expected-QALY Gap Views: The smaller clinical gap, G = 0.5, generally produces larger movements and more flips toward Person B than G = 1.25.This pattern is most visible for GPT-5-mini, especially on caregiver and wealth, while most other axes decay as the gap increases.
- F.3. Expected-QALY Gap Views: Flip-rate differences by severity and gap are threshold crossings to Person B, with smaller gaps and lower severity generally permitting more crossings.The severity and expected-QALY gap views provide threshold-based companions to the continuous-shift analyses.
- F.4. Moralized-Framing Flip-Rate View: Under moralized wording, wealth, race, and caregiver exceed 88% flip rate for GPT-5-mini; race rises from 0.61 to 0.92 and wealth from 0.68 to 0.94.For the other three models, moralization produces visible but smaller threshold-scale movements because neutral baselines are farther below the threshold.
F.5. Paired vs. Independent Flip-Rate View · F.6. Mirror Attribute Assignment
Paired and independent inference produce sharply different flip-rate changes for GPT-5-mini, while mirror assignment generally reverses probability shifts toward the person receiving the target attribute. The reversal is strongest for GPT-5-mini and caregiver, but occurs across models and attributes at varying rates.
- F.5. Paired vs. Independent Flip-Rate View: +0.812, +0.637, and +0.450 paired B-rate changes for caregiver, wealth, and race fall to +0.025, +0.025, and +0.050 under independent inference for GPT-5-mini.The collapse of paired flips under independent inference is most visible in the threshold analysis.
- F.5. Paired vs. Independent Flip-Rate View: DeepSeek V4 Flash shows comparable or larger threshold movement under independent inference, consistent with the continuous-shift view.This contrasts with GPT-5-mini’s paired-flip collapse.
- F.5. Paired vs. Independent Flip-Rate View: Figure 13 compares paired and independent inference by model and attribute using changes in the rate of recommending Person B.It serves as the threshold companion to Figure 2.
- F.6. Mirror Attribute Assignment: The mirror-assignment diagnostic reverses the contextual contrast so the target attribute favors Person A, testing dependence on attribute contrast rather than position or added context.It uses the same paired-inference setup while changing which person receives the target attribute.
- F.6. Mirror Attribute Assignment: Across all 32 mirror cells, GPT-5-mini’s probability shift favors Person A.The matrix spans minor and severe conditions, G = 0.5 and G = 1.25, eight attributes, neutral D, and 20 runs per cell.
- F.6. Mirror Attribute Assignment: 13/32, 18/32, and 23/32 mirror cells reverse for GPT-5.2, DeepSeek V4 Flash, and Kimi K2.5, respectively.Caregiver shows the strongest mirrored structure across all models; age also reverses in several DeepSeek and Kimi cells.
- F.6. Mirror Attribute Assignment: Figure 14 reports mean paired probability shifts by model and attribute, pooled over severity and expected-QALY gap, contrasting original Person-B assignment with mirror Person-A assignment.Person A remains clinically favored in the mirror comparison.