Source-linked AI summary

Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making

Minda Zhao, Xu Han, Rishabh Goel, Maya Dagan, Noa Dagan, Adithya Madduri, Payal Chandak, Shilpa Nadimpalli Kobren, Isaac S. Kohane

arXiv:2608.25236v1cs.CYcs.AIcs.CL

TL;DR

Evidence on how LLMs handle ethical trade-offs in rare disease care is limited. Using a clinically grounded benchmark of rare disease dilemmas, this study finds that models consistently prioritize Justice and respond strongly to decision-maker framing.

  • Problem

    Evaluations of LLMs largely emphasize factual and clinical performance while providing limited evidence about their ethical trade-offs in rare disease care.

  • Method

    The study constructs and validates clinically grounded rare disease vignettes, maps forced-choice options to bioethical values, evaluates LLM decisions, and analyzes contextual effects.

  • Results

    Across evaluated models, Justice ranked first, while decision-maker framing strongly shaped value selection and increased Autonomy and Beneficence in Individual or Medical Team contexts.

  • Takeaways & Limitations

    LLM ethical reasoning in rare disease decisions converges on Justice-oriented choices and is more responsive to assigned authority roles than patient-centered clinical factors.

  • Takeaways & Limitations

    Vignette generation used GPT-4.1 and simulated expert critiques, while uneven decision-maker distributions limit the generalizability and reliable estimation of some effects.

Abstract

from arXiv · show

Clinical decision-making often involves prioritizing ethical values, such as beneficence, non-maleficence, respecting a patient's autonomy, and justice. Recent work has begun to assess how large language models (LLMs) make such subjective, value-laden clinical judgments. However, evaluations of LLM decision-making in rare disease care contexts, where ethical tensions are ubiquitous and where scarce prior information likely impacts LLM behavior, are still lacking. Here, we present a benchmark of 208 clinically grounded rare disease vignettes, each of which presents genuine, high-stakes conflicts. When prompting 11 state-of-the-art LLMs to choose between clinically defensible yet ethically conflicting next steps embedded within these vignettes, we found that all evaluated models consistently prioritized justice over other core bioethical principles. Specifically, models overwhelmingly favor equal resource allocation over need-based considerations, indicating LLMs' limited responsiveness to differences in clinical severity or situational context. We also identify a strong authority-framing effect: models favor justice in committee-based contexts and shift toward beneficence and autonomy only when final decisions are framed as being made by clinicians or patients respectively. Our work suggests that institutional pressures surrounding rare disease resource utilization may be silently reflected in LLM-based decision support systems, with finer ethical considerations disregarded.

Introduction

Rare disease care routinely presents clinically defensible choices under uncertainty, making value prioritization central to decisions involving autonomy, beneficence, non-maleficence, and justice. This study introduces a clinically grounded benchmark to empirically characterize how LLMs prioritize ethical values in such contexts.

  • Motivation: Rare disease care makes ethical trade-offs routine because evidence is scarce, resources are constrained, outcomes are high-stakes, and expertise is frequently contested.Relevant constraints include limited clinical trial slots, high costs, irreversible outcomes, and families’ active roles in care decisions.
  • Motivation: LLM evaluations have predominantly emphasized epistemic performance, including correctness, harmfulness, and bias, rather than subjective value prioritization.The introduction identifies medical explanation, diagnostic reasoning, triage, and clinical decision support as expanding LLM applications.
  • Motivation: Patients and caregivers increasingly act as information brokers in rare disease settings because diagnostic odysseys, fragmented expertise, and uneven specialist access limit reliance on clinical authority.The passage also notes substantial willingness to use ChatGPT for self-diagnosis and health-related decision-making.
  • Motivation: Rare disease care continuously implicates justice because substantial resources serve small populations amid uncertain benefit and disproportionate healthcare utilization and financial burden.Examples include prolonged diagnostic workups, repeated specialist encounters, hospital use, high-cost interventions, and delayed diagnosis.
  • Contributions: The study presents the first clinically grounded, large-scale framework for probing ethical value alignment in rare disease contexts.Its benchmark uses 208 distinct vignettes derived from authoritative Orphanet and OMIM data to evaluate 11 state-of-the-art LLMs and characterize latent value prioritization.

Methods

The study used a five-phase pipeline combining rare-disease data curation, vignette generation and validation, forced-choice LLM evaluation, and statistical analysis. Methods emphasized clinical grounding, reproducible ethical-value coding, and diverse model selection.

  • Study workflow: The methodology comprised five phases: rare-disease curation, vignette construction, validation, LLM evaluation, and downstream statistical analysis.The workflow was designed to support clinical grounding, ethical validity, and analytical interpretability.
  • Rare-disease curation: 2,444 diseases with complete gene, phenotype, and age-of-onset data were filtered by removing 181 with unresolved or unknown inheritance, yielding 2,263 diseases.Disease entities were curated from Orphanet and OMIM, with harmonized gene symbols and inheritance patterns.
  • Vignette generation and validation: Draft vignettes were generated from structured disease metadata and passed through diversity screening, simulated clinician and bioethicist review, and human feasibility auditing.Retained vignettes underwent value assignment and review by six medical or biomedical-background reviewers.
  • Ethical-value coding: Ethical annotations treated Autonomy, Beneficence, Nonmaleficence, and Justice as action-level values, with Justice further divided into four allocation logics.The Justice subtypes included greatest clinical need, correcting disadvantage, equal distribution, and maximizing aggregate benefit under scarcity.
  • LLM evaluation: Each of 11 LLMs evaluated all 208 vignettes using standardized JSON inputs and a forced choice between two candidate actions, A or B.The model set was diversified across providers to improve generalizability.
  • Statistical analysis: Cramér’s V was used descriptively for contextual associations, while regressions reported odds ratios comparing Medical Team and Individual vignettes with Committee vignettes.The authority-framing regressions used the GPT closed-source subgroup, selected for its lowest inter-model heterogeneity (Cramer’s V = 0.0305, p = 0.979).

Results

Across all 11 evaluated models, Justice dominated ethical value selection, with normalized win rates of 57%–70% and a strong skew toward equality-based allocation. Decision-maker framing was the strongest contextual factor: committee contexts elicited Justice, whereas individual and clinical-team contexts increased Autonomy and Beneficence selection.

  • Overall value preferences: 57%–70%: Justice had the highest normalized win rate across all 11 evaluated models, spanning proprietary and open-weight systems.DeepSeek-V3 and GPT-OSS-20B reached 70%, while GPT-4o, GPT-4.1, and Claude 4 were near 69%.
  • Secondary value preferences: 31%–58%: Nonmaleficence showed the widest secondary-value range across models, while Autonomy was consistently preferred over Beneficence in direct conflicts.Nonmaleficence ranged from 31% in Qwen3-14B to 58% in DeepSeek-V3.
  • Justice subcategories: Equality, defined as equal resource distribution, formed the largest component of Justice-oriented selections, exceeding Need- and Equity-based Justice.The decomposition included Need, Equity, Equality, and Maximum Overall Benefit.
  • Contextual factors: 0.504: Decision-maker framing had the strongest association with selected ethical value, with a large Cramér’s V effect exceeding associations from patient-related factors.The association was statistically significant at p < 0.001, while patient decisional ability and age showed smaller associations.
  • Authority framing: 5.71-fold: Individual decision-makers increased the odds of selecting Autonomy relative to Committee contexts, while Medical Team contexts increased Autonomy odds 3.72-fold.Beneficence was also elevated in Individual and Medical Team contexts, whereas Nonmaleficence did not vary significantly and Justice results were unreliable because of severe sample imbalance.

Discussion

The study’s generalizability is limited by GPT-4.1-generated vignettes, simulated expert critique, starting-distribution constraints, and the absence of repeated sampling. Future work should test alignment with human ethical reasoning using clinical ethicists and rare disease specialists.

  • Limitations: GPT-4.1 generated all vignettes, and clinician and bioethicist critique stages were simulated rather than conducted by actual domain experts.Human review nevertheless showed promising clinical feasibility and correct value labels for the generated vignettes.
  • Limitations: The starting distribution of vignettes may preclude reliable analysis of Justice preferences across authority structures.Filtered-out vignettes presented infeasible or uninteresting ethical dilemmas.
  • Limitations: Repeated sampling could provide a robustness check by quantifying within-prompt variability around the reported preference patterns.
  • Future work: Future work should assess whether LLM preferences align with appropriate clinical ethical reasoning and steer models toward decisions resembling patients, families, and rare disease care teams.The proposed evaluation would recruit clinical ethicists and rare disease specialists to assess the same vignettes.

Conclusion

The analysis finds that LLMs converge on justice-oriented reasoning in rare disease decisions, emphasizing equal resource distribution potentially at the expense of disease severity and medical necessity. It warns that deployment may reproduce institutional power asymmetries and constrain ethical deliberation.

  • Core findings: LLMs converge on justice-oriented ethical reasoning across rare disease clinical decision-making.This convergence persists regardless of model architecture or training paradigm.
  • Core findings: Models’ focus on equal resource distribution may come at the expense of disease severity and medical necessity.
  • Benchmark implications: The rare disease benchmark reveals striking similarity across models when evaluating high-stakes ethical value tradeoffs.This contrasts with benchmarks designed around factual accuracy, clinical question-answering, or physician-rubric performance.
  • Risks and deployment: LLM-based decision support may reproduce existing institutional power asymmetries rather than provide stable, principled ethical guidance.
  • Risks and deployment: High-stakes clinical deployment must account for latent biases so AI assistance enhances rather than constrains ethical deliberation.
  • Data availability: The final dataset will contain 208 clinical vignette JSON files and associated disease, symptom, decisional-ability, decision-maker, age-category, and ethical-value information.The dataset is slated for release upon publication acceptance.

Vignette Generation Pipeline

The vignette generation pipeline uses multi-agent iterative refinement to create clinically realistic rare-disease cases featuring genuine ethical conflicts. Simulated clinician and bioethicist reviewers assess clinical defensibility and ethical validity, with refinement continuing until convergence.

  • Pipeline: The pipeline has three stages: initial generation, dual-perspective critique by simulated clinician and bioethicist agents, and iterative refinement until convergence.
  • Vignette design: Each vignette presents a binary A-vs.-B choice in which both options are clinically defensible and support one value while harming another.The design requires irreconcilable stakeholder obligations and excludes obviously good or bad medicine.
  • Vignette design: Vignettes specify the decision-maker, avoid ethical labels, use no more than 5 sentences, and end with “What should he/she/they do?”
  • Refinement: Up to five refinement cycles incorporate both agents’ critiques to preserve value conflict, neutralize clinical considerations, avoid trivial decisions, and maintain realism.A diversity gatekeeper rejects cases too similar to existing vignettes in setting, intervention type, or ethical structure.

Example Vignette (Actual Data)

A 35-year-old woman with KIF1A-related Autosomal Spastic Paraplegia type 30 requests experimental gene therapy after counseling despite concerns that it could accelerate decline or cause lasting harm. The dilemma pits patient autonomy and deliberation against nonmaleficence, with supportive care offering no disease-modifying effect.

  • Ethical dilemma: The ethical conflict is autonomy–deliberation versus nonmaleficence because therapy might accelerate decline or cause lasting harm.The vignette identifies these values as being in conflict and contrasts the request with clinicians’ concern about harm.
  • Clinical scenario: The patient has worsening leg stiffness, frequent falls, and increasing coordination difficulty.She is diagnosed with Autosomal Spastic Paraplegia type 30 caused by a KIF1A mutation.
  • Clinical scenario: After thorough counseling, she requests experimental gene therapy and is described as capable of understanding its complex uncertainties.The neurologist and clinical ethics committee recognize her decision-making capacity.
  • Decision options: The choices are to proceed with the requested therapy or recommend against it and continue supportive care only.Supportive care remains available but will not alter disease progression.

Model Group Analysis Supporting Forest Plot · Model Selection

The model-group analysis selected groups for logistic regression by assessing inter-model variation in ethical choices with Cramér’s V. Lower V indicates greater within-group consistency and more reliable aggregated analysis.

  • Model Selection: Cramér’s V evaluated inter-model variation for logistic regression model-group selection.The analysis assessed consistency among models within each group.
  • Model Selection: Lower Cramér’s V indicated greater consistency among models within a group.Greater consistency makes aggregated analysis more reliable.
  • Model Selection: Aggregated analysis was considered more reliable for groups with greater within-group consistency.Consistency was operationalized through lower Cramér’s V values.
  • Model Group Analysis Supporting Forest Plot: The model-group table included group size, Cramér’s V, interpretation, and p-value fields.The supplied table passage provides the column labels but no corresponding values.
  • Model Selection: Cramér’s V was computed between model identity and selected ethical value category.This quantified the association between model identity and ethical-category choice.
Loading 2608.25236v1…