Source-linked AI summary
Position: Medical AI Neglects Real Treatment Outcomes
Shiva Kaul, Anjum Khurshid
TL;DR
Medical AI largely learns from published text and human judgments rather than real treatment outcomes, limiting its understanding of treatment. This position paper argues for incorporating outcome data into training and evaluation, with improving real treatment outcomes as medical AI’s long-term goal.
Problem
Medical AI training and evaluation largely neglect real treatment outcomes, relying instead on published syntheses and human opinions despite their central role in understanding treatment effects.
Method
The paper develops a position through analysis of medical AI training corpora, treatment-information sources, and benchmark deficiencies.
Results
The paper concludes that real treatment outcomes from observational databases and randomized experiments should be substantially incorporated into medical AI training and evaluation.
Takeaways & Limitations
Medical AI’s long-term goal should be to improve real treatment outcomes and help write clinical practice guidelines rather than merely read and reference them.
Takeaways & Limitations
Most observational medical databases are not yet ready for large-scale AI training because they are fragmented, difficult to access, and semantically interoperable only with additional infrastructure.
Abstract
from arXiv · showhide
Medical AI has rapidly improved its ability to perform diagnostic and prognostic tasks that lead to treatment decisions. But understanding of treatment itself is still inadequately trained and evaluated, using human opinions and syntheses (especially texts such as biomedical publications and clinical practice guidelines) rather than actual underlying data on treatment outcomes. This neglect seriously limits the potential of medical AI, and is already causing deficiencies in both frontier models and major benchmarks, as argued in this position paper. Real treatment outcomes, drawn from sources such as observational databases and randomized experiments, should be substantially incorporated into both training and evaluation. Improving these outcomes should be reemphasized as the downstream goal of all medical AI.
1. Introduction
Medical AI focuses on diagnosis and prognosis while inadequately learning or evaluating real treatment outcomes. The paper argues for incorporating observed patient outcomes into training and evaluation, with improving those outcomes as medical AI’s long-term goal.
- Medical AI emphasizes intermediate predictive tasks such as diagnosis and prognosis without analyzing downstream improvements in patient outcomes.
- Language models lack a rich understanding of treatment effects because causal reasoning is outsourced to coarse preexisting works that can mask uncertainty and heterogeneity.
- Real treatment outcomes—actual outcomes of real patients observed with their prior treatment—are neglected from medical AI training through evaluation.
- The paper recommends using more real treatment-outcome data and frames improving those outcomes as medical AI’s long-term purpose.
- The paper examines outcome data in training, limitations of human-written syntheses as ground truth, and benchmarks’ emphasis on ancillary or intermediate tasks.
2. Doctors Train From Clinical Observation of Treatment; AI Does Not
Physicians develop clinical competence by observing treatment consequences during residency, whereas modern medical AI largely trains on text and generally lacks real treatment outcomes. The key distinction is whether training includes longitudinal treatment outcomes, with electronic-health-record models identified as the main foundation-model exception.
- 2. Doctors Train From Clinical Observation of Treatment; AI Does Not: Physicians observe patients’ treatment consequences during clinical residency, making real treatment outcomes a crucial part of physician training.Residency typically lasts as long as academic medical schooling and directly exposes physicians to intervention consequences.
- 2. Doctors Train From Clinical Observation of Treatment; AI Does Not: Most medical AI training uses textual materials, especially biomedical publications and clinical practice guidelines.This pattern partly reflects models’ origins in base language models trained primarily on web data.
- 2. Doctors Train From Clinical Observation of Treatment; AI Does Not: Real patient data are common in medical AI training, but subsequent evaluations often use synthetic cases rather than real patient data.Examples of real patient data include clinical notes and radiological images, while many models are fine-tuned on medical question-answering datasets.
- 2. Doctors Train From Clinical Observation of Treatment; AI Does Not: The major training watershed is inclusion versus exclusion of real treatment outcomes, with electronic-health-record foundation models the only major class known to include them.EHR models represent sequences of medical events rather than language, and treatment outcomes are inherently longitudinal.
3. The Use of Opinions and Syntheses as Ground Truth About Treatments
Medication labels, clinical practice guidelines, and clinician agreement are convenient but imperfect proxies for treatment ground truth. They omit important clinical variation, may be stale or oversimplified, and can produce high agreement without establishing accuracy.
- Medication labels: Medication labels constrain manufacturers’ marketing rather than define standards of care, omit off-label use, and may misrepresent treatment ground truth.Off-label use constitutes 30-80% of medication use in some fields.
- Medication labels: An incorrect answer passed both automated RxQA correctness checks, demonstrating that label-based validation can fail even on a real treatment case.The example involved ivacaftor resolving a patient’s pancreatitis, while the incorrect answer passed both verifications.
- Clinical practice guidelines: Guidelines are nonbinding professional standards that oversimplify complex, stochastic care and are updated infrequently.The supplied examples include successful treatment by reducing pill burden and an OpenEvidence plan that ignored important precision laboratory measurements.
- Clinician opinion: Because clinicians do not observe treated patients’ counterfactual outcomes, independently accumulating treatment expertise is fundamentally limited.This problem is especially acute for evidence-based methodology in medicine because of causal inference.
- Clinician opinion: Agreement with clinicians is not ground truth: 55,546 labels (91.2%) came from examples labeled by exactly two clinicians, and agreement can coexist with low accuracy.For 5,105 HealthBench factual-accuracy metaexamples, alternative ground-truth assignments produce materially different clinician and GPT 4.1 accuracy rates.
- Clinical practice guidelines: Real-world clinician guideline concordance generally averages only 55% to 60%, underscoring the gap between written recommendations and treatment practice.The supplied passage notes that discordance may reflect outdated training, lack of knowledge, or more nuanced personalization.
4. Treatment Is the Focus of Medicine, But Not of AI Evaluations
Medical care exists to improve outcomes through treatment, yet medical AI evaluations largely assess biomedical knowledge, clinical reasoning, or treatment-related proxies rather than causal understanding grounded in real treatment outcomes. Even benchmarks using clinical-trial data often prioritize therapeutic development or operational endpoints over comparative effectiveness and patient-centered benefit.
- Medicine’s downstream goal: Medical care’s purpose is improving outcomes through treatment, whose initiation is reserved to licensed physicians and increasingly tied to value-based reimbursement.This framing is reflected in law and health economics, including US policies after the Affordable Care Act and MACRA.
- What evaluations should measure: Treatment-outcome evaluation asks realistic cases for recommendations or quantitative judgments about treatment effects, effectiveness, or comparative quality.These formats include free-form next steps, regression, ranking, and classification.
- Current evaluation gap: 41.5% of 1277 verifiable JAMA Clinical Challenge questions ask for next steps of treatment, but only approximately 18.1% involve some assessment of treatment outcomes.The repository contains 1700+ questions, with 530 asking about treatment next steps and approximately 232 assessing outcomes.
- Current evaluation gap: Only 16/4609 reusable benchmarks involve real outcomes, while textual syntheses or opinions remain pervasive as ground truth even in nominally treatment-related evaluations.A systematic review found 789/4609 clinical evaluations nominally involved treatment planning or recommendation, but scrutiny found far fewer reusable benchmarks with real outcomes.
- Limits of clinical-trial benchmarks: Clinical-trial benchmarks use crucial outcome data but only partially measure treatment understanding because they emphasize therapeutic-development costs rather than comparative effectiveness.They lack focus on determining when one treatment is superior to another, and some outcome definitions reflect enrollment, stock-price, or administrative factors.
- Downstream outcome misalignment: Screening examples show that optimizing intermediate predictions such as diagnosis or sensitivity can conflict with patient-centered benefit when downstream harms are ignored.Randomized trials found several screening tests less worthwhile than anticipated, while overdiagnosis and data-acquisition risks can cause harm.
5. Recommendations (Call to Action)
The paper calls for medical AI training to rely more on longitudinal patient records and evaluation to rely more on randomized experiments. It also recommends quantifying how intermediate predictive improvements translate into treatment outcomes and patient-oriented priorities.
- Training: Training should rely more on longitudinal patient records, including electronic health records and insurance claims, than on textual syntheses.Large-scale longitudinal data pipelines are common in real-world evidence, and datasets such as CRITICAL are increasingly available to researchers.
- Evaluation: Evaluation should rely more on randomized experiments and predict results of previously conducted experiments rather than regulatory documents, guidelines, or clinician opinion.New benchmarks can be formulated as target trial emulations, while involving AI in new experiments may create ethical and logistical challenges.
- Intermediate predictive tasks: Intermediate predictive tasks should continue, but their contribution to improving treatment outcomes should be quantified at least roughly.Ideally, deployments of predictive models would be analyzed as causal interventions while reporting outcome-oriented metrics.
- Intermediate predictive tasks: Intermediate evaluation should connect algorithmic metrics to clinician performance, clinical experience, and practitioner well-being.For ambient AI summarization, NLP metrics such as Word Error Rate and Mean Text Recurrence can be complemented by downstream evaluations.
6. Alternative Views
The paper addresses concerns about privacy, data readiness, evidence quality, and clinical deployment by arguing for ethical, gradual use of treatment outcomes. It maintains that emphasizing outcomes can improve medical AI without diminishing intermediate tasks.
- Privacy: Treatment outcomes are not inherently more sensitive than existing individual-level medical data and can be learned ethically through group-level inference and differential privacy.Medical AI already trains on sensitive patient data while respecting prevailing biomedical standards.
- Data readiness: Observational medical databases remain fragmented and semantically non-interoperable, so most are not yet ready for large-scale AI training despite standardized formats and API-access requirements.Current provisions ensure syntactic, not semantic, interoperability.
- Implementation: The recommendation is a gradual shift toward real treatment outcomes rather than an extreme replacement of publications and other syntheses.The transition may occur over the course of years.
- Evidence quality: Randomized controlled trials have internal- and external-validity limitations, but trial-design biases can be disentangled from underlying effects, while individual patient data can support individual-effect inference.Identified limitations include low power, blinding violations, dropouts, and highly artificial inclusion criteria.
- Intermediate tasks: Emphasizing treatment outcomes need not diminish intermediate medical-AI tasks, paralleling evidence-based medicine’s historical shift toward patient-centered outcomes.The paper states that this shift did not imperil in-silico or in-vitro biomedical investigation.
7. Conclusion
The paper concludes that medical AI’s reliance on published text and human judgments reflects practices borrowed from other AI rather than medicine’s customs or ambitions. Its long-term priority should be real treatment outcomes.
- Conclusion: Medical AI should reemphasize real treatment outcomes as what truly matters for its long-term development.The authors argue that published text and human ground truth do not embody medicine’s customs or ambitions.
A. Appendix · A.1. Related Work
Related work supports emphasizing real clinical outcomes in medical AI while distinguishing this paper’s focus on learning treatment effects from work on safely deploying predictive systems. Prior research also highlights causal modeling, evaluation validity, observational databases, trial emulation, and health-policy metrics as relevant foundations.
- A.1. Related Work: Joshi et al. (2025) also emphasize clinical outcomes, but study causal analysis of predictive-AI deployments rather than training AI to understand existing treatment effects.Their example concerns deployment risks when a heart-disease model is trained on a smaller, biased population.
- A.1. Related Work: This paper argues that substitutes for real treatment-outcome data, including clinical guidelines, do not facilitate training medical AI to understand treatment effects.The target systems include reasoning language models and the treatments include pharmaceuticals.
- A.1. Related Work: The two perspectives are complementary: safe-deployment frameworks could govern the advocated medical AI, while treatment-effect understanding could help manage deployment problems such as ECG distribution shift.The paper explicitly states that it sees no conflict between the approaches.
- A.1. Related Work: Causal machine learning has been broadly advocated in medicine, and several deep-learning models predict treatment effects, although relatively few language models do so.Causal reasoning can also support diagnosis through pathophysiology, because diseases cause symptoms.
- A.1. Related Work: Only roughly 5% of external, published medical-AI evaluations were found to involve real patient data, amid broader criticism of unrealistic benchmarks and weak construct validity.The cited critiques question whether benchmarks measure what they claim to measure.
- A.1. Related Work: Because patient data are private and geographically scattered, large-scale observational learning has relied on federated learning, collaborative research networks, and regulatory postmarketing-surveillance systems.OHDSI is cited as an example of a collaborative research network, and FDA Sentinel as a regulatory example.
- A.1. Related Work: Prior benchmarks assess treatment-outcome prediction through target-trial emulation, propensity-score methods on observational data, or prediction of future trials from previous meta-analyses.These approaches examine whether algorithms can reproduce or anticipate randomized-trial outcomes.
- A.1. Related Work: Health-policy and health-economics research uses outcome-oriented metrics including quality-adjusted life years, event-free life years, mortality indices, and incremental cost-effectiveness ratios.Such holistic measures may be preferable to formulaic aggregation measures such as mean win rate.
A.2. Experiments with RxQA Prompts
The RxQA experiments show that automated verification grounded primarily in FDA product labels can accept incorrect answers and reject correct ones, across multiple Gemini models. These counterexamples are scientifically uncertain and borderline, but they target the asymmetric requirement that benchmark questions and treatment plans be correct and effective.
- A.2.1. BYPASSING AUTOMATED RXQA CORRECTNESS VERIFICATION: RxQA prompts ask a model to verify whether an answer derived from medication information correctly answers a question.The prompt explicitly presents medication information, a question, and a proposed answer for verification.
- A.2.1. BYPASSING AUTOMATED RXQA CORRECTNESS VERIFICATION: Palepu et al. (2025) used automated correctness and clarity checks before sending passing RxQA questions to pharmacists for manual revision.The workflow combined automated verification with an additional clarity assessment and pharmacist review.
- A.2.1. BYPASSING AUTOMATED RXQA CORRECTNESS VERIFICATION: Gemini 1.5 Flash falsely verified an incorrect answer about successful ivacaftor treatment, with Gemini 3.5 Flash and Gemini 3.1 Pro Preview also failing.The correct answer was (C), but the model accepted an incorrect choice, indicating a verification-method problem based primarily on the FDA product label.
- A.2.1. BYPASSING AUTOMATED RXQA CORRECTNESS VERIFICATION: The prescribing information incorrectly led verification to classify answer (D), “Ineffective,” as correct while rejecting the clinically relevant alternative.The explanation treated pancreatitis as insufficient for contraindication because the label did not list it as a contraindication.
- A.2.1. BYPASSING AUTOMATED RXQA CORRECTNESS VERIFICATION: Gemini 1.5 Flash falsely eliminated the correct answer (C), and Gemini 3.5 Flash and Gemini 3.1 Pro Preview also failed to identify it.The experiment used a question about a patient successfully treated with ivacaftor.
- A.2.2. GENERAL DISCLAIMER ABOUT CASE EXAMPLES: The examples involve substantial scientific uncertainty and borderline cases, but the authors argue that benchmarks must generate correct answers and models must offer treatment plans that work well.The authors emphasize an unusually and asymmetrically high burden of proof for these claims.
A.2.3. ABLATION: VERIFICATION SUCCEEDS WITH ADDITIONAL INFORMATION
Verification succeeds when models receive additional patient-specific information alongside treatment evidence: they distinguish a successful ivacaftor case from a different patient for whom the drug is ineffective.
- Verification with additional information: Gemini 1.5 Flash, Gemini 3.5 Flash, and Gemini 3.1 Pro Preview correctly answered the positive case and the negative case when the case report was provided.The negative-case result indicates that the successful case report did not overwhelm the models’ clinical reasoning.
- Verification with additional information: The correct answer for the negative case is (D) Ineffective.The case describes a 42-year-old man with chronic calcific pancreatitis, rather than the published successful-treatment patient.
- Verification with additional information: The published case involved idiopathic chronic pancreatitis, methylmalonic acidemia, and a CFTR carrier whose drug-responsive variant supported successful ivacaftor treatment.The additional methylmalonic acidemia risk factor likely contributed to pancreatitis, while the CFTR variant responded to ivacaftor.
A.2.4. ABLATION: SIMILAR RESULTS FOR COMBINATION THERAPY … A.5. Systematic Literature Analysis Methodology and Results
The paper finds that medical AI can verify incorrect treatment answers, miss clinically important medication reasoning, and rely on stale evidence. Its benchmark analysis further shows limited reusability and systematic classification of treatment evaluations by ground-truth source.
- A.2.4. ABLATION: SIMILAR RESULTS FOR COMBINATION THERAPY: The Trikafta ablation suggests incorrect verification was not specific to ivacaftor, because Gemini 1.5 Flash again accepted a likely incorrect answer (D).Trikafta combines elexacaftor, tezacaftor, and ivacaftor; the authors describe the ablation as uncertain but informative.
- A.2.4. ABLATION: SIMILAR RESULTS FOR COMBINATION THERAPY: Gemini 1.5 Flash also rejected the likely correct answer (C) in a Trikafta variant, despite the medication guide indicating only one correct answer.The variant replaced ivacaftor with Trikafta.
- A.3.1. PRESCRIBING CASCADE EXAMPLE: ChatGPT for Clinicians recognized medication adherence but missed the underlying prescribing cascade and did not recommend reducing furosemide.The case involved an 87-year-old woman taking furosemide and potassium chloride among several medications.
- A.3.1. PRESCRIBING CASCADE EXAMPLE: ChatGPT for Healthcare likewise missed the prescribing cascade and explicitly suggested temporarily increasing furosemide instead of reducing it.It addressed medication adherence but recommended the opposite loop-diuretic adjustment.
- A.3.2. STALENESS EXAMPLE: An OpenEvidence example based primarily on older, cheaper tests ignored some modern laboratory results relevant to the clinical reasoning underlying its treatment plan.The authors state that their scrutiny targets the reasoning, not necessarily the correctness of the final treatment plan.
- A.4. Analysis of HealthBench Professional: 881/(1135−150) ≈89% of HealthBench Professional criteria were classified as grounded, with manual verification of 150 filtered cases and an assumed maximum 3% error rate.The grounding estimate depends on correctly filtering non-clinical criteria and grounding clinical criteria.
- A.5. Systematic Literature Analysis Methodology and Results: The systematic review classified 789 nominally treatment-related studies by treatment focus, benchmark reusability, and six categories of ground truth, with classifications manually adjudicated or verified.Ground-truth categories included guidelines, expert opinion, product labels, observational outcomes, randomized-trial outcomes, and other or unsure.
- A.5. Systematic Literature Analysis Methodology and Results: 509 studies were not reusable, comprising 111 treatment-focused studies and 398 studies in the other category shown in the reported table.The accompanying ground-truth table distinguishes all studies from reusable treatment benchmarks.