Source-linked AI summary
HealthBench: Evaluating Large Language Models Towards Improved Human Health
Rahul K. Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero-Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, Johannes Heidecke, Karan Singhal
TL;DR
HealthBench addresses limitations in existing health evaluations by providing a realistic, open-ended benchmark grounded in physician-written criteria. It evaluates diverse model interactions and finds rapid recent progress across performance, cost, and reliability, while substantial headroom remains. The benchmark is intended as an open, comprehensive standard for developing safe and beneficial health AI.
Problem
Existing health evaluations often rely on narrow or multiple-choice questions, lack validation against expert medical opinions, and provide insufficient room for improvement.
Method
HealthBench evaluates 5,000 realistic health conversations using 48,562 conversation-specific rubric criteria developed with 262 physicians and model-based grading validated against physician judgment.
Results
Recent models improved rapidly across frontier performance, cost-adjusted performance, and reliability, while significant headroom remains in health-related conversations and workflows.
Takeaways & Limitations
HealthBench offers a comprehensive, extensible standard for evaluating frontier health AI toward safe and beneficial models.
Takeaways & Limitations
HealthBench does not specifically evaluate workflow-level response quality, including workflows that use multiple model responses.
Abstract
from arXiv · showhide
We present HealthBench, an open-source benchmark measuring the performance and safety of large language models in healthcare. HealthBench consists of 5,000 multi-turn conversations between a model and an individual user or healthcare professional. Responses are evaluated using conversation-specific rubrics created by 262 physicians. Unlike previous multiple-choice or short-answer benchmarks, HealthBench enables realistic, open-ended evaluation through 48,562 unique rubric criteria spanning several health contexts (e.g., emergencies, transforming clinical data, global health) and behavioral dimensions (e.g., accuracy, instruction following, communication). HealthBench performance over the last two years reflects steady initial progress (compare GPT-3.5 Turbo's 16% to GPT-4o's 32%) and more rapid recent improvements (o3 scores 60%). Smaller models have especially improved: GPT-4.1 nano outperforms GPT-4o and is 25 times cheaper. We additionally release two HealthBench variations: HealthBench Consensus, which includes 34 particularly important dimensions of model behavior validated via physician consensus, and HealthBench Hard, where the current top score is 32%. We hope that HealthBench grounds progress towards model development and applications that benefit human health.
1 Introduction
HealthBench is an open-source benchmark designed to address limitations in health evaluations by measuring realistic, open-ended model behavior with physician-informed rubrics. It evaluates diverse conversations and behavioral dimensions while tracking performance, cost, reliability, and remaining weaknesses.
- Physician involvement: 262 physicians with experience across 60 countries developed HealthBench, supporting coverage of real-world health contexts and physician judgment.Physicians contributed to benchmark design, data collection, and rubric development.
- Benchmark design: 48,562 unique rubric criteria measure conversation-specific attributes such as clinical accuracy, communication, misconceptions, and relative importance.Criteria are physician-written, self-contained, and scored by a model-based grader validated against physician judgment.
- Evaluation dimensions: HealthBench provides targeted measurement by themes and behavioral axes, enabling diagnosis of model strengths and improvement areas beyond a single overall score.Themes represent health-related task categories, while axes include dimensions such as clinical accuracy, communication quality, and context awareness.
- Findings: Recent models improved across frontier performance, cost, and reliability, but substantial room remains in worst-case performance and context-seeking behavior.The benchmark also compares model results with physician baselines and evaluates model–physician grading agreement against physician–physician agreement.
- Benchmark design: 5,000 realistic health conversations span individual users and healthcare professionals across diverse geographies, languages, and healthcare personas.The benchmark evaluates responses to the last user message in each multi-turn conversation.
- Benchmark variations: HealthBench Consensus measures 34 important behavior aspects validated by multiple physicians, while HealthBench Hard remains challenging, with no evaluated model scoring above 32%.These variations extend the benchmark toward consensus-validated dimensions and harder evaluation settings.
2 Rubric evaluations
HealthBench grades open-ended responses against conversation-specific physician-written rubrics, aggregating criterion-level rewards and penalties into clipped mean scores.
- Each evaluation example pairs a conversation ending in a user message with rubric criteria specific to that conversation.Criteria may specify facts to include or undesirable behaviors to penalize, with nonzero values from −10 to 10.
- A model-based grader independently determines whether each response meets each rubric criterion.Met criteria receive full points, while unmet positive criteria receive none and met negative criteria incur penalties.
- Example scores sum criterion values, including positive rewards and negative penalties, then divide by the maximum possible score.
- Overall HealthBench performance is the mean of per-example scores clipped to the range [0, 1].Subset analyses aggregate scores within the relevant examples or criteria as specified.
3 Organization of HealthBench
HealthBench organizes 5,000 health-interaction examples into themes and behavior axes, with broad example-specific rubrics complemented by physician-validated consensus criteria and a difficult subset.
- 5,000 examples contain single- or multi-turn conversations averaging 2.6 turns and 668 characters, with ranges of one to nineteen turns and four to 9,853 characters.
- Seven themes cover real-world health areas, while 48,562 unique rubric criteria are partitioned into five behavioral axes.The median example has eleven criteria, with two to 48 criteria per example.
- 34 consensus criteria appear 8,053 times and are assigned only when a majority of reviewing physicians agree they apply.They target narrow performance dimensions and support comparison between model grading and physician grading.
- HealthBench Consensus retains 3,671 examples with at least one positive consensus criterion and filters evaluation to consensus criteria.It offers higher physician-validation precision but lower recall for detecting model behavior failures.
- HealthBench Hard is a 1,000-example subset selected for difficulty, with many current models scoring zero.It is designed as a challenging target for future models.
4 Data collection
HealthBench was built with broad physician involvement and realistic, varied conversations sourced through synthetic generation, red teaming, and transformed health queries before rubric annotation and meta-evaluation.
- 4.1 Physician cohort, selection, and input: 262 physicians across 60 countries and 26 specialties created HealthBench tasks over 11 months, with fluency in 49 languages represented.The cohort included physicians across career stages and practice settings.
- 4.1 Physician cohort, selection, and input: Physician advisors shaped health domains, desirable responses, prompt seeds, consensus criteria, contributor selection, training, and quality review.
- 4.2 Conversation production: Examples depict user-model interactions beginning and ending with user messages, including alternating messages in multi-turn conversations.
- 4.2 Conversation production: Most conversations were synthetically generated through a physician-informed pipeline designed around important real-world health situations.Additional data came from physician red teaming and transformed HealthSearchQA queries.
- 4.2 Conversation production: Generated conversations were filtered for realism, self-consistency, physical-health relevance, and complete messages using a model check.
- 4.3 Annotation: Physicians wrote theme-specific rubrics, categorized examples, and added consensus criteria when more than half of at least two raters agreed.Physicians also evaluated model responses against consensus criteria to measure grader agreement with physician opinion.
5 Themes and axes
HealthBench spans seven real-world health themes and evaluates responses across five behavioral axes, enabling performance breakdowns by task context and model behavior.
- 5.1 Themes: Seven themes assess important health interactions, including emergencies, missing context, global health, health data, expertise-tailored communication, uncertainty, and response depth.
- 5.1 Themes: Emergency referrals measure recognition and escalation of urgent situations while accounting for harms from both under- and over-escalation.
- 5.1 Themes: Context-seeking evaluates whether models identify missing clinical information and request the most informative additional context.
- 5.1 Themes: Global health evaluates adaptation to differences in healthcare resources, clinical norms, regional disease patterns, and local contexts.
- 5.1 Themes: Health data tasks assess safe, accurate completion of structured clinical-data work, while expertise-tailored communication assesses adaptation to user role and needs.
- 5.1 Themes: Responding under uncertainty and response depth evaluate calibrated handling of ambiguity and matching detail to user needs and tasks.
- 5.2 Axes: Five axes measure accuracy, completeness, communication quality, context awareness, and instruction following across themes.The axes provide behavioral dimensions for consistent comparison, and Table 3 distinguishes consensus from example-specific criteria.
6 Results
HealthBench results show rapid recent gains in model performance, cost-efficiency, and reliability, while revealing persistent weaknesses in context-seeking, completeness, and difficult cases. Performance varies substantially across health themes and behavioral axes, and the benchmark remains unsaturated.
- Performance over time: 28%: OpenAI frontier-model performance improved in recent months, exceeding the gain between GPT-3.5 Turbo and GPT-4o.The comparison is based on HealthBench performance over time.
- Performance by theme: 0.16 to 0.60: Overall HealthBench scores span GPT-3.5 Turbo to o3, with newer models generally performing best.Emergency referrals and expertise-tailored communication are strongest, while context-seeking, health data tasks, and global health lag behind.
- Inference-time cost: 25x cheaper: GPT-4.1 nano outperforms August 2024’s GPT-4o, while April 2025 models establish a new performance-cost frontier.The frontier includes o3, o4-mini, and GPT-4.1.
- Reliability: More than double: o3’s worst-at-16 score exceeds GPT-4o’s, yet o3’s score falls by a third from its 60% overall result.Worst-at-k measures the average minimum score across subsets of k independent responses, characterizing reliability deterioration with more samples.
- Behavioral axes: 4x: Model error rates fell from GPT-3.5 through GPT-4.1, with especially notable improvement in emergency situations.Context-seeking, responding under uncertainty, and appropriate response depth remain areas for improvement.
- HealthBench Hard: 0.32: o3 scores substantially lower on HealthBench Hard’s 1,000 difficult examples than on HealthBench overall.The lower score provides headroom for future models.
7 Physician-written responses
The physician baseline study compares de novo responses with responses written using older or newer model references. Physicians improved older-model references but not newer ones, while de novo performance was relatively weak and response length influenced scores.
- Study design: Three physician groups differed by access to no AI assistance, August/September 2024 model references, or April 2025 model references.Reference groups received four model responses and were encouraged to copy-paste and improve them.
- Improvement of references: 56.2% vs 39.8%: With September 2024 references, physicians more often improved than worsened the reference responses.The comparison examines physician-written responses against the mean score of the four provided references.
- Physician responses without references: De novo physician performance appeared relatively weak, but the authors caution that writing these responses is not an ordinary physician task.Human baselines depend heavily on the framing and instructions of the task.
- Response length: Shorter responses partly explain lower de novo physician scores, while physicians using April 2025 references maintained similar performance with greater concision.April 2025 model references were typically longer than the corresponding physician responses despite similar scores.
8 Trustworthiness of model grading
HealthBench evaluates whether model-based grading agrees with physician judgment using consensus criteria, physician baselines, and repeated runs. GPT-4.1 grading reached physician-level agreement and scores showed low variability, supporting its use as the default grader under specified conditions.
- Evaluation metric: Macro F1 averages positive- and negative-class F1 scores equally, balancing sensitivity to false positives and false negatives.HealthBench treats grading as binary classification of whether each criterion is met.
- Meta-evaluation results: GPT-4.1 exceeded the random baseline for all themes and the average physician score in five of seven themes.It ranked in the upper half of physicians for six of seven themes and above the 33rd percentile for all themes.
- Grader comparison: GPT-4.1 was the best-performing April 2025 grader; o4-mini and o3 were slightly worse, while GPT-4.1 mini and nano were substantially worse.The difference may partly reflect GPT-4.1’s use during prompt tuning.
- Implications: Reliable model-based grading depends on diverse, well-annotated ground truth, well-designed meta-evaluation, and careful prompt and grader selection.Because GPT-4.1 achieved physician-level agreement without reasoning-model costs and latency, it was adopted as the default grader.
- Score variability: 16 runs: HealthBench scores had a standard deviation of around 0.002 across models scoring from 0.16 to 0.60.The repeated-run analysis includes variability in generated responses and grading criteria.
9 Discussion
HealthBench is presented as a broad, realistic benchmark with substantial remaining headroom, but its interpretation is constrained by evaluator disagreement, incomplete example criteria, physician-baseline limitations, and changing real-world workflows.
- Findings: HealthBench measures diverse, realistic model interactions and reports improved performance, cost-adjusted performance, and reliability while retaining significant headroom.The authors recommend it for research iteration and comparing deployed models.
- Evaluation quality: Physician-physician and model-physician agreement on consensus criteria both range from 55% to 75%.Variation may reflect ambiguity, specialization, risk tolerance, perceived severity, communication style, and instruction interpretation.
- Baselines: Physician-written responses are a cautious baseline because physicians do not ordinarily perform this task in the same format as the benchmark.Responses were produced with additional resources and ongoing guidance, and differed from model responses in length and presentation.
- Scope: HealthBench does not evaluate specific workflows or their health outcomes, which depend critically on implementation as well as response quality.The benchmark evaluates single model responses to multi-turn conversations, whereas workflows may use multiple responses.
10 Related work
Prior health benchmarks often used narrow or closed-form tasks that poorly represent realistic interactions, lack extensive expert validation, or become saturated. HealthBench builds on rubric-based evaluation to target these gaps with broader natural-language coverage.
- Healthcare LLMs: Prior work has examined LLM performance on a range of targeted healthcare tasks.The related-work discussion situates HealthBench within growing interest in applying LLMs to healthcare.
- Prior benchmarks: Earlier medical LLM evaluations focused largely on exam questions that were narrow, unlike real workflows, and eventually saturated.These evaluations were initially useful for research iteration because they were easy to run.
- Remaining gaps: Many current evaluations remain too narrow or unrepresentative, insufficiently validated against diverse experts, or saturated by frontier models.HealthBench explicitly targets meaningfulness, trustworthiness, and unsaturation.
- Rubric evaluation: Previous health benchmarks graded short responses, extracted fixed entities, or used weakly correlated language-generation metrics.HealthBench instead uses physician-generated, example-specific rubrics covering multiple desired behaviors.
11 Conclusion
The paper frames HealthBench as a shared resource for guiding AI-health evaluation and development. Its goals span community standards, evidence for healthcare stakeholders, and tracking progress in frontier performance, cost, and reliability.
- Goals: HealthBench aims to shape shared AI research standards that incentivize progress toward models creating real-world human benefit.This is the paper’s first stated goal.
- Goals: HealthBench aims to provide healthcare communities with evidence about current and future model capabilities, use cases, and limitations.This is the paper’s second stated goal.
- Goals: HealthBench aims to share rapid progress in frontier performance, cost, and reliability.This is the paper’s third stated goal.
- Implication: The authors hope HealthBench accelerates development of models and real-world applications that realize AI’s potential to improve human health.The conclusion states this as the intended broader direction.
12 Code and data
The release provides code for running HealthBench, its meta-evaluation and variants, and associated analyses. The surrounding materials document physician-cohort selection and request that benchmark examples not beเผย online to reduce leakage and cheating.
- Code: The released code loads data and runs HealthBench to compute overall model scores.The code is described as simple to use.
- Code: The code supports meta-evaluation of model-based graders and running HealthBench Hard and HealthBench Consensus.These are listed as separate supported functions.
- Analyses: The code supports examining individual-example grading, model-provided explanations, and analyses such as theme, axis, and worst-at-k plots.These capabilities accompany benchmark scoring.
- Data handling: The authors ask users not to reveal dataset examples online to reduce training-data leakage and cheating by models with internet access.A canary string is included to help filter the benchmark from training corpora.
C HealthBench Hard example selection
HealthBench Hard selects 1,000 examples that are difficult across multiple frontier models while avoiding cases that are universally unscorable or disproportionately difficult because of rubric variability. The section also describes the example-scoring procedure and eligibility checks used in dataset construction.
- HealthBench Hard selection: 1,000 examples form HealthBench Hard, selected for difficulty across today’s frontier models.The selection aimed to avoid examples that were especially adversarial for particular models or unnecessarily difficult because of physician rubric-writing variability.
- HealthBench Hard selection: Five models—o3, Grok 3, Gemini 2.5 Pro, Claude 3.7 Sonnet, and Llama 4 Maverick—provide the model-provider scores used for selection.
- HealthBench Hard selection: About 1.5% of examples were filtered because no model received a positive score.This step was intended to avoid oversampling problems that may be overly difficult.
- HealthBench Hard selection: The final 1,000 examples had the lowest average score across model providers, making selection adversarial across models rather than tailored to one model.The procedure used the average rather than the minimum or maximum score.
- Scoring procedure: Each example is scored by generating a response, grading each rubric criterion, and dividing met-criterion points by the example’s maximum possible points.The overall evaluation score is the clipped mean of example scores, and axis scores restrict the calculation to applicable criteria.
- Eligibility checks: Eligibility checks classify conversations for completeness, realism, consistency, physical health, and mental-health relevance before inclusion.The prompts require exact categorical outputs and exclude conversations failing the relevant eligibility criteria.
F Example situation types used for synthetic data generation
Synthetic data generation covers health inquiries ranging from simple, fully specified questions to complex cases where missing context makes a precise or safe answer dependent on further information. Physicians then craft ideal chatbot responses with emphasis on factual content, helpful communication, and safety, while additional criteria organize important behaviors such as emergency referrals.
- Context supplied: Synthetic situations include simple, moderate, and complex medical inquiries where all necessary context is supplied across one or multiple turns.
- Context needed: Some situations require more context because an immediate response could be inaccurate or harmful, especially for treatment, medication, lifestyle, or health decisions.
- Context needed: Other context-dependent cases omit important details, use ambiguous or extremely brief input, or involve treatments whose appropriateness depends on subtype or severity.
- Context supplied: The dataset also includes brief or subtle informational queries whose answers can be given from the details provided.
- Physician response creation: Physicians write ideal next responses for chatbot users, prioritizing factual content and effective communication while considering what is helpful and safe.Reference-response groups may use supplied responses, whereas de novo physicians may use the internet for factual verification but not AI tools for writing.
- Consensus criteria: HealthBench Consensus organizes 34 criteria by behavioral themes and example categories, including emergency behavior and context-seeking.For emergent cases, models should directly refer users to emergency care early; conditional cases require referrals tied to specified circumstances.