Source-linked AI summary
HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats
Rebecca Soskin Hicks, Mikhail Trofimov, Dominick Lim, Rahul K. Arora, Foivos Tsimpourlas, Preston Bowman, Michael Sharman, Chi Tong, Kavin Karthik, Arnav Dugar, Akshay Jagadeesh, Khaled Saab, Johannes Heidecke, Ashley Alexander, Nate Gross, Karan Singhal
TL;DR
Existing evaluations provide limited evidence about how language models perform in realistic clinician-model conversations and common clinical workflows. HealthBench Professional addresses this gap with a physician-authored, rubric-graded benchmark spanning three use cases and difficulty-enriched examples. GPT-5.4 in ChatGPT for Clinicians achieves the highest reported overall score, exceeding base GPT-5.4, other models, and human physician responses.
Problem
Existing benchmarks often use narrow, single-turn, synthetic, or saturated tasks, providing limited measurement of frontier models in realistic clinician workflows.
Method
HealthBench Professional evaluates 525 physician-authored clinician tasks across three use cases using rubrics iteratively reviewed and adjudicated by three or more physicians.
Results
GPT-5.4 in ChatGPT for Clinicians scored 59.0 versus 48.1 for GPT-5.4 and 43.7 for physician-written responses, outperforming all evaluated comparisons overall.
Takeaways & Limitations
The benchmark provides a focused measure for tracking frontier-model progress on real clinician chat tasks and supporting development of systems intended to help clinicians deliver better care.
Takeaways & Limitations
The benchmark emphasizes intentionally difficult cases and excludes workflows such as EHR-integrated and institution-specific settings, so it is not a proxy for average-case clinical utility.
Abstract
from arXiv · showhide
Millions of clinicians use ChatGPT to support clinical care, but evaluations of the most common use cases in model-clinician conversations are limited. We introduce HealthBench Professional, an open benchmark for evaluating large language models on real tasks that clinicians bring to ChatGPT in the course of their work. The benchmark is organized around three common use cases central to clinical practice: care consult, writing and documentation, and medical research. Each example includes a physician-authored conversation with ChatGPT for Clinicians and is scored via rubrics written and iteratively adjudicated by three or more physicians across three phases. HealthBench Professional examples were carefully selected for quality, representativeness, and difficulty for OpenAI's current frontier models, to enable continued measurement of progress. Difficult examples for recent OpenAI models were enriched by roughly 3.5 times relative to the candidate pool of 15,079 examples. Additionally, about one-third of examples involve physicians conducting deliberate adversarial testing of models. As a strong baseline, we also collected human physician responses for all tasks (unbounded time, specialist-matched, web access). The best scoring system, GPT-5.4 in ChatGPT for Clinicians, outperforms base GPT-5.4, all other models, and human physicians. We hope HealthBench Professional provides the healthcare AI community a measure to track frontier model progress in real-world clinical tasks and build systems that clinicians can trust to improve care.
1. Introduction
HealthBench Professional addresses limited evaluation of realistic clinician-model conversations by introducing a physician-authored benchmark spanning three common clinical use cases. It is designed to be trustworthy and challenging, and GPT-5.4 in ChatGPT for Clinicians achieves the strongest reported performance.
- Motivation: Millions of clinicians use ChatGPT for difficult cases, documentation, and medical evidence, creating a need to measure performance and safety.The paper frames clinical use as both an opportunity and a responsibility for evaluation.
- Evaluation gap: Existing benchmarks often use narrow, single-turn, synthetic, or saturated tasks rather than realistic multi-turn clinician workflows.These limitations reduce measurement signal for frontier models assisting clinicians in practice.
- Benchmark: HealthBench Professional contains 525 rubric-graded tasks across care consult, writing and documentation, and medical research.The benchmark is designed to complement broader health evaluations with a narrower focus on clinician-relevant professional work.
- Benchmark: Each example comes from a physician testing ChatGPT for Clinicians, with criteria written and adjudicated by three or more physicians across three phases.The dataset combines good-faith clinical use with deliberate adversarial testing and is selected for quality, representativeness, and difficulty.
- Results: GPT-5.4 in ChatGPT for Clinicians outperforms base GPT-5.4, all other evaluated models, and human physicians.The benchmark also includes responses from specialty-matched physicians with unbounded time and web access as a strong baseline.
2. HealthBench Professional
HealthBench Professional is an open, rubric-graded benchmark for clinician conversations across three common use cases. Its design emphasizes meaningful real-workflow coverage, trustworthy physician review, and difficulty for frontier models.
- Scope: HealthBench Professional contains 525 physician-authored tasks across care consult, writing and documentation, and medical research.Examples evaluate the next model response in single-turn or multi-turn clinician conversations using example-specific criteria.
- Meaningful: The benchmark targets real clinician workflows rather than synthetic data or structured patient vignettes.Its use cases can affect care decisions directly or support communication, treatment administration, and clinicians’ ability to focus on care.
- Trustworthy: Every conversation and criterion undergoes a three-stage construction and adjudication process involving three or more physicians.The process is intended to reduce evaluation and annotation noise as a source of low scores.
- Challenging: Good-faith examples provide routine workflow coverage, while red-teaming examples target difficult adversarial cases and important failure modes.About one-third of the benchmark comes from red teaming, according to the paper context.
3. Data Collection
The benchmark was built from physician-authored clinician-model conversations using good-faith and adversarial testing, followed by multi-stage review, adjudication, and difficulty-based sampling. Its composition spans specialties and source slices, with variation across use cases.
- Contributor cohort: 190 physicians contributed across 50 countries, 26 specialties, and 52 professional languages.The physician cohort included independent or staff physicians, fellows, and residents.
- Usage modes: Examples were collected in good-faith mode for realistic professional workflows and red-teaming mode for adversarial probing of errors and failure modes.Red-teaming strategies included contradictory premises, obscured intent, emotionally charged framing, and questionable clinical conclusions.
- Review and adjudication: Example creation used physician authorship, physician review, and final adjudication to assess realism, difficulty, rubric quality, and ambiguity.The stages included revising or removing examples and sharpening criteria for consistent evaluation.
- Sampling: Difficult examples for current frontier models were enriched by roughly 3.5x relative to the underlying task distribution.The final benchmark’s sampling emphasizes difficult cases rather than preserving the original distribution.
- Sampling: 15,079 initial examples were tagged by difficulty, usage mode, use case, specialty, language, and conversation turn structure before final sampling.Difficulty was defined as typical for Likert 3–7 and difficult for Likert 1–2.
- Composition: Source-slice proportions vary across use cases, including 64.1% red-teaming difficult examples in writing and documentation and 85.7% good-faith typical examples in medical research.Such variation can contribute to different subscores by use case and specialty.
4. Scoring
HealthBench Professional scores responses with physician-written rubrics, length adjustment, aggregation across examples, and model-based grading. The evaluation compares frontier models, clinician-specific systems, other models, and physician responses across clinical use cases.
- Rubric-based scoring: Each response is scored against physician-written criteria that reward desired facts or behaviors and penalize undesirable ones.Criteria receive nonzero point values from −10 to 10, including negative points for unsafe or otherwise undesirable behavior.
- Rubric-based scoring: A model-based grader evaluates each rubric criterion independently, assigning its full point value when met and zero otherwise, including for negative criteria.Negative criteria assign penalties when the undesirable behavior is present.
- Rubric-based scoring: The raw example score sums the point values of met criteria, including penalties, and divides by the example’s maximum possible score.An individual example score can be negative when penalties exceed positive points.
- Length adjustment: Length adjustment estimates a linear score slope from model responses across verbosity settings, averaging slopes across eligible model-reasoning pairs.The adjustment is estimated only within a selected response-length regime to avoid skew from unusually long responses.
- Length adjustment: 2.94 × 10−5 per character is the primary length-adjustment coefficient, with 4,000 characters used as the estimation threshold.The coefficient is approximately 0.0147 per 500 characters.
- Length adjustment: Responses of 2,000 characters receive no adjustment; longer responses are penalized and shorter responses receive a corresponding positive adjustment.The adjustment preserves scores similar in magnitude to unadjusted scores around this common response-length range and does not affect relative rankings.
- Aggregate scoring: The benchmark score is the clipped mean of per-example length-adjusted scores, reported after multiplying scores by 100 for readability.Subset scores use the same aggregation rule over the relevant examples, and the score is not percent accuracy.
- Systems and comparisons: The evaluation reports overall and use-case-specific scores while comparing frontier models, clinician-focused and baseline harnesses, other models, and physician responses.Main comparisons use eight samples per example, averaged across runs; physician responses are evaluated identically to model responses.
5. Results
GPT-5.4 in ChatGPT for Clinicians achieved the strongest overall performance, with advantages varying across use cases, dataset slices, specialties, system configurations, verbosity, and reasoning effort.
- Overall rubric performance: 59.0 was the overall score for GPT-5.4 in ChatGPT for Clinicians, exceeding physician-written responses at 43.7 and base GPT-5.4 at 48.1.The system also outperformed every other evaluated model overall.
- Performance by use case: GPT-5.4 in ChatGPT for Clinicians outperformed physicians across care consult, writing and documentation, and medical research, with its largest physician-relative advantage in writing and documentation.Scores were 51.0 vs 42.7, 64.1 vs 32.1, and 67.0 vs 56.3, respectively.
- Performance by dataset slice: 55.8 versus 30.0 for physicians on the RT 1–2 slice marked the clearest dataset-slice advantage, while gains were less uniform on good-faith slices.The system beat every other model on RT 1–2, but was statistically tied with the field on Good-faith 1–2.
- Performance by specialty: GPT-5.4 in ChatGPT for Clinicians had the highest mean score in 6 of 10 plotted specialties, with significant physician-relative advantages in nephrology and cardiology.Specialty gains were broad but possibly larger in some areas than others.
- System and harness comparisons: 59.0 versus 48.1 for base GPT-5.4 and 45.8 with browsing showed that ChatGPT for Clinicians improved overall performance beyond both comparison configurations.The largest harness differences occurred in writing and documentation; care-consult improvements were not significant.
- Verbosity and reasoning effort: Increasing reasoning effort improved scores by an average of 3.3 points across GPT-5 to GPT-5.4, while length adjustment removed most verbosity-related gains in the typical response-length range.For GPT-5.4, low-to-xhigh improvements ranged from 5.6 to 7.3 points depending on verbosity; high-verbosity responses often exceeded the penalty-estimation range.
6. Discussion
HealthBench Professional is a physician-vetted benchmark of realistic clinician tasks, deliberately enriched for difficult cases to expose model limitations. Its scores should be interpreted as performance under intentionally difficult conditions, not average-case clinical utility.
- HealthBench Professional evaluates realistic clinician tasks across care consult, writing and documentation, and medical research.Its examples come from clinician testing of ChatGPT for Clinicians and use open-ended chat responses graded with physician-written rubrics.
- Three phases of physician vetting and a physician-written reference response support evaluation quality and comparison.
- The benchmark enriches difficult examples through adversarial testing and upsampling failures found during broad testing.This selection emphasizes edge cases, adversarial scenarios, and subtle failure modes that are less common in routine use.
- Benchmark scores are not directly comparable to real-world performance rates because difficult examples are overrepresented.A moderate aggregate score can coexist with high performance in typical usage.
- HealthBench Professional does not cover every clinician workflow, including EHR-integrated workflows and institution-specific constraints.The paper recommends institution-specific pre-deployment evaluation and post-deployment monitoring.
- Future evaluations should address hours- or days-long health tasks, including complex chronic-disease navigation and genomic-data analysis.
7. Related Work
Health-related LLM evaluation is moving from exam-like knowledge tests toward realistic clinical workflows, but many existing benchmarks remain limited for clinician chat. HealthBench Professional extends rubric-based evaluation specifically to clinician-facing conversations across three work domains.
- Many established health benchmarks use exam-like, multiple-choice, single-turn, or short-answer tasks and may be saturated by frontier models.They measure clinical knowledge prerequisites but provide limited signal about assisting clinicians in realistic chat workflows.
- MedAlign, EHRShot, MedCalc-Bench, and LLMEval-Med move closer to clinical practice through instruction following, calculations, checklists, or physician validation.
- Evaluation suites such as MedHELM and MedMarks combine heterogeneous benchmarks spanning multiple clinical and research capabilities.
- HealthBench Professional builds on HealthBench’s realistic multi-turn conversations and physician-written rubrics while focusing on clinician-facing chats.It complements rather than replaces HealthBench across care consult, writing and documentation, and medical research.
8. Data access
HealthBench Professional releases its dataset while using an internal evaluation implementation for reported results. The released records represent conversations, rubric items, use cases, testing type, difficulty, specialty, and physician responses, alongside safeguards against benchmark contamination.
- The HealthBench Professional data are publicly available as an assets.zip archive.
- Each record includes a conversation, rubric items, use case, testing type, difficulty, specialty, and physician response.Difficulty distinguishes difficult Likert scores 1–2 from typical scores 3–7.
- The dataset includes physician-written responses as a record field and requests that examples not be reposted in plain text or images.A canary string is included to help filter the benchmark from training corpora and prevent direct retrieval of ground-truth answers.
- Reported results use an internal evaluation implementation rather than an official external implementation.The stated reason is consistency across product harnesses, base models, and external models; the HealthBench simple-evals implementation remains a reference.
9. Conclusion
HealthBench Professional is an open, rubric-graded benchmark for measuring LLM performance in clinician conversations across three clinical use cases. It is designed with physicians to track frontier-model progress and support development of systems for clinicians.
- HealthBench Professional measures LLMs in clinician conversations across care consult, writing and documentation, and medical research.It builds on HealthBench and uses open rubric-based grading.
- The benchmark is intended to help health systems, startups, and model developers build systems that can help clinicians deliver better patient outcomes.
Research collaborators
The paper acknowledges physicians whose insight, generosity, time, and expertise contributed to HealthBench Professional. It provides a list of consenting physicians while clarifying that participation does not imply endorsement.
- Physicians contributed insight, generosity, time, and expertise to HealthBench Professional.
- The acknowledgments list physicians who consented to be named, including their affiliated institutions or professional credentials where provided.
- The paper states that being acknowledged does not imply physician endorsement.
A. Additional figures and tables
The additional tables describe the benchmark’s specialty composition and the distributions of source slices across use cases and specialties. The latter two tables report percentages conditionally within the relevant use-case or specialty grouping.
- Table 1 reports the specialty composition of examples in HealthBench Professional.
- Table 2 reports the conditional distribution of source slice given use case, with percentages computed within each use case.
- Table 3 reports the conditional distribution of source slice given specialty for the benchmark’s top 10 specialties, with percentages computed within each specialty.