Source-linked AI summary
A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making
Zhang Sheng, Jinming Li, Wangyang Chen, Zhiwei Bao, Yu YoSean Wang
TL;DR
Medical knowledge scores do not establish reliable oncology decision-path navigation, especially when models must choose among pathways under uncertainty. This study builds and applies ODBB with deterministic, clinician-validated scoring to nine frontier LLMs, finding a substantial shared failure boundary and implications for deployment architecture.
Problem
Existing clinical LLM benchmarks largely test factual recall and do not establish whether frontier models share decision-path blind spots or whether combining models can overcome them.
Method
The study evaluates nine frontier LLMs on 2,005 NCCN-guideline and colorectal-cancer case decision points using a deterministic scorer with 14 typed failure labels and clinician validation.
Results
42% of items were answered correctly by none of the nine models, indicating a collective boundary concentrated in cross-pathway navigation rather than a single-model shortcoming.
Takeaways & Limitations
Clinical deployment should not treat any single model as the sole basis for a decision; systems should detect competence boundaries and route decisions to clinicians.
Takeaways & Limitations
The evaluation is a single May 2026 snapshot using nine models and United States NCCN standards, so boundaries may shift and generalizability requires study.
Abstract
from arXiv · showhide
Large language models (LLMs) achieve high scores on medical knowledge examinations, yet real-world oncology is not a knowledge test--it is a sequence of guideline-pathway choices, escalation judgments, and commitments under uncertainty. Existing benchmarks largely measure factual recall, leaving open whether frontier LLMs share decision-path blind spots that combining models cannot fix. We built the Oncology Decision Boundary Benchmark (ODBB)--2,005 oncology decision points across NCCN guidelines and colorectal cancer cases--and evaluated nine frontier LLMs (four closed-source, five open-weight families) released between June 2025 and April 2026. A fully deterministic scorer (zero LLM inference) classified outputs into 14 failure types, independently validated by two oncologists (Cohen's weighted $κ$ = 0.939 and 0.790) on a 225-item stratified sample. Treating the nine as a pooled super-model, 42.1% (Wilson 95% CI 40.0--44.3%) of all items--35.7% of the 1,586 NCCN items and 66.4% of the 419 colorectal-cancer cases--were answered correctly by none, with failures concentrated in choosing between guideline pathways before reasoning within any: a consistent blind spot in clinical meta-judgment that likely requires architectural intervention rather than more training data. Two models tuned for decisiveness (GPT-5.5, Gemini 3.1 Pro Preview) made unsafe commitments three to five times more often than the seven cautious models without scoring higher. In 3--9% of items, models stated the correct next clinical step yet did not commit to it--failures of decision, not knowledge. Model quality is no longer the primary bottleneck for clinical LLM deployment; the binding constraint is the assumption that any single model can be the sole basis for a clinical decision. Progress requires architectures that detect when a model reaches its competence boundary and route the decision to a clinician.
1 Introduction
High performance on medical knowledge tests does not establish reliable oncology decision-path navigation. The study therefore examines shared capability boundaries, using ODBB to test whether model diversity can overcome them.
- 1 Introduction: Medical knowledge benchmarks have encouraged optimism about LLM clinical utility and clinical-grade reasoning.Models routinely exceed passing thresholds on USMLE-style assessments and perform highly on curated medical benchmarks.
- 1 Introduction: Oncology practice requires selecting the applicable guideline pathway, identifying missing information, routing alternatives, and distinguishing action from restraint.These tasks demand meta-decision competence: choosing the reasoning framework before reasoning within it.
- 1 Introduction: Existing evaluations emphasize factual recall, often use LLM-as-judge scoring, and rarely disaggregate safety-relevant failure modes.They also leave collective blind spots across frontier models insufficiently characterized.
- 1 Introduction: The study asks where the collective capability boundary lies, what it comprises, and whether model diversity can breach it.This shifts the focus from identifying the single best LLM to characterizing shared decision-making limits.
- 1 Introduction: ODBB contains 2,005 oncology decision points across 1,586 NCCN guideline items and 419 colorectal-cancer case items.Nine frontier LLMs—four closed-source and five open-weight families—were evaluated with a deterministic pipeline and 14 typed failure labels.
- 1 Introduction: The benchmark and scorer are designed to be regenerable, reproducible, clinician-validated, and suitable for analyzing collective failure patterns.The scorer uses zero LLM inference and was validated against two oncologists on a 225-item stratified sample.
2 Methods
The methods combine a two-track oncology benchmark, automated item generation, deterministic scoring, typed failure attribution, and ensemble analyses. The design separates normative NCCN pathway decisions from reproduction of treatments documented in colorectal-cancer case reports.
- 2 Methods: ODBB evaluates 2,005 scorable items across structured NCCN decisions and colorectal-cancer case-based therapy recommendations.The NCCN track covers 1,586 items from 69 guideline workspaces.
- 2 Methods: An eight-stage pipeline converts raw NCCN PDFs into structured, schema-validated decision items and reruns automatically after guideline updates.Artifacts are hash-locked, with explicit invalidation matrices; five stages use one auxiliary LLM for semantic judgments.
- 2 Methods: Six question types test distinct oncology decision flavors, including option disambiguation, missing-information requests, and in-guide handoffs.Parallel option disambiguation is the modal task at n = 920, while missing-information requests number n = 348 and in-guide handoffs n = 136.
- 2 Methods: The NCCN track tests guideline-conformant next-step decisions, whereas the CRC track tests narrative case interpretation and specific regimen recommendation.Together they represent normative guideline mastery and free-text reproduction of case-report treatment decisions.
- 2 Methods: CRC was selected because its literature supports sample size, regimen normalization, biomarker-driven branching, and specialist clinical auditing.Extension to lung, breast, and hematologic malignancies is reserved for future work.
- 2 Methods: CRC gold answers are the regimens actually administered in published cases, not necessarily NCCN-recommended regimens.Because case reports over-represent atypical, refractory, or experimental management, NCCN and CRC results are reported separately.
- 2 Methods: The scorer deterministically validates, repairs, aligns, and scores structured outputs while assigning typed clinical and technical failures.Decision alignment separates stop-versus-proceed intent from partial content scoring; every verdict is bit-for-bit reproducible and auditable.
- Quantifying the collective capability boundary.: Collective capability was quantified through zero-correct items, greedy-optimal ensemble coverage, and pairwise verdict concordance across 36 model pairs.Low coverage growth combined with high concordance indicates redundancy and a shared blind spot.
3 Results
ODBB evaluates nine frontier LLMs using an automated, reproducible benchmark and reveals statistically separated performance tiers, with model rankings that do not alone capture the broader clinical decision landscape.
- Benchmark and evaluation: 2,005 clinical decision items were generated through an automated, reproducible pipeline spanning 1,586 NCCN guideline items and 419 colorectal cancer case-based items.The pipeline alternates LLM-augmented semantic judgment with deterministic structural assembly and can be rerun after NCCN guideline updates.
- Headline ranking: Claude Sonnet 4.6 ranked first on NCCN strict concordance at 0.372, followed by GLM-5.1 at 0.325 and a cluster of models through GPT-5.5.The ranking is based on the proportion of NCCN items receiving a correct deterministic-scorer verdict.
- Statistical separation: Permutation testing identified five statistically distinct performance tiers, with Claude Sonnet 4.6 alone in Tier 1 and five models forming an indistinguishable second-place plateau.The plateau comprised GLM-5.1, GLM-5, Qwen 3.6-Plus, Minimax M2.7, and GPT-5.5.
- Cost-performance landscape: The cost-performance frontier contained three models: GLM-5, GLM-5.1, and Claude Sonnet 4.6.Two frontier models were open-weight, while Claude Sonnet 4.6 was the only closed-source model on the frontier.
3.3 The 42% wall: a collective capability boundary
Across nine models, 42.1% of ODBB items were solved by none, and the failures were concentrated in cross-pathway navigation rather than within-pathway reasoning.
- The collective boundary: 42.1% of all items, or 845 of 2,005, were answered correctly by none of the nine models.The zero-correct rate was 35.7% for NCCN items and 66.4% for colorectal cancer case items.
- The collective boundary: 56.9% of zero-correct items received partial verdicts from all nine models, but the mean content score was only 0.09.These partial responses generally touched peripheral information rather than approaching the correct clinical reasoning path.
- The collective boundary: Nine-model ensemble coverage reached only 57.9%, after marginal gains declined sharply beyond the first three models.The best single model covered 34.2%; adding two models increased coverage to 50.7%, while models 4 through 9 added 7.2 percentage points collectively.
- Cross-pathway navigation: In-guide handoff resolution had a 77.2% zero-correct rate, while upstream routing had a 64.5% zero-correct rate.Both task types require recognizing that a case should move to a different guideline pathway.
- Cross-pathway navigation: The best model achieved only 14.0% correct on in-guide handoff and 19.1% on upstream routing, whereas seven of nine models exceeded 40% on parallel option disambiguation.This contrast separates cross-pathway navigation from within-pathway reasoning as distinct decision difficulties.
3.5 The cost of climbing the wall: decisiveness versus safety
Greater decisiveness did not improve performance and instead carried substantially higher safety risk. GPT-5.5 and Gemini 3.1 Pro Preview were more willing to act, but their hard-failure rates exceeded those of the seven conservative models by approximately three- to fivefold.
- Decisiveness versus safety: 20.6% and 20.9% hard-failure rates for GPT-5.5 and Gemini 3.1 Pro Preview compared with 4.2%–7.7% for the seven conservative models.Hard failures combined unsafe overreach and premature commitment on 1,586 NCCN items.
- Decisiveness versus safety: Gemini 3.1 Pro Preview’s NCCN strict concordance fell 31% from Gemini 2.5 Pro’s 0.230 to 0.158 while unsafe overreach and premature commitment increased.Its “stop—need evidence” decisions collapsed from 155 to 4, showing that greater decisiveness was not a safety improvement.
3.6 The knowledge-to-action gap
The knowledge-to-action gap captures cases where models identify the correct next clinical step in their reasoning but fail to make that decision. It was largest in cautious models, while aggressive models showed smaller gaps partly because they committed more readily, often incorrectly.
- Definition and prevalence: 3.1%–9.1% of items showed a knowledge-to-action gap between reasoning that mentioned the correct step and a final decision that omitted it.The deterministic scorer identified this as a decision failure despite relevant clinical knowledge appearing in the reasoning.
- Cautious-model failures: 9.1%, 8.4%, and 8.3% were the largest knowledge-to-action gaps, occurring in Qwen 3.6-Plus, GLM-5, and Minimax M2.7.Common triggers included declining to choose among parallel regimens and listing options instead of selecting one after identifying biomarker discriminators.
- Aggressive-model contrast: Gemini 3.1 Pro Preview’s 3.1% gap reflected more frequent commitment, often to incorrect actions, rather than better translation of knowledge into decisions.The findings distinguish action rate from accuracy and note that they may be inversely correlated.
- Training implication: The paper proposes treating when not to commit as an explicit training capability because wrong clinical commitments can cost more than recommending further consultation.Expected-cost-aware reward shaping is presented as one possible intervention, with the existing gap defining improvement reachable without new clinical knowledge.
3.7 Within the wall: rank reversals and irreplaceability
Model rankings and strengths varied across NCCN guideline navigation and colorectal-cancer cases, showing that overall rank does not capture all useful capabilities. The strongest model also displayed a distinct safety paradox involving contraindications and contradiction sensitivity.
- Rank reversals: GPT-5.5, Gemini 2.5 Pro, and Gemini 3.1 Pro Preview improved their ranks from NCCN to colorectal-cancer evaluation.The reverse pattern was seen for GLM-5, which fell from NCCN rank 3rd to CRC rank 9th.
- Irreplaceability: 57 unique solves for Minimax M2.7 and 52 for GPT-5.5 made them the least replaceable ensemble members despite neither ranking in the overall top three.Models also had distinct question-type specialties, so ensemble value was not determined by overall rank alone.
- Safety paradox: Claude Sonnet 4.6 combined the highest overall score, 0.372, with the highest CRC contraindicated-recommendation rate, 6.0%.It also had the lowest contradictory-output rate, 1.4%, versus 10.6% for GPT-5.5.
- Contradiction sensitivity: Claude’s results suggest that output coherence and commitment confidence may coincide with reduced sensitivity to contradictions in clinical inputs.The paper presents this as a mechanistic hypothesis, not a definitive causal finding.
- Deployment interpretation: High accuracy and high safety are treated as independent dimensions, motivating separate evaluation of contradiction sensitivity and commitment strength in low-uncertainty settings.The paper argues that procurement should not use one dimension as a proxy for the other.
3.8 Cost-effectiveness and the open-weight Pareto frontier
The cost-performance frontier included two open-weight models and one closed-source model, with open-weight systems offering competitive scores at much lower evaluation cost. Open-weight models also clustered behaviorally around more conservative and lower-risk decisions.
- Pareto frontier: GLM-5, GLM-5.1, and Claude Sonnet 4.6 occupied the cost-performance Pareto frontier.Their cost-score pairs were $2.07 and 0.322, $3.73 and 0.325, and $20.68 and 0.372, respectively.
- Cost-performance tradeoff: Claude Sonnet 4.6 achieved a 14.5% score advantage over GLM-5.1 at 5.5× the cost.It also held a 15.5% advantage over GLM-5 at 10.0× the cost.
- Behavioral clustering: Pairwise verdict-level concordance exceeded 0.77 for three open-weight model pairs, reaching 0.801 for GLM-5.1/Qwen 3.6-Plus.The open-weight cluster also showed lower hard-failure rates of 4.2%–7.7% versus 20.6%–20.9% for aggressive closed-source models.
3.9 Clinician validation
Clinician adjudication supported the deterministic scorer’s clinical defensibility, including the incorrect verdicts underlying the reported collective boundary.
- Agreement: 93.1% of 245 evaluable scorer–clinician pairs agreed with the scorer’s verdict.Two oncologists independently reviewed a stratified sample; five unsure or blank judgments were excluded.
- Agreement: Weighted κ was 0.939 for Doctor A and 0.790 for Doctor B.Both estimates used linear weights and bootstrap confidence intervals.
- Boundary validity: Incorrect verdicts received unanimous clinician endorsement in the adjudication sample, 9/9.This supports interpreting the 42% boundary as genuine model failure rather than scorer false positives.
- Item validity: The study separately scrutinized whether generated items and gold answers on the zero-correct subset were clinically defensible.The resulting bounds placed gold-error risk below the level needed to explain the 42% boundary as a generation artifact.
4 Related work
Prior clinical LLM evaluations emphasize factual recall, rubric-scored responses, or binary task outcomes; ODBB addresses decision-path validity with typed, deterministic scoring.
- Existing paradigms: Medical multiple-choice benchmarks inherit exam-authority validity but mainly measure fact retrieval rather than multi-step decision-path navigation.Examples include MedQA, PubMedQA, USMLE-derived benchmarks, and MedBench.
- Existing paradigms: Rubric-based evaluations use clinician judgments of open-ended responses across axes such as consensus, harm, bias, and helpfulness.HealthBench scales this physician-graded paradigm across specialties while retaining response-level rubric scoring.
- Existing paradigms: Structured clinical benchmarks move closer to practice but commonly collapse partial credit and safety tradeoffs into binary correct-versus-incorrect labels.This obscures decisiveness–safety tradeoffs and knowledge-to-action gaps relevant to deployment risk.
- Existing paradigms: Reasoning-model evaluations improve MedQA-style scores but remain tied to single-answer multiple-choice scoring and lack systematic decision-path assessment.They therefore do not directly test commitment among multiple internally permissible clinical pathways.
- ODBB’s contribution: ODBB adds structured verdicts with 14 typed failure labels and removes LLM-as-judge inference for bit-for-bit reproducibility.Its design targets the failure distributions and reproducibility gaps left by prior evaluation approaches.
5 Discussion
The discussion reframes clinical LLM evaluation from ranking models to locating a collective capability boundary and designing deployment systems around it.
- Boundary versus ranking: 42.1% of items formed a collective failure boundary, while ensemble coverage saturated at 57.9%.The result indicates that substantial decision-making lies beyond the evaluated model class rather than merely beyond one model.
- Deployment architecture: A boundary-aware architecture would combine per-item boundary classification, specialist routing, and clinician escalation.The proposed components are intended to identify out-of-envelope items and route in-envelope items to suitable models.
- Deployment architecture: The deterministic scorer, clinician adjudication, and regenerable benchmark provide infrastructure for training and validating a boundary classifier.The typed failure distribution supplies a candidate feature space, while weighted κ supports scorer defensibility.
- Failure modes: All nine models struggled disproportionately with cross-pathway navigation, a meta-judgment task not optimized by standard within-context training objectives.The failure concerns selecting which pathway applies before reasoning within a pathway.
- Failure modes: Decisive models had 27%–33% decisive rates versus 5.2%–9.0% for the conservative majority, although about 84% of NCCN items required requesting information rather than acting.Greater decisiveness therefore conflicts with the stopping behavior required by most benchmark items.
- Failure modes: In 3%–9% of items, reasoning contained the correct next step while the final decision failed to commit to it.This knowledge-to-action gap was largest in the most cautious open-weight models.
- Failure modes: The three failure modes interact: reducing non-commitment through greater agency can close the knowledge-to-action gap by increasing wrong commitments.The discussion therefore treats them as structurally coupled rather than independently solvable.
- Model selection: The best open-weight model scored 0.325 versus 0.372 for the best closed-source model, alongside a 5.5× cost difference.The comparison is presented with the closed model’s higher contraindicated-recommendation rate and open-weight deployment advantages.
6 Conclusion
Across 2,005 oncology decisions and nine models, the paper finds a structural collective boundary: many items remain unsolved even after pooling models. It therefore favors boundary-aware routing and separate safety evaluation over selecting a single more capable or more agentic model.
- Headline finding: 42% of 2,005 oncology decision items were answered correctly by none of nine models, and pooling all nine solved fewer than 58%.The evaluation used a deterministic, clinician-validated scoring framework.
- Interpretation: The collective boundary was primarily composed of cross-pathway navigation rather than reasoning within an already selected framework.The paper characterizes this as a meta-judgment that current training paradigms do not optimize for.
- Deployment implications: Agentic tuning increased hard-failure rates three to five times without increasing accuracy, while 6%–9% of failures were knowledge-to-action gaps.The latter pattern is framed as a reward-shaping target rather than a need for more training data.
- Deployment implications: The proposed path forward is to recognize competence boundaries, route work through fit-for-purpose systems, and evaluate accuracy and safety independently.The conclusion rejects relying on larger base models, more agentic policies, or higher leaderboard scores alone.
- Resources: The released benchmark, scorer, model outputs, and adjudication data support cross-benchmark validation and adversarial replication.The reported 42% wall is presented as a measurement of the current frontier, not an ultimate ceiling for clinical AI.
Declarations
The paper reports no specific external funding or competing interests, uses no human subjects or identifiable patient data, and openly releases its benchmark, scoring, outputs, adjudication data, and analysis materials.
- No specific grant supported the research, and the authors declare no competing interests.
- ODBB items derive from published NCCN guidelines and colorectal cancer case reports without human subjects or identifiable patient data.
- The complete 2,005-point benchmark, deterministic scorer source, nine-model raw outputs, and clinician adjudication data are openly released.
- The scoring engine, benchmark-generation pipeline, and analysis scripts for the paper’s figures and tables are also openly released.