Source-linked AI summary

Large Language Models Show Metacognitive Sensitivity in Medical Reasoning

Ahmad Nazzal

arXiv:2608.14552v1cs.AI

TL;DR

Medical-LLM evaluations rarely isolate whether confidence tracks evidence quality and uncertainty. This study uses a controlled psychophysics-inspired benchmark and finds partially evidence-sensitive confidence alongside localized overconfidence in conflicting Alzheimer-type cases.

  • Problem

    Medical-LLM research needs controlled evidence on whether confidence reflects evidence strength and uncertainty, not merely diagnostic accuracy.

  • Method

    The study evaluates diagnostic choice and confidence in a medical LLM using controlled psychophysics-inspired vignettes varying evidence strength, conflict, and missing information.

  • Results

    Confidence increased with evidence strength, decreased when key information was missing, and remained higher on correct than incorrect trials, but errors and overconfidence clustered in conflicting AT-NCD cases.

  • Takeaways & Limitations

    Controlled evidence manipulations can distinguish diagnostic sensitivity from second-order confidence behavior more clearly than aggregate benchmark accuracy alone.

  • Takeaways & Limitations

    The synthetic benchmark centers on one narrow diagnostic contrast, so its results should not be generalized to clinical reasoning as a whole.

Abstract

from arXiv · show

Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidence tracks evidence quality and uncertainty. We developed a controlled, psychophysics-inspired clinical benchmark to test diagnostic choice and confidence behavior in a medical LLM. The benchmark focused on probable Alzheimer-type neurocognitive disorder (AT-NCD) versus depression-related cognitive impairment (DRCI). We generated 45 synthetic vignettes varying evidence strength, conflicting evidence, and missing information. Each vignette was presented under three prompt variants, yielding 135 trials. In a pilot run with gpt-4.1-nano, all trials produced valid structured outputs. Across forced-choice trials, diagnostic accuracy was 93.5%, mean confidence was 78.4%, and AUROC2 was 0.876. Confidence increased with evidence distance from the diagnostic boundary, decreased when information was missing, and remained higher on correct than incorrect trials after adjustment for evidence strength and prompt format. These findings indicate partial metacognitive sensitivity rather than globally uninformative confidence. However, errors clustered in moderate, conflicting AT-NCD cases, where the model shifted toward DRCI and retained more confidence than empirical accuracy justified. Model comparison suggested that confidence quality should be measured directly rather than inferred from benchmark accuracy or model capability alone. This study establishes a reproducible framework for evaluating evidence sensitivity, metacognitive sensitivity, and localized calibration failure in medical LLMs.

1 Introduction

The introduction frames clinically meaningful confidence as metacognitive sensitivity: confidence should track evidence quality, information sufficiency, and correctness rather than accuracy alone. The study therefore introduces a controlled AT-NCD-versus-DRCI benchmark to test these relationships behaviorally and establish a reproducible evaluation framework.

  • Motivation: Clinical reasoning requires recognizing when evidence is insufficient for a confident decision, enabling caution, information gathering, or judgment deferral.Metacognition monitors and evaluates one’s own judgments, supporting safer decisions under uncertainty.
  • Problem: Medical-LLM usefulness cannot be judged by answer accuracy alone because safety, factuality, bias, reliability, and uncertainty recognition also matter.Prior evaluations often test whether models answer correctly without testing whether confidence appropriately represents uncertainty or missing knowledge.
  • Conceptual framework: Metacognitive sensitivity is assessed by whether confidence tracks evidence quality and correctness, not by mean confidence alone.Confidence is treated as a second-order estimate of the probability that a choice is correct.
  • Methodological rationale: Psychophysics provides a controlled alternative by systematically manipulating evidence strength and analyzing threshold, slope, bias, and uncertainty.This approach helps separate confidence effects due to evidence strength, prompt framing, prior bias, and task-specific artifacts.
  • Study aims: The study tests whether diagnostic choices and explicit confidence respond to graded evidence, missing or conflicting information, and correctness in AT-NCD versus DRCI cases.Its narrower behavioral goal is to evaluate confidence tracking of clinical evidence strength, information quality, and correctness, rather than broad philosophical claims about true metacognition.

2 Methods · 2.1 Study design · 2.2 Diagnostic framing

The study used a controlled, psychophysics-inspired benchmark to test whether a medical LLM’s diagnostic choices and confidence respond systematically to clinical evidence. It framed synthetic cases around probable AT-NCD versus DRCI while treating probable AT-NCD as an operational, non-biomarker-confirmed label.

  • 2.1 Study design: The benchmark tested whether diagnostic choices and reported confidence varied systematically with the strength and quality of clinical evidence.The model was evaluated as a behavioral observer rather than through broad medical-knowledge questions.
  • 2.1 Study design: The design manipulated the amount of evidence favoring one of two competing diagnostic interpretations.This followed psychophysics, which studies observer performance under controlled variation in stimulus evidence.
  • 2.1 Study design: The benchmark focused on the clinical differential between probable Alzheimer-type neurocognitive disorder (AT-NCD) and depression-related cognitive impairment (DRCI).This contrast was selected as clinically plausible and suitable for controlled manipulation.
  • 2.1 Study design: The cases were designed to manipulate ambiguity, conflicting evidence, and missing information within the AT-NCD-versus-DRCI differential.These manipulations supported controlled evaluation of diagnostic and confidence behavior.
  • 2.2 Diagnostic framing: Probable Alzheimer-type neurocognitive disorder served as an operational syndrome-level label for generating synthetic vignettes.The label was used to preserve the contrast between progressive Alzheimer-type decline and depression-related cognitive impairment.
  • 2.2 Diagnostic framing: The probable AT-NCD label did not imply biomarker-confirmed Alzheimer’s disease, reflecting a clinically cautious diagnostic framing.The AT-NCD side was informed by the NIA-AA framework for dementia due to Alzheimer’s disease.

2.3 Benchmark construction · 2.4 Feature bank and evidence levels · 2.5 Information-quality conditions

The benchmark used 45 synthetic clinical vignettes with three prompt variants, systematically manipulating diagnostic features, evidence levels, and information quality. Conditions represented coherent, conflicting, or incomplete evidence to test medical-LLM reasoning beyond closed-book recall.

  • 2.3 Benchmark construction: 45 unique synthetic clinical vignettes were paired with 3 prompt variants, yielding 135 total trials.The vignettes formed a factorial benchmark for controlled evaluation.
  • 2.3 Benchmark construction: The predefined feature bank contained clinically plausible cues favoring AT-NCD or DRCI, plus neutral, conflict, and missing-information cues.Internal feature weights supported construction and analysis but were not shown to the model.
  • 2.4 Feature bank and evidence levels: AT-NCD cues included gradual progression, collateral decline reports, impaired delayed recall, poor cue benefit, reduced insight, and instrumental functional decline.These features represented progressive amnestic and functional impairment.
  • 2.4 Feature bank and evidence levels: DRCI cues included low mood, anhedonia, sleep disturbance, fatigue, subjective cognitive distress, variable effort, and concentration difficulties.Additional patterns favored slowed processing or executive dysfunction over stable amnestic storage failure.
  • 2.5 Information-quality conditions: Each evidence level was crossed with three information-quality conditions: clear, conflicting, and missing.Clear vignettes coherently favored one interpretation, conflicting vignettes supported both diagnoses, and missing vignettes omitted informative elements.
  • 2.5 Information-quality conditions: Missing-information vignettes omitted elements such as collateral history, longitudinal information, cognitive testing, or structured mood assessment.The manipulation modeled uncertainty from incomplete evidence rather than weak evidence alone.
  • 2.5 Information-quality conditions: The design addressed calls for medical-LLM evaluation under incomplete or ambiguous clinical information rather than relying only on closed-book multiple-choice recall.This motivation was linked to realistic clinical uncertainty and prior evaluation work.

2.6 Prompt design · 2.7 Model and inference settings · 2.8 Outcome measures

The study standardized prompts, constrained gpt-4.1-nano outputs through Structured Outputs, and separated diagnostic-choice outcomes from confidence-based metacognitive outcomes. Three prompt variants tested format effects while preserving required diagnostic, confidence, and information-need responses.

  • 2.6 Prompt design: Each vignette asked the model to select between two diagnostic interpretations, estimate numerical confidence, and indicate whether more information was needed.
  • 2.6 Prompt design: Three prompt variants tested ordering and naming effects: P01 listed AT-NCD first, P02 listed DRCI first, and P03 used named options with AT-NCD first.
  • 2.7 Model and inference settings: The primary pilot used gpt-4.1-nano through the OpenAI Responses API with stateless calls and no retained conversation history.
  • 2.7 Model and inference settings: Structured Outputs enforced a strict JSON schema containing three required fields, and raw and parsed responses were saved for every trial.
  • 2.8 Outcome measures: First-order outcomes included forced-choice diagnostic accuracy for non-equivocal cases and evidence-sensitive diagnostic choice across the clinical evidence gradient.
  • 2.8 Outcome measures: Second-order outcomes comprised metacognitive bias, metacognitive sensitivity, and exploratory metacognitive efficiency.
  • 2.8 Outcome measures: Metacognitive bias related mean confidence to empirical accuracy, whereas metacognitive sensitivity used AUROC2 primarily and confidence–correctness relations secondarily.

2.9 Calibration and confidence metrics · 2.10 Statistical analysis

The study assessed calibration through confidence–correctness metrics and quantified metacognitive sensitivity primarily with AUROC2. Analyses used condition-level summaries, evidence curves, and a penalized logistic robustness model for diagnostic choice.

  • 2.9 Calibration and confidence metrics: Global calibration was assessed with mean confidence, Brier score, expected calibration error (ECE), and reliability diagrams.These measures quantified the relation between stated confidence and empirical correctness.
  • 2.9 Calibration and confidence metrics: AUROC2 primarily quantified metacognitive sensitivity by measuring confidence discrimination between correct and incorrect responses independently of raw confidence magnitude.The benchmark elicited explicit probability-like confidence judgments.
  • 2.9 Calibration and confidence metrics: Confidence was modeled as a function of correctness after adjustment for evidence strength, information quality, and prompt format.This tested whether confidence tracked correctness beyond these task factors.
  • 2.9 Calibration and confidence metrics: Exploratory signal-detection-theory-based analyses were also performed.These analyses supplemented the primary calibration and metacognitive-sensitivity measures.
  • 2.10 Statistical analysis: Analyses first computed condition-level summaries for parsed response rate, mean confidence, forced-choice accuracy, more_information_needed = true, and information-sufficiency accuracy.All analyses were performed in Python.
  • 2.10 Statistical analysis: Diagnostic choice was summarized with condition-level results and an accuracy-by-evidence curve, with a penalized logistic model fitted as a robustness check.The model predicted AT-NCD choice from evidence score and information-quality condition.

2.11 Equivocal cases · 2.12 Reproducibility and ethics

Equivocal cases were treated as underdetermined, so evaluation focused on whether the model requested additional information and reduced confidence appropriately. The benchmark was reproducible, synthetic, version-controlled, and intended only for model evaluation, not clinical use or medical advice.

  • 2.11 Equivocal cases: Equivocal vignettes were designed to be underdetermined.
  • 2.11 Equivocal cases: Forced-choice accuracy was treated as not applicable for equivocal cases.
  • 2.11 Equivocal cases: The principal equivocal-case outcomes were requesting additional information and reducing confidence appropriately.
  • 2.12 Reproducibility and ethics: The benchmark was designed as a reproducible pilot.
  • 2.12 Reproducibility and ethics: The feature bank, vignette generator, prompt templates, trial manifest, raw outputs, parsed outputs, and analysis notebooks were versioned and frozen after the run.
  • 2.12 Reproducibility and ethics: No patient data were used; all cases were synthetic and intended solely for model evaluation, not clinical use or medical advice.

3 Results

In 135 trials, gpt-4.1-nano showed high diagnostic accuracy and partially informative confidence, with sensitivity to evidence distance and missing information. However, confidence was locally fragile, especially for moderate conflicting AT-NCD cases, and confidence–correctness discrimination varied independently of diagnostic accuracy across models.

  • Overall performance: All 135 trials produced valid structured outputs, including 108 forced-choice and 27 equivocal or underdetermined cases.The remaining equivocal cases assessed whether the model indicated that more information was needed.
  • Overall performance: 93.5% accuracy and 78.4% mean confidence were observed across forced-choice trials, with Brier score = 0.077, ECE = 0.151, and AUROC2 = 0.876.Information-sufficiency judgment accuracy was 83.7%.
  • Localized calibration failure: 44.4% accuracy and 72.8% mean confidence in moderate conflicting AT-NCD cases produced a 28.3-percentage-point overconfidence gap.The corresponding moderate conflicting DRCI condition was classified correctly on all trials.
  • Metacognitive sensitivity: Confidence remained higher on correct than incorrect forced-choice trials after adjustment for evidence strength, information quality, and prompt format (coefficient for correctness = +5.15, p = 0.001).This supports second-order sensitivity beyond confidence tracking evidence magnitude alone.
  • Model comparison: AUROC2 varied across models—0.919 for gpt-5, 0.873 for gpt-4.1-nano, 0.786 for gpt-4.1-mini, and 0.644 for gpt-5-nano—without increasing monotonically with diagnostic accuracy or nominal capability.gpt-5.5 achieved 1.000 forced-choice accuracy, but AUROC2 was undefined because it made no errors.

4 Discussion

This pilot establishes a controlled, psychophysics-inspired framework showing that medical-LLM confidence tracks evidence and correctness, while revealing partial, uneven metacognitive sensitivity and localized diagnostic failures. The discussion emphasizes methodological caution, operational limits of metacognition claims, and broader evaluation and extension priorities.

  • Methodological contribution: The benchmark provides a controlled clinical-evidence paradigm for studying diagnostic choice, uncertainty behavior, and confidence calibration in medical LLMs.Its psychophysics-inspired design reveals patterns that aggregate benchmark scores may obscure.
  • Metacognitive findings: Confidence increased with evidence distance, decreased with missing information, and remained higher on correct than incorrect trials, indicating partial rather than absent metacognitive sensitivity.The model was globally underconfident, yet confidence discriminated correct from incorrect responses; exploratory efficiency analysis suggested weaker second-order signaling than first-order performance alone would predict.
  • Failure pattern: Errors concentrated in Alzheimer-type cases with competing depressive features, with equivocal cases showing a DRCI-leaning default.The design could not distinguish whether depressive cues were overweighted or DRCI served as a lower-commitment default under uncertainty.
  • Limitations and interpretation: The findings should not be generalized broadly because the study used a small synthetic vignette benchmark centered on one diagnostic contrast, with some conditions near ceiling.Confidence was explicit and prompt-elicited, and metacognition was assessed behaviorally rather than as evidence of human-like reflective awareness.
  • Future work: Future studies should expand the AT-NCD versus DRCI dataset around the ambiguity boundary and test more models and diagnostic contrasts.Larger datasets would support more stable calibration, AUROC2, and meta-d′/d′ estimation while enabling comparisons across model families, scale, architecture, and training regime.
  • Evaluation framework: Medical-LLM evaluation should report complementary measures rather than rely on a single score, treating AUROC2 as one component of artificial metacognitive evaluation.Recommended quantities include diagnostic accuracy, calibration error, Brier score, information-sufficiency performance, evidence-linked confidence modulation, and confidence–correctness discrimination.

5 Conclusion

The study introduced a controlled clinical-evidence benchmark for evaluating diagnostic choice, confidence calibration, and information-sufficiency judgments in a medical LLM. Results showed evidence-sensitive diagnostic choices and confidence that was not globally flat or arbitrary.

  • The benchmark evaluated diagnostic choice, confidence calibration, and information-sufficiency judgments in a medical LLM.It was designed as a controlled clinical-evidence benchmark.
  • Confidence increased with evidence strength and decreased when key information was missing.
  • Confidence remained higher on correct than incorrect trials after accounting for evidence quality and prompt format.The findings indicate that confidence was not globally flat or arbitrary.

A Supplementary analyses

Supplementary analyses used penalized logistic modeling and exploratory signal-detection efficiency estimates to support the main interpretation. They were not primary because the pilot sample was small, some cells approached ceiling, and confidence values had a restricted range.

  • Supplementary analyses: Supplementary analyses included a penalized logistic model of AT-NCD choice as a robustness check of evidence-sensitive first-order behavior.The model was used to assess whether the main behavioral interpretation was robust.
  • Supplementary analyses: Exploratory SDT-based efficiency estimates included d′, meta-d′, and meta-d′/d′.These signal-detection measures were exploratory rather than primary outcomes.
  • Supplementary analyses: The supplementary analyses supported the main interpretation but were not primary because the pilot sample was small, several cells were at or near ceiling, and confidence values occupied a restricted range.These limitations reduced the suitability of the supplementary analyses as primary evidence.

B Supplementary analyses

Supplementary analyses confirmed strong evidence sensitivity and substantial confidence–correctness discrimination, while exploratory SDT results indicated that confidence captured only part of the information supporting diagnostic decisions. Model comparisons further suggested that confidence quality did not increase monotonically with nominal model capability.

  • Penalized logistic model: Each onepoint increase in AT-NCD evidence approximately doubled the odds of choosing AT-NCD.The DRCI-first prompt and missing-information cases reduced the odds of an AT-NCD response, while small interactions suggested a shifted decision boundary rather than reduced evidence sensitivity.
  • SDT analyses: Exploratory SDT estimates showed strong diagnostic discrimination, substantial confidence–correctness discrimination, and an M-ratio of approximately 0.31.Meta-d′ was substantially lower than d′, suggesting that confidence captured only part of the information available to the primary diagnostic process.
  • Model comparison: Confidence–correctness discrimination did not improve monotonically with nominal model capability or newer model family.In this benchmark, gpt-4.1-nano showed the highest AUROC2, followed by gpt-4.1-mini, whereas gpt-5-nano showed weaker confidence discrimination; the comparison was exploratory and limited by the small benchmark and one diagnostic contrast.
  • Model comparison: For gpt-5.5, AUROC2 could not be estimated because the model made no forced-choice errors, although confidence still varied across evidence levels and information-quality conditions.This distinguishes confidence modulation by evidence from confidence–correctness discrimination.
Loading 2608.14552v1…