Source-linked AI summary

Do Large Language Models Know What They Don't Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty

Ali Şenol, H. Russell Bernard, Huan Liu

arXiv:2609.07879v1cs.AI

TL;DR

The paper asks whether LLMs appropriately acknowledge the boundaries of their knowledge, a behavior conventional correctness benchmarks do not directly measure. It introduces EHQ and the EHQ-3000 benchmark to evaluate restraint, false-assertion avoidance, and substantive-answer calibration across defined knowledge and context boundaries. Across the analysed panel, EHQ varies substantially despite near-ceiling document-grounded capability, while restraint and calibration show different patterns under important scope and measurement limitations.

  • Problem

    Correctness-centric evaluation does not distinguish confident fabrication from appropriately uncertain answers, leaving epistemic boundary behavior insufficiently measured.

  • Method

    The paper introduces EHQ’s three behavioral sub-scores and evaluates models on EHQ-3000, a 3,000-question benchmark spanning fabricated, post-cutoff, hyper-niche, and context-conditioned questions.

  • Results

    EHQ ranges from 0.3145 (GPT-4o-mini) to 0.8131 (Claude-5-Sonnet), separating models despite the near-ceiling document-grounded capability probe.

  • Takeaways & Limitations

    EHQ reveals behavioral differences not visible in conventional correctness assessment, with restraint, calibration, and boundary type showing distinct structure.

  • Takeaways & Limitations

    EHQ is limited to English, single-turn prompts with verbalized confidence, does not observe internal knowledge states, and has incomplete category-specific controls and review coverage.

Abstract

from arXiv · show

Large Language Models (LLMs) are frequently confident, eloquent, and well versed. A natural question arises: do they know what they don't know? To answer this question, we borrow the concept of epistemic honesty and develop a novel metric to systematically evaluate whether an LLM appropriately acknowledges the boundaries of its knowledge. In this work, we introduce the Epistemic Honesty Quotient (EHQ), which reports three observable sub-scores across two operational axes (epistemic restraint and substantive-answer calibration), and construct EHQ-3000, a 3,000-question benchmark spanning Fabricated Entity, Post-Cutoff Event, Hyper-Niche True, and Context-Conditioned Questions. From a frozen registry of 21 model API routes, 15 completed the protocol after endpoint and eligibility checks; 14 entered the confirmatory analysis because severe provider-side truncation made one route's score indeterminate. The study reveals substantial variation across models, including a difference that can not be explained by their capability to extract explicitly available information. Composite EHQ ranges from 0.31 to 0.81 across the analysed panel, despite near-ceiling performance on the document-grounded capability probe. The two restraint criteria overlap strongly under the present category composition, whereas substantive-answer calibration varies across models and does not reliably co-vary with restraint; however, the small panel leaves substantial uncertainty. Thus, EHQ reveals behavioral differences that are not visible to conventional correctness-based assessment, while also showing why dataset composition, provider behavior, and confidence elicitation must remain part of the interpretation.

1 Introduction

The paper argues that correctness-centric evaluation overlooks whether LLMs should answer at all, motivating epistemic honesty as a distinct behavioral target. It introduces EHQ, EHQ-3000, and a reproducible evaluation design to measure restraint, false-assertion avoidance, and confidence calibration.

  • Motivation: Correctness-centric benchmarks overlook whether models should answer, treating honest abstention and confident fabrication alike when both responses are marked incorrect.The paper frames this as a deployment-risk distinction, especially in high-stakes medical and legal settings.
  • Conceptual foundation: Epistemic honesty concerns asserting what is warranted, qualifying uncertainty, and withholding what the model does not know, rather than factual accuracy alone.The paper distinguishes this behavioral property from correctness and links it to regulating assertion under uncertainty.
  • Metric: EHQ reports three observable sub-scores across restraint, unqualified-false-assertion avoidance, and confidence–correctness alignment.EHQ1 and EHQ2 operationalize control, while EHQ3 measures calibration on substantive answers.
  • Scope: EHQ is an operational composite: its response labels are mutually exclusive, but its sub-scores are not and do not constitute a psychometrically identified decomposition.The paper explicitly limits the construct’s scope beyond the measured response behaviors.
  • Contributions: The evaluation contributes a 21-route registry, deterministic checkpointed protocol, and prespecified statistical plan covering correlations, generational comparisons, uncertainty, permutation inference, and multiplicity control.The protocol verifies endpoint identity, disables retrieval and conversation history, enforces non-reasoning conditions, and hashes scientific artifacts.

2 Related Work

The related-work foundation connects EHQ to metacognitive monitoring and control, calibration theory, and research on representations of cognitive limits. The paper uses these traditions to motivate, rather than presume, a dissociation between capability, restraint, and confidence calibration.

  • Metacognition: Metacognition distinguishes monitoring knowledge states from controlling behavior, which EHQ maps onto confidence alignment and answer-or-abstain decisions.EHQ1 and EHQ2 measure epistemic control, while EHQ3 measures monitoring through stated-confidence alignment.
  • Knowledge limits: A model may produce confident errors when it fails to represent structural unavailability, temporal exclusion, withheld context, or low accessibility appropriately.This motivates examining epistemic boundaries rather than relying on overall accuracy.
  • Calibration: EHQ adapts calibration theory to verbalized confidence on substantive responses within a benchmark-defined boundary set.The restriction targets consequential miscalibration without claiming direct access to internal knowledge states.
  • Research framing: The study tests whether capability, restraint, and substantive-answer calibration dissociate, while treating EHQ1 and EHQ2 as conceptually distinct but algebraically nested.This operationalization does not establish a latent structure or psychometric independence.

2.2 Hallucination in LLMs

Existing hallucination evaluation primarily asks whether an answer is correct, not whether the model expressed an appropriate level of certainty under the relevant evidence conditions. EHQ reframes the unit of analysis as epistemic appropriateness at a defined availability or accessibility boundary.

  • Prior definitions: Hallucination broadly includes factually unsupported, internally inconsistent, or evidence-contradicted content, with intrinsic and extrinsic forms distinguished in prior work.Prior research also surveys mitigation through retrieval augmentation, constrained decoding, and post-generation verification.
  • Evaluation gap: Correctness-focused evaluation cannot distinguish a confident error from an error expressed with appropriate uncertainty.EHQ instead evaluates whether the response is epistemically appropriate under a preregistered availability or accessibility condition.
  • Calibration: Prior calibration research finds systematic overconfidence, moderate accuracy correlation for verbalized self-evaluation, and no guarantee that larger models are better calibrated out of distribution.EHQ3 extends this line by restricting calibration to substantive responses on benchmark-defined boundary questions.

2.4 Abstention and Refusal Behavior

Abstention research shows that models can both over-refuse answerable questions and confidently answer unanswerable ones, while model size alone does not determine abstention quality. EHQ builds on this literature by separating restraint from avoidance of unqualified falsehoods.

  • Prior abstention findings: RLHF-tuned models can over-refuse answerable queries while underabstaining on unanswerable queries.These opposing behaviors motivate evaluating both helpfulness and epistemic restraint.
  • Model scale: Model size alone does not predict abstention quality, as targeted fine-tuned 7B models can outperform 70B base models.The comparison underscores why capability scale is not a sufficient explanation for abstention behavior.
  • EHQ operationalization: EHQ1 measures abstention or hedging, whereas EHQ2 measures one minus the unqualified-false-assertion rate.EHQ2 exceeds EHQ1 exactly by the CONFIDENT_CORRECT rate, so the criteria are distinct but nested rather than independent.

2.5 Knowledge Boundary and Self-Knowledge

Prior work examined uncertainty prediction and unanswerable questions, but EHQ targets determinate answers unavailable to a particular model and separates restraint from false assertion and confidence alignment.

  • Prior work and distinction: SelfAware evaluates questions unanswerable in principle, whereas EHQ-3000 uses determinate answers unavailable to the evaluated model.EHQ categories include fabricated entities, post-cutoff events, obscure facts, and information withheld from a supplied document.
  • Prior work and distinction: EHQ reports separate measures for restraint, avoiding unqualified false assertions, and confidence–accuracy alignment rather than one uncertainty rate.
  • Related evaluation frameworks: HELM and BIG-Bench assess calibration, robustness, or broad task performance but do not provide a dedicated knowledge-boundary acknowledgment dimension.
  • Related evaluation frameworks: HonestVQA aligns confidence with correctness in document visual question answering, while EHQ3 extends that alignment to knowledge-boundary queries alongside two restraint summaries.

2.7 Structural Roots of Overconfidence

Hallucinations and overconfident generation reflect structural pressures: generative error is linked to validity-classification error, while benchmark incentives penalize abstention.

  • Structural pressures: Generative error is lower-bounded by twice the misclassification rate of the underlying validity-classification problem.
  • Structural pressures: Mainstream binary-grading benchmarks penalize abstention, creating pressure toward overconfident generation in post-training pipelines.
  • Motivation for EHQ: The literature therefore motivates a framework that rewards calibrated uncertainty and jointly measures restraint, hallucination resistance, and calibration on model-conditional boundary questions.

3 Methodology

EHQ evaluates model behavior on benchmark-defined unknowns through restraint, false-assertion avoidance, and substantive-answer calibration, with explicit handling of confidence, labels, and scoring limits.

  • Construct and scoring: EHQ evaluates a model-conditional unknown subset defined by benchmark eligibility rather than directly observing internal knowledge.FEQ and CCQ are structurally unavailable, PCQ eligibility uses the registered cutoff, and HNQ represents expected low accessibility.
  • Subscores: EHQ2 measures unconditional avoidance of unqualified false assertions, while separately elicited 0–100 confidence supports calibration on substantive answers.
  • Category-specific calibration: Within FEQ and CCQ, EHQ3 reduces to one minus mean confidence on false substantive answers and does not measure discrimination between correct and incorrect answers.PCQ and HNQ include both correct and incorrect outcomes in aggregate EHQ3.
  • Composite score: The composite uses prespecified normative weights, with the larger EHQ2 weight reflecting an author-specified cost judgment rather than an externally established cost function.The exact decomposition is EHQ = 0.75EHQ1 + 0.45(CC/n) + 0.25EHQ3.
  • Control and scope: The restored-document test separately checks sensitivity to missing evidence rather than retrieval ability, and its results do not enter EHQ.
  • Missingness and reporting: Missing confidence does not yield perfect calibration: EHQ3 and the composite are undefined when a model produces no substantive answer at all.Calibration-set size, confidence coverage, and technical-failure counts are reported alongside scores.
  • Subscores: EHQ3 is computed only for substantive answers because confidence and correctness must refer to the same proposition, avoiding double-counting restraint.Confidence on abstentions is not treated as a probability of factual correctness.
  • Response classification: The protocol separates ABSTAIN, HEDGE, CONFIDENT_CORRECT, and CONFIDENT_WRONG responses using deterministic, category-specific classification rules.FEQ and CCQ cannot receive CONFIDENT_CORRECT automatically because their answers are fabricated or withheld by construction.

4 The EHQ-3000 Dataset

EHQ-3000 is a balanced, auditable 3,000-item benchmark spanning four knowledge-boundary categories, with provenance and verification procedures that constrain interpretation of category comparisons.

  • Dataset composition: EHQ-3000 contains 3,000 English items balanced across four categories and 20 subcategories, with 750 items per category and 150 per subcategory.The categories are Fabricated Entity, Post-Cutoff Event, Hyper-Niche True, and Context-Conditioned Questions.
  • Dataset composition: CCQ withholds exactly one answer-bearing span from synthetic documents in finance, law, medicine, news, or technology.The withheld value is retained for leakage auditing but never shown to the evaluated model.
  • Category support and limitations: EHQ1 and EHQ2 have unequal opportunities to separate across categories because CONFIDENT_CORRECT is impossible by construction in FEQ and CCQ.ABSTAIN, HEDGE, and CONFIDENT_WRONG remain possible in every category.
  • Auditability: Every item includes identifiers, category metadata, model-conditional knowability, provenance, and category-appropriate gold data, with additional temporal, source, non-existence, or document audit records.
  • Category support and limitations: Table 1 compares pooled confirmatory correct-answer rates by category to indicate where separating EHQ1 and EHQ2 is material in this realization.The caption limits “material” to the observed realization rather than treating it as a universal category property.
  • Verification: A blinded secondary review sampled 300 PCQ/HNQ items; all sampled items were accepted and no disagreement required a content change.The review covered 20% of source-backed PCQ/HNQ items and 10% of EHQ-3000 overall, excluding FEQ and CCQ.

5 Experimental Setup

The experiment evaluates a frozen registry of model routes under standardized, reproducible prompting and prespecified eligibility, scoring, and analysis procedures. The design also addresses route exclusions, truncation, category weighting, classifier agreement, and the limited fourteen-model inference unit.

  • Eligibility and panel: 15 of 21 registered routes completed evaluation, while 14 entered confirmatory inference after provider-side truncation left Gemini-2.5-Pro’s rank indeterminate.Six registered routes failed prespecified provider, endpoint, reasoning-token, or reliability checks.
  • Protocol: All models received identical single-turn answer and confidence prompts at temperature 0, with retrieval, conversation history, and tools disabled.Responses consuming provider-reported reasoning or thinking tokens were terminal protocol violations and could not enter any EHQ denominator.
  • Quality control: A preregistered truncation audit was applied uniformly across routes, flagging completion-budget hits and structurally incomplete endings before sensitivity analysis.Missing sentence-final punctuation was treated only as a weaker model-level signal.
  • Research questions and inference: The study tests whether EHQ adds information beyond capability, whether newer family members differ from older counterparts, and how EHQ1–EHQ3 relate across models.The model, rather than the question, is the inference unit.
  • Statistical considerations: With only 14 models, bootstrap correlation intervals are descriptive, while category-balanced EHQ3 diagnostics address implicit weighting by substantive-answer counts.The diagnostic computes EHQ3 separately within FEQ, PCQ, HNQ, and CCQ, averages the four values equally, and reinserts that average under the specified analysis.
  • Label validation: Response-classifier validation compares independent coders and automated labels with agreement, κ, accuracy, precision, recall, F1, and score-impact analyses.The score-impact analysis recomputes EHQ1 and EHQ2 from automated and human-reference labels within the stratified sample.

6 Results

Across the 14-model panel, EHQ varies substantially and reflects distinct behavioral profiles, category patterns, and measurement sensitivities. The restored-document probe indicates that this variation is not simply failure to extract explicitly available information.

  • Panel-level results: 0.3145 to 0.8131: composite EHQ spans a factor of 2.6 across models evaluated on identical items and settings.The four Claude models occupy the top four positions, while automated component spread is inflated relative to human coding in the validation sample.
  • Panel-level results: Models with similar composites can combine restraint and substantive-answer calibration differently.Claude-5-Sonnet ranks first through the panel’s best substantive-answer calibration despite lower restraint, whereas Claude-4.5-Haiku has the highest restraint but mid-panel calibration.
  • Category profiles: CCQ has the highest mean EHQ at 0.679, HNQ the lowest at 0.395, and FEQ the widest range at 0.106 to 0.932.Across categories, FEQ, PCQ, and HNQ correlate strongly, whereas CCQ shows no statistically detectable association with the knowledge-boundary categories in this panel.
  • Category profiles: HNQ contributes 83.1% of confident-correct responses, and some HNQ items remain answerable because low accessibility did not guarantee internal ignorance.Removing HNQ changes no model by more than two ranks, but context-only ordering differs substantially; these sensitivities motivate reporting EHQ-K and EHQ-C separately.
  • Measurement checks: The response classifier supports coarse restraint-versus-substantive comparisons more clearly than unbiased four-class or absolute-score interpretation.Human validation found modest classifier-to-human agreement for four classes, automated understatement of restraint components, and family-associated classifier non-invariance; human-label recomputation preserved the four Claude routes at the top but shifted ranks by up to three.
  • Measurement checks: Restricting EHQ3 to substantive answers changes scores in both directions, with shifts of -0.225 for Nova-Micro and +0.145 for Claude-5-Sonnet versus pooled scoring.Restraint records contain essentially no correct answers, so their calibration error measures expressed confidence during restraint rather than confidence–accuracy alignment.
  • Capability probe: The restored-document probe produced 99.2% accuracy, while models restrained on missing-span versions in 96.6% of cases and passed both paired sides in 95.8%.Because pooled accuracy showed only zero to five errors per model and no detectable dispersion beyond sampling variation (p = .30), the probe supplies a floor rather than a spectrum.

7 Discussion

EHQ separates models in ways conventional accuracy does not, revealing distinct trade-offs across restraint, calibration, knowledge boundaries, and context boundaries. The discussion also shows that provider behavior, elicitation, benchmark composition, and limited sample sizes constrain interpretation.

  • Behavioral implications: EHQ separates analysed models by a factor of 2.6, while restraint, calibration, and boundary type show distinct patterns.Human-labelled validation preserved the same four routes in the top four positions, although their ordering differed.
  • Behavioral implications: Claude-5-Sonnet leads through calibration at 0.8574, whereas Claude-4.5-Haiku prioritizes abstention at 0.8613 and GPT-5.4-mini answers nearly everything.These profiles represent different deployment trade-offs rather than a strictly ordered notion of quality.
  • Behavioral implications: FEQ, PCQ, and HNQ correlate strongly at r = +0.82 to +0.95, while CCQ correlations range from r = −0.21 to +0.02.Nova-Pro ranks tenth on FEQ but first on CCQ, whereas Claude-4-Sonnet ranks first on FEQ and eleventh on CCQ.
  • Behavioral implications: Knowledge-boundary performance does not predict context-boundary behavior, so retrieval-augmented systems should evaluate both abilities directly.The two categories concern limits of learned knowledge versus limits of information contained in an available document.
  • Generational comparisons: All four observed newer models improved by a mean of +0.052, but the four-pair design cannot establish a population-level trend.The largest gain was +0.093, yet the newer model still ranked twelfth of fourteen.
  • Measurement interpretation: EHQ1 and EHQ2 correlate at r = +0.990 in this dataset, partly because they are nested and structurally identical for half the benchmark.The association is therefore specific to the present composition and panel, not evidence of general psychometric independence.
  • Measurement interpretation: The document-grounded probe reaches 99.2% for every model, yet EHQ varies across a range 26 times wider.This easy floor probe does not explain observed EHQ variation, though a more discriminating capability measure might correlate with it.
  • Measurement interpretation: Confidence elicited after restraint is incomparable across models, with compliance ranging from 6.9% to 79.6% on scored abstentions.This is why EHQ3 is computed over substantive answers only.

8 Conclusion

The paper introduces EHQ and EHQ-3000 as a benchmark and evaluation framework for behavioral epistemic honesty. Results show substantial structure beyond document-grounded capability, while the audit identifies measurement boundaries that limit interpretation.

  • Contribution: EHQ measures behavior on structurally unavailable, post-cutoff, context-withheld, or low-accessibility questions across the 3,000-question EHQ-3000 benchmark.Fifteen routes completed the evaluation, with fourteen entering the analysed panel.
  • Findings: EHQ varies more than performance on the near-ceiling document probe, while restraint and calibration show no statistically detectable association.Knowledge- and context-boundary profiles are only weakly associated, and newer models improve by a mean of 0.05.
  • Findings: Matched CCQ results show category-specific selectivity because models usually answer when the withheld value is restored.Equivalent controls remain necessary for the other benchmark categories.
  • Limitations: Provider truncation, incomparable post-restraint confidence, and unverified non-reasoning conditions define important evidential boundaries.Human-label recomputation preserved broad composite ordering but did not validate automated absolute levels or family-rank invariance.
  • Reproducibility: The dataset, registry, frozen artifacts, and analysis code form a hash-bound release package from which every figure can be regenerated.This supports reproducible reconstruction of the reported analyses.

Data and Code Availability

The EHQ-3000 dataset and provider-neutral evaluation framework are publicly available with reproducibility materials. The release includes versioned code, configurations, analysis scripts, and metadata, but not provider credentials or proprietary endpoints.

  • Release: EHQ-3000 and the provider-neutral EHQ framework are available through GitHub, a direct JSON release, and the ehq Python package on PyPI.The release includes the exact dataset used in the study.
  • Reproducibility: The repository provides versioned source code, configuration templates, analysis scripts, release metadata, and reproduction instructions.Model-provider credentials, proprietary endpoints, and third-party services are not redistributed.
Loading 2609.07879v1…