Source-linked AI summary
HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
Weiqi Zhai, Zhihai Wang, Jinghang Wang, Boyu Yang, Xiaogang Li, Xander Xu, Bohan Wang, Peng Wang, Xingzhe Wu, Anfeng Li, Qiyuan Feng, Yuhao Zhou, Taolin Han, Wenjie Luo, Yiyuan Li, Xiang Zheng, Yaxuan Wang, Ruixiang Luo, Guojie Lin, Peiyao Xiao, Chengliang Xu, Ben Wang, Zeyu Wang, Zichao Chen, Jianan Ye, Yijie Hu, Jialong Chen, Zongwen Shen, Yuliang Xu, An Yang, Bowen Yu, Dayiheng Liu, Junyang Lin, Hu Wei, Que Shen, Bing Zhao
TL;DR
HLE evaluation can be affected by noisy or defective items, raising questions about the reliability of model comparisons. HLE-Verified addresses this gap with a structured verification-and-revision framework and explicit uncertainty categories. The resulting benchmark contains verified, revised, and uncertain subsets for more transparent evaluation.
Problem
HLE contains potential annotation defects whose prevalence and impact on evaluation outcomes had not been systematically characterized.
Method
HLE-Verified uses a transparent, component-wise verification and revision protocol that preserves evaluation objectives and records item validity and uncertainty.
Results
1,143 items were corrected and re-verified, alongside 668 validated items and 689 items retained as uncertain.
Takeaways & Limitations
HLE-Verified provides benchmark infrastructure for more rigorous, interpretable, and reproducible model comparisons.
Takeaways & Limitations
689 items remain uncertain when validity requires non-standard assumptions, external references, unresolved expert disagreement, or expertise beyond the verification scope.
Abstract
from arXiv · showhide
Humanity's Last Exam (HLE) has become a widely used benchmark for evaluating frontier large language models on challenging, multi-domain questions. However, community-led analyses have raised concerns that HLE contains a non-trivial number of noisy items, which can bias evaluation results and distort cross-model comparisons. To address this challenge, we introduce HLE-Verified, a verified and revised version of HLE with a transparent verification protocol and fine-grained error taxonomy. Our construction follows a two-stage validation-and-repair workflow resulting in a certified benchmark. In Stage I, each item undergoes binary validation of the problem and final answer through domain-expert review and model-based cross-checks, yielding 668 verified items. In Stage II, flawed but fixable items are revised under strict constraints preserving the original evaluation intent, through dual independent expert repairs, model-assisted auditing, and final adjudication, resulting in 1,143 revised-and-certified items. The remaining 689 items are released as a documented uncertain set with explicit uncertainty sources and expertise tags for future refinement. We evaluate eight state-of-the-art language models on HLE and HLE-Verified, observing an average absolute accuracy gain of 7--10 percentage points on HLE-Verified. The improvement is particularly pronounced on items where the original problem statement and/or reference answer is erroneous, with gains of 30--40 percentage points. Our analyses further reveal a strong association between model confidence and the presence of errors in the problem statement or reference answer, supporting the effectiveness of our revisions. Overall, HLE-Verified improves HLE-style evaluations by reducing annotation noise and enabling more faithful measurement of model capabilities. Data is available at: https://huggingface.co/datasets/skylenage/HLE-Verified
1 Introduction
HLE is a difficult, broad benchmark whose annotation noise can distort model comparisons and capability measurements. HLE-Verified addresses this reliability problem through transparent verification, structured revision, and documented uncertainty.
- Annotation errors in HLE can make measured performance reflect dataset artifacts rather than genuine capability differences.
- Systematic post-release validation is needed because unverified benchmark defects can undermine interpretability, reproducibility, and measurement reliability.
- HLE-Verified applies a component-level, two-stage validation-and-revision protocol focused on problem statements and final answers, with rationales as secondary contradiction signals.
- HLE-Verified is positioned as dataset infrastructure that supports more rigorous, interpretable, and reproducible model comparisons.
- 668 items were verified without modification, 1,143 were revised and re-verified, and 689 remained in a documented uncertain set.
- +7–10 accuracy points were observed overall, with +30–40 points on items containing erroneous problems or answers.
2 Background
HLE is widely used for challenging multi-domain evaluation, but annotation defects and unstable evaluation artifacts may compromise its interpretability. HLE-Verified reframes these defects as explicitly characterized sources of validity, error, and uncertainty.
- Flawed benchmark items can distort aggregate metrics, especially when errors cluster by domain, reasoning pattern, or difficulty.
- Ill-posed questions and incorrect answer keys weaken the correspondence between model correctness and confidence.
- Recurring defects include ambiguous statements, mismatched answer keys, rationale-answer inconsistencies, and unstated assumptions.
- Because defects can occur in statements, final answers, or rationales, item validity is multi-dimensional rather than binary.
- Prior work raised concerns about systematic benchmark errors, but their prevalence and evaluation impact had not been rigorously characterized for HLE.
- HLE-Verified makes correctness, revision status, and epistemic uncertainty explicit at the dataset level.
3 Dataset Verification Process and Methods
HLE-Verified constructs a consolidated benchmark through component-wise verification, constrained repair, re-verification, and explicit retention of uncertain items. The resulting release separates validated, revised, and indeterminate cases while recording structured defect and revision metadata for auditing.
- Pipeline overview: The 2,500-item HLE collection is processed through Stage I verification and Stage II repair, with indeterminate items retained rather than discarded.Each item is decomposed into problem, final answer, and rationale components; problem and answer are primary correctness targets, while rationale provides diagnostic support.
- Stage I: component-wise verification: 668 items form the gold subset after validation without modification.Stage I produces a high-confidence subset, while other items proceed to repair or the uncertain pool.
- Stage II: systematic revision: 1,143 items form the revision subset after correction and re-verification under preserved evaluation objectives.Repairs use independent expert proposals, auxiliary model-assisted checks, and final expert adjudication; revised items include structured audit metadata.
- Uncertain subset: 689 items form the uncertain subset because their validity remains indeterminate under available expertise and evidence.Uncertainty may reflect non-standard assumptions, unresolved expert disagreement, authoritative references, or knowledge beyond the verification scope.
- Annotation framework: The release records component-wise labels, error types, revision traces, adjudication notes, and uncertainty descriptors, including a 19-category defect taxonomy.The taxonomy contains 5 problem-level, 10 rationale-level, and 4 answer-level categories, supporting prevalence analysis, repair analysis, and public metadata.
- Verification protocol: Validity is assessed separately for problem statements, final answers, and rationales using expert review, model-assisted replication, and internal adjudication.Gold inclusion requires both problem and answer to be judged unproblematic without high-risk ambiguity.
4 Dataset Statistical Analysis
HLE’s reliability problems are concentrated in answer and rationale components, vary substantially by discipline, and are driven mainly by correctness, specification, and representation defects rather than conceptual invalidity.
- Cross-domain variation: Answer- and rationale-level defects account for most reliability degradation, while domains differ between explicit incorrectness and epistemic indeterminacy.These differences support component-wise, epistemically explicit verification rather than aggregate item-level labeling.
- Component-wise defect distribution: Answer-level defects are dominated by Incorrect Answer across all subjects, with within-component proportions ranging from 69.4% in Chemistry to 97.2% in Biology/Medicine.All subjects exceed 70%, indicating primarily deterministic answer-key failures.
- Component-wise defect distribution: Rationale instability is heterogeneous: structural incompleteness dominates Mathematics and Biology/Medicine, while format-induced semantic errors lead Chemistry and Computer Science.Type 3 accounts for 40.1% and 45.0%, whereas type 10 accounts for 40.0% and 63.3%, respectively.
- Cross-domain variation: Problem-level defects are domain-sensitive, with format semantic errors dominating Mathematics, Chemistry, and Computer Science, while Physics and Biology/Medicine are led by other defect types.The dominant type-5 shares are 40.0%, 56.5%, and 94.6%; Physics is led by type 1 at 57.1%, and Biology/Medicine by type 2 at 33.3%.
- Defect taxonomy: Most HLE defects arise from specification or answer-key instability rather than fundamentally incorrect task concepts.Recurring patterns include structural validity failures, representation errors, constraint omissions, and numerical or sign inconsistencies.
- Case studies: Component-wise auditing is necessary because a well-formed problem can still be unsuitable when its answer key is incorrect or its rationale is structurally unverifiable.Canonical revisions address both theoretical–implementation confusion and domain-specific incoherence.
5 Experimental Results
Experiments compare eight frontier models on full and revised subsets using accuracy and calibration error, showing larger gains on corrected items and consistent improvements across categories. Confidence also rises after repairing statement-level errors, while unchanged full-set items dilute the shift.
- Experimental setup: The evaluation reports Accuracy and Calibration Error for eight frontier models on full-set and revised-subset comparisons.The revised subset isolates items edited or flagged during verification, while the full set measures end-to-end benchmark effects.
- Main evaluation: Revised-subset accuracy increases by 29.94 to 39.58 percentage points across the reported models.Reported shifts include GPT-5.2 (+38.04), Claude-Opus4.5 (+32.94), and DeepSeek-V3.2 (+39.58).
- Main evaluation: Calibration error decreases after revision on the revised subset, including GPT-5.2 from 63 to 28 and DeepSeek-V3.2 from 70 to 28.The results indicate that flawed items distort confidence-based evaluation as well as accuracy.
- Main evaluation: Full-set accuracy increases by 7.58 to 10.79 percentage points, while calibration error also decreases across models.Smaller full-set gains are expected because most items remain unchanged.
- Category-level results: Revised-subset accuracy increases across every subject category, with the largest jumps in Physics and Biology/Medicine and smaller gains in Chemistry and Computer Science/AI.The category pattern indicates nonuniform measurement noise in raw HLE.
- Confidence diagnostic: Confidence rises by roughly 1.83 to 11.08 absolute points on the Problem-Error Subset, whereas Full-Set shifts are near zero or slightly negative.Unchanged items dilute the full-set effect, while statement-level repair consistently increases confidence.
6 Conclusion
HLE-Verified strengthens the scientific reliability of HLE-based evaluations by making benchmark flaws measurable, providing transparent corrections, and quantifying their effects on accuracy and calibration. Its disputed set supports continued community-driven improvement.
- HLE-Verified strengthens the scientific reliability of HLE-based evaluations.
- The disputed set provides a roadmap for community-driven improvements.
A.1 LLM Judge Prompt
The LLM Judge Prompt appendix standardizes answer generation and extraction in Stage I, then supports structured repair, conservative adjudication, and post-repair validation in Stage II.
- A.1.1 LLM Judge Prompt in Stage I: Stage I prompts standardize solver outputs by requiring detailed reasoning and boxed final answers.This formatting facilitates downstream answer extraction.
- A.1.1 LLM Judge Prompt in Stage I: Answer extraction prompts isolate clean final answers from mathematical solutions, preserving symbols and separating multiple answers.The procedure reduces evaluation noise by separating reasoning from the final answer.
- A.1.1 LLM Judge Prompt in Stage I: Answer-equivalence prompts compare reference and model answers across equivalent forms, ordering, and formatting differences.Outputs are limited to Correct, Incorrect, or Uncertain and serve as diagnostic replication signals.
- A.1.2 LLM Judge Prompt in Stage II: Stage II repair prompts review the problem, answers, rationale, and historical model responses before producing a corrected answer and reasoning.Uncertain repairs can be marked empty rather than guessed.
- A.1.2 LLM Judge Prompt in Stage II: Repair candidates are adjudicated by selecting a self-consistent, verifiable answer matching the problem, with EMPTY preferred over fabrication.The adjudication output records the choice, answer, rationale, and justification.
- A.1.2 LLM Judge Prompt in Stage II: Final adjudication implements multi-model consensus and conservative arbitration to minimize false corrections.
- A.1.2 LLM Judge Prompt in Stage II: Post-repair auditing checks question, answer, and rationale correctness using conservative false decisions under uncertainty.This structured audit closes the verification loop.
- A.1.2 LLM Judge Prompt in Stage II: The post-repair audit records separate Boolean judgments and reasons for the question, answer, and rationale.
A.2 Quality control and decision principles
Quality control prioritizes conservative inclusion because flawed items can bias evaluation and distort cross-model comparisons, while model checks remain auxiliary to expert judgment.
- Including a flawed item can introduce systematic evaluation bias and distort cross-model comparisons.Excluding a potentially valid item primarily reduces coverage without inducing measurement error.
- Gold or revision inclusion requires positive evidence that the problem and final answer are well-posed, correct, and stable under expert scrutiny.Rationale defects alone do not automatically disqualify an answer-based evaluation item.
- Model replication outcomes are diagnostic signals rather than adjudicative authority.Systematic solver failures may trigger expert re-audit, but high agreement is not proof of correctness.
- Verification and revision rely on domain experts recruited through two independent supplier teams and supported by internal adjudication specialists.Experts are predominantly Master’s- and Ph.D.-level researchers across relevant scientific disciplines.
A.3 Case Study
The case studies show that benchmark errors arise from omitted constraints, violated mathematical invariants, and incorrect anatomical localization, with revisions restoring domain-consistent answers.
- Database Systems: A database case omitted a no-materialization constraint, altering the computational model and invalidating the original answer.
- Database Systems: 465 I/Os is the corrected minimum obtained by applying the canonical BNLJ formula.
- Mathematics: A mathematics case found that a closed-form answer contradicted Euler-sequence invariants, with small-n enumeration exposing systematic divergence.
- Mathematics: Invariant validation is essential because symbolic plausibility does not guarantee theoretical correctness.
- Biology/Medicine: A medical case incorrectly localized complete oculomotor palsy to the reticular formation instead of the midbrain.The revision restores clinicopathological coherence and prevents domain-level misinformation in medical evaluation.
A.4 Component-wise Defect Distribution Across Subjects
Figure S1 organizes HLE error types by subject.
- Figure S1 presents the distribution of HLE error types across subjects.
- The figure uses subject as the organizing dimension for comparing error distributions.
- Its focus is cross-subject variation in HLE error types.
A.5 Cross-subject proportional differences.
Error profiles vary substantially across subjects, with different domains concentrating defects in different answer, rationale, and problem categories.
- Mathematics: 89.4% of mathematics answer-level defects are Incorrect Answer, while type 3 leads rationale defects at 40.1%.
- Physics: 80.0% of physics answer-level errors are type 1, while its rationale defects are more evenly distributed than in other domains.
- Chemistry: 71.4% of chemistry answer-level defects are incorrect answers, while format semantic errors dominate problem- and rationale-level defects at 56.5% and 40.0%.
- Biology/Medicine: 97.2% of Biology/Medicine answer-level defects are incorrect answers, with type 3 leading rationale defects at 45.0%.
- Computer Science: 94.6% of Computer Science problem-level defects are format semantic errors, while type 10 dominates rationale defects at 63.3%.