Source-linked AI summary
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation
Xue Liu, Xin Ma, Yuxin Ma, Yongchang Peng, Duo Wang, Zhoufutu Wen, Ge Zhang, Kaiyuan Zhang, Xinyu Chen, Yida Ding, Tianci He, Jiani Hou, Liang Hu, Ziyun Huang, Yongzhe Hui, Jianpeng Jiao, Chennan Ju, Yingru Kong, Yiran Li, Jiashuo Liu, Mengyun Liu, Luyao Ma, Fei Ni, Yiqing Ni, Pengbo Niu, Yueyan Qiu, Yanle Ren, Xinyu Shen, Zilin Shi, Zaiyuan Wang, Wenjie Yue, Chun Zhang, Shiyu Zhang, Xinyi Zhang, Kaiwen Zhao, Zhenwei Zhu, Shanshan Wu, Qi Zhao, Wenhao Huang
TL;DR
XpertBench addresses the limited ability of conventional benchmarks to evaluate open-ended, expert-level cognition across authentic professional domains. It combines practitioner-curated tasks and granular expert rubrics with ShotJudge, an expert-calibrated LLM judging paradigm. Evaluations reveal a performance ceiling, substantial domain specialization, and no single universally dominant expert model.
Problem
Existing benchmarks provide limited evidence about open-ended expert cognition because they often rely on narrow domains, generalist tasks, or potentially self-evaluating judges.
Method
XpertBench combines practitioner-sourced tasks across seven domains with expert-anchored rubrics and ShotJudge few-shot calibration for scalable evaluation.
Results
Models show domain-specific capability divergence, with GPT-5.4-high scoring 84.65% in Finance but 42.84% in STEM, while evaluation identifies retrieval interference and principle hallucinations.
Takeaways & Limitations
The findings indicate that a singular omni-capable expert model does not yet exist and that current systems retain an expert gap across professional workflows.
Abstract
from arXiv · showhide
As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks characterizing genuine expert-level cognition. Existing frameworks suffer from narrow domain coverage, reliance on generalist tasks, or self-evaluation biases. To bridge this gap, we present XpertBench, a high-fidelity benchmark engineered to assess LLMs across authentic professional domains. XpertBench consists of 1,346 meticulously curated tasks across 80 categories, spanning finance, healthcare, legal services, education, and dual-track research (STEM and Humanities). These tasks are derived from over 1,000 submissions by domain experts--including researchers from elite institutions and practitioners with extensive clinical or industrial experience--ensuring superior ecological validity. Each task uses detailed rubrics with mostly 15-40 weighted checkpoints to assess professional rigor. To facilitate scalable yet human-aligned assessment, we introduce ShotJudge, a novel evaluation paradigm that employs LLM judges calibrated with expert few-shot exemplars to mitigate self-rewarding biases. Our empirical evaluation of state-of-the-art LLMs reveals a pronounced performance ceiling: even leading models achieve a peak success rate of only ~66%, with a mean score around 55%. Models also exhibit domain-specific divergence, showing non-overlapping strengths in quantitative reasoning versus linguistic synthesis.. These findings underscore a significant "expert-gap" in current AI systems and establish XpertBench as a critical instrument for navigating the transition from general-purpose assistants to specialized professional collaborators.
1 Introduction
XpertBench targets the gap between conventional benchmark scores and practical expert workflows through open-ended, multi-domain tasks curated by experts and evaluated with granular rubrics. Its evaluation of frontier models reveals substantial domain-specific weaknesses and no universally capable expert model.
- Benchmark motivation: XpertBench evaluates LLMs on end-to-end, real-world expert workflows rather than isolated benchmark questions.The benchmark is designed to improve ecological validity and measure practical AI utility.
- Benchmark motivation: Open-ended tasks require navigating ambiguity, synthesizing domain literature, and resolving conflicting constraints beyond point-estimate metrics.These tasks are modeled on deep research and ill-structured expert problem-solving.
- Benchmark scope: XpertBench covers seven high-stakes professional domains, including underrepresented areas such as Education and Humanities & Social Sciences.Education accounts for 24.4% and Humanities & Social Sciences for 8.6% in the cited domain distribution.
- Benchmark construction: Over 1,000 experts contributed 1,346 testable scenarios, each evaluated with a rubric containing 15–40 granular checkpoints.The curation process included two-stage qualification and multi-stage peer review.
- Empirical findings: Evaluation of 12 state-of-the-art models found weaker performance in STEM and Education, alongside retrieval interference, principle hallucinations, and domain specialization.These failure modes affect end-to-end usability and reasoning coherence.
- Empirical findings: The evidence indicates that a singular omni-capable expert model does not yet exist.The conclusion follows from the observed capability gaps and domain-specific divergence.
2 Related Work
Prior benchmarks improve difficulty, retrieval, or domain specialization but often miss open-ended professional reasoning and cross-domain synthesis. XpertBench addresses these gaps with practitioner-sourced tasks and expert-anchored, granular rubric evaluation through ShotJudge.
- Existing benchmarks: Specialized benchmarks assess fields such as medicine, STEM, law, and finance but remain isolated from cross-domain synthesis.Their disciplinary focus limits measurement of adaptable reasoning for versatile AI assistants.
- Existing benchmarks: Difficulty-focused and agentic benchmarks still commonly reduce evaluation to knowledge recall, retrieval, or short reference answers.They therefore miss ill-structured problem-solving involving ambiguity and conflicting constraints.
- XpertBench: XpertBench sources 1,346 tasks from active practitioners across seven high-demand domains to prioritize ecological validity.The benchmark anchors evaluation in authentic professional workflows rather than academic proxies.
- Rubric evaluation: Traditional Exact Match and ROUGE metrics are increasingly unsuitable for complex generative artifacts, motivating granular rubric-based evaluation.Rubrics decompose holistic quality into interpretable dimensions.
- Evaluation gaps: LLM-as-a-judge methods offer scalability but face circularity, self-enhancement bias, and weak reliability on challenging comparisons.Existing human-centric approaches provide a contrasting evaluative extreme.
- ShotJudge: ShotJudge uses expert-anchored rubrics with 15–40 weighted checkpoints and few-shot human assessments to ground automated judgments in professional standards.This hybrid design aims to reconcile evaluative rigor with scalability.
3 Overview of XpertBench
XpertBench is a multi-domain benchmark for high-value, open-ended, long-horizon professional tasks, constructed from expert-authored scenarios and evaluated with granular rubrics. Its design emphasizes broad domain coverage, expert qualification, and reproducible checkpoint-based scoring.
- XpertBench comprises 1,346 complex tasks spanning seven professional domains and capabilities from strategic planning to cultural interpretation.
- The benchmark targets knowledge-intensive sectors selected for economic contribution, cognitive complexity, and societal impact.
- Experts contribute authentic professional scenarios after proficiency testing, trial annotation, standardized training, and senior review.
- A multi-stage selection process filters for challenging, typical, high-frequency tasks rather than edge cases or overly specialized scenarios.
- Checkpoint weights combine Essential, Important, and Optional categories with numerical values from 1 to 10 for score calibration.
4 Evaluation
ShotJudge makes automated scoring more expert-aligned by calibrating an LLM judge with expert-validated exemplars and weighted rubric criteria. Its reliability is assessed against human experts on a curated 245-task subset.
- ShotJudge uses expert-annotated exemplars to calibrate an automated judge against professional evaluative standards.
- The expert anchoring and meta-evaluation process produces the XpertBench-Gold subset of 245 stratified tasks for empirical evaluation.
- The judge receives the task, expert-designed rubric, and baseline response paired with expert-validated scores and rationales as one-shot context.
- Each rubric criterion receives a binary score, and the final metric aggregates these scores using expert-assigned criterion weights.
- ShotJudge achieves a CDR of 52.0%, outperforming standard zero-shot LLM-as-a-judge baselines.
5 Experiments and Results
Evaluation on XpertBench-Gold shows that leading models reach only about 65–66% overall while displaying pronounced domain-specific specialization. Model strengths vary across law, humanities, finance, education, and STEM.
- Claude-Opus-4.6-thinking achieves the highest overall score at 66.20% on the 245-task XpertBench-Gold subset.
- Claude-Opus-4.6-thinking leads Law at 65.54% and Humanities at 83.02%, while GPT-5.4-high leads Finance at 84.65%.
- GPT-5.4-high leads Education at 59.29%, whereas GPT-5-high slightly exceeds GPT-5.2-high in STEM at 48.20% versus 46.13%.
- Leading models achieve only a ∼65–66% success rate even with retrieval/search capabilities.
- GPT-5.4-high scores 84.65% in Finance but 42.84% in STEM, illustrating specialization rather than universal dominance.
6 Contributions
The supplied contribution passages identify the paper’s core and contributing authors and note corresponding-author and internship affiliations.
- Xue Liu, Xin Ma, Yuxin Ma, Yongchang Peng, Duo Wang, Zhoufutu Wen, Ge Zhang, and Kaiyuan Zhang are listed as core contributors.
- Xinyu Chen and the listed additional contributors are identified in the contributor roster.
- A dagger marks corresponding authors, while several contributors are identified as ByteDance Seed interns during the work.
7 Xpert Platform
Xpert is presented as ByteDance’s expert-level data service platform for specialized training data and evaluation, focused on converting expert knowledge into high-quality data and AI utility.
- Xpert is an expert-level data service platform under ByteDance that provides specialized training data and evaluation solutions.
- The platform aims to transform experts’ industry knowledge and experience into high-quality data for AGI and broader commercial and social value.
- Xpert’s approximately 3,000 selected experts include university scholars and professionals with 2–10 years of practical experience across multiple industries.
- Its leaderboard evaluates AI on complex expert-level real-world tasks rather than mainstream exam-oriented questions.
A Example Tasks and Scoring Rubrics
The appendix presents representative Finance tasks and rubrics that require rigorous, quantitative comparison of Lockheed Martin and Northrop Grumman during 2022–2023. The example evaluates revenue visibility, profitability, cash conversion, and evidence-based synthesis without predictions or investment advice.
- A Example Tasks and Scoring Rubrics: The appendix provides representative example tasks and scoring rubrics across Finance, Law, Education, STEM, and Humanities & Social Sciences.
- A.1 Finance: The Finance category covers corporate strategy analysis, financial reporting, market research, and business case studies.
- A.1.1 Example Task: The example task asks a senior credit-rating analyst to compare Lockheed Martin and Northrop Grumman against 2022–2023 geopolitical and defense-budget conditions.
- A.1.1 Example Task: The analysis requires locating net sales and net orders or backlog changes and calculating each company’s FY2023 Order-to-Sales Ratio.
- A.1.1 Example Task: It compares operating margins across business segments and identifies each company’s highest-margin core segment for FY2023.
- A.1.1 Example Task: It assesses cash-flow generation by comparing operating cash flow, capital expenditures, and net earnings to calculate FY2023 Free Cash Flow Conversion Rate.
- A.1.1 Example Task: The final synthesis connects revenue visibility, profit engines, and cash-flow efficiency to divergent operational and financial performance.
- A.1.2 Scoring Rubric: The Finance scoring rubric is presented in Table A.1 and its continuation.
A.2 Law
The appendix illustrates expert-task formats across Law, Education, STEM, and associated scoring rubrics. Tasks range from legal classification and liability allocation to philosophical scriptwriting and experimental plasmid analysis.
- A.2 Law: The Law category covers legal drafting, complex reasoning, evidence analysis, and legal argumentation.
- A.2 Law: The Law example asks whether agreements constitute loan or factoring relationships, assesses agreement validity, and allocates potential liability.
- A Scoring Rubrics: Scoring rubrics for the Law, Education, and STEM examples are provided in Tables A.2, A.3, and A.4.
- A.3 Education: The Education category addresses instructional design, curriculum development, pedagogical assessment, student learning analysis, and educational research synthesis.
- A.3 Education: The Humanities example requires a Modern Chinese philosophical dialogue between Confucius and Socrates that preserves their styles, beliefs, biographies, and script-format conventions.
- A.4 STEM: The STEM example asks researchers to infer empty-vector rates and foreign-fragment sizes from restriction-digestion electrophoresis patterns.
- A.4 STEM: It also asks why MboI and Sau3AI digestion results differ for the same plasmid, linking interpretation to the displayed experimental figures.
A.5 Humanities & Social Sciences (HSS)
The HSS materials emphasize critical analysis, theory application, academic writing, and research methods through a historically grounded philosophical dialogue task. The task requires stylistic, philosophical, factual, linguistic, and formatting fidelity.
- A.5 Humanities & Social Sciences (HSS): The HSS category covers critical analysis, theory application, academic writing, and research methods.
- A.5 Humanities & Social Sciences (HSS): The example asks for a philosophical dialogue script featuring Confucius and Socrates in their final years, centered on life and death.
- A.5 Humanities & Social Sciences (HSS): The script must reflect the sages’ textual styles, philosophical views, and relevant biographical facts while allowing moderate artistic license.
- A.5 Humanities & Social Sciences (HSS): The final output must be in Modern Chinese, preserve historical and philosophical accuracy, and follow standard script formatting.
- A.5 Humanities & Social Sciences (HSS): The HSS scoring rubric is presented in Table A.5 and its continuation.