Source-linked AI summary
\$OneMillion-Bench: How Far are Language Agents from Human Experts?
Qianyu Yang, Yang Liu, Jiaqi Li, Jun Bai, Hao Chen, Kaiyuan Chen, Tiliang Duan, Jiayun Dong, Xiaobo Hu, Zixia Jia, Yang Liu, Tao Peng, Yixin Ren, Ran Tian, Zaiyuan Wang, Yanglihong Xiao, Gang Yao, Lingyue Yin, Ge Zhang, Chun Zhang, Jianpeng Jiao, Zilong Zheng, Yuan Gong
TL;DR
Existing benchmarks often underrepresent the context-heavy, constrained work required in professional settings. $OneMillion-Bench addresses this gap with 400 expert-curated tasks and rubric-based evaluation, finding that current agents frequently lack consistent evidence grounding and professional reliability. The benchmark therefore measures practical readiness through domain-intensive, economically grounded scenarios.
Problem
Existing benchmarks remain largely structured or exam-style, while professional tasks require retrieval, evidence resolution, domain rules, and constraint-aware reasoning.
Method
The paper builds a 400-task benchmark across five professional domains and evaluates outputs with expert-defined rubrics measuring factual, logical, practical, and compliance criteria.
Results
Current models often fail to maintain the consistency and evidence grounding required for autonomous professional labor.
Takeaways & Limitations
$OneMillion-Bench shifts agent evaluation from surface-level correctness toward grounded, compliant performance in economically meaningful professional work.
Takeaways & Limitations
Rubric evaluation remains less objective than checking a single expression or number and still relies on model-judge capabilities, limiting scalability of manual scoring.
Abstract
from arXiv · showhide
As language models (LMs) evolve from chat assistants to long-horizon agents capable of multi-step reasoning and tool use, existing benchmarks remain largely confined to structured or exam-style tasks that fall short of real-world professional demands. To this end, we introduce \$OneMillion-Bench \$OneMillion-Bench, a benchmark of 400 expert-curated tasks spanning Law, Finance, Industry, Healthcare, and Natural Science, built to evaluate agents across economically consequential scenarios. Unlike prior work, the benchmark requires retrieving authoritative sources, resolving conflicting evidence, applying domain-specific rules, and making constraint decisions, where correctness depends as much on the reasoning process as the final answer. We adopt a rubric-based evaluation protocol scoring factual accuracy, logical coherence, practical feasibility, and professional compliance, focused on expert-level problems to ensure meaningful differentiation across agents. Together, \$OneMillion-Bench provides a unified testbed for assessing agentic reliability, professional depth, and practical readiness in domain-intensive scenarios.
1 Introduction
$OneMillion-Bench addresses the gap between exam-style benchmarks and professional work by evaluating economically consequential, expert-level tasks. It combines broad domain coverage, fine-grained rubric analysis, and an economic interpretation of reliable agent work.
- Motivation: Professional tasks require specialized knowledge, sustained multi-step reasoning, and strict constraints beyond ordinary prompt answering.The motivation includes actuarial, legal, and investment workflows whose deliverables depend on context and professional requirements.
- Benchmark scope: $OneMillion-Bench evaluates 400 open-ended tasks across Finance, Law, Healthcare, Natural Science, and Industry, curated through more than 2,000 expert hours.Each task is assigned monetary value based on senior-professional completion time and prevailing hourly wages, with total estimated value exceeding $1 million.
- Evaluation design: The benchmark analyzes search, reasoning, verbalization, and instruction-following skills alongside aggregate performance across professional domains.This skill-oriented structure is intended to avoid overfitting evaluation to one skill cluster or domain-specific idiosyncrasy.
- Evaluation design: Rubric-based evaluation is designed to reduce reward hacking and align scoring with domain-specific policies and expert expectations.The mechanism supports multi-dimensional, fine-grained scoring of open-ended agent performance.
- Implication: The benchmark frames agent capability as the amount of reliable professional work delivered and its corresponding economic value.This shifts evaluation toward practical readiness, trustworthiness, and economically meaningful performance.
2 How does $OneMillion-Bench Measure?
$OneMillion-Bench measures economic value and professional expertise through expert-cost estimates, wage anchoring, rubric-based scores, pass rates, and balanced score aggregation. Its scoring captures both graded rubric fulfillment and whether outputs meet a minimum professional standard.
- Economic value: Benchmark economic value equals senior-expert task completion time multiplied by the relevant hourly wage.Completion-time estimates are collected from two to three senior experts to reduce subjective bias.
- Economic value: Hourly wages incorporate regional and sectoral references, including U.S. OEWS data, industry reports, and Chinese tier-1-city wage guidelines.Reported annual salaries are standardized using a 2,080-hour work year; a 1.3 overhead multiplier applies only to wage-only figures.
- Expertise measurement: Expert Score measures how strongly a generation fulfills the positive-weighted professional rubrics for a question, with the result clipped to [0, 1].Each rubric contributes according to its assigned score and predefined weight.
- Expertise measurement: Pass Rate is a binary success measure based on whether generations meet the benchmark’s minimum professional standard.The metric uses an indicator criterion applied to question-level performance.
- Score aggregation: Overall scores average domain-level scores, while rubric-type aggregation supports analysis across functional capability dimensions.This produces a balanced estimate across domains rather than relying only on a single aggregate.
3 Constructing the $OneMillion-Bench
$OneMillion-Bench is constructed through expert-designed, peer-reviewed, adversarially validated, and consensus-resolved tasks spanning professional domains and capability tags. Its workflow-realistic design uses specialized rubrics, negative penalties, and source-aware evaluation to distinguish meaningful agent performance.
- Data Curation Pipeline: The curation pipeline prioritizes economically valuable, workflow-realistic problems with objective rubrics and diverse rubric structures.A multi-expert annotation process is used to support objectivity and professional integrity.
- Data Curation Pipeline: Experts create specialized tasks, reference answers, and detailed scoring rubrics before independent peer review and consensus-based revision.A third expert audits cases involving risk or unresolved disagreement.
- Data Curation Pipeline: Frontier-agent validation retains tasks only when several agents fail a predefined expert-rubric threshold, preserving discriminative difficulty.Tasks are also filtered at both extremes to remove trivial items and review potentially mission-impossible items.
- Data Overview: Tasks require deep reasoning, real-time information integration, precise retrieval, traceable justification, and constraint satisfaction.Correct final answers are intentionally made inseparable from correct processes in many tasks.
- Data Overview: Rubrics encode criteria, weights, capability tags, and official-source citations for transparent, domain-tailored evaluation.The four capability classes include web search, while source and citation checks support factual faithfulness.
- Data Overview: Negative rubric weights range from -20 to 10 and penalize professional violations, unsafe outputs, hallucinations, and foundational competency lapses.This asymmetric design is intended to reflect domain requirements and operational robustness.
- Data Overview: The benchmark includes 200 English and 200 Chinese instances, with the Chinese collection purpose-built around local linguistic and cultural contexts.The design addresses variation in regulations and practical scenarios across environments.
4 Benchmarking Frontier Agents on $OneMillion-Bench
Evaluation across 35 models shows strong variation in performance, search effects, domain and rubric difficulty, and test-time scaling. Claude-Opus-4.6 leads overall, while search benefits depend on model robustness and task requirements.
- Evaluation Setup: 35 models are evaluated across vanilla, search-enabled, and deep-research categories on Global and CN subsets.The benchmark reports overall, domain, rubric-type, and test-time scaling analyses.
- Main Results: Claude-Opus-4.6 achieves the best overall performance among vanilla models and remains the top performer with search enabled.Search amplifies its advantage in both Expert Score and Pass Rate.
- Search Effects: Search can hurt weaker systems: HUNYUAN-2.0 falls from 34.7 to 30.2 Expert Score and from 8.5% to 3.0% Pass Rate on Global.The reported regressions are attributed to noisy or conflicting evidence and weak evidence identification.
- Agent Comparisons: Deep-research agents generally lag behind the strongest search-enabled generalists in Expert Score, Pass Rate, and Economic Value.The results suggest robust rubric coverage and compliance matter more than longer research pipelines under this evaluation.
- Evaluation Metrics: Models often achieve moderate Expert Score but much lower Pass Rate, revealing partial rubric satisfaction below the Expert Score(q) ≥0.7 threshold.Pass Rate therefore distinguishes broad but shallow gains from improvements that cross the competence boundary.
- Domain and Rubric Patterns: Finance remains consistently difficult, while Healthcare and Law are often easier for top systems; this domain profile is broadly stable across Global and CN.Factual Information and Analytical Reasoning are harder than Structure and Formatting and Instructions Following, while search most consistently helps evidence-centric rubrics.
- Search Effects: Search amplifies underlying capabilities: stronger models often improve across rubrics, whereas weaker models may degrade in reasoning, formatting, or instruction following.Using tools under constraints requires planning and compliance control, and longer citation-heavy responses can destabilize formatting.
- Test-time Scaling: As test-time sample size k increases, pass@k gains logarithmically while pass^k decays toward zero.On the Finance subset, Claude-Opus-4.6 leads pass@k and plateaus near 30%.
5 Discussions
The discussion connects benchmark performance to economic value, retrieval failure modes, and persistent weaknesses in complex professional reasoning. It also identifies coverage, rubric automation, and evolving real-world information as continuing challenges.
- Economic Value: Search agents can deliver higher economic value than the same base models, but the benchmark exposes a trade-off between outcome value and inference cost.Figure 9 presents Pareto frontiers for base models, search agents, and deep-research agents.
- Scope: The benchmark currently covers five representative domains, leaving fields such as energy, climate science, and public policy for future expansion.The authors also propose a live benchmark incorporating real-time or frequently updated information.
- Evaluation Limitations: Rubric-based evaluation is less objective than exact expression or number checking and remains dependent on model judges and difficult-to-scale manual scoring.The authors suggest automated scoring of reasoning chains, evidence citations, and compliance checks.
- Retrieval Failure Modes: Search is a conditional gain: it helps when retrieved results supply rubric-critical knowledge but can harm reasoning or structure-sensitive tasks.Failures include outdated or conflicting evidence, incompatible medical guidelines, and blurred screening hierarchies.
- Domain Bottlenecks: Finance tasks expose errors in extracting statement figures, performing arithmetic, and completing multi-step derivations required for cross-company comparisons.The discussion highlights turnover days, earnings quality, and cash conversion cycle calculations.
- Domain Bottlenecks: Legal and compliance failures commonly involve mapping fact patterns to the correct provisions, case holdings, and local standards.The bottleneck combines normative coverage, pinpoint retrieval, and faithful rule application.
- Complex Reasoning: Agents remain weak on deep understanding, multi-step deduction, long-horizon exploration, and exhaustive test-case coverage.Machine-learning tasks also show skipped contextual diagnosis, intermediate findings, contextual adaptation, and causal analysis.
- Domain Bottlenecks: Healthcare rubrics demand operational clinical details such as staging, treatment choice, follow-up, contraindications, and patient education, which models frequently omit.These omissions contribute to low scores on clinically detailed questions.
6 Related Work
Prior work spans harder static question answering, task-oriented agent environments, and evaluations grounded in external reality. OneMillion-Bench positions itself between exam-style benchmarks and unconstrained deployment for professional evaluation.
- Hard Question Answering: Hard-question benchmarks such as GPQA, LiveBench, and MMLU-Pro increase domain difficulty, freshness, reasoning demands, or resistance to shortcut guessing.These approaches primarily target static question-answering evaluation.
- Agentic Workflow Benchmarks: Agentic workflow benchmarks evaluate profession-aligned productivity, software engineering, planning under constraints, and tool-agent-user interactions.Examples include XBench, SWE-bench Verified, TravelPlanner, and τ-bench.
- Reality-Grounded Evaluation: Reality-grounded evaluations test agents where outcomes depend on external environments, including streaming financial markets and dynamic competition settings.Reported examples include LiveTradeBench and Alpha Arena.
- Positioning: OneMillion-Bench occupies a middle ground between static exam-style benchmarks and unconstrained real-world deployment.This framing aims to support more reliable and economically meaningful evaluation in professional settings.
7 Conclusion
OneMillion-Bench bridges exam-style evaluation and high-stakes professional deployment through expert-curated workflows and rubric-based assessment. Its findings indicate that current agents often lack the consistency and evidence grounding required for autonomous professional labor.
- Conclusion: OneMillion-Bench evaluates logical coherence, factual grounding, and professional compliance through expert-curated workflows and rubric-based assessment.The benchmark is designed to assess economic readiness rather than surface-level correctness alone.
- Conclusion: The findings highlight a reliability gap: current models often fail to maintain the consistency and evidence grounding required for autonomous professional labor.The paper prioritizes grounded, compliant, and economically consequential decision-making as indicators of agentic maturity.
A.1 Data Contributors
The benchmark’s contributors supported coverage across five high-level domains and a broad set of specialized subdomains. Its taxonomy includes 92 unique third-level subdomain tags.
- Domain coverage: 92 unique third-level subdomain tags span the benchmark’s five high-level domains.Non-synonymous tags across languages are retained as separate entries.
- Healthcare and Medicine: Healthcare and Medicine contains 25 tags, including emergency care, cardiovascular medicine, laboratory science, and endocrinology.
- Industry and Law: The taxonomy includes specialized industry and legal areas such as semiconductors, telecommunications, contract disputes, cybersecurity, patents, and labor law.
- Economics and Finance: Economics and Finance contains 18 tags covering equities, bonds, mergers and acquisitions, accounting, risk management, and quantitative finance.
- Natural Sciences: Natural Sciences contains 18 tags spanning physics, chemistry, mathematics, microbiology, molecular biology, genetics, and ecology.
C.3 Evaluation Cost Estimation
Evaluation cost is estimated by combining token, tool-call, cache, and infrastructure costs, then applying a buffer for operational overhead. The procedure uses α = 1.2 in practice.
- Cost equation: Ceval = α (cin · Tin + cout · Tout + ctool · Ntool + ccache · Tcache + cinfra · teval) estimates total evaluation cost.The equation aggregates API billing, cached-token usage, tool calls, and infrastructure overhead.
- Cost components: Input and output token quantities, tool calls, cached tokens, and evaluation time each contribute through their corresponding unit costs.Tool calls include web searches, while infrastructure captures compute and storage overhead beyond API billing.
- Pricing inputs: Token and tool-call pricing are queried through the OpenRouter cost and stats API.
- Operational buffer: α ≥ 1 buffers underestimation from retries, rate-limit handling, and other incidental overhead.The evaluation sets α = 1.2.
D.1 Economics and Finance
The economics and finance material analyzes 2025 yen depreciation through monetary policy, exchange-rate phases, capital flows, and structural current-account weaknesses. It emphasizes that a rate hike alone may not strengthen the yen when guidance and external flows remain unfavorable.
- Structural outflows: Structural services outflows, especially digital-services deficits, are described as persistent sources of yen-selling pressure.The benchmark rubric specifically requests analysis of digital-services outflows and their structural character.
- Analytical caveat: The analysis cautions against explaining post-hike depreciation solely through market pricing or a buy-the-rumor, sell-the-fact narrative.It calls for considering real interest rates, policy credibility, and capital flows.
- Exchange-rate phases: The 2025 yen cycle is divided into phases linking policy milestones, macroeconomic catalysts, and exchange-rate movement.The described phases include an early-year BoJ hike, a spring safe-haven yen bid, and a December hike followed by depreciation.
- December policy shock: The December 19 BoJ hike from 0.50% to 0.75% was followed by USD/JPY breaching 157 as the yen weakened.The passage characterizes the move as a hike delivered with disappointing guidance.
- External accounts: Japan recorded a JPY 31.8799 tn current-account surplus and a JPY 41.5903 tn primary-income surplus, alongside goods and services deficits.The goods deficit was JPY 0.8487 tn and the services deficit was JPY 3.3928 tn.
- Current-account composition: Primary-income surpluses may not create immediate yen demand when overseas earnings are retained abroad through reinvested earnings.
- Recovery triggers: A sustained yen recovery is associated with clearer BoJ guidance, persistent Japanese wage and inflation growth, faster Fed cuts, or global carry unwinding.These are presented as variables that could trigger trend reversion rather than as guaranteed outcomes.
What would trigger a trend reversal against JPY (USD/JPY back toward highs)
The cited scenarios identify several conditions that could push USD/JPY higher by reviving yen-funding demand or increasing risk premia. These include renewed BoJ pauses, fiscal stress, persistent risk-on carry demand, and absent intervention deterrence.
- Renewed BoJ slowing or pausing could re-anchor markets on the yen as a funding currency.
- Rising Japanese fiscal-stress concerns could lift yields while widening the country’s risk premium.
- A persistent risk-on environment could keep carry demand dominant and sustain yen-selling pressure.
- The absence of intervention deterrence during sharp moves could leave further yen depreciation less constrained.
D.4 Natural science
Inelastic neutron scattering evidence indicates that Herbertsmithite’s spin correlations extend beyond nearest-neighbor models. These longer-range correlations help explain its disordered quantum-spin-liquid ground state and broad excitation continuum.
- INS measurements show longer-range spatial correlations inconsistent with a nearest-neighbor singlet model.
- Beyond-nearest-neighbor correlations further suppress local magnetic ordering on the geometrically frustrated kagome lattice.
- A broad 2–11 meV excitation continuum, rather than sharp spin-wave peaks, supports fractionalized excitations in the spin-liquid state.
- Narrower reciprocal-space features in the energy-integrated structure factor imply correlations extending beyond nearest neighbors.
2. Impact of beyond-nearest-neighbor correlations on the ground state spin liquid
Beyond-nearest-neighbor correlations further suppress local-moment ordering tendencies while contributing to the unresolved structure of the gapless spin-liquid state. The supplied passages also describe MSE-based image enhancement as producing smooth predictions that sacrifice high-frequency detail.
- Longer-range correlations enhance competing quantum fluctuations and frustration pathways, further suppressing local-moment ordering in the kagome lattice.
- The gapless spin-liquid mechanism remains unresolved: longer-range correlations must be related to spinon dispersion across the full momentum range.
- Experiments suggest short-range static correlations coexist with long-range quantum coherence, but their quantitative coupling to entanglement and excitation spectra is unclear.
- Multiple extended-Hamiltonian candidates can match parts of the data, yet falsifiable high-precision predictions for the full S(Q, ω) remain difficult.
- MSE minimizes the conditional expectation in ill-posed enhancement tasks, averaging plausible solutions into smooth predictions that suppress edges and textures.
- Downsampling can discard high-frequency information, while MSE provides no incentive to reconstruct those details through skip connections or texture synthesis.
5. Metric implications (PSNR/SSIM divergence)
PSNR and SSIM can remain poor for different reasons when MSE-based enhancement favors blurry mean predictions over structural coherence. The passages also include benchmark table captions without quantitative table contents.
- Low PSNR relative to SOTA is attributed to convergence toward blurry mean predictions and a higher absolute MSE than methods using better inductive biases or combined losses.
- Poor SSIM follows from optimizing pixel-wise similarity rather than luminance, contrast, and structural correlations, destroying local contrast and coherence.
- Spatially varying attenuation and backscatter make MSE’s globally uniform L2 penalty unsuitable for the local contrast enhancements required in underwater restoration.
- Tables 6–8 are identified as sampling parameters and cross-domain Expert Score comparisons, but the supplied captions provide no numerical values.