Source-linked AI summary

xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations

Kaiyuan Chen, Yixin Ren, Yang Liu, Xiaobo Hu, Haotong Tian, Tianbao Xie, Fangfu Liu, Haoye Zhang, Hongzhang Liu, Yuan Gong, Chen Sun, Han Hou, Hui Yang, James Pan, Jianan Lou, Jiayi Mao, Jizheng Liu, Jinpeng Li, Kangyi Liu, Kenkun Liu, Rui Wang, Run Li, Tong Niu, Wenlong Zhang, Wenqi Yan, Xuanzheng Wang, Yuchen Zhang, Yi-Hsin Hung, Yuan Jiang, Zexuan Liu, Zihan Yin, Zijian Ma, Zhiwen Mo

arXiv:2506.13651v1cs.LG

TL;DR

Existing benchmarks often emphasize isolated technical capabilities rather than the productivity and commercial value of agents in professional work. xbench addresses this gap with dynamic, profession-aligned evaluations built from real-world domain tasks and metrics intended to relate scores to value, predict TMF, and track products over time. Its initial Recruitment and Marketing benchmarks establish baseline evaluations for professional agents.

  • Problem

    Existing AI benchmarks focus on technical capabilities and may not predict agents’ real-world economic impact or productivity value.

  • Method

    xbench constructs profession-aligned evaluation sets from real-world professional demands, using domain-oriented metrics and long-term updated evaluations.

  • Results

    The initial xbench batch covers Recruitment and Marketing and reports evaluations of specialized and general-purpose agents.

  • Takeaways & Limitations

    xbench is intended to support productivity-value assessment, Technology-Market Fit prediction, and tracking of competition and capability growth among agent products.

  • Takeaways & Limitations

    LLM-Judge scoring uses Gemini-2.5-Flash, with metric stability, human alignment, and cross-judge consistency left for subsequent analysis.

Abstract

from arXiv · show

We introduce xbench, a dynamic, profession-aligned evaluation suite designed to bridge the gap between AI agent capabilities and real-world productivity. While existing benchmarks often focus on isolated technical skills, they may not accurately reflect the economic value agents deliver in professional settings. To address this, xbench targets commercially significant domains with evaluation tasks defined by industry professionals. Our framework creates metrics that strongly correlate with productivity value, enables prediction of Technology-Market Fit (TMF), and facilitates tracking of product capabilities over time. As our initial implementations, we present two benchmarks: Recruitment and Marketing. For Recruitment, we collect 50 tasks from real-world headhunting business scenarios to evaluate agents' abilities in company mapping, information retrieval, and talent sourcing. For Marketing, we assess agents' ability to match influencers with advertiser needs, evaluating their performance across 50 advertiser requirements using a curated pool of 836 candidate influencers. We present initial evaluation results for leading contemporary agents, establishing a baseline for these professional domains. Our continuously updated evalsets and evaluations are available at https://xbench.org.

1 INTRODUCTION

xbench introduces profession-aligned evaluations for domain-specific agents, targeting real-world productivity value, Technology-Market Fit, and long-term product competition. Its initial benchmarks cover Recruitment and Marketing, with continuously updated evaluations of specialized and general-purpose agents.

  • xbench evaluates domain-specific agents on real-world demands using domain-oriented metrics designed to assess professional performance.Its construction selects domains based on market size and technological maturity and incorporates domain-expert input.
  • Profession-aligned metrics aim to correlate agent scores with real-world value and predict Technology-Market Fit through market pricing and technology-feasibility curves.The intersection of the curves signifies TMF achievement, while their intersection area represents value generated per unit task.
  • xbench-Index uses long-term updated evaluations to track competing agent products and identify products whose capabilities improve rapidly.The framework is intended to capture dynamic growth in agent application capabilities.
  • The first xbench batch covers Recruitment and Marketing, evaluating professional tasks in commercially significant domains.Recruitment includes 50 headhunting tasks spanning company mapping, information retrieval, and talent sourcing.
  • The paper presents xbench as a framework for discovering, defining, and predicting domain-specific agent products while analyzing competition and technological barriers.It reports evaluations of specialized and general-purpose agents and plans continuous reporting.

2 AI CAPABILITY CENTRIC V.S. PROFESSION-ALIGNED

Capability-centric benchmarks commonly use short, simulated tasks that can saturate and fail to distinguish deeper model capabilities from real domain-agent performance. Profession-aligned evaluations instead prioritize commercially valuable professional work, dynamic real-world environments, and business-aligned feedback.

  • Most existing evaluations focus on short, simulated tasks centered on individual capabilities, despite AI systems performing increasingly longer tasks.Task capability length is reported to approximately double every seven months.
  • Saturating indicators make it difficult to distinguish the deep capabilities of different models, leaving a gap with real domain-specific agent performance.The stated gap separates AI capability-centered evaluation sets from professional-agent performance.
  • Evaluation Direction: Profession-aligned evaluations prioritize scenarios offering significant productivity improvements and commercial value.Capability-centric evaluations instead create scenarios targeting deficiencies or areas where models lack mastery.
  • Task Distribution: Profession-aligned tasks represent domain-expert demands, with each task intended to deliver productivity or business value and metrics aligned to practical outcomes.Domain experts collaborate in co-designing the evaluation system and tasks.
  • Environment: Profession-aligned evaluations seek alignment with real work by requiring interaction with dynamic tools, websites, or systems.A temporally consistent index is proposed to preserve metric comparability across periods.
  • Feedback Mechanism: Profession-aligned feedback aligns scores with key business indicators and uses professional-rubric LLM judges for objectively difficult tasks.Capability-centric metrics focus primarily on task completion and correctness.

3 BUILDING XBENCH

xbench builds profession-aligned evaluations from live expert business demands, prioritizing feasible, testable work and dynamic real-world settings. Its initial Recruitment and Marketing benchmarks focus on information-intensive tasks assessed through domain-specific designs and scoring.

  • Evaluation principles: xbench co-constructs evaluation tasks with professional headhunters and marketing enterprises using historical business data.
  • Evaluation principles: Experts map weekly work time and task importance, then prioritize task components that are feasible and testable with current technology.
  • Evaluation principles: Evaluation tasks are collected live from ongoing business operations and dynamically updated as real demands and environments change.
  • Recruitment benchmark: Recruitment focuses on quantifiable information collection because communication interaction lacks an automated testing foundation.Information collection includes candidate search, resume screening, and talent persona construction, typically occupying over 50% of recruiter time and effort.
  • Recruitment benchmark: Recruitment tasks cover company mapping, People-to-Info, and Info-to-People, assessing industry knowledge, talent search, profile completion, screening, and information trustworthiness.Company Mapping identifies suitable schools, companies, or teams from a job description; People-to-Info completes a target person’s professional history.
  • Recruitment benchmark: The recruitment benchmark scores open-ended responses with an LLM Judge on a 1–5 scale, linearly mapped to 0–100, with hallucination-free accuracy earning the highest score.Figure 3 depicts search results or gathered public experience being matched against annotations and checked for hallucinations.
  • Recruitment benchmark: The recruitment set contains 50 cases: Company Mapping, Info-to-People, and People-to-Info comprise 44%, 30%, and 26%, while 38% take over 40 minutes.The remaining task durations are 12% at 0–5 minutes, 16% at 5–20 minutes, and 34% at 20–40 minutes.
  • Marketing benchmark: Marketing focuses on Influencer Search because client communication, matching, monitoring, and strategy adjustment remain difficult to evaluate due to dynamic and delayed effects.

4 EVALUTIONS

xbench evaluates contemporary agents through web-based interfaces and reports Recruitment and Marketing leaderboards. o3 ranks first on both benchmarks, while comparisons also reveal judge-model and version-related caveats.

  • Setup: All tasks require web access, and the evaluation assesses final results without restricting solution architecture.The evaluated systems include specialized search products and general-purpose agents and models.
  • Setup: Testing occurred in May 2025 using web interfaces with internet search enabled, before significant product version updates.
  • Caveats: All LLM Judge scoring uses Gemini-2.5-Flash, with future analysis planned for metric stability, human alignment, and cross-judge consistency.
  • Results: xbench reports Recruitment and Marketing results in Tables 8 and 9, with theme-specific score leaderboards in Figure 7.
  • Results: o3 ranks first on both benchmarks, while GPT-4o ranks last and Gemini-2.5-Pro performs comparably to Gemini-2.5-Flash.
  • Results: Perplexity-Search outperforms Perplexity-Research on Recruitment, while Deepseek R1 performs lower because it lacks adaptation for search-centric tasks.

5 DISCUSSION

The discussion focuses on evaluating agent capability growth despite changing products, environments, and test sets, then connects performance-cost trade-offs to Technology-Market Fit and evolving human roles.

  • Lifecycle evaluation: Dynamic benchmarks can rank concurrent agents but do not capture capability growth across evaluation periods when environments and tasks change.The paper frames this as the central challenge for tracking agent capabilities over time.
  • Lifecycle evaluation: IRT estimates agent capability from incomplete score matrices using ability, item difficulty, and item discrimination parameters.The model predicts a subject’s probability of answering a test item correctly.
  • Lifecycle evaluation: IRT capability scores can reflect continuous model growth, including advances in Gemini after October 2024 and improvements from DeepSeek-V2 and DeepSeek-r1.Future evaluations will report these scores to track developmental velocity and breakthroughs beyond rankings.
  • Technology-Market Fit: Real-world agent evaluation should balance performance, cost, and latency rather than optimize inference scaling for performance alone.The proposed reporting includes demand, human capability, and optimal supply curves on performance-cost graphs.
  • Technology-Market Fit: TMF is represented by overlap between market-acceptable and technologically accessible regions, with the intersection indicating incremental AI value.The framework distinguishes stages before TMF, after TMF, and specialized-agent development.
  • Technology-Market Fit: Progression toward specialized agents depends on domain experts constructing evaluation systems and guiding agent iteration.Experts shift from directly delivering results toward building professional evaluation and training frameworks and offering AI services at scale.
  • Technology-Market Fit: AI may transform established business processes and production relations by introducing novel approaches to existing needs.The discussion also anticipates re-pricing human contributions according to the relative scarcity of expertise and computational power.

6 RELATED WORKS

Related work spans real-world and vertical-domain evaluations, rubric-based LLM judging, and scaling laws, while highlighting limited guidance for agent evolution and productivity measurement.

  • Real-world evaluation: Existing evaluations cover browser-use, GUI, coding, customer service, database operations, and healthcare agents.These vertical benchmarks have supported domain-specific agent development and commercially successful companies.
  • Real-world evaluation: Profession-aligned evaluation centers tasks designed with domain experts to reflect practical demands and business value.This contrasts with capability-centric evaluation, which prioritizes task variety across capabilities and domains.
  • LLM judging: Rule-based judgments are limited for open-ended outputs, motivating rubric-based LLM judges across domains.The paper situates rubric-based judging within prior work on alignment and evaluation.
  • Scaling laws: Traditional scaling laws relate model performance to compute, parameters, and data, but agent-evolution scaling laws remain relatively scarce.METR reports that the effective task-completion time for AI approximately doubles every seven months.

7 SUMMARY

xbench measures AI-agent productivity through live, expert-defined tasks aligned with professional workflows, beginning with Recruitment and Marketing benchmarks and continuously updated evaluation.

  • 7 SUMMARY: xbench evaluates agents on live, expert-defined tasks from commercially significant fields rather than abstract technical skills.Its stated aim is to assess whether agents deliver tangible business value.
  • 7 SUMMARY: Initial benchmarks cover Recruitment headhunting tasks and Marketing influencer identification for real-world campaigns.The paper establishes baseline results for leading contemporary agents in these domains.
  • 7 SUMMARY: xbench uses continuous updating and IRT to track capability changes across evolving agents and environments.This design addresses the dynamic nature of agent products and their operating environments.

A.1 PROMPTS FOR RESPONSE COLLECTION AND EVALUATION

The response-collection appendix specifies prompts, geographic or type constraints, result truncation, and a fixed output format for recruitment information-search tasks.

  • Recruitment prompts: Recruitment prompts ask agents to identify target search objects from job-requirement descriptions.The prompt frames the agent as a recruitment expert and specifies attention points for the analysis.
  • Recruitment prompts: Unless otherwise specified, recruitment searches consider candidates within China.This is an explicit default assumption in the response-collection prompt.
  • Recruitment prompts: Returned results are truncated to the specified demand when they exceed the requested number of objects.The prompt also instructs agents not to over-search.
  • Recruitment prompts: Recruitment outputs must use a Search Results format listing each search object and its associated entries.The appendix provides the required template for response collection.
  • Recruitment prompts: Info-to-People prompts ask talent-information specialists to identify target people from reference information.Their default scope restricts individuals by country or person type unless otherwise specified.

A.2 COMPLETE EXAMPLES OF THE EVALUATION TASKS

This section presents prompt materials for evaluating recruitment and marketing responses, alongside complete examples of company mapping and People-to-Info tasks.

  • Recruitment response evaluation prompts are provided across three parts.
  • Marketing evaluation materials include an agent-response collection prompt and response-evaluation prompts in two parts.
  • A complete company mapping task example is included.
  • A complete People-to-Info task example is included.
Loading 2506.13651v1…