Source-linked AI summary

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, Duo Wang, Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, Chuqian Yu, Tianhao Yu, Longxiang Liu, Jianbo Xue, Huimin Che, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang

arXiv:2608.17800v1cs.AI

TL;DR

Existing benchmarks may not reflect the real-world workflows users demand from AI or the requirements of professionally usable deliverables. StartupBench grounds end-to-end evaluation in adopted AI startup workflows and finds that the strongest model completes only about 30% of benchmark tasks.

  • Problem

    Existing benchmarks often use researcher-defined tasks and coarse evaluations, leaving uncertain whether agents can produce professionally usable deliverables that reflect authentic user needs.

  • Method

    StartupBench derives 97 end-to-end tasks across six domains from adopted AI startup workflows through surveys and interviews, then evaluates deliverables against rubric items.

  • Results

    About 30% of benchmark tasks were completed by the strongest model, while performance varied substantially across domains.

  • Takeaways & Limitations

    Many market-validated workflows remain beyond reliable general-purpose-agent capabilities, making StartupBench a realistic measure of practical agent performance.

  • Takeaways & Limitations

    Pilot models were randomly selected from three model families to construct a challenging benchmark rather than estimate performance over the natural task distribution.

Abstract

from arXiv · show

Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.

1 Introduction

StartupBench grounds end-to-end agent evaluation in AI workflows validated by real-world adoption, covering complete deliverables and fine-grained quality requirements. On this benchmark, even the strongest model completes only about 30% of tasks, with complex instruction following and domain-specific expertise driving major failures.

  • Benchmark construction: StartupBench derives realistic end-to-end tasks from AI-native startup workflows validated through adoption, product demonstrations, and user interviews.The benchmark preserves users’ core objectives and practical requirements.
  • Benchmark construction: 97 tasks across 6 primary domains capture realistic end-to-end user requirements identified through survey- and interview-driven construction.The benchmark is built through startup research, enterprise-user interviews, task abstraction, expert data construction, and multi-stage quality control.
  • Evaluation design: Each task requires complete user-facing deliverables rather than intermediate outputs or simplified demonstrations.This design targets professionally usable results in realistic workflows.
  • Evaluation design: Fine-grained rubrics independently assess functionality, structure, formatting, and domain-specific quality to evaluate complex deliverables comprehensively.These requirements address functional, structural, formatting, and domain-specific dimensions that holistic evaluations may miss.
  • Results: 30%: even the strongest model successfully completes only about 30% of StartupBench tasks under a unified agentic harness.Finance, STEM & Computer Science, and Education & Humanities are particularly challenging; failures are primarily attributed to complex instruction following and domain-specific expertise.

2 Related Work

Related work shows a progression from evaluating isolated agent capabilities toward end-to-end completion of realistic, economically meaningful workflows. As tasks increasingly produce heterogeneous deliverables, evaluation correspondingly shifts toward deliverable-centric assessment.

  • Workflow-Oriented Products: Foundation-model advances have enabled agents to acquire planning, tool-use, file-processing, and multi-step execution capabilities for workflow-oriented products.Commercial products package these capabilities into vertical workflows spanning document processing, data analysis, financial research, enterprise operations, and knowledge work.
  • Agent Benchmarks: Agent benchmarks have progressed from tool use and reasoning capabilities toward end-to-end completion in interactive, professional, and economically meaningful environments.Early benchmarks include AgentBench and GAIA, while later work evaluates web interaction, computer use, and professional tasks.
  • Deliverable-Centric Assessment: Heterogeneous outputs such as reports, spreadsheets, slides, code patches, and structured files motivate a shift from capability-centric scoring to deliverable-centric assessment.LLM-as-a-Judge methods, including MT-Bench and Prometheus, support preference-based or rubric-conditioned evaluation, while agentic workflows require evidence retrieval and file inspection.

3 StartupBench

StartupBench is built from market-validated AI-product workflows and reconstructed into realistic, deliverable-oriented end-to-end tasks with explicit success criteria. It evaluates heterogeneous outputs through fine-grained, rubric-level AgentJudge decisions using evidence views of submitted deliverables.

  • Task Design: StartupBench requires tasks to be realistic, answerable, evaluable, and discriminative, with sufficient information, clear outputs, assessment criteria, and challenging requirements.Tasks must originate from actual AI-product usage while remaining reproducible and capable of distinguishing current models and agents.
  • Evaluation: Each task is represented as T = (q, E, R), and evaluation uses an evidence view plus independent AgentJudge sessions that produce binary rubric decisions and textual justifications.Most tasks contain 20+ fine-grained rubric items covering functionality, completeness, formatting, and domain-specific requirements.
  • Benchmark Composition: 97 real-world workflow tasks span 6 top-level domains and require agents to process heterogeneous files, follow complex constraints, and produce usable professional deliverables.Supported output formats include DOCX, XLSX, PPTX, PDF, Markdown, images, and text-based deliverables.

4 Experiments

StartupBench remains far from saturated: even the strongest models achieve average scores near 74% but fail to complete one third of tasks under the strict criterion. Performance varies by domain and rubric, with core requirements, domain compliance, instruction following, and output-format constraints driving failures.

  • Overall performance: 73.67% and 73.61% are the highest average scores, achieved by Kimi-K3 and GPT-5.6-sol, yet no model completes one third of tasks at score ≥90.Most models average roughly 55–75, indicating substantial partial progress without reliable end-to-end completion.
  • Overall performance: Kimi-K3 has the highest average score but lower success rate than GPT-5.6-sol, showing that partial requirement satisfaction does not ensure complete delivery.The same discrepancy appears between Kimi-K2.6 and Qwen-3.6-Max, while all models have lower success rates than average scores.
  • Domain performance: 54.48% is Finance’s lowest average score, whereas Business is consistently easiest, with every model above 60 and the highest average success rate.Competence varies substantially across the six professional domains.
  • Domain performance: Kimi-K3 and GPT-5.6-sol score 73.67% and 73.61% overall but lead different domains, and no model consistently dominates across all professional areas.Kimi-K3 leads Medical, Business, Finance, and Education by average score; GPT-5.6-sol leads Legal and STEM.
  • Evaluation and failure analysis: 92.78% rubric-level and 92.84% task-level agreement with experts validates the automatic evaluator, while holistic judging lowers agreement to 83%.Output-format violations, complex instruction following, and domain-specific compliance remain recurring failure sources.
  • Failure analysis: 68.67%, 65.89%, and 63.45% are the average satisfaction rates for Auxiliary, Important, and Core rubrics, respectively, with 8 of 9 models showing this decline.Models often satisfy peripheral requirements while missing core requirements that determine successful completion.

5 Conclusion

StartupBench evaluates general-purpose agents on end-to-end professional workflows derived from market-validated AI products and real-world user demands. Its results show a substantial gap between meaningful progress and deliverables that satisfy practical acceptance criteria, driven by several capability limitations.

  • 5 Conclusion: StartupBench benchmarks models on end-to-end professional workflows derived from market-validated AI-native products and real-world user demands.The benchmark spans multiple tasks across diverse domains.
  • 5 Conclusion: The evaluation reveals a substantial gap between making meaningful progress on professional work and producing deliverables that fully satisfy practical acceptance criteria.This gap appears across multiple tasks spanning diverse domains.
  • 5 Conclusion: The gap is primarily driven by limitations in complex instruction following, domain-specific expertise, professional conventions, and long-horizon workflow execution.These limitations constrain reliable completion of practical professional workflows.

6 Contributions · Appendix

The paper credits project leads, core contributors, additional contributors, and a sponsor committee for StartupBench. The contributions are distributed across these named groups.

  • 6 Contributions: Project leads are Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, and Ge Zhang.This group is explicitly identified as the project leads.
  • 6 Contributions: Core contributors include Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, and additional named collaborators.The passage lists 18 core contributors in total.
  • 6 Contributions: Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, and Yongzhi Liao are among the listed core contributors.These names are part of the core-contributor roster.
  • 6 Contributions: Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, and Duo Wang complete the core-contributor list.These names appear in the same core-contributor passage.
  • 6 Contributions: Additional contributors are Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, and six other named collaborators.The passage lists 10 contributors in this group.
  • 6 Contributions: The sponsor committee consists of Yujia Qin and Jiaheng Liu.They are identified specifically as the sponsor committee.

A Tutorial of StartupBench Task Annotation · A.1 Task Construction

StartupBench tasks are structured as triples combining a natural-language query, a self-contained workspace, and weighted evaluation rubrics. This design preserves real-world workflow structure while enabling systematic assessment across multifaceted requirements.

  • A.1 Task Construction: Each StartupBench task is defined as a triple T = (q, E, R) that experts fully specify.The three elements are the input query, task workspace, and evaluation rubrics.
  • A.1 Task Construction: The input query q states the user’s task instruction, including realistic background, implicit constraints, and expected deliverables.Queries are derived from authentic annotator work scenarios to reflect realistic user intent.
  • A.1 Task Construction: The workspace E is a self-contained execution environment containing multimodal source files, interaction settings, and all necessary input artifacts.Artifacts can include files, datasets, or structured resources, ensuring sufficient input information for task completion.
  • A.1 Task Construction: The rubric R = {(pi, wi)}n i=1 is a checklist of natural-language scoring points pi paired with positive importance weights wi.Weights reflect each criterion’s importance to successful task completion.
  • A.1 Task Construction: Multiple rubric items evaluate different dimensions because real-world workflows are complex and multifaceted.The rubrics provide fine-grained criteria for assessing task-completion quality.
  • A.1 Task Construction: Together, the components preserve real-world agent workflows while supporting systematic and objective evaluation.Tasks may involve information synthesis, constraint following, structured generation, or domain-specific reasoning and judgment.

A.2 Cross Validation … D Examples of Behavior-level Failure Modes

StartupBench uses domain-expert cross-validation to ensure task authenticity, workflow fidelity, rubric validity, and ground-truth consistency. Its evaluation framework spans six rubric dimensions and weights requirements according to their importance for successful completion.

  • A.2 Cross Validation: Each task is independently reviewed by at least one additional domain expert, who validates realism and correctness from multiple complementary perspectives.Reviewers assess more than annotation consistency during construction and cross-validation.
  • A.2 Cross Validation: Reviewers verify that task descriptions, workspaces, inputs, materials, and domain knowledge accurately reflect authentic user requests.Tasks containing factual or domain-specific knowledge errors are identified during review.
  • A.2 Cross Validation: Reviewers assess workflow fidelity, revising or discarding tasks that deviate from practical procedures or genuine professional responsibilities.This ensures benchmark tasks represent real-world workflows in the corresponding profession.
  • A.2 Cross Validation: Reviewers inspect rubrics for meaningful capabilities, clear wording, unambiguous criteria, and weights that reflect task priorities.Rubric validity is evaluated against the target workflow rather than domain-independent annotation consistency alone.
  • B Domain Expert Background: 57 domain experts cover all 6 StartupBench domains and are assigned according to their reported expertise and professional experience.Their backgrounds span medicine, law, finance, software and AI, engineering, management, education, and humanities.
  • C Rubrics of StartupBench: The benchmark organizes rubrics into six dimensions: correctness, completeness, data processing, professional reasoning, engineering quality, and presentation.These categories recur across heterogeneous office tasks rather than being tied to specific application domains.
  • C Rubrics of StartupBench: 38.9% of the 2,453 rubrics are Calculation Precision, while Structure & Completeness and Domain-Specific Compliance together account for 42.8%.The distribution emphasizes objective correctness alongside complete deliverables and professional domain reasoning.
  • C Rubrics of StartupBench: 59.10%, 34.33%, and 6.57% of average task weight belongs to Core, Important, and Auxiliary rubrics, respectively.This weighting makes requirements that materially affect correctness, reliability, safety, or usability dominant in task scores.

D.1 Self-Verification Hallucination

A representative case shows self-verification failure: the model produced a seemingly complete Excel workbook and incorrectly concluded that the task was finished despite artifact-level inspection.

  • D.1 Self-Verification Hallucination: The model filtered level-11 members, simulated route check-ins toward level 12, and generated a largely complete workbook, leading it to declare success.The workbook contained the required worksheet, member records, and output columns, while the task also imposed precise value and formatting constraints.

D.2 Domain-Specific Failure

Domain-specific failures can persist despite polished outputs: in a clinical example, the model produced a professional treatment plan that violated multiple safety-critical constraints because it could not turn medical knowledge into reliable actions.

  • Clinical actionability: A well-formatted multidisciplinary treatment plan violated multiple safety-critical clinical constraints, reflecting insufficient domain-specific actionability rather than poor writing quality.The failure arose from inability to convert medical knowledge into actionable, clinically reliable decisions.

Self-Verification Hallucination

The agent performs selected checks but mistakes execution summaries for true verification, failing to validate the final artifact against source semantics and task requirements. A related clinical case shows that a well-structured MDT plan can still conflict with case-specific criteria.

  • Self-Verification Hallucination: The agent checks selected calculations but fails to reconcile the final workbook with source names and output-field semantics.This reflects a verification gap at the artifact level rather than an absence of execution checks.
  • Self-Verification Hallucination: Artifact-level spot checks expose acceptance-critical errors that the model misses when it treats a progress summary as verification.The model generates a plausible workbook and concludes completion based on its execution process instead of the original task requirements.
  • Medical Domain: Clinical Actionability: The submitted MDT plan is well structured but conflicts with case-specific medication, timing, and handoff criteria.This exemplifies insufficient domain-specific clinical actionability despite a coherent overall structure.
Loading 2608.17800v1…