Source-linked AI summary
GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
Jize Wang, Xuanxuan Liu, Yining Li, Songyang Zhang, Yijun Wang, Zifei Shan, Xinyi Le, Cailian Chen, Xinping Guan, Dacheng Tao
TL;DR
Existing tool-use benchmarks provide limited evidence about authentic, multimodal, long-horizon agent work and system-level execution. GTA-2 addresses this with a hierarchical benchmark and recursive checkpoint evaluation, finding a sharp workflow capability gap and substantial benefits from advanced execution frameworks.
Problem
Existing benchmarks often rely on AI-generated queries, dummy tools, and simplified settings, while providing limited insight into realistic workflows and execution-framework effects.
Method
GTA-2 combines GTA-Atomic and GTA-Workflow with recursive checkpoints that decompose open-ended objectives into verifiable sub-goals and evaluate both models and execution harnesses.
Results
Frontier models perform below 50% on atomic tasks, while workflow performance falls to 14.39% root success rate; advanced frameworks substantially improve workflow completion.
Takeaways & Limitations
GTA-2 highlights execution-harness design, alongside model capability, as important for completing realistic long-horizon workflows.
Takeaways & Limitations
GTA-2 is a capability benchmark rather than a complete characterization of workflow distributions, harness causality, deployment safety, or workflow-failure causality.
Abstract
from arXiv · showhide
The development of general-purpose agents requires a shift from executing simple instructions to completing complex, real-world productivity workflows. However, current tool-use benchmarks remain misaligned with real-world requirements, relying on AI-generated queries, dummy tools, and limited system-level coordination. To address this, we propose GTA-2, a hierarchical benchmark for General Tool Agents (GTA) spanning atomic tool use and open-ended workflows. Built on real-world authenticity, it leverages real user queries, deployed tools, and multimodal contexts. (i) GTA-Atomic, inherited from our prior GTA benchmark, evaluates short-horizon, closed-ended tool-use precision. (ii) GTA-Workflow introduces long-horizon, open-ended tasks for realistic end-to-end completion. To evaluate open-ended deliverables, we propose a recursive checkpoint-based evaluation mechanism that decomposes objectives into verifiable sub-goals, enabling unified evaluation of both model capabilities and agent execution frameworks (i.e., execution harnesses). Experiments reveal a pronounced capability cliff: while frontier models already struggle on atomic tasks (below 50%), they largely fail on workflows, with top models achieving only 14.39% success. Further analysis shows that checkpoint-guided feedback improves performance, while advanced frameworks such as Manus and OpenClaw substantially enhance workflow completion, highlighting the importance of execution harness design beyond the underlying model capacity. These findings provide guidance for developing reliable personal and professional assistants. Dataset and code will be available at https://github.com/open-compass/GTA.
1 Introduction
GTA-2 addresses the gap between simplified atomic tool-use benchmarks and realistic, long-horizon productivity workflows with a hierarchical benchmark and checkpoint-based evaluation. Experiments show a sharp performance decline on workflows and highlight execution-harness design as an important factor.
- Motivation: Existing benchmarks often use AI-generated queries, dummy tools, and text-only settings, limiting assessment of authentic multimodal problem solving.These simplifications may explicitly provide solution steps or tool choices and do not test genuine end-to-end execution.
- Benchmark: GTA-2 unifies GTA-Atomic for short-horizon, closed-ended tool-use precision with GTA-Workflow for long-horizon, open-ended productivity tasks.GTA-Workflow is an independent setting targeting end-to-end completion under realistic constraints rather than a simple extension of atomic tasks.
- Evaluation: Recursive checkpoint-based evaluation decomposes workflow objectives into verifiable sub-goals without requiring predefined trajectories.The mechanism supports consistent and interpretable evaluation of open-ended deliverables.
- Results: 14.39% success was achieved by Gemini-2.5-Pro on workflows, while checkpoint-guided feedback moderately improved performance and Manus and OpenClaw substantially enhanced completion.The evaluation describes a pronounced capability cliff from atomic tasks to workflow settings.
- Evaluation: GTA-Workflow jointly evaluates LLM capabilities and execution harnesses, enabling analysis of how system design affects final outcomes.The benchmark treats execution frameworks as part of the evaluated agent system.
2 Related Work
Related work has expanded from isolated tool calls toward complex workflows, but existing benchmarks still leave gaps in realism, multimodality, and evaluation of execution frameworks. GTA-2 responds with a general, hierarchical, and jointly diagnostic framework.
- Tool-use agents: Tool-use research established agents that interpret intent, select tools, and generate executable actions through API invocation and interleaved reasoning.Toolformer and ReAct exemplify these foundational paradigms.
- Execution frameworks: Recent execution frameworks increasingly integrate tools, memory, and coordination, moving beyond fixed pipelines for long-horizon interactions.The cited systems include Agent Operating Systems, OpenClaw, and MiniMax Agent.
- Long-horizon workflows: Open-ended workflow agents differ from atomic-task systems because success depends on final deliverables across long action horizons and flexible solution paths.Examples include Claude Code, Kortix, and Manus.
- Benchmark gap: Existing benchmarks often rely on synthetic queries or simulated environments and provide limited insight into how execution frameworks influence end-to-end completion.This limitation is especially pronounced for long-horizon workflows.
- GTA-2: GTA-2 combines authentic tools, user queries, and multimodal contexts with general-purpose tasks, an atomic-to-workflow hierarchy, and verifiable deliverable evaluation.It also supports unified evaluation of LLMs and execution frameworks.
3 Hierarchical Design of GTA-2 Benchmark
GTA-2 is designed as a hierarchy that extends realistic atomic tool-use evaluation to independent, open-ended workflow completion. The two levels enable systematic analysis from precise tool execution to complex deliverables.
- Design motivation: GTA-2 builds on realistic evaluation using real user queries, executable deployed tools, and multimodal environments.The design preserves real-world fidelity while extending beyond atomic tasks.
- GTA-Workflow: GTA-Workflow evaluates long-horizon, open-ended productivity tasks by their final deliverables rather than predefined execution trajectories.Its flexible solution processes create distinct benchmark and evaluation requirements.
- Hierarchical benchmark: Together, GTA-Atomic and GTA-Workflow form a unified hierarchy for analyzing agent capabilities from precise tool execution to complex workflow completion.The hierarchy spans multiple levels of task complexity.
3.2 Design Principles of GTA-Workflow
GTA-Workflow evaluates realistic, open-ended productivity tasks through authentic inputs and deliverables rather than prescribed execution trajectories. Its checkpoint hierarchy specifies verifiable outcomes while allowing diverse execution strategies and supporting analysis across models and execution systems.
- Real-world authenticity: Real-world authenticity grounds GTA-Workflow in genuine user queries, deployed tools, and multimodal contexts.These elements are intended to make workflow tasks resemble practical scenarios rather than synthetic constructions.
- Deliverable-oriented formulation: Workflow evaluation centers on final deliverables such as reports, code, and multimedia artifacts instead of intermediate execution steps.The design avoids trajectory-level evaluation because open-ended tasks can have multiple valid solution paths.
- Goal-driven decomposition with flexible execution: A hierarchy of goal-oriented checkpoints decomposes each objective into verifiable sub-goals without prescribing actions.This decouples evaluation from specific execution procedures and permits diverse agent strategies.
- Goal-driven decomposition with flexible execution: The checkpoint design enables consistent evaluation across diverse LLMs and execution systems, supporting analysis of model capability and harness design.The same task can therefore be assessed across different combinations of models and execution frameworks.
3.3 Open-ended Workflow Evaluation
GTA-Workflow is constructed from real-world needs and evaluated through multimodal, tool-rich, checkpoint-driven workflows. A semi-automatic generation and verification pipeline refines tasks and recursively scores final deliverables against weighted sub-goals.
- Evaluation framework: The workflow benchmark is organized around task sourcing, multimodal tools, checkpoint formulation, task construction, and checkpoint-based evaluation.These five components structure the benchmark’s workflow evaluation process.
- Task sourcing: A dual-source strategy combines cases from agent platforms with refined high-engagement requests from Reddit and Stack Exchange.The sources are selected to reflect current deployment scenarios and authentic human demands.
- Multimodal ecosystem and tools: GTA-Workflow supports images, documents, audio, and video, while expanding the executable tool set from 14 to 37 tools.The expanded environment supports operations such as audio processing, document editing, and video manipulation.
- Checkpoint-driven formulation: Each open-ended objective is decomposed into a tree of verifiable sub-goals, with checkpoints specifying target states rather than action sequences.Subtasks form a task-to-sub-task hierarchy and carry associated weights for fine-grained assessment.
- Task construction: Task construction combines LLM generation, controlled refinement or augmentation, automatic validation, and human verification.Validation rejects action-oriented criteria or predefined execution steps, while annotators check correctness, feasibility, and alignment.
- Task construction: 154 raw tasks were collected and 132 retained after filtering and rewriting; 67 underwent augmentation, 62 refinement, and 3 passed unchanged.Augmentation added averages of 3.57 constraints, 1.18 deliverable requirements, and 3.48 tools per task, while refinement added 4.45 structural constraints and 1.81 deliverable requirements on average.
- Checkpoint-based evaluation: Recursive checkpoint scoring evaluates leaf deliverables with an LLM judge and aggregates child scores using normalized checkpoint weights.Leaf scores range from 0 to 10, and the judge assesses final artifacts against goal-oriented requirements rather than reasoning or tool usage.
3.4 Dataset Statistics
GTA-2 combines short-horizon atomic tool-use tasks with broader long-horizon workflow tasks. GTA-Workflow increases tools, subtasks, checkpoints, modalities, and deliverable diversity while removing fixed execution trajectories.
- Benchmark composition: GTA-2 integrates GTA-Atomic and GTA-Workflow into a hierarchical benchmark spanning structured tool use and open-ended workflow completion.The two components jointly cover complementary evaluation settings.
- GTA-Atomic: GTA-Atomic contains 229 tasks, 728 total steps, and 14 executable tools, with each task using 1–4 tools across 2–8 steps.Its tasks emphasize perception-grounded reasoning and structured text or image outputs.
- GTA-Workflow: GTA-Workflow contains 132 open-ended tasks, 37 tools, and 1156 subtasks organized into 3–19 checkpoints per task.It broadens inputs to documents, audio, and video and covers outputs such as reports, code, and structured files without fixed trajectories.
4 Experimental Setup
Experiments evaluate atomic precision and workflow completion across multiple models, execution frameworks, and task metrics. Atomic evaluation uses step-by-step and end-to-end modes, while workflows are scored through checkpoint-based success and efficiency measures.
- Models: The experiments evaluate 8 representative LLMs on GTA-Atomic and 13 frontier models on GTA-Workflow.The model sets include both closed-source and open-source systems.
- Agent frameworks: The default setup uses Lagent with ReAct, alongside OpenClaw, Manus, and Kortix harnesses that provide planning, memory, and tool coordination.The framework comparison is designed to assess execution-harness effects beyond the base model.
- Evaluation modes: GTA-Atomic uses step-by-step prediction without execution and end-to-end execution assessed by tool selection and final outcomes.Step-by-step mode aligns predictions with ground truth, whereas end-to-end mode measures dynamic tool invocation.
- Atomic metrics: Atomic metrics include InstAcc, ToolAcc, ArgAcc, SummAcc, AnsAcc, and category-specific tool-selection F1 scores.AnsAcc w/ ImgGen additionally evaluates image-generation parameters.
- Workflow metrics: Workflow evaluation uses Root Score, Root SR, Leaf SR, Tool SR, and capability-specific success rates across Perception, Operation, Logic, and Creativity.Root SR counts tasks whose root score exceeds the default threshold k = 7.
- Efficiency metrics: Workflow framework efficiency is measured by Total Time, Total Cost, and Score-to-Cost Ratio.These metrics capture runtime, API expenditure, and performance relative to computational cost.
5 Main Results
GTA-2 reveals substantial difficulty in both atomic tool use and long-horizon workflow completion. Results also show that execution harness design materially affects workflow outcomes, while failures concentrate in execution, integration, and deliverable realization.
- GTA-Atomic: Fewer than 50% of problems are correctly solved by GPT-4 and GPT-4o on GTA-Atomic, while other models solve fewer than 25%.Real-world tasks combine implicit steps, deployed tool calls, and multimodal inputs.
- Model performance: 14.39% Root SR is achieved by Gemini-2.5-Pro on GTA-Workflow despite a 91.20% Tool SR.The gap indicates that correct tool invocation does not ensure sustained workflow completion.
- Harness comparison: 50.0% Root SR is achieved by OpenClaw versus 0.0% for Lagent with the same Claude-Sonnet-4.5 base model.OpenClaw also raises Root Score from 2.49 to 6.82 and Leaf SR from 10.14% to 73.55%.
- System-level comparison: Advanced systems achieve Root Scores of approximately 6.8–6.9 and Success Rates above 50%, while their Leaf SR exceeds 65%.These system-level results combine model, harness, and product engineering effects.
- Failure analysis: Over 40% of failures are formatting-related across every harness, while reasoning errors remain below 4%.Advanced harnesses reduce content synthesis failures from 29.4% to approximately 20%–23%, but data extraction failures become relatively more prominent.
- Failure analysis: 77.78% and 80.56% are the deliverable-level failure rates for Gemini-2.5-Pro and Claude-Sonnet-4.5 with default Lagent, respectively.Composition-level failures are also around 70%, showing difficulty translating partial sub-goal completion into correct final outputs.
6 Additional Analysis on GTA-Workflow
Additional analyses show that workflow performance declines with greater operational depth and varies by deliverable and domain. They also examine efficiency, metric sensitivity, evaluation reliability, and checkpoint-guided feedback.
- Model capability: 11.36%–14.39% success rates are achieved by frontier models, while smaller or earlier-generation models fall below 1%.Leading open-source models form a second tier at around 10.61%.
- Task complexity: Leaf success generally declines as checkpoint-tree complexity increases from Short workflows with 3–7 nodes to Long workflows with 13–19 nodes.Performance converges across models at high complexity, making operational depth and sub-goal coordination key bottlenecks.
- Deliverable types: Text deliverables such as PDF, plain text, and HTML receive the best performance, while multimedia outputs average 3.48.The analysis indicates that final-artifact type influences task difficulty in addition to reasoning depth.
- Workflow domains: No single model leads across all workflow domains: Gemini-2.5-Pro leads Retrieval & QA, whereas Claude-Sonnet-4.5 performs slightly better in Creative Design.The results indicate complementary domain-specific strengths.
- Evaluation metrics: 89.85% Tool SR corresponds to only 8.33% Root SR for Kimi-K2, showing that tool-call correctness alone misses deliverable-level efficacy.GTA-Workflow therefore evaluates whether tools achieve the final goal across sustained interactions.
- Efficiency: 14.39% Root SR is achieved by Gemini-2.5-Pro with relatively moderate step consumption, placing it at the apex of the reported efficiency frontier.Grok-4 uses more redundant steps, while leading open-source agents approach the frontier.
- Evaluation reliability: Pearson 0.966 and ICC 0.928 measure agreement between the LLM judge and human evaluation at task level.A cross-model validation study reports root-level Pearson correlations above 0.92 for all sampled model sources.
- Feedback: 12.03% improvement over the initial attempt is obtained with checkpoint feedback, compared with 4.05% from coarse feedback.Fine-grained checkpoint diagnostics provide a stronger correction signal than generic retry instructions.
7 Conclusion
GTA-2 presents a hierarchical benchmark for atomic tool use and long-horizon workflows, using deliverable-centric checkpoint evaluation. Its reported workflow results emphasize execution harness design, while the benchmark remains bounded by construction, safety, and causal-interpretation limitations.
- Conclusion: GTA-2 spans atomic tasks and long-horizon workflows through a deliverable-centric benchmark for realistic tools and multimodal settings.GTA-Workflow evaluates complex deliverables with a checkpoint-based framework.
- Conclusion: 14.39% root success rate is achieved by frontier models on workflows, while Manus and OpenClaw substantially improve performance.The conclusion attributes these results to the critical role of execution harness design beyond model capability.
- Limitations: GTA-2 is a capability benchmark rather than a complete characterization of workflow distributions, isolated harness causality, deployment safety, or failure causality.Workflow tasks are partially constructed through LLM-based reformulation, which may introduce construction bias.
- Limitations: High benchmark scores do not necessarily imply safe or reliable real-world deployment.Safety, authority control, privacy protection, and governance remain outside the current evaluation scope.
Additional GTA-2 Information
The additional materials describe GTA-2’s tools, agent implementation, benchmark construction, failure categories, and task-rewriting procedures. GTA-Atomic focuses on realistic short-horizon tool use, while GTA-Workflow expands evaluation toward structured, open-ended workflows.
- Agent system: The agent system uses Lagent with ReAct as the action-and-planning schema and 37 tools implemented through AgentLego.Different LLMs are substituted into the Lagent framework for evaluation.
- Quality control: All tasks undergo manual quality inspection, while checkpoint regeneration and validation prompts support task quality and realism.The materials list prompts for raw-query construction, classification, refinement, augmentation, rewriting, validation, and checkpoint regeneration.
- Failure taxonomy: Failure analysis distinguishes data extraction, reasoning, content synthesis, and formatting errors.Formatting includes file format, filename, layout, embedding, export, packaging, and delivery-artifact compliance.
- GTA-Atomic construction: GTA-Atomic samples contain inputs, real-world queries, involved tools, reference tool chains, and final answers.The benchmark evaluates short-horizon, closed-ended tool-use tasks in realistic settings.
- Tools: GTA-Atomic uses 14 tools across perception, operation, logic, and creativity, while GTA-Workflow adds an extended tool set.The tool categories organize the benchmark’s executable capabilities.
- Task construction: The construction pipeline classifies tasks as DELETE, REFINE, AUGMENT, or PASS before applying targeted modifications.Refinement clarifies requirements and output formats, while augmentation expands complexity and tool diversity.
- Checkpoint construction: A hierarchical checkpoint tree is required, with result-oriented checkpoints and clearly defined success criteria.Checkpoint modification accompanies task complexity enhancement.
- Task construction: Task rewriting removes explicit tool-call phrasing so models must reason and plan tool use from required outcomes.The rewritten task preserves requirements and deliverables in a coherent narrative.
1. Task-level Quality Control
Task-level quality control combines clear objective criteria, attachment-aware inspection, multimodal deliverable checks, and checkpoint-based failure classification across GTA-Atomic and GTA-Workflow examples.
- Quality criteria: Quality control requires task objectives to be clear, unambiguous, and independently understandable in terms of the expected deliverable.Inspection must also be conducted alongside task attachments and input files.
- Multimodal inspection: The evaluation prompt inspects generated files and images, converts videos into image frames, and checks audio files for existence.This procedure is applied using the original task as evaluator context.
- Failure classification: Failure decomposition distinguishes leaf-level failures from mid-level composition failures across checkpoint outcomes.Leaf failures affect local sub-goals or substeps, whereas composition failures concern integration or coordination across mostly completed sub-goals.
- GTA-Atomic task types: GTA-Atomic examples cover objective, subjective, and image-generation queries with distinct answer requirements and tool combinations.Objective queries seek uniquely determined answers, subjective queries accept descriptive responses, and image-generation queries do not directly evaluate the generated image.
- Checkpoint-based evaluation: Checkpoint-based evaluation decomposes workflow deliverables into verifiable components, such as physical modeling, video synthesis, system integration, and workbook generation.Examples include a responsive HTML5 presentation with three labeled videos and an Excel workbook with dedicated game worksheets and a summary sheet.
- GTA-Workflow examples: GTA-Workflow examples span education, retrieval and question answering, and planning and decision tasks with condensed queries and checkpoints.Representative deliverables include a physics presentation, an Italian lottery workbook, and a Japan travel handbook.