Source-linked AI summary

Agents' Last Exam

Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang, Tianyu Wang, Yuhan Cao, Yixiao Huang, Chris Duroiu, Haoyun Zhang, Jeffrey Lin, Weishu Zhang, Tyler Zeng, Ying Yan, Bo Liu, Hanson Wen, Mingyang Xu, Xiaoyuan Liu, Zimeng Chen, Weiyan Shi, Amanda Dsouza, Vincent Sunn Chen, Patrick Bryant, Carl Boettiger, Yamini Rangan, Bradley Rothenberg, Kyle Steinfeld, Arvind Rao, Tapio Schneider, Georgios Yannakakis, Laure Zanna, Kaan Ozbay, Ida Sim, Tarek Zohdi, George Em Karniadakis, Jack Gallant, Teresa Head-Gordon, Yushan Li, Wenxi Deng, Tao Sun, Huiqi Wang, Zhun Wang, Justin Xu, Chris Yuhao Liu, Yafei Cheng, Rongwang Hu, Aras Bacho, Shengcao Cao, Zengyi Qin, Yixiong Chen, Hengduan Fan, Hao Liu, Lin Zeng, Shashank Muralidhar Bharadwaj, Litian Gong, Yingxuan Yang, Maojia Song, Ruheng Wang, Zongzheng Zhang, Honglin Bao, Shuo Lu, Jianhong Tu, Zhonghua Wang, Zheng Zhang, Zijiao Chen, Yanqiong Jiang, Zhendong Li, Bohan Lyu, Chang Ma, Peiran Xu, Benran Zhang, Shangding Gu, Haoyue Hua, Haoyang Li, Wanzhe Liao, Chengzhi Liu, Junbo Peng, Haoran Sun, Zechen Xu, Bo Chen, Jiayi Cheng, Yi Jiang, Keying Kuang, Yuan Li, Youbang Pan, Ziyan Rao, Alexander Schubert, Yifan Shen, Vincent Siu, Xiatao Sun, Kangqi Zhang, Xiaopan Zhang, Yuchen Zhu, Ishaan Singh Chandok, Lei Ding, Jingxuan Fan, Andrew Glover, Jiaming Hu, Yiran Hu, Wenbo Huang, Zixin Jiang, Haoran Jin, Lukas Kim, Ming Liu, Yang Liu, Alireza Rafiei, Xuhuan Shen, Kunyang Sun, Sophia Sun, Ting Sun, Eric Wang, Yixin Wang, Hanwen Xing, Sihan Xu, Yuzheng Xu, Zhongxing Xu, Zhiling Yan, Boqin Yuan, Ruiqi Zhang, Yifan Zhang, Zibo Zhao, Liana, Santanu Bosu Antu, Haoyue Bai, Carlo Bosio, Joseph Cavanagh, Patricia Cavazos-Rehg, Tianxing Chen, Xuewen Chen, Yipu Chen, Chenyu Zhu, Chen Dai, Stefano De Castro, Yunfu Deng, Kaustubh Dhole, Jiayuan Ding, Chenchen Du, Zhehang Du, Hao Fan, Run-Ze Fan, Hengyu Fu, Shi Gu, Yifan Gu, Charlie Guo, Baihe Huang, Baixiang Huang, Rimika Jaiswal, Zhihan Jiang, Ran Jin, Erin Kasson, Xin Lan, Joseph Lee, Deren Lei, Chenyu Li, Daofeng Li, Haitao Li, Hongwei Li, Jingyan Li, Xiao Li, Yi Li, Yinsheng Li, Yuangang Li, Zhixu Li, Wenyu Liang, Longtai Liao, Kevin Qinghong Lin, Andy Zeyi Liu, Che Liu, Jiaming Liu, Kaiyuan Liu, Xuan Liu, Pan Lu, Wenbo Lv, Yicheng Lyu, Qiuyang Mang, Kyle Montgomery, Yuzhou Nie, Ruoxi Ning, Jorin Overwiening, Xu Pan, Layna Paraboschi, Core Francisco Park, Justin Purnomo, Swati Rajwal, Scott Rankin, Bixuan Ren, Yiren Rong, HaoYang Shang, Ventus Shaw, Fiona Shen, Jiawei Shen, Minqi Shi, Shi Qiu, Huaxiu Yao, Tianneng Shi, Jonah So, Vladislav Susoy, Hannah Szlyk, Haocheng Wang, Jialu Wang, Wei Wang, Xinyu Wang, Zehao Wang, Dowling Wong, Angela Wu, Dehao Wu, Fangyu Wu, Mengyuan "Millie" Wu, Yu Wu, Yuchen Wu, Yuhao Wu, Qingpo Wuwu, Weihang Xiao, Yongyi Xiong, Fan Xu, Ruiling Xu, Mingxuan Yan, Benjamin Yang, Jirong Yang, Sen Yang, Xiaoli Yang, Yushi Yang, Haoran Ye, Xiaohu Yu, Zhengming Yu, Chenlong Zhang, Chi Zhang, Hanning Zhang, Hanwen Zhang, Junge Zhang, Kunpeng Zhang, Song Zhang, Wenjin Zhang, Wenshuo Zhang, Ying Zhang, Yizhi Zhang, Brian Zhao, Qijian Zhao, Yimin Zhao, Yuhaohua Zheng, Liwei Zhou, Tianyue Zhou, Sichen Zhu, Siqi Zhu, Yan Zhu, Yishu Zhu, Jierui Zuo, Chonghao Cai, Helena Casademunt, Wenjia Chen, Cheng Cheng, Nawen Deng, Rao Fu, Tianfu Fu, Yifan Han, He Ren, Zhenyu He, Qiao Jin, Langlang Li, Yuetai Li, Sylvia Liu, Lu Lu, Luqing Zhou, Subhabrata Mukherjee, Yunqi Ouyang, Yin Ren, Dawei Shi, Haoran Wu, Zhiyue Wu, Hannah Yao, Zhuoran Yi, Jenny Yu, Rhea Zhan, Hang Zhou, Blake Zhu, Junfan Zhu, Alan Yuille, Yang Liu, Russell Alan Poldrack, Jiachen Li, Zhenglu Li, Molei Tao, Jing Huang, Wenqi Shi, Costas Spanos, Lichao Sun, Chenguang Wang, Orson Xu, Zhen Dong, Hector Gomez, Aylin Caliskan, Ali Emami, Haimin Hu, Zhi Li, Lihui Liu, Murphy Niu, Yi Shao, Jianxin Sun, Mikko Tolonen, Ting Wang, Sanjiv Das, Yanjun Gao, Wenbo Guo, Erika J Schneider, Zhiyong Lu, Yian Ma, Mark Mueller, Radha Poovendran, Somayeh Sojoudi, Yinglun Zhu, Dawn Song

arXiv:2606.05405v2cs.AIcs.CLcs.LG

TL;DR

Existing benchmarks measure abstract competence but not sustained performance on economically valuable professional workflows. ALE addresses this gap with expert-authored, long-horizon task workflows, finding that frontier agents currently clear only a small fraction of them.

  • Problem

    Widely used benchmarks underdevelop sustained evaluation of long-horizon, economically valuable work in real professional environments.

  • Method

    ALE evaluates agents on 1,490 expert-authored task instances across 55 digital industries using deterministic checks and structured rubrics.

  • Results

    Frontier agents currently clear only a small fraction of ALE’s long-horizon, tool-intensive professional workflows.

  • Takeaways & Limitations

    ALE is intended as an instrument for closing the gap between benchmark success and GDP-relevant impact.

Abstract

from arXiv · show

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long horizon, economically valuable, real world tasks with verifiable outcomes. Developed in collaboration with 250+ industry experts, ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy). It is organized around a task taxonomy with 55 sub fields grouped into 13 industry clusters covering 1K+ tasks. Current results show that the hardest tier remains far from saturated: across mainstream harness and backbone configurations, the average full pass rate is below 1%. ALE is designed as a living benchmark: its task pool grows continuously as new workflows and industries are onboarded. More broadly, ALE is intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP relevant impact.

1 Introduction

ALE addresses the gap between benchmark achievements and muted economic impact by evaluating agents on authentic, long-horizon, economically valuable professional workflows. It spans broad industry coverage with expert-built tasks and uses verifiable deliverable- or milestone-based checks to assess generalist computer-use agents.

  • Motivation: Benchmark victories have accumulated faster than measurable transformation in core industries, exposing a utility problem for AI evaluation.The introduction frames economic output as the metric that ultimately matters.
  • Motivation: Benchmarks shape research attention, engineering targets, and which domains become tractable for rapid improvement.The paper argues that verifiable, widely used evaluations can accelerate progress and support subsequent deployment.
  • Benchmark design: Authentic long-horizon workflows are expensive to collect, while broad coverage of economically valuable industry workflows is structurally difficult.These constraints distinguish ALE’s construction challenge from shorter computer-use tasks, synthetic environments, and question-answering setups.
  • Benchmark rationale: “Last” denotes both a competence threshold for sustained professional work and a difficulty frontier at the boundary of current reliable capability.The name reflects ALE’s aspiration that passing an industry exam demonstrates readiness beyond answering questions.
  • Benchmark design: ALE covers 1K+ task instances across 55 subfields and 13 industry clusters, developed with 250+ domain experts.Its industry organization is anchored in the O*NET / SOC 2018 occupational taxonomy and focuses on economically meaningful workflow families from professional practice.
  • Verification: ALE standardizes heterogeneous-output evaluation through structured deliverable- or milestone-based checks against expert-provided references and rubrics.This design aims to make real-world outputs verifiable without human judges.
  • Evaluation target: ALE evaluates Generalist Computer-Use Agents that combine visual perception, code execution, tool use, and long-horizon planning in one action loop.Its task surface is constructed as a superset of GUI-only and CLI-only benchmarks, including OSWorld and Terminal-Bench.
  • Purpose: ALE is intended as an instrument for closing the gap between benchmark success and GDP-relevant impact, rather than merely another leaderboard.The paper links progress on passing this exam with the possibility of real economic transformation.

2 Benchmark Design and Dataset Construction

ALE admits professional workflows that are representative, complex, and verifiable, organizes them into an O*NET/SOC 2018-grounded taxonomy, and constructs them through staged expert review. To limit contamination, it releases only a small public subset while maintaining a rotating private pool.

  • Design requirements: ALE requires workflows to reflect real professional practice and use the software domain experts would actually use.For example, architectural workflows would use SolidWorks or Rhino rather than AutoCAD for converting a 2D blueprint into a 3D model.
  • Design requirements: Tasks must be end-to-end deliverables requiring substantial expert time, rather than isolated UI actions.A suitable video-editing workflow may combine tracking, rotoscoping, compositing, and color matching.
  • Design requirements: Outputs must support deterministic checking or an unambiguous rubric tied to observable artifacts.Game reproduction can be checked by comparing map geometry, character attributes, and event states against a reference under identical user-operation trajectories.
  • Taxonomy: 13 domains and 55 subdomains are derived by clustering software-mediated occupations from SOC 2018 and O*NET while excluding sectors whose core work is not meaningfully digital.The taxonomy is intended to avoid ad hoc or purely economic-ranking-based industry selection.
  • Task construction: Tasks are sourced from domain professionals and screened through a five-gate construction protocol covering expert sourcing, submission, first-pass review, engineering implementation, and final quality control.The protocol is designed to ensure authenticity, complexity, and technical executability.

3 Evaluation Pipeline

ALE’s evaluation pipeline decouples task specifications, agents, and environments through defined interfaces, then executes and scores runnable benchmark instances. It supports broad agent capabilities and heterogeneous deliverables using task-specific, preferably deterministic evaluation methods.

  • Pipeline architecture: The pipeline uncouples the task specification, agent, and environment so each component can be interchanged through defined interfaces.A benchmark instance is executed through coordinated task, agent, and environment components.
  • Pipeline architecture: Each task specification is an executable main.py exposing load(), start(), and evaluate() across a remote virtual-machine environment.The specification encodes the description, input assets, target software, reference assets, and evaluation criteria.
  • Agent capabilities: ALE targets agents that combine GUI reading, dialog interaction, shell commands, coding, API use, and long-running session management.No single existing agent family natively covers this operational surface.
  • Agent capabilities: The agent capability taxonomy has five layers: Brain, Eyes, Body, Hands, and Feet; generalist CUA-agents cover all five, while CLI- and GUI-agents have narrower coverage.Brain denotes reasoning and planning; Eyes, Body, Hands, and Feet cover perception, orchestration, tool invocation, and runtime execution.
  • Evaluation and scoring: ALE evaluates heterogeneous deliverables along two orthogonal axes rather than imposing one scoring metric, using artifact-specific comparison forms and scoring approaches.Examples include CAM toolpaths, financial workbooks, 3D meshes, game states, screenshots, structured filings, and free-text reports.

4 Experiment

The experiment evaluates agents on authentic professional workflows in a unified GCUA configuration combining shell, GUI, file, and web operations. Results show uneven domain performance, limited GUI use, knowledge-dominated failures, and foundation-model effects larger than harness effects.

  • Evaluation setup: Authentic workflows interleave shell commands, GUI applications, file manipulation, and web research, so all systems are evaluated as Generalist CUA-agents across five functional layers.Each agent is brought into GCUA configuration through GUI-as-Tool mode.
  • Evaluation setup: 3.8% overall timeout rate is observed under a five-hour cap, ranging from ∼1% for lightweight harnesses to 5.7% for OpenClaw.Results report mean scores, full pass rates, costs, and wall-clock times across grouped harness–backbone configurations.
  • Domain-level performance: ∼55–85% scores in computational mathematics and agriculture/environment exceed ∼50–55% in business and legal, while education remains below 25%.Claude Fable 5 and GPT-5.5 exhibit similar domain profiles when averaged over harnesses.
  • Tool usage: 34% of public task instances designate graphical software as the primary tool, yet GUI usage remains small as agents execute GUI tasks through Bash/CLI substitutes.Both harness and model shape the tool-call mix.
  • Failure taxonomy: Roughly three quarters of failed Claude Code + Opus 4.7 tasks are Understanding or Approach failures, indicating domain knowledge as the dominant bottleneck.Agents lacking specialized knowledge often use ad-hoc scripts instead of intended domain software, reinforcing GUI underuse. Foundation-model choice accounts for roughly 3× the performance spread of harness choice among well-engineered systems.

5 Related Work

ALE is positioned against prominent benchmarks using industry coverage mapped to its 55-industry taxonomy. Existing knowledge and agentic benchmarks differ in focus, with the latter adding interaction and tool use but remaining software-centric.

  • Benchmark positioning: Table 2 compares ALE with prominent benchmarks using industry coverage mapped onto its 55-industry taxonomy.The “Breadth” column reports coverage across the taxonomy’s 55 industries.
  • Knowledge benchmarks: Knowledge and exam-style benchmarks such as MMLU, GPQA, and HLE test what models know rather than what they can do.These benchmarks are described as topically broad but focused on knowledge and exam-style evaluation.
  • Agentic benchmarks: Agentic benchmarks including SWE-bench, OSWorld, WebArena, and GAIA add multi-step interaction and tool use but cover only a handful of software-centric domains.Their coverage is contrasted with broader industry mapping through ALE’s taxonomy.

6 Conclusion … B.1.6 Robotics & Autonomous Systems

ALE is introduced as a benchmark for economically valuable, expert-authored digital workflows, with a taxonomy grounded in SOC/O*NET and extended to frontier work. Its construction spans 55 workflow-level subdomains across 13 domains, while example industries emphasize artifact-producing workflows and handoff consistency.

  • 6 Conclusion: ALE contains 960 expert-authored task workflows and 1,490 task instances across 55 digital industries, scored with deterministic checks and structured rubrics.Frontier agents currently clear only a small fraction of tasks; ALE is intended to help close the gap between benchmark performance and GDP-relevant impact.
  • B.1.1 Taxonomy Definition: ALE evaluates generalist computer-use agents on professional workflows whose computer-produced outputs depend on domain expertise and are objectively evaluable.The scope includes digital interfaces, software tools, files, and APIs.
  • B.1.1 Taxonomy Definition: SOC 2018 and O*NET anchor ALE’s occupational taxonomy, while screening and consolidation organize records into workflow-level subdomains.The screening applies to 1,016 O*NET 30.2 entries and leaves 117 unique SOC base codes after variant consolidation.
  • B.1.1 Taxonomy Definition: 55 workflow-level subdomains are grouped into 13 domains, including four frontier subdomains and seven extensions to SOC-anchored subdomains.The frontier supplement covers emerging digital workflows absent or under-specified in SOC 2018 but present in current research and professional practice.
  • B.1.2 Industry Landscape Review: Industry landscape records describe work settings, roles, inputs, dependencies, and output artifacts, using references, workflow documentation, LLM-assisted research, and expert review.These landscapes organize candidate task families and check that collected tasks specify real workflows.
  • B.1.3 Manufacturing & Industrial Operations: Manufacturing workflows transform CAD, bills of materials, process specifications, and production targets into routings, toolpaths, layouts, control logic, and inspection plans.Handoff consistency is tracked across CAM, controls, industrial engineering, and quality review, including controller dialects, fixture geometry, alarm priorities, and tolerance stacks.
  • B.1.4 Biomolecular Structure & Design: Biomolecular design workflows produce ranked ligand poses, structure models, sequence designs, codon-optimized constructs, plasmid maps, and design registry entries before wet-lab execution.Recorded assumptions can help identify computational layers to revisit when binding, predicted poses, or pathway titers fail.
  • B.1.5 3D, Animation & Interactive Media; B.1.6 Robotics & Autonomous Systems: 3D media and robotics workflows convert creative or operational specifications into rendered frames, runtime states, robot descriptions, controllers, planners, safety logic, and simulation assets.Both areas emphasize preserving cross-tool handoff information: media tracks scale, coordinate frames, import settings, color spaces, and timing, while robotics tracks sim-to-real parameters from digital twin to hardware.

B.2 Task Construction Pipeline Details · B.3 Industry Task Cards and Metadata

ALE converts practitioner-submitted workflows into benchmark instances through five staged gates, combining expert sourcing and editing with review, implementation, and final quality control. The process preserves source provenance and checks executable evaluation environments for reproducibility and integrity before acceptance.

  • B.2 Task Construction Pipeline Details: Five gates—expert sourcing, task submission and editing, first-pass review, task implementation, and final QC and acceptance—convert workflow submissions into benchmark instances.The appendix expands the staged construction protocol summarized in Section 2.3 and depicted in Figure 4.
  • B.2 Task Construction Pipeline Details: An advisory committee of industry professionals and expert practitioners anchors recruitment of specialists who perform complex software workflows daily.
  • B.2 Task Construction Pipeline Details: Practitioners submit past projects through a dedicated web portal, with AI-assisted tools helping refine proposals while keeping structural overhead low.Submitted projects historically took experts days or weeks to complete.
  • B.2 Task Construction Pipeline Details: Task specifications define the natural-language description, input files, target software and tools, expected deliverable, and evaluation specification.Source-grounded tasks may additionally retain papers, datasets, standards, workflow documentation, label sets, and reference artifacts as provenance records.
  • B.2 Task Construction Pipeline Details: First-pass review uses conference-style decisions—major or minor revision, borderline accept, accept, and strong accept—and sends revision cases back for further editing.
  • B.2 Task Construction Pipeline Details: Engineering teams translate written specifications into runnable assets, software containers, and evaluation logic, then conduct engineer review and dry-run execution.Detected logic gaps or missing dependencies trigger automatic email notification.
  • B.2 Task Construction Pipeline Details: Final peer review checks reference-output correctness, calibrated evaluation bounds, reproducibility, and evaluation integrity before benchmark admission.The final quality-control gate is overseen by the expert committee and also assesses whether the problem context provides sufficient information.

B.3.1 Representative Task Cards

Representative task cards make ALE’s benchmark construction and evaluation protocol concrete at the individual task-instance level. Each card documents the request, materials, deliverables, rubric, score, and outcome for an executed instance.

  • B.3.1 Representative Task Cards: Representative task cards illustrate how ALE task instances are specified, executed, and scored.They make the benchmark construction and evaluation protocol concrete at the instance level.
  • B.3.1 Representative Task Cards: Each card summarizes the agent-facing request, input materials, expected deliverables, evaluation rubric, observed score, and observed outcome.
  • B.3.1 Representative Task Cards: All shown trajectories were produced by the Claude Code harness running Claude Opus 4.7.

Task Card 1: Injection Mold-Flow Analysis PARTIAL

This partial Moldex3D mold-flow task required configuring and running a four-cavity injection-mold simulation, then extracting solver-derived metrics into results.json. The agent achieved the setup workflow but earned only partial credit because numerical outputs were not reliably verified.

  • Task requirements: The agent had to apply process_spec.json to project 230057, run fill, pack, cooling, and warp analyses, and write solver-derived values to output/results.json.Required fields included pressure, force, time, volume, and weight metrics.
  • Evaluation: Scoring compared each numeric results.json field with a hidden solver reference within 1 percent relative tolerance, while rejecting missing, malformed, or all-null files.Generated project artifacts could provide limited partial credit when main metrics were weak.
  • Result: 0.4762 score: the agent reached the correct Moldex3D setup workflow and submitted results.json, but received only partial numeric/result credit.The observed score was 0.476, while the detailed score was 0.4762.
  • Failure analysis: The failure arose because submitted metrics were not reliably extracted from completed solver outputs, with some values estimated rather than measured CAE results.The task therefore required closed-loop numerical verification: waiting for result files, inspecting or exporting solver data, and populating the JSON.

Task Card 2: Orchestral Music Transcription FAIL

The orchestral transcription task received a score of 0.0 because the submission did not confirm the complete required PDF, MIDI, and notation screenshot bundle. Although the agent exported transcription.mid from a matching Dorico project, the hard gates prevented musical accuracy from being evaluated.

  • Task requirements: The task required converting an audio recording into a readable full score and multitrack MIDI under a specified Dorico, tempo, instrumentation, and output-file contract.Required outputs were transcription.pdf, transcription.mid, and overview.png.
  • Scoring: After the gates pass, pitch and rhythm each contribute 30 percent, dynamics 20 percent, instrument assignment 10 percent, and score layout 10 percent.The MIDI is compared with a reference MIDI after track pairing and quantization.
  • Result: 0.0 score: the agent exported transcription.mid, but transcription.pdf and overview.png were not confirmed, triggering the complete-bundle hard gate.Missing PDF, MIDI, or invalid notation screenshot produces a hard zero before musical accuracy is evaluated.
  • Interpretation: The failure demonstrates that partial tool use did not produce reliable multi-artifact delivery for a creative-production task.The submission contained part of the required package, but musical accuracy was not scored.

Task Card 3: Reference-Based Chroma Key Compositing FAIL … C Evaluation Pipeline Details

The evaluated tasks test reference-grounded visual production, motion reproduction, and radiology adjudication, with outcomes showing failures in task fidelity despite artifact or file-level compliance. Their evaluation pipelines combine validity gates with reference-based visual, replay, or annotation checks, while the executed-task inventory is available online.

  • Task Card 3: Reference-Based Chroma Key Compositing FAIL: The chroma-key task required removing the green screen, preserving foreground edges, matching a reference composition, and exporting output/output.mp4.Inputs included a bird foreground clip, a reference image, and a DaVinci Resolve environment.
  • C Evaluation Pipeline Details: Evaluation pipelines combined hard validity gates with hidden-reference checks, including visual keying confirmation, replay consistency, motion matching, exact case-level agreement, and box IoU at least 0.50.These checks distinguish file or artifact presence from substantive task success.
  • Task Card 3: Reference-Based Chroma Key Compositing FAIL: 0.0 score: the chroma-key submission exported an MP4 but failed to preserve the intended foreground/background relationship from the reference materials.The failure reflected visual task grounding rather than software operation alone.
  • Task Card 4: Skeletal Animation Reproduction PARTIAL: The animation evaluation required a rigged character, 19 exact bone names, an editable final.blend, and a replayable preview.mp4 matching reference motion.Success depended on timing, pose range, replay consistency, and non-trivial animation rather than a single valid-looking frame.
  • Task Card 4: Skeletal Animation Reproduction PARTIAL: 0.4285 score: the skeletal-animation run earned partial credit because both required artifacts existed and contained animation, but motion fidelity remained incomplete.The preview was about 4 seconds versus a reference of about 13.6 seconds, and motion was damped to reduce mesh distortion.
  • Task Card 5: MicroDicom Chest X-Ray Adjudication PARTIAL: The radiology task required selecting the better atelectasis annotation for each case and submitting adjudicated_boxes.tsv, adjudication_log.tsv, and final_impressions.tsv.The inputs comprised nine NIH CXR DICOM cases, competing reader boxes, reader notes, and fixed TSV schemas.
  • Task Card 5: MicroDicom Chest X-Ray Adjudication PARTIAL: 0.3333 score: the radiology run produced all three required TSV files for nine cases but achieved only partial agreement with hidden references.Later decisions relied on coordinates and heuristics rather than systematic visual adjudication.
  • B.3.2 Executed Task Inventory: The full inventory of 150 selected public tasks, including names, identifiers, domains, and descriptions, is browsable online under the Benchmark Splits (ALE-V1, 2026/06) tab.The listed location is https://agenthle.org/demo.

C.1 Pipeline Architecture Details … C.3.4 Judge Types and the LLM-Judge Helper Layer

ALE uses a decoupled, reproducible pipeline in which executable task specifications, agents, and controlled environments interact through standardized interfaces. Evaluation supports multiple artifact and scoring modes, favoring deterministic verification and constraining unavoidable LLM judging to targeted, evidence-anchored probes.

  • C.2 Task Specification Protocol: A task specification encodes its description, input assets, required software, reference assets, and evaluation criteria in main.py through load(), start(), and evaluate() lifecycle functions.load() declares the task, start() prepares the environment, and evaluate() scores the agent’s outputs against references or rubric criteria on a normalized [0, 1] scale.
  • C.1 Pipeline Architecture Details: The agent repeatedly observes the environment, selects and executes actions, and terminates when it decides the task is complete.Actions include mouse clicks, keystrokes, shell commands, file edits, and API calls; the environment partitions inputs, software, outputs, and references in a remote virtual machine.
  • C.1 Pipeline Architecture Details: ALE separates each evaluation into a task specification, an agent, and an environment, allowing agents and task specifications to be reused across compatible interfaces and backends.The action interface supports shell commands, GUI interactions, and file I/O, while task specifications can run on cloud VMs or local containers without modification.
  • C.3.1 Execution Locale: Scoring runs on the host by default when artifacts and tooling can be transferred, while VM-side verifiers handle artifacts requiring specialized software and return scores through stdout JSON.Host-side scoring is preferred because it is easier to review, version-control, and rerun offline; VM-side verifiers never write into output/.
  • C.3.2 Artifact Modes: ALE supports exact or hashed, structured tabular or numeric, geometric or spatial, visual, behavioral, free-text, and executable artifact modes.These modes compare normalized strings or hashes, tolerant fields, spatial distances, images, replayed system states, rubric criteria, or held-out execution outcomes.
  • C.3.4 Judge Types and the LLM-Judge Helper Layer: ALE rejects LLM judges when deliverables can be evaluated through bytes, fields, geometry, world state, or executable behavior, reserving them mainly for creative or perceptual outputs without objective code-based references.The helper layer provides targeted yes/no, binary checklist, graded, and JSON-structured visual judging functions.
  • C.3.4 Judge Types and the LLM-Judge Helper Layer: LLM-judged workflows use narrow, evidence-anchored probes targeting identifiable artifacts, with gates that can force a zero before later quality checks.Recurring probes assess items such as a game-map feature, DAW tracks, UV seams, silhouettes, or animation poses rather than asking whether a deliverable is broadly good.

C.3.5 Empirical Distribution of Judge Types … D.1 Public-Subset Representativeness

ALE’s reference workflows predominantly use deterministic, host-side scoring, with structured safeguards for missing outputs and references. Its harnesses support long-horizon tool use, delegation, context management, and GUI interaction, while the public subset closely reflects the full benchmark’s cluster-level performance.

  • C.3.5 Empirical Distribution of Judge Types: 93.2% of task workflows use code-based deterministic judges, versus 6.8% using LLM-as-judge evaluation.These proportions come from static analysis of open-sourced main.py files and their scoring scripts.
  • C.3.5 Empirical Distribution of Judge Types: 88.5% of scoring code runs host-side, while 11.5% uses VM-side verifiers for deliverables requiring specialized software stacks.VM-side workflows include CAD/CAM, licensed financial workbooks, and headless 3D rendering; their verifier JSON output is parsed on the host.
  • C.3.5 Empirical Distribution of Judge Types: Most evaluate() bodies use early [0.0] returns before a continuous-score path, while explicit weighted-rubric expressions appear in only a small fraction of main.py files.In most workflows, weighting is encoded inside per-task-workflow scripts rather than directly in main.py.
  • C.3.6 Reference Isolation and Robustness: Reference isolation and output-shape checks make missing references, absent outputs, and malformed outputs yield score 0 rather than inflated scores or crashes.Code-based judges are deterministic, while LLM-judged results record the model and prompt so scores can be re-derived from saved artifacts.
  • C.3.7 Workflows and Task Instances: A workflow exposes multiple VARIANTS instances that share one evaluate() function, with instance scores averaged into the workflow score.The manufacturing/gcode example contains 18 workpiece instances scored by the same collision-gate-then-STL pipeline.
  • C.4 Agent Harness Internals: The harness uses a six-phase loop covering initialization, context building, LLM calls, action routing, tool-result collection, and overflow checking.The loop repeats through context building until the model delivers, triggering compaction when accumulated context exceeds a threshold.
  • C.4.1 Tool Surface and Terminology: The tool surface includes file operations, shell execution, web retrieval, sub-agent management, and 14 GUI desktop-action tools exposed through a CUA MCP bridge.Sub-agents can operate in isolated contexts, while the context manager uses microcompaction, LLM summarization, and truncation for long-horizon tasks.
  • C.4.1 Tool Surface and Terminology: ALE-Claw ports OpenClaw’s agent loop to Python, enabling component ablations and a shared scaffold for comparing GUI models on the same task suite.The port adds CUA-specific adaptations while removing production assistant components unnecessary for single benchmark tasks.

D.2 Timeout Analysis … D.6 Per-Task Instance Score Heatmaps

The appendix analyzes timeout effects, classifies failure causes, compares model and harness contributions, evaluates resource efficiency, and visualizes per-task performance across tiers. Results show that foundation-model choice dominates harness choice, while cost, time, and token use are only loosely related to performance.

  • D.2 Timeout Analysis: 3.8% of evaluated runs reached the five-hour cap, with mean scores of 20.7 for capped runs versus 33.2 for earlier-ending runs.The harness stops the agent at the cap and scores artifacts remaining in the output directory.
  • D.3.1 Stage 1: Trajectory analysis: The failure analysis uses a two-stage pipeline: Codex generates structured cards from run artifacts, then GPT-4o classifies them with temperature 0.Cards include verdicts, task descriptions, correct and incorrect behaviors, and artifact-specific evidence; full transcripts are excluded to bound generation cost.
  • D.3.2 Stage 2: Taxonomy classification: The taxonomy organizes failures into Understanding, Approach, Execution, and Infrastructure categories, with subtypes covering knowledge gaps, strategy, implementation, formatting, GUI, and resource failures.Each analysis card is classified using the hierarchy defined in the classification prompt.
  • D.3.3 Distribution: 47% of classifiable failures are Approach errors, compared with 31% Understanding and 22% Execution errors; timeout and resource-exhaustion cases are excluded.Approach failures comprise wrong strategy (30%) and premature abandonment (17%), while Understanding includes domain knowledge gaps (25%) and hallucination or fabrication (6%).
  • D.4 Model vs. Harness Effect: 16.8 percentage points is the overall pass-rate spread from changing backbones under fixed OpenClaw, versus a 4.9–7.2 pp spread from changing harnesses under fixed backbones.The fixed-harness range runs from 4.3% for Grok 4.3 to 21.1% for GPT-5.5, indicating a larger backbone effect among evaluated configurations.
  • D.5 Cost, Time, and Token Efficiency: 45.8% is the highest overall mean score at $326 total API cost, while a 3.6× costlier configuration scores 5.3 percentage points lower; time and token use likewise do not predict performance.The fastest configuration takes 24 hours and scores 25.7%, whereas another takes 466 hours for 37.2%; 1 373M tokens scores below 460M tokens at 40.5% versus 41.8%.
  • D.6 Per-Task Instance Score Heatmaps: Figures 14–16 report mean scores for every task instance, sorting rows by cross-system average and columns by within-harness average, with gray cells marking missing runs.Task labels are colored by taxonomy domain; the heatmaps cover Near-Term, Full-Spectrum, and Last-Exam tiers.
  • D.6 Per-Task Instance Score Heatmaps: The three heatmaps cover 67 Near-Term, 55 Full-Spectrum, and 38 Last-Exam task instances.These counts correspond to Figures 14, 15, and 16 respectively.
Loading 2606.05405v2…