Source-linked AI summary

Remote Labor Index: Measuring AI Automation of Remote Work

Mantas Mazeika, Alice Gatti, Cristina Menghini, Udari Madhushani Sehwag, Shivam Singhal, Yury Orlovskiy, Steven Basart, Manasi Sharma, Denis Peskoff, Elaine Lau, Jaehyuk Lim, Lachlan Carroll, Alice Blair, Vinaya Sivakumar, Sumana Basu, Brad Kenstler, Yuntao Ma, Julian Michael, Xiaoke Li, Oliver Ingebretsen, Aditya Mehta, Jean Mottola, John Teichmann, Kevin Yu, Zaina Shaik, Adam Khoja, Richard Ren, Jason Hausenloy, Long Phan, Ye Htet, Ankit Aich, Tahseen Rabbani, Vivswan Shah, Andriy Novykov, Felix Binder, Kirill Chugunov, Luis Ramirez, Matias Geralnik, Hernán Mesura, Dean Lee, Ed-Yeremai Hernandez Cardona, Annette Diamond, Summer Yue, Alexandr Wang, Bing Liu, Ernesto Hernandez, Dan Hendrycks

arXiv:2510.26787v1cs.LGcs.AIcs.CL

TL;DR

Existing benchmarks do not clearly show how AI progress translates into economically valuable work across the diverse remote labor economy. The paper introduces RLI, a multi-sector benchmark of real end-to-end freelance projects, and finds that frontier agents achieve near-floor performance, with a highest automation rate of 2.5%. RLI provides an empirical basis for monitoring AI automation, while its coverage excludes several types of remote work.

  • Problem

    It remains unclear how progress on knowledge, reasoning, and agent benchmarks translates into the ability to perform diverse, economically valuable remote work.

  • Method

    RLI evaluates AI agents on complete projects sourced from freelance platforms, spanning diverse remote-work categories and comparing outputs with human gold-standard deliverables.

  • Results

    2.5% is the highest automation rate achieved by evaluated agents on RLI, indicating near-floor performance across the benchmark.

  • Takeaways & Limitations

    RLI provides an economically grounded empirical basis for monitoring AI capabilities and navigating potential labor-market impacts.

  • Takeaways & Limitations

    RLI excludes remote work requiring client interaction, teamwork, or other unmet project requirements, so complete automation on RLI would not establish automation across all remote work.

Abstract

from arXiv · show

AIs have made rapid progress on research-oriented benchmarks of knowledge and reasoning, but it remains unclear how these gains translate into economic value and automation. To measure this, we introduce the Remote Labor Index (RLI), a broadly multi-sector benchmark comprising real-world, economically valuable projects designed to evaluate end-to-end agent performance in practical settings. AI agents perform near the floor on RLI, with the highest-performing agent achieving an automation rate of 2.5%. These results help ground discussions of AI automation in empirical evidence, setting a common basis for tracking AI impacts and enabling stakeholders to proactively navigate AI-driven labor automation.

1 Introduction

The paper introduces RLI as a standardized benchmark grounded in diverse, economically valuable remote-work projects. Frontier agents perform near the floor, with the best achieving only 2.5% automation, while RLI provides evidence for tracking AI labor automation.

  • Motivation: Existing benchmarks provide limited insight into labor automation because they emphasize specialized skills or simplified tasks rather than remote work’s diversity and complexity.The paper identifies a lack of standardized empirical metrics grounded in real economic activity.
  • Contribution: RLI evaluates AI agents on complete, economically valuable projects sourced from real freelance work and compared against human deliverables.Its projects span diverse remote-work domains and preserve the original brief, inputs, and gold-standard output.
  • Results: 2.5% is the highest automation rate achieved by current AI agents on RLI, with most projects remaining beyond their capabilities.The evaluation uses manual comparison against human gold standards, and agents fail to complete most projects at an acceptable commissioned-work level.
  • Contribution: RLI aims to ground discussions of AI automation in empirical evidence and provide a common basis for stakeholders to monitor and navigate its effects.

2 Related Work

Prior benchmarks increasingly evaluate economically valuable work but often remain domain-specific or simplified. RLI extends this literature by measuring end-to-end automation on diverse projects drawn from real remote labor markets.

  • Prior benchmarks: Existing work has expanded from closed-ended academic knowledge to agentic tasks, but many benchmarks still represent only a small fraction of remote work.
  • Prior benchmarks: Domain-specific benchmarks measure areas such as software engineering and ML engineering, while some broader task benchmarks approach human parity on shared professional tasks.
  • RLI’s distinction: RLI differs from prior benchmarks by evaluating end-to-end projects sourced from real remote-labor transactions.It targets automation capacity across diverse specializations rather than general cognitive ability or isolated skills.

3 Remote Labor Index

The Remote Labor Index is a 240-project benchmark of end-to-end freelance work, designed to reflect the diversity, complexity, and economic value of real remote labor. It uses professionally produced deliverables, extensive dataset cleaning, and manual evaluation against human standards.

  • Dataset Description: RLI contains 240 end-to-end remote freelance projects sourced directly from professionals on freelance platforms.The dataset is intended to ground evaluation in practical, economically valuable work.
  • Dataset Description: Each project includes a brief, input files, and a professional human deliverable that serves as the gold standard.The brief and inputs come from the professional who produced the deliverable, supporting self-contained task completion.
  • Dataset Description: RLI spans 23 work categories and includes design, operations, marketing, administration, data/BI, audio–video production, and other remote-labor domains.This broader coverage contrasts with benchmarks concentrated on software engineering, research, and writing.
  • Dataset Description: RLI projects require substantially more effort than previous benchmarks, averaging 28.9 hours to complete with a median cost of $200.Human completion times exceed previous benchmarks by more than 2×, while average project cost is $632.6.
  • Metrics: Automation rate measures the percentage of projects whose AI deliverable completes the work at least as well as the human deliverable.Evaluations are manual and reliable, with 94.4% inter-annotator agreement for automation rate.

4 Experiments

RLI evaluates frontier AI agents on diverse, economically valuable projects using absolute and relative performance measures, complemented by qualitative failure analysis. Agents achieve near-floor automation, while Elo scores detect steady relative progress and failures commonly involve unusable, incomplete, inconsistent, or low-quality deliverables.

  • 4.2 Quantitative Results: 2.5% is the highest Automation Rate achieved, showing that current agents complete very few RLI projects at human-accepted quality.Automation Rate measures projects completed at or above the human gold standard; Manus achieved the highest rate.
  • 4.2 Quantitative Results: Economic-impact metrics, including Dollars Earned and Autoflation, also remain close to the floor.The results indicate that most systems do not complete projects at a realistic commissioned-work standard.
  • 4.2 Quantitative Results: Elo rankings show steady improvement between models even when agents fail to complete most projects.Pairwise comparisons capture relative progress through deliverable quality and closeness to successful completion.
  • 4.3 Qualitative Findings: Across roughly 400 evaluations, failures predominantly involve technical or file-integrity problems, incomplete deliverables, poor quality, and inconsistencies.Examples include corrupt files, truncated videos, missing assets, unprofessional outputs, and mismatched deliverable files.
  • 4.3 Qualitative Findings: Failure analysis includes concrete mismatches such as 8-second videos requested as 8 minutes and web games whose graphics fall below professional standards.Evaluators assigned one or more categories to each deliverable; categories were not mutually exclusive.
  • 4.3 Qualitative Findings: A small subset of creative, audio, image, writing, retrieval, and visualization projects matched or exceeded human baselines.Successful examples include audio editing, ad and logo creation, report writing, and interactive data visualization code.

5 Discussion

The discussion contrasts AI's potentially general cognitive capabilities with historically task-specific automation and notes that RLI does not cover all remote work. It also cautions that reported project costs are historical and unadjusted for inflation.

  • 5 Discussion: Simple web visualizations that only require code are within current agents' capabilities, but they represent a small share of remote labor.The discussion uses a successful Sonnet 4.5 project as an example.
  • 5 Discussion: Unlike task-specific automation, AI is being developed to automate human intelligence and may generalize to new jobs.The paper argues that such generalization would require many of the general cognitive skills humans use across tasks.
  • 5 Discussion: RLI excludes client-interaction, teamwork, and other remote-work projects, so a 100% RLI automation rate would not establish human-level performance everywhere.The excluded categories include tutoring and project management.
  • 5 Discussion: Reported project costs reflect completion-time prices and are not inflation-adjusted, likely understating current economic value.Most projects with known dates were completed within the past five years.

6 Conclusion

RLI provides an economically grounded benchmark spanning 240 projects across 23 digital-freelance domains. Frontier agents remain near the automation floor, while the benchmark is intended to support empirical monitoring of AI's labor-market implications.

  • 6 Conclusion: RLI spans 240 projects across 23 digital-freelance domains, grounding automation measurement in demonstrated market value.The benchmark is designed to reflect economically valuable work rather than only isolated capabilities.
  • 6 Conclusion: Frontier AI agents achieve an automation rate below 3% on RLI, revealing a large gap between computer-use evaluations and economically valuable work.The conclusion characterizes frontier performance as near the floor.
  • 6 Conclusion: RLI aims to give stakeholders an empirical foundation for monitoring capabilities, forecasting labor-market impacts, and navigating AI-driven automation.The stated audience includes researchers, policymakers, and the public.

A.1 Full Results

The full results report model-specific Elo scores, automation rates, and earnings, alongside the RLI project-cost reduction measured as autoflation.

  • Table 3 reports the precise Elo scores and automation rates for every evaluated model, including GPT-5 CLI and CUA scaffolds.
  • Table 4 reports dollars earned for all evaluated models, which represent a small fraction of the dataset projects’ total costs.
  • Autoflation measures the reduction in completing the fixed RLI project bundle when AI achieves acceptable deliverables at lower effective cost than humans.
  • For each project, cost reduction is measured against the human deliverable using the lowest-cost acceptable method, and is zero when no AI method is cheaper.

A.3 Effect of Agent Scaffolds

Agent scaffolds materially affect measured performance, while autoflation tracks reductions in the effective cost of the fixed RLI project bundle.

  • GPT-5’s CLI scaffold outperforms its CUA setup on both Elo score and automation rate: 436.7 versus 431.6, and 1.7% versus 0.8%.
  • Autoflation is the percentage decrease in the cost of completing the fixed RLI project bundle with AI agents at lower effective cost than humans.

B Evaluation Details

The evaluation combines sampled pairwise preferences, majority voting, Bradley-Terry utilities, normalized Elo scores, and time-bounded review procedures.

  • Model pairs and projects are randomly sampled, with stratification ensuring each model pair is compared on at least 10 projects, with a median of 25.
  • Two independent evaluations and a tie-breaking third produce majority-voted project preferences, with indifference coded as 50/50.
  • Global Bradley-Terry fitting converts sampled preference edges into Elo scores, with 100 bootstrap samples used for 95% confidence intervals.
  • Utilities are scaled so the human baseline scores 1,000 and a 400-point difference corresponds to 10 : 1 odds of winning.
  • Evaluators received soft limits of 20 minutes for model-versus-human comparisons and 30 minutes for model-versus-model comparisons.

B.3 Evaluation and Generation Budgets

Evaluation uses trained annotators, human reference deliverables, explicit quality criteria, and capped review times for automation-rate and pairwise assessments.

  • Automation-rate evaluations allow 20 minutes per project, while Elo evaluations allow 30 minutes because they require inspecting two AI deliverables.
  • Annotators judge deliverables holistically from a reasonable-client perspective using the human deliverable as the reference for acceptable completion.
  • The human reference defines a zone of acceptable error, so AI work is not penalized for similar minor flaws or non-critical omissions.
  • Training highlighted common AI failures including rasterized graphics, nonsensical image text, and inconsistent spatial or visual structure across files.
  • The final evaluation instructions achieved 94.4% inter-annotator agreement after iterative refinement, auditing, and quality checks.
  • Automation rate counts projects rated 2 or 3, meaning the AI deliverable matches or exceeds the human reference and would be accepted by a reasonable client.

B.5 Evaluation verification

RLI evaluation uses standardized agent scaffolds, multimedia tools, and manual verification to assess deliverable quality and benchmark performance. The setup compares multiple environments while auditing annotation errors and project-specific evaluation conditions.

  • AI deliverables judged good or better than human work were manually audited to reduce false positives.All such cases could be audited because only a small number received that annotation.
  • Manual audits found no false negatives in 50 randomly sampled human–model comparisons, bounding the false-negative rate at ≤5.8% with 95% confidence.Two co-authors independently evaluated the sampled pairs to estimate missed cases where the human deliverable was incorrectly preferred.
  • The evaluation uses three agent environments: integrated agents, a Scale AI computer-use environment, and the CLI-based OpenHands environment.Models supporting computer use defaulted to the computer-use scaffold, while others used OpenHands.
  • GPT-5 is reported with the CLI scaffold because it outperformed its computer-use setup in the main results.The two configurations are labeled GPT-5 (CLI) and GPT-5 (CUA), with the CLI version selected for reporting.
  • Agents receive specialized multimedia tools for image, speech, and video generation alongside standardized input and output directory instructions.OpenHands was extended with tools including gpt-image-1, openai/tts-1, and veo-3.0-generate-preview.
  • The evaluation platform supports diverse source, document, spreadsheet, media, design, 3D, database, and interactive formats, with project-specific notes guiding exclusions.The platform is open-source and includes fallback rendering for unsupported file types; evaluators may be instructed to ignore features outside a project brief.

C.4 Data Collection Details

RLI projects are collected, anonymized, and documented as self-contained remote-work tasks with human cost and time measures. The dataset spans varied deliverables, including multimedia, design, software, and interactive projects.

  • The collection excludes or anonymizes sensitive material: rights-verified freelance work is included, while PII and copyrighted details may be redacted or replaced synthetically.Long-tail projects were purchased or included with permission from the original author.
  • Human project costs and completion times are self-reported, with midpoint values used when professionals provide ranges.Cost is measured in USD earned or fairly estimated for recreating the work, while completion time is measured in hours.
  • Project cost and completion time have a Pearson correlation of 0.785 on a log-log comparison.The relationship is shown for projects where both values are available.
  • The Mega Merge project requires responsive desktop and mobile gameplay, touch and mouse controls, physics-based object behavior, and a file size below 5 MB.Its core mechanic is combining matching falling objects to create higher-level items.
Loading 2510.26787v1…