Source-linked AI summary

AI Evaluation Should Work With Humans

Jan Kulveit, Gavin Leech, Tomáš Gavenčiak, Raymond Douglas

arXiv:2608.13577v1cs.AIcs.LG

TL;DR

Current AI evaluation emphasizes autonomous performance against humans, leaving human–AI team performance insufficiently measured. This position paper advocates collaborative evaluation and argues that it can reveal interaction effects and guide AI toward systems that enhance human agency and collective intelligence.

  • Problem

    AI benchmarks predominantly compare solo AI systems with solo humans, leaving the effectiveness of human–AI teams insufficiently evaluated.

  • Method

    The paper advocates evaluating human–AI collaborations, proposing metrics and research directions centered on team performance, human agency, and collective intelligence.

  • Results

    The paper argues that team-centered evaluation can capture interaction effects missed by isolated testing, including cases where standalone AI accuracy reduces team utility.

  • Takeaways & Limitations

    AI evaluation should prioritize systems that enhance human agency and collective intelligence to support more beneficial and humane AI development.

  • Takeaways & Limitations

    Human surrogate models may underrepresent cognitive biases, fatigue, emotional states, and user diversity, so they should not replace evaluation with real participants.

Abstract

from arXiv · show

This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) is guiding AI development in the wrong direction. Instead, the AI community should pivot to evaluating the performance of human--AI teams. We argue that this collaborative shift will foster AI systems that act as true complements to human capabilities and therefore lead to far better societal outcomes than will the current process.

1. Introduction

The paper argues that AI evaluation should move beyond autonomous systems surpassing human baselines, because replacement-focused metrics steer development toward substituting for human labor and cognition. It advocates evaluating how effectively humans or teams achieve goals in concert with AI, including capabilities such as shared understanding, creativity, and insight.

  • Current evaluation paradigm: ML research typically treats a problem as solved when the best autonomous system surpasses a human baseline, reinforcing evaluations that pit solo AIs against solo humans.This convention implicitly expects successful systems to work alone and replace human effort.
  • Current evaluation paradigm: Human-replacement metrics can steer AI development toward systems that substitute for human labor and cognition, risking greater economic inequality and devalued human skills.The paper also warns that this trajectory could leave humanity disempowered.
  • Current evaluation paradigm: Replacement-focused evaluation often misses capabilities important to real-world productivity, including asking questions, building shared understanding, and fostering human creativity or insight.These capabilities concern progress that collaborative systems could support but autonomous-competitive benchmarks may not capture.
  • Proposed paradigm shift: The paper advocates evaluating human–AI collaborations by asking how effectively a human or team achieves a goal with AI rather than whether AI can perform the task instead of a human.The paper will develop benefits, metrics, research directions, and responses to counter-arguments for this collaborative focus.

2. The Limits of the Replacement Paradigm

The replacement paradigm’s focus on autonomous, superhuman performance steers AI toward substituting for human labor and narrows progress criteria. It overlooks collaboration-oriented capabilities, interaction risks, and synergistic benefits that emerge in human–AI teams.

  • Replacement-oriented evaluation: Human-centered benchmarks increasingly celebrate autonomous systems that outperform people on discrete tasks, steering AI evaluation toward replacement as the primary model of progress.This paradigm has accelerated algorithmic capabilities across language understanding and complex game-playing, but treats human performance mainly as a yardstick.
  • Replacement-oriented evaluation: Benchmarking AI on tasks performed instead of humans can create an “automation race” aimed at matching human output at lower cost rather than empowering people.The paper links this direction to potential labor displacement without corresponding human-centric roles or adequate social safety nets.
  • Missing collaborative capabilities: Replacement-focused benchmarks neglect collaboration capabilities including useful explanations, uncertainty expression, adaptation to human intent, and support for human learning.Current evaluations rarely assess whether explanations help collaborators understand, trust, or debug systems, and often reward confident answers over clarifying questions or expressed uncertainty.
  • Missing collaborative capabilities: Maximizing standalone AI accuracy can decrease human–AI team utility when people cannot predict or calibrate to complex models, producing systems that are difficult to integrate into workflows.Such systems may be “brilliantly competent” in isolation but “socially inept” in settings requiring collaboration and shared understanding.
  • Interaction effects: Team-centered evaluation captures interaction-specific risks and benefits that isolated testing misses, including misuse risks, creative breakthroughs in Hanabi, and large-scale collective intelligence.Dual-use and dangerous-capability evaluations already assess human–AI teams because evaluating the AI alone would underestimate risks.

3. What Human+AI Evaluation Could Look Like

The paper proposes evaluating human–AI teams as combined systems, complementing rather than replacing standalone AI evaluation. This shift emphasizes synergy, richer collaboration outcomes, and a tiered process combining benchmarks, surrogates, and real-participant evaluations.

  • Team Evaluation: Human–AI team evaluation treats humans and AI agents as an integrated system and asks whether their collaboration achieves capabilities neither can achieve alone.The evaluation considers performance, interaction dynamics, and emergent capabilities toward a shared objective.
  • Evaluation Metrics: Team evaluation requires metrics beyond task accuracy or speed, including user satisfaction, skill enhancement, learning gains, mental-effort management, appropriate reliance, agency, and collaborative fluency.Collaborative fluency includes communication overhead, common ground, adaptability, and recovery from mistakes or misunderstandings.
  • Team Evaluation: The proposed shift targets capabilities that standalone benchmarks undervalue, including fruitful prompting, scaffolding learning, stimulating exploration, and mutual explainability.These criteria aim to guide development toward systems that are effective, trustworthy, and empowering partners.
  • Evaluation Scope: Standalone evaluation remains necessary because it answers a different question from team evaluation, as illustrated by autonomous driving’s need for robust standalone assessment.Deployment can still involve human–AI systems in which people monitor, collaborate with, or take over from AI systems.
  • Evaluation Tiers: The proposed three tiers are standalone benchmarking, surrogate team evaluation for collaboration design, and real-participant team evaluation for deployment-critical validity.Tier 3 is reserved for systems approaching deployment or for establishing baseline team performance in a new domain.

4. Related Work

Prior work shows that human–AI collaboration can outperform either humans or AI alone, but its benefits and risks depend on collaboration mode, evaluation design, and team dynamics. This paper extends that literature by arguing that benchmarks should measure human value and synergy to steer research toward augmentative innovation.

  • Collaborative systems: Hybrid-intelligence systems dynamically routing subtasks between algorithms and crowd workers can outperform either party alone.Examples include Legion’s real-time crowdsourced control and interactive machine-teaching loops.
  • Collaboration modalities: Collaboration outcomes vary by modality: LLM sounding boards help non-experts approach expert-level content quality, while ghostwriter modes assign LLMs the main generation role.This contrast comes from creative-work experiments comparing feedback on human-created content with AI-led content generation.
  • Evaluation methods: Human–AI collaboration has reached 89-96% accuracy in systematic-review assessment, outperforming either humans or AI alone.New frameworks also use structured decision trees to select metrics for distinct collaboration modes.
  • Benchmark incentives: Benchmarks still rank agents in isolation, rarely optimizing user learning, calibrated trust, or human agency, even though benchmark design channels frontier ML research incentives.The paper proposes benchmarks measuring AI’s synergy and value to humans as a concrete lever connecting teaming research with economic debates over augmentative versus automating innovation.
  • Existing cooperative evaluation: The Anthropic Economic Index reports that 52% of tasks are classified as human-augmenting, up from 47% in 2025, reflecting performance of a human+AI system.Users select suitable tasks, correct outputs, solve subtasks, and steer the AI during these sessions.

5. Alternative Views

The proposed shift to human–AI team evaluation faces greater logistical costs, proxy-validation challenges, human variability, and difficulties defining collaboration quality. The paper argues these challenges can be addressed through shared infrastructure, surrogate models, experimental design, multidisciplinary AI research, and collaboration-specific metrics.

  • Logistical costs: Human evaluations are slower and more expensive than static benchmarks because they require IRB approval, participant recruitment, coordination, and performance measurement.The paper proposes shared interactive platforms, crowdsourced labor, human surrogate models, and phased evaluations to reduce these burdens.
  • Logistical costs: Human surrogate models can accelerate collaborative-AI iteration, but current proxies may underrepresent cognitive biases, fatigue, emotional states, and user diversity.The paper recommends using surrogates for rapid screening and preliminary iteration, not as substitutes for real participants.
  • AI research scope: Human–AI collaboration is a multidisciplinary problem with substantial ML research content, including adaptive, explainable, aligned, and representationally aligned AI.The paper rejects treating collaboration research as merely a shift toward HCI or human-factors engineering.
  • Proxy validation: Scaling HCI methods requires validating proxy measures, especially when humans enter the AI training loop and log-derived signals are optimized.Strong benchmarks for human+AI team performance could instead incentivize collaboration-promoting training, model scaffolding, inference steering, or dataset curation.
  • Human variability: Human variability can make benchmarks noisy, while sampled users’ skill at working with AI is a critical confounder that must be measured.The paper points to robust experimental designs, AI adaptability, validated surrogate models, and multifaceted metrics as ways to address variability.
  • Collaboration metrics: Although collaboration quality is subjective and harder to quantify than accuracy, observable behaviors and combined objective–subjective measures can yield actionable evaluations.Metrics should be selected and weighted according to the collaboration type, such as real-time co-creation, where fluency, turn-taking, and flow matter.

6. Call to Action

The paper calls for infrastructure and reporting practices that make human–AI collaboration measurable, including datasets, shared evaluation platforms, collaborative metrics, and human surrogate models. It proposes piloting teaming benchmarks in software engineering across solo and synchronous or asynchronous human–AI conditions.

  • Research infrastructure: Publish annotated logs of successful and failed human–AI interactions to enable research on team dynamics.The datasets should capture collaborative sessions rather than only outcomes.
  • Research infrastructure: Create shared platforms and APIs, or “Human–AI Interaction Gyms,” to reduce the overhead of running collaborative benchmarks.Standardized infrastructure would support repeatable evaluations across studies.
  • Evaluation practice: Report collaborative metrics alongside solo benchmarks when releasing models, measuring how effectively humans can work with them.The proposal explicitly rejects reporting autonomous performance alone.
  • Research infrastructure: Develop validated human surrogate models to enable rapid iteration before costly human studies.These surrogates are intended to simulate human collaborators.
  • Example teaming benchmark: A software-engineering pilot could adapt standalone benchmarks such as Terminal-Bench into interactive tasks based on real GitHub issues.The proposal also suggests enriching existing uplift studies for collaborative evaluation.
  • Example teaming benchmark: Participants could compare human alone, AI alone, human+AI synchronous, and human+AI asynchronous conditions using task performance and efficiency metrics.Suggested outcomes include issues resolved, code quality via test coverage and review acceptance, and time.

7. Conclusion

The paper concludes that the challenges of developing AI that works with humans are real but surmountable. It frames these challenges as exciting research opportunities and argues that the imperative outweighs the difficulties of changing evaluation practices.

  • 7. Conclusion: The challenges involved in developing AI that works with humans are real but surmountable.The authors characterize these challenges as comparable in scale to those the field has overcome previously.
  • 7. Conclusion: These challenges constitute exciting research opportunities for the field.The conclusion presents them as opportunities rather than merely obstacles.
  • 7. Conclusion: The imperative to develop AI that works with humans outweighs the difficulties of evolving the evaluation paradigm.The authors treat the collaborative evaluation shift as necessary despite the associated challenges.
Loading 2608.13577v1…