Source-linked AI summary
When combinations of humans and AI are useful: A systematic review and meta-analysis
Michelle Vaccaro, Abdullah Almaatouq, Thomas Malone
TL;DR
Human-AI collaboration can outperform either humans or AI alone in some settings, but the field lacks a broad account of when this occurs. This meta-analysis synthesizes 370 effect sizes from 106 experiments and finds that combinations generally underperform the better standalone option, with task-dependent exceptions.
Problem
Despite extensive human-AI research, the conditions under which combined systems outperform humans or AI alone remain insufficiently understood.
Method
The authors conducted a meta-analysis of 370 effect sizes from 106 recent experimental studies to quantify human-AI synergy and identify influencing factors.
Results
Human-AI systems generally underperformed the better standalone option, with losses in decision tasks but significantly greater gains in creation tasks.
Takeaways & Limitations
Human-AI synergy is heterogeneous, and creation tasks may be an especially fruitful area for further study.
Takeaways & Limitations
The quantitative results apply only to studies identified by the systematic review that measured humans alone, AI alone, and their combination.
Abstract
from arXiv · showhide
Inspired by the increasing use of AI to augment humans, researchers have studied human-AI systems involving different tasks, systems, and populations. Despite such a large body of work, we lack a broad conceptual understanding of when combinations of humans and AI are better than either alone. Here, we addressed this question by conducting a meta-analysis of over 100 recent experimental studies reporting over 300 effect sizes. First, we found that, on average, human-AI combinations performed significantly worse than the best of humans or AI alone. Second, we found performance losses in tasks that involved making decisions and significantly greater gains in tasks that involved creating content. Finally, when humans outperformed AI alone, we found performance gains in the combination, but when the AI outperformed humans alone we found losses. These findings highlight the heterogeneity of the effects of human-AI collaboration and point to promising avenues for improving human-AI systems.
1 Introduction
This systematic review and meta-analysis quantified human-AI synergy and augmentation across 370 effect sizes from 106 experiments. Human-AI systems improved on humans alone on average but did not outperform both humans and AI alone, while task type and relative human-versus-AI performance significantly shaped outcomes.
- Main findings: On average, human-AI systems demonstrated human augmentation by outperforming humans alone, but not human-AI synergy by failing to outperform both humans and AI alone.The results indicate that using either humans or AI alone would have been better than the studied combinations on the performance dimensions examined.
- Moderating factors: Providing suggested decisions together with confidence levels or explanations did not significantly affect human-AI system performance.The study contrasted these commonly examined design features with task type and relative human-versus-AI performance.
- Moderating factors: Task type and the relative performance of humans alone and AI alone significantly affected human-AI performance.These factors were identified as promising directions for designing systems with greater synergy.
2 Results
The review synthesized 370 effect sizes from 106 experiments across 74 included papers and found that human-AI systems were, on average, better than humans alone but worse than the better standalone performer. Effects varied substantially by task and by whether humans or AI performed better alone.
- Study scope: 370 unique effect sizes from 106 experiments across 74 included papers were synthesized in the meta-analysis.The initial search yielded 5126 papers, of which 74 met the inclusion criteria.
- Overall performance: g = −0.23 indicated that human-AI systems performed significantly worse than the better of humans alone or AI alone.The pooled effect was small and statistically significant: t(92) = −2.89, two-tailed p = 0.005, 95% CI −0.39 to −0.07.
- Overall performance: g = 0.64 showed that human-AI systems performed significantly better than humans alone, but not better than both humans alone and AI alone.This human-augmentation effect was medium to large: t(98) = 11.87, two-tailed p = 0.000, 95% CI 0.53 to 0.74.
3 Discussion
The meta-analysis found that human-AI systems generally underperformed the better standalone performer, although outcomes varied by task and relative human/AI performance. The authors therefore call for process-focused research while emphasizing substantial limitations and unexplained heterogeneity.
- Performance Losses from Human-AI Collaboration: Human-AI groups performed worse than either humans or AI working alone on average, despite the AI helping humans outperform humans alone.The authors distinguish this human augmentation from synergy with the best standalone performer.
- Moderating Effects: Task type significantly moderated synergy: decision tasks were associated with performance losses, whereas creation tasks were associated with performance gains.The authors identify task type as a significant moderator of human-AI collaboration effectiveness.
- Moderating Effect of Relative Human/AI Performance: When AI alone outperformed humans, human-AI systems showed substantial losses; when humans outperformed AI alone, the combinations showed performance gains.The authors suggest that relative baseline performance may relate to humans’ ability to decide when to rely on their own judgments or the algorithm’s.
- Surprisingly Insignificant Moderators: Across 300+ effect sizes, explanations, AI confidence information, and participant type did not impact human-AI collaboration effectiveness on average.The authors recommend shifting attention toward baseline performance, task type, and division of labor instead.
- Limitations: The results are limited to selected studies reporting standalone and combined performance, chosen tasks and populations, varying study quality, possible selection bias, and high unexplained heterogeneity.The analysis excludes tasks that humans or AI cannot perform alone and may not represent practical human-AI configurations outside laboratories.
- A Roadmap for Future Work: Finding Human-AI Synergy: Although combinations produced average performance losses, the authors argue that future work should identify effective processes for integrating humans and AI, especially in creation tasks.They present the overall result as a reason to study integration processes more specifically, not as evidence that combining humans and AI is inherently undesirable.
4 Methods · 4.1 Literature Review
The meta-analysis followed established systematic-review and PRISMA standards and applied preregistered eligibility criteria. Researchers searched multiple databases, coded comparative performance data and moderators, and used supplementary calculations or author contact when necessary.
- 4 Methods: The meta-analysis followed Kitchenham’s systematic-review guidelines and PRISMA standards.
- 4.1.1 Eligibility Criteria: Eligible studies were original experiments testing a human-AI system on a task and reporting quantitative performance for humans alone, AI alone, and their combination.
- 4.1.1 Eligibility Criteria: The review excluded studies lacking either human-alone or AI-alone performance, along with meta-analyses, reviews, theoretical work, qualitative analyses, commentaries, opinions, and simulations.
- 4.1.2 Search Strategy: The search covered ACM DL, AISeL, and WoS across computer, information, social, and other fields, using terms for human, AI, collaboration, and experiments.
- 4.1.2 Search Strategy: The search targeted studies published between January 1, 2020 and June 30, 2023, and included forward and backward searches of eligible studies.
- 4.1.3 Data Collection and Coding: Researchers recorded averages, standard deviations, and sample sizes for human-alone, AI-alone, and human-AI conditions to calculate the effect of combining human and artificial intelligence on task performance.
- 4.1.3 Data Collection and Coding: When studies lacked directly reported values, researchers derived standard deviations, computed statistics from public raw data, digitized plots, or contacted authors; unavailable effect sizes were excluded.
- 4.1.3 Data Collection and Coding: The coding captured 11 potential moderators, including publication date, preregistration, design, data and task characteristics, AI features, participant type, and performance metric.
4.2 Data Analysis · 4.3 Bias Tests
The analysis used standardized effect sizes and multilevel models to accommodate heterogeneous, dependent evidence. Bias tests found no evidence of publication bias for human-AI synergy but did find evidence favoring human augmentation studies.
- 4.2.1 Calculation of Effect Size: Hedges’ g measured human-AI effects against either the better-performing human-alone or AI-alone baseline for strong synergy, and against humans alone for augmentation.Because Hedges’ g is unitless, it enables comparisons across performance metrics and corrects upward bias in Cohen’s d.
- 4.2.2 Meta-Analytic Model: A random-effects model accounted for sampling error and true variability across experiments with different tasks, participants, and experimental designs.The model was selected because substantial between-experiment heterogeneity was expected.
- 4.2.2 Meta-Analytic Model: A three-level meta-analytic model nested effect sizes within experiments, accommodating multiple treatment groups and performance measures.This addressed dependencies that standard models assuming independent effect sizes cannot handle.
- 4.2.2 Meta-Analytic Model: Robust variance estimation accounted for dependent sampling errors from overlapping samples, while Knapp-Hartung adjustments supported inference using t-distribution-based tests and intervals.Moderator analyses used separate meta-regressions, and heterogeneity was quantified with I2 and multilevel variants.
- 4.3 Bias Tests: Publication-bias diagnostics combined funnel-plot inspection with Egger’s regression and rank-correlation tests, assessing asymmetry and missing nonsignificant effects.In the absence of publication bias, funnel-plot points should fall roughly symmetrically around the y-axis.
- 4.3 Bias Tests: For human augmentation, both Egger’s regression and the rank-correlation test indicated publication bias favoring studies where human-AI systems outperform humans alone.Egger’s regression: β = 1.96, t(104) = 3.24, two-tailed p = 0.002, 95% CI 0.76 to 3.16; rank correlation: τ = 0.19, two-tailed p = 0.000.
- 4.3 Bias Tests: The researchers did not correct for potential publication bias because adjustment methods can overcorrect and distort the original data.They preserved transparency by reporting the original data without adjustment.
4.4 Sensitivity Analysis
Sensitivity analyses consistently reproduced the main findings for human-AI synergy and human augmentation. Results were robust to clustering at the paper level, excluding outliers, omitting individual effect sizes, experiments, or papers, and removing estimated effect sizes.
- Paper-level clustering: Paper-level clustering produced comparable effects for human-AI synergy (g = −0.22) and human augmentation (g = 0.65).The estimates remained statistically significant: human-AI synergy, 95% CI −0.41 to −0.04; human augmentation, 95% CI 0.52 to 0.78.
- Outlier sensitivity: Excluding 11 human-AI synergy outliers and 9 human augmentation outliers yielded similar effects for both outcomes.Human-AI synergy was g = −0.25, while human augmentation was g = 0.60 after exclusion.
- Leave-one-out analyses: Leave-one-out analyses showed human-AI synergy estimates ranging from −0.28 to −0.19 and human augmentation estimates ranging from 0.61 to 0.66.Both effects remained significant in every analysis, demonstrating robustness to any single effect size, experiment, or paper.
- Estimated-effect-size sensitivity: Removing effect sizes estimated from figures or author-provided information produced almost identical results for human-AI synergy (g = −0.21) and human augmentation (g = 0.64).The corresponding confidence intervals were [−0.36, −0.05] and [0.53, 0.75], respectively.
5 Data Availability
The analysis data were compiled from studies identified in the systematic literature review and made available through the project’s Open Science Framework repository.
- The analysis data were compiled from studies identified in the systematic literature review.
- The collected data are available through the project’s Open Science Framework repository.
8 Author Contribution Statement
MV conceived the study idea with feedback from AA and TM, collected the data, and performed the statistical analysis with their feedback. MV, AA, and TM wrote the manuscript.
- MV conceived the study idea with feedback from AA and TM.
- MV collected the data and performed the statistical analysis with feedback from AA and TM.
- MV, AA, and TM wrote the manuscript.
10 Figures
The figures visualize 370 meta-analysis effect sizes and summarize three-level meta-regression results for moderator subgroups. They encode effect-size direction, uncertainty, subgroup sample sizes, and statistical differences in human-AI synergy and augmentation.
- Figure 1: 370 effect sizes appear in forest plots, with point positions showing effect-size values and bars showing 95% confidence intervals.Negative effect sizes are colored red and positive effect sizes green.
- Figure 1: The forest plots mark Hedges’ g = 0 with a black dotted line, representing the reference point for effect sizes.The supplied caption identifies this line as corresponding to an effect size of Hedges’ g = 0.
- Figure 2: Three-level meta-regression models report moderator-subgroup sample sizes and estimated effect sizes with corresponding 95% confidence intervals.N denotes the number of included effect sizes for each moderator subgroup level.
- Figure 2: Symbols before moderators indicate statistically significant differences between subgroups for human-AI synergy (*) and human augmentation (∧).The caption distinguishes the significance symbols by outcome.
Supplementary Information · S1 Supplementary Methods
The supplementary methods define human–AI outcome categories, describe the systematic review and data-coding workflow, and specify how task performance and effect sizes were calculated. The procedures include extracting or reconstructing performance data, coding moderators, and interpreting Hedges’ g consistently across metrics.
- S1.1 Types of Outcomes in Human-AI Systems: Human–AI synergy is defined as the combination outperforming both the human and AI alone.
- S1.1 Types of Outcomes in Human-AI Systems: Human augmentation means the human–AI group outperforms the human alone, whereas AI augmentation means it outperforms the AI alone.
- S1.1 Types of Outcomes in Human-AI Systems: Negative synergy is defined as the human–AI group outperforming neither the human nor AI alone.
- S1.2 Systematic Literature Review: The literature review used database search strings and a PRISMA-based screening and study-inclusion process.
- S1.3 Data Collection and Coding: In July 2023, records were downloaded as .ris files, deduplicated in Zotero, exported to Excel, screened by abstract, and assessed against inclusion criteria.
- S1.3 Data Collection and Coding: The primary outcome required means, standard deviations, and sample sizes for human-alone, AI-alone, and human–AI conditions.
- S1.3 Data Collection and Coding: Missing numerical data were reconstructed from confidence intervals, standard errors, public datasets, author responses, or digitized plots; studies without computable effect sizes were excluded.
- S1.4 Calculation of Effect Size: Effect sizes were oriented so positive Hedges’ g indicates combination gains over the baseline, negative values indicate losses, and larger absolute values indicate larger effects.
S2 Supplementary Results
The supplementary results show that most effect sizes concern decision tasks, while few concern creation tasks or predetermined human–AI divisions of labor. Predetermined task allocation was associated with positive synergy in a small set of experiments, and human–AI synergy showed suggestive but slow progress over time.
- Descriptive statistics: 34 effect sizes (9%) came from creation tasks, while 4 (1%) involved a pre-defined division of labor between humans and AI.The vast majority of effect sizes came from decision-task experiments.
- Division of labor: Only 3 of the 100+ experiments explored predetermined delegation, so differences from experiments without division of labor were not statistically significant.The 4 effect sizes from these experiments showed positive synergy, whereas experiments without this feature showed significantly negative effects.
- Division of labor: g = 0.22 for experiments with predetermined division of labor, compared with g = −0.24 without it.The predetermined-division estimate was not statistically significant (two-tailed p = 0.494), while the estimate without division was significantly negative (two-tailed p = 0.004).
- Examples: Human-alone summaries scored 5.8/7, exceeding AI-alone summaries at 4.7/7 and the human–AI combination at 5.5/7.This predetermined task allocation did not produce human–AI synergy.
- Evolution over time: Human–AI synergy showed potential progress over the years, but any progress was suggestive and not rapid relative to AI-system development.The contrast was especially notable given the quick development of large language models.