Source-linked AI summary
Can large language models replace humans in the systematic review process? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages
Qusai Khraisha, Sophie Put, Johanna Kappenberg, Azza Warraitch, Kristin Hadfield
TL;DR
Systematic reviews are important but slow, and evidence on whether LLMs can match human reviewers—particularly GPT-4 across diverse literature—has been limited. This pre-registered study evaluated GPT-4 in screening and data extraction across tasks, literature types, and languages. Results were mixed: adjusted performance was often low to moderate, but highly reliable prompts yielded almost perfect full-text screening agreement and improved further when false rejections were weighted.
Problem
Systematic reviews are slow and labour-intensive, while prior LLM evidence was limited and had not comprehensively tested GPT-4 across grey and non-English literature.
Method
The pre-registered study evaluated GPT-4 autonomously in title/abstract screening, full-text screening, and data extraction across peer-reviewed, grey, and non-English literature using reliability and agreement metrics.
Results
GPT-4 showed mixed performance: adjusted agreement was generally none to moderate, while highly reliable prompts produced almost perfect agreement (.91), increasing to .97 with weighted kappa.
Takeaways & Limitations
LLMs may rival humans for certain systematic-review tasks under reliable prompts, but their use currently requires substantial caution.
Takeaways & Limitations
The relatively small sample and efforts to balance non-English texts may limit generalisability and introduce bias.
Abstract
from arXiv · showhide
Systematic reviews are vital for guiding practice, research, and policy, yet they are often slow and labour-intensive. Large language models (LLMs) could offer a way to speed up and automate systematic reviews, but their performance in such tasks has not been comprehensively evaluated against humans, and no study has tested GPT-4, the biggest LLM so far. This pre-registered study evaluates GPT-4's capability in title/abstract screening, full-text review, and data extraction across various literature types and languages using a 'human-out-of-the-loop' approach. Although GPT-4 had accuracy on par with human performance in most tasks, results were skewed by chance agreement and dataset imbalance. After adjusting for these, there was a moderate level of performance for data extraction, and - barring studies that used highly reliable prompts - screening performance levelled at none to moderate for different stages and languages. When screening full-text literature using highly reliable prompts, GPT-4's performance was 'almost perfect.' Penalising GPT-4 for missing key studies using highly reliable prompts improved its performance even more. Our findings indicate that, currently, substantial caution should be used if LLMs are being used to conduct systematic reviews, but suggest that, for certain systematic review tasks delivered under reliable prompts, LLMs can rival human performance.
1 Introduction
Systematic reviews are important but slow and labour-intensive, while growing literature and complex questions intensify these challenges. AI tools may improve efficiency, yet prior evidence on whether they can match humans remains limited, especially for GPT-4, grey literature, and non-English studies.
- Systematic reviews support practice, research, and policy but can be too slow for their findings to remain current.
- The expanding scientific literature and increasingly specific research questions add to the workload of systematic reviews.
- AI tools have been developed for screening, data extraction, and study-quality assessment, but their performance deteriorates without human decision-making.
- No prior study had tested grey and non-English literature comprehensively or evaluated GPT-4 across the systematic-review tasks examined here.
- Earlier LLM studies were limited by contaminated datasets, inadequate metrics, narrow task coverage, and reliance on human-in-the-loop evaluation.
2 Methods
The study evaluated GPT-4 autonomously across screening and extraction tasks using varied literature types and languages, while testing prompt reliability and correcting for dataset imbalance and chance agreement. Performance was assessed against human decisions with sensitivity, specificity, accuracy, and agreement metrics.
- The study tested GPT-4 on peer-reviewed, grey, and non-English literature from an ongoing systematic review using four inclusion and exclusion criteria.
- Prompt formats were iteratively revised by supplying complete abstracts, separating criteria, and reducing batch sizes to address focus, complexity, volume, and hallucination concerns.
- Test-retest reliability was measured across 10 studies per criterion, with four prompts tested five times to assess how prompt reliability affected accuracy.
- Full texts were segmented and each snippet was paired with its tested criterion, with Include or Exclude decisions determining progression through criteria.
- 2.3 Analysis metrics: Sensitivity, specificity, and accuracy quantified GPT-4 performance, while Cohen’s kappa compared observed agreement with agreement expected by chance.
- 2.3 Analysis metrics: PABAK and weighted kappa addressed prevalence, bias, and the severity of false rejections in imbalanced screening data.
3 Results
GPT-4 reliability varied by criterion, and dataset balance differed across literature types, languages, and stages. Although specificity and raw accuracy often approached human performance, adjusted agreement was generally low to moderate, except under highly reliable prompts.
- 100% reliability occurred for empirical-data and refugee-status prompts, compared with 50% for parenting behaviour and 70% for protracted refugee situations.
- The high-reliability prompt group contained 23 studies divided approximately equally among English peer-reviewed, English grey, and non-English literature.
- English peer-reviewed data were balanced except during extraction (.03), whereas grey and non-English data became irrelevance-skewed during full-text screening (.11 and .09) and extraction (.24 and .20).
- Specificity exceeded .80 across categories except English peer-reviewed full-text screening, while extraction sensitivity reached .75 for English peer-reviewed and .65 for grey literature but only .36 for non-English studies.
- Adjusted kappa scores ranged from none to moderate, but highly reliable prompts produced almost perfect agreement (.91), rising to .97 with weighted kappa.
4 Discussion
GPT-4 showed mixed performance across systematic-review tasks, languages, and literature types. Its apparent parity with humans was affected by chance agreement and dataset imbalance, while prompt reliability and dataset composition constrained interpretation.
- GPT-4’s accuracy was influenced by chance agreement and dataset imbalance, often substantially underperforming humans after adjustment.
- Under entirely highly reliable prompts, GPT-4 achieved ‘almost perfect’ full-text screening performance comparable to humans.
- GPT-4 showed moderate performance in non-English and grey literature but very poor performance with English peer-reviewed texts for full-text screening and extraction.
- The unusual balance of English peer-reviewed studies, small sample size, and absence of advance prompt-engineering planning limit generalisability.
5 Conclusion
The study uses the Hans analogy to emphasize GPT-4’s dependence on clear human input while generating outputs autonomously. Reliable prompts were associated with almost perfect screening performance, suggesting conditional potential for systematic-review automation.
- GPT-4’s screening performance rose to almost perfect when it received a reliable prompt.
- Unlike Hans, GPT-4 does not require human guidance for each output, but it remains heavily influenced by the prompts it receives.
- The findings suggest that reliable prompting could support a transformative role for LLMs in systematic reviews.