Source-linked AI summary
Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency
Nikol Figalová, Lynn Huestegge, Anne Böckler-Raettig
TL;DR
Evidence remains limited on how LLM screening workflows perform in conceptually complex reviews and how processing configuration affects their results. This study compared complete human and LLM workflows, finding that GPT-5.4 file-batch runs approached human workload–recall profiles, whereas all-at-once configurations recovered substantially fewer verified eligible records.
Problem
Evidence is limited on LLM screening in conceptually complex tasks and on comparisons with both individual and distributed human screening workflows.
Method
The study evaluated complete human and LLM screening workflows, distinguishing processing configurations and measuring recovery, workload, agreement, and consistency.
Results
GPT-5.4 file-batch runs produced workload–recall profiles close to human workflows, whereas all-at-once configurations recovered substantially fewer verified eligible records.
Takeaways & Limitations
LLMs are better suited to documented, auditable, human-supervised screening workflows than autonomous replacements for human screening judgment.
Takeaways & Limitations
Operational recall covered only the known verified eligible set because final eligibility was not independently established for all 1,131 benchmark records.
Abstract
from arXiv · showhide
Background. Large language models (LLMs) are increasingly used for screening in evidence synthesis, where false negatives can remove relevant studies before full-text assessment. We compared human and LLM title-and-abstract screening workflows in a preregistered study embedded in a conceptually complex scoping review. Methods. After a conservative title-only screen, 1,131 records were screened by one review lead, four trained assistants screening non-overlapping subsets, and seven complete LLM runs using different models and processing configurations, including a nominally identical repeat run. We compared retained workload, operational recall against 316 verified eligible records, agreement, run-to-run consistency, and procedural burden. Because eligibility was verified only for records advanced and assessed in the parent review, recall estimates were operational. Results. No workflow recovered all verified eligible records. The human workflows and two GPT-5.4 file-batch runs retained 42.2-45.0% of records while achieving 82.3-82.9% recall. Gemini 3.1 file batches achieved the highest recall (83.9%) but retained 56.7% of records. All-at-once configurations recovered fewer eligible records than corresponding file-batch configurations. Two nominally identical GPT-5.4 file-batch runs agreed on 91.7% of records but differed on 94 records, including 29 verified eligible records retained by only one run. Discussion. LLM screening performance depended on the implemented workflow, not model identity alone. Processing configuration, workload, record-level variation, and human-LLM decision integration are therefore substantive properties of deployed systems. For high-recall tasks, LLMs are better suited to validated, auditable, human-supervised workflows than autonomous exclusion.
Introduction
LLM screening for evidence synthesis should be evaluated as a complete, implemented workflow rather than by model identity alone. This evaluation addresses recall, retained workload, agreement, conditional reference-based measures, run-to-run consistency, and records missed or recovered by different workflows while accounting for partial verification bias.
- Motivation: False-negative title-and-abstract decisions can permanently remove relevant evidence, making screening a demanding task with asymmetric consequences.Large numbers of short documents must be classified against natural-language eligibility criteria before subsequent assessment.
- Workflow framing: An implemented LLM screening system includes prompts, uncertainty rules, record preparation, processing configuration, interaction procedures, output handling, and quality-control steps.The relevant unit of evaluation is therefore the complete workflow, not necessarily the underlying model alone.
- Workflow variability: Processing configuration and repeated execution can change classification performance and individual record decisions even when the model, records, and instructions appear unchanged.Batch size may be a substantive workflow parameter, and similar aggregate recall or workload does not guarantee record-level reproducibility.
- Evaluation framework: Recall, agreement, and downstream workload capture distinct properties, so no single metric fully characterizes a screening workflow.High agreement can reflect concordant exclusions, while high recovery can result from retaining a large proportion of the dataset.
- Evaluation framework: Reference-based recall and related measures are conditional because final eligibility is usually verified only for records advanced to full-text assessment.Records excluded by all advancement workflows or lacking retrievable full texts have unknown final eligibility and cannot automatically be treated as true negatives.
- Study scope: The study compares a solo review lead, four trained assistants, and seven complete LLM runs on the same benchmark while varying models, processing configurations, and nominally identical repetition.It defines workflow as the complete screening procedure and run as one execution on the benchmark records.
Materials and Methods · Study design, preregistration, and open materials · Benchmark dataset and analytical sets
This preregistered comparative methodological study evaluated human and LLM title-and-abstract screening workflows within a scoping review. It used a benchmark of 1,131 records and distinguished analytical sets for workload, agreement, consistency, classification measures, and operational recall.
- Study design, preregistration, and open materials: The study was preregistered on the Open Science Framework on February 24, 2026, after searching and deduplication but before preliminary title-only and title-and-abstract screening.The preregistration specified the principal human and LLM screening comparisons and analytical approach.
- Study design, preregistration, and open materials: Open materials included outcome definitions, statistical analyses, the screening manual, input data, complete LLM prompts, screening outputs, and comprehensive analytical results.Deviations from the preregistered plan and exploratory additions were documented in Supplementary Material S1.
- Study design, preregistration, and open materials: The study reported only parent-review procedures needed to define the screening task, benchmark dataset, and reference outcomes, while complete search strategies and PRISMA-ScR reporting appeared in a companion manuscript.The present study focused on comparative human and LLM title-and-abstract screening workflows.
- Benchmark dataset and analytical sets: The parent-review searches covered PubMed, PsycINFO, Web of Science Core Collection, and OSF Preprints, yielding 5,291 deduplicated records.A supplementary Web of Science search was conducted on February 23, 2026, after the initial searches on February 16, 2026.
- Benchmark dataset and analytical sets: A conservative preliminary title-only screen removed 4,160 records while retaining records whose eligibility could not be determined reliably from titles; it only constructed the benchmark.This preliminary screen was not evaluated as a screening workflow in the present study.
- Benchmark dataset and analytical sets: The benchmark comprised 1,131 records, which all human workflows and seven LLM runs independently screened at the title-and-abstract level.Backward-citation records were excluded, so performance estimates concerned records remaining after the preliminary title-only screen.
- Benchmark dataset and analytical sets: Three analytical sets comprised the complete benchmark dataset (n = 1,131), full-text-assessed set (n = 859), and verified eligible set (n = 316).They supported workload/agreement/consistency analyses, conditional reference-based classification measures, and operational recall with missed eligible records, respectively.
Data preprocessing · Screening task and decision structure
The study standardized 1,131 post-screen records for human and LLM title-and-abstract screening using minimally modified bibliographic inputs and structured output checks. Eligibility was operationalized through a shared manual, four mutually exclusive decisions, and a binary retained/not-retained mapping for primary analyses.
- Data preprocessing: 1,131 records received stable identifiers after the conservative preliminary title-only screen.LLM inputs contained only record identifiers, titles, and abstracts, with batching or file division determined by processing configuration.
- Data preprocessing: LLM outputs were structurally checked, linked by stable identifier, and formatting-standardized without substantive manual changes to decisions.Checks covered missing or duplicated identifiers, missing decisions, and malformed rows.
- Screening task and decision structure: Include and Unclear mapped to retained, whereas Exclude and Exclude–citation seed mapped to not retained for primary quantitative analyses.The binary mapping represented whether a record would proceed toward full-text assessment; primary exclusion reasons were collected for transparency but not included in the supplied passage’s unfinished sentence.
- Screening task and decision structure: A screening manual operationalized the parent review’s eligibility criteria and was refined through human calibration before formal screening.The manual provided a common decision framework for human and LLM workflows.
- Screening task and decision structure: Records were potentially eligible when gaze or related eye behaviour was substantively examined as conveying or shaping meaning within a human social context.Eligible phenomena included direct or averted gaze, eye contact, mutual gaze, gaze shifts, and gaze avoidance.
- Screening task and decision structure: Records were excluded for primary focuses including developmental or clinical populations, psychiatric classification, non-social attention, methods-only gaze research, or non-central technology evaluation.Studies involving robots, avatars, or virtual agents remained eligible when they examined human social interpretations of gaze.
- Screening task and decision structure: Each record received one of four mutually exclusive title-and-abstract decisions, including Exclude–citation seed.The supplied passages explicitly identify Exclude–citation seed as the fourth decision category.
- Screening task and decision structure: Unclear indicated insufficient title-and-abstract information for a confident decision, while Exclude–citation seed marked an otherwise ineligible record useful for backward citation searching.The citation-seed category preserved records for supplementary search despite ineligibility.
Human screening workflows
Human screening comprised a single-reviewer workflow covering all 1,131 benchmark records and a distributed workflow in which four assistants screened non-overlapping subsets. A 28-record calibration exercise informed the finalized manual, but formal distributed screening used no ongoing overlap or adjudication.
- Single-reviewer workflow: 1,131 benchmark records were screened by the review lead using the final screening manual and four predefined decisions.The single-reviewer workflow was conducted from March 2 to March 16, 2026.
- Distributed-team workflow: Four psychology student assistants screened non-overlapping subsets whose decisions were combined into one distributed-team workflow.Assistants were enrolled in a master’s-level psychology programme.
- Calibration and formal screening: 28 calibration records were independently assessed before discussion and manual finalization, with no second calibration round.Calibration decisions were not carried forward and the records were reassessed during formal screening.
- Calibration and formal screening: Formal distributed screening used no additional overlap subset, ongoing double-screening, or formal adjudication procedure.The calibration records remained in their original benchmark positions and were reassessed by the responsible screeners.
- Distributed-team workflow: Because assistants screened different non-random subsets, assistant-level differences were interpreted descriptively rather than as direct comparisons of reviewer performance.Files were divided sequentially, not randomized, and selected on a first-come-first-served basis.
LLM-based screening workflows
The study evaluated seven implemented LLM screening workflows spanning models and processing configurations, using a structured prompt aligned with the human screening manual. Workflow estimates reflect hosted-interface implementations whose proprietary infrastructure, updates, and possible prior publication exposure could not be independently assessed.
- Study design: Seven complete LLM runs screened all 1,131 benchmark records across file-batch, all-at-once, repeat, interactive, and model–configuration combinations.The runs used ChatGPT 5.4 Thinking, Gemini 3 Thinking, and Gemini 3.1 Pro through standard hosted web interfaces.
- Prompt development: The LLM prompt translated the final human screening manual into explicit eligibility, decision, uncertainty, exclusion, and output requirements.It was iteratively developed with ChatGPT 5.4 Thinking and reviewed by NF for alignment with the human manual.
- Prompt development: The 28-record procedural test checked execution and output structure, while human decisions and full-text outcomes remained unavailable to models.The same 28 records stayed in the benchmark and were reassessed during every formal run.
- Limitations: The calibration set was not used to select prompt variants by agreement, recall, or eligibility, and no systematic prompt-sensitivity analysis was conducted.Because ChatGPT 5.4 Thinking assisted with prompt development, estimates were interpreted as evaluations of complete implemented workflows.
- Limitations: Proprietary hosted services prevented assessment of training-corpus exposure or undocumented updates, so model-specific effects could not be isolated from hosted workflow properties.The inferential target was the implemented model–prompt–processing-configuration workflow under the described conditions, not external validation of the underlying models.
Advancement to full-text assessment and reference outcomes
Full-text advancement used a prospectively designated liberal union of two human workflows and two file-batch LLM workflows, retrieving records marked Include or Unclear by any one workflow. Full-text eligibility was assessed by a single reviewer, so reference-based outcomes reflected the implemented review pathway rather than a fully independent standard.
- Advancement rule: A record advanced when any designated workflow assigned Include or Unclear, without requiring agreement, majority voting, or consensus adjudication.Records assigned Exclude or Exclude–citation seed by all four workflows were not selected for retrieval.
- Advancement rule: The retrieval set was fixed before later comparative runs, so those runs could not retrospectively change which records underwent full-text retrieval.The later runs included the nominally identical ChatGPT repeat, Gemini 3.1 Pro, and interactive sequential configurations.
- Full-text assessment: 859 retrieved full-text reports were assessed by NF alone, without displaying workflow-specific title-and-abstract decisions during full-text assessment.Because NF had previously completed the single-reviewer screen, the assessment was not fully blinded to prior screening experience.
- Reference outcomes: Reference-based measures quantified performance within the implemented review pathway because eligibility was verified only among records advanced by the four-workflow union and successfully retrieved.Eligible records missed by all four workflows could remain unidentified among 209 non-advanced records, while 63 unretrieved records had unknown eligibility.
Outcome measures and statistical analysis
The analysis jointly evaluated retained workload and operational recall, while separately assessing conditional classification performance, agreement, run-to-run consistency, record-level complementarity, and liberal unions. Confidence intervals used resampling, with no paired hypothesis tests or between-workflow contrast intervals.
- Primary screening trade-off: Retained workload counted Include or Unclear decisions among 1,131 benchmark records, while operational recall measured retained verified eligible records among 316.Recall was explicitly operational because final eligibility was not verified across the complete benchmark.
- Conditional classification: Conditional classification analyses covered 859 records with verified full-text outcomes and included precision, specificity, F1, and recall.These measures did not extend to 209 non-advanced or 63 unretrieved records with unknown final eligibility.
- Agreement: All 36 pairwise comparisons among nine screening outputs assessed retained/not-retained agreement using overall agreement, Cohen’s κ, Gwet’s AC1, and positive and negative agreement.AC1 was included because κ can be strongly influenced by category prevalence and marginal decision distributions.
- Consistency and complementarity: The two nominally identical ChatGPT 5.4 file-batch runs were compared for aggregate performance, record retention consistency, and discordant decisions involving verified eligible records.Record-level analyses also examined missed-eligible overlap, uniquely recovered records, workflow complementarity, and liberal unions trading recovered eligible records against retained workload.
- Uncertainty estimation: 95% confidence intervals used 10,000 resamples, with stratified resampling for classification measures and multinomial resampling of observed 2 × 2 tables for agreement.Percentile intervals were obtained from the resulting empirical distributions.
- Inference and sensitivity analysis: No paired hypothesis tests or between-workflow contrast intervals were calculated, and a sensitivity analysis excluded 28 human-calibration and procedural-prompt-testing records.The study used the complete available benchmark descriptively rather than conducting an a priori power calculation.
Software and computational reproducibility
The analysis environment and random seed were specified, while code, inputs, prompts, outputs, definitions, and results were archived to support reconstruction. Re-execution of proprietary hosted LLM workflows was not guaranteed to reproduce the original outputs.
- Computational environment: Analyses used R 4.6.1, RStudio 2026.06.0, Windows 11, tidyverse 2.0.0, knitr 1.51, and ggrepel 0.9.8.The random seed for resampling was 20260703.
- Open materials: Open materials archived the analysis code, benchmark input data, LLM prompts, screening outputs, statistical definitions, and analytical results.The archived screening outputs allow the reported comparisons to be reconstructed.
- Re-execution limitations: Re-execution of proprietary hosted LLM workflows could not be guaranteed to reproduce the original outputs.The passage attributes this limitation to the underlying models and web interfaces being externally controlled.
Results
Screening workflows showed recall–workload trade-offs, with human workflows and GPT-5.4 file batches achieving similar operational performance, while Gemini configurations and all-at-once processing produced different workload and recall profiles. Agreement and aggregate metrics did not guarantee reproducible record-level classifications, although combining workflows increased recovery at the cost of retained workload.
- No output recovered all 316 verified eligible records, and greater retention did not consistently correspond to greater recovery.
- 42.2–45.0% retained workload and 82.3–82.9% operational recall characterized the two human workflows and two GPT-5.4 file-batch runs.The single reviewer retained the fewest benchmark records; interactive GPT-5.4 batches-of-10 achieved 44.5% retained workload and 79.7% operational recall.
- 83.9% was Gemini 3.1 file batches’ highest operational-recall point estimate, alongside 56.7% retained workload; Gemini 3 file batches retained 63.1% while recovering 77.2%.GPT-5.4 all-at-once processing recovered 66.8% versus 82.6% for file batches, while Gemini 3 all-at-once recovered 66.1% versus 77.2%.
- 72.8% agreement separated the single reviewer and distributed team, compared with 71.2–74.0% for GPT-5.4 file batches and humans and 55.0–55.7% for Gemini 3 file batches.Agreement also differed across configurations: GPT-5.4 file batches versus all-at-once reached 82.5%, while Gemini 3 configurations reached 62.4%; agreement did not directly map onto recovery.
- 91.7% agreement between nominally identical GPT-5.4 file-batch runs concealed 94 record-level disagreements, including 29 verified eligible records retained by only one run.The runs agreed on 1,037 of 1,131 records, while aggregate workload, recall, precision, specificity, and F1 remained closely aligned.
- 98.1% recovery from combining the two human workflows required retaining 57.2% of benchmark records; among pairs recovering 310 eligible records, this was the lowest workload.A human–Gemini 3.1 pair matched 310 recovered records but retained 67.4%, while combining GPT-5.4 runs recovered 276 records at 48.5% retained workload.
Discussion
Screening performance depended on the implemented workflow, not model identity alone, with processing configuration, workload, complementarity, and record-level consistency shaping operating points. The findings support validated, auditable, human-supervised LLM decision support rather than unvalidated autonomous exclusion in conceptually complex, high-recall tasks.
- Workflow-level performance: Screening performance was a property of the implemented workflow rather than the displayed model alone.The evaluated procedures differed in recovery, retained workload, agreement, conditional classification performance, complementarity, and run-to-run consistency.
- Workload–recall trade-offs: 42–45% retained workload paired with 82–83% recovered verified eligible records for human workflows and GPT-5.4 file-batch runs.Gemini 3.1 file batches achieved the highest operational-recall point estimate but retained substantially more records, while all-at-once configurations recovered only about two thirds.
- Processing configuration: File-batch processing recovered more verified eligible records than all-at-once processing for both GPT-5.4 and Gemini 3.The comparisons do not establish a causal effect of processing configuration because model, interface, date, interaction structure, and configuration were not varied independently.
- Complementarity: 98.1% recovery resulted from the liberal union of the two human workflows, with 57.2% of the benchmark retained.Combining outputs changes the operating point: higher recovery generally requires more downstream assessment, and the value depends on pairing and the relative cost of missed evidence.
- Run-to-run consistency: 91.7% agreement between nominally identical GPT-5.4 file-batch runs concealed disagreement on 94 records, including 29 verified eligible records retained by only one run.Aggregate workload, recall, and conditional classification measures were closely aligned, but record-level consistency remained a separate operational property.
- Deployment implications: Unvalidated LLM runs should not serve as autonomous exclusion components in conceptually complex, high-recall screening tasks.A more defensible role is auditable decision support, such as supplementary screening, prioritisation, or identifying uncertain cases; operating points depend on error costs, resources, expertise, and conceptual complexity.
Ethics statement
This methodological study analysed bibliographic records and screening outputs without collecting or analysing patient or directly identifiable personal data; formal ethics approval was therefore not required.
- Ethics statement: No patient data or directly identifiable personal data were collected or analysed, and formal ethics approval was not required.The study used bibliographic records and screening outputs from an evidence-synthesis project.
Data availability
The study’s preregistration and supporting materials—including the screening manual, prompts, data, outputs, results, and analysis code—are openly available through the Open Science Framework. Records of the AI-assisted research workflows and manuscript-preparation uses were retained and provided in supplementary materials and the research repository.
- Data availability: The preregistration is available through the Open Science Framework.https://osf.io/cnsr2/overview?view_only=697b0ae9f0234f069bc18173626ba954
- Data availability: The screening manual, prompts, input data, screening outputs, detailed results, and analysis code are available through the Open Science Framework.https://osf.io/v2q4r
- Data availability: Full records of prompts and methodological procedures for the LLM screening workflows were retained and provided in supplementary materials and the research repository.The workflows used OpenAI ChatGPT 5.4 Thinking, Google Gemini 3 Thinking, and Google Gemini 3.1 Pro.
- Data availability: During manuscript preparation, ChatGPT 5.4 and ChatGPT 5.6 assisted with editing, analysis checks, LaTeX, and R-code development.Figure 1 was generated from study data using R, not an AI image-generation model.
Related work
The methodological study was embedded in a scoping review reported separately by Figalová et al., whose companion report presents the review’s substantive findings.
- Related work: The present manuscript evaluates human and LLM-based screening workflows, while the companion report presents the scoping review’s substantive findings.The relationship between the reports and any overlapping methods or data is disclosed in the manuscript.