Source-linked AI summary
PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews
Miguel Zabaleta, Baihan Lin
TL;DR
AI-assisted systematic reviews increasingly affect consequential evidence-synthesis decisions, but workflow details and evaluation evidence are reported inconsistently. This study analyzes 888 SciLitBench papers and uses the observed patterns to introduce PRISMA-LLM, whose key result is that reporting varies by workflow context while positive assessments often coexist with unmet reliability requirements.
Problem
AI-assisted review workflows require detailed reporting to make automation, error propagation, independent checking, and reuse auditable.
Method
The study analyzes SciLitBench’s 888 papers, comparing automation trends, review-stage use, evaluation and limitation reporting, and LLM workflow complexity.
Results
52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement, while reporting coverage increased with workflow complexity and software/product papers reported less evaluation detail.
Takeaways & Limitations
PRISMA-LLM separates implementation disclosure from consequence-sensitive evaluation and makes human oversight, evaluation evidence, and failure modes more visible.
Takeaways & Limitations
SciLitBench is concentrated in life sciences and medicine, covers publications through June 2025, and uses reporting richness as a descriptive rather than validated study-quality measure.
Abstract
from arXiv · showhide
Large language models (LLMs) and AI-enabled software increasingly participate in systematic-review decisions, yet the information needed to audit these workflows is reported inconsistently. We analyze SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, to characterize changes in methods, review-stage use, evaluation and reported limitations. Automation has shifted toward LLM- and software-facing workflows, including stages that can alter the evidence base. Since 2023, 38.0% of software/product papers reported no evaluation, compared with 9.3% of LLM papers. Reporting coverage increased with LLM workflow complexity, yet 52% of positive-only LLM evaluations still reported an unmet reliability or performance requirement. From these patterns, we introduce PRISMA-LLM, an empirically grounded framework separating implementation disclosure from consequence-sensitive evaluation and limitation reporting.
1 Results
SciLitBench shows rapid growth in review automation, a shift toward LLM/software-facing workflows and consequential review stages, and uneven evaluation reporting. These patterns motivate PRISMA-LLM’s focus on implementation disclosure, human oversight, evaluation evidence, and failure modes.
- 1.1 Growth and shifts: 4.7% monthly growth was estimated for review-automation publications from January 2020 through June 2025, with 61.7% of records from life sciences and medicine.Publication counts rose around 2018–2019 and steepened after ChatGPT’s release.
- 1.1 Growth and shifts: Review automation shifted toward LLMs and software products and increasingly entered screening, selection, and evidence-construction stages that can alter the evidence base.Recent automation therefore extends beyond discovery and navigation.
- 1.1 Growth and shifts: 84.1% of LLM usage relied on proprietary systems, while open-weight usage accounted for 4.9% of papers.Mixed proprietary/open-weight usage was 11.0%.
- 1.2 Evaluation reporting differs by automation context and method complexity: Since 2023, software/product papers averaged 3.3 reporting-richness points and 38.0% reported no evaluation, compared with 6.3 points and 9.3% for LLM papers.Reporting richness is a descriptive index, not a study-quality measure.
- 1.2 Evaluation reporting differs by automation context and method complexity: By June 2025, prompt-only LLM workflows comprised 34.0% of the latest window, compared with 29.2% software/product-only and 8.5% retrieval, adapted, or agentic workflows.Engineered or structured LLM workflows comprised 14.2%.
- 1.3 Usefulness and adequacy are distinct judgments: 52% of positive-only LLM evaluations still reported a concern that the workflow fell below the reliability or performance bar required for intended use.The corresponding shares were 67% for positive-with-caveats papers and 79% for mixed or negative papers.
2 Discussion
Across 888 papers, automation is entering evidence-determining review stages while evaluation information remains uneven, motivating PRISMA-LLM’s combination of implementation disclosure with consequence-sensitive evaluation. The framework also distinguishes disclosure complexity from risk and requires validation and human verification when workflows can change the evidence base.
- More than half of positive-only LLM evaluations still reported an unmet performance or reliability requirement, separating usefulness from adequacy for delegation.The paper therefore emphasizes residual failure modes and responsibility for consequential decisions.
- The proposed disclosure levels reconstruct workflow choices but do not measure error consequences, human oversight, or proprietary-system opacity.Workflows that can change the evidence base therefore require task-specific validation and human verification regardless of disclosure level.
- SciLitBench is concentrated in life sciences and medicine, covers publications through June 2025, and uses a descriptive richness index with limited within-dimension depth.The complexity gradient is descriptive and may be confounded, while paper-level absence of evaluation does not prove validation is absent elsewhere.
- PRISMA-LLM is proposed in this study and has not yet undergone formal consensus development, independent usability testing, or prospective evaluation.Those studies could test and refine the framework as AI-assisted review practice evolves.
- PRISMA-LLM translates inspectability requirements into a checklist, implementation-disclosure scheme, and consequence-sensitive evaluation framework.It is intended for use alongside PRISMA 2020 rather than as a replacement.
3 Methods
The study performs a secondary analysis of 888 SciLitBench papers, combining longitudinal, stage-based, approach-based, and reporting-richness analyses. It uses descriptive paper-level measures to characterize automation and reporting patterns without treating richness as study quality.
- Study design: The analysis covers 888 full-text papers with six annotated fields and 14,726 annotation items, using each paper as the unit of analysis.Fields include publication year, domain, review stage, computational approach, evaluation results, and limitations.
- Temporal analysis: The study combines recovered publication dates, rolling three-month windows, and log-linear growth fitting to analyze temporal trends through June 2025.The reported growth rate is descriptive, and the extrapolation is not treated as a calibrated forecast or causal break.
- Review-stage analysis: Review stages are grouped into discovery and navigation, screening and selection, and evidence construction, with papers allowed to contribute to multiple stages.Evidence construction includes data extraction, quality assessment, and claim verification.
- Method classification: Approach, model-family, and access analyses use multi-label annotations, while an exclusive orientation classifies papers as custom, LLM/software-facing, hybrid, or other.Model families and access categories are normalized from approach annotations.
- Reporting measures: Reporting richness is a descriptive index from 0 to 15 that sums performance, comparison, modification, resource, and limitation dimensions.Each dimension contributes 0 to 3 points according to the number of reported items, and the score is not a study-quality appraisal.
- Framework construction: PRISMA-LLM translates observed reporting patterns into disclosure tiers and separates implementation complexity from methodological risk.The framework links implementation choices to minimum expectations for evaluation, limitations, and reporting detail.
Code and materials availability
The study makes its analysis code, PRISMA-LLM instructions, checklist workbook, and checklist evidence table available online.
- Materials: Analysis code, reporting instructions, a fillable checklist workbook, and the checklist evidence table are available in the PRISMA-LLM GitHub repository.Access to the underlying SciLitBench corpus and source-derived materials follows the resource’s stated terms.
Supplementary Information
The supplementary material documents calculation procedures, reporting analyses, model and limitation profiles, and supporting figures. It also reports substantial variation in evaluation and performance evidence across approaches and stages.
- Calculation details: Supplementary analyses use papers as denominators and retain multi-label fields, so percentages across approach, stage, limitation, and model-family groups need not sum to 100%.The contributing paper count is reported when needed to interpret an estimate.
- S2 Supplementary analyses: 52% of positive-only LLM evaluations still reported at least one unmet high-bar concern, showing that positive assessments and reliability requirements can diverge.The share was 67% for positive evaluations with caveats and 79% for mixed or negative evaluations.
- S2 Supplementary analyses: LLM papers reported unmet high bars, prompt sensitivity, limited validation, and narrow data more often than software/product papers.The corresponding LLM versus software/product prevalences were 54% versus 26%, 40% versus 15%, 32% versus 24%, and 45% versus 34%.
- S2 Supplementary analyses: Software/product papers reported fewer evaluation items than LLM papers within review-stage groups, with the largest gap in screening and selection.In screening, software/product papers averaged 3.2 items and had a 39.3% no-evaluation rate, versus 7.3 items and 3.8% for LLM papers.
- S2 Supplementary analyses: Reported precision and recall varied substantially across approach groups for both screening and selection and data extraction.For screening, LLM median recall was 0.884 and median precision was 0.695; for extraction, the corresponding medians were 0.890 and 0.881.
- S3 Supplementary Figures: The supplementary figures show model-family distributions, recurring limitation prevalence, reporting richness over time, and stage-specific evaluation reporting.Figures use within-group percentages or means and include cautions about multi-label groups, heterogeneous studies, and small family sizes.