Source-linked AI summary
ARISMA: Guidelines for AI- and LLM-Assisted Systematic Reviews, Scoping Reviews, and Mapping Studies
Mahyar Tourchi Moghaddam, Mina Alipour
TL;DR
AI is entering evidence synthesis while existing standards do not fully specify appropriate use, validation, human sign-off, or disclosure. ARISMA develops a literature-informed and expert-refined framework that treats AI as a bounded, benchmarked, documented, and reversible assistant. It concludes that responsible use requires auditable governance and continued human responsibility, especially where performance is mixed or outputs are synthesis-critical.
Problem
Existing review standards remain essential but do not answer when AI use is justified, which outputs require validation, which decisions require human sign-off, or what AI involvement readers should be told.
Method
ARISMA combines a lifecycle taxonomy, stepwise review guidance, calibration, provenance tracking, governance controls, reporting instruments, and structured expert consultation.
Results
ARISMA provides a practical, auditable guideline that keeps consequential decisions human-accountable while permitting conditional AI assistance across evidence-synthesis tasks.
Takeaways & Limitations
AI-assisted review should be inspected, benchmarked, logged, and reversible, with permissions earned through task- and corpus-specific calibration rather than assumed from the model.
Takeaways & Limitations
Current evidence is mixed across tasks, and synthesis-critical values and final appraisal judgments still require explicit human verification.
Abstract
from arXiv · showhide
Systematic reviews, scoping reviews, mapping studies, and related evidence syntheses are increasingly difficult to conduct with fully manual workflows as search volumes, update cycles, and synthesis requirements continue to expand. At the same time, artificial intelligence, machine learning, and large language models are rapidly entering review practice across query formulation, screening, extraction, categorization, appraisal support, and reporting. Yet the empirical evidence remains uneven, task-dependent, and insufficient to justify unconstrained automation. Existing standards such as PRISMA 2020, PRISMA-S, PRISMA-ScR, PRISMA-P, PRESS, and SWiM remain essential, but none provides an end-to-end operational standard for when AI use is methodologically appropriate, how it should be validated, which review decisions must remain human-led, and how AI involvement should be reported so that readers can audit it. This paper proposes ARISMA, an AI Reporting and Integration standard for Systematic Methods and Analysis. ARISMA treats AI as an inspected, benchmarked, logged, and reversible assistant rather than an autonomous reviewer. It is built around one governing principle: every consequential scientific decision must remain human-interpretable, human-auditable, and human-accountable. The paper contributes a lifecycle taxonomy, process guidance, stepwise recommendations across the review pipeline, a governance and provenance model, a tool-support framework, an AI-integrated reporting checklist, and a validation matrix. It also addresses legal, privacy, infrastructure, and sustainability considerations. The framework was iteratively refined through structured expert consultation. The result is a practical and auditable guideline for responsible AI-assisted evidence synthesis.
1 Introduction
ARISMA addresses the gap between established review-reporting standards and the governance requirements introduced by AI-assisted evidence synthesis. It proposes bounded, validated, auditable human oversight rather than unconstrained automation.
- Evidence base: Evidence for AI-assisted review is mixed: GPT-4 screening studies reported 0.00 estimated false-exclusion rate in one test and 100% sensitivity with 65%–85% simulated workload reductions in another, while other tasks performed poorly.The mixed evidence does not justify permissive claims about autonomous reviewing.
- Reporting gap: Existing standards guide review, search, protocol, scoping, and synthesis reporting but do not specify when AI use is justified or how it should be validated and disclosed.The missing operational questions include task risk, benchmark validation, mandatory human sign-off, and reader-facing disclosure of AI involvement.
- Motivation: AI-assisted evidence synthesis can redistribute epistemic authority across reviewers, models, vendors, platforms, and infrastructures, affecting inclusion, visibility, privacy, accountability, and credibility.ARISMA frames this redistribution as an ethics-of-methods problem requiring transparent, accountable, privacy-respecting, and scientifically justified use.
- Governing principle: ARISMA treats AI as an inspected, bounded, and reversible assistant whose inputs, outputs, constraints, validation, and human accountability must be documented.Humans retain override power and responsibility for consequential decisions.
- Contributions: The framework contributes an end-to-end lifecycle taxonomy, stepwise review guidance, governance and provenance controls, reporting instruments, and validation logic.These instruments let editors, reviewers, and readers inspect where AI entered the process and whether its role remained scientifically legitimate.
- Implication: ARISMA seeks to make AI-assisted reviews more auditable, governable, and defensible rather than more automatic.The paper retains established review standards while adding AI-specific accountability and oversight requirements.
2 Development Basis and Expert-Informed Refinement of ARISMA
ARISMA was developed through literature-informed design and structured consultation with evidence-synthesis researchers and information specialists. Feedback expanded validation, documentation, revision, and human-oversight requirements while remaining formative rather than statistically representative.
- Development rationale: The authors combined established evidence-synthesis guidance, AI-assisted review literature, and structured expert consultation to develop ARISMA.The design began by identifying shared commitments such as transparent searching, explicit eligibility, reproducible selection, traceable extraction, and defensible synthesis.
- Structured expert consultation: Twenty-one researchers and information specialists participated in purposive 45–60-minute video consultations focused on relevant methodological expertise rather than statistical representation.Participants were recruited through recent methodological and AI-assisted evidence-synthesis publications and further recommendations.
- Feedback: Consultation feedback addressed benchmark validation, model-version and prompt logging, model instability, stopping rules, and human oversight for high-consequence tasks.The exercise was formative and intended to identify omissions and improve clarity and practical usability.
- Resulting refinements: The revised framework expanded calibration and pilot-testing requirements, added model metadata and prompt documentation, clarified when to revise or disable AI support, and strengthened human review responsibilities.These refinements were considered alongside published methodological and empirical literature, which remained ARISMA’s principal evidential basis.
- Consent and confidentiality: Experts provided informed consent, feedback was anonymized, identifying information was removed, and no identifiable quotations were reported.These procedures governed the use of consultation records during manuscript preparation.
3 ARISMA Framework and Operational Process
ARISMA organizes AI-assisted evidence synthesis across review types and lifecycle phases, using bounded intervention points, validation, provenance, and explicit human control. Its operational guidance covers retrieval, screening, extraction, categorization, language coverage, and error logging.
- Lifecycle framework: ARISMA organizes evidence synthesis into six linked phases spanning framing, protocolization, retrieval, selection and enrichment, evidence structuring, and synthesis and reporting.The framework accounts for different evidentiary products across scoping, mapping, and effect reviews.
- AI-use taxonomy: AI use is classified as assistive, adjudicative, or generative, with stricter validation and human oversight required as epistemic consequences increase.Assistive uses include ranking and tagging; adjudicative uses include eligibility judgment; generative uses include drafting and interpretation.
- Governance and control: ARISMA places AI in limited intervention zones whose legitimacy depends on mandatory human control, auditable artefacts, and reversible sign-off.The framework applies this control logic across the review lifecycle and retains explicit human accountability.
- Retrieval and enrichment: Search and enrichment are treated as one controlled subsystem combining versioned strategies, multi-source retrieval, deduplication, benchmark validation, and backward and forward snowballing.AI-generated search strings must be tested against benchmark studies and archived with all AI contributions; missed known studies constitute failure.
- Screening: Screening AI is calibrated against human-labeled pilots and assigned conditional trust tiers that determine whether it may rank records, recommend exclusions, or act only as a second reviewer.Production use requires explicit sensitivity thresholds; failure returns the process to revision and recalibration.
- Validation and auditability: ARISMA requires calibration before major tasks and after prompt or setting changes, while monitoring language coverage and logging model failures, amendments, and output shifts.The error ledger supports reconstruction of how AI affected the workflow and strengthens reproducibility for peer evaluation.
- Evidence structuring: Extraction guidance distinguishes relatively reliable textual retrieval from less reliable numerical extraction, especially for event counts and continuous-data elements.Reported accuracies for continuous-data elements were as low as 24% to 56%, depending on model and variable.
4 Governance, Evidence, and Provenance
ARISMA treats AI as a governed methodological intervention rather than a software convenience, requiring predefined validation, auditability, accountability, and oversight. Its evidence pipeline preserves human control through provenance layers, checkpoints, override rules, and verified evidence tables.
- Governance and validation logic: ARISMA requires teams to define an AI role, acceptable boundaries, validation criteria, decision permissions, and disabling conditions before deployment.The protocol-first cycle integrates validation, deployment, auditing, and reassessment throughout the review lifecycle.
- Evidence flow and provenance: AI-assisted evidence flow must record machine involvement alongside numerical record counts, including validation, human control, and provenance from sources to synthesis claims.ARISMA extends conventional evidence-flow diagrams with provenance layers, validation checkpoints, override rules, and audit traces.
- Evidence flow and provenance: Provenance records model recommendations, human overrides, AI-prepopulated fields, source confirmation, and whether synthesis prose derives from verified tables.The framework treats provenance as cumulative across deduplication, screening, extraction, and synthesis.
- Evidence flow and provenance: Descriptive fields may be AI-prepopulated subject to verification, whereas interpretive fields require human adjudication and critical quantitative fields require stronger confirmation.Field triage is based on consequences for later categorization and narrative synthesis.
- Legal, privacy, and infrastructure governance: AI-assisted review materials require explicit governance for lawfulness, minimization, access, retention, confidentiality, and infrastructure choice.ARISMA advises against uploading sensitive full texts or manuscripts to external services unless legally, contractually, and institutionally permissible.
5 Tools and Trends
ARISMA frames existing review tools as bounded assistants whose acceptability depends on task-specific validation and increasing human verification for higher-consequence decisions. The framework combines pilots, calibration, logging, human sign-off, and reversible use with efficiency-aware infrastructure guidance.
- Tool support: AI tools support search construction, translation, synonym expansion, deduplication, snowballing, screening, extraction, appraisal, and workflow management at different levels of maturity.Some tools improve transparency or reduce manual effort, while others remain organizational scaffolds rather than AI systems.
- Retrieval and snowballing: Deduplication and citation-chasing can reduce review burden, but ambiguous matches, seed-set completeness, and infrequently cited studies require human checking.Citation-chasing should follow benchmarking against sentinel studies.
- Screening and prioritisation: Screening evaluations reported Work Saved over Sampling at 95% recall values of approximately 49% to 87% under benchmark conditions.The framework still requires calibration, stopping criteria, and continued human verification to reduce false-negative exclusions.
- Extraction and appraisal: Extraction and appraisal systems show mixed, item-dependent performance and are most defensible for assistive workflow support rather than synthesis-critical values or final judgments.Final appraisal judgments and critical extracted values require explicit human verification.
- Governance principles: ARISMA requires labeled pilots, sensitivity and precision thresholds, complete action logs, and experienced human sign-off for high-consequence tasks.AI may prioritize work but may not make final eligibility, risk-of-bias, or synthesis judgments.
- Infrastructure and sustainability: Compute escalation is discouraged unless additional computation measurably improves validation or reduces human burden without increasing epistemic risk.Large reasoning workflows, repeated sampling, and long prompts can increase inference cost without proportionate methodological gain.
6 Reporting and Validation Checklist
ARISMA converts its governance principles into submission-ready reporting and validation instruments. The checklist makes AI use inspectable, while the validation matrix ties permission to step-specific evidence, audit trails, and revocable thresholds.
- Checklist and validation matrix: Tables 1 and 2 operationalize ARISMA by specifying what AI use must be disclosed and what must be validated before influencing a review.Together they turn the framework into a reporting and validation instrument.
- ARISMA item checklist: Table 1 requires reporting the system, inputs, task, restrictions, validation evidence, and human override rule rather than merely naming automation.This complements PRISMA-style transparency for search, eligibility, screening, and synthesis methods.
- ARISMA item checklist: AI reporting should include prompts, model metadata, validation scripts, audit logs, availability information, and vendor or licensing disclosures when legally and ethically permitted.The checklist also addresses data, code, extraction forms, appendices, and governance constraints.
- Validation matrix: Passing a validation threshold permits proceeding but does not replace routine human supervision, and permissions must be recalibrated when models, prompts, or corpora change.The matrix treats authorization as revocable rather than permanent.
- Validation matrix: Table 2 links each review step to potential failures, validation requirements, and the minimum audit trail retained after completion.Examples include known-study recall for search-term generation, false-positive and false-negative balance for deduplication, and sensitivity for screening.
- Validation matrix: Editors and peer reviewers can use the validation matrix to assess whether methodological claims are supported by validation artefacts and human scientific control.This makes ARISMA useful during submission and peer review as well as during review conduct.
7 Threats to Validity
ARISMA identifies threats arising from poor review design, incomplete retrieval, context-sensitive screening, extraction and appraisal errors, generative overreach, reproducibility limits, underreporting, and limited framework evaluation. These threats constrain AI use and reinforce validation, disclosure, and human responsibility.
- Design threat: AI cannot correct a poorly chosen review design and may make an inappropriate product appear more polished or precise.A focused causal question and a conceptually immature field may require different review designs.
- Retrieval threat: LLM-assisted query generation may miss unusual relevant studies, while citation-chasing yields depend on database coverage and the seed set.ARISMA therefore requires benchmark sets, PRESS-style review, and complementary retrieval methods.
- Screening threat: Screening performance may not transfer across domains, languages, prevalence structures, or eligibility regimes because encouraging studies used constrained, benchmarked conditions.Teams should report language distributions before and after AI-influenced filtering.
- Extraction and appraisal threat: Extraction and appraisal errors can involve numeric inaccuracies, misread tables, arm or time-point confusion, inferred values, and judgments requiring domain knowledge.Current evidence supports assistance rather than abdication.
- Generative reasoning threat: LLM-generated synthesis can omit qualifiers, smooth contradictions, and overstate what source evidence supports, making discussion text a threat surface.The risk is epistemic drift rather than merely formatting error.
- Reproducibility threat: Closed-model updates and hidden interface details limit exact replay even when teams log model, date, prompt, and settings.AI-assisted reproducibility therefore requires extensive provenance without guaranteeing identical outputs.
- Reporting and ethics threat: Underreporting AI use prevents readers from assessing automation, interpretation, confidentiality, copyright, governance, and responsibility.The review team, not the AI system, retains authorship responsibility.
- Framework-development limitation: The purposive expert consultation improved ARISMA’s clarity but did not establish formal consensus or demonstrate effectiveness across disciplines and settings.Broader prospective evaluation remains necessary.
8 Declarations
The expert-informed refinement of ARISMA involved voluntary consultation with experienced researchers and information specialists, with informed consent and anonymized feedback. Manuscript preparation and all scientific decisions remained under human supervision, verification, and authorship responsibility.
- Consultation involved researchers and information specialists experienced in evidence synthesis, who participated voluntarily and provided informed consent.
- Consultees consented to the use of anonymized, non-identifiable feedback, and no identifiable participant information is reported.
- Any AI-assisted manuscript support remained under full human supervision and verification.
- Authors reviewed, verified, and approved all scientific judgments, methodological decisions, interpretations, validations, and final manuscript content.
- No AI system is listed as an author or assumes responsibility for the manuscript.
9 Conclusion
ARISMA frames AI-assisted evidence synthesis as a governance problem requiring validation, provenance, transparency, and human accountability. It permits bounded assistance under validated conditions while rejecting unconstrained automation of consequential review decisions.
- ARISMA treats AI as a governance issue rather than merely a productivity tool for evidence synthesis.
- The framework permits bounded AI assistance for retrieval, prioritisation, screening, extraction, categorization, and drafting under constrained and validated conditions.
- AI performance remains task-dependent, unstable across contexts, and insufficient to justify unconstrained automation of consequential review decisions.
- ARISMA asks whether AI’s influence on evidence selection, interpretation, and synthesis remains transparent, auditable, and scientifically controllable.