Source-linked AI summary
Artificial Intelligence for Literature Reviews: Opportunities and Challenges
Francisco Bolanos, Angelo Salatino, Francesco Osborne, Enrico Motta
TL;DR
SLRs are time-consuming and prior surveys provide only limited coverage of AI features, motivating a broader assessment. This paper evaluates 21 SLR tools across 23 general and 11 AI-specific features and examines 11 LLM-based applications. Existing tools can be powerful but often lack usability, while newer LLM-based tools remain immature and face hallucination challenges.
Problem
SLRs are time-consuming, while previous surveys provide limited analysis of the expanding AI-enhanced SLR-tool ecosystem.
Method
The survey evaluates 21 SLR tools across 23 general and 11 AI-specific features and additionally analyses 11 LLM-based applications.
Results
Existing SLR tools can be powerful when used effectively, but usability and user-friendliness limit broader adoption; LLM-based tools remain in their infancy.
Takeaways & Limitations
Future SLR tools should pursue knowledge injection and retrieval-augmented generation to support robust and verifiable information.
Takeaways & Limitations
The review excluded inaccessible tools and tools not updated in the past 10 years, potentially restricting generalisability.
Abstract
from arXiv · showhide
This manuscript presents a comprehensive review of the use of Artificial Intelligence (AI) in Systematic Literature Reviews (SLRs). A SLR is a rigorous and organised methodology that assesses and integrates previous research on a given topic. Numerous tools have been developed to assist and partially automate the SLR process. The increasing role of AI in this field shows great potential in providing more effective support for researchers, moving towards the semi-automatic creation of literature reviews. Our study focuses on how AI techniques are applied in the semi-automation of SLRs, specifically in the screening and extraction phases. We examine 21 leading SLR tools using a framework that combines 23 traditional features with 11 AI features. We also analyse 11 recent tools that leverage large language models for searching the literature and assisting academic writing. Finally, the paper discusses current trends in the field, outlines key research challenges, and suggests directions for future research.
1 Introduction
Systematic literature reviews rigorously integrate prior research but are time-consuming and resource-intensive. This survey examines how AI can semi-automate screening and extraction while assessing current tools, LLM-based applications, challenges, and future directions.
- SLRs rigorously identify and appraise relevant literature for specific research questions while following protocols intended to minimise bias.
- SLRs can extend beyond a year and require domain experts, paid resources, and periodic updates as publication volumes grow.
- AI-enhanced tools primarily target screening and data extraction, with advances in NLP and LLMs creating potential for further automation.
- The survey evaluates 21 prominent SLR tools using 23 general and 11 AI-specific features, focusing on semi-automation in screening and extraction.
- The paper discusses AI integration, usability, and standardised evaluation as major research challenges and proposes evaluation best practices.
- The study also analyses 11 LLM-based applications for literature searching and academic writing as potential sources of features for future SLR tools.
2 Background
SLRs proceed through six stages, from planning and search to screening, extraction, quality assessment, and reporting. Current AI support is concentrated on task-specific assistance, especially screening and extraction, while newer approaches extend toward natural-language search and reporting.
- Current SLR tools use weak or narrow AI for task-specific activities such as classification, clustering, and named-entity recognition.
- The SLR methodology comprises six stages: planning, search, screening, data extraction and synthesis, quality assessment, and reporting.
- Planning: Planning establishes precise research questions and a protocol that supports consistency, bias reduction, transparency, and reproducibility.
- Search: Search identifies relevant papers through query-based strategies, snowballing, or hybrid approaches combining both methods.
- Screening: Screening applies inclusion and exclusion criteria first to titles and abstracts, then to full texts for more thorough assessment.
- Data Extraction and Synthesis: Data extraction and synthesis systematically collect information from selected studies, with techniques varying by research field and objective.
- Quality Assessment: Quality assessment evaluates the rigour and validity of selected studies, informing the review’s overall strength and trustworthiness.
- Reporting: Reporting presents findings in a structured paper format, while recent LLM-based systems begin extending AI support into literature assistance and reporting-related work.
3 Methodology
The survey uses PRISMA-informed, multi-source procedures to identify and evaluate AI-enhanced SLR tools. It applies explicit inclusion and exclusion criteria before analysing a final set of 21 tools across their functions and accessibility.
- The review adopts PRISMA methodology to conduct and report the systematic review and meta-analysis.
- Selection Criteria: Tools had to use AI for semi-automating screening or extraction while preserving the user’s final decision-making capacity.
- Selection Criteria: Additional inclusion criteria required a user interface and installation and execution without advanced technical expertise.
- Selection Criteria: The review excluded tools under maintenance or without updates during the previous 10 years.
- Sources and Search: Tool identification combined prior surveys, the SLR Toolbox, CRAN, and snowballing through Semantic Scholar.
- Tool Selection: The process produced 21 analysed tools after deduplication, exclusions, and consolidation of RobotReviewer with RobotSearch.
- Tool Overview: Nineteen tools support screening and four support extraction; 17 primarily target screening, while two focus exclusively on extraction.
- Tool Overview: Most tools are web applications, but only four release their code under an open license.
4 Meta-review of previous surveys
Earlier surveys examined AI in SLR tools through a small set of features and covered the ecosystem unevenly. This survey identifies that limitation and responds with 11 AI features applied to 21 tools.
- The meta-review examines four previous surveys that analysed AI features in SLR tools.
- Features Examined: Earlier surveys considered approach, text representation, human interaction, input, and output as their principal AI-related features.
- Features Examined: Approach was the most examined feature, whereas human interaction and input received the least attention among prior surveys.
- Coverage Limitations: Prior analyses varied substantially in breadth: one survey used all five features for seven tools, while another considered only approach.
- Coverage Limitations: The five reported features provide only a narrow perspective on how AI can support SLRs.
- Survey Contribution: The survey addresses this gap by introducing 11 AI features and applying them to 21 identified SLR tools.
5 Survey of SLR Tools
The survey examines how AI supports SLR screening and extraction, comparing 21 tools across AI and general features. Screening tools vary in tasks, input representations, model execution, interfaces, and researcher support.
- Evaluation framework: The framework evaluates SLR tasks, minimum training requirements, model execution, research fields, and other tool characteristics.The survey groups features to compare how AI systems support screening and extraction.
- AI-supported screening: 19 tools use AI for screening, with 15 performing one task and four performing two AI-related tasks.Most single-task systems classify papers as relevant or irrelevant; the four two-task systems add another classification or categorisation function.
- AI approaches: Screening tools use traditional classifiers and evolving text representations, ranging from SVM and Bag of Words to word and sentence embeddings.SVM is the most adopted classifier; most tools use titles and abstracts, while Colandr requires full text and uses word2vec and GloVe embeddings.
- Human interaction: Interfaces support binary classification, user-defined categories, topic maps, keyword filtering, clustering, and PICO identification.Sixteen tools provide graphical interfaces for relevance classification, while Colandr enables flexible category assignment and Iris.ai supports topic-map navigation.
- Minimum requirements: Most screening classifiers require 1–15 relevant and irrelevant seed papers, while Colandr requires 10 and SysRev requires 30.ASReview, SWIFT-Active Screener, and SWIFT-Review can begin with one relevant and one irrelevant paper.
- Model execution: Thirteen tools execute models in real time, whereas SysRev and SWIFT-Active Screener use delayed execution at predetermined intervals.SysRev runs operations overnight, while SWIFT-Active Screener updates its model after each specified interval.
5.3 Outstanding SLR tools
The survey finds substantial development among SLR tools, with functionality features becoming standard. Tool strengths differ across research fields and article-selection needs.
- SLR tools have advanced significantly, particularly through standard functionality for project tracking, auditing, multiple users, and reference management.
- Non-biomedical fields: ASReviewer offers multiple classifiers for selecting relevant articles in non-biomedical research, including Logistic Regression, Random Forest, Naive Bayes, and Neural Networks.
- Non-biomedical fields: Iris.ai and Colandr provide flexibility through semantic document clustering and user-defined categories, respectively.
- Biomedical field: Covidence, PICOPortal, EPPI-Reviewer, RobotReviewer/RobotSearch, and Rayyan are identified as reliable tools for biomedical research.
- Biomedical field: Covidence, PICOPortal, and EPPI-Reviewer identify Randomised Controlled Trials using predefined classification models.
5.4 Threats to validity
The survey addresses validity threats through systematic procedures, multiple tool sources, feature development, collaborative extraction, and supplementary checks. Remaining limitations concern coverage, construct completeness, software evolution, and possible replication variation.
- Internal Validity: The review used a prespecified protocol, PRISMA guidelines, multiple repositories, manual survey searches, snowballing, and staged selection to support internal validity.
- External Validity: Search-engine choices and search-string formulation may have reduced the completeness of tool identification, limiting generalisability across domains.
- External Validity: Excluding tools without user interfaces, unavailable tools, or tools not updated for ten years may have omitted otherwise eligible systems and restricted generalisability.
- Construct Validity: The 34-feature framework may omit relevant characteristics despite adding 11 AI-specific features to address gaps from previous studies.
- Conclusion Validity: Collaborative extraction checks supported consistency, but evolving software and rapidly emerging Generative AI and LLM features may make findings change over time.
6 Research Challenges
The paper identifies research challenges for AI-enhanced SLR tools involving advanced NLP, interpretability, semantic technologies, social impact, usability, evaluation, transparency, and trustworthiness. It proposes best practices centred on performance, usability, and transparency while noting that comprehensive evaluation frameworks remain beyond the paper’s scope.
- AI techniques: Current SLR tools often rely on outdated classifiers and bag-of-words representations, motivating research into advanced NLP technologies such as LLMs.Recent tools have begun adopting word and sentence embeddings, while LLM integration introduces additional challenges because models are trained on general data.
- Interpretability: Interpretability remains a challenge because screening classifiers typically operate as black boxes; fact-checking, argument mining, and prompting techniques are proposed as possible mechanisms.The paper links improved explanations with greater reliability and credibility of screening tools.
- Semantic technologies: Knowledge graphs are proposed as semantic technologies for enhancing the characterisation and classification of research papers.They represent entities and relationships using machine-readable information organised according to domain ontologies.
- Social impact: Incorrect or biased outputs can affect authors, readers, and policy-relevant populations, making transparency, inspection, interpretation, override mechanisms, validation, and correction important safeguards.The paper warns that inaccurate scientific information may enter research papers and influence policy development.
- Usability: Usability is a critical challenge: among 81 surveyed researchers, poor usability accounted for 43% of tool discontinuations, while insufficient functionality and workflow incompatibility each accounted for 37%.Evaluated tools had SUS scores from 66 to 77, corresponding to satisfactory but not outstanding usability.
- Evaluation and best practices: Tool evaluations lack standard frameworks and comparable benchmarks, and should assess performance, usability, and trustworthiness rather than performance alone.The paper recommends replicable performance studies, comprehensive usability assessment, and transparency through disclosures, shared models or knowledge bases, and explainable AI.
7 Emerging AI Tools for Literature Review
The survey examines 11 emerging LLM-based tools that support literature searching and academic writing, although they do not directly implement specific SLR phases. These systems mainly divide into search engines and writing assistants, with proprietary implementations and variable output quality limiting current reliability.
- Tool landscape: 11 tools were selected to analyse emerging systems for retrieving research papers and supporting academic writing.The researchers screened 164 tools identified through TopAI Tools and selected 11 after a two-stage process.
- Tool categories: Seven tools were classified as search engines and three as writing assistants, while Textero.ai fit both categories.
- Tool categories: Search engines accept natural-language queries and return related papers with summaries, whereas writing assistants generate text from document descriptions for iterative refinement.Jenni.ai exemplifies interactive, real-time collaborative editing between users and AI.
- Search engines: Tool coverage differs by bibliographic source and document access: Scite and Consensus process full texts, while Elicit and EvidenceHunt use titles and abstracts.Sources include Semantic Scholar, PubMed, and publisher collections, but some databases are undocumented.
- Search engines: Six of eight search-oriented tools are cross-disciplinary; EvidenceHunt targets biomedicine, while Elicit covers biomedicine and social sciences.
- Implementation and quality: Most systems appear to use OpenAI APIs, often with RAG, but proprietary implementations and fine-tuning practices remain undisclosed.Elicit and Perplexity explicitly report using OpenAI GPT technology; whether writing tools are fine-tuned remains unclear.
- Implementation and quality: Generated-text quality varies substantially, and the surveyed writing tools may currently suit brief student essays better than research writing.
8 Conclusion
The survey evaluates 21 SLR tools across general and AI-specific features and examines 11 LLM-based applications for literature retrieval and writing. It finds powerful but usability-limited existing tools, while newer LLM systems remain promising yet immature because of hallucination risks.
- Scope and contribution: 21 SLR tools were evaluated across 23 general features and 11 AI-specific features, alongside 11 LLM-based research applications.
- Scope and contribution: The survey identifies strengths, weaknesses, suitable use cases, research challenges, and emerging opportunities in AI-enhanced SLR tools.
- Findings: Existing SLR tools can be highly powerful when used effectively, but limited usability and user-friendliness constrain broader adoption.
- Findings: LLM-based tools are developing rapidly but remain in their infancy and face the documented problem of hallucinations.
- Future direction: Knowledge injection and retrieval-augmented generation strategies are identified as needed for robust and verifiable information generation.
- Future direction: The survey anticipates AI-enabled research assistants that could generate literature reviews, identify hypotheses, and foster innovation within five years.
Appendix A Systematic Literature Review Tools analysed through AI and Generic Features
The appendix provides detailed tables describing the 21 analysed SLR tools across generic and AI-based features. Because of space constraints, only summarized tables appear in the paper, with full versions available online.
- Appendix contents: Three appendix tables describe 21 SLR tools using generic and AI-based features for screening and extraction.A third table reports the tools’ generic features.
- Appendix contents: Only summarized tables are included because of space constraints; full versions are available on GitHub and the Open Research Knowledge Graph.
A.1 Screening Phase of Systematic Literature Review Tools analysed through AI Features
The screening-phase appendix records AI features, tasks, text representations, inputs, and minimum requirements for the analysed tools. Its entries include classification and clustering systems using bag-of-words or embedding representations, with requirements varying from none to specified relevant and irrelevant papers.
- Feature schema: The appendix organizes screening tools by field, SLR task, text representation, input, and minimum requirement.
- Tool entries: Abstrackr, Covidence, DistillerSR, FAST2, LitSuggest, Nested Knowledge, Rayyan, Research Screener, RobotAnalyst, ASReview, Screener, SWIFT-Review, and SysRev.com are listed with screening-task entries.
- Representations: Screening representations include bag of words, ngrams, embeddings, GloVe, SentenceBERT, doc2vec, and SciBERT.
- Minimum requirements: Reported minimum requirements range from no stated requirement to specified relevant and irrelevant papers, including 30 relevant papers for SWIFT-Review.
- Reporting completeness: Some appendix entries report unavailable or not-applicable information for tasks, requirements, or representations.
A.2 Extraction Phase of Systematic Literature Review Tools analysed through AI Features
The extraction-phase approaches include machine-learning classifiers, biomedical named-entity recognition, rule-based detection, and entity linking. Some implementations use SVMs, BI-LSTM-CRF architectures, convolutional neural networks, n-grams, and manually annotated sentences, while one implementation is explicitly unknown.
- CNN-based models are trained on manually annotated sentences stating the level of risk.
- Classifiers use SVMs, neural networks based on bidirectional LSTM with a Conditional Random Field, and models combining a linear model with a CNN.
- Extraction tasks include identifying control-related sentences, biomedical entities, CONSORT categories, risks, and animal-study entities.
- Rule-based detection identifies the 21 CONSORT categories in a second task.
- One approach uses a knowledge graph to represent relations among entities in papers or tables, but its technical implementation is unknown.
A.3 Systematic Literature Review Tools analysed based on General Features
The general-features analysis lists tools with differing support for search sources, multi-user workflows, and reference labelling or comments. The listed tools include both single- and multi-user systems, with sources ranging from none to PubMed and other repositories.
- Search-source entries range from none or NA to PubMed, CORE, Europe PMC, DOAJ, ClinicalTrials.gov, CORDIS, and the US Patent Office.
- The table includes tools marked for reference labelling and comments, including RobotAnalyst, PICO Portal, DistillerSR, EPPI-Reviewer, and Knowledge.
- Several tools are listed with no search source, including SWIFT-Review, FAST2, ASReview, Screener, Dextr, ExaCT, Abstrackr, and Colandr.