Source-linked AI summary

LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs

Muhammad Ali Chaudhry, Xinyuan Hao, Haifa Alwahaby

arXiv:2608.26145v1cs.AIcs.HCcs.IR

TL;DR

The paper examines how short and long LLM context windows affect AI-generated literature reviews and how humans should oversee their use. It evaluates twenty reviews generated from Semantic Scholar and arXiv sources, finding that larger contexts broaden information and coherence while increasing repetition, omissions, and weak synthesis. The results support AI as a drafting aid rather than a fully automated replacement for expert review.

  • Problem

    The study investigates the impact of LLM context-window size on literature-review quality and the role of human oversight in using AI-generated academic reviews.

  • Method

    The researchers generated twenty AI-in-education literature reviews from Semantic Scholar and arXiv sources, using Gemini 1.5 Pro with long or short context windows and evaluating them across fifteen dimensions.

  • Results

    Larger context windows incorporated broader information and maintained coherence across extended inputs but increased repetition, omission of critical works, and lack of analytical synthesis, while smaller windows reduced redundancy but limited comprehensiveness.

  • Takeaways & Limitations

    AI-generated reviews can provide foundational overviews and initial drafts, but human expertise is needed to refine them to academic scholarship standards.

  • Takeaways & Limitations

    The study evaluated Gemini 1.5 Pro reviews on AI and education using two open-access databases and two domain experts, so quality may vary across models, domains, databases, and evaluators.

Abstract

from arXiv · show

Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs) to investigate the impact of context window on the quality of AI-generated literature reviews and the role of AI in supporting literature review writing. Twenty AI-generated literature reviews based on research sources from Semantic Scholar and Arxiv were evaluated by two researchers across 15 dimensions. Our findings reveal that AI-generated literature reviews require human oversight to meet academic publishing standards. As context windows increase, LLMs can incorporate broader information and maintain coherence across longer inputs, but they also exacerbate issues such as content repetition, omission of critical work, and a tendency towards descriptiveness over synthesis. Our work shows that AI-generated reviews can provide foundational overviews, but their output must be critically evaluated and refined by domain experts. Future research should consider integrating other LLMs and fine-tuned models in different domains with hybrid approaches that combine human expertise with AI capabilities to address the limitations identified in this study.

1. Introduction

The paper examines how context-window size affects LLM-generated literature reviews and the responsibilities of humans who evaluate them. Larger windows broaden information processing and support coherence across extended inputs, while expanding the practical use of LLMs for long-form academic review tasks.

  • Context windows determine how much sequential input an LLM can consider during inference, affecting its ability to understand and maintain coherence across extended text.
  • Transformer attention mechanisms historically constrained context length because their quadratic scaling imposed high computational costs.
  • Larger context windows broaden applications including long-form generation, summarization, multi-document processing, retrieval, knowledge integration, and time-efficient text analysis.
  • The research compares long-context literature reviews with retrieval augmented generation, citing evidence that long context generally outperforms RAG for document-based question answering.
  • The study investigates the quality of AI-generated academic literature reviews and the human-in-the-loop responsibilities involved before their use.

2. Literature Review

The literature review situates LLMs within established academic review practices and prior research on AI-assisted searching, synthesis, and evaluation. It emphasizes that systematic reviews require structured critical work, while LLM tools provide support across several component tasks.

  • Academic literature reviews require identifying gaps, defining questions, searching and evaluating sources, organizing evidence, synthesizing patterns, writing clearly, and citing appropriately.
  • Prior research reports that LLMs can summarize multiple abstracts and cluster research articles thematically, supporting literature-review workflows.
  • LLM-based reviewing systems can reduce inflated scores, overconfident ratings, and skewed score distributions by incorporating review forms, guides, and ethical documents.
  • Semantic search retrieves papers according to contextual meaning rather than only keyword matches, which is useful for interdisciplinary research.
  • Researchers’ trust in LLM use varies with usage time and frequency, while academic AI-tool providers encourage human oversight before publication.

3. Methodology

The study compares twenty AI-generated literature reviews produced with long and short context windows, using consistent source and prompting procedures. Two AI-education researchers scored the reviews across a multidimensional rubric with an inter-rater reliability check.

  • Twenty literature reviews addressing ten AI-in-education questions were generated, split evenly between long-context and short-context conditions.
  • Two AI-in-education researchers rated twenty reviews, with each comparing paired versions of five research questions.
  • Weighted Cohen’s Kappa = 0.743 indicated substantial agreement on a parallel 10% sample, with disagreements resolved through discussion.
  • Each review was evaluated across 15 dimensions scored from 1 to 5, producing a maximum total of 75.
  • Sources came from Semantic Scholar and arXiv, while Gemini 1.5 Pro generated both formats using one common prompt and differing context-window conditions.
  • The rubric prioritised coverage of major themes and subthemes, structure, organization, and progression from general overview to specific focus.

4. Findings

Both context settings produced generally strong reviews, but longer contexts were associated with more repetition and descriptiveness, alongside broader citation coverage. Reviews also omitted important works and showed stylistic and grammatical weaknesses, including irrelevant content.

  • 4. Findings: All AI-generated literature reviews reached at least the good rating, indicating potential as supporting tools for literature-review writing.
  • 4. Findings: Both long- and short-context systems performed well, but the reviews showed limitations and differences potentially related to context-window size.
  • 4.1 Repetition of Content: Larger-context reviews repeated themes, concepts, and verbatim content more frequently, although they tended to cite more relevant prior literature.
  • 4.2 Omission of Key Works: Some reviews omitted essential, highly cited works, including Wayne Holmes and Ryan Baker in one review on educational AI ethics and algorithmic bias.
  • 4.3 Descriptive Nature: Longer reviews were often descriptive summaries rather than analytical syntheses, adding articles without substantive critical evaluation or meaningful connections between sources.
  • 4.4 Stylistic and Grammatical Issues: Reviews also contained repetitive transitions, inconsistent tense, irrelevant content, and broader AI material unrelated to the specific review focus.

5. Discussions

The findings show that LLM-generated literature reviews balance broader coverage against repetition, fragmentation, and limited analytical depth. Human oversight remains necessary to refine content, citations, and style for academic use.

  • Analytical depth: AI-generated reviews tend to prioritize length and breadth over quality, depth, and critical synthesis.Models can summarize individual papers effectively but struggle to combine sources into coherent analytical frameworks.
  • Context-window trade-offs: Larger context windows incorporate more information and maintain clearer coherence but can increase redundancy through ineffective prioritization and structuring.Smaller windows mitigate repetition by constraining scope, but may produce fragmented narratives and omit necessary details.
  • Citation selection: LLMs may omit influential or seminal studies because they rely on training data and do not independently discern the relative importance of cited works.The black-box nature of citation choices further obscures why particular sources are selected.
  • Human oversight: All evaluated reviews received at least a good mark from human experts, indicating potential as drafting tools when paired with human-in-the-loop revisions.Human editing can address redundancy, omissions, stylistic inconsistencies, and insufficient analytical depth.
  • Human oversight: Repetitive transitions and inconsistent tense usage reinforce the importance of human editing and reviewing.The authors also note that prompting formality may influence output readability and that human writers can make similar inconsistencies.

6. Limitations and Future Work

The study presents itself as an early examination of LLM-generated literature reviews and identifies model, domain, expert, database, and retrieval choices as boundaries for future work.

  • Scope: This research is an early step in examining the efficacy and limitations of LLM-generated academic literature reviews.The authors frame the work as being in an early stage of development.
  • Models: The evaluation used only Gemini 1.5 Pro, leaving other closed- and open-source models and academically fine-tuned models for future investigation.The proposed fine-tuning target is an academic corpus containing millions of open-access papers.
  • Data and evaluation: The reviews focused on AI and education and were assessed by two domain experts, so quality may vary across domains and evaluators.The authors link variation across domains to publication quality and LLM analysis.
  • Data and evaluation: Academic database selection, search quality through APIs, and the number of relevant papers retrieved can affect review quality and length.These factors constrain how broadly the study’s findings should be generalized.

7. Conclusion

The conclusion finds that context size creates a trade-off: larger windows broaden information coverage and coherence while increasing repetition, omissions, and insufficient synthesis. AI-generated reviews can support academic drafting, but human expertise is needed to achieve scholarly rigor.

  • Conclusion: Larger context windows broaden information coverage and maintain coherence across extended inputs, but introduce repetition, omissions of critical works, and weak analytical synthesis.Smaller windows reduce redundancy but often lack the comprehensiveness required for academic rigor.
  • Conclusion: AI-generated reviews also show descriptive and stylistic weaknesses, including repetitive transitions, inconsistent tense usage, omitted seminal works, and irrelevant content.Despite these shortcomings, every review achieved at least a good rating.
  • Human expertise: Human-in-the-loop approaches are essential for addressing redundancy, stylistic inconsistencies, and content gaps, particularly for non-expert users.Domain knowledge is needed to evaluate generated literature critically.
  • Future work: Future work should examine advanced and fine-tuned domain-specific models, different databases, hybrid approaches, and user-friendly tools.The authors position these directions as ways to improve the literature-review process.
  • Human expertise: LLMs should complement rather than replace human judgment and critical thinking in academic research.The conclusion limits their role to supporting scholarly work alongside human expertise.

8. Impact Statement

The study highlights both the potential and limitations of LLMs for academic literature reviews. Comparing short and long context windows supports a hybrid approach that combines AI efficiency with human expertise to preserve rigor and reliability.

  • Impact Statement: Comparing short and long context windows shows advantages in handling extensive information and maintaining coherence, alongside redundancy, shallow analysis, and omissions of key works.The study advocates combining AI efficiency with human expertise and reports that its code and data are publicly available.

Appendix

The appendix reports varied performance across organization, introduction coverage, synthesis, critical evaluation, relevance, and academic tone. Reviews range from logically structured and well connected to poorly organized, summary-driven, and incompletely developed.

  • Organization: Logical organization ranged from clear structures with transitions to reviews lacking connections between sections and ideas.The appendix includes both organized reviews with clear transitions and reviews described as not logically organized.
  • Introduction: Introduction coverage ranged from complete background, significance, and structure to reviews missing one, two, or all key elements.Some reviews contained all introduction elements, while others omitted varying numbers of required elements.
  • Synthesis: Synthesis varied from interconnected studies supporting ideas to reviews consisting mostly or entirely of descriptive summaries.Several assessments noted weak or absent connections among studies, although some reviews used studies to support a paragraph’s main idea.
  • Analysis: Critical evaluation ranged from strong and appropriately integrated assessment to no evaluation or evaluation that was not well integrated.The appendix records intermediate cases in which evaluation was present but incompletely integrated.
  • Conclusion: Conclusions ranged from synthesizing all ideas with a clear research gap to addressing only some ideas before closing the review.Other assessments described satisfying or logical closings that synthesized most ideas or generally addressed the review’s main ideas.
  • Relevance and tone: Relevance and academic tone were generally positive but imperfect, with assessments reporting at least 80% relevant content and occasional tone lapses.The appendix records content relevance at >= 80% and >= 90%, alongside generally academic tone with some lapses.

Presentation

The presentation evaluation considers grammar and conjunction variety. Reviews ranged from readable despite a few grammar errors to reviews with too many or pervasive errors, while conjunction use varied from repetitive to more diverse.

  • Grammar: Some reviews contained no more than two grammar errors, whereas others had too many or were riddled with grammatical errors.
  • Conjunctions: Conjunction use ranged from five or fewer similar conjunctions to seven or more varied conjunctions.
  • Conjunctions: Some reviews used less than seven conjunctions with variation, while another category contained zero to two conjunctions.
Loading 2608.26145v1…