Source-linked AI summary
Scientific production in the era of Large Language Models
Keigo Kusumegi, Xinyu Yang, Paul Ginsparg, Mathijs de Vaan, Toby Stuart, Yian Yin
TL;DR
The paper examines how LLM adoption is changing scientific production and whether polished, complex writing still signals research quality. Across large-scale analyses, LLM use increases manuscript output, reverses the link between writing complexity and quality, and broadens cited prior literature, challenging traditional evaluation practices.
Problem
The paper asks whether LLM-generated polished writing reveals or conceals research quality and how LLMs affect discovery of prior literature.
Method
The study analyzes large-scale preprints, peer reviews, accesses, and citations using writing-complexity measures, publication outcomes, and event-study comparisons around LLM adoption.
Results
LLM use reverses the positive relationship between writing complexity and scientific merit while steering authors toward more diverse, younger, and less impactful cited work.
Takeaways & Limitations
Traditional language-based quality signals are becoming unreliable as scientific output rises, requiring stronger quality assessment and methodological scrutiny.
Takeaways & Limitations
The study does not provide causal identification because LLM use is imperfectly measured and adoption is non-random.
Abstract
from arXiv · showhide
Large Language Models (LLMs) are rapidly reshaping scientific research. We analyze these changes in multiple, large-scale datasets with 2.1M preprints, 28K peer review reports, and 246M online accesses to scientific documents. We find: 1) scientists adopting LLMs to draft manuscripts demonstrate a large increase in paper production, ranging from 23.7-89.3% depending on scientific field and author background, 2) LLM use has reversed the relationship between writing complexity and paper quality, leading to an influx of manuscripts that are linguistically complex but substantively underwhelming, and 3) LLM adopters access and cite more diverse prior work, including books and younger, less-cited documents. These findings highlight a stunning shift in scientific production that will likely require a change in how journals, funding agencies, and tenure committees evaluate scientific works.
LLM use, scientific writing, and publication outcomes
LLM assistance produces more complex scientific language while reversing its relationship with publication success, making linguistic complexity an unreliable quality signal. It also broadens literature discovery and citation toward books, younger work, and less-cited scholarship.
- Writing complexity and publication outcomes: LLM-assisted manuscripts have significantly higher writing-complexity scores than natural-language papers across all three archives.The difference is significant at P < 0.001 using two-tailed t-tests in all repositories.
- Writing complexity and publication outcomes: For LLM-assisted papers, greater writing complexity correlates negatively with publication success, reversing the positive relationship observed in human-written papers.The reversal is replicated for lexical complexity and morphological complexity.
- Literature discovery and citation behavior: 26.3% higher access to books occurs among Bing users, while LLM-referred visits reach manuscripts with median ages estimated at 0.18 years younger.Book access is significant at P < 0.001; LLM users do not increase citations to well-cited works.
- Literature discovery and citation behavior: 11.9% greater likelihood of citing books accompanies citations to documents 0.379 years younger and with 2.34% lower citation impact among LLM adopters.The increased likelihood of citing books is not statistically significant in SSRN.
- Implications: LLMs shift search and citation toward a more diverse knowledge base, including books and younger, less-cited scholarship, while obscuring authorial-effort signals.The study cautions that language characteristics are becoming uninformative signals for reviewers and editors as scientific communication increases.
Supplementary Materials
The supplementary materials track monthly preprint productivity and LLM adoption among 168,553 authors across arXiv, bioRxiv, and SSRN from January 2022 through July 2024. A text-based detector identifies whether and when each author adopted LLMs in scientific writing.
- Study design: 168,553 authors were tracked between January 2022 and July 2024 across three scientific preprint platforms.The dataset includes 109,965 arXiv authors, 43,218 bioRxiv authors, and 15,370 SSRN authors.
- Study design: Monthly scientific productivity was measured by the number of preprints published by each author.The analysis tracks productivity dynamics over time.
- LLM adoption measurement: A text-based detector determined whether and when authors adopted LLMs in scientific writing.The detector was applied to authors’ preprints.
A B C
LLM-associated productivity gains vary across races, ethnicities, and home geographies, while post-event users access more books and recent, less highly cited works.
- Heterogeneity across races, ethnicities and home geographies: 43.0% (arXiv), 70.9% (bioRxiv), and 70.1% (SSRN) productivity gains occurred for authors with East Asian names.Gains remained meaningful for authors with Caucasian names: 25.7% (arXiv), 41.5% (bioRxiv), and 46.5% (SSRN).
- Heterogeneity across races, ethnicities and home geographies: 37.7% (arXiv) and 89.3% (bioRxiv) productivity boosts occurred for authors with Asian names affiliated with institutions in Asia.The figure describes this effect as most pronounced for this group.
- LLM usage and references to prior works: Users accessed more books, more recent works, and less highly cited works after the release of Bing Chat.The comparison examines online accesses redirected from Google and Bing before and after the February 2023 release.