Source-linked AI summary
More than half of recent astronomy papers are written with language-model assistance
Serat M. Saad, Yuan-Sen Ting
TL;DR
The paper asks how much astronomy writing carries a language-model trace when the unassisted background after 2020 cannot be observed directly. It counts marker vocabulary and fits a hierarchical Bayesian mixture model with alternative background assumptions, finding traces in more than half of recent astro-ph papers.
Problem
Estimating language-model use from marker vocabulary requires modeling how frequently those words would appear without models, because that background rate is unobservable after 2020.
Method
The authors model per-paper marker counts as a length-scaled mixture of calibrated assisted and unassisted components, with subject, time, and fading-excess effects, fitting three post-2020 background variants.
Results
54+8_-8% of 2025 papers carry a language-model trace under the primary specification, never falling below 36% across variants, while the first half of 2026 reaches 87+8%.
Takeaways & Limitations
The estimated trace is widespread, whereas disclosure is rare: roughly 66 papers carry a trace for every one that declares it.
Takeaways & Limitations
The estimand detects traces in analyzed prose, so use that leaves no lexical trace is missed and every reported fraction is a floor.
Abstract
from arXiv · showhide
Language models leave a distinctive vocabulary in the prose they help write, and we measure how much of the astronomy literature now carries it. From the full text of 207,111 astro-ph papers spanning 2015 to mid-2026, we count those words in each paper and model the counts, in proportion to paper length, as a mixture of assisted and unassisted writing in a hierarchical Bayesian model. Papers from before 2020 calibrate the unassisted rate, and the 392 papers that disclose model use calibrate the assisted one. Our answer depends on how often these words would appear today if nobody used a model, a rate that must be modeled rather than observed, so we extend it past 2020 under three assumptions and report all three. For 2025 that gives $54^{+8}_{-8}\,(\mathrm{stat},\,95\%)\,^{+26}_{-0}\,(\mathrm{sys,\ background})$% of papers, the second error being the spread across the three. The estimate stays at or above 36% when we vary that choice, the calibration, and the requirement that adoption only rises. A word list built from the astro-ph corpus, keeping only words that rose across every subfield, leaves 2025 in the same range. Assisted writing is also getting harder to see, since authors adapt to the words that reveal it and the marker excess more than halves between 2023 and 2026. Our model allows for that fading, so it can separate a fainter trace from reduced use. More than half of recent astro-ph papers therefore carry a language-model trace, while only 0.81% of 2025 papers disclose it, one declaration for every $\sim$66 papers with a trace.
1. INTRODUCTION
Language-model assistance has left distinctive marker words in scientific prose, motivating estimates of undisclosed use in astronomy. The paper addresses limitations of earlier estimators with a full-text corpus and a hierarchical Bayesian model that accounts for paper length, changing background rates, and fading markers.
- 1. INTRODUCTION: Marker words are stylistic choices favored by language models, including “delve,” “underscore,” “intricate,” and “pivotal.”Human writers before November 2022 used these words rarely.
- 1. INTRODUCTION: The key inferential challenge is estimating how often marker words would appear without models, because that post-2020 background rate is unobservable.Observed frequencies are precise; uncertainty comes from modeling the missing counterfactual frequency.
- 1. INTRODUCTION: Earlier estimators fix or extrapolate the background and omit uncertainty in that extension, while fixed excess assumptions cannot distinguish adoption from fading detectability.These approaches also reduce papers to binary flags and use one yearly background rate.
- 1. INTRODUCTION: The study uses the AstroMLab 5 catalog to analyze astronomy full text, addressing the prior lack of a clean corpus for comparable measurements.The catalog curates 408,590 astro-ph papers from 1992 through July 2025 under standardized OCR normalization.
- 1. INTRODUCTION: A hierarchical Bayesian model counts marker words relative to analyzed text length and calibrates assisted and unassisted rates using declarations and pre-2020 writing.The model lets the relevant rates vary with topic and time and allows the difference between assisted and background rates to change over time.
- 1. INTRODUCTION: The paper evaluates the assisted fraction, marker fading, an astronomy-derived vocabulary, sensitivity to assumptions, and comparison with declared use.These analyses are organized across the paper’s results and discussion sections.
2. METHODS
The analysis counts marker vocabulary in 207,111 astro-ph papers and models length-adjusted counts with a hierarchical Bayesian mixture calibrated by pre-2020 writing and declared model use. It tests vocabulary, declarations, identification, background assumptions, and recovery of known injected prevalences.
- 2.2. Vocabulary: The pipeline counts three baskets: 38 marker words, ten neutral astronomy controls, and fifteen null hedges and intensifiers.Controls and nulls pass through the same pipeline to quantify false positives and systematic error.
- 2.3. Declarations of model use: 392 authorial declarations calibrate assisted writing, whose marker vocabulary occurs at 2.6–6.8 times the pre-2020 background while controls remain near background.After screening, 513 statements of use remain corpus-wide; 392 are in the fitted cohort.
- 2.4. The mixture model: The model treats each paper’s marker count as a length-scaled negative-binomial mixture, with topic and time effects for background, fading assisted excess, and prevalence.The latent assistance probability is the reported quantity, while hierarchical partial pooling shares information across eight subject classes.
- 2.5. Identification and the missing background: The assisted component is not identified by counts alone, because the same aggregate lift can arise from many weakly marked papers or few strongly marked papers.The analysis anchors the excess with declared papers and uses controls to expose the prevalence–excess trade-off.
- 2.5. Identification and the missing background: Post-2020 background vocabulary is unobserved, so the analysis fits three continuations, reports their spread separately, and also tests a nondecreasing-adoption constraint.Injection–recovery tests recover prevalence within two or three percentage points at excess 2.5 or higher; anchored recovery remains within six points for prevalence of at least 15% at excess 1.5.
3. RESULTS
The estimated fraction of astro-ph papers carrying a language-model trace exceeds half by 2025, while the lexical signal fades as adoption rises and remains detectable with an astronomy-derived vocabulary. The estimate is far above declared use, and the paper defines it narrowly as prose consistent with calibrated assisted writing.
- 3.1. The assisted fraction: 54% of astro-ph papers show a language-model trace in 2025 under the primary specification, with a 95% statistical interval of ±8% and a one-sided background spread of +26%.The linear background continuation is the conservative case; alternative assumptions shift the estimate separately.
- 3.1. The assisted fraction: The estimand counts papers whose analyzed prose is more consistent with calibrated assisted writing than background, not all model use or proof that a machine wrote the paper.It excludes assistance leaving no lexical trace, including code or idea assistance and text edited to remove markers.
- 3.2. Time evolution of the marker excess: The fitted excess of assisted prose falls from 3.5 times background in 2023 to 1.5 by 2026 while prevalence rises.This allows the model to represent a fainter lexical trace without treating it as reduced use.
- 3.3. Marker vocabulary derived from the corpus: 80+10/-12% of 2025 papers are estimated assisted under the corpus-derived vocabulary and linear drift, compared with 67+12/-11% for the imported basket.Thirty-two of 34 selected words independently exceed 20% excess in 2025–2026, while neutral controls remain near unit excess.
- 3.4. Comparison with declared use: 0.81% of 2025 full texts contain a model-use declaration versus an estimated 54+8/-8% prevalence, roughly one declaration for every 66 papers with a trace.An audit found no missed declaration in 150 undeclared papers, bounding the false-negative rate below 2%.
4. DISCUSSION AND CONCLUSIONS
The paper estimates that more than half of recent astro-ph papers carry a language-model trace, while the trace is fading as authors adapt and disclosures remain rare. Robustness checks support the broad conclusion, but the estimate remains a floor limited to detectable lexical traces and dependent on assumptions about undeclared users.
- DISCUSSION AND CONCLUSIONS: The estimate targets papers whose analyzed prose carries marker vocabulary at the strength observed in declared papers, not the fraction of text generated by a model.A per-token marker excess independently implies 61% assisted papers in 2025, above half.
- DISCUSSION AND CONCLUSIONS: The 2025 full-text estimate is 47% with the standard extrapolation estimator, but that method returns about 51% even for neutral controls and null hedges.The discrepancy is attributed to applying an abstract-oriented document-frequency estimator to much longer full texts.
- DISCUSSION AND CONCLUSIONS: The estimate is a floor because the detector cannot see model use that leaves no lexical trace, and the assisted calibration uses a selected minority of declaring authors.The primary analysis also uses OCR text from the latest revision, so the pre-2020 unassisted calibration may contain a small admixture of later revisions.
- DISCUSSION AND CONCLUSIONS: The marker excess falls from 3.5 to 1.5 within three years while use rises, making detection harder as authors adapt to revealing vocabulary.The model allows the assisted excess to fade separately from prevalence, while the control drift is absorbed into the background.
- DISCUSSION AND CONCLUSIONS: At roughly 66 trace-carrying papers per declaration, disclosure is far rarer than detectable use, while the fading trace makes future measurement less durable.The paper argues that disclosure is more tractable than enforcement and that reducing stigma could align reporting with policies permitting language editing.
DATA AND CODE AVAILABILITY
The paper provides analysis code and derived data while not redistributing the underlying OCR corpus. The release includes per-paper features, frozen vocabularies, and declaration information for reproducing the analysis.
- DATA AND CODE AVAILABILITY: The AstroMLab 5 OCR full-text corpus is available on request but is not redistributed; analysis code and derived data are released online.The catalog update extends coverage through August 2026.
- DATA AND CODE AVAILABILITY: The release contains per-paper word counts, section lengths, declaration flags, and frozen marker, control, and null baskets.It also includes the declaration table with evidence sentences, as described in the release passage.