Source-linked AI summary

Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation

Yixuan Liu, Lin Chen, Zhuoqi Liu, Jianglin Lu, Dakota Murray

arXiv:2609.01432v1cs.DLcs.CLcs.CYcs.SI

TL;DR

The paper asks whether LLM-assisted citations preserve human rhetorical intent and social selection patterns. It compares masked-citation reconstructions with human choices in aligned contexts and finds that LLMs cite less critically, favor popular and older work, and reach farther across coauthorship networks.

  • Problem

    LLMs increasingly generate scientific citations, but it remains unclear whether they differ from humans in citation intent and selection across rhetorical roles.

  • Method

    The study uses masked citation reconstruction to create position-aligned LLM counterfactuals, labels intent with an LLM judge, and analyzes cited-author distance in a coauthorship network.

  • Results

    LLMs cite less critically, over-cite popular and older work, and select more socially distant authors than humans.

  • Takeaways & Limitations

    LLM citation may broaden scholarly reach beyond close collaborators while muting critical engagement and amplifying visibility bias.

  • Takeaways & Limitations

    Intent labels come from LLM judges not validated against human judgment, so absolute rates require caution despite replicated directional comparisons.

Abstract

from arXiv · show

Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We introduce a masked-citation task to compare human and LLM-generated citation behavior. For each citation context, an LLM generates a replacement citation sentence, producing a counterfactual corpus directly comparable to human citation. We analyze what, whom, and how models cite, using an LLM-as-a-judge to classify citation intent and a 20-million-edge coauthorship network to measure social distance between cited authors. Across six popular LLMs and 1,746 top NLP conference papers (63k+ contexts, 132k+ citations), three patterns emerge: (1) Compared with human citation, LLMs cite significantly less critically; (2) LLMs over-cite popular and older papers, a tendency amplified for contrasting citations where human writing more often draws on recent, niche work; (3) Whereas humans often cite within their close social network, especially for supporting citations, LLMs tend to draw on more socially distant authors. Together, these differences are double-edged: LLM citation reaches beyond a scholar's close collaborators while being less critical and amplifying visibility bias, reshaping the rhetoric and reach of scientific citation.

1 Introduction

The paper asks whether LLM-generated citations preserve the rhetorical intent and social patterns of human citation. It introduces an aligned counterfactual benchmark and finds less critical, popularity-amplifying, and more socially distant citation behavior.

  • Research questions: The study compares supporting, contrasting, and mentioning citations to test whether LLMs reproduce human citation intent and coauthorship proximity.
  • Research gap: LLM citation research has examined reliability, selection, and author demographics, but not how citation choices vary by rhetorical role.
  • Contributions: The position-aligned benchmark makes each LLM citation a counterfactual to the human citation at the same rhetorical slot.
  • Contributions: LLMs under-produce contrasting citations and amplify popularity and recency biases along the intent dimension.
  • Contributions: Humans tend to cite within close coauthorship neighborhoods, especially for supporting citations, whereas LLMs cite more socially distant papers.
  • Implications: LLM citation may broaden selection beyond a scholar’s narrow social circle while citing older, more popular work in a less critical tone.

2 Methodology

The methodology reconstructs masked citation sentences under matched conditions, labels citation intent, links references to bibliometric records, and evaluates six LLMs on a corpus of 1,746 NLP papers.

  • 2.1 Task definition: The reconstruction task masks each citation sentence while preserving local context, section heading, citing-paper metadata, and the original citation count.
  • 2.1 Task definition: The benchmark aligns human and LLM outputs by position, time, and citation count, leaving citation choice as the remaining variation.
  • 2.2 Labeling citation intent: An LLM judge assigns supporting, contrasting, or mentioning labels from each citation sentence and its local context without seeing the cited paper.
  • 2.3 Bibliometric grounding: References are matched to Dimensions by DOI and then normalized title for downstream bibliometric analysis.
  • 2.4 Data and models: The corpus contains 1,746 ACL, EMNLP, and NAACL 2025 main-track papers, 63,944 citation contexts, and 132,913 citation slots.
  • 2.4 Data and models: The experiment spans six LLMs, with Gemini-3-Flash-Preview as the primary intent judge and DeepSeek-V4-Flash as a robustness check.
  • 2.4 Data and models: Human references match Dimensions at 86.7%, while LLM matching varies from 39.5–81.9% across models.

3 LLMs cite less critically (RQ1)

LLMs systematically make scientific citation rhetoric warmer than humans, producing fewer contrasting citations and often rewriting criticism as support or mention. This warming varies by paper section but persists across models and evaluation methods.

  • 21% of human citations were supporting, 19% contrasting, and 60% mentioning, whereas LLM-filled sentences produced only 10.7–17.2% contrasting citations.Five of six models exceeded the human supporting baseline of 21%, reaching up to 40.6%.
  • ∆cont = −2.4 to −8.6%, as LLM-filled sentences persistently move citation rhetoric toward mentioning or supporting.Contrasting human sentences became supporting in 12.5–31.1% of cases, whereas supporting-to-contrasting changes occurred in only 3.6–7.6%.
  • Contrasting citations preserved their original intent only 34.6–50.6% of the time, compared with 60–80% preservation for mentioning citations.Human–LLM intent agreement was at best fair, with Cohen’s κ = 0.32–0.38.
  • LLM self-reports were warmer than external-judge classifications, with four models reporting 54–75% supporting intent and almost no contrasting intent.The externally judged warming effect persisted, so it was not attributable only to model self-reports.
  • LLMs under-produced contrasting citations in Background, Method, and especially Discussion, while showing more section-specific support in Experiments and Conclusion.Humans used Discussion for both support and contrast, whereas every LLM generated contrasting citations below its own baseline there.
  • The consistent warming pattern mutes the critical engagement that human citation encodes.The result held despite section-level variation and was supported by independent judging and manual annotation.

4 Citation biases moderated by LLMs’ intent (RQ2)

Citation intent moderates how LLM-selected references differ from human selections. LLMs favor highly cited work, older work in contrasting contexts, and smaller teams, with the strongest divergences depending on intent.

  • 2.25 years is the human recency advantage for contrasting citations, where humans cite more recent and less-cited work than in other intent categories.Human contrasting citations averaged 322 citations, compared with 583 for supporting and 670 for mentioning citations.
  • 1.3–4.5× citation-count ratios show that LLMs select more highly cited papers than humans for the same citation type.The strongest gaps occurred for supporting citations, while contrasting gaps were smaller or below parity for some models.
  • 1.6–3.3 years older is the LLM selection gap for contrasting citations, the largest recency divergence across intents in every model.This contrasts with the citation-count peak, which occurs most strongly for supporting citations.
  • 0.30–0.61× team-size ratios show the largest LLM under-selection of large teams for mentioning citations.The highlighted intent was significantly different from the other two combined in all six models, with p < 0.01.
  • Figure 3 compares intent-matched human–LLM gaps in citation count, recency, and team size, with parity at 1 for ratios and 0 for recency differences.Each focal paper contributes one observation per intent, and uncertainty reflects between-paper variation.

5 Attenuated social proximity in LLM citation (RQ3)

LLMs cite authors who are more socially distant from the focal authors than humans do, and they largely ignore the rhetorical-intent gradient in human social proximity. This conclusion is bounded by limited reachable coverage in the coauthorship network.

  • The analysis averages four first/last-author dyads per citation slot using BFS distances in a 2015–2024 coauthorship graph spanning 2.1M+ nodes and 20M+ edges.Multi-citation contexts contribute one slot-level observation so over-cited slots do not dominate.
  • 26.5% of 133k+ human citations had at least one reachable author-role dyad in the coauthorship graph.LLM citations had reachable dyads for 14.5–25.5% of slots, depending on the model.
  • 3.40 hops is the human mean author-dyad distance, versus 3.65–3.89 hops for all six LLMs.No LLM recovered the human bias toward socially proximate citations.
  • 3.31 hops is the human supporting-citation distance, compared with 3.45 for contrasting and 3.43 for mentioning citations.The supporting-versus-other-intents difference was statistically significant at p < 0.001.
  • ±0.05 hops describes the tight clustering of LLM distances across rhetorical intents, unlike the strong human gradient.The flat LLM pattern persisted within the same research field.
  • 7–10% of human citation slots were in-network at d ≤1, compared with only 0.5–1.6% for every LLM.Human in-network citation reached 9.8% for supporting citations, while LLM rates showed no intent gradient and nearly absent self-citations.

6 Discussion

The masked-citation framework enables direct, intent-conditioned comparisons between human and LLM citation behavior. Results suggest that LLMs warm critical rhetoric, draw on socially broader sources, and retain selection biases that vary by intent.

  • Discussion: The framework aligns each LLM reconstruction with the human citation at the same slot, enabling direct comparison of citation selection and rhetoric.It uses six LLMs and more than 1,700 ACL, EMNLP, and NAACL papers.
  • Discussion: LLMs under-produce contrasting citations and often rewrite human criticism as supporting language.The observed warming may relate to model positivity and agreeableness, but its mechanism remains unresolved.
  • Discussion: LLM citation biases vary by intent: models favor highly cited work when supporting, older work when contrasting, and smaller author teams when mentioning.The paper proposes training-data composition as one possible explanation and calls for further mechanistic study.
  • Discussion: Humans tend to cite within close social networks, whereas LLMs select more socially distant authors across intents.The authors relate this difference to humans’ social-network exposure and models’ lack of equivalent social proximity.
  • Discussion: LLMs may broaden scholarly exposure beyond close collaborators while obscuring critical divides through warmer citation rhetoric.The paper frames these effects as a trade-off rather than an unqualified improvement.
  • Discussion: The authors recommend continued auditing and researcher awareness as machine-generated citations become embedded in scientific writing.They argue that preserving critical engagement requires understanding these systems’ tendencies and limits.

Limitations

The study’s conclusions are constrained by judge validation, uneven matching coverage, limited coauthorship-network reach, and a narrow conference-and-language corpus.

  • Limitations: Intent labels were not validated against human judgment, so absolute rates require caution even though directional comparisons replicate under a second judge.Both judges may share biases inherited from LLM training distributions.
  • Limitations: Matching rates vary across models from 39.5–81.9%, although a shared-context analysis reproduces all three intent-amplified patterns.This supports the conclusion that coverage variation is not the source of the reported bias.
  • Limitations: The coauthorship network reaches only 26.5% of human and 14.5–25.5% of LLM citation dyads.Early-career researchers, non-Western institutions, and industry practitioners are systematically underrepresented in the reachable subgraph.
  • Limitations: The corpus contains English-language ACL, EMNLP, and NAACL main-track papers, so findings may be specific to this conference series.Citation norms differ across disciplines, geographies, and languages.

Ethical considerations

The study uses public scholarly and bibliometric records for aggregate comparisons while identifying risks from LLM-assisted citation generation. Its controlled prompts and withheld citation information support the comparison, but the ethical implications include bias, privacy, and uncritical adoption concerns.

  • Ethical considerations: The study uses public publication records and authorship metadata only to compute aggregate, model-comparative statistics.It does not evaluate, rank, profile, or re-identify individual researchers, and it infers no protected or sensitive attributes.
  • Ethical considerations: Unresolved authors and unmatched citations are excluded rather than imputed, and raw bibliometric records and networks are not redistributed.The planned release is limited to code and aggregate, de-identified outputs.
  • Ethical considerations: The paper warns that LLM-generated citations may favor highly cited and older work while moving away from contrasting citations and close collaborators.Uncritical reliance could therefore entrench a canonical scholarly record.
  • Ethical considerations: The reconstruction task masks the citation sentence while preserving local context and required citation count, preventing direct access to the removed sentence and cited work.The setup makes each model output a counterfactual to the human citation at the same slot.
  • Ethical considerations: The generation prompt fixes the output schema, citation count, one-distinct-work rule, and three intent categories across reconstructions.These constraints reduce variation unrelated to the citation behavior being compared.
  • Ethical considerations: The judge classifies sentences as supporting, contrasting, or mentioning using local context while withholding cited-paper titles and alternative sentence versions.The same judging procedure is applied independently to human and reconstructed citations.
  • Ethical considerations: Supporting citations align with evidence, methods, or findings, contrasting citations dispute or compete, and mentioning citations provide background without clear support or contrast.These operational categories distinguish rhetorical role rather than treating citations as homogeneous retrieval outputs.
  • Ethical considerations: LLMs’ self-reported intent is warmer than judge-assessed intent, which is itself warmer than the human baseline.Self-supporting labels reach 28–75%, while self-contrasting labels are only 3–17%.

C Robustness check with DeepSeek as judge

Using DeepSeek-V4-Flash instead of Gemini as judge preserves the paper’s central robustness patterns: LLMs are less contrasting, show intent-dependent citation biases, and cite more socially distant papers than humans.

  • Robustness overview: DeepSeek-V4-Flash reproduces the finding that LLM-generated citations are less contrasting than human citations and that citation biases vary by rhetorical intent.The alternative judge labels human and LLM-filled sentences under the same taxonomy, and the replicated patterns are described as qualitatively unchanged.
  • Judge agreement: The two judges reach Cohen’s κ = 0.52 on human originals and 0.45 on pooled LLM fills, with disagreements concentrated on mentioning boundaries.They differ on supporting versus contrasting labels in only 2–9% of cases.
  • Intent preservation: Contrasting is least preserved for every model at 20.3–32.4%, while mentioning is most preserved at 70.3–83.6%.Supporting preservation is intermediate at 25.3–41.9%; the ordering matches the Gemini-judge analysis despite lower absolute rates.
  • Recency bias: For five of six models, the human–LLM recency gap is largest for contrasting citations, with significant gaps of +1.63 to +3.08 years in four models.Llama-4 shows similar contrasting and mentioning gaps, with the mentioning gap significant; Claude-3.5-Haiku has the same contrasting direction without significance.
  • Social proximity: Humans cite closer than every LLM, with mean distance 3.33 for supporting citations versus 3.65–3.92 across LLMs.Human in-network citation rates are 7.6–9.7%, compared with 0.5–1.6% for every LLM, and the human intent gradient remains significant for supporting versus mentioning (p < 0.001).
  • Human validation: Human annotators support the warming pattern: contrasting falls and supporting rises in GPT-5.1 reconstructions for every annotator.The primary judge matches human-majority labels on 73% of validation sentences, with F1 = 0.79 for supporting and F1 = 0.73 for contrasting.

F Robustness to context window

The intent distribution of LLM-generated citations changes little when GPT-5.1 receives progressively wider manuscript context, indicating that the warming pattern is robust to context-window size.

  • Context-window comparison: Contrasting remains low and supporting high as GPT-5.1’s input expands from a ±1-sentence window to the full paragraph and paragraph plus abstract.The experiment uses a stratified 600-slot subset with 200 slots per human intent and DeepSeek-V4-Flash judging.
  • Robustness result: The intent distribution barely shifts across context conditions, so the observed distribution is robust to the amount of context provided.The full paragraph supplies about 2.6× the baseline context.

G Robustness of the recency gradient

The recency gradient persists when analysis is restricted to cited works published before 2024, showing that the pattern is not limited to access to very recent papers.

  • Pre-2024 restriction: Contrasting remains the peak-recency intent in every model and stays significant when both human and LLM citations are restricted to works published before 2024.The restriction targets works within every model’s stated knowledge access.

H Robustness check with slot intersections

Analyses on identical context intersections preserve the intent-amplified citation patterns, while retrieval augmentation and field controls leave the main conclusions unchanged.

  • Slot intersection: The shared-context intersection retains 12,556 contexts across 1,695 of 1,746 papers, enabling an apples-to-apples comparison across human and all six LLM sources.The design reduces context-depth coverage differences, though confidence intervals widen with the smaller sample.
  • Intent-amplified bias: Citation-count gaps peak at supporting, recency gaps at contrasting, and team-size gaps at mentioning on the shared context set.Representative gaps include 4.74× for DeepSeek citation counts, +3.06 years for Gemini-2.0 recency, and 0.38× for Qwen team size.
  • Interpretation: The persistence of these gaps on identical contexts indicates that they reflect citation choices rather than differential slot coverage.Every model contributes at least one Dimensions-matched citation in each retained context.
  • Retrieval augmentation: With live web search, GPT-5.1 still produces about 10% contrasting citations versus 33% for humans.The retrieval-augmented test uses a balanced subset of 599 slots judged by DeepSeek-V4-Flash.
  • Field control: Within and between fields, humans cite closer than GPT-5.1, with gaps of +0.31 to +0.39 hops and p < 10^-6.The field-stratified result indicates that the proximity gap is not explained by field structure.
Loading 2609.01432v1…