Source-linked AI summary

LLM assisted writing deserves empirical evaluation

Xuan Zhong Feng, Yi Lin, Yiye Zhang, Chunhua Weng, Yifan Peng

arXiv:2608.22124v1cs.CL

TL;DR

LLM-assisted writing is usually framed as a detection and integrity problem, leaving its effects on scientific communication empirically underexamined. An analysis of 69,209 Health Informatics articles found associations with more focused presentation, broader citation practices, and more globally distributed authorship.

  • Problem

    LLM-assisted writing is often treated as a detection problem, leaving its effects on scientific communication as an empirical and social question.

  • Method

    The study analyzed 69,209 PubMed-indexed Health Informatics articles and compared bibliometric, semantic, reference, and authorship metadata across LLM-use groups.

  • Results

    LLM-assisted papers were associated with more focused presentation, broader citation practices, and greater participation from non-Anglophone settings and intercontinental collaborations.

  • Takeaways & Limitations

    Manuscripts should be evaluated by communication quality, scholarly accountability, citation practices, and access rather than tool use alone.

  • Takeaways & Limitations

    Whether these patterns generalize across disciplines, time horizons, and modes of LLM use remains an open question.

Abstract

from arXiv · show

LLM-assisted writing is often treated as a detection problem, as it raises questions about clarity, integrity, equity, and evaluation. An analysis of 69,209 Health Informatics papers links it to more focused presentation, broader citation practices, and more globally distributed authorship. These patterns do not prove better science, but they support evaluating manuscripts by scholarly quality and accountability rather than by tool use.

Main text

The study examines LLM-assisted writing in 69,209 PubMed-indexed Health Informatics articles, comparing pre-LLM, LLM-independent, and LLM-assisted papers. It reports more focused topic and semantic presentation, greater referenced-knowledge diversity, broader global representation, and similar visibility and impact for LLM-assisted papers.

  • Study design: 69,209 PubMed-indexed Health Informatics articles were analyzed to compare pre-LLM, LLM-independent, and LLM-assisted papers.The study used ChatGPT’s public release in late 2022 as the pre-LLM/post-LLM boundary and classified post-LLM papers with a SciBERT classifier.
  • Observed patterns: LLM-assisted papers showed lower topic and semantic spread than the control groups.These patterns indicate more focused presentation, without establishing that the underlying science was better.
  • Observed patterns: LLM-assisted papers showed higher referenced-knowledge diversity than the control groups.The comparison used bibliometric, semantic, reference, and authorship metadata.
  • Observed patterns: LLM-assisted papers showed higher global representation than the control groups.The observed authorship pattern was more globally distributed.
  • Observed patterns: LLM-assisted papers had similar visibility and impact to the control groups.The reported associations concern scientific communication and do not by themselves prove better science.

Semantic concentration in scholarly writing

LLM-assisted papers showed more focused topical and semantic presentation than comparison groups, with fewer topics per paper and embeddings closer to topic centroids. This concentration may improve clarity and accessibility, but semantic closeness to field norms does not establish novelty or contribution quality.

  • Semantic concentration in scholarly writing: LLM-assisted papers were assigned fewer topics per paper and had title-and-abstract embeddings closer to their topics’ semantic centroids.The pattern was more pronounced for LLM-assisted papers than for post-era LLM-independent and pre-LLM papers.
  • Semantic concentration in scholarly writing: A more focused presentation may clarify a manuscript’s contribution, situate it within existing work, and help readers assess its relevance.The passage presents these as possible benefits of concentration, not established causal effects.
  • Semantic concentration in scholarly writing: LLM-assisted writing may align language with disciplinary conventions, potentially improving clarity, discoverability, and accessibility.The passage highlights particular relevance for authors writing across linguistic or disciplinary boundaries, though its final wording is truncated.
  • Semantic concentration in scholarly writing: Semantic closeness to field norms does not by itself indicate either novelty or lack of novelty.A semantically conventional paper may still contain a novel method, dataset, application, or empirical result, while fluent writing may accompany a limited contribution.

Reference patterns and knowledge breadth

LLM-assisted papers cited more references, covered more diverse topics, and drew from broader bibliographic ranges than post-era LLM-independent papers. These patterns may reflect broader literature exploration, but they do not establish better selection or deeper synthesis and underscore the need to evaluate citation quality.

  • Reference patterns and knowledge breadth: LLM-assisted papers cited more references and showed higher reference-topic diversity than post-era LLM-independent papers.Their bibliographies also drew from a broader range of topics.
  • Reference patterns and knowledge breadth: LLMs can help researchers summarize literature, identify adjacent concepts, generate search terms, and reorganize related work.These practices may lower the cost of exploring neighboring literatures and expand the papers authors consider relevant.
  • Reference patterns and knowledge breadth: Broader retrieval can increase meaningful interdisciplinary connections or include citations weakly related to the scientific argument.A more extensive and diverse bibliography does not necessarily indicate deeper synthesis; citations may be superficial, misplaced, or loosely connected to claims.
  • Reference patterns and knowledge breadth: Human-written papers can contain decorative citations and incomplete literature reviews, while LLMs may amplify, reduce, or leave these practices unchanged.Automated evaluation of citation accuracy, balance, and relevance may become increasingly important as LLM-assisted writing becomes more common.

Authorship and global participation

LLM-assisted papers showed broader global participation, including more non-Anglophone first authors, modestly greater representation from low- and middle-income economies, and more intercontinental collaboration. The passages suggest language assistance may reduce publishing barriers, while unequal access could create new advantages.

  • Authorship and global participation: LLM-assisted papers had more first authors affiliated with non-Anglophone settings, modestly higher representation from low- and middle-income economies, and more intercontinental collaboration.These differences were observed relative to both post-era LLM-independent papers and pre-LLM papers.
  • Authorship and global participation: LLMs may reduce barriers to English-language scientific publishing by providing on-demand language assistance and helping authors align manuscripts with publication norms.Researchers outside Anglophone settings often face additional editing, translation, and conformity costs.
  • Authorship and global participation: Unequal access to advanced writing assistants may create new advantages for researchers who can produce and submit manuscripts more rapidly.Greater access could reduce time spent on literature review, drafting, revision, and manuscript preparation in competitive publication environments.

Visibility and recognition

LLM-assisted papers had lower raw citation counts, largely because they were published more recently, but they did not show weaker early citation acquisition within a fixed post-publication window. These findings address visibility and recognition rather than intrinsic scientific quality, and their robustness across contexts remains open.

  • Visibility and recognition: Lower raw cited-by counts for LLM-assisted papers were largely attributable to more recent publication dates and less time to accumulate citations.The analysis did not support the simple expectation that LLM-assisted writing necessarily receives less recognition.
  • Visibility and recognition: Within a fixed post-publication window, LLM-assisted papers did not show weaker early citation acquisition.
  • Visibility and recognition: Citation counts and journal impact factors are imperfect indicators of scientific value, shaped by timing, field norms, editorial selection, and topic popularity.The findings therefore speak to visibility and recognition, not intrinsic scientific quality.
  • Visibility and recognition: Whether these patterns hold across disciplines, time horizons, and modes of LLM use as writing assistance remains an open question.

LLM-assisted writing vs agentic research

The study distinguishes LLM-assisted writing from human intellectual collaboration and agentic research, treating it as non-agentic support for scientific communication. Its effects are presented as contingent on adoption, regulation, and integration into academic practices rather than inherently beneficial or detrimental.

  • LLM-assisted writing vs agentic research: LLM-assisted writing is distinguished from co-authorship and co-research, which involve human intellectual collaboration.The study focuses on writing assistance rather than shared scientific responsibility.
  • LLM-assisted writing vs agentic research: The tools transform language, summarize content, and restructure writing without generating scientific intent or formulating research questions.They also do not design studies or assume responsibility for result validity.
  • LLM-assisted writing vs agentic research: LLM-assisted writing is neither inherently detrimental nor inherently beneficial; its consequences depend on adoption, regulation, and integration into academic practices.The paper frames it as a shifting component of scholarly production rather than a fixed moral category.
Loading 2608.22124v1…