Source-linked AI summary
How LLMs Distort Our Written Language
Marwa Abdulhai, Isadora White, Yanming Wan, Ibrahim Qureshi, Joel Z. Leibo, Max Kleiman-Weiner, Natasha Jaques
TL;DR
The paper asks whether widespread LLM assistance changes the meaning of human writing, beyond documented effects on style. Using a randomized user study, counterfactual essay revisions, and real peer reviews, it finds consistent shifts in stance, semantics, voice, and evaluation criteria. The authors conclude that AI writing assistance can alter intended meaning even when used for editing, with implications for cultural and scientific institutions.
Problem
The effect of LLM writing assistance on intended meaning remains underexplored despite evidence of stylistic homogenization and growing use across political and scientific institutions.
Method
The paper combines a randomized study of 100 writers, counterfactual revisions of 86 pre-LLM essays using human feedback, and analysis of LLM-generated ICLR peer reviews.
Results
LLMs consistently alter human writing’s intended meaning, voice, and style, including a 68.9% increase in neutral essay stances and changed criteria in AI-generated peer reviews.
Takeaways & Limitations
LLM writing assistance can steer human text and institutional evaluation toward different conclusions than originally intended, even when models are asked to make minimal edits.
Takeaways & Limitations
The user-study findings may differ for people residing in other countries or writing in other languages because participants were native English speakers in the United States.
Abstract
from arXiv · showhide
Large language models (LLMs) are used by over a billion people globally, most often to assist with writing. In this work, we demonstrate that LLMs not only alter the voice and tone of human writing but also consistently alter the intended meaning. First, we conduct a human user study to understand how people actually interact with LLMs when using them for writing. Our findings reveal that extensive LLM use led to a nearly 70% increase in essays that remained neutral in answering the topic question. Significantly more heavy LLM users reported that the writing was less creative and not in their voice. Next, using a dataset of human-written essays that was collected in 2021 before the widespread release of LLMs, we study how asking an LLM to revise the essay based on the human-written feedback in the dataset induces large changes in the resulting content and meaning. We find that even when LLMs are prompted with expert feedback and asked to only make grammar edits, they still change the text in a way that significantly alters its semantic meaning. We then examine LLM-generated text in the wild, specifically focusing on the 21% of AI-generated scientific peer reviews at a recent top AI conference. We find that LLM-generated reviews place significantly less weight on clarity and significance of the research, and assign scores that, on average, are a full point higher. These findings highlight a misalignment between the perceived benefit of AI use and an implicit, consistent effect on the semantics of human writing, motivating future work on how widespread AI writing will affect our cultural and scientific institutions.
1. Introduction
LLM use in writing can reshape both the style and intended meaning of human text, while its institutional consequences remain underexplored. Across user, editing, and peer-review analyses, the paper finds shifts toward neutral stances, shared semantic styles, and changed evaluation criteria.
- Motivation: Over 1 billion people use LLMs weekly, making their effects on written language and cumulative culture a central open question.
- Research gap: The effect of LLM use on intended meaning remains underexplored despite prior evidence of stylistic homogenization.
- User study: 70% change in argumentative stance occurred in the user study, with extensive AI use shifting essays from for-or-against positions toward neutrality.Heavy users also reported less creativity and weaker adherence to their own written voice.
- Editing analysis: LLM editing produced larger, more globally directed changes than human revision, altering semantics even when applied to human-written drafts.The edits shifted essays away from both the original drafts and counterfactual human edits.
- Editing analysis: LLMs increased both argumentative and analytical language and emotional language, roughly doubling positive and negative sentiment usage.
- Institutional analysis: In peer reviews, LLM-generated text changed not only average scores and response diversity but also the criteria used to evaluate research.These reviews focused less on clarity, relevance, and impact and more on reproducibility, scalability, and practical application.
- Implication: The findings identify a capability deficiency: LLMs do not reliably assist with writing without altering meaning or reducing creative expression of human voice.
2. Related Work
Prior work documents stylistic convergence, reduced diversity, and possible cognitive effects of LLM-assisted writing, but this paper extends the concern to expressed views and institutional judgment.
- LLM homogenization: LLMs tend to converge stylistically across models and tasks, producing similar responses and narrowing communicative styles.Prior work characterizes this convergence as a structural property rather than a surface-level effect.
- LLM homogenization: The paper extends homogenization findings by showing that transformed text can have different lexical and emotional characteristics and argue for different conclusions.
- Effects on humans: LLM writing assistance has been linked to reduced creativity and less diverse writing, alongside reported productivity benefits and possible cognitive costs.
- Effects on humans: The paper uses randomized controlled trials to examine whether LLMs influence not only writing style but also the views and judgments people express.
- Effects on institutions: As LLMs enter scientific institutions, prior work raises concerns that individual productivity gains may coincide with a contraction in the range of scientific topics studied.
3. Experiment Methodology
The paper combines a human writing study, counterfactual revision analysis, institutional text analysis, and multiple measures of semantic, lexical, affective, and linguistic change.
- Overview: The methodology quantifies semantic, lexical, grammatical, and affective differences between human-written and LLM-written or LLM-edited text.
- Human study: The human study randomly assigned 100 participants to write an argumentative essay with or without access to an embedded LLM.The study also collected questionnaires about attitudes, creativity, and alignment with writing preferences.
- Human study: A pilot study identified participants who used the LLM extensively versus those who abstained or used it only for peripheral information-seeking or critique.
- ArgRewrite-v2: ArgRewrite-v2 contains 86 university students’ argumentative revisions, based on essays written in 2021 before widespread LLM writing assistants.Each initial draft received human expert feedback before a human-produced revision was created.
- ArgRewrite-v2: Three production LLMs were prompted with the original draft and human feedback to generate comparable revisions for counterfactual comparison with human edits.
- Peer reviews: The peer-review analysis examines ICLR reviews to study LLM use in a setting where AI-generated reviews were prohibited and reviewer reputation was visible to area chairs.
- Metrics: Semantic change was measured in embedding space using PCA, where distances represent similarity in essay meaning or semantic content.
- Metrics: Lexical shifts were measured with unigram distributions and Jensen-Shannon Divergence, with higher values indicating larger differences from human writing.
4. Results
Across the user study and ArgRewrite-v2 analyses, LLM assistance consistently changed essays’ semantics, vocabulary, grammar, emotional tone, and argumentative style, often making writing less personal and more homogeneous.
- User perceptions: Heavy LLM users reported that their essays were less creative and less in their voice.They did not report significantly less writing struggle or greater satisfaction.
- Semantic change: LLM-generated essays formed a tight semantic cluster distinct from the broader distribution of human-written essays.Information-seeking use was more similar to human writing and preserved semantic meaning, creativity, and voice.
- Semantic change: LLM revisions produced larger, consistently aligned semantic shifts than human revisions, even for minimal or grammar-only edits.The shifts frequently changed the human writer’s conclusion and were largest when models completed or expanded essays.
- Lexical change: LLM editing replaced vocabulary more extensively than human editing, while the unique lexical fingerprint of each writer was overwritten by model-preferred vocabulary.Human edits made modest substitutions, whereas LLM revisions shifted lexical distributions away from the human baseline.
- Grammar and style: LLM writing reduced pronouns by 50% and increased nouns by 14%, shifting essays toward impersonal language.Across editing conditions, models also increased adjectives and coordinating conjunctions while reducing pronouns and determiners; human POS changes were typically under 5%.
- Argumentation: LLM-assisted writing increased analytical, expert-opinion, and statistical language, while LLM editing also substantially altered the content and meaning of original essays.These changes occurred alongside the shifts toward more emotional language described in the editing analyses.
5. Discussion
Across three studies, the paper finds that LLMs consistently alter the meaning, style, and voice of human writing, with implications for cultural and scientific institutions. The authors argue that writing assistance should better preserve users’ preferences and intended meaning.
- LLMs significantly alter the meaning, style, and voice of human writing across user interactions, editing comparisons, and real-world institutional use.The effects persist across model types and even when LLMs are prompted to make minimal edits.
- 70% of essays changed argumentative stance with extensive AI use, while heavy users more often judged the resulting writing less creative and unlike their voice.
- LLM use increases both emotional language and logical or analytical argumentation, potentially making text more broadly convincing without preserving individual preferences.The paper links this possibility to reinforcement learning from human feedback and the absence of mechanisms for maintaining users’ intended meaning.
- The authors call for writing systems that infer and preserve users’ underlying preferences and intended meaning, alongside research on effects across cultural institutions.
- LLM-generated peer reviews shift scientific evaluation criteria toward reproducibility, scalability, and practical application, away from clarity, relevance, and impact.The authors describe these changes as having unknown downstream effects on decisions about what scientific work is considered valid and incentivized.
Ethics Statement
The ethics statement describes participant protections and identifies language and cultural variation as an important boundary of the user-study findings.
- The user study received IRB approval, obtained participant consent, provided compensation, and followed protocols to anonymize personally identifiable information.
- Because participants were native English speakers living in the United States, the findings may differ across languages and cultural settings.The paper identifies effects on other languages and cultural norms as a key question for future research.
- The ArgRewrite-v2 and academic peer-review data were publicly available and contained no personally identifiable information.
Reproducibility Statement
The reproducibility statement points readers to detailed study materials, prompts, and full analyses covering the paper’s major evaluation dimensions.
- The paper provides user-study instructions, pre-study and post-study questions, and other recruitment and consent materials in Section A.1 and related appendices.
- The authors provide the prompts used for ArgRewrite-v2 analyses and full results for semantic, lexical, emotional, and part-of-speech measurements.
B. ArgRewrite-v2 Analysis
The ArgRewrite-v2 analysis appendix documents the procedures, prompts, and statistical significance materials supporting the paper’s essay-editing analysis.
- Appendix B.1 documents how LLM drafts were generated for the ArgRewrite-v2 analysis.
- Appendices B.2–B.4 report semantic-shift analyses across embeddings, settings, and models.
- Appendix C lists the prompts used for LLM-as-a-judge analyses of ICLR reviews and for the user study.
- Appendix D provides statistical significance scores for the ICLR review categories.
E. Human User Study Essay Samples
The human user study recruited participants to write an argumentative essay with or without an embedded LLM, while collecting attitudes, writing experiences, and usage patterns. The study also recorded participants’ evaluations of the resulting essays and their perceived degree of LLM involvement.
- Participants: Participants were recruited through Prolific and were required to be native English speakers residing in the United States.The study paid participants 8 US dollars for an estimated 35-minute session, with a one-hour maximum.
- Study procedure: Participants completed pre-study questions about AI and writing habits before writing, then answered post-study questions about their experience.The study instructions covered writing behavior, attitudes toward AI, and reflections on the completed essay.
- Control condition: A separate control instruction prohibited LLM, AI-assistant, and internet use while participants wrote their essays.Responses were reviewed for indications of AI-generated content, and participants found to have used such tools were not compensated.
- Study measures: The study measured perceived essay quality, voice, creativity, writing difficulty, effort, learning, and control over the writing process.Participants rated statements about satisfaction, whether the essay reflected their voice, creativity, and the LLM’s initiative and collaboration.
B. ArgRewrite-v2 Analysis
The ArgRewrite-v2 analysis compares how different LLM revision modes and models shift human-written essays relative to their initial drafts. It examines semantic, lexical, grammatical, and emotional changes under revisions with and without expert feedback.
- Revision settings: Five revision modes were evaluated: expert, minimal, grammar, completion, and expansion.The prompts respectively asked models to revise, preserve a similar word count, correct grammar, finish a first paragraph, or expand the draft.
- Lexical analysis: Lexical divergence was analyzed with Jensen-Shannon Divergence distributions for the five revision modes, separately for revisions with and without expert feedback.The corresponding figures cover general, grammar, minimal, completion, and expansion revisions.
- Grammatical analysis: Grammatical shifts were represented through part-of-speech distributions for general, grammar, minimal, completion, and expansion revisions.The figures compare revisions with expert feedback against revisions without feedback.
- Emotional analysis: Emotional shifts were measured for the same five revision modes by comparing revisions with expert feedback and revisions without feedback.The figures present these comparisons using left-right layouts rather than the top-bottom layouts used for several other analyses.
F.1. Sample 1 — Placing conditions on when money leads to happiness
This sample develops a conditional argument that money can support happiness by meeting needs and enabling freedom, while financial resources alone cannot resolve relational or personal problems. The LLM-influenced version emphasizes stress reduction, experiences, autonomy, and financial freedom.
- Conditions for happiness: Money is presented as potentially happiness-producing when it meets basic needs such as food, shelter, healthcare, security, and safety.The argument distinguishes need satisfaction from using money to address insecurity, status concerns, or appearance.
- Limits of money: The sample argues that money cannot by itself fix personal problems, create meaningful relationships, or guarantee happiness.It uses wealthy people’s continuing experiences of divorce, loneliness, and emotional difficulty as supporting examples.
- Human perspectives: One human essay takes the narrower position that money does not buy happiness but can make daily life easier under some circumstances.Its examples contrast the practical burdens faced by poor people with the convenience available to wealthy people.
- LLM-influenced argument: The LLM-influenced essay reframes money as a contributor to happiness when supportive relationships and other foundations are already present.It says financial freedom can ease stress and anxiety, enable travel and hobbies, and expand available choices.
- Freedom and constraint: The revised argument presents financial freedom as a form of autonomy because monetary constraints can limit access to food, shelter, healthcare, and meaningful experiences.It links reduced financial worry to greater freedom in choosing how to live.
G.2. Sample 2 — A neutral perspective, striking a balance
This sample balances the view that relationships and experiences generate happiness with the claim that financial stability may be needed to access or enjoy them. The resulting position treats money as a possible condition for happiness rather than its direct cause.
- Non-material happiness: The sample argues that relationships, experiences, nature, and other non-material sources can produce happiness without financial investment.It uses time with loved ones, awe-inspiring moments, and walking outside as examples.
- Financial stress: It then argues that severe financial stress can prevent people from having the mental space to prioritize or enjoy such experiences.The reasoning focuses on stress about bills and meeting immediate needs.
- Revised conclusion: The revised conclusion shifts the question from whether money causes happiness to whether financial confidence or stability is required for happiness.It preserves uncertainty by making the answer dependent on individual circumstances.
- LLM feedback: The LLM’s feedback explicitly recommends clearer transitions, local rephrasing, and a stronger conclusion before finalizing the essay.Its suggested revisions sharpen the contrast between happiness from relationships and the possible need for money under financial stress.
G.3. Sample 3 — Money does not lead to happiness
The LLM preserves the writer’s central argument while affirming and elaborating it, extending the essay’s treatment of wealth, social barriers, and inner healing.
- The draft describes extreme wealth as burdensome because it creates financial-management stress and barriers to meaningful relationships.
- The final essay argues that money cannot create happiness without prior inner healing and a foundation of genuine happiness.
- The LLM responds by expanding the author’s ideas about insecurity, wealth-related social isolation, purpose, and happiness as a foundation.
- The author resubmits the same draft without changes, indicating that the interaction retains the original content rather than documenting a revised human draft.
H.1. Sample 1 — A neutral balanced take
The LLM-generated essay presents a balanced account: money supports basic needs and reduces financial stress, but relationships, purpose, and fulfilment remain important to happiness.
- The writing was assembled through guided prompts requesting ideas, an introduction, paragraphs on financial stress and research, and a conclusion.
- The essay frames money as necessary for basic needs while arguing that it alone cannot provide deeper happiness.
- Financial strain is presented as a source of stress, depression, dissatisfaction, and reduced access to opportunities that may support happiness.
- A cited account states that happiness gains from income diminish beyond an annual threshold of $75,000.
- The essay considers universal basic income and AI-driven changes to work as developments that could reshape money’s relationship with happiness.
H.2. Sample 2 — A neutral take supported by statistics
The essay combines statistical and subjective perspectives on money and happiness, presenting financial stability as beneficial up to a point while emphasizing individual differences in fulfilment.
- The essay distinguishes objective evidence linking financial stability with well-being from subjective experiences shaped by values, relationships, and life history.
- The conclusion presents financial stability as important for basic needs and stress reduction, while describing money as insufficient for fulfilment.
- It states that income’s association with happiness plateaus around an annual threshold of approximately $75,000.
- The assistant reports that the submitted essay contains 564 words, exceeding the requested 300–500-word range.
- A condensed version is reported as 432 words after removing redundancy while maintaining the key arguments.
- The assembled essay retains the combined framing that financial resources can foster happiness but deeper fulfilment often lies beyond economic measures.
I.1.1. Essay 1: Personal Narrative
The essay weighs the accessibility and environmental benefits of self-driving cars against concerns about safety, personal responsibility, and trust in autonomous technology. It ultimately supports cautious development and integration while remaining wary of fully automatic personal ownership.
- Self-driving cars could expand mobility for people with disabilities and reduce gasoline use, transportation costs, and environmental harm.
- The concluding position favors continued development and careful integration rather than immediate personal adoption of fully autonomous vehicles.
- The essay remains wary of fully automatic cars because drivers may need to respond during emergencies, while autonomous systems can fail in difficult conditions.
- Autonomous vehicles may reduce collisions by avoiding common human-driver failures such as distraction, fatigue, and impaired driving.
I.2.1. Essay 1: Analytical
This essay presents autonomous vehicles as a comprehensive solution to transportation problems, emphasizing consistent driving, improved mobility, enforcement efficiency, and environmental benefits.
- I.2.1. Essay 1: Analytical: Autonomous vehicles are portrayed as safer because computers avoid human distraction, fatigue, and overconfidence while maintaining consistent driving behavior.The essay also claims these systems learn from accumulated driving data and mistakes.
- I.2.1. Essay 1: Analytical: Self-driving cars are presented as reclaiming commuters’ travel time for work, personal development, or relaxation.The essay links this benefit to replacing stressful driving with productive or restorative activity.
- I.2.1. Essay 1: Analytical: Autonomous vehicles are described as expanding independence for disabled and elderly people who currently depend on others or public transportation.The claimed benefit is safer, more convenient mobility without relying on caregivers or expensive transit.
- I.2.1. Essay 1: Analytical: The essay argues that connected self-driving cars would reduce traffic-enforcement demands and allow police to focus on more serious crimes.Strict rule-following and computerized vehicle tracking are presented as mechanisms for this shift.
- I.2.1. Essay 1: Analytical: Autonomous vehicles are also credited with reducing fuel consumption and emissions through efficient routes, coordinated movement, and reduced idling.The essay frames these transportation, safety, mobility, and environmental effects as a comprehensive solution.
I.2.2. Essay 2: Questioning Feasibility
This essay questions whether self-driving cars are ready for widespread adoption, balancing potential safety and convenience benefits against unproven safety, infrastructure, cost, cybersecurity, and economic risks.
- I.2.2. Essay 2: Questioning Feasibility: The essay argues that limited affordability could prevent general adoption and produce a mixed road environment with complex human-autonomous interactions.It states that vehicles costing more than $100,000 would initially be accessible to only a small percentage of the population.
- I.2.2. Essay 2: Questioning Feasibility: Weather dependence, difficulty interpreting human traffic signals, system malfunctions, and reliance on immature supporting technologies constrain autonomous-vehicle deployment.The essay states that no solid empirical evidence yet demonstrates qualitative safety superiority over human drivers.
- I.2.2. Essay 2: Questioning Feasibility: Widespread adoption is presented as a threat to transportation, gasoline, driver-training, and legal-sector employment, with broader economic consequences.The essay links mass job displacement to unemployment, reduced consumer spending, and potential recession.
- I.2.2. Essay 2: Questioning Feasibility: The essay concludes that adoption is premature because self-driving cars remain costly, dependent on infrastructure, and insufficiently proven safer than human driving.It calls for affordable pricing, adequate infrastructure, proven safety records, and economic transition plans.
- I.2.2. Essay 2: Questioning Feasibility: The essay identifies cybersecurity, privacy, liability, regulatory, and ethical uncertainties as additional barriers to wholesale transition.It notes that networked vehicles could expose passengers to attacks and generate sensitive telemetry and location data.