Source-linked AI summary

Co-Writing with Opinionated Language Models Affects Users' Views

Maurice Jakesch, Advait Bhat, Daniel Buschek, Lior Zalmanson, Mor Naaman

arXiv:2302.00560v1cs.HCcs.AIcs.CL

TL;DR

Opinionated language models may influence users’ views, but the scale of this effect is uncertain. This study tests the issue in a controlled writing experiment and finds that model-preferred opinions affected participants’ writing and later survey attitudes.

  • Problem

    The study asks whether language models that generate some opinions more often affect what users write and think.

  • Method

    In an online experiment, 1,506 participants wrote about social media with GPT-3 assistance configured to support either side; judges evaluated their writing and participants completed an attitude survey.

  • Results

    Opinionated model assistance affected participants’ expressed writing opinions and shifted their subsequent attitudes toward social media.

  • Takeaways & Limitations

    The findings support monitoring and more careful engineering of opinions built into AI language technologies.

  • Takeaways & Limitations

    The study tested one topic and one GPT-3 writing-assistant implementation, so generalization to other topics, models, and applications remains open.

Abstract

from arXiv · show

If large language models like GPT-3 preferably produce a particular point of view, they may influence people's opinions on an unknown scale. This study investigates whether a language-model-powered writing assistant that generates some opinions more often than others impacts what users write - and what they think. In an online experiment, we asked participants (N=1,506) to write a post discussing whether social media is good for society. Treatment group participants used a language-model-powered writing assistant configured to argue that social media is good or bad for society. Participants then completed a social media attitude survey, and independent judges (N=500) evaluated the opinions expressed in their writing. Using the opinionated language model affected the opinions expressed in participants' writing and shifted their opinions in the subsequent attitude survey. We discuss the wider implications of our results and argue that the opinions built into AI language technologies need to be monitored and engineered more carefully.

1 INTRODUCTION

Large language models may influence opinions by making some views easier to express, creating a form of latent persuasion that can be difficult to detect. This study tests whether opinionated writing assistance affects users’ writing and attitudes.

  • Language models may influence opinions when they generate some views more often than others during everyday communication.
  • Unlike visible choice architectures, opinion preferences embedded in language models may be opaque to users, policymakers, and developers.
  • Latent persuasion operates through users’ writing, which can then be read by other people.
  • The experiment asked 1,506 participants to write about whether social media is good or bad, using assistance configured to support either side.
  • The model affected both participants’ expressed writing opinions and their later survey attitudes toward social media.

2 RELATED WORK

Related work frames language-model influence within research on social influence, online algorithms, text-entry systems, and the risks of increasingly capable generative models. Writing assistants are becoming active co-writing partners whose suggestions can affect written output.

  • Social influence and persuasion: Social influence research examines shifts in thoughts, feelings, attitudes, or behaviors resulting from interaction with others.
  • Social influence and persuasion: Online users can be influenced by human sources and algorithmic entities, with influence depending partly on perceived trustworthiness and decision uncertainty.
  • Writing assistants: Traditional text-entry systems prioritize efficiency by suggesting likely words or short continuations based on language context.
  • Writing assistants: Newer language models support broader goals such as inspiration, story writing, revision, and creative writing, positioning them as active writing partners or coauthors.
  • Writing assistants: Suggested text is associated with shorter, more predictable writing, more standard phrases, and sentiment shifts in written text.
  • Risks and societal effects: Generative language models offer beneficial applications while raising ethical and social risks, including effects on public opinion and political life.

3 METHODS

The study used an online Reddit-like writing experiment in which GPT-3 suggestions were configured to support opposing views about social media. Participants’ writing, interactions with suggestions, and subsequent attitudes were measured.

  • Experiment design: The online experiment recruited 1,506 participants to reply to a simulated social-media discussion using a writing assistant.
  • Experiment design: Social media was selected as a politically relevant topic with mixed views, while avoiding controversy likely to produce entrenched positions.
  • Conditions: Participants were randomly assigned to write without assistance or receive suggestions configured to argue that social media was good or bad for society.
  • Writing assistant: The assistant generated continuations after writing pauses, displayed them progressively, and allowed participants to accept suggestions word by word.
  • Model manipulation: The study used strongly opinionated model prompts to test whether language models could shift users’ views.
  • Measures: Outcomes included judge-coded sentence opinions, keystroke-level interaction data, a post-task attitude survey, and treatment-group experience measures.

4 RESULTS

The results section analyzes participants’ expressed opinions first, then examines suggestion acceptance, later survey attitudes, and perceptions of the model’s opinion and influence.

  • The analysis proceeds from written opinions to convenience-based suggestion acceptance, subsequent survey attitudes, and participants’ perceptions of the model.

4.1 Did the interactions with the language model affect participants’ writing?

The model’s preferred opinion affected what participants wrote: supportive suggestions increased pro-social-media statements, while critical suggestions increased anti-social-media statements.

  • Control participants argued that social media was bad in 38% of sentences and good in 28%.About 28% of control sentences argued both positions, while 11% argued neither or were unrelated.
  • Figure 4 compares sentences written by participants, co-written with the model, and fully accepted from suggestions by whether model opinions matched participants’ likely pre-task views.The figure covers 6,142 sentences from 1,000 participants.
  • 2.04 times more likely, supportive-model participants argued that social media is good than control participants.The comparison was statistically significant (p<0.0001, 95% CI [1.83, 2.30]).
  • 2.0 times more likely, critical-model participants argued that social media is bad than control participants.This difference was statistically significant (p<0.0001, 95% CI [1.79, 2.24]).

4.2 Did participants accept the model’s suggestions out of mere convenience?

Participants generally co-wrote rather than blindly accepted suggestions, but convenience likely contributed to stronger opinion differences among those who completed the task quickly.

  • 63% of sentences were written by participants without accepting model suggestions, while 25% were co-written and 11.5% were fully accepted.About one in four participants accepted no suggestions, while one in ten had more than 75% of their post written by the model.
  • Participants whose views likely aligned with the model accepted more suggestions, whereas participants with opposing views accepted fewer.This pattern indicates that acceptance was related to participants’ personal views, not merely indiscriminate uptake.
  • Figure 5 plots mean expressed opinion against writing time, with opinions scaled from -1 for bad to 1 for good.The figure’s right panel shows how treatment differences vary with time, while the left panel aggregates across writing times.
  • 0.38 was the treatment-group opinion difference for participants who took less than 160 seconds, compared with 0.29 across writing times.The overall difference had d=0.5; the quick-completion difference was statistically significant (p<0.001, 95% CI [0.31, 0.45]).
  • 0.20 remained the opinion difference for participants who took four to six minutes, corresponding to d=0.34.This difference remained significant (p<0.001, 95% CI [0.13, 0.27]), indicating that convenience did not explain all observed differences.

4.3 Did the language model affect participants’ opinions in the attitude survey?

The model’s stance during writing was reflected in participants’ later social-media attitude reports: supportive interactions increased pro-social-media responses, while critical interactions increased anti-social-media responses.

  • Figure 6 displays subsequent survey answers by whether participants received supportive or critical suggestions during writing.Orange indicates saying social media was good, blue indicates not good, and white indicates undecided responses.
  • 45% of participants exposed to a supportive model later said social media was good for society, compared with 35% in the control group.The difference corresponded to d=0.22 (p<0.001).
  • Participants exposed to a critical model were more likely to say afterward that social media was bad for society, with d=0.19.The difference was significant at p<0.005.
  • Participants interacting with a supportive model evaluated social media more favorably throughout their statements, while critical-model participants were more critical from their first sentence.The writing assistant augmented rather than replaced participants’ narratives, although the study could not ascertain the persuasion mechanism.

4.4 Were participants aware of the model’s opinion and influence?

Participants often failed to recognize the model’s opinion or its influence, even when the model was perceived as knowledgeable and balanced. Awareness varied with whether the model supported participants’ views.

  • 84% of participants whose views were supported by the model said the assistant was knowledgeable and had expertise.
  • The model affected participants’ writing equally across earlier and later sentences in their posts.
  • When the model contradicted participants’ views, only 15% said it was not knowledgeable or lacked expertise.
  • Only 10% of participants whose opinions were supported by the model noticed that its suggestions were imbalanced.
  • 30% noticed the model’s skew when it contradicted their opinions, although more than half still considered its suggestions balanced and reasonable.
  • Only about 20% of participants who disagreed with the model believed it influenced their writing, compared with 34% among those aligned with it.

4.5 Robustness and validation

The study validated that the prompting manipulation produced opinionated suggestions and examined whether participants’ behavior could instead reflect experimenter demand effects. The authors report evidence against demand effects threatening the results.

  • 86% of full sentences suggested by the model supported the intended view, while 8% were labeled balanced.
  • Only about 14% of participants identified the assistants’ effect on people’s opinions as a possible study purpose.
  • Most participants believed the model did not affect their argument and were unaware of its opinion, reducing concern that they adapted views to satisfy experimenters.

5 DISCUSSION

The discussion frames opinionated language models as a form of latent persuasion that can shift users’ writing and subsequent attitudes, while emphasizing unresolved mechanisms, generalizability limits, and misuse risks.

  • Findings: Participants assisted by an opinionated model were more likely than controls to support the model’s view in their posts, and later differed in attitude surveys.The survey shifts suggest that writing differences were associated with changes in personal attitudes.
  • Conceptual framing: The authors call this influence latent persuasion because model preferences can be opaque and embedded in ordinary writing assistance.Unlike visible choice architectures, opinion preferences built into language models may be difficult for users, policymakers, and developers to identify.
  • Mechanisms: Secondary findings suggest the influence was at least partly subconscious and may have changed participants’ opinion-formation process behaviorally, but its mechanisms remain uncertain.The authors discuss informational, normative, and behavioral pathways while calling for further research.
  • Implications: The results raise concerns that commercial and political actors could use opinionated language technologies for targeted influence across writing assistants, keyboards, smart replies, and voice assistants.The same persuasive power could also support beneficial interventions such as reducing polarization or countering harmful false beliefs.
  • Limitations and generalizability: The experiment tested one topic and one GPT-3 writing-assistant implementation, so effects may differ for entrenched opinions, other models, applications, and repeated real-world interactions.The authors also note that influence magnitude and persistence over time remain to be established.
  • Limitations and generalizability: The strongly opinionated, single-interaction design may underestimate cumulative effects in deployments where many people repeatedly encounter and reproduce a model’s preferred views.Widespread reuse could increase the prevalence of those views in future training data.
Loading 2302.00560v1…