Source-linked AI summary
Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models
Paramveer S. Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub, Shaochun Zheng, Lionel P. Robert
TL;DR
The paper asks how different levels of AI scaffolding affect co-writing and which writers benefit from each level. Using a within-subjects Latin square field experiment with 131 participants across no-AI, next-sentence, and next-paragraph conditions, it finds a U-shaped pattern: high scaffolding improves writing outcomes, while scaffolded tools reduce ownership and satisfaction. The authors therefore emphasize balancing effective assistance with user agency and extending evaluation across more diverse settings.
Problem
Existing research offers limited insight into how varying intensities of AI input influence writing outcomes and user experience, despite concerns that too little or too much assistance may be inadequate.
Method
A within-subjects Latin square field experiment compared no AI assistance, next-sentence suggestions, and next-paragraph suggestions among participants completing argumentative writing tasks.
Results
The study found a U-shaped relationship: low-level scaffolding did not significantly improve writing quality or productivity, whereas high-level scaffolding significantly benefited these outcomes, especially for non-regular and less tech-savvy writers.
Takeaways & Limitations
AI writing tools should balance scaffolding that improves productivity and quality with active user involvement and preservation of agency.
Takeaways & Limitations
The authors call for studies spanning more diverse conditions, genres, populations, timeframes, and naturalistic settings to improve generalizability.
Abstract
from arXiv · showhide
Advances in language modeling have paved the way for novel human-AI co-writing experiences. This paper explores how varying levels of scaffolding from large language models (LLMs) shape the co-writing process. Employing a within-subjects field experiment with a Latin square design, we asked participants (N=131) to respond to argumentative writing prompts under three randomly sequenced conditions: no AI assistance (control), next-sentence suggestions (low scaffolding), and next-paragraph suggestions (high scaffolding). Our findings reveal a U-shaped impact of scaffolding on writing quality and productivity (words/time). While low scaffolding did not significantly improve writing quality or productivity, high scaffolding led to significant improvements, especially benefiting non-regular writers and less tech-savvy users. No significant cognitive burden was observed while using the scaffolded writing tools, but a moderate decrease in text ownership and satisfaction was noted. Our results have broad implications for the design of AI-powered writing tools, including the need for personalized scaffolding mechanisms.
1 INTRODUCTION
The paper examines how different granularities of AI assistance shape writing outcomes and user experience. It compares sentence- and paragraph-level scaffolding to identify which support benefits which writers.
- Research motivation: AI writing tools can provide richer, context-sensitive collaboration than traditional word-level autocomplete.Large language models generate coherent text spanning multiple sentences or paragraphs.
- Research motivation: Writers may need different scaffolding levels because minimal assistance can be inadequate while extensive generation may reduce critical engagement and ownership.The paper motivates research into support that helps writers while preserving their control over the text.
- Research motivation: The study asks which granularity of AI support is most effective and for which groups of writers.It focuses on next-sentence and next-paragraph suggestions as natural compositional levels.
- Study approach: The experiment uses a within-subjects Latin square design in which participants complete argumentative writing tasks under no-AI, next-sentence, and next-paragraph conditions.The design exposes each participant to every treatment while mitigating order effects.
- Key findings: Low-level next-sentence scaffolding did not significantly improve writing quality or productivity, whereas high-level next-paragraph scaffolding produced marked improvements.The findings describe a U-shaped relationship between scaffolding level and writing quality and productivity.
- Key findings: Paragraph-level benefits were especially pronounced among non-regular writers and users less familiar with advanced technology.The paper links these benefits to more holistic support that can reduce cognitive load and focus attention.
- Key findings: Scaffolded writing tools moderately reduced perceived text ownership and overall satisfaction.The finding motivates adaptive and personalized scaffolding that balances assistance with user agency.
2 RELATED WORK
Prior work progresses from lightweight writing assistance to generative systems that actively collaborate with writers. This paper extends scaffolding research by varying AI support granularity and evaluating both written products and the collaborative process.
- AI writing tools and collaboration: Writing assistance has evolved from spelling correction and phrase completion toward open-ended generative collaboration.Generative models support longer, more inventive responses across writing and creative tasks.
- AI writing tools and collaboration: Earlier studies show that generative AI can help organize thoughts, establish narrative frameworks, and contribute directly to text generation.Examples include systems for co-writing stories and developing overarching themes.
- Research contributions: The paper contributes quantitative, human, and automated evaluation of co-authored texts, including measures of cohesion and linguistic sophistication.TAACO and TAALES provide automated assessments alongside human evaluation.
- Research contributions: Its evaluation examines multiple levels of AI scaffolding and considers both final text quality and the writing process.This broadens analysis beyond the end product to the dynamics of human-AI collaboration.
- AI writing tools and collaboration: HCI research also investigates how co-authoring tools alter creative writing processes and how users with different language proficiencies use them.These concerns situate the paper within broader human-AI collaboration research.
- Scaffolding strategies: Scaffolding is temporary instructional support intended to enhance task performance and help learners attain skills difficult to achieve independently.Related work emphasizes adapting support to individual skill levels and needs.
- Scaffolding strategies: Existing education research studies scaffolding through high- and low-structure support, feedback, process writing, revision, and mediated learning.This literature provides the conceptual foundation for studying AI-based writing guidance.
- Scaffolding strategies: The paper introduces AI-based scaffolding through next-sentence and next-paragraph suggestions, varying assistance granularity in a contextual manner.This approach extends scaffolding literature into human-AI co-writing.
3 METHODS
The study is a field experiment of AI-assisted argumentative writing using three scaffolding conditions and controlled prompt sequences. Participants interact with a custom editor that records writing behavior and supports selecting, rejecting, and editing suggestions.
- Study design: A screened participant pool of N=131 completed argumentative writing tasks with varying levels of AI assistance.The custom tool collected user inputs and evaluation metrics during the field experiment.
- Procedure: The study included live virtual writing sessions overseen by facilitators who introduced instructions and answered participant questions.Sessions accommodated between 1 and 10 participants.
- Study design: Participants experienced no AI assistance, next-sentence suggestions, and next-paragraph suggestions as the three treatment levels.These conditions represent control, low scaffolding, and high scaffolding.
- Study limitations: Many selected participants could not complete the main study because of scheduling conflicts.This recruitment issue constrains the realized sample and should be considered when interpreting generalizability.
- Study design: The within-subjects Latin square design exposed each participant to all conditions while distributing order sequences to reduce order bias.Participants were assigned to three ordered sequences and received unique prompts to limit response learning.
- Tasks and prompts: Participants responded to argumentative prompts selected from The New York Times prompt set, with each main task requiring at least 250 words.The study used 10 accessible and balanced prompts.
- Interface and implementation: The custom tool combined a React interface, the Lexical editor, and GPT-3 DaVinci text completion.The editor supported standard text interaction, timing, and word-count tracking.
- Interface and implementation: In AI conditions, participants could request five suggestions, select or dismiss them, and modify or partially adopt selected text.Suggestions were displayed in a panel and inserted at the cursor when selected.
3.3 Participant Recruitment
The study recruited participants from a large public-university health system, using screening and live-session oversight to support engagement and data integrity. The final main-study sample included 131 participants with varied demographic, writing, and technology-use profiles.
- Recruitment: Participants were recruited from a health-system population of around 50,000 individuals at a large public university in the United States.The pool included healthcare workers, administrative and support staff, and some former or current patients.
- Recruitment: The researchers avoided Mechanical Turk and Prolific because remote participation could compromise representativeness, data quality, and oversight.They instead conducted more than 100 hours of live Zoom sessions to monitor engagement and prevent outside AI use.
- Screening: After screening interested volunteers for writing effort, scheduling constraints reduced the final main-study sample to 131 participants.Initially, 890 people expressed interest, 453 completed the invitation response, and roughly 200 passed screening before the final reduction.
- Participant profile: All participants were U.S. adults, and main-task compensation ranged from $10 to $19.Participants received a fixed $15 for pre-screening; the main-task total included quality-based bonuses.
- Participant profile: The study collected demographics, English and writing proficiency, prior AI experience and attitudes, and related participant characteristics.These measures were intended to characterize the sample and assess how language and writing skills relate to interaction with the AI tool.
4 EMPIRICAL ANALYSIS
The analysis evaluated writing quality, investment, efficiency, and persuasion outcomes across the experiment’s AI-scaffolding conditions. Generalized linear models accounted for condition, sequence position, participant expertise, technology use, and prior condition exposure.
- Outcome variables: The study analyzed output quality, emotional and cognitive investment, task efficiency, and persuasion as broad outcome categories.Measures included text quality, satisfaction, ownership, NASA cognitive load, productivity, and AI influence.
- Outcome variables: Productivity was defined as total words written per unit time, while influence measured how much the AI suggestions shaped the user response.Both measures belonged to the task-efficiency and persuasion outcomes.
- Statistical models: Model 1 estimated each outcome using condition and sequence position, with no-AI support serving as the reference condition.The analysis used generalized linear models, clustered standard errors by user, and applied false-discovery-rate correction.
- Statistical models: Models 2 and 3 tested whether writing expertise and technology use interacted with scaffolding condition for text quality and errors.Expertise categories were not regular, regular, and proficient writers; technology use was modeled separately.
- Statistical models: Model 4 examined whether the previously encountered condition and sequence position predicted text quality.The previous condition could be none, baseline, sentence, or paragraph.
5 RESULTS
AI scaffolding produced contrasting effects across writing outcomes: paragraph-level assistance improved quality and productivity, whereas sentence-level assistance did not consistently help and reduced quality relative to baseline. Scaffolding also reduced satisfaction and ownership without significantly changing cognitive load, with effects varying by writer expertise and technology use.
- Editing and errors: -334.1 edits for sentence-level and -795.04 edits for paragraph-level suggestions indicated fewer edits under both AI-assisted conditions.Both reductions were significant at p<0.0001, consistent with refining pre-generated text rather than composing from scratch.
- Editing and errors: No statistically significant difference in spelling errors appeared across the three conditions.The AI-assisted drafts did not noticeably change the types of errors made during editing.
- Writing quality: -0.29 quality points for sentence-level suggestions contrasted with +0.18 quality points for paragraph-level suggestions relative to baseline.Both effects were statistically significant at p=0.02, producing a U-shaped pattern in writing quality.
- User experience: Satisfaction decreased by -2.14 with sentence scaffolding and -1.89 with paragraph scaffolding relative to control, while ownership decreased by -3.65 and -5.65, respectively.Paragraph scaffolding improved quality but produced the larger ownership decrease; cognitive load did not significantly change across conditions.
- Productivity and influence: +0.07 words per unit time occurred in the paragraph condition, while sentence-level assistance showed no significant productivity advantage over baseline.Paragraph-level suggestions also had greater influence on the final product: 5.91 versus 5.07 for sentence-level suggestions.
- Writer differences: Paragraph-level suggestions improved text quality by 0.53 for non-regular writers and 0.07 for regular writers, with no significant quality effect for professional writers.AI conditions slightly increased errors among regular and professional writers by around three compared with no AI assistance.
- Writer differences: Paragraph-level assistance increased text quality by 0.55 for users with no technology expertise, whereas sentence-level assistance decreased quality by -0.38 for users with basic technology expertise.No significant error variation was found across technology-use groups.
- Sequence effects: Text quality increased by 0.37 as participants progressed through conditions, while quality dropped when transitioning from paragraph assistance back to the no-AI baseline.The authors relate these patterns to interface familiarity and possible reliance on AI support.
6 DISCUSSION
The discussion finds that AI scaffolding can improve writing outcomes, but its effects depend on scaffolding level, user characteristics, and the balance between assistance and human agency. It therefore emphasizes adaptive, human-centered systems while noting important limits on generalizability.
- Summary of Key Results and Potential Mechanisms: A U-shaped relationship linked scaffolding level with writing quality and productivity: low scaffolding adversely affected both, whereas high scaffolding produced significant benefits.The discussion frames these trends as reflecting different mechanisms of support across scaffolding levels.
- Summary of Key Results and Potential Mechanisms: Paragraph-level scaffolding improved writing outcomes but reduced users’ satisfaction and sense of ownership, creating a trade-off between efficiency and agency.The authors connect this trade-off to the amount of AI-generated content and users’ perceived effort and creativity.
- Summary of Key Results and Potential Mechanisms: Moving from sentence-level to paragraph-level suggestions improved writing quality, suggesting that initial simpler interactions may help users learn to use more sophisticated support.The discussion interprets this sequence effect as a possible learning effect rather than as a general guarantee of improved performance.
- Implications for Human-AI Co-writing: The authors advocate adaptive and personalized scaffolding that adjusts assistance to users’ evolving skills, preferences, and writing experience.They specifically argue that beginners or less frequent writers may need more support, while experienced writers may need less to preserve independence.
- Limitations and Generalizability: The study’s generalizability is constrained by its limited scaffolding levels, argumentative test prompts, predominantly proficient English-speaking sample, short-term sessions, and artificial experimental setting.The authors call for broader conditions, genres, populations, timeframes, and naturalistic settings in future research.
- Implications for Human-AI Co-writing: Responsible deployment also requires attention to plagiarism, over-reliance, bias, transparency, and the balance between augmenting and replacing human effort.The discussion places human well-being and dignity at the center of ethical considerations for AI writing assistants.
A.1 Pre-screening HAI Collaborative Writing Task
The pre-screening task asks participants to read two opposing opinions about whether U.S. wood companies should pursue eco-certification, then summarize how the second opinion challenges the first. The task sets a 20-minute response period.
- Participants read an initial passage about eco-certification and then a second expert opinion on the same issue.The materials present the first opinion before directing participants to read the second expert’s view.
- The second expert argues that U.S. consumers may value independent certification claims, environmental protection, and internationally competitive products.The opinion challenges assumptions about advertising, price, and international competition.
- Participants must explain how the second expert’s points cast doubt on specific claims in the first expert’s opinion.The response is evaluated for presenting the second opinion and its relationship to the first.
- 20 minutes are allotted for the response.
A.2 Pre-screening Task Rubric
The rubric evaluates whether responses accurately and coherently connect the second expert’s arguments to relevant claims in the first opinion. Scores range from 5 for strong, well-organized integration to 0 for responses that are copied, blank, or unrelated.
- The rubric treats the second expert’s opinion as the lecture and the first expert’s opinion as the reading.
- Score 5 requires selecting important information from the second opinion and accurately connecting it to relevant information from the first.Responses should be coherent, well organized, and precise despite only occasional language errors.
- Scores 4 and 3 reflect progressively greater omissions, vagueness, imprecision, or language problems in presenting the connection between the two opinions.A score-3 response conveys some relevant connection but may be globally unclear or omit a major point.
- Scores 2 and 1 indicate substantial omission, inaccuracy, disconnection, or difficulty conveying meaningful content from the second opinion.These responses may significantly misrepresent or omit the overall relationship between the lecture and reading.
- Score 0 applies when a response copies the reading, rejects the topic, uses a foreign language, consists of keystrokes, or is blank.
B THE CUSTOM-BUILD INTERFACE: SURVEYS AND INTERACTION PATTERNS
Figure 3 presents the pre-task and post-task surveys, while Table 9 documents the user-interface interactions and controls for the AI-assisted sentence and paragraph modes.
- Figure 3 shows the surveys administered before and after the task.
- Table 9 records interface interaction details and controls for the sentence and paragraph AI-scaffolding modes.
C MAIN WRITING STUDY ARGUMENTATIVE PROMPTS
The main writing study uses argumentative prompts covering varied subjects, with examples addressing technology, school policies, stereotypes, audiobooks, college athletes, and extreme sports. The materials also identify the study and its interface documentation.
- 10 argumentative writing prompts were selected to cover a wide range of subjects.
- The study is identified as CHI ’24, and the documentation includes the custom-build interface and user-AI interaction actions.
- The prompt collection was derived from materials used in the Co-Author paper.
- Example prompts ask about screen time during the pandemic, technology’s effects on relationships, and whether schools should provide free pads and tampons.
- Other prompts address stereotypical characters, audiobook listening versus reading, whether college athletes should be paid, and risky extreme sports.
D GPT PARAMETERS
The experiments used GPT-3 davinci as the backend large language model because it was state-of-the-art when the study was conducted.
- GPT-3 davinci served as the backend large language model for the experiments.
- The authors selected GPT-3 because it was state-of-the-art when the study was performed.
- The GPT-3 parameter settings used in the experiments are reported in Table 10.