Source-linked AI summary

Can Good Writing Be Generative? Expert-Level AI Writing Emerges through Fine-Tuning on High-Quality Books

Tuhin Chakrabarty, Paramveer S. Dhillon

arXiv:2601.18353v1cs.AIcs.CLcs.HC

TL;DR

The paper examines how experts and lay judges assess AI-generated literary writing and how AI preferences affect writers’ understanding of good writing. In a controlled comparison, expert preference for human writing under in-context prompting reversed after fine-tuning, while lay readers preferred AI writing.

  • Problem

    The study asks how experts and lay judges define good writing and how discovering an AI preference affects writers’ professional identity and aesthetic understanding.

  • Method

    The authors conducted blind pairwise evaluations of human and LLM-generated excerpts by expert and lay judges, supplemented by interviews with participating writers.

  • Results

    Expert preference for human writing reversed with fine-tuning, while lay readers preferred AI; interviews found writers reframed writing’s value around process and intention.

  • Takeaways & Limitations

    The findings challenge assumptions that good writing cannot be generative and raise questions about how creative work’s value is understood.

  • Takeaways & Limitations

    The study evaluates excerpts of up to 450 words, which may not capture plot, character development, or broader narrative structure.

Abstract

from arXiv · show

Creative writing has long been considered a uniquely human endeavor, requiring voice and style that machines could not replicate. This assumption is challenged by Generative AI that can emulate thousands of author styles in seconds with negligible marginal labor. To understand this better, we conducted a behavioral experiment where 28 MFA writers (experts) competed against three LLMs in emulating 50 critically acclaimed authors. Based on blind pairwise comparisons by 28 expert judges and 131 lay judges, we find that experts preferred human writing in 82.7% of cases under the in-context prompting condition but this reversed to 62% preference for AI after fine-tuning on authors' complete works. Lay judges, however, consistently preferred AI writing. Debrief interviews with expert writers revealed that their preference for AI writing triggered an identity crisis, eroding aesthetic confidence and questioning what constitutes "good writing." These findings challenge discourse about AI's creative limitations and raise fundamental questions about the future of creative labor.

1 Introduction

The study examines what counts as “good writing” when human experts and LLMs emulate the styles and voices of 50 acclaimed authors. It combines blind evaluations of writing quality and stylistic fidelity with analyses of judges’ rationales and writers’ responses to AI authorship.

  • Study design: 28 expert writers and three LLMs were assigned to produce 200–450-word excerpts emulating the styles and voices of 50 critically acclaimed authors.The authors represented diverse cultural backgrounds and age groups.
  • Study design: The study compared in-context prompting with fine-tuning, which modifies models by training them on authors’ books and requires more resources.Figure 1 presents the writing, evaluation, and debrief phases.
  • Evaluation: 131 lay participants and the 28 writers evaluated excerpts blindly on writing quality and stylistic fidelity through planned human-versus-AI pairwise contrasts.The writers also served as expert judges, while lay participants were recruited through Prolific.
  • Research questions and implications: The research also investigates how writers reconcile AI preferences with their judgments, professional identities, and understandings of good writing, alongside implications for education, publishing, and literary platforms.The discussion connects these questions to AI style emulation, fair use, labor-market dilution, hidden AI authorship, and the future of creative work.
  • Judgment criteria: Lay judges primarily cited surface qualities such as flow, organization, clarity, and emotional impact, whereas experts analyzed narrative voice, character interiority, imagery, syntax, and how technical elements serve.These rationales address the study’s question about how different judges justify what qualifies as good writing.

2 RELATED WORK

Related work shows that LLMs are increasingly used across professional, scientific, and creative writing, while book-based training and style reproduction raise ethical concerns. This study extends prior work by testing whether AI-generated writing can surpass professional human writing for expert readers and by examining writers’ psychological responses.

  • LLM writing applications: 10-24% of content in consumer complaints, business communications, job listings, and UN press statements involved LLM assistance by late 2024.LLMs are also used in scientific research, collaborative fiction, and other artistic text generation.
  • Books as training data: GPT-3 and Llama were trained on substantial book collections, establishing a longstanding precedent for using books as LLM training data.GPT-3 used BooksCorpus alongside Common Crawl and WebText, while Llama reported 177 GB of book data from Project Gutenberg and Books3.
  • Ethical and social concerns: Writers’ interviews identify support for creative chains and respect for human creativity while highlighting tensions involving limited control, industry-scale effects, and power imbalance.These concerns arise in response to contentious practices surrounding LLM training data.
  • Study contribution: Unlike prior work on training-data attitudes, copyright, and stakeholder perspectives, this study tests whether AI writing leads experts to prefer it over professional human writing and examines writers’ psychological reconciliation with those preferences.The study therefore connects AI writing quality with writers’ aesthetic self-understanding.
  • Creative-work ramifications: Prior studies show that generative AI can reduce collective diversity and reproduce individual artists’ styles, raising concerns about homogenization in creative work.Style-transfer outputs have also been analyzed as “boundary objects.”

3 Methodology

The study compared expert human writers with LLMs using controlled style-emulation tasks across 50 authors, evaluating both in-context prompting and fine-tuning. It used short excerpts, balanced author assignments, and expert and lay judges to assess the outputs.

  • Human writers: 28 writers—27 with completed or ongoing MFAs at top programs—produced style-emulation excerpts, representing established and emerging literary voices.Sixteen identified as female, 11 as male, and one as non-binary; nearly all had fiction published in prestigious literary magazines.
  • Writing task: 150 human excerpts were collected for 50 authors, with each author assigned to exactly three writers and excerpts constrained to 200–450 words.Writers selected the authors they wished to emulate and were compensated $75 per writing task.
  • AI conditions: Three LLMs—GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro—performed the same task under in-context prompting, while GPT-4o was used for fine-tuning.The prompting condition used long-context versions of the identical writing prompt provided to human writers.
  • AI conditions: Fine-tuning used complete works from 30 living authors, segmented into context-independent excerpts of 250–650 words.Authors were selected for cultural relevance and the feasibility of acquiring and training on their books.
  • Scope and limitations: The study focused on excerpts up to 450 words to isolate voice, style, and quality, while acknowledging that plot, character development, and structure may be underrepresented.This length was also chosen because current LLMs cannot automatically balance coherence, quality, thematic consistency, and pacing in long-form narratives.
  • Evaluation: Evaluation used 28 expert writer judges and 131 college-educated lay judges from the USA and UK.The expert judges were the recruited writers, all of whom had teaching experience in US undergraduate writing classes.

Participant privacy and Advocacy Statement

The study received IRB approval and informed consent, protected participant data through restricted access, and did not publicly share writers’ contributions. Although the researchers used legally obtained copyrighted books for training, they do not advocate this practice and emphasize its risks.

  • Participant privacy: The study was approved by the University of Michigan IRB, and all participants provided informed consent.Approval identifier: HUM00264127.
  • Participant privacy: No writer-provided data was shared publicly, while evaluation data and interview files were stored on secure servers.
  • Advocacy statement: The researchers trained models on legally obtained copyrighted books but explicitly stated that they do not advocate this practice.The authors said their results highlight risks associated with fine-tuning models on high-quality data.

4 Findings

Expertise shaped evaluation before fine-tuning: experts strongly preferred human writing, whereas fine-tuning led both evaluator groups to prefer AI without a significant evaluator-type difference. Experts justified preferences through detailed, line-level reasoning, while AI’s performance unsettled writers’ criteria, confidence, and professional identity.

  • Evaluator Preferences: 82.7% for Writing Quality and 72.0% for Stylistic Fidelity: expert judges preferred human writing in the In-context condition.Lay judges showed weaker or no clear preference, at 44.0% and 52.7%, respectively; both within-condition differences were statistically significant (p< 0.001).
  • Evaluator Preferences: After fine-tuning, both evaluator groups preferred AI-generated writing, with no significant difference between evaluator types (p> 0.05).The expert-lay gap shrank from 38.7% to 6.7% for Writing Quality and from 19.3% to −1.1% for Stylistic Fidelity.
  • Evaluation Rationales: Lay judges primarily rewarded flow, organization, clarity, conciseness, straightforwardness, and general emotional impact rather than the mechanisms producing those effects.Their reactions could conflict with an author’s style, as one judge criticized David Foster Wallace’s deliberately maximalist, very long sentences.
  • Evaluation Rationales: Experts wrote 126 versus 79 words on average and more often used quotations, line-level evidence, contrastive frames, and analysis of narrative perspective.Experts also converged on shared judgments, identifying the same flaws or effective imagery in particular emulations.
  • Writer Reactions: Revealed AI successes prompted criteria reframing, process attribution, technical sensemaking, expectation violation, and capability reassessment among expert writers.Participants contrasted AI’s polished first attempts with human revision, and explained performance through training data, model scale, patternmatching, or computational resources.
  • Writer Reactions: Choosing AI over human writing destabilized writers’ understanding of “good writing,” eroded aesthetic and diagnostic confidence, and unsettled their professional identity.Several writers reported that fine-tuned outputs no longer matched their familiar “feel” of AI writing, producing surprise, fear, and diminished confidence in distinguishing authorship.

5 Discussion

The discussion frames AI writing within longstanding human influence and intertextuality while raising unresolved questions about copyright, disclosure, detection, and human steering of fine-tuned systems. It argues that AI’s speed and capacity to emulate styles create challenges for creative practice and regulation.

  • Human influence and intertextuality: Intertextuality treats every text as a “mosaic of quotations,” emphasizing that writers are continually influenced by what they read.Harold Bloom likewise described poets’ creative processes as shaped by their relationships with precursor poets.
  • Copyright and fair use: AI outputs may closely mirror source styles without directly copying works, while copyright law does not grant authors exclusive rights to writing style.The passage also notes that no one has a monopoly on producing high-quality literature.
  • AI’s speed and human steering: Fine-tuning a model on 20 novels takes 3-4 hours, far faster than an average human can consume comparable information.The authors foresee humans steering fine-tuned models to produce high-quality books or novellas.
  • Disclosure and detection: Creators may conceal generative AI involvement because disclosure could reduce their prospects for copyright protection and profit.The authors call for updated beliefs about AI detection, noting that some commercial detectors are inaccurate without concluding that detection does not work.

6 Limitations and Future Work

The study’s credibility and generalizability were limited by practical recruitment constraints, including a small expert pool and primarily U.S.-based creative writing programs.

  • Limitations: A larger pool of experts would have added more credibility, but recruiting additional writers was challenging.The authors attribute this difficulty partly to their affiliation as AI researchers, which deterred writers from working with them.
  • Limitations: Recruitment was mostly restricted to American creative writing programs.This geographic and institutional concentration limits how broadly the findings can be generalized.
  • Future Work: Further study is needed across creative writing programs outside the United States.Expanding recruitment internationally is identified as a direction for future research.

7 Conclusion

A controlled experiment found that lay readers preferred AI writing, while experts preferred human writing under in-context prompting but reversed that preference after fine-tuning. These results challenge fundamental assumptions about whether good writing can be generated by AI.

  • The study compared professionally trained human writers and LLMs emulating critically acclaimed authors in a controlled experiment.
  • Lay readers, who represent a large share of the consumer base, preferred AI over human writing.
  • Experts strongly preferred human writing over AI under in-context prompting, but this preference reversed drastically after fine-tuning.
  • The findings challenge fundamental assumptions about whether good writing can be generated by AI.

A Appendix · A.1 Themes for Writing

The writing tasks covered themes from intimate personal struggles to complex societal examinations and philosophical explorations. Figure 8 indicates that several themes were complex.

  • A.1 Themes for Writing: The excerpts addressed intimate personal struggles, including loneliness, grief, and routine.
  • A.1 Themes for Writing: The tasks examined societal issues such as racial identity, post-colonialism, and class consciousness.
  • A.1 Themes for Writing: The writing included philosophical explorations of time, mortality, and truth versus fiction.

A.2 Fine-grained Human vs AI performance

Claude 3.5 Sonnet performs best for experts under in-context prompting, while GPT-4o performs best for lay judges. GPT-4o’s expert win rate rises fivefold after fine-tuning, with no significant performance difference for lay judges.

  • Claude 3.5 Sonnet is the best-performing model for experts in in-context prompting.
  • GPT-4o is best for lay judges, achieving a 12% winning rate in the in-context setup.
  • 5X: GPT-4o’s in-context expert winning rate increases fivefold after fine-tuning.For lay judges, GPT-4o shows no significant performance difference after fine-tuning.

A.3 Preference Evaluation Examples

Figures 10–12 provide examples of preference evaluations, pairing model conditions with detailed expert rationales for writing quality and stylistic fidelity.

  • Preference Evaluation Examples: Figures 10 and 11 show preference-evaluation examples for In-Context Prompting and Fine-tuned GPT-4o, with expert rationales for writing quality.The examples include detailed rationales from Experts for Writing Quality.
  • Preference Evaluation Examples: Figure 12 shows a Fine-tuned GPT-4o preference-evaluation example with expert rationales for stylistic fidelity.The example focuses on detailed rationales from Experts for Stylistic Fidelity.

A.4 User Interface Design

The study presents interfaces for evaluating writing quality and stylistic fidelity, using paired AI–human excerpts that emulate acclaimed authors. The materials also organize content themes supplied to writers and LLMs.

  • Evaluation interfaces: Figures 13 and 14 show separate interfaces for quality evaluation and stylistic fidelity evaluation.
  • Content and evaluation materials: Figure 8 presents the themes for all content provided to writers and LLMs.
  • Quality evaluation: The quality interface compares AI and human excerpts emulating Maya Angelou’s style from content based on Letter to My Daughter.
  • Stylistic fidelity evaluation: The stylistic-fidelity interface compares AI and human excerpts emulating Ben Lerner’s style from content based on an original excerpt.
Loading 2601.18353v1…