Source-linked AI summary
Co-Writing Screenplays and Theatre Scripts with Language Models: An Evaluation by Industry Professionals
Piotr Mirowski, Kory W. Mathewson, Jaylen Pittman, Richard Evans
TL;DR
Longform creative writing requires coherence that language models struggle to maintain across distant text. The paper introduces Dramatron, which hierarchically generates scripts through structural context and prompt chaining, and evaluates it through co-writing with industry professionals. The authors report that explicit narrative structures and characters help generate more coherent text while positioning the system as an interactive co-writing tool.
Problem
Language models have limited long-range semantic coherence, constraining their usefulness for longform creative writing.
Method
Dramatron hierarchically generates scripts from a log line using explicit narrative structures, characters, prompt chaining, and structured generation.
Results
Explicit narrative structures and characters help Dramatron generate more coherent text, and the system can produce complete scripts and screenplays.
Takeaways & Limitations
Dramatron is designed as an interactive, augmentative co-writing tool in which human writers can intervene during generation.
Takeaways & Limitations
Participants observed logical gaps, insufficient common sense, and limited nuance and subtext in Dramatron’s storytelling.
Abstract
from arXiv · showhide
Language models are increasingly attracting interest from writers. However, such models lack long-range semantic coherence, limiting their usefulness for longform creative writing. We address this limitation by applying language models hierarchically, in a system we call Dramatron. By building structural context via prompt chaining, Dramatron can generate coherent scripts and screenplays complete with title, characters, story beats, location descriptions, and dialogue. We illustrate Dramatron's usefulness as an interactive co-creative system with a user study of 15 theatre and film industry professionals. Participants co-wrote theatre scripts and screenplays with Dramatron and engaged in open-ended interviews. We report critical reflections both from our interviewees and from independent reviewers who watched stagings of the works to illustrate how both Dramatron and hierarchical text generation could be useful for human-machine co-creativity. Finally, we discuss the suitability of Dramatron for co-creativity, ethical considerations -- including plagiarism and bias -- and participatory models for the design and deployment of such tools.
1 INTRODUCTION
Dramatron addresses LLMs’ limited long-range coherence by hierarchically generating scripts from a log line, while keeping writers involved throughout the process. The paper evaluates this approach with theatre and film professionals and reflects on its creative and ethical implications.
- Motivation: LLMs struggle with long-range dependencies because their context windows limit access to information from many pages earlier.The cited state-of-the-art context window is at most 2048 tokens, or about 1500 words.
- Approach: Dramatron uses hierarchical story generation, prompt chaining, and structured generation to improve coherence across an entire script.The approach combines generated structural context with carefully designed prompts rather than relying on flat sequential generation.
- Approach: From a user-provided log line, Dramatron can generate a title, characters, plot beats, location descriptions, and dialogue, producing scripts sometimes tens of thousands of words long.Users can intervene at each stage by editing, rewriting, soliciting alternatives, or continuing generation.
- Evaluation: The system was evaluated through two-hour co-writing sessions with 15 theatre and film industry professionals rather than non-expert crowd raters.Participants provided feedback on interactive co-authorship and the outputs, and their input informed iterative system improvements.
- Evaluation: The study included scripts staged at the Edmonton International Fringe Theatre Festival and reflections from creative teams and professional reviewers.The authors frame these materials as critical reflections on human-machine co-creativity.
- Implications: The paper presents Dramatron as a pathway toward human-machine co-creativity intended to uplift human writers and artists while raising ethical questions such as plagiarism and bias.It also discusses participatory models for designing and deploying such tools.
2 STORYTELLING, THE SHAPE OF STORIES, AND LOG LINES
The paper situates dramatic writing within modular narrative structures and uses plot beats to support hierarchical generation. It begins generation from log lines, while acknowledging that its chosen structure is culturally specific rather than universal.
- Narrative structures: A plot is a sequence of actions in which each point is coherent with and consequential to previous points.Aristotle’s simple structure divides plot into beginning, middle, and end.
- Narrative structures: Freytag’s pyramid organizes dramatic progression through exposition, inciting incident, rising action, climax, falling action, resolution, and dénouement.The study chose this structure because it fits dramatic scripts and was considered familiar to participating playwrights.
- Narrative structures: Narrative theories describe stories as compositions of large, modular elements that have informed computational narratology and automated storytelling.The cited traditions include structuralist narratology and personal-experience narratives.
- Limitations: The authors state that their chosen narrative structure is not universal and is characteristically Western, while alternative story shapes remain possible.The Hero’s Journey is presented as one such alternative.
- Narrative structures: The study also illustrates the Hero’s Journey or Monomyth as an alternative narrative structure.Dramatron leverages plot beats in its hierarchical generation for narrative coherence.
- Log lines: A log line summarizes a screenplay or theatre script in a few sentences, typically covering setting, protagonist, antagonist, conflict or goal, and sometimes the inciting incident.Dramatron uses log lines as starting inputs because they encode answers to who, what, when and where, how, and why.
3 THE USE OF LARGE LANGUAGE MODELS FOR CREATIVE TEXT GENERATION
Language models can generate text probabilistically from prompts but struggle with long-term semantic coherence. Dramatron addresses this through hierarchical generation and prompt chaining, enabling interactive co-writing of scripts and screenplays.
- 3.1 Language Models: Language models generate text by sampling probabilistically from conditional token distributions, so different random seeds produce different outputs.Figure 3 illustrates prompts concatenated with prefixes and decorated with tags such as <end>.
- 3.1 Language Models: LLMs appear coherent locally but have difficulty maintaining long-term semantic coherence, partly because context windows restrict the amount of text available.The described context limit is 2048 tokens, with memory requirements of O(n^2).
- 3.2 Hierarchical Language Generation to Circumvent Limited Contexts: Dramatron generates stories hierarchically across log line, character and plot descriptions, locations, and final dialogue.Content is generated top-down, with lower-level prompts chained to outputs from higher levels.
- 3.2 Hierarchical Language Generation to Circumvent Limited Contexts: A compressed middle plot layer helps the entire plot fit within the model context window and supports narrative arcs and dramatic closure across scenes.The method is designed to enable long-term semantic coherence without requiring human intervention.
- 3.3 The Importance of Prompt Engineering: Prompt chaining combines user inputs and earlier generations with hard-coded few-shot prompts for titles, characters, plots, locations, and dialogue.Dramatron primarily used prompt sets based on Medea and science-fiction films.
- 3.4 Interactive Co-Writing: Dramatron is designed as an interactive, augmentative co-writing tool in which humans and the system both contribute to script authorship.Users can intervene at any stage and modify the script at every level of the hierarchy.
4 EVALUATING TEXT GENERATED BY LARGE LANGUAGE MODELS
The paper reviews automated and human-centered approaches for evaluating generated text, emphasizing that long-form scripts require careful human assessment. It therefore studies Dramatron with theatre and film professionals through co-writing sessions, interviews, ratings, and edit analysis.
- 4.1 Evaluation Methods: Automated metrics assess similarity, prompt consistency, or linguistic diversity, but were not designed for screenplay- or theatre-script-length generations.This motivates a focus on human-centric evaluation.
- 4.2 Crowdsourced Evaluation: Crowdsourced evaluations can be affected by workers’ personal opinions, demographics, cognitive biases, and limited reading attention.The paper identifies these as quality and bias concerns for subjective, open-ended text assessment.
- 4.3 Expert Evaluation: The study engages 15 theatre and film professionals with AI-writing experience who co-write a screenplay or theatre script alongside Dramatron in two-hour sessions.Most participants completed both a full script and an open discussion interview within the allotted time.
- 4.3 Expert Evaluation: Participants answer adapted questions on a 5-point Likert-type scale, while interviews provide qualitative analysis of the co-writing experience.The questions are detailed in Section 6 and follow the interactive sessions.
- 4.4 Edit Analysis: The researchers track absolute and relative word edit distance and Jaccard word similarity to compare generated suggestions with human-edited drafts.These measures assess whether Dramatron contributes new ideas or mainly expands existing writer ideas.
- 4.4 Edit Analysis: Grammatical correctness is not evaluated because the paper treats the few observed LLM errors as fixable by the human co-writer.The evaluation instead emphasizes co-writing and human-centered assessment.
5 PARTICIPANT INTERVIEWS
Participants saw Dramatron as useful for inspiration, exploration, content generation, and interactive co-writing, while identifying persistent problems with coherence, nuance, stageability, and bias. They also described practical possibilities ranging from rough drafts and writers’ rooms to staged productions after editing.
- Positive uses: Participants valued hierarchical generation for working on the narrative arc, either interactively or by letting Dramatron generate material.They also identified inspiration, world building, and content generation as potential applications.
- Structural and expressive problems: Participants identified logical gaps, inconsistent character dialogue, repetition, weak motivation, limited nuance and subtext, and poor awareness of stageability.They also noted that characters could become prescriptive, dialogue overly verbalised, and inputs constrained.
- Staging and evaluating productions: Several participants thought generated scripts could be staged after substantial editing, and Participant p1 staged several Dramatron-generated scripts.Participants described the outputs as rough drafts requiring work, while p1’s productions were part of the Plays By Bots performances.
- Inspiration for the Writer: All participants found Dramatron useful for inspiration, especially for overcoming writers’ block and stimulating new ideas.Participants described outputs as indirectly stimulating creativity or providing suggestions that could be refined through back-and-forth editing.
- Exploration and content generation: Participants saw Dramatron as a way to explore characters, relationships, literary styles, and ideas beyond their initial expectations.Some also viewed analysing and editing its outputs as a potential learning activity.
- Co-writing and professional applications: Several participants considered Dramatron useful for producing material, reviving stalled projects, and supporting formulaic television writing rooms.They discussed possible uses for long-running series, synopses, dramaturgical support, and rapid script generation.
- Ethical and representational concerns: Participants reported stereotyped and biased outputs, including gender bias, ageism, misogynistic content, and concerns about cultural appropriation.Some participants responded by using names that were not gender- or ethnicity-specific.
6 PARTICIPANT SURVEYS
Thirteen of 15 participants completed the post-session survey, which assessed collaboration, usability, creative expression, ownership, and reactions to outputs. Responses were generally positive about surprise, enjoyment, helpfulness, collaboration, and creative expression, while ownership and ease of use were more divided.
- Survey results: 77% agreed that they enjoyed writing with the AI, 69% that the scripts felt unique, and 69% strongly agreed that they enjoyed writing.Uniqueness referred to lines of dialogue or connections between characters.
- Survey results: 84% found the AI helpful, while 77% felt they were collaborating with it and 77% could express their creative goals.The corresponding strong-agreement rates were 38%, 54%, and 23%.
- Survey results: Participants were more divided on ease of use and pride in the output: 61% agreed about ease and 46% agreed about pride.These answers were highly correlated at r = 0.9, and one participant strongly disagreed about pride.
- Ownership and use: More than half felt they did not own the result, viewing Dramatron more as a learning and inspiration tool than a complete script-writing tool.Participants described outputs as starting points or provocations that writers still needed to develop.
7 DISCUSSION AND FUTURE WORK
The discussion frames Dramatron as a useful but constrained co-creative system whose hierarchical, formulaic structure supports long-form generation while fitting some writing practices better than others. It also identifies ethical risks involving bias, copyright, creative labor, and authorship.
- Future work towards coherent story generation: The authors propose character arcs, beginning-middle-end dialogue beats, and genre-oriented prompting as ways to improve coherence and stylistic consistency.These are presented as future enhancements for complex characters, scene conclusions, and thematic outputs.
- Future work towards coherent story generation: Hierarchical generation helps address long-range coherence, but each later generation step depends on the preceding steps.Future work could allow generation steps to occur out of order through creative prompting.
- Film, theatre, and formulaic hierarchy: Participants found Dramatron’s top-down hierarchy more aligned with screenwriting than playwriting, where characters or story summaries may emerge late.Playwrights also described a more investigative, theme-led process rather than beginning with a fixed story.
- Ethical questions and risks: Participants identified three ethical risks: biased or offensive output, automation that displaces creative work, and copyright infringement involving training data.The proposed mitigations keep human artists involved and maintain transparency about generated text’s origins.
- Ethical questions and risks: Participants reported stereotypical outputs, questioned dataset provenance and plagiarism, and raised concerns that generative tools could replace writing opportunities.They nevertheless found the study’s mitigation strategies satisfactory, and none reported distress about model outputs.
- Co-creativity and participatory design: The study raises questions about authorship, cultural appropriation, and how co-creative tools should fit artists’ expected forms of collaboration.Participants connected these concerns to ownership of machine co-created text and to interactions resembling co-creation, subcontracting, or inspiration.
8 CONCLUSIONS
The conclusion presents Dramatron as an interactive co-writing tool that combines explicit narrative structure with human participation to support long-form theatre and screenwriting. The study documents professional engagement with the system and invites further examination of co-creativity and ethics.
- Conclusions: Dramatron lets writers generate scripts from a provided log line through an interactive co-writing process.The system is presented as a tool for human authors rather than an autonomous replacement for them.
- Conclusions: Hierarchical story generation with explicit narrative structures and characters helps produce more coherent text for theatre scripts and screenplays.The conclusion specifically highlights long-form outputs such as these scripts.
- Evaluation: The evaluation involved 15 theatre and film industry professionals, qualitative interviews, and a short post-session survey.The paper also reports feedback from a creative team whose co-written scripts were publicly performed and reviewed.
- Conclusions: Dramatron can be used as a co-creative writing tool for human authors writing screenplays and theatre scripts alongside language models.The conclusion frames this use as the system’s principal practical contribution.
- Conclusions: The work leaves open questions about the nature of co-creativity and the ethics surrounding language models.These questions accompany the conclusion rather than being resolved by the study.
A RELATED WORK ON AUTOMATED STORY GENERATION AND CONTROLLABLE STORY GENERATION
The related-work section situates the paper within automated plot and story generation and controllable language generation. It identifies these as intersecting fields relevant to the study’s approach.
- Related work: The paper reviews automated plot and story generation as background for its approach.
- Related work: Together, these fields provide the stated background for the paper’s work on hierarchical script generation.
- Related work: It also reviews controllable language generation as an intersecting field.
A.1 Automatic Story Generation
Automatic story generation models narratives as sequences of causally related elements, but generating coherent plots remains difficult, especially for intertwined plot lines and character consistency.
- A.1 Automatic Story Generation: Automatic story generation seeks to generate sequences of story elements that collectively tell a coherent narrative.Narrative elements include actions, beats, scenes, or events, with each plot event affecting the next.
- A.1 Automatic Story Generation: Generative plot systems have supported human authors with creative material and randomness, and have been adapted for web-based and theatrical interaction.Such systems have been developed for nearly a century, with computerized versions existing for decades.
- A.1 Automatic Story Generation: Multiple subplots can combine into single plots, while intertwined plot lines create complex narratives common in contemporary stories.The literature identifies computational generation of multiple interacting plot lines as an underexplored area.
- A.1 Automatic Story Generation: Earlier systems used symbolic planning and hand-engineered heuristics, whereas recent approaches use machine learning, large datasets, deep models, compute, and prompt engineering.These approaches demonstrate both successful and failed generation of unique, coherent stories.
- A.1 Automatic Story Generation: Coherence remains challenging to measure for causal narrative events, common-sense knowledge, and consistency in characters.Dialogue coherence has been studied, but story-level coherence remains difficult to evaluate across these dimensions.
A.2 Symbolic and Hierarchical Story Generation
Prior story-generation systems have decomposed narratives into events, cards, storylines, or coarse-to-fine stages, but have rarely synthesized coherent scenes and dialogue for longform scripts.
- A.2 Symbolic and Hierarchical Story Generation: Prior methods modeled events from text, expanded plot events into sentences, or simulated causal events before transforming them into prose.Other systems represented stories with character and challenge cards or simulated social practices between autonomous agents.
- A.2 Symbolic and Hierarchical Story Generation: Other approaches separated storyline planning from story writing or decomposed generation into coarse-to-fine processes.Rescoring methods were also introduced for character and plot events.
- A.2 Symbolic and Hierarchical Story Generation: These methods did not focus on synthesizing coherent stories by generating scenes and dialogue.The distinction concerns producing a final script rather than relying primarily on reading-comprehension or preference-based evaluation.
- A.2 Symbolic and Hierarchical Story Generation: THEaiTRE’s hierarchical play-generation process began with a title or story prompt, generated a synopsis, and then generated dialogue.Unlike the present approach, it did not generate characters alongside the synopsis and began dialogue from manually entered two-line exchanges.
- A.2 Symbolic and Hierarchical Story Generation: THEaiTRE’s system served production of a specific theatrical play rather than evaluation within a diverse community of writers.Its performance documentation was associated with the Prague-based company THEaiTRE.
A.3 Controllable Story Generation
Controllable story-generation research has used cues, keywords, genres, themes, narrative arcs, and staged prompting to guide generated stories and scripts.
- A.3 Controllable Story Generation: Prior systems controlled story generation with user cues, keywords, topics, genres, themes, or character-specific language models.These approaches included genre-controlled short stories, multi-user dialogue, and chatbot-assisted character creation.
- A.3 Controllable Story Generation: Tale Brush generated story arcs using the protagonist’s changing fortune, while related work used the narrative arc to produce creative dialogue.Dramatron also uses a narrative arc textually in its prompts.
- A.3 Controllable Story Generation: Controlled Cue Generation for Play Scripts used language models to generate the next line and a stage direction.This provides control at the level of individual play-script continuations.
- A.3 Controllable Story Generation: Dramatron decomposes a log line into a synopsis, using language-model output as the prompt for the next generation stage.The authors describe this as analogous to Chain of Thought prompting and refer to the approach as prompt chaining or language engineering.
A.5 Interactive Authorship
Interactive authorship research has explored tools for generating, revising, and extending text, while Dramatron places writer interventions within a hierarchical generation structure for longform scripts.
- A.5 Interactive Authorship: Research on interactive LLM authorship has examined multiple interaction modalities, including idea generation, copy editing, and scene interpolation.These modalities have been studied alongside user evaluations.
- A.5 Interactive Authorship: Creative-writing systems have supported revision, summarization, descriptive rewriting, writer’s-block assistance, and suggestions for science writers.These systems operated on drafts, short interactions, or suggested continuations and details.
- A.5 Interactive Authorship: Dramatron allows writer interventions within a hierarchical generation structure.This distinguishes its interaction model from systems centered on turn-by-turn continuation or revision.
- A.5 Interactive Authorship: Automatic story generation has been used in games, interactive narratives, virtual worlds, artistic performances, improvised theatre, short films, lyrics, and playwriting.The cited examples include Sunspring, Beyond the Fence, and work at the Young Vic.
- A.5 Interactive Authorship: Before this work, few techniques had generated long-range coherent theatre scripts or screenplays, and none had used few-shot learning and prompt engineering to prime LLM generation.The authors identify this gap while noting related work by THEaiTRE and DialogueScript.
A.6 Review of Automated and Machine-Learned Metrics for the Evaluation of Story Generation
Automated and machine-learned metrics for story generation compare generated text with prompts or reference stories, while other measures assess diversity, repetition, plausibility, and linguistic properties.
- Ground-truth and prompt-based evaluation: Story-generation evaluations may compare generated stories with ground truth using metrics such as perplexity, prompt-ranking accuracy, ROUGE, and coherence judgments.These approaches evaluate similarity, prompt alignment, continuation quality, or coherence against reference material.
- Ground-truth and prompt-based evaluation: Prompt-based evaluations also measure consistency, style matching, lexical cohesion, entity coreference, grammaticality, diversity, frequency, and syntactic complexity.Some metrics are story-dependent, while others assess general linguistic properties.
- Prompt and text-property metrics: Generated text can be compared with prompts using N-gram similarity, sentence-embedding similarity, unique-word counts, verb diversity, entity-name diversity, rare-word usage, and sentence length.These measures target lexical overlap, semantic proximity, vocabulary, and structural characteristics.
- Prompt-independent evaluation: Without reference comparisons, evaluations can measure vocabulary-token ratios, entities per plot, unique verbs, verb diversity, and inter- or intra-story n-gram repetition.These metrics describe originality, diversity, and repetition within generated content.
- Prompt-independent evaluation: Other approaches assess sentence diversity with self-BLEU or train classifiers to judge short-story plausibility.These methods provide alternative evaluations of diversity and perceived plausibility.
B ADDITIONAL DISCUSSION FROM PLAYS BY BOTS CREATIVE TEAM
The Plays by Bots creative team described Dramatron’s outputs as glitchy, repetitive, and sometimes nonsensical, but often enjoyable and productive for performers. Their reflections also emphasized human agency, collaborative interpretation, and the usefulness of co-written scripts.
- Discussion themes: Four themes structured the creative-team discussions: glitch style, repetition, agency and expectations, and the fun of participating.These themes were discussed alongside supporting quotations from the production team.
- Glitch style: Dramatron’s glitch style could be nonsensical, vague, passive, or internally conflicted, prompting performers to make concise and specific interpretive choices.Performers described the system as sometimes working against itself and said flawless text might be less useful.
- Repetition: Repetition could function either as a mistake or as a playful authorial choice, with performers using delivery, physicality, proximity, and acting choices to add meaning.The cast could treat repeated lines as choices to honor rather than simply errors to correct.
- Agency and expectations: Team members discussed Dramatron’s apparent choices and the limits of its understanding, treating it as a system that was trying without fully understanding the world.These exchanges shaped expectations about what the system could and could not do.
- Participation and co-creativity: Most performers found participation fun, and one actor described it as liberating because the platform and world creation were already provided.The reflections connect this experience with the usefulness of co-written scripts and with professional improvisers interpreting generated and edited text.
- Interaction metrics: Relative Levenshtein distance was used as a proportional measure of interaction, distinguishing active editing from passive acceptance of generated text.Positive scores indicate one aspect of active writing, while negative scores indicate one aspect of passive acceptance; choosing among seeds is not captured.
- Interaction metrics: The interaction metric does not account for choosing among generated text seeds, and distance comparisons become difficult to interpret when string sets differ in length.The authors normalize distances by length to compare participant editing across structural levels.
- Jaccard similarity: Jaccard similarity measures overlap between sets of lemmatised vocabularies and was used descriptively to compare word choice in original and edited outputs.The preprocessing tokenizes text and generates lemma sets by probable part of speech.
D SUPPLEMENTARY FIGURES
The supplementary material documents the participant grouping and Likert-scale evaluation, and provides prompt-prefix examples for generating titles and character descriptions from writer-supplied log lines.
- Quantitative evaluation: Participants were grouped by experience with AI writing tools and by primary expertise in improvisation, scripted theatre, or film and television.The first grouping is binary, while the second has three classes.
- Prompt sources: The Medea prompts draw on the tragedy by Euripides, public-domain dialogue, summaries, and adapted locations, while other prompts use material from Antigone, The Bacchae, The Frogs, Star Wars, and Plan 9 from Outer Space.The source materials include public-domain works and adaptations from reference websites and summaries.
- Title generation: The prompt examples use a writer-provided <LOG_LINE> to generate titles for Medea and science-fiction scenarios.The examples include alternative, original, and descriptive titles for known plays and films.
- Character generation: Character-generation prompts use the log line and previously generated character descriptions to produce roles such as Medea, Jason, Luke Skywalker, Ben Kenobi, and Darth Vader.The examples specify character roles, traits, relationships, and narrative functions.
G CO-WRITTEN SCRIPTS
The supplementary material includes four scripts co-written by a human playwright and Dramatron, including works presented at the 2022 Edmonton International Fringe Theatre Festival. The excerpts illustrate varied premises, characters, settings, and dialogue from these scripts.
- Supplementary scripts: Four scripts co-written by a human playwright and Dramatron are included as supplementary material.The scripts were produced and presented at The 2022 Edmonton International Fringe Theatre Festival as described in Section 5.9.
- The Day The Earth Stood Still: The Day The Earth Stood Still follows mechanic Miranda and her sister Beth in a world dominated by machines, culminating in Miranda’s sacrifice to save the world.The script includes government agent Ford, sentient military machine Mech, and a setting of devastated San Diego and Miranda’s garage.
- Cheers: Cheers centers on waitress Ella and teacher Allen, whose relationship changes when Allen forms friendships across a different social class and Ella turns to food.The supplementary description identifies the central relationship and its subsequent divergence.
- The Black Unicorn: The Black Unicorn follows Gretta, a peasant from Bridge-End, and her pet dragon Nugget as they confront a foul-smelling wizard.The description says Gretta gains the upper hand through brains and courage.
- The Man at the Bar: The Man at the Bar features lounge singer Teddy, who is in love with patron Rosie and later puts out a fire to save the day.Rosie regularly attends the club with her husband Gerald.
- The Day The Earth Stood Still: The script’s dialogue stages conflict over whether machines should be opposed, with Miranda defending machines and Beth arguing that humans must choose a side.The exchanges also establish the government’s dependence on Miranda and the machines’ attacks on humans.