Source-linked AI summary
How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling
Katrin Rohrbacher, Björn Nieth, Emmanuelle Salin, Bjoern Eskofier, Michaela Mahlberg
TL;DR
The paper asks how LLMs construct fictional worlds through narrative setting, an underexplored spatial dimension of AI-generated fiction. It analyzes 8,000 model-generated stories and matched human fiction using five setting categories and classifier-based corpus comparisons. LLMs consistently overrepresent perceived space and underrepresent action space relative to human fiction, with model-specific and language-sensitive variation; interpretation is bounded by chapter-based prompting, classifier performance, and the historical English- and German-language corpora.
Problem
The spatial dimension of worldbuilding in AI-generated fiction remains largely unexplored despite studies of creativity, style, embodiment, emotion, temporal structure, and plot.
Method
The study compares 1,000 stories per model and language from four LLMs with human-authored fiction, classifying sentences into five spatial categories and modeling distributions across narrative sections.
Results
LLMs systematically overrepresent perceived space and produce less action space than human authors across models and narrative sections, with model-specific and cross-linguistic variation.
Takeaways & Limitations
LLM-generated fiction differs consistently from human-authored fiction in narrative-space construction, revealing model- and language-sensitive worldbuilding patterns.
Takeaways & Limitations
The findings are bounded by chapter-boundary effects, lower classifier performance on Mistral 3.2 outputs in German, and English and German public-domain corpora.
Abstract
from arXiv · showhide
In this paper, we analyze how Large Language Models (LLMs) employ worldbuilding strategies, focusing on setting as one measurable dimension of storyworld construction. We compare 1,000 AI-generated stories per model in English and German with human-authored fiction from Project Gutenberg. Building on prior work, we operationalize setting through five types of narrative space: "action", "perceived," "visual," "descriptive" and "no space", identified using fine-tuned BERT classifiers for German and English. We generate narratives using GPT 4.1, LlaMA 3.3, Mistral 3.2, and Gemma 3 and compare their spatial distributions to a human-authored baseline. We find that human-authored texts predominantly employ "action space," grounding narratives in embodied character-environment interaction, whereas LLMs systematically overproduce "perceived space," emphasizing atmosphere and affect. This divergence remains stable across narrative time. Overall, our findings show that LLMs exhibit worldbuilding patterns that differ consistently from human-authored fiction in ways that are both model-specific and language-sensitive.
1 Introduction
This study examines setting as a measurable dimension of worldbuilding in AI-generated fiction, addressing the largely unexplored spatial dimension of LLM narratives. It compares model-generated and human-authored fiction using narratology-informed categories and finds systematic differences in spatial construction.
- Research gap: The spatial dimension of worldbuilding in AI-generated fiction remains largely unexplored despite research on creativity, style, embodiment, emotion, time, and plot.
- Research focus: The study treats setting as a central textual mechanism through which fictional worlds are constructed, alongside time and events.Its framework is grounded in lived space: space experienced and inhabited by a perceiving subject.
- Main finding: LLM storytelling is systematically more atmospheric and less embodied/action-driven than human fiction.
- Contribution: The study compares English and German corpora to show that deviation patterns are jointly shaped by model and language context.
- Resources: The dataset contains 8,000 AI-generated stories, with 1,000 stories per model across four LLMs and two languages, matched with a human-authored corpus.
- Resources: An English setting classifier and annotation resources extend an existing German classifier for corpus-scale, narratology-informed evaluation of AI-generated fiction.
2 Related Work
Prior research has established worldbuilding as a narratological concern, while machine-learning studies have only limitedly incorporated narratological concepts into analyses of LLM-generated narratives. Existing work uses creativity judgments and contextual transformer classifiers, but long-form narrative research remains limited.
- Research gap: Narratological research provides concepts for analyzing how stories construct worlds, but relatively little machine-learning research applies those concepts to LLM narratives.
- Existing approaches: Existing evaluations of LLM-generated narratives commonly use creativity tests adapted from psychological frameworks or human judgments of stylistic quality.
- Analytical methods: Fine-tuned transformer classifiers can capture contextual and long-range semantic dependencies in complex narratological phenomena beyond n-gram overlap.
- Long-form generation: Long-form narrative generation research remains limited because many studies focus on short passages or synopses, although improved context windows and prompting have made extended generation more feasible.
3 Method
The study generates long-form continuations from human fiction with four LLMs in English and German, classifies their narrative-space distributions, and compares them with matched human excerpts across narrative time. Statistical models test author and section effects while accounting for story variation and genre.
- Corpora: The corpus samples 1,000 English and German works from approximately 4,300 public-domain fiction texts, primarily sourced from Project Gutenberg collections.
- Narrative generation: Four models—LlaMA 3.3, Mistral 3.2, Gemma 3, and GPT 4.1—generate four-chapter continuations averaging 6,500–9,400 words.
- Narrative generation: The iterative chapter-based prompting strategy conditions each chapter on prior output to support long-form generation and reduce repetition, meta-commentary, premature endings, and user-directed dialogue.
- Setting classification: A fine-tuned BERT classifier assigns each sentence to action, perceived, visual, descriptive, or no space, with an English model replicated from the German annotation scheme.
- Setting classification: Category proportions are computed for generated and human texts after matching literary excerpts to the lengths of corresponding generated stories.
- Analysis design: Texts are analyzed at the micro level using the first 15 sentences and at the macro level using ten equal-length narrative sections.
- Analysis design: Normalized frequency denotes the proportion of sentences assigned to each spatial category within a narrative section.
- Statistical analysis: A GLMM models each spatial category by author, narrative section, their interaction, and genre, with a by-story random intercept.
4 Results
Human-authored fiction most often uses action space, whereas LLM-generated fiction systematically overuses perceived space. These differences vary by model and language, persist across narrative sections, and are not explained by prompt formulation alone.
- Classifier validation: 78.0% accuracy and mean Cohen’s κ = 0.725 show that the setting classifier closely approached the human inter-annotator ceiling.Human annotation achieved κ = 0.736 and 79.2% raw agreement; performance was consistent across languages and models except Mistral 3.2.
- Overall differences: Action space was the most frequent category in human fiction, while LLMs systematically deviated from the human distribution through perceived-space overuse.Gemma 3 differed most from the human action-space baseline, whereas descriptive space was reproduced most closely by the models.
- Narrative time: All four LLMs produced significantly more perceived space than human authors across every narrative section, with GPT 4.1 showing the largest average deviations.GPT 4.1 exceeded baseline by approximately 0.14–0.25, LlaMA 3.3 by approximately 0.10–0.23, and Gemma 3 remained closest at approximately 0.06–0.10.
- Narrative time: Action-space trajectories differed by model: GPT 4.1 tracked the human baseline, LlaMA 3.3 declined, and Gemma 3 and Mistral 3.2 stabilized below it.For descriptive space, all models broadly followed the human baseline’s declining trend, with model-specific elevations or deficits.
- Language comparison: Cross-linguistic comparisons showed model- and language-shaped deviations, including a reversal in perceived-space differences relative to the English human baseline.Most models fell at or below the German baseline, except GPT 4.1; action-space deviations replicated across languages but were more pronounced in German.
- Robustness: Prompt formulation changed absolute category frequencies somewhat, but trajectories across narrative sections remained consistent across conditions.The authors report that the main findings were not an artifact of the specific prompt used.
5 Discussion
LLMs differ from human fiction in narrative-space composition, systematically favoring perceived space while showing model- and language-specific variation. Chapter prompting explains some section-level fluctuation, but not the elevated overall perceived-space pattern.
- Human–LLM differences: Human fiction foregrounds embodied action, whereas LLM openings favor perceived space and atmospheric mood over concrete character activity.This produces prose rich in affect and ambience but less grounded in embodied action and spatial concreteness.
- Alternative explanations: Human-corpus spatial distributions show no systematic directional trend across the main historical period, weakening temporal drift as an explanation for the contrast.The authors therefore characterize the human–LLM difference as systematic rather than a consequence of historical change in narrative-space composition.
- Alternative explanations: Chapter-level prompting elevates perceived space at chapter boundaries and explains part of the fluctuation across narrative sections, but not the overall skew already present in chapter 1.A different prompting strategy could change the boundary profile without removing the underlying perceived-space tendency.
- Cross-linguistic variation: GPT 4.1 maintains elevated perceived-space distributions in both languages, while open-source models converge toward or below the human baseline in German.Their action-space deficit widens in German, demonstrating language-sensitive variation in the divergence.
- Implications: Genre significantly affects several spatial categories, although matched genre conditions preserve the human–model comparison.Genre-stratified corpora could test spatial composition differences more directly.
- Implications: Spatial category distributions identify the generating model above chance, indicating that setting is a reliable stylistic marker of AI-generated fiction.The paper proposes space-conditioned generation as one way to mitigate the reported differences and reader studies to test links with immersion or quality.
6 Conclusion
The paper introduces a narratology-informed framework for measuring narrative space in AI-generated fiction and applies it across four LLMs, two languages, and human-authored texts. The results show systematic, model-specific, and language-sensitive differences from literary norms.
- Framework and scope: The framework combines five theory-grounded spatial categories, an English BERT classifier extending a German classifier, and corpus-scale analysis across four LLMs.The study examines generated fiction in English and German against human-authored fiction.
- Findings: LLMs systematically overrepresent perceived space and underproduce action space relative to human authors across models and narrative sections.The magnitude varies by model and language, but the distributional contrast is consistent.
- Findings: GPT 4.1 shows the largest and most consistent perceived-space overproduction across both languages while remaining near the human baseline for action and descriptive space.
- Implications: The proposed spatial metrics offer a path toward richer evaluation of AI-generated narratives.
Limitations
The study’s conclusions are bounded by the generation procedure, classifier performance, corpus scope, and operational labeling rules. These constraints affect how stable, generalizable, and interpretable the reported spatial patterns are.
- Scope and stability: Long-form generation remains challenging, so patterns across narrative time may reflect tendencies rather than fully stable narrative strategies.
- Generation procedure: Chapter-based prompting makes every chapter start behave like an opening, elevating perceived space at boundaries and contributing to temporal variance.The authors re-bin texts by position within chapter to examine this effect.
- Classifier performance: The classifier performs comparatively worse on Mistral 3.2 outputs, especially in German, so findings for that model require caution.
- Corpus scope: The corpus contains public-domain English and German fiction from 1800–1920 and 1780–1940, limiting generalizability to contemporary fiction and other languages.
- Operationalization: Each sentence receives one category, with the most prominent space type selected when multiple types co-occur.
- Operationalization: Action space denotes character movement or goal-directed interaction with the environment, whereas perceived space denotes atmospheric, mood-laden environmental experience.
- Operationalization: Visual space captures static observation, descriptive space provides character-independent localization, and no space excludes merely imagined or remembered settings.
- Operationalization: When movement and atmosphere co-occur, the category that predominates determines whether a sentence is labeled action or perceived space.
A.3.1 Inter-annotator agreement
Agreement analyses assess both annotator consistency and classifier performance for the five-category setting scheme. Disagreements concentrate around perceived space, no space, and visual space, while classifier errors are especially associated with no-space cases and Mistral 3.2.
- Agreement measures: Table 9 reports per-category inter-annotator agreement on a double-coded sample of 600 sentences using IoU and chance-corrected κ.The table defines IoU as overlap divided by the union of assignments and κ as one-vs-rest Cohen’s κ.
- Inter-annotator agreement: Perceived space has the lowest raw agreement, with IoU = 0.543, but its κ = 0.646 still indicates substantial chance-corrected agreement.
- Classifier performance: After adjudication, perceived-space classifier F1 is 0.780, comparable to the other categories’ F1 values of 0.720–0.826.
- Confusion structure: The main annotator confusions involving perceived space are with no space and visual space, reflecting threshold judgments about concrete spatiality and affective observation.
- Validation: Only 8 of 600 sentences were disputed between annotators, compared with 4 of 300 English and 2 of 300 German sentences for classifier evaluation.
- Classifier performance: Classifier errors concentrate in the no-space row, whose recall is 0.611 in English and 0.564 in German.Misassigned cases most often go to perceived space, with over half originating from Mistral 3.2 outputs.
- Robustness: Prompt variants change absolute category frequencies but preserve the overall trends across narrative sections.
E Model Identifiability from Spatial Distributions
A random forest could identify the source of texts from per-section spatial proportions in both languages. Human-authored English texts were easiest to recognize, while German LLM distributions converged toward the human baseline.
- A random forest trained on per-section spatial category proportions identified text sources above chance in both English and German.
- English human-authored texts had the highest recall at 0.97, whereas Mistral 3.2 was hardest to identify at 0.56.
- LlaMA 3.3 and Mistral 3.2 were the main confusion pair, consistent with closely aligned setting distributions.
- German human-text recall fell to 0.79 as LLM distributions converged toward the baseline.
- Human-corpus setting distributions were broadly stable across the main publication periods, supporting comparisons without historical-subperiod confounding.
G Full GLMM Results
The paper reports full Type III Wald χ2 tests from generalized linear mixed models fitted separately for each setting category and language. These models provide the test statistics reported in Section 4.3.
- Type III Wald χ2 tests were obtained from GLMMs fitted separately for each setting category and language.
- The GLMM specification includes author × section, genre, and a story-level random intercept.
- Genre was coded as novel versus other fiction and matched across author conditions by construction.
H Chapter-Boundary Position
Perceived-space profiles were examined by quarter within chapter in both languages. The chapter-boundary pattern was specific to perceived space and inversely mirrored by action space.
- Figures 12 and 13 profile perceived-space frequency by chapter position for English and German, with corresponding means reported in Table 12.
- Perceived space changes at chapter boundaries relative to chapter middles, while action space shows the inverse pattern.
- Visual and descriptive space decline slightly and steadily across chapters, with a total spread of at most 0.006 and no boundary peak.
- The boundary effect is therefore specific to perceived space and, inversely, action space.
I Generation Length
Generation used chapter-level token limits and produced four-chapter narratives. Story lengths were summarized in words by model and language for 1,000 stories per condition.
- Chapters were generated at 2,500–3,000 tokens, with a hard maximum of 3,000 tokens.Generation stopped at an end-of-sequence token or the cap.
- Each narrative comprised four chapters totaling 10,000–12,000 tokens per story.
- Prompt-sensitivity ablations used three variants that changed phrasing and verbosity while retaining metadata and each source text’s opening sentence.The procedure was repeated for each subsequent chapter.
- Table 13 reports story length in words separately by model and language, with n = 1,000 each.