Source-linked AI summary
SAGE: A Hierarchical Framework for Evaluating Interpretive Literary Quality in Narratives
Tianyu Wang, Nianjun Zhou
TL;DR
Existing NLG metrics do not measure key interpretive dimensions of literary quality, motivating SAGE’s six-layer, theory-grounded framework with iterative LLM evaluation and independent cross-validation. Across 600 evaluations of 100 short stories, SAGE identifies a capability boundary: emotional-psychological representation is closer to human levels, while cultural critique and philosophical depth show approximately double the gap.
Problem
Existing NLG metrics cannot measure cultural representation, emotional depth, and philosophical engagement, leaving these interpretive dimensions without reliable evaluation.
Method
SAGE separates rule-based assessment of observable textual properties from theory-grounded LLM evaluation of interpretive dimensions using multi-round iteration and independent cross-validation.
Results
Emotional-psychological representation approaches human competence, whereas cultural critique and philosophical depth exhibit approximately double the gap; LLM-generated narratives score below commercial genre fiction on all three layers.
Takeaways & Limitations
The findings distinguish pattern-reproducible literary capacities from stance-requiring capacities demanding cultural positioning and philosophical engagement.
Takeaways & Limitations
Generalizability is limited by the corpus of 100 English-language short stories, and inter-rater agreement reflects LLM-to-LLM consistency rather than agreement with professional literary critics.
Abstract
from arXiv · showhide
Assessing the literary quality of narratives requires evaluating interpretive dimensions (cultural representation, emotional depth, and philosophical engagement) that existing NLG metrics cannot measure. We introduce SAGE, a six-layer evaluation framework that separates rule-based assessment of observable textual properties from LLM-based evaluation of interpretive qualities drawn from cultural theory, affect theory, and existentialist philosophy. Each interpretive layer is assessed through multi-round iterative LLM evaluation with independent cross-validation, achieving measurement-grade reliability (98.8% convergence, >94% inter-rater agreement) stable across evaluator models. Validated on 600 evaluations across 100 short stories, our central finding is a systematic capability boundary: emotional-psychological representation approaches human levels, while cultural critique and philosophical depth exhibit approximately double the gap. LLM-generated narratives score below even commercial genre fiction on all three layers. We interpret this as a boundary between pattern-reproducible literary capacities learnable from training corpora and stance-requiring ones demanding cultural positioning and philosophical engagement that pattern matching alone cannot provide.
1 Introduction
SAGE addresses the lack of reliable evaluation for cultural, emotional-psychological, and philosophical dimensions of narrative quality. It separates observable textual assessment from theory-grounded LLM evaluation and reveals a capability boundary across these interpretive dimensions.
- Existing NLG metrics measure surface similarity poorly correlated with human narrative-quality judgments, leaving interpretive dimensions unaddressed.These dimensions include cultural representation, emotional-psychological depth, and philosophical engagement.
- SAGE decomposes literary quality into six analytically distinct layers, separating rule-based textual assessment from LLM-based interpretive evaluation.The interpretive layers are grounded in literary and cultural theory and assessed through iterative self-reflection with independent cross-validation.
- Emotional-psychological representation is substantially closer to human levels than cultural representation and philosophical depth, whose gaps are approximately double.The paper characterizes this difference as a boundary between pattern-reproducible and stance-requiring literary capacities.
- LLM-generated narratives score below even commercial genre fiction on all three interpretive layers.The authors interpret this deficit as more fundamental than insufficient literary training data.
- SAGE achieves >94% inter-rater agreement and 98.8% convergence on interpretive literary dimensions.These results support the paper’s claim that theory-driven multi-round LLM evaluation can provide measurement-grade reliability.
2 Related Work
Prior narrative-evaluation work has expanded beyond surface metrics toward structural and computational-literary analysis, but interpretive literary qualities remain largely unmeasured. Existing NLP and LLM-as-judge approaches address related problems without providing a unified framework for cultural, emotional, and philosophical depth.
- BLEU, ROUGE, and BERTScore assess surface similarity and correlate poorly with human judgments of narrative quality.Narrative-specific frameworks improve coverage of structural properties such as coherence and character consistency.
- Existing narrative frameworks remain focused on observable structural properties rather than cultural positioning, inner-life representation, or philosophical inquiry.These interpretive dimensions are described as distinguishing canonical literature from competent prose.
- Bias detection, sentiment analysis, and topic modeling address harmful representation, broad affect, and recurring themes without measuring literary sophistication in those domains.The cited approaches do not capture cultural-engagement sophistication, emotional granularity and interiority, or philosophical depth.
- LLM-as-judge methods offer scalable evaluation but require mitigation for systematic biases such as position, verbosity, and self-preference.Fine-grained decomposition is presented as improving evaluation precision.
3 Methodology
SAGE organizes narrative-quality assessment into independent rule-based and LLM-based layers, with L4–L6 operationalizing cultural, emotional-psychological, and philosophical interpretation. Its dual-track procedure combines iterative evaluation, independent validation, bias mitigation, and contrasting information modes.
- 3.1 Framework Overview: SAGE’s six layers progress from observable textual properties assessed by rules to increasingly interpretive qualities assessed by LLMs.The layers are independent rather than dependent, so no layer’s scores are computed from another layer’s scores.
- Interpretive Layers: L4 measures cultural engagement and power structures, L5 measures emotional-psychological representation, and L6 measures philosophical depth and existential engagement.Each interpretive layer is grounded in a distinct set of cultural, affective, psychological, existential, moral, or hermeneutic theories.
- Construct Validity: Dimension non-redundancy is tested through theoretically motivated dissociations among canonical literary exemplars.Examples include high emotional complexity with near-zero psychological interiority and opposing cultural-voice positions despite similarly high power-dynamics scores.
- Evaluation Modes: Content-limit mode supplies only story text, whereas title-limit mode supplies only title and author information.Comparing the modes probes reliance on textual analysis, memorized critical opinions, or both.
- Evaluator Architecture: The dual-track pipeline pairs a five-round iterative evaluator with an independent validator that assesses the same texts without access to the first evaluator’s outputs.Agreement supports inter-rater reliability, while systematic disagreements identify less reliable evaluations; Figure 1 summarizes the architecture.
4 Experiments
The experiments evaluate 100 short stories across three quality categories, three interpretive layers, two information modes, and two evaluator roles. The design produces 600 complete evaluations and uses statistical tests tailored to reliability, discrimination, robustness, and dimensional independence.
- Corpus: 100 short stories form the evaluation corpus, with all stories constrained to 2,000–8,000 words.The corpus contains canonical literature, pulp fiction, and LLM-generated stories.
- Corpus Composition: The corpus includes 50 canonical, 30 pulp-fiction, and 20 LLM-generated stories.The LLM-generated subset contains high-, average-, and low-quality examples for testing deliberate quality variation.
- Evaluation Matrix: 600 complete evaluations result from assessing each story across L4–L6 in both modes with both a five-round iterative evaluator and a one-round independent validator.This matrix varies interpretive layer, information condition, and evaluator role.
- Implementation: GPT-5-mini performs all LLM-based evaluations using theory-grounded prompts, evidence citation requirements, strict layer boundaries, and layer-specific bias mitigation.The prompts draw on frameworks including Bourdieu, Said, Geertz, Sedgwick, Barrett, Wood, Heidegger, Levinas, and Ricoeur.
- Analysis: Reliability, discriminative validity, robustness, and dimensional independence are assessed with convergence and inter-rater MAD, t-tests, ANOVA with Bonferroni correction, effect sizes, and correlations.The statistical procedures correspond to RQ1–RQ4.
5 Results
Across 600 evaluations, SAGE produced highly stable assessments and a statistically significant genre hierarchy. The results show stronger emotional-psychological performance than cultural or philosophical performance, with distinct cross-layer patterns and consistent rankings across evaluator models.
- Reliability and robustness: 98.8% convergence and >94% inter-rater agreement demonstrate measurement-grade reliability across 600 evaluations.Scores stabilized by Rounds 3–4, with mean MAD <0.3.
- Genre discrimination: Canonical narratives scored 4.02, pulp fiction 3.85, and LLM-generated narratives 2.83 across the three layers.All canonical-vs.-LLM comparisons were significant at p < 0.001, while canonical-vs.-pulp comparisons were not significant.
- Genre discrimination: L4 Cultural (d=2.68) and L6 Existential (d=2.40) showed roughly 60% larger effects than L5 Emotional (d=1.68).This pattern indicates qualitatively different capability profiles rather than uniform underperformance.
- Genre discrimination: Commercial pulp fiction approached canonical literature emotionally, while LLM-generated stories scored lower than pulp fiction on all three layers.Emotional scores were 4.04 for pulp fiction versus 4.15 for canonical literature, while LLM-generated stories scored 3.36 versus pulp fiction’s 4.04.
- Cross-model validation: GPT-5-mini and Claude rankings correlated from r=0.636–0.812 despite Claude scores being approximately 0.5 points lower.The replication included 59 of 60 story-mode pairs; the absolute offset reflected calibration differences rather than ranking disagreement.
- Dimensional independence: Cross-layer correlations ranged from r=0.649–0.683, remaining below the r > 0.7 redundancy threshold and supporting distinct dimensions.Emotional scores were highest overall at 3.96, followed by cultural at 3.64 and existential at 3.59.
6 Discussion
The discussion identifies a capability boundary: LLMs more closely reproduce emotional-psychological structures than cultural critique or philosophical depth, while reliable, theory-grounded evaluation makes this asymmetry measurable.
- The Capability Boundary: LLM-generated narratives reproduce emotional-psychological structures more closely than cultural critique or philosophical depth, whose gaps are approximately twice as large.The proposed explanation distinguishes pattern-reproducible capacities from stance-requiring capacities demanding cultural positioning and philosophical engagement.
- The Capability Boundary: 3.36 versus 4.15, the LLM mean remains below the canonical mean on emotional-psychological representation.Emotional complexity, psychological interiority, and emotional granularity are described as densely represented and learnable from narrative corpora.
- The Capability Boundary: LLM-generated stories score below pulp fiction on all three layers, including emotional representation at 3.36 versus 4.04.The cultural and philosophical gaps against pulp are +1.28 and +1.09 points, respectively.
- The Capability Boundary: Cultural and existential dimensions correlate most strongly at r=0.683, whereas cultural and emotional dimensions show the weakest reported link at r=0.649.The authors argue that aggregating the dimensions would obscure the differentiation revealing the capability boundary.
- Why Reliable Evaluation Matters: Existing NLG metrics capture fluency, surface coherence, and emotional mimicry more readily than cultural critique or philosophical depth.The resulting feedback is rich where improvement is least needed and absent where the gaps are largest.
- Why Reliable Evaluation Matters: 98.8% convergence and >94% inter-rater agreement support SAGE as a potentially usable evaluation signal for interpretive dimensions.Cross-model validation reports overall r=0.775 with Claude Sonnet, supporting the architecture rather than a single evaluator model.
- Scope and Limitations: The evaluation corpus contains 100 English-language short stories, so generalizability to other languages, narrative lengths, and cultural traditions requires separate validation.The framework validates interpretive layers L4–L6, while full integration with rule-based L1–L3 metrics remains future work.
- Scope and Limitations: 13 of 15 story-pair rankings matched scholarly critical consensus, but validation against professional literary critics remains essential.The reported 87% directional consistency is presented as a proxy for external validity rather than direct agreement with literary experts.
7 Conclusion
SAGE evaluates interpretive literary quality through theory-grounded decomposition and multi-round LLM assessment with independent cross-validation. Its central finding is a capability boundary: emotional-psychological representation approaches human competence, while cultural critique and philosophical depth show larger gaps.
- Conclusion: SAGE evaluates interpretive literary quality by decomposing narratives into theory-grounded dimensions assessed through multi-round LLM evaluation and independent cross-validation.The framework is hierarchical and separates interpretive assessment from observable textual-property assessment.
- Conclusion: 600 evaluations support measurement-grade reliability across convergence, inter-rater agreement, and mode invariance.The reliability is reported as a property of the multi-round architecture rather than any single model.
- Conclusion: LLM-generated narratives approach human competence on emotional-psychological representation but exhibit approximately double the gap on cultural critique and philosophical depth.The authors characterize this as a boundary between pattern-reproducible capacities and stance-requiring capacities.
- Limitations: The corpus is restricted to 100 English-language short stories, and cross-lingual and cross-cultural validation remains necessary.External validation against professional literary critics is also identified as essential.
Ethical Considerations
The study uses publicly sourced literary materials and acknowledges ethical and methodological limits in automated LLM evaluation. It provides evaluation data and materials while cautioning that residual bias prevents treating results as ground-truth judgments.
- Data Use: The corpus contains 100 short stories, with public-domain works sourced from Project Gutenberg, Wikisource, and the Internet Archive.Copyrighted works published between 1928 and 1990 are used under Fair Use for noncommercial academic research.
- Data Use: The 20 LLM-generated stories come from the publicly available lars76/story-evaluation-llm dataset under its stated license.The dataset contains no personally identifiable information.
- Evaluation Risks: LLM-based evaluation risks systematic bias, including cultural projection, Western-centric norms, and hallucinated textual evidence.SAGE addresses projection bias and hallucination through explicit checks in its multi-round architecture.
- Evaluation Risks: Residual bias cannot be fully eliminated, so the results should be interpreted as evidence from a structured automated evaluation system rather than ground-truth literary judgments.This qualification applies despite the framework’s structured bias-mitigation checks.
- Transparency: Evaluation scores, prompt templates, analysis scripts, and the story catalog are publicly available, excluding copyrighted text.The materials are released with metadata and evaluation resources.
A Exemplar Matrices
The exemplar matrices demonstrate that SAGE’s theoretically motivated dimensions capture distinct aspects of literary quality. Literary works frequently score highly on one dimension while remaining low or contrasting on another.
- Construct Validity: Tables 7–9 use exemplar matrices to demonstrate that all four dimensions within each layer are non-redundant.Works can score high on one dimension while scoring low on another, supporting construct validity.
- L4 Exemplar Matrix: IPD and CSP dissociate: Animal Farm maximizes power visibility while reducing cultural specificity, whereas The Great Gatsby renders Jazz Age New York precisely but centers aspiration.The examples show that cultural specificity and power visibility need not co-occur.
- L4 Exemplar Matrix: CVP and IPD dissociate: Conrad and Achebe share comparable IPD while occupying opposite CVP positions in their portrayals of colonial encounter.Conrad presents a European observer’s consciousness, while Achebe presents the encounter from within Igbo life.
- L4 Exemplar Matrix: CPC can be high while IPD is low, as Borges’s “Tlön” stages incompatible epistemic systems without organizing them through power hierarchy.The example separates epistemic complexity from ideological positioning.
- L5 Exemplar Matrix: AC can be high while PI is near zero in Hemingway’s “Hills Like White Elephants,” whereas Woolf’s Mrs. Dalloway achieves both through stream of consciousness.The contrast separates emotional complexity from access to characters’ inner states.
- L5 Exemplar Matrix: Carver’s “Cathedral” combines flat emotional vocabulary with high ENC because the narrator’s transformation is fully motivated by established relationships and events.Emotional vocabulary and emotional-narrative coherence therefore dissociate.
- L6 Exemplar Matrix: LP and MR dissociate in Camus’s The Stranger and Dostoevsky’s Crime and Punishment, which share high LP but differ in whether moral evaluation is refused or central.The examples distinguish a coherent life-position from moral reflection.
- L6 Exemplar Matrix: HC and ME dissociate: Zola’s deterministic naturalism forecloses meaning, while Hemingway’s old waiter constructs meaning through maintaining order against nada.Both works share maximal HC but diverge in meaning engagement.
B Evaluation Protocol and Prompt Templates
The evaluation protocol standardizes scoring and prompt design across SAGE’s layers and modes. The accompanying templates and figures document the iterative procedure and convergence behavior.
- Five-Round Iterative Protocol: Table 10 presents the five-round iterative evaluation protocol applied across all layers.The protocol is the common procedural structure for the evaluations.
- Scoring Rubric: Table 11 presents the scoring rubric applied uniformly across all layers.Uniform application supports comparability among layer-specific assessments.
- Prompt Design Summary: Table 12 summarizes prompt-design components and their layer-specific instantiations.The prompt architecture adapts shared components to individual layers.
- Prompt Design Summary: Full prompt templates and JSON schemas for all three layers and both modes are publicly available.The release also identifies Table 12 as a summary of the instantiated design decisions.
- Convergence Trajectories: Figure 3 shows score trajectories across five iterative rounds for each layer and genre category.The figure is the protocol’s convergence visualization.
- Convergence Trajectories: Figure 4 shows the corresponding effect sizes for the canonical-versus-LLM-generated comparison.The figure complements the convergence trajectories with between-category effect sizes.
D Cross-Model Validation Results
The cross-model validation materials compare evaluator consistency and visualize convergence and group differences. They include prompt-design documentation, effect-size reporting, and five-round score trajectories.
- Prompt Design Summary: Prompt-design documentation summarizes the components and layer-specific instantiations used in the evaluation.The documentation is provided in Table 12 and accompanying prompt templates.
- Effect Sizes: Figure 4 reports Cohen’s d for canonical versus LLM-generated stories across the three layers.All values exceed the large-effect threshold of d = 0.8.
- Cross-Model Consistency: Table 13 reports cross-model consistency between GPT-5-mini and Claude Sonnet on a 30-story subset.The comparison includes 59 pairs per layer.
- Convergence Trajectories: Figure 3 displays score-convergence trajectories across five rounds for each layer and genre category.The visualization compares canonical, pulp-fiction, and LLM-generated categories.
- Convergence Trajectories: Scores stabilize by Rounds 3–4, with canonical and pulp fiction consistently scoring higher than LLM-generated stories.This pattern is shown across the displayed layers and genre categories.