Source-linked AI summary

Operationalizing Narrative Entropy (Sn): A Two-Scene Registered Pilot Report and Pre-Validation Protocol

Levent Bulut

arXiv:2608.18109v1cs.CL

TL;DR

Narrative Entropy has been theoretically defined but not yet computed from real texts. This report operationalizes the candidate measure in a two-scene pilot and finds that the monologue scored higher than the dialogue scene.

  • Problem

    Narrative Entropy has been stated theoretically but never applied to actual text, leaving its status as a measure untested.

  • Method

    The report manually codes two narrative scenes, applies Sn = If × Cb × t, and preregisters tests and decision rules for the next stage.

  • Results

    Sn was 30.0 for the single-voice monologue and 18.8 for the nine-character dialogue scene, producing the report’s central divergence.

  • Takeaways & Limitations

    The pilot computes Sn from real texts and makes the competing explanations and construct-validity questions answerable without post-hoc formula fitting.

  • Takeaways & Limitations

    With n = 2, a single rater, judgment-dependent coding, and unresolved construct validity, the pilot does not validate Sn.

Abstract

from arXiv · show

Narrative Entropy ($S_n$) is a proposed quantitative descriptor within the Bulut Doctrine, intended to capture the rate at which a narrative text imposes processing load on a reader. To date the construct has been defined theoretically but not operationalized against real texts. This report documents the first such operationalization (the v2.0 pilot): two narrative scenes -- the opening restaurant scene of Tarantino's Reservoir Dogs and the opening interior-monologue block of Carver's Cathedral -- were coded manually by a single rater and scored with the candidate formula $S_n = I_f \times C_b \times t$. The result was a divergence from the author's naive intuition: the single-voice monologue ($S_n = 30.0$) scored higher than the nine-character dialogue scene ($S_n = 18.8$). We treat this not as a result to be explained away but as the central finding, and we refuse post-hoc adjustment of the formula. Three competing interpretations are presented -- formula incompleteness, genuine high-load prose, and measurement error -- and the design that would discriminate among them is pre-registered. This v2.1 revision adds: (i) explicit acknowledgement that the divergence is consistent with the pre-existing architectural framework which privileges inferential reconstruction over surface declaration, and that what was called "contrary to expectation" in v2.0 reflected the author's anticipatory intuition rather than the methodology's own predictions; (ii) a pre-registered construct validity test for $I_f$, motivated by the observation that $I_f$ values were nearly equal across the two scenes (1.71 vs 1.58) despite the headline $S_n$ divergence. The document functions simultaneously as a pilot report ($n=2$) and as a pre-registration of the next-stage protocol. It does not claim that $S_n$ has been validated.

1 Introduction

Narrative Entropy (S_n) is proposed as a measurable account of narrative interpretive and informational load, but its candidate formula had previously remained theoretical. This registered pilot operationalizes S_n on two real-text scenes while pre-registering the next validation stage, without claiming validation from n = 2 and a single rater.

  • 1 Introduction: Narrative Entropy (S_n) is proposed to measure the rate at which narrative imposes interpretive and informational load on readers.It supports a broader six-layer architecture whose higher-order phenomena assume narrative load can be measured.
  • 1 Introduction: The candidate formula S_n = I_f × C_b × t had been stated but never applied to an actual text.I_f denotes Information Friction, C_b Causal Branching, and t elapsed time.
  • 1 Introduction: The pilot responds to an objection that narrative “load” remained an intuitive label rather than a measurement procedure by attempting to count something.The stated methodological response is empirical operationalization rather than rhetorical defense.
  • 1 Introduction: The registered pilot documents the first operationalization of S_n on real texts (n = 2 scenes) and pre-registers the next validation stage’s design, hypotheses, and decision rules.Its two functions are transparent reporting of the pilot and advance specification of the subsequent data-collection protocol.
  • 1 Introduction: n = 2 and a single rater make validation impossible, so the report makes no validation claim and instead treats the pilot’s surfaced problem as its value.The underlying raw data is published as an open laboratory notebook.

2 Operational Definitions

The pilot used provisional, judgment-dependent coding definitions for information units, uncertainty, topic shifts, elapsed time, and spatial compactness. Timing differed by medium, with prose duration estimated from word count under an explicitly consequential reading-rate assumption.

  • Pilot-stage definitions: Definitions were explicitly provisional and judgment-dependent, intended to invite criticism and revision.The paper presents these as pilot-stage operational choices rather than settled constructs.
  • Core coding variables: New information units are first appearances of characters, locations, events, or concepts; the uncertainty ratio measures introduced units not fully explained at introduction.t denotes elapsed time in minutes.
  • Spatial compactness: Spatial Matrix was split into Mp for physical compactness and Mn for narrative compactness because the Cathedral narrator remains stationary while memory traverses multiple locations.The split prevents one spatial measure from conflating bodily constraint with narrated movement.
  • Timing assumptions: 200 WPM was assumed for the Cathedral excerpt, although 160–180 WPM may better fit dense literary prose, and the choice directly scales Sn.Screen-material duration comes from scene duration, whereas prose duration is estimated from word count; the reading rate remains an open parameter.

3 Method and Results

A single-rater pilot coded two contrasting opening scenes using direct enumeration and the candidate formula S_n = I_f × C_b × t. The single-voice Cathedral monologue scored higher than Reservoir Dogs’ nine-character dialogue scene, contrary to intuitive prediction.

  • Scene selection: Two opening scenes contrasted rapid dialogue among many characters with a single-voice interior monologue.The scenes were Reservoir Dogs’ opening restaurant scene and Cathedral’s opening narrative block.
  • Coding procedure: Each scene was read and coded once by a single rater using direct enumeration, with only the reading-rate parameter not directly observed.An intuition-based estimate of 3.0 for Scene A was replaced by direct counting, which yielded 18.8.
  • Interpretation: 30.0 Narrative Entropy (S_n) for the single-voice monologue exceeded 18.8 for the nine-character dialogue scene, contrary to intuitive prediction.The formula discriminated between the scenes, but in the unexpected direction.

4 Discussion: Five Open Problems

The pilot did not validate S_n; it exposed five open problems, including a dimensionally unclear formula, single-rater measurement uncertainty, and an unresolved interpretation of the scene divergence. The v2.1 revision situates the divergence within the Bulut Doctrine while pre-registering tests of competing explanations and I_f construct validity.

  • Scope: Five open issues, including four documented in v2.0 and one added in v2.1, remain unresolved because the pilot did not produce clean validation.The report treats these issues as a problem definition for the next-stage protocol rather than as grounds for post-hoc revision.
  • Formula structure: S_n = I_f × C_b × t is dimensionally unclear because I_f and C_b are per-minute rates, so elapsed time is divided out twice and multiplied back once.The report notes that using underlying totals would be defensible but deliberately does not adopt that alternative after a single pilot.
  • Measurement: A single rater judged categories such as “new information unit,” “uncertainty ratio,” and “topic shift,” leaving inter-rater agreement unestablished.Until agreement is tested, Section 3 values are one rater’s careful counts rather than objective measurements.
  • Interpretation: 18.8 vs 30.0 separated the scenes, but the pilot cannot distinguish formula incompleteness, genuinely higher-load Carver prose, or measurement error.Discriminating among the three interpretations requires more data, including attention to coding error and the reading-rate assumption.
  • Architectural interpretation and validity test: 1.71 vs 1.58 was the near-equal I_f difference, while S_n = 30.0 > 18.8 was driven primarily by C_b = 2.53 vs 1.57 and t = 7.5 vs 7.0.Because the framework privileges inferential reconstruction over surface declaration, the revision pre-registers whether that rationale appears primarily in I_f rather than relying on post-hoc rationalization.

5 Pre-Registered Protocol for the Validation Stage

The validation stage preregisters a four-scene design and tests whether I_f measures inferential load through a parallel Suppressed Information Index. It also fixes reliability, physiological, and anti-post-hoc decision rules before data collection.

  • Scene design: Four scenes will span a 2 × 2 grid crossing physical energy and branching, adding high-energy/low-branching and low-energy/low-branching cases.The expanded grid tests the hypothesized space beyond Scenes A and B.
  • Construct validity: The construct-validity test compares I_f with SI, defined as information units reconstructed from indirection per minute of elapsed time.SI units must be implied but unstated, locally required for coherence, and explicitly paraphrasable by a second reader.
  • Construct validity: Spearman ρ > 0.7 is the preregistered criterion for treating I_f and SI as strongly correlated measures of inferential load.If ρ ≤0.7, I_f will be judged invalid for inferential load regardless of whether the pilot’s S_n ordering is reproduced.
  • Reliability: Cohen’s κ or Fleiss’ κ above 0.60 is targeted for “new information unit,” “uncertainty ratio,” and “topic shift” before any S_n value is treated as established.Two to three independent raters will code the full scene set.
  • External validation: HRV and EDA will test whether the Cathedral condition produces higher physiological load than the Reservoir Dogs condition under the S_n = 30.0 > 18.8 ordering.The protocol prohibits revising the formula merely because it contradicts intuition; changes must follow data and independent theoretical motivation.

6 Limitations

The report’s limitations determine its status as a pilot rather than a validated measure: n = 2, one rater, judgment-dependent coding, an unresolved dimensional problem, and unestablished construct validity remain.

  • Limitations: n = 2 means two scenes cannot validate the measure.The report explicitly treats the small sample as defining its pilot status.
  • Limitations: A single rater provides no inter-rater agreement, while judgment-dependent coding remains unshown to be reproducible.Several coding categories depend on rater judgment.
  • Limitations: The formula has an unresolved dimensional problem, and construct validity for I_f is not yet established.The pre-registered test addresses whether I_f may produce correct orderings through a different mechanism than purported.
  • Limitations: The report states these limitations plainly because concealing them would make the pilot less useful, not more.The limitations define the report’s status rather than being incidental shortcomings.

7 Conclusion

The pilot operationalized S_n on real texts, producing an informative divergence that v2.1 reframes within the methodology’s inferential-load framework rather than treating it as a failed prediction. The report preserves the result without post-hoc formula changes and pre-registers tests to determine what it means and whether I_f measures the claimed construct.

  • Conclusion: The pilot computed S_n from real narrative texts, meeting its narrow aim of turning the doctrine’s intuitive constructs into measures.The authors characterize this operationalization as an informative surprise rather than a validation claim.
  • Conclusion: The result was contrary to the author’s naive expectation, but consistent with the methodology’s framework because inferential friction was privileged over sensory density.The v2.1 revision separates the author’s expectation from the methodology’s own prediction.
  • Conclusion: The report rejects post-hoc formula changes and instead pre-registers a design to distinguish formula incompleteness, genuine high-load prose, and measurement error.v2.1 adds a construct-validity test for If because its values were nearly equal across the two scenes despite the headline divergence.
  • Conclusion: v2.1 adds explicit architectural-framework acknowledgment and a pre-registered construct-validity test for If, while retaining v2.0 as the public pilot record.The revision states that no v2.0 claims are retracted.

How to Cite

The paper is cited as Bulut (2026), with its full title, v2.1 designation, Zenodo publication venue, and DOI provided.

  • The citation identifies the author as Bulut, L., and the publication year as 2026.
  • The cited work is titled “Operationalizing Narrative Entropy (Sn): A Two-Scene Registered Pilot Report and Pre-Validation Protocol (v2.1).”
  • The paper is published through Zenodo and is identified by DOI 10.5281/zenodo.20362901.
Loading 2608.18109v1…