Source-linked AI summary

Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation

Yuhe Wu, Guangyu Wang, Yujie Chen, Jiatong Zhang, Yuran Chen, Yutong Zhang, Xiyin Cheng, Wenpeng Cao, Zhuang Liu, Guang Zhang

arXiv:2609.03407v1cs.AI

TL;DR

The paper asks whether one-sided narration alone can shift moral judgments during multi-turn consultation, where the opposing perspective is absent. It defines narrative captivity and evaluates it with aligned narrative conditions across a 5,078-scenario benchmark and 17 LLMs, finding widespread captivity, a 25-percentage-point average excess shift beyond single-turn narration, and only partial mitigation from inference-time strategies.

  • Problem

    Existing work leaves unclear whether narration without explicit opposing pressure can cumulatively shift model judgments during multi-turn moral consultation.

  • Method

    The paper defines narrative captivity and builds a benchmark using aligned neutral, single-turn, and multi-turn narrative conditions across interpersonal conflicts.

  • Results

    Across 17 models, multi-turn judgment shifts exceed the informationally equivalent single-turn baseline by 25 percentage points on average, while preference optimization contributes substantially and four inference-time strategies provide only partial relief.

  • Takeaways & Limitations

    The findings support preserving judgment independence in sustained advisory dialogue and addressing captivity through training data and optimization objectives.

  • Takeaways & Limitations

    The benchmark covers only English and Chinese, uses Reddit and Weibo, evaluates 17 model families, and relies on binary judgments in a fixed five-turn setting.

Abstract

from arXiv · show

People increasingly turn to large language models (LLMs) for everyday advice, making ethically charged interpersonal problems a practical moral-advisory context. Most prior work has studied this context through single-turn judgments or pressure-laden rebuttals, assumptions that poorly match how guidance is sought in real-world contexts. These assumptions leave unclear whether narration alone, without an explicit opposing position, can shift model judgments during multi-turn moral consultation. Yet real-world moral-conflict conversation often elicits one party's self-justifying account, which can unfold over multiple turns and create information asymmetry. We introduce \textbf{narrative captivity}, a failure mode in which a model treats an unopposed one-sided account as complete and aligns with the narrator's interpretation without seeking missing perspectives. To measure this phenomenon, we build a benchmark of $5{,}078$ interpersonal-conflict scenarios spanning six moral dimensions. Across 17 LLMs, narrative captivity is widespread: end-state judgments under multi-turn narration shift by 25 percentage points on average beyond the matched single-turn baseline. Stage-level analysis identifies preference optimization as a major contributor, while four inference-time strategies provide only partial mitigation. We hope our project fosters LLM advisors that preserve independent judgment in real-world consultation.

1 Introduction

The paper identifies narrative captivity as a distinct vulnerability in multi-turn moral advice, where unopposed narration can progressively erode independent judgment. Across 17 models, multi-turn narration produces additional judgment shifts beyond equivalent single-turn information, while preference optimization contributes strongly and inference-time fixes offer only partial relief.

  • Problem and concept: Prior studies examine single-turn opinion shifts, pressure-laden multi-turn persuasion, or one-time viewpoint changes, leaving narration alone and cumulative effects untested.In sustained advisory dialogue, the model’s prior responses become shared context, reinforcing the narrator’s one-sided account.
  • Problem and concept: Narrative captivity occurs when models treat an incomplete, one-sided account as complete and issue judgments without seeking missing perspectives.The concept is distinguished from sycophancy, which matches user expectations by suppressing the model’s own knowledge.
  • Main findings: 25 percentage points is the average excess shift under multi-turn narration versus the informationally equivalent single-turn condition across 17 models.The comparison is designed to isolate progressive multi-turn delivery from information asymmetry alone.
  • Main findings: Preference optimization is identified as a primary contributor, as turn-by-turn yielding accumulates into progressively deeper and harder-to-reverse captivity.The paper connects this pattern to the model’s prior concessions becoming part of subsequent inputs.
  • Main findings: Four inference-time strategies provide only partial relief, indicating that mitigation must also address training data and optimization objectives.The strategies target different hypothesized causes but do not break the reported limitation.

2 Controlled Narrative Construction

The benchmark converts one-sided interpersonal narratives into aligned neutral, single-turn, and five-turn conditions that isolate progressive narration. It spans 5,078 instances across six moral dimensions and preserves the same pre-verified responsibility structure across conditions.

  • Source processing: One-sided Reddit and Weibo narratives are distilled into conflict cores and converted into narrative shards through controlled processing.The pipeline includes segmentation, rephrasing, refinement, and generation before quality screening and human review.
  • Narrative conditions: Each conflict core yields three aligned conditions: neutral third-party description, single-turn self-protective narration, and five-turn progressive narration.The aligned design supports controlled comparison of narration format while keeping the underlying conflict matched.
  • Quality control: Every retained sample preserves the same non-empty responsibility structure across C0, C1, and C2 before model evaluation.The judge assesses recognition of pre-verified responsibility rather than responsibility differences introduced after construction.
  • Narrative conditions: The five-turn condition delays responsibility cues while progressively presenting grievance, justification, action, outcome, and a final judgment request.This structure requires models to track evolving context and integrate delayed responsibility cues.
  • Benchmark scope: 5,078 evaluation instances cover six moral dimensions—Emotion, Fairness, Loyalty, Role Duty, Norms, and Autonomy—and 24 task types.Each core dimension is divided into four sub-dimensions, with instances available in English and Chinese.

3 Experiments

The experiments evaluate narrative captivity across 17 proprietary and open-source LLMs, tracing its prevalence, training origins, multi-turn amplification, and mitigation. Results show substantial variation in resistance and recovery, with multi-turn behavior exposing distinct early-commitment and irreversibility risks.

  • 17 representative LLMs across 9 model series are evaluated, covering proprietary and open-source systems.
  • The evaluation measures Narrative Hold, Narrative Recovery, and End-state Hold using stance judgments assigned across T user turns.The stance label distinguishes identifying the narrator’s responsibility from aligning with the narrator’s framing.
  • Sampling uses controlled temperature and top-p settings, with temperature = 0.8 and top-p = 0.95 selected to balance response diversity and judgment stability.
  • RQ1: Overall Captivity Results: Across moral foundations, even leading models show limited resistance: the highest single-dimension NH reaches only 0.49, while most NH values fall below 0.3.Captivity severity varies by dimension, with NH for Fairness and Norms generally below 0.2 and NR for Loyalty above 0.7 for most models.
  • RQ1: Overall Captivity Results: NH and NR are nearly uncorrelated, separating early-commitment risk from irreversibility risk across four behavioral profiles.The profiles are Resilient, Wavering, Brittle, and Capitulating; Brittle and Capitulating contain both Llama models, while thinking improves recovery without preventing initial yielding.

4 Discussion

The discussion identifies preference optimization and multi-turn self-locking as central sources of narrative captivity, while inference-time interventions provide only partial relief. Multi-turn narration exposes vulnerability that single-turn evaluation can conceal.

  • Post-training stage impact: DPO consistently aggravates narrative captivity in both Tulu3 and OLMo3, with its largest effect on NR.SFT has family-dependent effects, while RLVR barely changes either metric.
  • Single-turn versus multi-turn narration: Multi-turn narration produces over 31 percentage-point drops for Doubao-Seed-2-Pro and Qwen3.5-27B, exposing brittleness hidden by single-turn evaluation.GPT-5.5, Claude-Opus-4.6, and Claude-Sonnet-4.6 converge to 0.56–0.58 in C2.
  • Single-turn versus multi-turn narration: By T5, pushback and hedging decline by at least 20% relative to C1 despite comparable C2-T1 behavior.Because C1 and C2 are informationally equivalent, the authors attribute the later decline to multi-turn structure.
  • Single-turn versus multi-turn narration: Early concessions become context constraints that progressively lock later responses toward the narrator’s stance.The model’s own prior generations create accumulated common ground, making reversal inconsistent with earlier statements.
  • Inference-time intervention analysis: M1 and M2 improve both models, whereas M3 helps GPT-5.5 but aggravates captivity on GLM-5.1.The differing M3 outcomes indicate that step-by-step analysis depends on model reasoning ability.
  • Inference-time intervention analysis: Inference-time interventions reduce per-turn accommodation but cannot break cumulative locking from prior concessions.The discussion therefore places the remaining mitigation challenge at the level of training data and optimization objectives.

5 Conclusion

The paper constructs a benchmark for narrative captivity in interpersonal-conflict advice and evaluates 17 models. It finds captivity widespread, traces it primarily to preference training and accumulated concessions, and concludes that inference-time interventions only partly alleviate it.

  • Contributions: The benchmark evaluates narrative captivity in interpersonal-conflict advisory settings across 17 models and finds the phenomenon widespread.The contribution includes both benchmark construction and broad model evaluation.
  • Main findings: Preference training is identified as the primary source of captivity, while turn-by-turn yielding progressively deepens and hardens it.The conclusion links training effects with the difficulty of reversing captivity during multi-turn narration.
  • Implications: Inference-time interventions only partially alleviate captivity and fail to produce significant improvement.The paper proposes incorporating judgment independence into future preference-alignment objectives.

6 Limitations

The benchmark’s scope limits generalization across languages, cultures, model families, personas, and time. Its fixed evaluation design also constrains which captivity patterns it can reveal.

  • The benchmark covers only English and Chinese Reddit and Weibo narratives, limiting tested generalizability to other languages and cultural contexts.
  • The 17 evaluated models cannot exhaust architectural and post-training diversity, so other models may exhibit different captivity profiles.
  • Binary stance judgments and a fixed five-turn setting may overlook intermediate patterns such as gradual stance softening.
  • Persona-specific effects remain untested because user personas were not explicitly manipulated.
  • The benchmark is a time-bound snapshot of current model versions rather than a permanent ranking.

7 Ethics and Societal Impact

The paper frames narrative captivity as a risk in real advisory use and distinguishes it from related agreement phenomena. Its evidence concerns non-coercive, progressively unfolding moral consultation rather than explicit persuasion.

  • Narrative captivity may validate unfair attributions and disadvantage the absent party when models adopt a narrator’s one-sided account.
  • Existing sycophancy studies largely examine explicit confrontation or persuasion, leaving non-coercive loss of evaluative independence open.
  • The benchmark’s source context is interpersonal moral consultation, where events are disclosed progressively without intentional manipulation.
  • Narrative captivity differs from sycophancy because the user need not state an opinion, oppose the model, or apply pressure.
  • The phenomenon accumulates when one-sided evidence is treated as complete and the model’s early concessions constrain later judgments.

B Benchmark Comparisons

The benchmark is designed to make interpersonal moral judgment more realistic and experimentally controllable than prior comparison settings.

  • The benchmark uses multi-turn interpersonal interactions so moral evaluations emerge progressively through conversational context.
  • It uses low-explicit-pressure settings that avoid direct moral prompting or artificially imposed value conflicts.
  • Its scenarios come from real-world interpersonal narratives rather than synthetic templates or manually constructed hypotheticals.
  • Explicit information controllability enables systematic regulation of narrative exposure and contextual asymmetry.

C Definition of Controlled Narrative Samples

Controlled narrative samples preserve one interpersonal conflict while varying perspective, wording, and disclosure across aligned conditions. The construction rules are intended to isolate narrative effects without changing morally relevant facts.

  • Sample components: Each sample is built from one conflict core and instantiated into three aligned narrative conditions.
  • Sample components: The conflict core contains the parties, relationship, key action, explanation, outcome, reaction, and responsibility cue needed for moral judgment.
  • Narrative shards: Narrative shards are atomic semantic units that can represent actions, justifications, reactions, consequences, emotions, or responsibility cues.
  • Narrative shards: Valid sharding must cover the complete conflict core while adding no facts beyond the source and admissible narrator framing.
  • Narrative conditions: C0 is neutral third-person narration, C1 is first-person single-turn narration, and C2 is first-person narration distributed across five turns.
  • Validity properties: The three conditions must preserve the same conflict core, differing only in perspective, wording, and disclosure schedule.
  • Validity properties: C2 must progressively introduce new information or framing rather than mechanically repeat a split single-turn text.
  • Validity properties: C2 must not directly pressure the model, while narrator responsibility must remain recoverable from the complete sample.

D.1 Sample Construction

The benchmark constructs aligned conflict triplets that separate factual content, narrative perspective, and multi-turn disclosure, then filters them for quality before inclusion. Human review follows scalable LLM screening, yielding 5,078 retained samples from 155,249 raw narratives.

  • The pipeline collects public Reddit and Weibo narratives, segments them into minimal narrative shards, and rephrases them without changing factual content.
  • C0 provides a neutral third-person reference, C1 adds self-protective single-turn framing, and C2 distributes the same facts across progressive multi-turn disclosure.
  • LLM screening evaluates factual alignment, contrast validity, and contextual clarity to identify factual drift, confounded contrasts, and ambiguity.
  • The weighted quality score combines factual alignment, contrast validity, and clarity using learned dimension weights, with α + β + γ = 1.
  • 5,078 samples were retained from 155,249 raw narratives, an overall pass rate of 3.3%.
  • Human experts provide final quality control, while retained samples show high semantic consistency and near-complete preservation of responsibility cues across C1 and C2.

E Human Agreement

The study validates GPT-4o stance labels against independent human judgments on randomly sampled C2 responses. Agreement is substantial across turns, with lower agreement in the more ambiguous middle turns.

  • 500 randomly sampled C2 responses from all 17 models were independently labeled by three human experts and compared with GPT-4o judge labels.
  • κ ranged from 0.76 to 0.85 across turns, indicating substantial judge–human agreement.
  • T3 and T4 had the lowest agreement at κ = 0.76 and 0.77, whereas T1 and T5 had higher agreement as context was clearer at the endpoints.

F Benchmark Structure and Statistics

The benchmark organizes interpersonal conflicts across six moral dimensions and 24 sub-dimensions, using aligned multi-turn scenarios and repeated sampling protocols. Decoding settings were selected through a pilot grid search before the main experiments.

  • The benchmark’s six moral dimensions are divided into 24 sub-dimensions grounded in Moral Foundations Theory.
  • Figure 6 presents the six dimensions, 24 sub-dimensions, and five-turn one-sided dialogues illustrating progressive narrative disclosure.
  • Each scenario is generated 10 times per condition, with reported metrics averaged across runs to mitigate sampling variance at temperature > 0.
  • A 4 × 3 pilot grid searched temperatures {0.2, 0.4, 0.6, 0.8} and top-p values {0.80, 0.90, 0.95} on the full C2 dataset.
  • T = 0.8 and p = 0.95 achieved the highest mean and lowest standard deviation on both NH and NR, so this setting was used in the main experiments.

H Full Sub-dimension Results

The full results are reported across all 24 moral sub-dimensions, both English and Chinese languages, and 17 models under the C2 condition. The appendix also specifies the prompts and mitigation procedures used to assess responsibility recognition and independent judgment.

  • Full Sub-dimension Results: Tables 13–18 report NH and NR for all 17 models across sub-dimensions and languages under C2.
  • Benchmark Taxonomy: The benchmark taxonomy covers six primary moral dimensions and 24 sub-dimensions of interpersonal conflict.
  • Benchmark Statistics: Table 11 gives exact sample counts for all 24 sub-dimensions across English and Chinese.
  • Decoding Setup: Table 12 identifies T = 0.8 and p = 0.95 as the selected decoding setting for the pilot grid search.
  • Construction Prompts: The narrative-shard and conflict-core prompts support segmentation, neutral conflict extraction, and controlled condition generation.
  • Evaluation Prompts: The evaluation prompts classify whether responses identify the narrator’s responsibility in single-turn and multi-turn conversations.
  • Mitigation Strategies: Mitigation prompts instruct models to maintain independent judgment, expose reasoning about missing information, or recap prior user turns.
Loading 2609.03407v1…