Source-linked AI summary
Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
Heather Renze
TL;DR
Systems that generate autobiographical scenes have not been audited against a documented record at the episodic level. This study conducts a scene-level audit of 366 LLM-generated daily narratives, finding 96.7% verification failure, while independent re-rating replicates the binary result and corpus grounding reduces failure to 83.3%.
Problem
Episodic scenes in systems that simulate people have not been audited against the documented record of the life described.
Method
The study mechanically analyzes pre-existing scene-level verdicts for 366 daily first-person narratives against an independent verification corpus using a fixed four-verdict rubric.
Results
96.7% of days fail positive scene-level verification, with grounded drift dominant; independent re-raters replicate the binary headline, while grounding reduces verification failure from 100% to 83.3%.
Takeaways & Limitations
Positive corroboration should be required for persona-system claims about real people, because unverified scenes cannot be treated as probably accurate.
Takeaways & Limitations
The original verdicts were assigned by an LLM auditor, and the four-way taxonomy shows only fair-to-moderate agreement, especially at the WEAK/UNVERIFIED boundary.
Abstract
from arXiv · showhide
When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are aware of, based on an unsystematic literature search. The subject and the author of this paper are the same person: a 366-day "page-a-day" book of first-person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day's quote - not her corpus - and every day was subsequently audited at the anecdote-scene level against an independent verification corpus using a four-level rubric fixed before analysis. We define the verification-failure rate as the share of days not rated VERIFIED (scene positively corroborated): 354 of 366 days fail, 96.7% (Wilson 95% CI 94.4-98.1%). Only 12 days contain a corroborated scene; 19 days (5.2%) assert claims actively contradicted by the record; the dominant failure mode is grounded drift - real people, employers, and settings inside invented scenes - though its measured share varies across raters. Independent re-rating replicates the headline (no evidence the original rate was inflated) while showing that the four-way taxonomy has only fair-to-moderate reliability. Regenerating the same days with current named models reproduces 100% verification failure under the same inputs; grounding generation in the subject's corpus significantly improves the verification rate while leaving substantial residual failure (83.3%). We contribute the measurement, a reusable audit instrument whose WEAK/UNVERIFIED boundary we show to be unreliable, and a grounding remedy with quantified effect.
1 Introduction
The paper audits whether LLM-generated autobiographical scenes match the documented record of the life described. It frames scene-level verification as an unresolved evaluation gap and reports a naturalistic self-audit designed to measure it.
- Motivation: Scene-level claims about a person’s life have not been audited against that person’s documented record, despite growing personal uses of simulation systems.Prior evaluation has emphasized attitudes, survey answers, and surface facts.
- Study overview: 366 first-person daily entries were drafted with a conversational LLM and fact-checked scene-by-scene against an independent ground-truth corpus using a pre-fixed rubric.The documented inputs were the template and each day’s quote rather than the subject’s corpus.
- Research questions: The study asks what fraction of generated memories survive verification, how failures manifest, where contradictions concentrate, and what remediation they suggest.
- Contributions: 96.7% of the 366 day-narratives failed positive scene-level verification under the paper’s definition.The paper contributes a quantified audit, reusable instrument, remediation workflow, and inter-rater analysis.
- Positionality: The author and audited subject are the same person, shaping the paper’s consent, privacy, and self-audit implications.
2 Related Work
Related work measures hallucination, long-form factuality, persona fidelity, false-memory risks, and character drift, but not autobiography scene-by-scene against a documented human-life corpus. This paper fills that stated gap.
- Long-form factuality: Existing factuality benchmarks check atomic claims or long-form text against encyclopedic knowledge, whereas this audit uses scenes and an independent corpus documenting one life.
- Person simulation: Person simulators have been evaluated mainly on attitudes, survey responses, and surface or deep persona fidelity rather than episodic scene accuracy.
- Episodic evaluation: The audit adds an episodic dimension: employers, dates, and named events fail verification in the large majority of generated days.
- Risk complement: False-memory research documents risks from consuming LLM-generated suggestions, while this study examines unsupported scenes produced when an LLM writes a person’s past.
- Gap: No published study identified by the paper audits LLM-generated autobiography scene-by-scene against the described person’s documented record.
3 Data and Study Design
The study analyzes a 366-entry AI-padded autobiography against independently assembled records of the subject’s life. Existing labels and provenance were fixed before the paper’s mechanical parsing and aggregation.
- Suspect corpus: The audited corpus contains 366 long-form first-person entries totaling 198,949 words, generated during July–August 2025.The derived edition’s compressed pages were not treated as the audited text.
- Ground truth: The ground-truth corpus combines memoirs and anecdote ledgers, published writing and talks, a professional digital-twin knowledge base, music records, and targeted web checks.
- Independence: Generation artifacts, the page-a-day derivative, and ChatGPT backups were excluded from verification because they were not independent corroborating sources.
- Labels: Every day received one scene-level verdict in pre-existing month tables, with source citations and evidence notes; four critical author confirmations took precedence.
- Timeline: The entries were generated in 2025, audited on July 13, 2026, and mechanically analyzed on August 21, 2026 without modifying audit labels.
4 Method
The method assigns one pre-specified scene-level verdict per day, parses all 366 records with fixed scripts, and defines failure as any verdict other than verified. It also addresses adjudication and single-rater limitations.
- Rubric: Each anecdote scene is classified as verified, weak, unverified, or contradicted/flagged according to its evidential relation to the record.Verified requires positive corroboration; contradicted requires positive disproof; weak denotes a real setting or person with an invented specific scene.
- Scene audit: The auditor searches each day’s scene against the ground-truth corpus for corroboration or contradiction, records evidence, and flags candidate false premises.Subject rulings settle claims only the subject can determine, under a precedence rule.
- Reproducibility: Fixed parsing and analysis scripts produced 366 rows with zero duplicate keys and zero unparsed verdicts, then tallied verdicts and recurring-premise screens.
- Headline measure: Verification-failed day ≡ any day whose verdict is not verified: weak ∪ unverified ∪ contradicted.
- Headline measure: 354 of 366 days failed verification, yielding 96.7% with Wilson 95% CI [94.4%, 98.1%].The paper calls this verification failure rather than fabrication because unverified scenes may be real but unrecorded.
- Reliability: The original audit had no second rater, so it lacked an inter-rater statistic and could not measure rater drift retrospectively.A blind independent re-rate was designed to address this circularity.
5 Results
Across 366 generated day-narratives, only 12 scenes were positively corroborated, while failures were dominated by grounded drift and replicated under independent re-rating. Monthly distributions, recurring false-premise screens, source coverage, and model regenerations further characterize the audit and its remedy.
- Headline distribution: 354 of 366 days (96.7%, Wilson 95% CI 94.4–98.1%) failed positive scene-level verification.Only 12 days (3.3%) contained a corroborated scene.
- Headline distribution: 19 days (5.2%, CI 3.3–8.0%) asserted claims actively contradicted by the record.These claims misstated documented facts about a documented person.
- Failure modes: Grounded drift was the largest failure category, placing real settings, employers, and people inside invented scenes, but its measured share ranged from 43–82% across raters.Weak was the largest category in all three independent ratings of sampled data.
- Monthly variation: Six months contained zero verified days, while May contributed 4 of 12 verified days and August, April, and December contained 14 of 19 contradicted days.The repository records no generation order or dates for month files, so no chronological conclusion is drawn.
- Recurring false premises: 31 distinct days screened under nine recurring false premises spanning invented relations, capabilities or events, misattributed facts, and inverted dispositions.Three dates screened under two premises, and the education screen was overturned by author adjudication.
- Source coverage: 252 sourced audit rows carried 333 multi-label source-class mentions, with the two curated memoirs and anecdote ledger providing over 80% of verification coverage.This source distribution identifies personal archives as the available ground truth for the subject’s life.
- Inter-rater reliability: Binary verification-failure rates in independent re-rating were 96.7% for rater B and 98.3% for rater C, with no evidence that the original rate was inflated.Rater C found zero corroborated scenes in its 60-day sample.
- Inter-rater reliability: Four-way agreement was 63.3% (κ = 0.387) for rater B and 81.7% (κ = 0.573) for rater C, with disagreement concentrated at the WEAK boundary.The WEAK cell accounted for 21 of rater B’s 22 disagreements.
6 Discussion
The discussion argues that scene-level verification exposes structural confabulation that surface-fact or attitude evaluations can miss, while grounding and targeted rewriting offer partial remediation.
- Grounded drift was the dominant failure class: real settings, employers, and people appeared inside invented scenes.Its measured share ranged from 43% to 82% across raters because the WEAK/UNVERIFIED boundary was unstable.
- 96.7% of days failed verification, making corroborated scenes the exception rather than evidence of a normally reliable process.The paper reports zero verified days for six months and recurring false premises across months.
- 19 contradicted days show why unverified scenes cannot be treated as probably accurate when the record lacks corroboration.The recommended audit stance requires positive corroboration rather than interpreting silence as reassurance.
- Author adjudication and lesson-only rewriting provide a remediation workflow for unsupported scenes.Critical personal facts receive written rulings from the subject, while unsupported scenes are replaced with reflections and verified anecdotes.
- For digital twins, fidelity claims should distinguish profile, attitudes, and episodes because episodic content remained the trust-critical layer.The paper notes that the subject’s curated digital-twin knowledge base was not supplied to the autobiography generator.
7 Limitations
The study’s findings are bounded by a single-subject design, LLM-produced labels, finite ground truth, scene-level analysis, and limited evaluation of remediation. These constraints qualify how verification failure and its categories should be interpreted.
- One life, one corpus, and one generator deployment provide deep measurement but do not estimate a population fabrication rate.
- The original verdicts were LLM-produced, while blind re-ratings replicated the binary headline but found only fair-to-moderate four-way agreement.
- 96.7% of days fail verification, but 5.2% are provably false; unverified scenes may be real but unrecorded.
- The audit grades each day’s central anecdote scene rather than every atomic claim, so individually true sentences can occur within a weak day.
- Current models reproduced 100% verification failure without grounding, whereas grounding reduced failure to 83.3% but left substantial residual failure.
- The lesson-only rewrite and adjudication workflow were designed during the project, and their effect on reader trust remains unmeasured.
8 Ethics
The audit involves a structural dual role: the author is also the subject, and the study uses pre-existing documents rather than newly collected human-subjects data. Privacy protections and disclosed self-audit risks frame the ethical approach.
- The audited subject and paper’s author are the same person, making consent structural rather than procedural and avoiding new human-subjects data collection.
- Sensitive source files were excluded, health information was limited to author-confirmed rulings, and third-party names appeared only from the subject’s published corpus.
- Self-audit could bias labels toward a preferred narrative, so the study uses a preregistered rubric, evidence notes, written rulings, and inspectable tables.
- The paper discloses residual self-audit risk rather than claiming that its structural mitigations eliminate it.
9 Reproducibility Statement
The reproducibility package publishes code, derived data, computed results, and the complete verdict distribution while withholding source narratives and per-day evidentiary details to protect the named individual.
- The public package includes analysis scripts, derived data, computed results, provenance, tables, figures, bibliography, and a 366-row verdict dataset.
- The package omits source entries, replication entries, and per-day anecdote and evidence columns because publication would perpetuate machine-generated claims about a private person.
- The public analysis reproduces n = 366, the VERIFIED/WEAK/UNVERIFIED/CONTRADICTED distribution, the 96.7% verification-failure rate, confidence intervals, and monthly cross-tabs.
- The draft uses the plain article class for arXiv and requires only a venue-specific template change for submission.
A All Contradicted Days
The appendix identifies every day rated contradicted and provides abbreviated audit-row gists, with trims marking omitted list tails.
- Table 6 lists every contradicted day using abbreviated gists from the audit rows; trims indicate elided list tails.
B Replication Methods
The replication methods specify the sampled days, generation arms, model settings, retrieval procedure, blinded rating, and failure handling. Published artifacts and access checks support auditability and reproducibility.
- Auditability: Published scripts and result files make the replication auditable, while raw generations and rating transcripts remain withheld under the privacy rule.The appendix states that the published materials back every item while preserving privacy restrictions.
- Sampling: 60 of 366 days were sampled once using a published deterministic random seed and day list.The sample covered all months, with monthly counts ranging from 2 to 8 days.
- Generation inputs: Arms B and B2 received only the archived prompt and fixed quote continuation, while Arm C additionally received retrieved ground-truth excerpts.This distinguishes ungrounded generation from retrieval-assisted generation under otherwise shared inputs.
- Models and retrieval: Arms B and C used openai/gpt-5.4, Arm B2 used anthropic/claude-sonnet-5, and all calls used temperature 0.7 with max_tokens 1400.Arm C retrieval ranked corpus lines by keyword overlap after extracting up to six filtered quote keywords.
- Rating: Raters evaluated 240 anonymized, shuffled entries in fresh sessions without access to arm assignments, prior ratings, results, or the paper.Rater D covered Arms B/B2/C and rater E covered A1; access logs verified zero prohibited reads.
- Failure handling: API calls were retried up to four times, only complete responses were accepted, and all 180 generations succeeded.The initial non-blind rating pass was discarded wholesale, with its divergence reported separately.
C Premise Screen Day Lists
The premise screen identifies day-level hits for nine recurring false premises, with one premise reversed during adjudication.
- P1: P1 was flagged on Mar 13, Mar 16, Aug 8, Aug 17, and Aug 29.
- P2-P4: P2 through P4 were flagged on four, four, and five listed days respectively.P2: Oct 10, Dec 3, Dec 5, Dec 23; P3: Sep 2, Oct 3, Oct 17, Oct 30; P4: Apr 4, Apr 14, Apr 19, May 30, Oct 3.
- P5-P9: P5 through P9 were flagged on eight, three, three, three, and two listed days respectively; P8 was reversed on adjudication.The listed hits span June through December, with P8 specifically marked as reversed.