Source-linked AI summary

Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization

Aarik Gulaya

arXiv:2605.28969v2cs.CLcs.AIcs.HC

TL;DR

The paper asks how faithfully AI systems represent a person’s interpretation when making decisions on their behalf. It operationalizes an interpretive layer as a compressed Behavioral Specification and evaluates held-out behavioral predictions across context conditions, finding that the Specification improves representational accuracy, especially when interpretation is required, while sometimes interfering with literal recall.

  • Problem

    The paper examines how to measure whether an AI system faithfully captures a person’s interpretation, because prediction on held-out reasoning is relevant to downstream alignment.

  • Method

    The study evaluates a compressed Behavioral Specification independently and alongside retrieval, raw-corpus, extracted-fact, and commercial memory contexts using a 5-judge primary panel.

  • Results

    The Specification lifts representational accuracy across subjects, recovers 75% of the corpus’s predictive benefit at 4% of the context cost, and produces its largest gains on interpretation-required questions.

  • Takeaways & Limitations

    Representational accuracy is distinct from recall: an interpretive layer helps when the model lacks an interpretive frame but can add little or hurt when literal recall already suffices.

  • Takeaways & Limitations

    The all-LLM pipeline leaves class-level circularity unresolved, so the panel supports directional claims but not absolute-quality claims without human validation.

Abstract

from arXiv · show

If an AI agent makes decisions on a person's behalf, those decisions must align with its user. We introduce representational accuracy to measure how faithfully a system captures a person's interpretation. An interpretive layer is operationalized as a Behavioral Specification. Our reference implementation aggressively compresses a person's data into interpretive patterns, served as context to a language model. We evaluate the Specification on a prototype benchmark of held-out behavioral predictions scored by a calibrated 5-judge LLM panel. We test it independently and in composition with a range of context conditions: full raw corpus, full extracted facts, and four commercial memory systems (Mem0, Letta, Supermemory, Zep). Across 14 public-domain autobiographical corpora, the Specification lifts representational accuracy in aggregate and nearly eliminates model hedging. It recovers most of what the raw corpus delivers, at ~25x less context cost. The Specification lifts subjects toward a common predictive level regardless of pretraining baseline; the lift in absolute points is therefore largest where the baseline is lowest, suggesting the population of relevance is anyone not adequately represented in pretraining. Lift is greatest on interpretation-required questions, where providing an interpretive layer enables model behavior that extracted facts or raw corpus do not. Conversely, on recall-required questions, this layer can interfere rather than help. We conclude that representational accuracy is distinct from recall and that human-AI alignment is dependent on how accurately the user is represented. Representational accuracy makes that alignment testable.

1. Introduction

The paper argues that AI memory must represent how a person interprets situations, not merely recall facts, and introduces representational accuracy and a Behavioral Specification to test this. Across held-out behavioral predictions, the Specification improves most when models lack an interpretive frame, while offering little benefit or sometimes hurting literal recall and refusal-triggering questions.

  • 1.1 Recall is not interpretation: AI memory benchmarks optimize recall, but recall does not measure whether a system captures a person’s interpretive patterns.The paper distinguishes representational accuracy from recall, preference matching, and persona consistency.
  • 1.1 Recall is not interpretation: Representational accuracy measures how well an AI system’s internal model captures a specific person’s interpretive patterns.The target is transfer of those patterns to new situations the system has never seen.
  • 1.2 What we tested: The benchmark generates behavioral-prediction questions from held-out autobiographical text that the response model never sees, enabling subject-by-subject scoring across context conditions.Training text generates the Specification, memory systems, and fact pool; held-out text supplies the prediction targets.
  • 1.2 What we tested: The study evaluates nothing, retrieved facts, extracted facts, raw corpus, Behavioral Specification, and combinations, including controlled and native commercial memory configurations.Each meaningful input combination is treated as its own experimental condition.
  • 1.3 What we found: A Behavioral Specification lifts representational accuracy, with the largest benefit where the model’s pretrained baseline is lowest.Every one of 9 low-baseline subjects improved when the Specification was added to All Facts, with a mean lift of +0.89 points on the 1–5 rubric and 78.6% of 351 questions improving.
  • 1.3 What we found: The Specification moves 55% of low-baseline questions across at least one rubric anchor, but its effect depends on question type.It helps interpretation-heavy questions, can drift past literal-recall answers, and may produce principled refusals that the content-match rubric penalizes.
  • 1.3 What we found: The paper frames inspectability and traceability as requirements for representations used by agents acting on a person’s behalf.Fact-attribution memory allows auditing stored facts, whereas reasoning-trace specifications allow auditing what the system believes.

2. Prior Work, Industry Benchmarks, The Fifth Target

Existing AI personalization benchmarks measure recall, persona fidelity, or preference alignment, but not whether a system represents how a specific person reasons. The paper positions representational accuracy as a fifth target, tested through held-out behavioral prediction and supported by an auditable reasoning trace.

  • The fifth target: Representational accuracy is proposed as a fifth measurement target for how faithfully an AI system captures a specific person’s interpretive patterns.It is distinct from recall, preference matching, and persona consistency.
  • Existing targets: Current personalization commonly uses stated preferences and biographical facts, whereas this paper targets the interpretive layer that organizes experience and reasoning across new situations.The paper treats preferences and facts as downstream artifacts of that deeper interpretive layer.
  • The fifth target: Held-out behavioral prediction tests whether a system can apply a person’s interpretive patterns to new situations drawn from unseen text.The test generates how the subject would respond to held-out scenarios and measures transfer beyond retrieved facts.
  • Existing targets: 70% to 93% recall scores reported by commercial memory systems do not establish accurate representation of a user’s reasoning.Recall benchmarks test fact retrievability, while representational accuracy concerns reasoning about those facts.
  • Operationalization: Reasoning-level traceability lets users audit what the system believes about them, beyond fact-level attribution of stored claims.The paper presents this auditability as the minimum bar for a representation acting on someone’s behalf.
  • Operationalization: A Behavioral Specification compresses autobiographical material into schema-like structures intended to carry reasoning signals without storing every fact.Its design is computationally analogous to reconstructive, schema-driven human memory.

The instrument we use to measure representational accuracy is behavioral prediction

The paper measures representational accuracy through held-out behavioral prediction: models predict how a subject would respond, and judges score those responses against verbatim autobiographical passages. The evaluation also tests whether Behavioral Specifications produce a consistent, subject-specific shift while auditing judge calibration and rubric limitations.

  • Behavioral prediction: Held-out autobiographical passages provide unseen situations against which models’ predictions of a subject’s response are scored.The prediction task jointly tests whether behavioral patterns are capturable, whether the Specification carries them, and whether the model uses it.
  • Behavioral prediction: The specification-effect claim is a measured shift toward the subject’s demonstrated behavioral patterns, not a claim of newly acquired prediction capability.Scores are compared against held-out passages from the same subject.
  • Scoring and calibration: The primary aggregate averages 1–5 scores from five judges spanning Anthropic and OpenAI models, with a seven-judge panel used for sensitivity analysis.The primary panel was selected because its calibration behavior was conservative for the Specification-effect direction, while Gemini judges served as a cross-provider robustness check.
  • Scoring and calibration: Primary judges converge strongly on condition rankings, with Spearman ρ from 0.86 to 0.93, while absolute agreement is lower at α = 0.659.This separates reliable directional comparisons from less stable absolute score magnitudes.
  • Validity limits: The evaluation cannot establish that higher-scoring responses are absolutely correct for the subject because human annotation against the subject’s writing was not performed.The panel supports a cross-provider directional claim, not definitive correctness.
  • Behavioral prediction: A Behavioral Specification encodes a person’s behavioral patterns in a structured document served as model context.The reference pipeline distills recurring reasoning patterns from the training-half corpus into a document of approximately 7,000 tokens.

4. Results

The Behavioral Specification improves representational accuracy most for subjects and questions where the model lacks an interpretive frame. Its gains come from changing response categories and enabling pattern-based inference, while offering little or negative value when baseline knowledge or factual recall already suffices.

  • Boundary conditions: On factual-recall questions, retrieval alone is often sufficient, while the Specification adds little or actively degrades responses.On high-baseline subjects, it adds little or mildly hurts across conditions.
  • The cross-subject gradient: +0.89 points: on 9 low-baseline subjects, the Specification’s mean per-subject lift over baseline was +0.89 on the 1–5 rubric.Every low-baseline subject improved over the no-context baseline; the Specification adds value through layering rather than replacing facts or raw corpus context.
  • Per-question categorical movement: 55.0% of low-baseline questions crossed at least one rubric anchor upward when the Specification was added.Two-or-more-anchor jumps occurred on 18% of low-baseline questions, with roughly 6% showing jumps of three or more anchors.
  • Mechanisms of improvement: Interpretive patterns let the model correct wrong referents and directionally wrong predictions when surface facts or generic defaults fail.Examples include resolving Katsu rather than Kimura and predicting that Cortés would refuse help because self-reliance structured his authority.
  • Worked response categories: A confident wrong answer, an honest abstention, and a wrong referent illustrate that categorical scores distinguish engagement from behavioral alignment.The abstention example received a 2.80 mean because judges disagreed whether refusing to predict deserved credit or penalty.
  • Mechanisms of improvement: The Specification can encode an organizing principle that complete extracted facts do not announce, enabling inference across a person’s recurring behavior.In the Cortés example, naming performative self-reliance allowed the model to override its generic “good leaders accept help” default.

Spec only

The Behavioral Specification delivers substantial predictive gains in a compact context, especially where baseline context is sparse, while its effects depend on person-specific content and can complement raw narrative detail.

  • The Specification produces the largest categorical movements where prior context is sparsest, with 9–15% multi-anchor crossings from the No-Context Baseline.When layered onto full facts or corpus, crossings fall to 2–3%, indicating substantial bidirectional per-question movement despite small mean deltas.
  • 70.9% of questions improve with Spec alone at roughly 25× less context than the raw corpus, while corpus plus Spec reaches 83.7%.All Facts + Spec matches the raw corpus’s improvement rate; median improvement is +1.00 rubric points, versus −0.40 when performance worsens.
  • Hamerton is the only low-baseline subject where Spec alone exceeds the raw corpus, scoring 2.63 versus 2.27.Corpus + Spec reaches 3.09, indicating complementary information from the structured Specification and autobiographical corpus.
  • Ebers shows the principal compression boundary: Spec alone scores 1.54 versus 2.18 for the raw corpus, a 0.64-point gap.The raw corpus contributes anecdotal specificity, including childhood incidents, named mentors, and direct autobiographical quotations, that the Specification abstracts into axioms.
  • 75% of the corpus’s predictive benefit is recovered by the Specification at 4% of the context cost.The remaining ~25% is associated with context sizes described as production-prohibitive.
  • Correct Specification content, rather than formatting alone, drives the effect: correct Spec yields ∆=+0.35, random mismatch +0.15, and adversarial mismatch −0.25.Models cited Specification-specific tags on 78.6% of correct-Spec responses versus 50.0% of wrong-Spec responses, and flagged content mismatch in 60.6% of wrong-Spec responses.

Wrong-Spec examples: Ebers Q7 (identity), Bernal Díaz Q16 (frameworks), Seacole Q2 (inference)

Wrong Specifications can produce large accuracy losses when their content mismatches the named subject, while coincidental convergence can preserve the surface action but reduce rationale precision. These examples show that representational accuracy depends on matching the person’s interpretive framework, not merely predicting behavior.

  • Ebers Q7 (identity): 2.00 points: a mismatched Specification scored 1.60 versus 3.60 for the correct Specification when the model detected that the content did not fit the named subject.The mismatch involved Ebers and interpretive content belonging to Equiano; the model named the served anchors and declined to predict about Ebers.
  • Bernal Díaz Q16 (frameworks): 0.20 points: the wrong Specification scored 4.60 versus 4.80 for the correct Specification when both predicted refusal through different moral frameworks.The correct framework produced a rationale matching Bernal Díaz’s battlefield memoir register, whereas the wrong devotional framework predicted the same surface action in an alien register.
  • Seacole Q2 (inference): 3.60 points: a wrong Specification scored 1.40 versus 5.00 for the correct Specification in a clean cross-century, cross-culture mismatch.The model detected the mismatch between Mary Seacole and the Spanish-conquest anchors and refused to apply the unrelated framework.
  • Aggregate interpretation: Wrong-Spec aggregates were −0.25 under adversarial pairing but +0.15 under random pairing because coincidental overlaps sometimes produce the same surface behavior.The near-tie in the coincidence case is real, but it reflects different logic and does not establish that the wrong Specification represents the subject accurately.
  • Cross-system retrieval context: 35.9% of system-pair/question instances shared zero top-10 facts, while mean pairwise overlap was 8.3%.Native retrieval produced zero exact-string overlap across the four systems, and semantic matching raised overlap only to 0.004 at cosine ≥0.85.

5. Discussion

The study argues that a Behavioral Specification adds a measurable interpretive layer to AI personalization: it helps models predict person-specific behavior beyond recall, especially when pretraining is thin, while sometimes interfering with recall. The findings also motivate user ownership and selective activation of the Specification.

  • Interpretive layer: Across 14 subjects and five memory-system configurations, the Specification increased representational accuracy on held-out behavioral predictions and moved models from generic or refused responses toward subject-specific predictions.The effect was observed independently and when layered with other context conditions.
  • Gradient and activation: The Specification’s benefit is greatest when the model lacks an interpretive frame, whereas it can interfere when pretraining already supports a substantive answer.This pattern supports selective rather than universal attachment of the Specification.
  • Recall versus interpretation: The paper distinguishes interpretive prediction from recall by showing that identical retrieved facts can support different plausible responses, while the correct interpretive pattern selects the subject-consistent one.The Fukuzawa-Cortes case illustrates this divergence on a held-out question.
  • Content specificity: Matched Specifications outperform adversarial wrong Specifications, although random subject pairings sometimes produce correct behavioral outcomes.The reported deltas are −0.25 for the wrong-Spec condition, +0.35 for the correct Specification, and +0.15 for random derangement.
  • Hedging: The matched Specification reduces broad-pattern hedging from 41.2% to 0.4%, but the mechanism behind this effect remains poorly understood.The paper identifies unresolved questions about whether models detect Specification mismatch through internal coherence or pretraining knowledge.
  • Compression and ownership: The Specification is roughly 25× smaller than the raw corpus while recovering about 75% of its signal, and the paper frames faithful user-held representations as central to personalization.Operational usefulness also depends on provenance verification and a trust network.

6. Limitations

The paper’s evidence is bounded by a selected historical-autobiography sample, an all-LLM measurement apparatus, and pipeline choices that were not systematically varied. These constraints limit claims about generalization, absolute scores, prompt sensitivity, and specification stability.

  • Sample: The 14 main-study subjects are a selected sample rather than a population, limiting external validity across source populations and living users.The sample is drawn from public-domain autobiographies and includes a single living-subject constraint.
  • Source limitations: Public-domain selection, self-presentation, translation, historical era, and autobiography genre constrain generalization beyond preserved historical narratives.The paper specifically notes that modern work, family, technical, and digital-native contexts are not sampled.
  • Measurement: All questions, responses, and judges are LLM-generated, so calibration checks do not resolve possible class-level LLM circularity with human evaluators.The 5-judge and 7-judge checks address within-provider circularity but not the broader concern.
  • Model coverage: The main-study response model is Claude Haiku 4.5, while cross-provider testing covered only a small subset, so absolute effects may differ across response models.The Specification-effect direction reproduced on 5 of 6 tested subject–response-model cells.
  • Evaluation stability: Prompt phrasing and inter-judge calibration were not fully invariant: rank order was stable, but absolute scores remained panel-specific.Pairwise Spearman ρ ranged from 0.86 to 0.93, while calibration differed across judges.
  • Reproducibility: Rerun variance was smaller than the cross-subject signal but still nonzero: pooled per-subject run-to-run SD of ∆_C4a was 0.10 versus cross-subject SD 0.59.The directional finding survived reruns, with 6 of 6 reruns positive across two low-baseline probe subjects.
  • Pipeline stability: Pipeline model choices and the tested pipeline version were not systematically varied, so different extraction, embedding, authoring, or composition models could produce different Specifications.The study used a frozen v0.2.0 pipeline for the 14 scored Specifications.

7. Future Work

Future work targets component-level understanding, cleaner memory comparisons, adaptive activation, user feedback, life-event updates, safety integration, and stronger validation of the study’s gradient and measurement claims.

  • Specification components: Ablating anchors, core, predictions, and brief layers separately could identify which components drive the observed response-pattern distributions and inform dynamic activation.The proposed study would serve each layer alone and in combinations.
  • Representation: Named-entity grounding should be studied alongside predicate structure because secondary analysis identified it as a contributing factor in Letta’s case-study lift.The Base Layer pipeline currently abstracts source text into structured predicates while retaining some personal detail in the brief.
  • Memory-system comparison: A cleaner Letta comparison would anonymize its source corpus and extend the corpus-size axis beyond the current B¯abur ceiling.These controls target naming asymmetry and corpus-scale effects.
  • Temporal validity: Snapshot Specifications need mechanisms for detecting canonical life events because a major shift can make pre-event reasoning patterns stale.The paper identifies career changes, conversions, losses, and stance reversals as examples.
  • Feedback: User corrections could update the affected predicate or anchor through re-extraction, re-authoring, and recomposition of the Specification.Edits, explicit corrections, and rejected answers are proposed as learning signals.
  • Safety: Safety integration remains open: 75 of 81 Spec-induced refusals were routine behavioral prediction rather than morally loaded, and malicious-intent users were untested.The paper assigns these questions to collaboration with AI safety researchers.
  • Gradient analysis: Future gradient work should use category-balanced batteries because Literal Recall composition explains part of the cross-subject gradient.Baseline contributed 63.6% of unique explained variance, compared with 6.9% for Literal Recall fraction.
  • Subject heterogeneity: Hamerton’s 15 of 60 extreme jumps warrant follow-up, but the present design cannot separate battery-generation effects from thin pretraining coverage.Hamerton’s C5 baseline was 1.26, the lowest in the main study.

B.8 Per-predicate ablation (Phase 2c)

The per-predicate ablation tests whether individual Specification sentences are load-bearing, but its null removal result is difficult to interpret because rerun stochasticity is substantial and the Specification may contain redundant evidence.

  • Method: The experiment removed or reversed heuristically identified predicates in 16 extreme-upward-jump cases using three temperature-0 response variants.Variants were original, ablated, and reversed, generated with Claude Haiku 4.5.
  • Results: +0.05 anchor points was the mean ∆_removal across 16 cases, with 95% CI [−0.35, +0.45].Only 2 of 16 cases showed ∆_removal ≥1 anchor, while 11 of 16 were below 0.5.
  • Results: −0.24 anchor points was the mean ∆_reversal, with 95% CI [−0.45, −0.02], indicating that reversed predicates changed responses more than simple removal.The reversal comparison is distinct from the removal effect.
  • Interpretation: Single-predicate removal did not measurably reduce response quality, but the paper attributes this null to possible redundancy rather than mechanistic inertness.Higher-level wrong-Spec evidence is cited as showing that the Specification as a whole does causal work.
  • Caveat: Original-condition reruns drifted by −1.44 anchors on average, with 9 of 16 cases shifting by more than 1 anchor, confounding ∆_removal with stochasticity.The extreme-upward-jump cases have higher pipeline variance than the per-subject mean grain.
  • Future tests: Future tests should use human predicate identification, larger N, irrelevant-predicate controls, and multi-predicate cluster ablations.The proposed expansion includes all 47 PATTERN_PREDICATE cases.
  • Interpretive grain: The broader study interprets behavior at the per-subject mean grain, so sentence-level ablation should not be treated as a direct test of the aggregate gradient.The canonical +0.89 figure averages per-subject ∆_C4a values across nine low-baseline subjects.
  • Context: The wider evidence suggests that Specification effects are mixed across individual questions, with both substantial increases and decreases rather than uniformly positive changes.For the Supermemory pool, 57 questions increased and 53 decreased by at least one anchor among 546 paired questions.

B.13 Memory-system Wilcoxon results (C1 vs C3) and low-baseline ∆_spec

The study compares retrieval-only conditions with retrieval plus Behavioral Specification across memory systems, using aggregate tests and low-baseline question-level analyses. The Specification adds value in several configurations, while facts plus Specification can outperform much larger corpus-plus-Specification context on a substantial subset of questions.

  • Wilcoxon results: Four system-configuration cells are significant at α = 0.01: Zep controlled, Letta controlled, Mem0 native, and Zep native.Mem0 controlled is significant at α = 0.05 but not α = 0.01; Letta native, both Supermemory configurations, and Base Layer are not significant at α = 0.05.
  • Wilcoxon results: Supermemory native has partial coverage, with N = 10 of 14 subjects and 7 of 9 in the low-baseline slice.
  • Low-baseline comparisons: 53.3% of low-baseline questions favored Raw corpus over Spec only, versus 30.8% favoring Spec only, with 56 ties.The comparison covers 351 questions across nine low-baseline subjects.
  • Low-baseline comparisons: 49.0% of low-baseline questions favored Corpus + Spec over Facts + Spec, versus 36.5% favoring Facts + Spec, with 45 ties.The 7K-token facts + Spec package therefore scores higher than the much larger corpus + Spec package on roughly one-third of questions.
  • Refusal-score audit: Memory-system refusals score +0.21 to +0.23 anchor points above pure No-Context refusals, regardless of whether they recite retrieved n-grams.The recitation comparison is ∆+0.027 with p = 0.67, indicating the effect is associated with the retrieval condition rather than visible quotation.

C.4 Pipeline models (specification generation)

The pipeline extracts constrained facts, authors layered behavioral representations, composes a unified brief, and generates held-out-question batteries. Evaluation uses multiple response and judge models, with controlled and native memory-system ingestion paths.

  • Specification pipeline: The extraction step uses claude-haiku-4-5-20251001 at temperature 0 to produce AUDN facts under a 46-predicate constrained vocabulary.
  • Specification pipeline: The authoring step uses claude-sonnet-4-6 to create three layers: anchors, core, and predictions.
  • Specification pipeline: The composition step uses claude-opus-4-6 to produce a unified brief, while battery generation uses claude-haiku-4-5-20251001 with backward design from held-out corpus material.
  • Response generation: GPT-5.4 independently regenerates responses for 13 global subjects, providing a separate response-model check.
  • Evaluation: Five calibrated primary judges score responses independently on a 1–5 scale after viewing the held-out passage, subject context, question, and response.Gemini judges are used only for sensitivity analysis, not the five-judge primary panel.
  • Memory-system evaluation: Controlled memory configurations re-ingest identical extracted facts, whereas native configurations ingest the raw corpus through each provider’s own pipeline.All five systems use top-k = 10 retrieval with the question text as the query.

C.7 Ingestion exclusions and failure cases

The evaluation includes provider-specific ingestion constraints and exclusions that limit comparability for some configurations. These boundaries are reported as part of the frozen condition matrix and study coverage.

  • Exclusions: B¯abur’s 422,772-word source corpus exceeds the Haiku context window, so the corpus-plus-Specification condition C9 is excluded for that subject.C9 numbers are therefore reported for 13 of 14 subjects.
  • Provider limitations: Letta native retrieval has a 0.34–0.47 deduplication ratio, so a top-10 list often contains only 3–5 unique facts.
  • Provider limitations: Zep graph retrieval can favor entity-dense chunks over behavior-dense chunks, creating a provider-specific retrieval bias.
  • Provider limitations: Mem0’s native ingestion may reformulate facts through infer=True, while controlled evaluation uses infer=False to preserve identical inputs.The native variant retains infer=True as the realistic deployment path.
  • Study scope: The condition matrix was frozen before scoring, with later additions such as Tier 2 replication and random-derangement draws reported separately.

D.2 Per-subject anchor-crossing on the low-baseline band

On the nine-subject low-baseline band, extracted facts plus Specification shift many paired responses upward across integer rubric anchors. The pattern is broad across subjects, with limited downward movement and an identifiable outlier.

  • Aggregate crossings: 55.0% of 351 low-baseline questions crossed upward from C5 to C4a, 6.8% crossed downward, and 38.2% stayed in the same anchor.Anchor-crossing rate is defined from paired C5 and C4a responses whose 5-judge means land in different integer rubric anchors.
  • Per-subject pattern: Eight of nine low-baseline subjects fall within a 48–74% upward-crossing range.Downward-crossing rates remain at or below 15% for every low-baseline subject.
  • Per-subject pattern: B¯abur is the low-baseline outlier below 48% upward crossings, with four downward crossings out of 39.The passage associates this with his 422K-word corpus and partial pretraining exposure.
  • Per-subject pattern: Sunity Devee has a 74.4% upward-crossing rate alongside a notably low C5 baseline of 1.03.
  • Judging reliability: The five primary judges show Spearman ρ of 0.86 to 0.93 in directional agreement, while absolute magnitude varies with Krippendorff α = 0.659.Including Gemini lowers Krippendorff α to 0.535, and Gemini judges show systematic +1-point inflation relative to the calibrated primary panel.

F.8 Persona-input depth comparison across benchmarks

Persona-input depth spans roughly three orders of magnitude across benchmarks, but input form matters more than size. The Behavioral Specification occupies a middle depth while encoding interpretive patterns rather than transcripts or surface descriptors.

  • Behavioral Specification: ≈7,000 tokens places the Behavioral Specification in the middle of the depth range, while its authored compression represents how the subject reasons.Its anchors, core, and predictions form a unified brief rather than a transcript or descriptor.
  • Compression: 5× to 80× compression ratios are reported between per-subject Specifications and their source corpora.Specification sizes are approximately 4,000–7,000 tokens depending on subject.
  • Depth comparison: ~50 to ~32,000 tokens span the reported persona-input depths, from PersonaGym’s descriptor to Twin-2K’s full transcript.The comparison characterizes persona-input depth as the subject-specific representation served at inference time.
  • What depth measures: Twin-2K’s deeper transcript measures response-distribution completion, not interpretive-pattern capture on unseen situations.Its task is Likert interpolation across the same survey instrument.
  • Input form: PersonaGym and Twin-2K use raw inputs in different senses: surface attributes versus serialized prior responses.The comparison therefore distinguishes descriptor content from transcript content, not merely context length.
  • Memory-system contrast: Letta’s memory block grows with corpus length and mixes verbatim text, paraphrases, and agent-generated synthesis up to an ingestion ceiling.This contrasts with the Specification’s fixed-shape, deterministic pipeline and schema.
  • Trade-off: The Specification trades source-text texture for aggressive, deterministic structural compression.Letta retains voice, vocabulary, and syntax, whereas the Specification preserves a consistent schema across subjects.

G.2 Headline result (5-judge primary)

In the exploratory three-subject comparison, Letta’s self-edited memory block outscored the full-stack Behavioral Specification, with the advantage varying by corpus size and question type. The comparison also identifies complementary strengths: named-entity grounding favors Letta, while principle-driven questions can favor the Specification’s hedging.

  • Headline result: +1.21 was the largest Letta-over-Specification gap, observed for Ebers; Letta scored higher on all three subjects.The gap was smaller at the Hamerton and B¯abur endpoints.
  • Context growth: 22.5K, 68.4K, and 335.3K characters were the Letta block sizes for Hamerton, Ebers, and B¯abur, respectively.The Specification sizes were 34.6K, 39.7K, and 37.1K characters for those subjects.
  • Breakdown: 25.4% of B¯abur’s Letta sentences were verbatim repeats, while 56.1% had a near-paraphrase elsewhere at the ingestion ceiling.The block reached approximately 333K characters and began rewriting previously written content.
  • Leakage audit: 0 of 119 questions showed a five-word verbatim overlap with held-out passages for either Letta or the Specification.Score-delta correlation with surface-text similarity was Pearson −0.046, arguing against surface-text leakage as the lift explanation.
  • Question-type effects: Named entities explain Letta’s advantage on entity-rich questions, while Specification hedging can stay closer on principle-driven questions.Ebers q30 illustrates entity grounding; Hamerton q27, q29, and q51 illustrate Specification gains on principle-driven decisions.
  • Failure modes: Letta’s commit-without-value-anchor and the Specification’s value-anchor-without-context each create a structural failure mode.Letta can confidently invert principle-driven answers, whereas the Specification can produce rubric-penalized principled refusals.
  • Abstention: 17.9% versus 0% abstention occurred for B¯abur, and 10.3% versus 5.1% for Ebers, comparing Letta with the Specification.At Hamerton, abstention was 7.7% for Letta versus 10.3% for the Specification.

G.7 Exploratory: stacking the Letta block with the Behavioral Specification

A post-hoc probe concatenated Letta’s memory block and the Behavioral Specification to test whether their distinct representations conflict or compose. Stacking exceeded both representations alone on all three subjects, but the result is exploratory and not panel-matched.

  • Comparison: +0.07 / +0.08 / +0.10 was the stacked condition’s advantage over Letta alone across the three subjects.The ordering follows the subject rows in the reported stacked comparison.
  • Comparison: +0.34 / +1.28 / +0.48 was the stacked condition’s advantage over the full-stack Specification alone.The combined context outperformed the Specification on all three subjects.
  • Composition: The longer combined context did not dilute either signal, including for B¯abur’s duplicated ~335K-character Letta block.The two representations were interpreted as complementary rather than redundant.
  • Caveat: The stacked probe was an exploratory N=3 post-hoc comparison scored by four judges, whereas comparison columns used five-judge primary means.GPT-5.4 returned errors, preventing a panel-matched five-judge stacked result.

Appendix H. Glossary

The glossary defines the paper’s measurement vocabulary, artifacts, conditions, and interpretation patterns. Together, these terms distinguish recall from interpretive representation and specify how behavioral-prediction effects are scored.

  • Core concepts: Representational accuracy measures whether a model’s predicted response matches a subject’s response on held-out situations.Behavioral prediction is scored against the subject’s verbatim response using a 1–5 interpretive rubric.
  • Core concepts: A Behavioral Specification is a static approximately 7,000-token document encoding behavioral patterns through anchors, core, predictions, and a unified brief.It is layered above memory-system retrieval as an interpretive structure.
  • Conditions: C1 is Retrieval only, C2a is Specification only, C4 is All Facts only, C4a is Facts + Specification, and C9 is Raw corpus + Specification.These condition codes identify the context combinations used in the experiments.
  • Scoring: A cross-anchor interpretation rule treats a score change crossing an integer rubric anchor as a behavioral prediction shift.Anchor crossing is distinguished from a weaker within-anchor shift.
  • Hedging: Matched-Spec conditions reduce baseline hedging from approximately 41% to near zero, while Wrong-Spec produces approximately 60% hedging.Refusal or abstention includes explicit refusal language or low-confidence qualification.
  • Interaction patterns: Pattern 1 describes interpretive supply, Pattern 2 over-theorization, and Pattern 3 principled refusal in Specification–retrieval interactions.The Specification can improve underdetermined retrieval, degrade already-determined answers, or trigger refusal.
  • Scope: The population of relevance is low-baseline subjects whose interpretive patterns are not already represented in pretraining.The paper distinguishes this from users represented only through demographic or factual information.
  • Specification layers: Anchors state how the subject reasons, Core connects those anchors into patterns, and Predictions derive forward-looking decisions.The unified brief integrates all three layers.
Loading 2605.28969v2…