Source-linked AI summary

Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation

Divya Vetticaden, Arya Gupta, Julian Nyarko, Megan Ma

arXiv:2608.13712v1cs.CYcs.AIcs.CL

TL;DR

Legal-AI evaluations rarely test whether simulated witnesses sustain realistic, pedagogically useful behavior during depositions. WitnessSim separates these dimensions and evaluates them through adversarial, expert-comparison, trajectory, and intervention-based tests, finding plausible and responsive behavior alongside longer-horizon fidelity gaps.

  • Problem

    Deposition training needs dynamic witness behavior, but evaluations rarely test both behavioral credibility and pedagogical usefulness during sustained questioning.

  • Method

    WitnessSim is evaluated for realism and usefulness using adversarial tests, blinded attorney comparisons, longitudinal trajectories, and legally grounded intervention tests.

  • Results

    WitnessSim generally produced plausible, intervention-responsive behavior consistent with assigned challenges, but its longer trajectories were smoother, more compressed, and distinguishable from real testimony.

  • Takeaways & Limitations

    The findings support evaluating legal simulations across response plausibility, behavioral boundaries, intervention response, and realistic change over sustained interactions.

  • Takeaways & Limitations

    The evaluation uses a single opioid-litigation context and does not establish attorney learning or measurable improvements from repeated simulator use.

Abstract

from arXiv · show

Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simulator driven by controllable legal personas. We use an evaluation framework separating behavioral realism from pedagogical usefulness. We assess realism through adversarial testing, blinded attorney comparison, and analysis of longitudinal behavioral trajectories. WitnessSim generally maintained plausible behavioral boundaries, and attorneys did not systematically prefer either original testimony or WitnessSim generated testimony. Pedagogical tests showed that witness behavior changed meaningfully in response to question form and attorney intervention without uniformly collapsing the assigned persona. Together, these results showcase a model of behavioral fidelity in legal simulations, and provide a framework for evaluating its performance.

1. Introduction

The paper introduces WITNESSSIM, a controllable deposition-training simulator for practicing behavioral and interpersonal skills underrepresented in legal-AI benchmarks. It evaluates the system through separate behavioral-realism and pedagogical-usefulness dimensions using legally grounded tests and transcript-based analysis.

  • Motivation: Depositions require attorneys to manage nonresponsive testimony, force specificity, identify inconsistencies, and handle other interpersonal and behavioral challenges.Junior attorneys often rely on time-intensive, inconsistent mock exercises or high-stakes real depositions, especially outside resource-rich organizations.
  • Motivation: Current legal-AI benchmarks largely overlook these skills by emphasizing factual accuracy, legal reasoning, or isolated-question performance.The paper positions LLM-based simulation as a way to provide repeated and configurable practice opportunities.
  • System contribution: WITNESSSIM is a controllable, state-based system that simulates dynamic witnesses with distinct behavioral personas during deposition questioning.The system is designed to model witness behavior rather than only generate superficially plausible individual responses.
  • Evaluation framework: The evaluation separates behavioral realism from pedagogical usefulness rather than treating plausible responses alone as sufficient evidence of an effective simulator.Realism concerns human plausibility and coherent behavioral identity, while usefulness concerns strategically difficult interactions for training.
  • Evaluation framework: 300 deposition and trial transcripts were randomly sampled from 1,169 candidate transcripts in the National Prescription Opiate Litigation corpus.The framework also draws on legal training materials and experienced-attorney feedback to test responses to pressure, question form, sensitive topics, contradictions, and archetype-specific strategies.

2. Related Work

Prior work establishes LLM-based persona and affective simulation and supports professional role-play, but leaves open how to evaluate longitudinal witness behavior under sustained adversarial interaction. WITNESSSIM addresses this gap by jointly evaluating behavioral realism and pedagogical relevance in deposition simulation.

  • Prior Work: LLM agents have been used to simulate coherent social behavior, assigned personas, affective dynamics, and professional role-play across changing contexts.This related work spans social and persona simulation, affective and conversational behavior modeling, and simulation for professional training.
  • LLM-Based Social and Persona Simulation: Persona benchmarks evaluate plausibility and consistency across contexts using persona adherence, identity consistency, and conversational naturalness.Persona-based simulation occupies a middle ground between open-ended social modeling and replication of a specific individual.
  • Legal Simulation: Legal simulators and benchmarks emphasize courtroom procedure, legal reasoning, rule recall, issue spotting, and legal application rather than sustained individual-witness behavior.WITNESSSIM adds witness behavior across sustained questioning as a complementary evaluation unit and dimension.
  • Research Gap and Contribution: Prior work leaves open how to evaluate longitudinal behavioral fidelity and training relevance during sustained adversarial interaction, beyond plausible individual responses or apparent realism.WITNESSSIM evaluates whether witnesses maintain plausible identities over time and create recognizable examination challenges that respond meaningfully to attorney questioning.

3. Methods

WitnessSim models deposition behavior as a six-dimensional dynamical system whose state responds to questioning while mean-reverting toward a fixed archetype baseline. The simulator distinguishes behavioral dimensions that conventional Big Five projections can obscure and anchors generated dialogue to the evolving state.

  • State representation: Witness behavior is represented by six dimensions—composure, knowledge, agreeableness, verbosity, rigidity, and performance—that are deterministically updated from attorney-question features before informing an LLM.Knowledge, verbosity, and performance capture deposition-specific behaviors not cleanly represented by conventional Big Five dimensions.
  • Persona construction: The study retained 10 archetypes from 14 expert-informed candidates after using pairwise cosine similarity and angular separation to remove behaviorally indistinct profiles.The retained archetypes are combative, cooperative, defensive, dogmatic, inventive, loquacious, nervous, neutral, overconfident, and overprepared.
  • Persona construction: The six-state representation preserved distinctions that the Big Five projection compressed: Inventive and Loquacious witnesses had projected cosine similarity of approximately 0.996 versus approximately 0.944 in six dimensions.Their projected angular separation fell to approximately 5° from about 19°, obscuring the distinction between not knowing and talking excessively.
  • Question dynamics: Question processing combines linguistic pressure markers, sensitive-topic overlap weighted by intrinsic sensitivity, and leading-question form into a composite stress signal.Pressure markers accumulate additively and are clipped to [0, 1], while thresholds suppress incidental topic overlap.
  • State dynamics: Coupled updates apply immediate question responses and gradual mean reversion toward each archetype’s attractor, preventing short-term interactions from immediately overriding the assigned persona.Stress can reduce composure and cooperation, increase rigidity, degrade recall, and alter verbosity depending on topic sensitivity and question form.
  • Dialogue generation: Threshold-triggered behavioral events and the current state are appended to the LLM system prompt, anchoring generated dialogue in the simulator’s underlying dynamics.This mechanism is intended to limit surface-level behavioral drift during generation.

4. Evaluation

The evaluation separates behavioral realism from pedagogical usefulness and tests both through adversarial, blinded-comparison, longitudinal, and intervention-based analyses. WITNESSSIM generally preserved persona boundaries and produced responsive practice behavior, while longer-horizon trajectories revealed measurable fidelity gaps.

  • Behavioral realism: Behavioral realism was assessed through adversarial testing, blinded attorney comparison with authentic testimony, and longitudinal analysis of affective and behavioral dynamics.These evaluations targeted behavioral boundaries, contextual believability, and coherence across longer interactions.
  • Adversarial testing: 297/300 Cooperative witnesses avoided full admission under direct blame pressure, while no witness accepted the complete five-premise accusatory chain.In the escalating-chain test, 27.3% resisted earlier than intended; Evasive witnesses resisted sensitive material more than trivial material (rrb = .984, p < .001).
  • Blinded comparison: The associate selected WITNESSSIM in 30.0% of exchanges versus 20.0% for authentic testimony, and the senior litigator selected it in 32.7% versus 28.6%.Both evaluators also recorded ties, indicating no systematic preference for authentic testimony in these comparisons; a technically experienced counsel instead identified authentic sequences in 45/50 exchanges and marked five ties.
  • Longitudinal dynamics: The mean real–simulated emotion-trajectory correlation was r = .148, exceeding all 1,000 temporally permuted comparisons (p < .001).Synthetic trajectories were smoother and more compressed, and the first principal component separated real and simulated transcripts while explaining 29.9% of variance.
  • Pedagogical usefulness: Open questions elicited 175.9 words on average, compared with 86.5 for closed and 68.4 for high-pressure questions, while targeted interventions changed responsiveness without uniformly erasing archetypes.The intended Evasive Pin-Down trajectory occurred in 93.0% of contexts; the loquacious runaway-to-focused trajectory occurred in 59.3%, while 40.3% were focused from the outset.

5. Discussion

WitnessSim generally produced plausible, intervention-responsive behavior consistent with assigned witness challenges, but its longer-horizon trajectories remained smoother, more compressed, and distinguishable from real testimony. The discussion frames simulation as a way to support repeated professional practice while emphasizing that actual learning benefits remain to be tested.

  • Findings: WitnessSim generally produced plausible behavior that responded to meaningful attorney intervention and remained consistent with the assigned behavioral challenge.Across evaluations, the simulator maintained plausible behavioral boundaries without uniformly collapsing the assigned persona.
  • Findings: Trajectory analysis revealed longer-horizon fidelity gaps: simulated testimony shared some temporal structure with real testimony but was smoother, more compressed, and distinguishable.The trajectory representation exposed differences not captured by immediate response plausibility.
  • Behavioral realism: Realistic professional behavior is inherently multi-turn because rapport, strategies, attitudes, personalities, emotions, and behavioral tendencies can change over interaction.Realism depends on the person’s characteristics as well as the immediate prompt and conversational history.
  • Training implications: Simulation-based learning can create repeated opportunities to practice and develop professional judgment rather than replace its exercise with AI.The discussion presents this as an alternative to concerns about deskilling and erosion of human capabilities.
  • Limitations and future work: The evaluation establishes whether simulations contain behavioral challenges and responses that enable meaningful practice, not whether they improve learning outcomes.Future controlled studies could test effects on questioning strategy, adaptation to difficult witnesses, and transfer to new scenarios.

Limitations and Future Work

The study’s evaluation is limited to a single opioid-litigation context, and its emotion-vector analysis remains a proof of concept rather than a validated measure of witness affect. Future work should test broader case contexts and compare representations with human demeanor annotations.

  • Limitations and Future Work: The evaluation corpus covers only a single opioid-litigation context, limiting conclusions about generalizability.Broader evaluation across different cases is proposed as a next step.
  • Limitations and Future Work: The emotion-vector analysis is a proof of concept, not a validated measure of witness affect.Future studies should compare these representations with human annotations of demeanor.

Conclusion

WITNESSSIM combines a controllable, state-based deposition simulator with an evaluation framework that separates behavioral realism from pedagogical usefulness. The framework assesses coherent behavioral boundaries, meaningful responses to intervention, and realistic change across interactions, while acknowledging that attorney learning and full human-witness equivalence remain unestablished.

  • Conclusion: WITNESSSIM pairs a controllable, state-based deposition simulator with an evaluation framework separating behavioral realism from pedagogical usefulness.The framework was applied across adversarial tests, blinded plausibility judgments, and legally grounded examination tasks.
  • Conclusion: The simulator generally maintained meaningful behavioral distinctions while responding to changes in a deposition interaction.The conclusion emphasizes that professional simulation requires more than plausible individual responses.
  • Conclusion: The framework measures coherent behavioral boundaries, meaningful intervention responses, and realistic behavioral change separately, but does not establish attorney learning or full equivalence to human witnesses.These measures are intended to identify where simulation fidelity succeeds or breaks down as LLM-based training and simulation become more common.

Impact Statement · Appendix

The paper advances machine learning through a legal generative system while emphasizing risks from inaccurate, biased, or hallucinated outputs and the need for particular caution in high-stakes legal contexts.

  • Impact Statement: The work aims to advance the field of Machine Learning.
  • Impact Statement: Legal-data-trained generative models carry inherent risks.
  • Impact Statement: These models may produce inaccurate outputs that do not reliably reflect legal authorities.The passage identifies statutes, case law, and jurisdictional nuance as potentially misrepresented.
  • Impact Statement: The models may also generate biased or hallucinated outputs.
  • Impact Statement: Legal-domain errors can affect due process, client outcomes, and access to justice.
  • Impact Statement: Because of these stakes and limitations, the methods warrant particular caution.

A. Archetype Representation and Big Five Comparison … G.4. Human–Judge Agreement

The paper specifies WitnessSim’s archetype representation, behavioral state dynamics, corpus construction, dialogue-generation and judging procedures, and human-validation framework. It also identifies limitations of the Big Five projection and defines how human–judge agreement is interpreted.

  • A. Archetype Representation and Big Five Comparison: Ten retained witness archetypes are represented by six-state attractor vectors spanning composure, knowledge, agreeableness, verbosity, rigidity, and performance.The simulator began with fourteen expert-informed candidates and retained ten for evaluation.
  • A. Archetype Representation and Big Five Comparison: The approximate Big Five projection omits knowledge, and this can collapse behaviorally important distinctions, including Inventive–Loquacious angular separation declining from approximately 19.2° to 5.2°.Negative changes indicate compression, whereas positive changes indicate greater separation under the projection.
  • B. Question-Encoding Details; B.1. Pressure-Marker Dictionary; B.2. Sensitive-Topic Tokenization: Question encoding uses prespecified pressure markers and sensitive-topic tokenization, with substring accumulation, clipping to [0, 1], and wording-sensitive Jaccard matching.High-pressure markers contribute twice the increment of medium-pressure markers, while unrelated words can dilute topic similarity.
  • C. Behavioral Justifications for State Dynamics; C.1. Fatigue Multiplier; E. Derived Behavioral Score Rationale: State dynamics combine immediate forcing with smaller mean reversion toward archetype attractors, while fatigue makes pressure increasingly costly across the first twenty turns.Immediate effects range from 0.12 to 0.30, mean-reversion coefficients from 0.03 to 0.08, and a high-pressure question scales from ϕ1 = 0.05 to ϕt = 1 at turn 20 or later.
  • D. Per-Topic Recall Quality: Global knowledge and topic-specific recall quality serve distinct roles: focused pressure can reduce recall by 0.25, while recovery is only 0.02 per untouched turn and bottoms at 0.1.Returning from 0.1 to 1.0 requires approximately 45 untouched turns, making severe degradation effectively persistent within a typical examination.
  • F. Corpus Sampling and Warm-Up Procedure; F.1. Candidate Pool Construction; F.2. Supplementary Case-Document Retrieval: The evaluation corpus preserves random sampling while manually validating deposition eligibility, then retrieves up to five supplementary case documents through prioritized witness- and topic-based searches.The candidate pool comprised 1,169 documents, and uncertain witness-name extraction triggered case-level search.
  • F.3. Case-Preparation Extraction; F.4. Warm-Up Question Selection; F.4. Warm-Up Question Selection: Case preparation extracts sensitive and control topics, aggravating facts, exhibits, and prior-position probes, while warm-up uses up to ten early questions before the earliest evaluation-field occurrence.Fewer than ten eligible questions are retained without padding, and the extracted fields construct case-grounded adversarial and pedagogical prompts.

G.4.1. PEDAGOGICAL EVALUATION RELIABILITY … H.1. Scoring and Human Validation

The paper validates only sufficiently reliable pedagogical and adversarial measurements, while using structured protocols, independent human checks, and test-specific interpretation to evaluate behavioral robustness. Several fine-grained outcomes are excluded or reframed when agreement is inadequate.

  • G.4.1. PEDAGOGICAL EVALUATION RELIABILITY: U1 tone and pushback agreement was limited (weighted Cohen’s κ = .294 and .298), so these subjective fields were excluded from validated quantitative evidence.The deterministic U1 response-length outcome remained the principal quantitative measure.
  • G.4.1. PEDAGOGICAL EVALUATION RELIABILITY: T2 turn1-to-turn4 trajectory classification was retained as lower-confidence evidence, supported by PABAK = .640 and Gwet’s AC1 = .732.The classification fell within moderate agreement under PABAK and below the .70 PABAK threshold used for stronger confidence.
  • G.4.1. PEDAGOGICAL EVALUATION RELIABILITY: Several finer-grained T2 outcomes were excluded, whereas Turn 3 narrowing compliance was reliable but omitted because it represented a meta-instruction rather than a substantive question.Stays focused had PABAK = .160, Turn 2 narrowing compliance had PABAK = .370, and Turn 3 had PABAK = .970 and AC1 = .980.
  • G.4.2. ADVERSARIAL EVALUATION RELIABILITY: For AC1, the reliable primary measure collapsed self-protection into FULL ADMISSION versus retained boundary, replacing an unreliable five-way classification.The five-way field had PABAK = .311, while the binary collapse achieved PABAK = .837 and Gwet’s AC1 = .911.
  • G.4.2. ADVERSARIAL EVALUATION RELIABILITY: AC2 and AC3 classifications showed strong agreement, while sensitive-question evasion was retained despite moderate PABAK because Gwet’s AC1 corroborated it.AC2 had PABAK = .923 and AC1 = .948; trivial-question evasion had PABAK = .816 and AC1 = .839; sensitive-question evasion had PABAK = .671 and AC1 = .716.
  • G.5. Adjudication: No universal disagreement-resolution rule was used; annotations remained independent, and borderline cases were reviewed against their transcripts and test-specific rubrics.This preserved independent reliability calculations while allowing substantive downstream interpretation where needed.
  • H. Complete Adversarial-Test Protocols and Results: The adversarial evaluation used four tests across 300 case-grounded contexts, with cloned sessions ensuring identical conversational starting states across conditions.AC1 and AC2 used the Cooperative archetype, AC3 used Evasive, and AC4 used Cooperative, Evasive, and Combative archetypes.
  • H.1. Scoring and Human Validation: Except for deterministic AC4 repetition checks, an LLM judge scored behavioral outcomes, while independent human subsets assessed measurement reliability without replacing automated scores.Judges received response text and explicit rubrics but not latent state variables; prevalence-imbalanced fields used PABAK and Gwet’s AC1 where available.

H.2. AC1: Cooperative Self-Incrimination … I.4. Evaluators and Instrument Versions

Adversarial tests found that WitnessSim largely preserved archetype-specific behavioral boundaries, while contextual plausibility evaluation used blinded comparisons of authentic and generated four-response sequences. The evaluation design treated evaluator background and instrument version as interpretation-relevant rather than pooling judgments.

  • H.3. AC2: Cooperative Runaway Leading Chain: AC2: 218 of 300 contexts (72.7%) followed the intended trajectory, while 82 (27.3%) resisted the first question and none accepted the complete chain.Human and LLM classifications agreed on 96.2% of 52 annotated contexts, with PABAK = .923 and Gwet’s AC1 = .948.
  • H.4. AC3: Selective Evasion: AC3: Mean evasion was higher for sensitive questions (M = 1.02, SD = .322) than trivial questions (M = .451, SD = .217) in 86.0% of contexts.The Wilcoxon test yielded W = 597.5, p = 7.39 × 10−45, rrb = .984; 282 of 300 contexts (94.0%) met the secondary trivial-evasion criterion.
  • H.5. AC4: Repetition Attack: AC4: Mean irritation onset occurred at 2.00 for Combative, 2.13 for Evasive, and 2.33 for Cooperative, with composite pass rates of 100.0%, 99.7%, and 98.7%.The deterministic repetition analysis found zero exact or near-verbatim duplicates across all 900 sequences.
  • I. Contextual Plausibility Evaluation: Contextual plausibility was defined as whether multi-turn generated testimony reached a minimum expert-judged plausibility threshold in its local deposition context, not indistinguishability from real testimony.The evaluation focused on contextual plausibility rather than factual or legal correctness.
  • I.2. Stimulus Construction: Stimuli paired authentic and WITNESSSIM-generated four-response sequences from the same questioning context, with randomized Transcript A/B labels and preceding dialogue supplied.Four-response sequences enabled assessment of naturalness, multi-turn consistency, resistance, responsiveness, and demeanor.
  • I.3. Evaluation Instructions: Evaluators judged which sequence was more plausible, allowed ties or neither judgments, and considered contextual fit, naturalness, knowledge, resistance, responsiveness, and demeanor.They were instructed not to assess question quality, factual or legal correctness, persuasion, or party advantage.
  • I.4. Evaluators and Instrument Versions: Evaluator results were reported separately because backgrounds and instrument versions differed; the senior litigator’s Version 2 judgments were not pooled with Version 1 evaluations.The senior arbitration counsel was especially familiar with legal-AI and simulation-tooling artifacts, so the three evaluators were not treated as interchangeable.

I.5. Pairwise Plausibility Results … J.2. T1: Evasive Pin-Down Test

Two practicing attorneys often judged WITNESSSIM testimony as at least approximately as plausible as authentic testimony, although a model-familiar evaluator readily distinguished it. Pedagogical tests showed strong question-form sensitivity and an Evasive witness that typically shifted from resistance to substantive responsiveness under explicit pin-down.

  • I.5. Pairwise Plausibility Results: Practicing attorneys selected WITNESSSIM as more plausible in 30.0% and 32.7% of exchanges, versus 20.0% and 28.6% for authentic testimony, with frequent ties.The senior arbitrator instead selected the authentic sequence in 90.0% of exchanges and never uniquely preferred WITNESSSIM.
  • I.6. Directional Preferences: Among directional decisions, WITNESSSIM was selected in 60.0% of the associate’s decisions and 53.3% of the senior litigator’s decisions.Directional-preference analysis excluded ties and neither judgments and was descriptive rather than a separate hypothesis test.
  • I.7. Interpretation: The plausibility results support a meaningful threshold of realism, not universal indistinguishability: two practicing attorneys accepted WITNESSSIM as plausible, while a model-familiar evaluator distinguished it.The evaluators’ differing judgments do not support causal explanations because evaluator backgrounds and stimulus exposure differed.
  • J. Pedagogical Evaluation: Detailed Protocols and Results: The pedagogical evaluation used four legal-training tests across 300 case-grounded contexts, targeting question form, evasive pin-down, runaway-witness control, and hostile-witness control.Outcomes distinguished challenge elicitation, conditional control, full intended trajectory, protocol pass rate, and test-specific behavioral measures.
  • J.1. U1: Question Form Sensitivity: In U1, open questions produced longer responses than closed questions in 289 of 300 contexts (96.3%), while question form had a large overall response-length effect, Friedman χ2(4) = 832.74, p < .001, Kendall’s W = .694.Open questions elicited the longest responses, followed by clarifying questions; closed, leading, and high-pressure questions elicited shorter answers.
  • J.1. U1: Question Form Sensitivity: U1 also showed question-form differences in hedging, but human validation was weak for tone and pushback, with weighted Cohen’s κ = .294 and κ = .298.The primary U1 evidence therefore rested on deterministic response-length analysis, while judge-derived fields were descriptive.
  • J.2. T1: Evasive Pin-Down Test: In T1, the evasive challenge appeared in 298 of 300 contexts (99.3%), 281 of 300 witnesses (93.7%) became substantively responsive, and 279 contexts achieved the full intended trajectory and 93.0% pass rate.The protocol and full-trajectory rates were identical because both required early evasion followed by final-turn resolution.
  • J.2. T1: Evasive Pin-Down Test: Resistance fell from 97.3% under the first narrowing attempt to 37.3% under the final explicit demand, while 33.1% of substantively responsive witnesses still hedged, qualified, or quibbled.The principal behavioral change occurred under explicit pin-down, and successful responses commonly remained recognizably evasive rather than becoming generically cooperative.

J.3. T2: Runaway Witness Control … K.13. Story-Generation Topics and Prompt

The pedagogical tests showed that WitnessSim responded to questioning interventions while preserving assigned behavioral challenges, whereas trajectory analyses found above-chance but limited alignment with real depositions. Synthetic trajectories also exhibited compressed magnitude and broader distributional separation from real testimony.

  • J.3. T2: Runaway Witness Control: T2 passed in 299 of 300 contexts (99.7%, 95% CI [98.1, 99.9]), with 178 contexts (59.3%, simultaneous 95% CI [52.4, 65.9]) reaching focused responses after initially unfocused behavior.Only one context (0.3%, [0.0, 2.5]) never resolved, while 121 contexts (40.3%, [33.8, 47.2]) were focused from the outset.
  • J.3. T2: Runaway Witness Control: Under T2’s final explicit demand, 239 of 300 witnesses provided a direct answer (79.7%, Wilson 95% CI [74.8, 83.8]), while many remained verbose but focused from the outset.The principal limitation was failure to elicit the targeted runaway challenge, not persistent failure to regain focus once it appeared.
  • J.4. T3: Hostile Witness Control Test: T3 achieved a 100.0% protocol-pass rate across 300 interactions, while substantive control under one-fact narrowing occurred in 293 of 300 contexts (97.7%, 95% CI [95.3, 98.9]).Among controlled interactions, 285 (97.3%, 95% CI [94.7, 98.6]) retained hostile or resistant presentation.
  • J.5. Summary of Pedagogical Findings: Across four pedagogical tests, WITNESSSIM responded systematically to legally meaningful questioning changes while preserving archetype-associated challenges, without establishing improved attorney performance.The evaluations measured simulated interaction behavior rather than attorney learning or skill transfer.
  • K.3. Neutral-Space Denoising: Denoising improved held-out correct-emotion ranking from the 12/23 chance expectation to 3.9/23, while 20-bin normalization provided comparability at the cost of coarse longitudinal resolution.The directions discriminate the constructed emotion concepts but are not established as ground-truth measures of human emotion.
  • K.6. Mean Real–Synthetic Arc Similarity: The observed mean trajectory correlation was r = 0.1477 and exceeded all 1,000 permuted correlations, yielding empirical p < .001, but its modest magnitude indicated limited shared temporal structure.The permutation preserved marginal score distributions while destroying temporal ordering.

K.14. Interpretive Scope and Limitations

The study’s interpretive scope is limited by exploratory emotion proxies, inherited representation choices, temporal normalization, and a small, single-litigation corpus. Accordingly, the trajectory findings are a proof of concept rather than a general estimate of real–synthetic behavioral fidelity across legal domains.

  • Measurement limitations: Emotion-vector analysis is exploratory and proxies longitudinal affective and interpersonal structure rather than measuring witness emotion directly.The directions derive from language-model representations of generated narratives, not ground-truth psychological labels on deposition testimony.
  • Methodological limitations: The extraction layer was inherited from prior methodology, and Layer 21 may not be optimal across model architectures or domains.Layer 21 was selected as a mid-to-late residual-stream representation rather than optimized specifically for this deposition corpus.
  • Methodological limitations: Temporal normalization to 20 bins enables comparison across examinations of different lengths but removes fine-grained timing information.It also treats equivalent relative positions as comparable even when the underlying questioning structures differ.
  • Corpus limitations: The real corpus contains only 55 transcripts from ten witnesses in a single multidistrict litigation, limiting generalizability across legal domains.The trajectory results therefore provide a proof of concept for longitudinal behavioral evaluation rather than a general fidelity estimate.
Loading 2608.13712v1…