Source-linked AI summary
When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors
Cho-Ying Wu
TL;DR
The paper asks how LLM-simulated jurors respond to defendant statements and whether ideology or background shapes common-law jury judgments. It introduces JuryBench and evaluates 20 frontier LLMs across controlled criminal-case simulations. The results show emotionally persuasive statements can backfire, while background fit and ideology strongly shape verdict severity.
Problem
LLM behavior as common-law jurors, especially their responses to emotional defendant statements and background- or ideology-linked bias, remains unexplored.
Method
The study introduces JuryBench and evaluates 20 frontier LLMs using controlled criminal cases, defendant backgrounds, courtroom statements, juror ideologies, and severity scoring.
Results
Emotional persuasion can be detrimental, while background fit and juror ideology strongly shape verdict severity; jurors are generally harsher toward opposite-background defendants and more lenient toward same-background defendants.
Takeaways & Limitations
JuryBench provides a controlled testbed for mock-jury analysis and highlights both the promise and risks of modeling jury reasoning with LLMs.
Takeaways & Limitations
The simulation simplifies real trials by combining stages, focusing testimony on defendant speech, and using text without courtroom body language or facial expressions.
Abstract
from arXiv · showhide
LLMs have been used to simulate human decision-making in professional settings, yet their behaviors in common-law jury trials remain unexplored. We study when and how a defendant's courtroom statement affects LLM-simulated jurors, focusing on persuasion, ideological bias, and background-based affinity. To support the analysis, we introduce JuryBench, a benchmark containing controversial criminal cases in U.S. criminal law. In each case, a defendant can claim various plausible justifications to support acquittal or reduced liability. We fix the base case and design defendants of different backgrounds, who give courtroom statements with varying emotional appeal or rebuttal. Jurors with diverse ideological profiles across the spectrum are simulated. We examine 20 frontier LLMs, resulting in a total of 432K decisions and rationales, and quantify changes in verdict severity. Our findings show that LLM-jury simulation echoes many human-jury findings. First, emotional persuasion can be detrimental, since jurors may perceive it as evidence of guilt or inconsistency. Next, we show that background fit between jurors and defendants is a stronger and significant factor than other isolated factors, and that jurors are in general harsher toward opposite-background defendants and lenient toward same-background ones. Finally, we find that juror ideology also strongly shapes severity judgments. These findings highlight both the promise and risks of using LLMs to model jury reasoning and call for careful evaluation. The data and code are available at https://github.com/choyingw/JuryBench
1 Introduction
This paper studies how LLM-simulated jurors respond to defendant statements, emotional persuasion, ideology, and background-based bias in common-law trials. It introduces JuryBench and evaluates 20 frontier LLMs across 500 controversial criminal cases, producing 432K decisions and rationales.
- Legal setting: Common-law juries evaluate evidence, credibility, justification, and verdicts using commonsense, life experience, moral conviction, and ideology.Unlike civil-law proceedings, common-law trials emphasize jury storytelling and emotional persuasion.
- Research gap: The study addresses the unexplored behavior of LLMs acting as common-law jurors, including responses to emotional defendant statements and background-linked bias.It examines juror behavior, ideological profiles, and defendant–juror background relationships.
- Benchmark: JuryBench contains 500 controversial U.S. criminal cases in which defendants present plausible justifications for acquittal or reduced liability.Cases include around 400 potential charges and controlled defendant and juror backgrounds.
- Experimental design: The framework fixes case backgrounds and evidence while varying defendant profiles and courtroom statements to isolate their effects on verdicts.The cases include different defendant backgrounds and statements designed around emotional persuasion and justification.
- Evaluation: 20 frontier LLMs generate 432K juror decisions and reasons for analyzing simulated jury behavior.The evaluation varies defendant backgrounds, juror ideologies, and courtroom statements.
2 Related Work
Prior work studies legal prediction, legal reasoning, courtroom simulation, jury bias, and persona simulation, but does not provide this paper’s targeted common-law LLM-jury analysis. Existing courtroom simulations also rely on already-concluded judgment documents that restrict substantive defendant arguments.
- Legal AI: Legal AI research has addressed judgment prediction and legal language understanding but generally not simulated jury decisions.Prior benchmarks also omit some distinctions such as mens rea and attempted offenses.
- Courtroom simulation: AgentsCourt simulates courtroom debates and judgments from conclusive judgment documents, leaving defendants limited room for substantive arguments.Those documents are written after judges have determined facts, verdicts, and sentences.
- Jury bias: Legal-psychology research links jury bias to early impressions, defendant background affinity, outsider harshness, ideology, and occasional black sheep effects.Similar-background defendants may receive mercy, while value-violating insiders can sometimes be judged more harshly.
- Persona simulation: LLM persona simulation implicitly establishes internal personas that guide responses without necessarily explicitly revealing the simulated background.This work extends persona-style simulation into a juristic setting.
3 Simulation Framework
The simulation framework collaboratively constructs controversial criminal cases, defendant statements, and juror profiles, then evaluates verdicts under controlled combinations of backgrounds, statements, and ideologies. It converts charge outcomes into a severity scale for quantitative comparison.
- Data construction: The dataset uses controversial, debatable criminal cases with detailed defendant backgrounds, case scenarios, evidence, and courtroom speech.The materials are collaboratively crafted by LLMs and human experts because judgment documents are conclusive and omit trial uncertainty.
- Case generation: GPT-5.4 generates 500 cases across around 400 U.S. criminal charges, including stereotype-conforming and stereotype-subverting defendant backgrounds.Each entry contains defendant background, case background, and evidence.
- Statement generation: Defendant statements combine sympathy appeals, justifications, remorse, or rebuttals and contain around 15 sentences.Statements are generated separately for Group-A and Group-B defendants from the case materials and profiles.
- Human review: Human experts review case feasibility and statement conformity, rate emotional contagion and remorse from 1 to 5, and verify affinity scores from −5 to 5.Affinity scores indicate conservative or liberal background appeal and are checked by a sociology-trained expert.
- Decision simulation: Juror profiles are supplied as system prompts, while user prompts request guilty or not-guilty decisions across cases and statement conditions.For multiple crimes, the model outputs the most severe crime of which the defendant is guilty.
- Severity measurement: Severity assigns each potential charge a score from 0 to 15, with 0 representing acquittal or valid justifications and higher values representing greater liability.Scores support quantitative analysis of changes after defendant statements.
4 Analysis and Findings
The analysis finds that defendant statements often leave verdict severity unchanged, but their effects depend on emotional appeal, remorse, background fit, and juror ideology. Background fit is the strongest isolated factor, while conservative jurors generally assign greater severity and show stronger statement sensitivity.
- 4.1 How do statements affect jury decisions?: About 70%-80% NR indicates that LLM-simulated jurors usually do not change verdict severity after hearing defendant statements.This resembles reported human-jury decision inertia after initial case impressions.
- 4.1 How do statements affect jury decisions?: Defendant statements can either reduce or increase severity, as jurors interpret remorse as relatable mitigation, implicit guilt, or evidence of inconsistency.Some rationales describe inconsistent statements as weakening credibility, while others treat remorse as mitigating.
- 4.2 Which latent factors affect jury decisions?: Background fit is more significant than emotional contagion and remorse, with a substantial majority of models showing significant positive coefficients.Jurors sometimes relate to defendants with similar backgrounds, whereas opposite-background defendants can receive harsher judgments.
- 4.3 Do Group-A and Group-B differ in verdict severity?: Nearly all models show lower relative severity for same-direction cases and higher relative severity for opposite-direction cases after statements.Matched backgrounds are associated with greater leniency toward Group-A, while mismatched backgrounds can prompt stronger criticism.
- 4.3 Do Group-A and Group-B differ in verdict severity?: Opposite-direction cases often have larger |G|, indicating stronger post-statement movement in the signed group-severity gap.Across models, the direction of gap change is roughly balanced, suggesting model-specific tendencies to enlarge or close the gap.
- 4.4 How does ideology affect jury’s decisions?: Conservative jurors assign higher severity than neutral jurors, who are harsher than liberal jurors, and around 70% of models show larger statement effects for conservatives.Around 60% of models show the same direction of severity change across ideologies.
5 Conclusion
This work presents a systematic study of LLM-simulated jurors in common-law trials and introduces JuryBench for controlled analysis of their behavior. The findings echo legal-psychology patterns involving early impressions, defendant statements, ideology, and affinity.
- The study systematically examines LLM-jury behavior under the common-law setting.
- JuryBench provides controversial criminal cases, diverse juror profiles, controlled defendant backgrounds, and emotionally persuasive statements.
- The findings echo legal-psychology findings on sticky early impressions, double-edged defendant statements, and ideology- or affinity-driven decisions.
- JuryBench enables controlled study of LLM-jury behavior and supports mock-jury analysis for revising trial strategies.
Limitations
The study uses a simplified, text-based simulation focused on emotional persuasion rather than a full courtroom process. Its empirical scope is primarily U.S. criminal trials, and observed model behavior may change as frontier models are updated.
- The simulation combines case background and evidence into one stage and omits full testimony, cross-examination, and jury group discussion.
- The study primarily examines decision-making in U.S. common-law criminal trials.
- The courtroom process is simulated using text alone, excluding body language, facial expressions, and other real-courtroom factors.
- The study uses 20 frontier LLMs, whose relative behaviors may change with updates to training data, safety policies, and inference systems.
Ethical Considerations
The work is intended as a research tool for studying LLM-jury behavior and supporting legal professionals’ mock trials, not for automating real legal judgment. Its synthetic data contain sensitive attributes and intentionally introduced stereotypical associations that require careful handling.
- The system is intended to help legal professionals plan or refine mock trials and trial strategies using simulated jurors.
- The work is not intended to automate real legal judgment or interfere with real cases.
- Synthetic cases, defendants, and jurors involve sensitive social attributes and criminal allegations that should be handled carefully.
- The benchmark intentionally introduces stereotypical associations for bias analysis and will include documentation, use restrictions, and warnings.
A Human Evaluation
A human evaluation compares verdict choices before and after defendant statements across 25 cases and 12 lay participants. Humans, like the LLM jurors, often increased severity when statements suggested guilt or performativity, while justification and empathy could reduce severity.
- The evaluation sampled 25 cases and asked 12 U.S. lay participants without legal training to choose charges before and after hearing defendant statements.
- Human decisions leaned toward increased severity after statements, with remorse interpreted as evidence of guilt and some statements perceived as performative.
- Participants reduced severity when they recognized valid grounds for justification, sometimes combined with empathy produced by emotional contagion.
- In the human regression, emotional contagion and remorse coefficients were 0.173 and -0.168, respectively, with both p-values below 0.05 but above 0.01.
B Joint-Effect Regression
The full interaction model shows that defendant–juror affinity and coherent combinations of emotional appeal and remorse are associated with reduced post-statement severity, while isolated rhetorical cues can backfire. Compared with the additive model, it explains how these effects depend on one another.
- Joint effects: Affinity M, coherent emotional-remorseful statements pq, and their joint term pqM are associated with reduced post-statement severity.The positive pqM term indicates that matched affinity and jointly expressed emotionality and remorse reinforce leniency.
- Isolated cues: Isolated emotionality p and remorse q are mostly negative, potentially appearing performative, insufficient, or indicative of responsibility.Their positive effects are captured by joint terms such as pq and pqM.
- Affinity interactions: Negative pM and qM interactions show that affinity does not monotonically amplify isolated emotional or remorseful appeals.Under stronger affinity, one-dimensional appeals may be discounted as strategic or excessive, whereas the combined pq appeal remains mitigating.
- Model comparison: The unary model estimates marginal sensitivity, whereas the interaction model explains how emotionality, remorse, and affinity combine.The two models therefore provide complementary behavioral profiles for interpreting LLM-specific sensitivities.
- Overall pattern: The interaction analysis suggests that LLM jurors reward coherent persuasive configurations rather than isolated rhetorical signals.This interpretation accounts for the instability of emotionality and remorse as standalone predictors.
C Prompts in Use
The study documents the prompts used to generate case scenarios, juror profiles, and juror verdicts.
- Prompt set: Prompts are provided for generating case scenarios, juror profiles, and verdicts from individual jurors.These materials are documented in Figures 6, 7, and 8.
D Instructions for Legal Expert and Manual Check
The generation pipeline uses expert review to check legal plausibility and affinity-score quality before analysis.
- Legal review: A legal expert reviewed generated case scenarios and statements, with 67 cases flagged for regeneration and renewed gating.The review covered the quality of the generated cases and statements.
- Affinity review: A sociology-trained expert reviewed generated defendant-background affinity scores, correcting 51 scores.The corrections were part of the affinity-score gating process.
E Common-Law and Civil-Law Settings
The paper situates its study in the common-law jury setting, where lay jurors evaluate facts and persuasive courtroom narratives rather than primarily applying codified legal rules. The documented prompts and review instructions operationalize case generation, juror simulation, verdict prediction, and affinity assessment for this setting.
- Legal systems: Civil-law systems center codified statutes and professional legal interpretation, whereas common-law trials emphasize adversarial argumentation, precedent, and trial-based fact-finding.The contrast concerns both legal authority and courtroom decision-making structure.
- Jury role: U.S. common-law criminal juries are generally laypeople who evaluate evidence, witness credibility, defendant statements, and whether guilt is proven.Jurors rely on commonsense, life experience, moral judgment, credibility, and emotion in reaching verdicts.
- Statistical notation: Figure 5 marks interaction-model coefficient significance with *, **, and *** for p-values below 0.05, 0.01, and 0.001, respectively.The notation provides the significance thresholds used in the full interaction-model analysis.
- Courtroom persuasion: Common-law courtroom arguments translate legal claims into narratives intended to be understandable and convincing to ordinary citizens.Persuasion can draw on commonsense, moral judgment, credibility, and emotion.
- Case generation: The case-generation prompt requests highly controversial jury-deliberation scenarios involving doctrines such as mens rea, self-defense, and necessity.These scenarios are aligned with the study’s focus on criminal-case jury evaluation.
- Juror simulation: The juror prompt asks models to choose the specific crime committed and return one verdict label for each case.The instructions emphasize case facts, evidence, defendant speech, credibility, remorse, intent, and less severe offenses when appropriate.
- Affinity scoring: The affinity-review instructions define affinity as a signed score indicating whether a defendant background is expected to elicit stronger affinity from conservative ...The supplied passage ends before specifying the comparison group or full scoring definition.