Source-linked AI summary
Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?
Nirav Patel, Emily Wenger, Christopher Buccafusco
TL;DR
Legal reasonableness is vague, context-dependent, and potentially variable across demographic groups, raising questions about whether AI can reproduce human legal judgments. The paper compares 500 human participants with twenty-six LLMs across twenty-five scenarios and finds broad alignment, alongside greater model homogeneity and systematic demographic and institutional differences. The authors therefore treat these findings as preliminary and call for more systematic research.
Problem
The paper addresses limited evidence about whether AI models track human judgments on vague legal reasonableness standards and whether differences are systematic.
Method
The study compares responses from 500 humans and twenty-six LLMs across twenty-five legally relevant reasonableness questions using distributional analyses.
Results
LLM responses generally track human judgments, but models are more homogeneous, more favorable to government and corporations, and more aligned with white, male, older, and more educated respondents.
Takeaways & Limitations
LLMs can approximate human judgments on open-ended legal reasonableness tasks, but their systematic response patterns warrant attention in legal and policy settings.
Takeaways & Limitations
The study uses simple low-context prompts, sacrifices nuance needed to evaluate legal reasoning, and relies on limited demographic and model coverage for preliminary findings.
Abstract
from arXiv · showhide
As people increasingly rely on artificial intelligence (AI) for guidance in their own lives, scholars, lawyers, and even judges have begun to consider the role of AI in legal decision-making. As "silicon sampling" -- the use of generative AI models in social science research -- is now impacting academia, "silicon jurors" could make an appearance in courtrooms. This study joins an emerging line of research on generative AI models' ability to simulate human legal judgments. In particular, we study how large language model (LLM)-powered chatbots respond to series of questions about legal reasonableness. When the law needs to judge the appropriateness of a behavior, it most often asks whether the behavior was "reasonable." Yet despite the ubiquity of reasonableness judgments, they are the site of constant vexation for lawyers, judges, and lay people. Reasonableness seems inherently vague and unpredictable, since it relies on variable context and implicit conceptual schemas. Moreover, many scholars caution that reasonableness judgments may vary along demographic lines. We compare the answers of human participants to those of twenty-six LLMs across twenty-five different legally relevant reasonableness judgments. Overall, our findings suggest that chatbot responses generally track those of human participants. Nonetheless, we find some suggestive -- and potentially concerning -- results. Compared to humans, LLMs generate more homogeneous responses and occasionally treat a variable standard as an invariant rule. And, compared to humans, LLMs tend to generate answers that are more favorable to the government and to corporations. Finally, our results indicate that LLMs' responses tend to align more closely with those of respondents who are white, male, older, and more educated. More systematic research is needed to confirm or reject these initial findings.
1 Introduction
The paper studies whether LLM chatbots can simulate human legal judgments about reasonableness, a vague and potentially demographically variable legal standard. It compares twenty-six models with human participants and finds broad alignment alongside systematic differences.
- Research focus: The study examines whether AI chatbots can simulate human legal judgments about reasonableness.Reasonableness is widely used across contracts, government actions, and torts.
- Why reasonableness matters: Reasonableness is a vague legal standard requiring judgment rather than identifying an explicit attribute.Scholars also argue that such judgments may vary along demographic lines.
- Study contribution: Twenty-six LLMs are compared with 500 human participants answering the same legally relevant scenarios.The models span Meta, Google, Anthropic, OpenAI, DeepSeek, and xAI.
- Main findings: LLM responses generally track human notions of reasonableness, often remaining within the empirical range of human responses.Their distributions are statistically distinguishable from human distributions in some analyses.
- Main findings: Compared with humans, models produce more homogeneous answers and responses more favorable to government and corporations.Their answers also align more closely with white, male, older, and more educated respondents.
2 Why and How Reasonableness Matters to Law
Legal reasonableness is a context-dependent standard used across doctrines, requiring normative judgment rather than a fixed rule. Its apparent neutrality may nonetheless conceal variation and demographic bias.
- Legal role: Reasonableness is pervasive across legal doctrines, including negligence, contracts, and government action.The law uses it to evaluate conduct without requiring extraordinary or minimal care.
- Standards versus rules: Unlike rules, reasonableness standards require judgment about appropriate conduct under particular circumstances.A fixed deadline or numerical threshold specifies compliance, whereas a standard permits contextual evaluation.
- Normative judgment: Reasonable care is a normative judgment, not simply a count of how typical behavior is.Evaluators may consider either average conduct or what conduct should be expected.
- Scope of the concept: Some legal uses of reasonableness do not closely reflect ordinary citizens’ empirical views.The passage identifies antitrust law’s prohibition on unreasonable restraints as an example.
- Variation and bias: Laypeople’s reasonableness judgments matter because the standard may project neutrality while varying across people and contexts.Prior work has examined whether judgments fall between average and ideal quantities.
3 Why Study AI Reasonableness Judgements?
The paper asks whether AI models match human judgments about an uncertain legal standard and whether divergences are systematic. This question is important because LLMs may be used in legal decision-making while also exhibiting response homogeneity.
- Research question: The central research question is whether AI responses match human judgments about legal reasonableness and diverge systematically when they do not.The issue matters for legal decision-making and research on variation between AI models.
- Potential significance: AI systems that generally mimic human judgments could potentially supplement legal decision-making and reduce litigation time and expense.The passage frames this as a motivation for studying AI’s capacity to track human judgment.
- Why the task is difficult: Reasonableness judgments cannot be obtained by searching legal documents because they rely on schemas, norms, and empirical criteria.The relevant inputs and their relative weights are often difficult to identify.
- Model homogeneity: Prior research reports an Artificial Hivemind effect in which models converge on similar responses to open-ended prompts.The reported homogeneity occurs both within models and, more strongly, across models.
- Human–model contrast: Human judgments remain heterogeneous because life histories, values, cultural backgrounds, and contextual priors introduce variation.The paper contrasts this diversity with models’ shared training and optimization pressures.
4 Study Design
The study compares human and LLM responses to twenty-five low-context legal reasonableness questions, using repeated model sampling and complementary distributional tests. It evaluates similarity in central tendency, spread, overall distribution, and directional differences.
- Design: The study covers twenty-five legally relevant scenarios across torts, contracts, immigration, employment, criminal procedure, and family law.Responses come from 500 humans and twenty-six LLMs.
- Prompt design: The researchers intentionally use simple, low-context prompts requesting numerical estimates rather than sophisticated legal reasoning scenarios.This choice aligns with prior empirical work on legal reasonableness and resembles questions laypeople may answer.
- Human survey: Human participants were randomly assigned to provide either reasonable or ideal quantities using standardized instructions.The ideal condition tests whether model responses move toward or away from human ideals.
- LLM sampling: Each model–question pair was sampled twenty times through a stateless API pipeline with constant prompting parameters.Models received the same context as humans and responses were post-processed consistently.
- Research questions: The study asks whether models are human-like at the question level and analyzes distributions rather than legal correctness.The design is motivated by the absence of uniquely correct numerical answers.
- Analysis: The analysis compares human and model response distributions using Brown–Forsythe, Kolmogorov–Smirnov, and Brunner–Munzel tests.These tests target dispersion, overall distributional equality, and directional stochastic dominance, respectively.
5 Key Findings
LLMs generally track human reasonableness judgments, but their responses are more homogeneous and show systematic pro-government, pro-corporation, and demographic alignment patterns.
- Human-model alignment: LLMs generally track human reasonableness judgments, although their response distributions are often statistically distinguishable from humans.The detected differences reflect systematic distributional shifts rather than artifacts of a few extreme outputs.
- Response homogeneity: Human responses vary far more than tightly clustered model outputs on Q7, with Brown–Forsythe W=453.42.The result indicates a substantial difference in within-group disagreement for profits-on-safety judgments.
- Human-model alignment: 20 of 25 questions show significant Brown–Forsythe differences, 23 show significant Kolmogorov–Smirnov differences, and 19 show significant Brunner–Munzel differences.Sixteen questions are significant under all three tests.
- Institutional alignment: Models tend to favor government and corporate interests, allowing longer detentions and restrictions while supporting less corporate spending on safety.Examples include immigrant detention of 10 versus 90 days and non-compete restrictions of 6 versus 12 months for human and LLM medians, respectively.
- Demographic alignment: Across the 25-question response space, models align more closely with men, white respondents, older respondents, and more educated respondents than with complementary groups.This pattern appears in both median-based and mean-based cosine-similarity analyses, though the absolute shifts are modest.
- Scope and future research: The study recommends broader models and larger, more demographically diverse human samples to test whether these alignment patterns generalize.Future work should also examine whether models preserve, amplify, or diminish demographic differences in human judgments.
6 Discussion and Implications
The discussion concludes that AI models broadly approximate human judgments about vague legal reasonableness standards, while exhibiting homogeneity and systematic demographic and institutional alignment differences.
- Overall performance: AI models broadly approximate human responses across legally relevant reasonableness judgments, despite statistically distinguishable distributions.Across the survey, model central tendencies and overall response distributions generally matched human responses.
- Model homogeneity: Models sometimes treat variable reasonableness standards like rules, producing more homogeneous answers than humans.For landlord notice, models almost uniformly answered 24 hours; similar concentration appeared in several other scenarios.
- Demographic comparisons: Models align more closely with male, white, older, and more educated respondents than with complementary demographic groups.The pooled analysis reports statistically detectable alignment differences across gender, ethnicity, age, and education.
- Implications and future research: The demographic and institutional findings are preliminary and warrant further research, including broader human samples and targeted questions.The authors plan to expand demographic coverage and add questions addressing pro-government and pro-corporate tendencies.
7 Limitations
The study’s limitations include reduced legal nuance and insufficiently targeted sampling and questioning for evaluating bias.
- Scope and design: The study sacrifices meaningful legal nuance because it focuses on models’ central tendencies and response distributions.Its questions were not designed specifically to detect pro-government or pro-corporate responses.
Appendix
The appendix includes a table describing demographic distributions.
- Demographic distributions: Table 4 presents the study’s demographic distributions.
Question Level Demographic and Model Family Regression Analysis
The analysis first screens demographic differences across 25 questions, then tests model alignment only for statistically significant cases using corrected regression analyses. Meaningful alignment effects appear for Attorneys’ Fees and Interest Rate, but not Permit Price.
- Regression specification: The regressions treat each AI prediction as a model-family baseline and use binary demographic indicators to estimate subgroup deviations.The resulting coefficients are evaluated through squared deviations from the model baseline.
- Screening for demographic differences: Across 25 questions, significant human subgroup differences emerge only for Attorneys’ Fees, Permit Price, and Interest Rate.The Brunner–Munzel screening identifies these three cases before the regression analyses proceed.
- Model alignment results: For Attorneys’ Fees, all model families predict values significantly closer to the higher-education group’s responses.This is the only reported alignment result for Q10, and the analysis applies Bonferroni correction across tests.
- Model alignment results: For Interest Rate, DeepSeek, Gemini, Llama, and Claude show significant alignment effects, whereas Permit Price shows no reliable gender alignment.The Permit Price result contrasts with the significant human gender difference detected during screening.
Ethical Statement
The human participant study received IRB approval and informed consent, with anonymized data stored securely. The LLM survey used benign responses and did not involve sensitive data.
- Ethical safeguards: The study obtained IRB approval, collected clearly written consent, anonymized participant data, and stored it on secure servers.The authors characterize other ethical risks as minimal because the LLM survey involved no sensitive data and elicited benign responses.