Source-linked AI summary

Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards?

Lorenzo Molfetta, Alessio Cocchieri, Luca Ragazzi, Ilaria Bartolini, Marco Patella, Gianluca Moro

arXiv:2608.21409v1cs.CYcs.AIcs.CL

TL;DR

LLM performance on medical and legal examinations motivates claims of generalized expert reasoning, but legal validity depends on jurisdiction, time, and authoritative relationships. The paper compares both domains through four diagnostic axes and LEGAL-LINK-EU, finding that legal models are more vulnerable to misleading authority and superficial format cues than medical models. The study concludes that legal LLMs can over-trust authoritative but false information when it conflicts with internal knowledge.

  • Problem

    Existing legal benchmarks often treat applicable law as identified and stable, leaving temporal validity, jurisdiction, and normative relationships insufficiently evaluated.

  • Method

    The paper compares legal and medical LLM reasoning across Knowledge Recall, Grounding, Confidence, and Format Perturbation using LEGAL-LINK-EU.

  • Results

    Legal LLMs show sharper citation sycophancy and structural fragility than medical models, while misleading citations can displace correct internal beliefs.

  • Takeaways & Limitations

    Legal LLM evaluation should account for whether authorities are applicable, temporally valid, and normatively related rather than relying on aggregate accuracy alone.

  • Takeaways & Limitations

    The study uses MCQA, zero-shot prompting, and oracle contexts, excluding open-ended argumentation, few-shot strategies, and end-to-end RAG pipelines.

Abstract

from arXiv · show

In medicine, claims remain valid when supported by empirical evidence grounded in stable biological reality. In law, by contrast, truth is contingent, defined by jurisdiction, temporal validity, and the hierarchy of authoritative sources. The recent success of large language models (LLMs) on medical licensing examinations has encouraged an expectation of comparable legal competence. This analogy, however, obscures a critical distinction between domains. Unlike in medicine, legal performance often depends less on inference than on determining when external authority is applicable, valid, and non-contradictory. We introduce a comparative diagnostic framework evaluating legal reasoning against medical baselines along four axes (knowledge recall, grounding, confidence, and robustness), uncovering a sharp domain asymmetry when applied to a new benchmark that encodes temporal validity and normative relationships. While medical LLMs reliably benefit from verified sources, legal LLMs struggle to assess when retrieved citations are useful or misleading, exhibiting overconfidence in perturbed contexts and sensitivity to superficial formatting cues. Increased model scale amplifies this tendency, revealing that stronger instruction following can coincide with weaker resistance to authoritative perturbations. These findings show that LLMs treat law as unstructured text rather than binding precedent, while revealing a tendency to over-trust authoritative but false information when external references conflict with a model's internal knowledge.

1 Introduction

LLM success on medical and legal exams can suggest generalized expert reasoning, but law differs because legal knowledge varies by jurisdiction and time. The paper introduces a four-axis framework and LEGAL-LINK-EU to diagnose failures that aggregate accuracy can hide.

  • Medical and legal examination success has encouraged viewing expert reasoning as a generalized capability across domains.
  • Legal knowledge varies by jurisdiction and over time, unlike medical knowledge grounded in a mostly stable physical reality.A statute applicable in 2021 may later become irrelevant after a binding ruling or legal change.
  • Existing legal benchmarks emphasize static comprehension and MCQA despite legal reasoning requiring retrieval grounded in jurisdiction, temporal scope, and citation fidelity.
  • The framework evaluates Knowledge Recall, Knowledge Grounding, Knowledge Confidence, and Format Perturbation to expose domain-specific fragilities.These axes test parametric recall, use of authoritative context, susceptibility to misleading citations, and reliance on exam-style artifacts.
  • LEGAL-LINK-EU uses EUR-Lex relationships to test legal validity across time and hierarchy rather than simple text overlap.
  • Legal LLMs show greater citation sycophancy and structural fragility, overtrusting manipulated references and exploiting formatting cues.

2 Related Work

Prior legal benchmarks commonly treat applicable law as identified and stable, while research on sycophancy and MCQA shows that superficial cues can overstate model robustness. Retrieval-focused benchmarks improve grounding measurement but generally omit whether authorities remain legally valid.

  • Scenario-grounded legal benchmarks provide fixed passages for classification, extraction, reasoning, summaries, or answers.
  • LegalBench-RAG measures retrieval precision and grounding quality but assumes the underlying corpus is normatively stable.
  • Static legal evaluations typically omit determining which provisions apply, when they entered force, and whether they remain valid.
  • RLHF-associated sycophancy can prioritize agreement or perceived helpfulness over factual truthfulness, a vulnerability termed context-memory conflict.
  • Legal systems can misground authentic case law to support fabricated holdings and disproportionately retrieve high-frequency precedents.
  • MCQA accuracy can rely on option ordering, symbol binding, positional biases, or options-only artifacts rather than question comprehension.

3 Method

The method compares legal and medical reasoning across four dimensions, using LEGAL-LINK-EU to isolate temporal and hierarchical effects in EU law. It also tests responses to misleading contexts and surface-format changes.

  • The framework compares legal performance with matched medical baselines to determine whether vulnerabilities concentrate in law or generalize across authoritative evidence.
  • LEGAL-LINK-EU tests legal effects induced by relationships between EU normative acts rather than direct document comprehension.
  • The benchmark derives instances from EUR-Lex document pairs linked by seven legally operative relationship types.These relationships encode normative consequences across time and hierarchical levels.
  • GEPA jointly optimizes generated legal MCQA items for normative relevance, legal soundness, and distractor quality.
  • 1,127 instances span 880 document pairs, seven relationship types, and EUR-Lex materials from 1953–2025.
  • Knowledge Recall uses context-free questions and matched MMLU subjects, supplemented by MedQA and L2-EU for independent domain retention.
  • Knowledge Grounding supplies authoritative context, while Knowledge Confidence tests whether models reject perturbed external information.
  • Legal perturbations alter dates, jurisdictional scope, normative relations, or contextual framing, while Format Perturbation varies labels and answer positions.

4 Experimental Setup

The experiments combine MCQA accuracy with normalized diagnostic metrics designed to measure grounding, resistance to adversarial context, citation sycophancy, and option-artifact reliance across diverse LLMs.

  • The evaluation uses proprietary and open-weight LLMs with separate generation, judging, and evaluation roles to reduce contamination.
  • Accuracy on MCQA tests is combined with complementary diagnostic metrics, all normalized to [0,1] for cross-domain comparison.
  • Grounding Inefficiency Index measures failure to benefit from authoritative retrieval over parametric recall, while Parametric Override Index measures adversarial displacement of internal knowledge.
  • Citation Sycophancy Index measures over-deference to adversarial context relative to authoritative grounding, and Artifact Exploitation Index measures reliance on option-level patterns.
  • Artifact Exploitation Index uses “Select Incorrect” and “None Provided” accuracies to detect preference for plausible options over recognizing that none is correct.
  • Models are accessed through official APIs with 16K-token maximum outputs and temperature 1.0, while open models use four RTX 3090 GPUs and vLLM.

5 Results and Analysis

The results reveal a sharp legal–medical asymmetry: legal models depend more on authoritative grounding yet are more vulnerable when that authority is misleading or structurally perturbed. Larger reasoning models and legal format sensitivity further expose weaknesses in temporal validation, citation skepticism, and semantic comprehension.

  • 5.1 Domain Asymmetry in Grounding: Medical models gain only marginally from authoritative context, whereas legal models show much larger grounding gains despite lower knowledge-recall accuracy.Gemini-2.5-Flash gains 2.8 percentage points on MedQA versus 27.0 percentage points on L2-EU; GPT-OSS 120B gains 2.3 versus 29.9 percentage points.
  • 5.1 Domain Asymmetry in Grounding: Legal models struggle with normative relationships, with complex temporal dependencies such as implicitly repeals causing severe performance drops.Llama-3.1 falls to 62.1% on the cited relationship condition, indicating difficulty determining whether retrieved provisions remain in force.
  • 5.2 Sycophancy and Citation Bias: As perturbation density rises from 20% to 100%, GPT-OSS 20B accuracy declines more steeply on L2-EU than MedQA.L2-EU accuracy decreases from 35.7% to 14.8%, while MedQA decreases from 65.2% to 50.5%, based on three independent runs with 95% confidence intervals.
  • 5.3 Structural Fragility and Heuristics: Legal models often perform better with options only than with no provided context, revealing reliance on option-level and distributional cues rather than validated legal comprehension.Medical models preserve the hierarchy Standard > None-Provided > Options-Only, whereas legal models show the opposite inversion between Options-Only and None-Provided settings.
  • 5.4 Scaling Laws and Model Profiles: Instruction-tuned models can outperform larger reasoning models under perturbation, while reasoning-oriented architectures more readily rationalize supplied context instead of questioning its validity.For example, legal POI is 78.2% for Llama-3.1 versus 43.1% for GPT-OSS 120B, and CSI is 46.9% for Llama-3.1 versus 13.4% for GPT-OSS 20B.

6 Conclusion

The study evaluates reference calibration in knowledge-intensive QA by combining internal knowledge with external evidence under perturbations. It finds a domain asymmetry: verified context helps or preserves medical performance, while misleading legal citations can override correct beliefs and scale increases this fragility.

  • The evaluation combines internal knowledge with external evidence under different perturbation scenarios.
  • In medicine, verified context preserves or mildly improves strong baselines.
  • In law, reliable context can repair legal QA, but misleading citations can pull models away from correct internal beliefs.
  • Larger models more readily rationalize false authority, making legal fragility scale-sensitive.
  • Retrieval makes the problem operational because useful evidence and over-trust in formal or institutionally styled text arise through the same mechanism.
  • The study recommends measuring whether LLMs can reject false authoritative references, not only whether context helps.

Limitations

The framework has scope constraints: its multiple-choice design abstracts away open-ended legal argumentation, while zero-shot oracle-context evaluation excludes few-shot and end-to-end retrieval settings. The authors also call for testing other legal traditions, languages, and tasks.

  • Multiple-choice formats enable scalable cross-domain comparison but abstract away open-ended legal argumentation.
  • The benchmark proxies doctrinal recall rather than full drafting capability.
  • Zero-shot prompting and oracle contexts isolate intrinsic model sycophancy but exclude few-shot strategies and end-to-end RAG pipelines.
  • Future research should test whether citation biases persist across diverse legal traditions and non-English jurisdictions.
  • The evaluation should broaden to additional tasks, including entity extraction.

A Dataset Documentation and Validation

LEGAL-LINK-EU contains 1127 balanced multiple-choice instances built from EUR-Lex document relationships. Its construction controls answer labels, relation types, context lengths, and document-pair coverage.

  • 1127 MCQ instances each contain exactly one correct answer and three distractors.
  • The benchmark is balanced across answer labels and seven EUR-Lex relation types.
  • Correct-answer positions remain near-uniform: A 22.8%, B 26.6%, C 24.3%, and D 26.3%.
  • Questions and answer options are compact, while original and perturbed contexts preserve comparable long-document legal evidence.
  • The average number of questions per document pair remains close to one, limiting dominance by repeated variants of a small set of legal acts.

A.2 LLM-as-a-Jury Protocol

The validation protocol uses independent jurors to assess sampled benchmark items across legal relation types and reasoning dimensions. The audit finds that items require multi-hop reasoning, especially temporal and relational inference, while perturbations can make legally plausible distractors appear supported.

  • A.2 LLM-as-a-Jury Protocol: Three independent jurors evaluate a stratified 100-item sample using questions, answer options, and paired EUR-Lex contexts.
  • A.2 LLM-as-a-Jury Protocol: The jury records best answers, legal coherence, distractor plausibility, and reasoning complexity using a defined rubric.
  • A.2 LLM-as-a-Jury Protocol: All audited questions require multi-hop reasoning or above, with most requiring temporal or relational reasoning over document-pair effects.
  • A.2 LLM-as-a-Jury Protocol: Implicit supersession relations receive the highest complexity scores because they depend on non-explicit temporal and relational effects.
  • A.2 LLM-as-a-Jury Protocol: The worked repeal item requires recognizing that saithe quotas are repealed while herring remains governed by Regulation (EEC) No 198/83.
  • A.2 LLM-as-a-Jury Protocol: Context perturbation suppresses the herring carve-out and makes distractor B appear textually justified, producing the described KR, KG, and KC diagnostic pattern.
  • A.2 LLM-as-a-Jury Protocol: The benchmark covers seven normative interactions, including repeals, corrections, implicit repeal, validity extensions, application extensions, completion, and obsolescence.

B Perturbation Matching Across Domains

Legal and medical context perturbations have different structural profiles: legal perturbations preserve more vocabulary but reorganize passages more aggressively, especially for implicit supersession and repeal.

  • Legal perturbations preserve more vocabulary while producing stronger passage reorganization than medical perturbations.Medical perturbations rely more on localized substitutions.
  • Relations involving implicit supersession and repeal show the lowest sequence similarity among legal perturbations.These relations therefore induce the most substantial structural rearrangement.

C Prompts

The appendix documents prompts for recall, grounding, confidence, format perturbations, adversarial context generation, reasoning-budget parity, and GEPA-based prompt optimization.

  • C Prompts: The prompt appendix organizes templates by evaluation, format perturbation, context perturbation, and dataset-generation functions.The evaluation functions are KR, KG, and KC.
  • C.1 Evaluation Prompts: KR uses a zero-shot multiple-choice prompt without external context.The prompt is shown in Figure 4.
  • C.1 Evaluation Prompts: KG and KC use the same context-augmented prompt, warning that contexts may be incomplete or incorrect.KG uses authoritative context, whereas KC uses perturbed context, enabling comparison under valid and adversarial retrieval.
  • C.2 Format Perturbation Prompts: Format perturbations include Select Incorrect, which inverts the task, and None Provided, which replaces the correct option with an explicit none-of-the-options choice.These variants test sensitivity to task and answer-format changes.
  • C.3 Context Perturbation Generation: Adversarial contexts are generated with subtle, implicit misleading signals designed to require multi-step reasoning for detection.Legal strategies alter temporal scope, jurisdictional scope, normative relations, or contextual framing.
  • C.3 Context Perturbation Generation: Medical perturbations mirror the legal structure while changing diagnostic, therapeutic, mechanistic, and contextual content.The domain-specific strategies alter symptoms, patient characteristics, pathophysiology, or history details.
  • D Thinking Budget: Reasoning effort is standardized by evaluating OpenAI models at low effort and assigning Gemini-2.5-Flash a 1024-token thinking budget.The configuration is treated as functionally equivalent across model families.

E.1 Optimization Configuration

The optimization configuration separates generation from evaluation, samples stratified EUR-Lex document pairs, scores MCQs with a weighted rubric, and applies hard validation constraints within a GEPA pipeline.

  • gemini-3-flash-preview generates MCQs while gpt-5-mini evaluates them, separating generation from quality assessment.This separation is intended to prevent contamination between the two roles.
  • The data use 150 document pairs with stratification across seven relationship types and 30 full evaluation loops.The pairs are divided into training and validation examples as specified.
  • The judge scores six MCQ dimensions on 1–5 scales, including multi-provision synthesis, relationship use, novel scenarios, distractor quality, legal specificity, and generic-reference avoidance.Each dimension has an assigned rubric weight.
  • The final score is a weighted sum of normalized rubric scores, while any generic document reference triggers a score of 0.The generic-reference rule is a hard constraint independent of other rubric values.
  • The generation module wraps a DSPy signature in ChainOfThought and takes paired documents, identifiers, and a relation as inputs.The pipeline returns a question, one correct answer, and three distractors.
  • The optimized prompt requires source-dependent questions with specific legal identifiers and substantive legal effects rather than procedural metadata.Figures 11–14 contain the complete prompt.
  • Algorithm 1 initializes the generation module, rubric weights, and judge before iteratively optimizing an output module π* under evaluation limits.The pseudocode defines the generation and judging functions used in the pipeline.
Loading 2608.21409v1…