Source-linked AI summary

Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases

Amogh Raina, Ilias Chalkidis, Daniel Hershcovich, Henrik Palmer Olsen

arXiv:2608.17168v1cs.CLcs.AI

TL;DR

Legal reasoning by LLMs in demanding case-forecasting tasks remains understudied. This study tests GPT-5.4 on ECtHR cases under alternative prompts and evaluations, finding structurally complete but substantively shallow reasoning, with accuracy not tracking reasoning quality.

  • Problem

    How LLMs reason in legally meaningful ways during ECtHR case forecasting remains substantially understudied despite reasoning’s relevance to explainability.

  • Method

    The study evaluates GPT-5.4 on ECtHR case forecasts across prompting strategies, assessing reasoning quality through law-student annotators and LLM judges.

  • Results

    GPT-5.4 reproduces doctrinal structure but remains substantively shallow; expert prompting improves comprehensiveness without improving accuracy, while LLM judges weakly align with human annotators.

  • Takeaways & Limitations

    Task accuracy should not proxy legal reasoning quality, and automated LLM evaluation should not replace human assessment.

  • Takeaways & Limitations

    Findings are limited to ECtHR cases involving one Convention article.

Abstract

from arXiv · show

Reasoning has become a standard technique and feature for contemporary LLMs; however, its application and quality in the context of demanding legal-oriented tasks, such as legal case forecasting, remain under explored. We investigate how LLMs reason in the context of legal case forecasting, using legal cases from the European Court of Human Rights (ECtHR) as a testbed. We evaluate OpenAI GPT 5.4, a recent top-tier LLM, by exploring alternative prompting strategies that are more or less suggestive of what counts as legally meaningful reasoning in the context of ECtHR jurisprudence. We present our findings derived from assessing the model's responses with both human and LLM evaluation. We find that the examined model scores far from ideal in legal reasoning, the model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet align only weakly with our trained annotators, i.e., reliable but not a valid substitute for human evaluation. Overall, the expert-curated prompt leads to more comprehensive reasoning, which does not result in more accurate predictions compared to the other examined settings. Based on our findings, we urge the community not to rely solely on automated LLM-based evaluation and to avoid using task accuracy as an appropriate proxy for reasoning quality.

1. Introduction

This study examines whether LLMs reason in a legally meaningful manner when forecasting ECtHR judgments, an understudied application where reasoning also supports explainability. It uses a small Article 10 case testbed to compare prompting strategies and evaluate reasoning quality through human and LLM-based assessment.

  • Motivation: Legal case forecasting requires sophisticated reasoning to interpret facts, connect them with legal knowledge, and predict decisions across legal fields and jurisdictions.The paper frames reasoning as more than a route to predictive accuracy: it also provides insight into model decision-making beyond brute-force pattern matching.
  • Motivation: The study tests whether LLMs can reason legally meaningfully by forecasting how the ECtHR would rule on applicants’ alleged violations.The ECtHR serves as a small-scale testbed in which the model assesses case facts and alleged violations.
  • Study design: The dataset contains 30 recent ECtHR cases officially published in HUDOC from April 2025 onwards.The cases concern Article 10 of the European Convention on Human Rights, concerning freedom of expression.
  • Study design: The study compares GPT 5.4 under minor task guidance, explicit step-by-step reasoning guidance, and implicit guidance through an official multi-page article guide.The examined model’s knowledge cut-off is August 31, 2025.
  • Evaluation: Model responses are evaluated by law students and LLM-as-a-Judge systems for reasoning quality rather than prediction decisions.The authors also release the human and LLM-based evaluation results for future comparison with other LLM-as-a-Judge schemes.

2. Examined Task and Data

The study uses 30 Article 10 ECtHR cases to test whether models can reason from curated case facts and predict violations. The setup approximates judicial reasoning but omits broader case files, verbatim precedent, and several components of the Court’s ruling.

  • Dataset: 30 Article 10 ECtHR cases from HUDOC provide facts, the relevant legal framework, and court decisions on whether Article 10 was violated.The cases concern alleged violations of freedom of expression by defendant states.
  • Task: The model receives curated facts and must generate legal reasoning and a prediction of the potential Article 10 violation.The input averages 4,993 words and reflects facts as presented in the Court’s decision.
  • Experimental setup: The input contains curated facts rather than the broader, unfiltered case files available to the ECtHR.The Court’s files include factual background, parties’ arguments, and evidence.
  • Experimental setup: The model receives the Court’s case-law references as an oracle and is asked to place them appropriately, rather than retrieving precedent through RAG.This avoids potential cascading retrieval failures in the assessment.
  • Experimental setup: The output covers Article 10 reasoning and violation prediction but excludes the ECtHR’s admissibility, damages, costs, and expenses determinations.The experiment therefore does not reproduce the Court’s full ruling process.
  • Evaluation concept: Legal reasoning is prioritized over predictive accuracy and must ground interpretation in case facts while applying the examined articles in a jurisprudentially meaningful way.The framework treats appropriate reasoning as the quality that legitimizes juridical decision-making, even when predictions are wrong.

3. Experimental Methodology

The study evaluates GPT-5.4 on ECtHR case assessments under three prompting settings, distinguishing formally disclosed reasoning from hidden reasoning. Human and LLM evaluators assess reasoning quality using structured criteria, with predictive accuracy reported separately.

  • Model and setup: The experiments use OpenAI’s GPT-5.4 with medium reasoning effort, selected because ChatGPT is widely used among legal professionals.The model’s knowledge cutoff was August 31, 2025.
  • Prompting settings: Three prompting settings vary the guidance supplied for assessing alleged Article 10 violations: Non-Curated, Curated (Expert), and Curated (Guide).The Expert setting supplies a legal-expert-curated step-by-step strategy, while the Guide setting first asks the model to infer a strategy from the ECtHR’s official guide.
  • Evaluation target: Evaluation targets the formally disclosed assessment rather than the model’s hidden reasoning, which is treated as a scratchpad.Evaluator models use high reasoning effort.
  • Human evaluation: Three senior law students independently evaluate responses using step occurrence, step comprehensiveness, and overall conciseness criteria.Step occurrence is binary; comprehensiveness and conciseness use 1–5 Likert scales, with missing steps receiving a comprehensiveness score of 0.
  • Comparison design: The mixed evaluation presents balanced pairs of settings, enabling both per-setting scores and pairwise win-rate comparisons.Each of the three setting pairs is used for exactly 10 cases, so every setting is assessed on 20 cases; the human evaluation covers 22 distinct cases as 30 case-pair instances.
  • Automated evaluation: LLM judges follow the human evaluation protocol, while factual reference overlap and predictive accuracy are reported as supplementary metrics rather than reasoning-quality measures.Predictive accuracy measures alignment between the model’s violation prediction and the court’s ruling.

4. Results & Discussion

GPT 5.4 usually performs the required doctrinal steps, but its analyses remain substantively incomplete, especially on lawfulness and proportionality. Evaluation rankings are robust yet weakly aligned with human judgments, and predictive accuracy provides essentially no signal about reasoning quality.

  • Step occurrence: Human evaluators found high step occurrence across settings (0.94–0.98), while curated prompts mainly reduced missed steps rather than teaching the reasoning process.Under setting A, curated prompts removed a residual 8–18% of missed steps.
  • Reasoning quality: All settings scored between 3 and 4 for step comprehensiveness and overall conciseness, with setting B ranked first, followed by C and A.The model’s reasoning was therefore structurally present but far from ideal compared with the Court’s assessment.
  • Predictive accuracy: 77–82% accuracy did not meaningfully distinguish settings because the always-violation baseline achieved 0.77, while comprehensiveness and correctness were essentially uncorrelated (r = 0.08, p = 0.55).Correctly and incorrectly predicted cases had nearly identical average comprehensiveness scores of 4.27 and 4.20.
  • Step-wise results: Step 2, assessing interference lawfulness, was most frequently missed, occurring in only ∼90% of setting-A cases by annotators and ∼82% by model judges.Later steps, particularly steps 3 and 4, showed poorer reasoning overall, with substantial deviations from the Court’s assessment.
  • Evaluation validity: All model judges agreed on the setting ranking (B>C>A), but judge–judge α was 0.41 versus human–human α of 0.09, while judge–human correlations were only ρ = 0.16–0.33.The judges were internally consistent but weakly aligned with the human consensus.

5. Limitations

The study’s conclusions are limited by its narrow case scope, small and potentially contaminated dataset, single-model design, and evaluator constraints. Future work should broaden the legal coverage and address these methodological limitations.

  • Limited Scope: The study examines ECtHR cases involving only one Convention article, limiting how broadly its findings generalize to ECtHR reasoning.Future work should explore a wider range of ECHR articles.
  • Small Dataset: The dataset contains only 30 Article-10 cases, which may not faithfully represent the model’s broader capability on that article.The cases were selected for recency, but their limited number constrains representativeness.
  • Small Dataset: Data contamination cannot be fully excluded because the model’s August 31, 2025 cutoff postdates part of the corpus published from April 2025 onward.Roughly half of the cases could fall within the training window, although memorization is expected to be limited for these recent, less-discussed judgments.
  • LLM-Judge models: The generator GPT-5.4 and one judge, GPT-5.5, share OpenAI as developer, creating a potential self-preference bias.Two additional judges from Anthropic and DeepSeek, together with per-judge reporting, were used to mitigate this concern and expose evaluator idiosyncrasies.
  • Human Annotators: Human annotations come from three senior law students under expert supervision rather than independent legal experts, and fine-grained 1–5 ratings are rater-dependent.They agree strongly on coarse judgments and setting rankings, while the study reports inter-annotator agreement and human–model alignment.

6. Ethical and Societal Implications

Deploying LLMs in legal settings poses transparency and evaluation risks because disclosed reasoning may not reflect actual computation. Such systems should support, rather than replace, human legal judgment, and automated evaluation should be validated against expert assessment.

  • Transparency: For closed models, billed hidden reasoning is not returned, so fluent explanations do not guarantee that stated reasons produced the decision.This creates a transparency concern for closed models used in high-stakes legal settings.
  • Human oversight: LLM-based legal systems should support, not replace, human legal judgment.The passage frames this as a practical safeguard for deployment in legal settings.
  • Evaluation: Automated evaluation should be validated against expert assessment rather than substituted for it.Expert assessment remains necessary when evaluating these systems in legal contexts.

7. Related Work

Related work has examined legal judgment prediction over ECtHR cases and stress-tested LLMs on free-text legal reasoning. Against this background, the study uses a narrow ECtHR setting while pairing expert-curated prompting with parallel human/model evaluation to separate reasoning quality from predictive accuracy.

  • Prior work: Prior work covers legal judgment prediction over ECtHR cases and recent LLM studies of free-text legal reasoning.Examples include step-wise verification and correction on Hong Kong court cases and a Greek Bar exam benchmark scored through a reasoning taxonomy.
  • Study positioning: The study examines one ECHR article, 30 recent cases, and one model, making its scope deliberately narrow.Within this setting, it combines an expert-curated reasoning prompt with parallel human and model evaluation.
  • Study positioning: Its step-wise analysis decouples reasoning quality from predictive accuracy and cautions against using task accuracy as a proxy for legal reasoning quality.The evaluation is designed to identify where reasoning quality and prediction accuracy diverge.

8. Conclusion and Future Work

GPT-5.4 reliably reproduces the doctrinal structure of Article-10 ECtHR cases, but its substantive reasoning remains shallow and correct predictions largely track majority outcomes rather than sound legal analysis. Future work will broaden the legal coverage, dataset, models, and annotator pool.

  • Conclusion: GPT-5.4 reliably reproduces doctrinal structure, but its substantive reasoning remains shallow.The study examined Article-10 ECtHR cases under three prompting strategies, with evaluation by three annotators and three LLM judges.
  • Conclusion: Correct predictions largely track the majority outcome rather than sound legal analysis.This finding qualifies the model’s apparent forecasting success.
  • Conclusion: The expert-curated prompt yields the most comprehensive reasoning.The supplied passage states this conclusion but is truncated before specifying further comparative details.
  • Future Work: Future work will extend the study to Articles 3 and 11, enlarge the dataset, evaluate several top-tier models, and recruit more senior annotators.The expanded annotator pool will support further study of inter-annotator agreement and expert reasoning quality.

A. Models and cvz … Case Facts

The study uses GPT-5.4 to generate ECtHR Article 10 case forecasts and three other recent models as LLM-as-a-Judge evaluators, comparing non-curated, expert-curated, and guide-curated prompting. Prompts require structured, stepwise legal assessments with factual and jurisprudential references, while evaluation measures whether four reasoning steps occur and how comprehensively they are addressed.

  • A. Models and cvz: GPT-5.4 generates forecasts, while GPT-5.5, Claude Opus 4.7, and DeepSeek V4 Pro evaluate responses as LLM-as-a-Judge models through OpenRouter.The paper provides exact model identifiers and reasoning-effort settings for reproducibility.
  • A. Models and cvz: The experiments compare Non-Curated, Curated (Expert), and Curated (Guide) prompting settings, denoted A, B, and C.These labels are defined in Table 4 as the study’s prompting settings.
  • A. Non-curated Instruction (Prompt): The non-curated prompt asks a legal assistant to forecast ECtHR outcomes by assessing case facts under Article 10 and ECtHR jurisprudence.It focuses specifically on whether freedom of expression has been violated.
  • A. Non-curated Instruction (Prompt): The non-curated format requires titled, step-separated paragraphs of up to 1000 words, factual paragraph references, and ECHR case-law references.The factual citation format is “(see paragraph P)”.
  • B. Curated (Step-by-step) Instruction (Prompt): The expert-curated prompt imposes a four-stage assessment strategy covering interference, lawfulness, legitimate aim, and proportionality-related reasoning.If no interference is found, the prompt instructs the model to predict no violation and skip the remaining steps.
  • B. Curated (Step-by-step) Instruction (Prompt): The expert-curated assessment must contain four titled, step-separated paragraphs of up to 1000 words with factual and ECHR case-law references.The structure is more prescriptive than the non-curated format because it specifies exactly four paragraphs.
  • C1. Inferring Strategy from Guide Instruction: The guide-curated approach first asks the model to extract a sequential Article 10 assessment methodology from the official ECtHR case-law guide.Each step must describe the relevant factual assessment and whether the court should proceed to later steps.
  • C2. Applying Inferred Strategy from Guide Instruction: The inferred guide strategy begins by identifying interference or a positive obligation, then testing Article 10 § 2 gateways including lawfulness and legitimate aim.The lawfulness test includes domestic legal basis, accessibility, foreseeability, and safeguards against arbitrariness or abuse.

C. Case Examples

The section presents two representative ECtHR cases, pairing case facts and the Court’s merits assessment with GPT-5.4 analyses under three prompting strategies. It also includes supplementary Mistral Medium 3.5 assessments under Prompts A and B.

  • Case presentation: Two representative cases are presented with the facts supplied to the model, the Court’s merits assessment as ground truth, and GPT-5.4 assessments under three prompting strategies.Mistral Medium 3.5 assessments under Prompts A and B are included as a supplementary comparison.
  • Case 1: Case 1, TERGEK v. T ¨URK˙IYE, concerned withheld letters and internet-printed enclosures sent to a detained applicant.The applicant was serving a prison sentence after conviction for membership of the FETÖ/PDY, and the letters included materials related to physiotherapy, education, and real-estate management.
  • Case 1: The first letter was withheld after prison authorities considered its enclosures potentially threatening to prison security and unclear in origin or purpose.The letter contained thirty-one pages of documents printed from the internet.
  • Case 1: Domestic proceedings produced divergent outcomes: an enforcement judge found the first withholding unlawful, whereas later authorities dismissed or upheld challenges concerning the letters.The Constitutional Court dismissed the applicant’s complaints as manifestly ill-founded, while the Assize Court endorsed the enforcement judge’s reasoning on the second letter.
  • Case 1: The second letter’s internet printouts were withheld under section 68(3), while a handwritten note and four pictures were authorised for delivery.The decision did not address the content of the withheld documents and relied on concerns about internet printouts and a prior administrative decision.

COURT’S ASSESSMENT (ARTICLE 10 — MERITS) · MODEL ASSESSMENTS

The Court treated the refusal to deliver the prisoner’s printed documents as an interference with Article 10 rights, but found it lawful, legitimate, and proportionate. It therefore concluded that retaining the documents did not violate Article 10.

  • COURT’S ASSESSMENT (ARTICLE 10 — MERITS): The refusal to deliver the documents interfered with the applicant’s right to receive information and ideas.
  • COURT’S ASSESSMENT (ARTICLE 10 — MERITS): The interference had a legal basis under either section 62 or section 68(3) of Law no. 5275.
  • COURT’S ASSESSMENT (ARTICLE 10 — MERITS): The interference pursued legitimate aims under Article 10 § 2, including national security, prevention of disorder, and prevention of crime.
  • COURT’S ASSESSMENT (ARTICLE 10 — MERITS): The Court assessed whether the authorities’ reasons were relevant and sufficient and whether the measure was proportionate to the legitimate aim pursued.
  • COURT’S ASSESSMENT (ARTICLE 10 — MERITS): For the general measure, the quality of judicial review was a practical alternative to case-by-case examination, while the measure’s application remained relevant to proportionality.
  • COURT’S ASSESSMENT (ARTICLE 10 — MERITS): The Constitutional Court had balanced prisoners’ access to information against prison authorities’ duties, workload, and risks associated with intra-organisational communication.
  • COURT’S ASSESSMENT (ARTICLE 10 — MERITS): The Court found no disproportionate restriction, accepted regulation of photocopied or printed documents within the authorities’ margin of appreciation, and held that no Article 10 violation occurred.

A. Non-Curated (GPT-5.4)

GPT-5.4 identified that withholding the detainee’s letters and internet printouts interfered with Article 10 rights, but domestic authorities relied on vague or categorical security concerns rather than individualized, content-based proportionality assessments. The analysis therefore concluded that both incidents failed to satisfy Article 10 requirements.

  • Article 10 framework: The measures concerned a detainee’s receipt of family letters and printed physiotherapy and distance-learning materials, engaging Article 10 despite imprisonment.Restrictions on incoming written material fall within Article 10.
  • First letter: For the first letter, authorities relied on general security assertions without adequately addressing the applicant’s concrete explanation that the materials concerned ankle rehabilitation and study.The reasoning lacked an individualized assessment of the material’s source, purpose, and content.
  • Second letter: For the second letter, authorities withheld sixty-one pages of internet printouts by invoking a general administrative decision without examining the documents’ content.The decision treated the internet printouts categorically rather than assessing their specific contents.
  • Proportionality: The proportionality analysis failed to identify any concrete security-threatening material or explain why less restrictive measures were unavailable, despite rehabilitative and educational purposes.Those purposes were described as ordinarily compatible with detention aims.
  • Overall conclusion: Overall, the authorities used vague security concerns and categorical objections, withholding the first materials unlawfully for a prolonged period and the second set without content analysis.Taken together, the incidents failed to provide the individualized examination required by Article 10.

B. Curated Step-by-step (GPT-5.4)

The curated step-by-step analysis identified an Article 10 interference with formal but deficient legal grounding, plausibly legitimate aims, and an apparent failure to satisfy the democratic-necessity requirement.

  • Curated Step-by-step (GPT-5.4): The authorities interfered with Article 10 rights by withholding a sister’s 31-page materials for one year and a wife’s 61-page materials while delivering only her note.The first materials were withheld from 22 October 2018 until 25 October 2019 after a successful objection.
  • Curated Step-by-step (GPT-5.4): The interference had a formal domestic-law basis, but the legal basis was deficient because authorities relied on general prison-security formulas rather than the case-specific material.The cited provisions included section 68(3) of Law no. 5275 and, for the second letter, section 62.
  • Curated Step-by-step (GPT-5.4): The measures plausibly pursued preventing disorder or crime and protecting prison security, given concerns that the materials could facilitate communication within FETÖ/PDY.These aims qualify as legitimate under Article 10 § 2 but do not by themselves establish necessity.
  • Curated Step-by-step (GPT-5.4): The interference did not appear necessary in a democratic society because domestic authorities failed to demonstrate a pressing social need with relevant and sufficient reasons.The first documents were later delivered in full after an enforcement judge found the original decision unlawful for relying only on a general inspection.

C. Curated Guide (GPT-5.4)

The curated guide identifies an Article 10 interference in the withholding of the applicant’s letters and enclosures, while finding the restrictions disproportionate because they relied on generic security concerns and categorical exclusions rather than specific content assessments. Domestic review did not adequately remedy these defects, particularly for the second letter.

  • Interference with freedom to receive information: The prison authorities’ withholding of two letters and their enclosures interfered with the applicant’s Article 10 right to receive information.The materials reportedly concerned physiotherapy exercises and study materials for a distance-learning course.
  • Legal basis and legitimate aim: The interference pursued legitimate prison-order and security aims under sections 68 and 62 of Law no. 5275, but its lawfulness was doubtful.The authorities also invoked preventing communication within FETÖ/PDY as a legitimate aim under Article 10 § 2.
  • Necessity and proportionality: The first withholding lacked a concrete, content-based justification because authorities cited no specific passage or identifiable security risk.The domestic court accepted that withholding material merely because its content was unclear after general inspection was unlawful.
  • Necessity and proportionality: The second withholding was harder to justify because sixty-one pages were excluded under a blanket policy against internet printouts despite stated medical and educational purposes.The materials were not linked in the facts to violence, propaganda, coded messages, or any concrete institutional risk.
  • Quality of domestic review: Domestic review did not cure the interference: the first remedy came after lengthy delay, while review of the second endorsed formalistic, category-based exclusion.The Constitutional Court’s dismissal was described as formulaic, and the overall restrictions were assessed as disproportionate even allowing a wider prison-security margin of appreciation.
Loading 2608.17168v1…