Source-linked AI summary
Can Legal AI Know When It Is Wrong? And Do Students Know When It Is?
Angel Mary John, Vipin Kumar Singh, Jerrin Thomas Panachakel
TL;DR
The paper asks how legal AI overconfidence and human verification interact during Indian statutory change. It combines a 60-case model audit with a survey of 380 law students, finding high-confidence errors alongside verification differences associated with hallucination exposure. It concludes that adversarial research practices and verifiable AI architectures are needed within the study’s supported scope.
Problem
Existing evaluations do not adequately explain whether legal-AI failures reflect ignorance of new law or historical precedent overfitting, nor how this interacts with user trust in India.
Method
The study conducts a dual-phase socio-technical audit: black-box testing of three LLM ecosystems on 60 Indian legal cases and a cross-sectional survey of 380 LLB students.
Results
Meta AI recorded the highest High-Confidence Error Rate at 31.7%, while students reporting fabricated-citation encounters had a mean verification score of 4.2/5 versus 2.8/5 without such encounters.
Takeaways & Limitations
The paper concludes that adversarial legal-research pedagogy and verifiable AI architectures should replace reliance on generic human-in-the-loop frameworks.
Abstract
from arXiv · showhide
Integrating Large Language Models (LLMs) into the Indian judiciary promises access to justice but introduces severe risks. We identify the 'inertia of confidence'--an overconfidence phenomenon analogous to the Dunning-Kruger effect where LLMs provide incorrect legal verdicts with near-maximum confidence, driven by a hypothesized 'precedent overfitting' bias. Phase I of our socio-technical audit tested ChatGPT (GPT-5.2), Meta AI, and Perplexity AI on a 60-case battery regarding the Indian Contract Act, 1872, and the shift toward statutory enforcement of specific performance. We introduce the High-Confidence Error Rate (HCER) to quantify incorrect verdicts delivered with dangerous certainty (>= 9 on a 1-10 scale). All models struggled with statutory updates. Meta AI proved most vulnerable (31.7% HCER), frequently misapplying pre-amendment rules with a 9.1/10 mean confidence, followed by Perplexity (15.0%) and ChatGPT (6.7%). Phase II investigated human vulnerability to this overconfidence via a survey of Indian law students (N=380). Verification often functions as a reactive adaptation to machine hallucinations: students encountering fabricated citations reported higher verification scores (4.2/5) than those with no such encounters (2.8/5). Furthermore, while 81.6% knew submitting hallucinated cases can lead to contempt-of-court, 71.1% received no formal training on ethical AI use. We propose shifting toward adversarial legal research pedagogy and implementing source-grounded verification architectures to prevent systemic professional negligence.
I. INTRODUCTION
The paper examines how legal AI’s adoption in India intersects with hallucination, historical bias, and human over-reliance. It addresses the gap between general LLM competency and socio-technical risks arising during statutory change.
- Indian courts and legal workflows are increasingly integrating AI for translation, research assistance, automation, and contract analytics.
- The study names “inertia of confidence” as near-maximum certainty accompanying incorrect legal verdicts during domain-specific failures.It links this failure to a hypothesized “precedent overfitting” bias favoring historically prevalent pre-amendment jurisprudence over newer statutory developments.
- The research conducts a dual-layered audit combining specialized model benchmarking with a survey of Indian LLB students.It argues for adversarial legal research and verifiable AI architectures rather than generic human-in-the-loop frameworks.
- Existing research reports strong general LLM performance, but Indian legal evaluations identify struggles with complex reasoning and authority discipline.
- The literature has not adequately examined whether models fail because they lack recent legal knowledge or because historical precedent overwhelms statutory overrides.
- The paper therefore links algorithmic overconfidence, user cognitive offloading, and institutional preparedness within the Indian legal ecosystem.It frames this as a socio-technical research gap concerning whether human users effectively filter high-confidence machine errors.
D. Ethical, Regulatory, and Judicial Frameworks
The paper situates legal AI within existing judicial safeguards and investigates whether Indian legal education can mitigate hallucination-related liability. Its audit separates machine behavior from student verification and institutional readiness.
- Unverified AI citations have already been associated with real-world judicial sanctions, increasing the urgency of examining legal-AI safeguards.
- The study identifies a lack of empirical research connecting LLM calibration problems and cognitive offloading within India’s legal ecosystem.
- RQ1 asks whether frontier LLMs show “inertia of confidence” when pre-amendment precedent conflicts with recent statutory overrides.
- RQ2 examines the association between exposure to hallucinated legal citations and Indian law students’ verification habits.
- RQ3 assesses institutional readiness and policy interventions addressing AI-induced liabilities and “unprotected accountability.”
- The methodology uses separate model evaluation and cross-sectional student survey phases to address algorithmic, behavioral, and institutional questions.
2) The 60-Case Judicial Agent Battery:
The technical audit uses a 60-case, black-box battery to test Indian contract-law reasoning under statutory and jurisprudential change. Models must provide verdicts, authorities, and confidence scores.
- The battery tests foundational contract law alongside the Specific Relief (Amendment) Act, 2018’s shift toward presumptive statutory specific performance.
- The 60 cases are divided into six categories spanning offer and acceptance, capacity and formation, consideration and privity, and discharge and frustration.
- The temporal stress test probes whether models recognize uncertainty surrounding the amendment’s prospective or retrospective application after the recalled 2022 ruling.
- Case-identifying names, dates, and locations are removed so systems must reason from factual narratives rather than recognizable authorities.
- A standardized Judicial Persona prompt requires each model to act as a senior Indian jurist and deliver a definitive conclusion.The framing may partly influence observed confidence levels.
- For every scenario, models provide a definitive verdict, supporting statutory authority, and self-assessed confidence from 1 to 10.Interactions are logged and manually aggregated before HCER calculation.
1) Participant Demographics and Ethical Considerations:
The study combines a diverse but non-probability sample of 380 LLB students with structured measures of hallucination exposure, verification, training, and liability awareness. The technical audit finds strong control-set accuracy but sharply lower performance on statutory updates.
- Participant Demographics and Ethical Considerations: N = 380 undergraduate law students were recruited through purposive convenience sampling across national, private, and state-affiliated institutions.The sample included students from first year through final year, while limiting generalized inference.
- Participant Demographics and Ethical Considerations: Participation was voluntary and anonymous, informed consent was required, and directly identifying information was not retained.
- Participant Demographics and Ethical Considerations: The 10-question survey covered tool use, hallucination encounters, institutional support, training, career outlook, and anxiety.
- Participant Demographics and Ethical Considerations: Institutional training status tracks whether law schools provide formal ethical-AI training, while liability awareness includes judicial penalties such as contempt of court.
- Category Accuracy and the 2018 Amendment Paradox: 70% ChatGPT, 60% Perplexity, and 50% Meta AI accuracy were recorded on Cases 51–60 concerning Specific Relief and the 2018 Amendments.By contrast, established doctrinal controls ranged from 80% to 100%, with GPT-5.2 reaching 100% in both reported control categories.
- Category Accuracy and the 2018 Amendment Paradox: The observed pattern is consistent with precedent overfitting, but the black-box design cannot establish the underlying causal mechanism.
B. Confidence–Correctness Mismatch: The High-Confidence Error Rate
The audit operationalizes confidence–correctness mismatch through HCER and finds high-confidence legal errors, including statutory and task-boundary failures. These results motivate assessing how algorithmic unreliability affects law students.
- High-Confidence Error Rate: HCER measures the percentage of cases producing an incorrect legal conclusion with self-reported confidence of at least 9/10.The metric uses correctness and confidence for each evaluated case.
- High-Confidence Error Rate: 31.7% HCER for Meta AI exceeded Perplexity AI's 15.0% and ChatGPT's 6.7%, while mean confidence remained high despite lower accuracy on modern statutory cases.Mean confidence ranged from 8.8/10 for Perplexity AI to 9.4/10 for ChatGPT (GPT-5.2).
- Failure Modes: The audit identified retrospective–prospective error and statutory fabrication as two qualitative failure modes associated with overconfidence.Examples included misapplying the 2018 Amendment and hallucinating a 7-day cure period instead of the prescribed 30-day notice period.
- Instructional Over-Compliance: Meta AI generated 30 additional synthetic legal scenarios after reaching the final prompt instead of signaling completion of the 60-case battery.The outputs followed the requested format but were not grounded in actual Indian statutes or reported case law.
- Instructional Over-Compliance: The synthetic outputs suggest that maintaining conversational continuity and formatting can take precedence over acknowledging factual boundaries in bounded legal tasks.The paper characterizes this behavioral pattern as instructional over-compliance and identifies an operational risk for legal practitioners.
- Human Impact: The human-impact phase surveyed 380 undergraduate law students across hallucination exposure, verification frequency, AI ethics training, and job-displacement anxiety.These variables were used to evaluate behavioral responses to algorithmic unreliability.
1) Hallucination Exposure and Verification Behavior:
The study links high-confidence legal-AI errors with human verification behavior and broader socio-technical vulnerabilities in Indian legal practice. Students with prior hallucination exposure reported more verification, while technical and institutional gaps leave room for overreliance.
- Hallucination Exposure: 42.1% of students reported multiple encounters with fabricated case-law citations, 36.8% occasional encounters, and 21.1% none.The survey classified hallucination exposure across all 380 respondents.
- Verification Behavior: Students reporting multiple fabricated-citation encounters had a mean verification score of 4.2/5, versus 2.8/5 among students reporting none.The cross-sectional design supports an association consistent with, but does not establish, a reactive verification response.
- Algorithmic Risk: Meta AI recorded the highest overall HCER at 31.7%, indicating a confidence–correctness mismatch in the tested legal-reasoning task.The audit identifies high-confidence errors as a technical vulnerability alongside human verification behavior.
- Algorithmic Risk: Historical data weighting may bias models toward pre-amendment rules, but the black-box design cannot establish the underlying causal architecture definitively.The observed error patterns are consistent with the precedent-overfitting hypothesis and reflect systematic bias toward pre-amendment rules.
- Socio-Technical Risk: The study describes a double blindspot combining algorithmic temporal lag with limited student preparation for AI-related legal accountability.The proposed vulnerability joins model misinterpretation of recent amendments with insufficient pedagogical support.
VII. POLICY INTERVENTIONS FOR THE INDIAN JUDICIARY
The paper proposes moving from generic oversight toward structured, adversarial verification of AI-assisted legal work. Its recommendations target both student training and institutional safeguards, while recognizing that the evidence remains limited to students rather than practicing legal professionals.
- Policy rationale: 71.1% institutional training gap motivates a policy framework for structured AI verification in Indian legal education.The framework responds to high-confidence technical failures and insufficient institutional preparedness.
- Curriculum reform: Verification appears partly reactive to prior hallucination exposure, so the paper recommends developing it as a proactive skill.The authors propose incorporating adversarial legal research into LLB practical training.
- Curriculum reform: The proposed curriculum evaluates students on red-teaming AI outputs rather than banning AI use.Suggested activities include hallucination-detection labs and temporal checks against the latest reported judgments.
- Technological safeguards: 78.9% self-reported fake-citation exposure supports requiring a Verifiable Authority Index that links outputs to verified legal reporters.The proposed source-grounding targets SCC, AIR, and the e-SCR portal.
- Institutional safeguards: An Internal AI Verification Protocol would provide a structured, source-grounded human-audit layer for AI-assisted submissions.The authors connect systematic verification with professional protection and institutional integrity, while describing student anxiety as possibly correlated with absent safeguards.
- Conclusion and future work: The conclusion links high-confidence statutory failures and reactive student verification to the need for adversarial pedagogy and verifiable AI architectures.Future research should test whether professional experience mitigates the automation bias observed in students.
APPENDIX A THE 60-CASE JUDICIAL AGENT BATTERY
The appendix defines a 60-case judicial-agent battery combining doctrinal controls with a temporal stress test on amended specific-performance law. It specifies the model prompt, a representative scenario, scoring criteria, and control cases spanning core Indian Contract Act principles.
- Battery design: 60 benchmark cases comprise doctrinal control cases 1–50 and a temporal stress test, Cases 51–60, focused on the Specific Relief (Amendment) Act, 2018.Comparative authorities provide factual or doctrinal grounding, but the gold standard remains Indian statutory law.
- Prompt: The system prompt requests a definitive verdict, supporting statutory authority, and self-assessed confidence from 1 to 10.This prompt supplies the output format used in the audit.
- Temporal stress test: Case 54 tests whether substituted-performance costs are recoverable without the amended statute’s required prior written notice.The scenario concerns a buyer who hires a third party after the seller’s breach and immediately seeks the replacement costs.
- Scoring rubric: The correct response rejects recovery because Section 20(2) requires written notice of not less than 30 days.This is the pass criterion for the representative amended-law scenario.
- Scoring rubric: The HCER trigger covers permission based on pre-amendment discretionary principles or a fabricated notice period when confidence is at least 9.The rubric treats either legal error, paired with very high confidence, as an incorrect high-confidence response.
- Formation and acceptance controls: The battery also includes contract-formation edge cases involving price quotations, tenders, cross-offers, and counter-offers.The relevant authorities include Harvey v. Facey, Union of India v. Maddala Thathiah, Tinn v. Hoffman & Co., and Hyde v. Wrench.
Category 2: Capacity, Consent, Formation & Related Contract Principles (Control Set)
This control-set category covers capacity, consent, formation, consideration, lawful object, and related contractual principles through statutory issues paired with Indian or comparative authorities. The cases range from minors and mental capacity to restraints, privity, wagers, and contractual enforcement limits.
- Category 2: Capacity, Consent, Formation & Related Contract Principles (Control Set): Capacity and consent cases address minors’ agreements, necessaries, restitution, lucid intervals, undue influence, and mutual mistake.The controls use Sections 11, 12, 16, 20, and 68 of the Indian Contract Act or related Indian restitution principles.
- Category scope: Together, these controls test statutory application across capacity, consent, formation, consideration, lawful object, and enforcement-related doctrines.The appendix pairs each issue with a specified statutory provision or Indian contractual principle and an authority.
- Category 2: Capacity, Consent, Formation & Related Contract Principles (Control Set): Formation-related controls test certainty of terms, family settlements, coercion through threats of suicide, and fraudulent misrepresentation of identity.These cases invoke Sections 29, 15, and 19, alongside Indian authorities addressing family settlements and identity fraud.
- Category 3: Consideration & Lawful Object (Control Set): Consideration controls examine privity of consideration, privity of contract, past voluntary service, charitable subscriptions, and consideration moving at the promisor’s desire.The cases draw on Section 2(d), Section 25, and Indian privity principles.
- Category 3: Consideration & Lawful Object (Control Set): Lawful-object controls cover wagering collateral agreements, negative employment covenants, restraints of trade, and marriage brokerage agreements against public policy.These cases apply Sections 26, 27, and 30 of the Indian Contract Act.
- Category 3: Consideration & Lawful Object (Control Set): The category includes contractual limitation and enforcement restrictions under Section 28 of the Indian Contract Act.The listed authority is Food Corporation of India v. New India Assurance.
Category 4: Discharge, Frustration & Restitution (Control Set)
This control set covers discharge, frustration, restitution, breach, mistake, payment, novation, and commercial hardship under the Indian Contract Act. Each case pairs a defined issue with a statutory provision and supporting authority.
- Discharge and frustration: Discharge and frustration controls test physical destruction, frustration, supervening impossibility, and commercial hardship under Section 56.The listed authorities include Taylor v. Caldwell, Satyabrata Ghose, and Naihati Jute Mills.
- Restitution: Restitution controls examine recovery for lawful non-gratuitous benefits and compensation after partial performance under Section 70.The cases cite Damodar Mudaliar and Sumpter v. Hedges.
- Breach: Breach-related controls address the immediate right to sue following anticipatory breach under Section 39.The supporting authority is Hochster v. De La Tour.
- Mistake and payment: Mistake and payment controls cover res extincta, mutual mistake of fact, and money paid under mistake or coercion.The cases invoke Sections 20 and 72 and cite Couturier v. Hastie and Kanhaiya Lal.
- Contract modification: The set also tests novation and alteration of contracts under Section 62, with Lata Construction as the listed authority.This case evaluates a change in contractual obligations through novation or alteration.
- Payment allocation: Appropriation of payments is tested under Sections 59–61 through Clayton’s Case.The case addresses the statutory rules governing allocation of payments.
Category 5: Damages, Contractual Terms & Enforcement (Control Set)
This control set tests damages, contractual terms, enforcement, and statutory updates across the Indian Contract Act and Specific Relief Act. It includes issues ranging from remoteness and penalties to notice, standard-form terms, substituted performance, and amendment-related enforcement rules.
- Damages: Cases 41–43 test remoteness of damage, recoverability of special damages, and whether stipulated sums are penalties or genuine pre-estimates.The cases invoke Section 73 or Section 74 of the Indian Contract Act alongside Hadley v. Baxendale, Victoria Laundry, and ONGC v. Saw Pipes.
- Contractual Terms & Enforcement: Cases 44–50 examine incorporation, notice, coercion, undue influence, discharge, standard-form terms, and contractual formation.Authorities include Parker, Olley, Pioneer Urban, Masters v. Cameron, Boghara Polyfab, Fateh Chand, and Bharathi Knitting.
- Statutory Updates: Case 51 evaluates whether models recognize the unsettled temporal status of the 2018 Specific Relief Act amendment rather than present a recalled ruling as an unqualified current rule.The rubric uses Katta Sujatha Reddy v. Siddamsetty Infra Projects, recalled through 2024 INSC 861, as the legal baseline.
- Specific Performance: The control set also tests whether injunctions would impede infrastructure projects and what relief is available to subsequent purchasers.These issues are assigned to Cases 59 and 60 under Sections 20A and 19(b) of the Specific Relief Act.
Sample Scenario, Prompt, and Scoring Rubric
The appendix provides a complete example of the benchmark’s prompt, factual scenario, and scoring rubric, then introduces the student survey measuring the Trust-Accuracy Gap. The example operationalizes failure as a confident departure from the amended statutory rule, while the survey examines students’ exposure to and verification of AI-generated legal content.
- Sample Scenario, Prompt, and Scoring Rubric: Due to space constraints, the paper presents one complete example of the prompt structure, factual scenario, and evaluation rubric.
- Sample Scenario, Prompt, and Scoring Rubric: The system prompt asks the model to act as a senior Indian jurist and provide a definitive verdict, statutory authority, and confidence score from 1 to 10.
- Sample Scenario, Prompt, and Scoring Rubric: The Case 54 scenario asks whether a buyer may recover substituted-performance costs after hiring a third party without prior written notice.
- Sample Scenario, Prompt, and Scoring Rubric: A passing answer denies recovery and cites the mandatory 30-day written notice requirement under Section 20(2) of the 2018 Specific Relief amendment.
- Sample Scenario, Prompt, and Scoring Rubric: An HCER-triggering failure permits recovery under pre-amendment principles or invents a notice period while reporting confidence of ≥9.
- APPENDIX B THE 10-QUESTION STUDENT SURVEY INSTRUMENT: The student survey measures the Trust-Accuracy Gap among N = 380 students through questions about AI use, fabricated citations, contempt-of-court awareness, and answer verification.