Paper deep dive
Can Legal AI Know When It Is Wrong? And Do Students Know When It Is?
Angel Mary John, Vipin Kumar Singh, Jerrin Thomas Panachakel
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 8/24/2026, 6:05:44 AM
Summary
This paper presents a dual-layered socio-technical audit of Large Language Models (LLMs) in the Indian legal context, identifying a phenomenon termed 'inertia of confidence' where models provide incorrect legal verdicts with high certainty due to 'precedent overfitting' on historical data. Phase I benchmarks ChatGPT (GPT-5.2), Meta AI, and Perplexity AI on a 60-case battery regarding the Indian Contract Act, 1872, and the Specific Relief (Amendment) Act, 2018, introducing the High-Confidence Error Rate (HCER) metric. Meta AI showed the highest vulnerability (31.7% HCER). Phase II surveys 380 Indian law students, revealing that while most are aware of contempt-of-court risks, the majority lack formal AI ethics training, and verification often serves as a reactive measure to prior hallucination encounters.
Entities (10)
Relation Signals (8)
ChatGPT (GPT-5.2) â hashighconfidenceerrorrate â 6.7%
confidence 98% ¡ and ChatGPT (6.7%)
Meta AI â hashighconfidenceerrorrate â 31.7%
confidence 98% ¡ Meta AI proved most vulnerable (31.7% HCER)
Perplexity AI â hashighconfidenceerrorrate â 15.0%
confidence 98% ¡ followed by Perplexity (15.0%)
Inertia of Confidence â isanalogousto â Dunning-Kruger effect
confidence 95% ¡ an overconfidence phenomenon analogous to the Dunning-Kruger effect
Indian Law Students â awareofconsequence â Contempt of Court
confidence 94% ¡ 81.6% knew submitting hallucinated cases can lead to contempt-of-court
Indian Law Students â lacksformaltraining â AI Ethics
confidence 94% ¡ 71.1% received no formal training on ethical AI use
Inertia of Confidence â causedby â Precedent Overfitting
confidence 93% ¡ driven by a hypothesized 'precedent overfitting' bias
Verification â isreactiveto â Machine Hallucinations
confidence 92% ¡ Verification often functions as a reactive adaptation to machine hallucinations
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Integrating Large Language Models (LLMs) into the Indian judiciary promises access to justice but introduces severe risks. We identify the 'inertia of confidence'--an overconfidence phenomenon analogous to the Dunning-Kruger effect where LLMs provide incorrect legal verdicts with near-maximum confidence, driven by a hypothesized 'precedent overfitting' bias. Phase I of our socio-technical audit tested ChatGPT (GPT-5.2), Meta AI, and Perplexity AI on a 60-case battery regarding the Indian Contract Act, 1872, and the shift toward statutory enforcement of specific performance. We introduce the High-Confidence Error Rate (HCER) to quantify incorrect verdicts delivered with dangerous certainty (>= 9 on a 1-10 scale). All models struggled with statutory updates. Meta AI proved most vulnerable (31.7% HCER), frequently misapplying pre-amendment rules with a 9.1/10 mean confidence, followed by Perplexity (15.0%) and ChatGPT (6.7%). Phase II investigated human vulnerability to this overconfidence via a survey of Indian law students (N=380). Verification often functions as a reactive adaptation to machine hallucinations: students encountering fabricated citations reported higher verification scores (4.2/5) than those with no such encounters (2.8/5). Furthermore, while 81.6% knew submitting hallucinated cases can lead to contempt-of-court, 71.1% received no formal training on ethical AI use. We propose shifting toward adversarial legal research pedagogy and implementing source-grounded verification architectures to prevent systemic professional negligence.
Tags
Links
- Source: https://arxiv.org/abs/2608.21089v1
- Canonical: https://arxiv.org/abs/2608.21089v1
Trouble viewing inline? Open PDF directly â
Full Text
58,711 characters extracted from source content.
Expand or collapse full text
Can Legal AI Know When It Is Wrong? And Do Students Know When It Is? Angel Mary John1, Vipin Kumar Singh1, and Jerrin Thomas Panachakel2 Affiliation: 1Department of Law Sunrise University Alwar, Rajasthan, India Affiliation: 2School of Electrical and Electronic Engineering Technological University Dublin Dublin, Ireland jerrin.panachakel@tudublin.ie Abstract The rapid integration of Large Language Models (LLMs) into the Indian judiciary promises unprecedented access to justice, yet it introduces severe, unquantified risks. This paper identifies and quantifies a phenomenon we term the âinertia of confidenceââan overconfidence phenomenon analogous to the Dunning-Kruger effect where LLMs provide incorrect legal verdicts with near-maximum confidence consistent with a hypothesized algorithmic bias that we term âprecedent overfitting.â We conducted a dual-layered socio-technical audit. In Phase I, we subjected ChatGPT (GPT-5.2), Meta AI, and Perplexity AI (Sonar) to a specialized 60-case battery testing the Indian Contract Act, 1872, and the shift from discretionary relief toward statutory enforcement of specific performance, subject to specified exceptions. Our technical audit reveals a pronounced pattern of high-confidence errors across the tested systems. To measure this, we introduce the High-Confidence Error Rate (HCER)âa novel metric quantifying the percentage of incorrect verdicts delivered with dangerous certainty (âĽ9⼠9 on a 1â101-10 scale). While all models struggled with statutory updates, Meta AI proved the most vulnerable with a HCER of 31.7%, frequently misapplying pre-amendment legal rules while asserting a near-perfect mean confidence score of 9.1/10, followed by Perplexity (15.0%) and ChatGPT (6.7%). In Phase I, we investigated the human vulnerability to this algorithmic overconfidence through a primary survey of Indian law students (N=380N=380). We empirically demonstrate that formal AI-related training appears limited, and manual verification frequently functions as a reactive adaptation following prior encounters with machine hallucinations; students reporting multiple prior encounters with fabricated citations reported a higher descriptive mean verification score (4.2/5) than students reporting no such encounters (2.8/5), providing preliminary cross-sectional evidence. Furthermore, our findings expose a severe institutional gap: while 81.6% reported awareness that submitting hallucinated cases to an Indian court can lead to contempt-of-court consequences, 71.1% reported receiving no formal training on the ethical use of AI. To prevent systemic professional negligence, we propose a shift toward adversarial legal research pedagogy and the implementation of mandatory, source-grounded verification architectures for legal AI systems. Index Terms: Artificial Intelligence, Metacognitive Calibration, High-Confidence Error Rate, Indian Contract Act, Legal AI, Algorithmic Bias, Socio-Technical Systems. I Introduction The legal landscape in 2026 has transitioned from a skeptical evaluation of artificial intelligence (AI) to its integration as a foundational utility. Globally, the adoption of large language models (LLMs) has revolutionized routine legal workflows, including document automation, e-discovery, and contract analytics [22]. In the Indian context, this shift is particularly pronounced. The Supreme Court of India has championed AI for âAccess to Justiceâ through initiatives like SUVAS (Supreme Court Vidhik Anuvaad Software) for judgment translation and SUPACE for judicial research assistance [18, 10]. However, while AI enhances efficiency, it introduces significant risks regarding the accuracy of legal reasoning, especially when models encounter the specific statutory nuances of the Indian legal system. A critical challenge in the deployment of LLMs for legal research is the phenomenon of âhallucination.â While foundational studies suggest that models generally possess good metacognitive calibration on standard benchmarks [13], this study reveals a severe domain-specific breakdown. We conceptualize this behavior as the âinertia of confidenceâ: an overconfidence phenomenon analogous to the Dunning-Kruger effect where a model maintains near-maximum certainty despite delivering factually incorrect legal verdicts. In the Indian legal scenario, this is most evident during major legislative shifts, such as the transition from discretionary to mandatory relief under the Specific Relief (Amendment) Act, 2018 [19]. To explain why models fail in these exact scenarios, we introduce the concept of âprecedent overfitting.â This failure does not occur because the training data lacks the new amendment; rather, it is an algorithmic bias where the sheer statistical volume of historical, pre-amendment judgments may exert disproportionate influence on model outputs relative to recent statutory developments, causing models to default to outdated principles while confidently asserting modern statutory compliance. The primary objective of this research is to perform a dual-layered audit of the current state of legal AI. First, we benchmark the performance of leading LLMs against a specialized 60-case batteryâthe âjudicial agent benchmarkââfocused on the Indian Contract Act, 1872, and Specific Performance. Second, we complement the technical audit with a primary survey of 380 LLB students in India to assess verification behavior and institutional readiness. By linking machine overconfidence with human over-reliance, this paper argues that generic âhuman-in-the-loopâ frameworks are insufficient; instead, it proposes a paradigm shift toward âadversarial legal researchâ and verifiable AI architectures to mitigate blind technological trust. I Related Work The academic inquiry into Large Language Models (LLMs) within the legal domain has transitioned from assessing general competency to identifying specialized failure modes. This section provides a comprehensive review of the literature, categorized into academic performance benchmarks, jurisdiction-specific audits, and the socio-cognitive impact of AI on legal pedagogy. I-A LLMs in Academic and Professional Legal Gatekeeping The foundational âbenchmarkingâ era began with evaluations of AI on standardized legal examinations. In the seminal study âChatGPT Goes to Law School,â Choi et al. [3] demonstrated that early models could achieve passing grades across core courses, albeit with significant struggles in complex reasoning. This was expanded by Katz et al. [15], who reported that GPT-4 outperformed 90% of human test-takers on the Uniform Bar Examination (UBE). In the Indian context, Tiwari et al. developed Aalap, a fine-tuned Mistral 7B model for Indian legal and paralegal tasks, reporting performance comparable to or better than GPT-3.5 on portions of their evaluation. However, Guha et al. [8] and rigorous calibration studies [9] suggest that while AI clears objective benchmarks, it fails to replicate the âdesirable difficultyâ required for high-level legal analysis, often performing below the average of top-tier law students in open-ended evaluations. I-B Jurisdiction-Specific Benchmarks and Indian Legal NLP Recent scholarship has moved toward âjurisdiction-groundedâ evaluations to address the nuances of non-Western legal systems. Juvekar et al. [12] introduced an exam-grounded, India-specific yardstick for LLM court-readiness, combining objective examinations with lawyer-graded long-form answers from the Supreme Courtâs Advocate-on-Record examination. Their findings indicate that while frontier models clear objective benchmarks, they fail in âauthority discipline.â This is supported by the development of BhashaBench V1 [4] and InLegalBERT [14, 20, 17], which represent a shift toward Indic-language legal understanding. Despite these advances, current literature predominantly attributes AI legal errors to a simple âknowledge cutoffâ or temporal lag, assuming models are merely unaware of recent laws [21, 7]. However, this framework fails to address scenarios where modern statutory updatesâsuch as the 2018 Amendments to the Specific Relief Actâare chronologically well within the modelsâ training corpora. Existing evaluations have yet to investigate the potential for âprecedent overfitting,â where the sheer statistical volume of historical, discretionary Common Law precedent might overwhelm recent legislative overrides. Consequently, there remains a critical gap in understanding whether models fail out of true ignorance or due to an algorithmic bias toward historical data, making it necessary to evaluate how they process these specific statutory shifts. I-C The Dunning-Kruger Effect and Cognitive Offloading The intersection of AI performance and human psychology is central to understanding the evolving dynamics of user trust in automated systems. Foundational evaluations of AI metacognition, such as Kadavath et al. [13], have generally concluded that LLMs âmostly know what they know,â demonstrating strong self-calibration on standard Q&A benchmarks. However, our research challenges this assumption within specialized legal domains. We hypothesize that when faced with âprecedent overfitting,â this baseline calibration collapses, resulting in an overconfidence phenomenon analogous to the Dunning-Kruger effect where models deliver incorrect jurisprudence with maximum confidence. Complementing this algorithmic risk is the human vulnerability. Fernandes et al. [5] identified that AI usage can paradoxically lead to a disconnect between actual performance and metacognition, where users blindly trust a single AI output without independent verification. This âcognitive offloadingâ is particularly dangerous in the legal profession, where Buçinca et al. [1] note a troubling overreliance on AI in assisted decision-making. Furthermore, broad demographic studies, such as the UNESCO global guidance [23] and FICCI-EY India reports [6], highlight that while the vast majority of students now use AI, nearly half feel unprepared to critically evaluate its outputs. I-D Ethical, Regulatory, and Judicial Frameworks Government reports from the PIB [18] and NITI Aayog [16] emphasize a âhuman-in-the-loopâ strategy, particularly as courts implement tools like SUVAS and SUPACE. However, as documented in the AI Hallucination Cases Database [2], the reliance on unverified AI citations has already led to real-world judicial sanctions, reinforcing the urgency of our dual-audit approach. I-E Research Gap and Objectives While existing literature in computer science and human-computer interaction establishes the isolated existence of LLM calibration issues and the dangers of cognitive offloading, there is a distinct lack of empirical research bridging these socio-technical phenomena within the Indian legal ecosystem. Specifically, it remains unclear how the algorithmic bias of âprecedent overfittingâ interacts with user trust, and whether human users act as an effective filter against high-confidence machine errors. To address these gaps, this study poses the following research questions: 1. RQ1 (Algorithmic): To what extent do frontier LLMs exhibit an âinertia of confidenceâ when evaluating Indian legal scenarios where historical pre-amendment precedents conflict with recent statutory overrides (e.g., the 2018 Amendments to the Specific Relief Act)? 2. RQ2 (Behavioral): How is prior exposure to AI-generated hallucinated legal citations associated with verification habits among Indian law students? 3. RQ3 (Institutional): What is the current state of institutional readiness in Indian legal education to mitigate AI-induced liabilities, and what policy interventions are required to address the resulting âunprotected accountabilityâ? I Methodology To investigate the dual dimensions of the âinertia of confidence,â this study employs a dual-phase socio-technical audit combining a model evaluation and a cross-sectional student survey. The methodology is bifurcated to independently address the algorithmic behaviors of the machines (RQ1) and the psychological and institutional perceptions of the human users (RQ2 and RQ3). I-A Phase I: Algorithmic Audit for Jurisprudential Inertia (RQ1) I-A1 Model Selection and Black-Box Access To evaluate the extent of Jurisprudential Inertia, the audit utilized a âBlack-Boxâ testing approach. Rather than accessing the models via backend Application Programming Interfaces (APIs) with artificially adjusted temperature parameters, the evaluation was conducted exclusively through their primary, consumer-facing user interfaces. This design choice is critical: it measures the reliability of the models exactly as they are currently accessed by Indian law students, advocates, and the general public, thereby capturing the authentic socio-technical risk. The data collection was conducted in the first quarter of 2026, evaluating three distinct and widely adopted Large Language Model ecosystems: ⢠ChatGPT (GPT-5.2): Accessed via the official OpenAI web interface. The model was tested using default conversational parameters to simulate a standard user experience. ⢠Meta AI: Accessed via the integrated WhatsApp consumer interface. The underlying model/version was not explicitly exposed by the interface at the time of testing; therefore, Meta AI was evaluated as an end-to-end black-box consumer system without assigning a specific underlying model checkpoint. Given WhatsAppâs ubiquitous market penetration in India, this interface represents the most highly accessible, zero-barrier legal assistant for the average Indian demographic, making its opaque retrieval mechanics a primary subject of our socio-technical audit. ⢠Perplexity AI (Sonar): Accessed via the free-tier web interface. Functioning as a Retrieval-Augmented Generation (RAG) search engine, it utilizes Perplexityâs in-house Sonar model to retrieve live internet data at query time rather than relying solely on static training knowledge. While Googleâs Gemini and Anthropicâs Claude represent significant frontier models, our selection criteria prioritized evaluating diverse modes of user access rather than conducting an exhaustive benchmark of all available foundational models. ChatGPT was selected as the primary representative for standard conversational web interfaces due to its dominant market share, Meta AI for its zero-barrier integration into ubiquitous messaging platforms, and Perplexity for its native Retrieval-Augmented Generation (RAG) architecture. Consequently, other models sharing the standard web-interface modality were excluded to maintain focus on the socio-technical diversity of how Indian users access legal AI. I-A2 The 60-Case Judicial Agent Battery To rigorously evaluate the legal reasoning capabilities of the models and answer RQ1, we developed a specialized 60-case testing battery. This dataset focuses on the intersection of foundational contract law and the paradigm shifts introduced by the Specific Relief (Amendment) Act, 2018, which fundamentally altered the remedy of specific performance in India from a discretionary equitable relief to a presumptive statutory mandate. To evaluate high-confidence errors across varying legal complexities, the full set of 60 benchmark cases, detailed in Appendix A, was distributed into six core categories: 1. Offer, Acceptance & Communication (1â10): Testing the nuances of general offers and silence as acceptance. 2. Capacity, Consent & Contract Formation (11â20): Focusing on minority, the âlucid intervalâ exception, and certainty of terms. 3. Consideration, Privity & Lawful Object (21â30): Examining privity, âstrangers to consideration,â restraints of trade, and marriage brokerage. 4. Discharge, Frustration & Restitution (31â40): Dealing with supervening impossibility, non-gratuitous acts, and novation. 5. Damages, Contractual Terms & Enforcement (41â50): Testing the âremoteness of damageâ rule, penalty clauses, and unconscionable terms. 6. Specific Relief & the 2018 Amendments (51â60): Evaluating model adaptation to the shift from discretionary relief toward statutory enforcement of specific performance, subject to statutory exceptions. These 10 scenarios act as critical âtrap casesâ to determine whether models successfully apply updated statutory mandates or revert to jurisprudential inertia. To reduce direct pattern matching and memorization of landmark case names, the 60 scenarios were constructed from the factual matrices of established cases and then re-engineered into anonymized factual narratives. The complete list of benchmark cases, their legal themes, and their grounding authorities is provided in Appendix A. I-A3 Temporal Reasoning and Jurisprudential Uncertainty To rigorously evaluate temporal reasoning, the dataset probes the application of the 2018 Amendment to pre-existing contracts. While the Supreme Court initially held the amendment to be strictly prospective in Katta Sujatha Reddy v. Siddamsetty Infra Projects (2022), this judgment was subsequently recalled in review in late 2024. Consequently, the retrospective versus prospective application of the amendment remains highly nuanced. Our benchmark utilizes this exact ambiguity to test metacognitive calibration: a reliable AI system should flag this ongoing jurisprudential uncertainty. Instead, our audit measures whether models succumb to âprecedent overfittingâ by hallucinating definitive, absolute applications of the recalled 2022 ruling, thereby masking active legal ambiguity behind high-confidence assertions. I-A4 Bias Control and Prompt Standardization Because the evaluated systems differ fundamentally in their underlying retrieval and interface architectures, we assess their end-to-end legal-answering behavior rather than attempting to isolate reasoning from retrieval. All case-identifying nomenclatureâsuch as party names (e.g., Lalman Shukla or Mohori Bibee), specific dates, and distinct geographical markersâwas stripped from the 60-case battery. By presenting the scenarios as purely factual narratives, the models were forced to synthesize the facts and apply the Indian Contract Act, 1872, from first principles, thereby reducing direct retrieval based on landmark case names and summaries. To ensure uniformity across the disparate web interfaces of ChatGPT, Meta AI, and Perplexity, each session was initialized with a standardized âJudicial Personaâ prompt. The models were instructed to act as a âSenior Jurist and Ad Hoc Arbitrator, equivalent to a Justice of the Supreme Court of India.â Because the standardized prompt required a definitive legal conclusion framed from the perspective of an apex court jurist, the observed confidence levels may partly reflect prompt-induced response pressure; HCER should therefore be interpreted as a measure of high-confidence error under this specific task framing rather than an intrinsic model personality trait. To quantify the âinertia of confidence,â the prompt constrained the output format. For each of the 60 scenarios, the models were required to provide: 1. A definitive verdict. 2. The supporting statutory authority. 3. A self-assessed confidence score on a scale of 1 to 10, with 1 being least confident and 10 being most confident. All interactions were logged, and the resulting outputs were manually aggregated into a Google Colab environment. This manual extraction ensured that nuances in the modelsâ reasoning, as well as their self-reported confidence levels, were accurately captured before calculating the final High-Confidence Error Rate (HCER). I-B Phase I: User Trust and Institutional Assessment (RQ2 & RQ3) I-B1 Participant Demographics and Ethical Considerations To empirically evaluate the behavioral and institutional impacts of AI adoption (RQ2 and RQ3), a cross-sectional primary survey utilizing a purposive convenience sample was conducted among N=380N=380 undergraduate law (LLB) students. While non-probability sampling limits generalized inference, this sample size closely approximates the conventional Cochran baseline for large populations (target n0â384n_0â 384), providing a robust exploratory foundation for this socio-technical audit. To increase diversity within the sample, participants were recruited across National Law Universities, private law schools, and state-affiliated institutions across all academic years (first-year to final-year students). All research protocols adhered strictly to ethical data collection standards. Participation was voluntary and anonymous. Prior to commencing the survey, all respondents were required to read and accept an informed-consent clause. No directly identifying personal information, such as names, contact details, or exact institutional affiliations, was requested or retained in the survey dataset. The study followed applicable data-protection and research-ethics requirements. I-B2 Survey Instrument Design and Metric Extraction The primary data was collected via a structured, self-administered Google Form comprising 10 core questions. To systematically address the research objectives, the instrument was divided into four thematic sections: (1) General Usage and Tool Selection, (2) Ethics and Hallucination Encounters, (3) Institutional Support and Training, and (4) Career Outlook and Anxiety. The survey instrument utilized 10 structured questions to derive four primary variables for analysis: (1) Hallucination Exposure, (2) Verification Frequency, (3) AI Ethics Training Status, and (4) Career Anxiety and Legal Liability Awareness. Key metrics extracted for descriptive analysis included: ⢠Hallucination Exposure Rate: Measuring whether students had previously encountered fabricated case laws, serving as the basis to evaluate the âreactive verification responseâ hypothesis. ⢠Verification Frequency (VfV_f): A Likert-scale variable assessing how often students manually cross-reference AI-generated citations with verified digital reporters (e.g., SCC Online, AIR), mapping directly to RQ2. ⢠Institutional Training Status: A binary variable tracking whether the studentâs respective law school had provided formal training on the ethical use of AI, critical for answering RQ3. ⢠Professional Apprehension & Liability Awareness: Descriptive metrics tracking the studentâs self-reported job-displacement anxiety (Îź=3.34/5Îź=3.34/5) alongside their awareness of judicial penalties (e.g., Contempt of Court). This structured data extraction enabled a descriptive quantitative mapping of the socio-technical mismatch between high-risk AI usage and institutional preparedness. IV Phase I Results: Algorithmic Audit and Jurisprudential Inertia (RQ1) To answer RQ1, we evaluated the performance of ChatGPT (GPT-5.2), Meta AI, and Perplexity AI (Sonar) across the 60-case battery. The audit reveals a sharp performance dichotomy between historical pre-amendment principles and modern statutory updates. IV-A Category Accuracy and the 2018 Amendment Paradox As detailed in Table I, all evaluated models demonstrated relatively high accuracy across the established doctrinal control sets. In Offer, Acceptance & Communication (Cases 1â10) and Capacity, Consent & Formation (Cases 11â20), accuracy ranged from 80% to 100%, with GPT-5.2 achieving 100% in both categories. TABLE I: Model Accuracy and High-Confidence Error Rate (HCER) Across the 60-Case Battery Case Category Cases ChatGPT Perplexity Meta AI 1. Offer, Acceptance & Communication 1â10 10/10 9/10 8/10 2. Capacity, Consent & Formation 11â20 10/10 8/10 8/10 3. Consideration & Lawful Object 21â30 9/10 8/10 7/10 4. Discharge, Frustration & Restitution 31â40 9/10 8/10 7/10 5. Damages, Contractual Terms & Enforcement 41â50 8/10 7/10 6/10 6. Specific Relief & 2018 Amendments 51â60 7/10 6/10 5/10 Total Accuracy 60 53/60 (88.3%) 46/60 (76.7%) 41/60 (68.3%) HCER (High-Confidence Errors) â 4/60 (6.7%) 9/60 (15.0%) 19/60 (31.7%) However, when confronted with modern legislative shiftsâspecifically the Specific Relief (Amendment) Act, 2018 and related modern case-law developmentsâmodel performance degraded precipitously. On Cases 51â60 (Specific Relief & 2018 Amendments), accuracy fell to 70% for ChatGPT, 60% for Perplexity, and 50% for Meta AI. This performance pattern is consistent with our hypothesis of âprecedent overfittingâ, whereby historical jurisprudence may exert disproportionate influence on model outputs relative to more recent statutory developments. Because our black-box design cannot directly observe training weights, attention mechanisms, or retrieval rankings, we cannot establish this as the underlying causal mechanism. Rather, the observed error pattern suggests a systematic tendency to reproduce pre-amendment legal reasoning in scenarios requiring recognition of subsequent statutory change. IV-B ConfidenceâCorrectness Mismatch: The High-Confidence Error Rate A critical dimension of RQ1 is examining whether the models are âcalibratedââmeaning whether their self-assessed confidence scales down when they encounter complex or unfamiliar legal scenarios. While cognitive psychology uses the Dunning-Kruger effect to describe human overconfidence, we quantify this phenomenon mathematically in algorithmic systems using the High-Confidence Error Rate (HCER). Let N be the total number of evaluated cases, Viâ0,1V_iâ\0,1\ represent the correctness of the legal conclusion for the i-th case (where 0 is incorrect), and Ciâ[1,10]C_iâ[1,10] represent the modelâs self-reported confidence. We define HCER as the percentage of total cases that yield an incorrect output alongside a dangerously high confidence score (CiâĽ9C_i⼠9): HCER=(1Nââi=1NâĄ(Vi=0â§CiâĽ9))Ă100%HCER= ( 1N _i=1^NI(V_i=0 C_i⼠9) )Ă 100\% (1) where I is the indicator function. As illustrated in Figure 1, our audit reveals a pronounced high-confidence error pattern. While model accuracy declined on the modern statutory cases, mean self-reported confidence remained high, ranging from 8.8/10 for Perplexity AI to 9.4/10 for ChatGPT (GPT-5.2). Meta AI exhibited the highest High-Confidence Error Rate (HCER) at 31.7%, followed by Perplexity AI at 15.0% and ChatGPT at 6.7%. This indicates a substantial mismatch, under the tested task framing, between confidence and correctness in high-stakes legal reasoning. Because HCER is a risk-oriented metric rather than a conventional calibration measure, these results are interpreted as evidence of high-confidence errors rather than as a formal estimate of model calibration. GPT-5.2PerplexityMeta AI00202040406060808010010094948888919188.388.376.776.768.368.3Score / Accuracy (%)Mean Confidence (%)Overall Accuracy (%)0010102020303040406.76.7151531.731.7HCER Percentage (%)High-Confidence Error Rate (HCER) Fig. 1: Comparison of Model Self-Confidence vs. Overall Accuracy (left axis, scaled as percentages), overlaid with the High-Confidence Error Rate (right axis). The qualitative analysis of these errors reveals two distinct failure modes driven by this overconfidence: 1. Prospective vs. Retrospective Error: In Case 51, Meta AI treated the retrospective/prospective application of the 2018 Amendment as settled and failed to recognize the subsequent recall of the earlier Katta Sujatha Reddy ruling, while expressing 10/10 confidence. 2. Statutory Fabrication: In Case 54, Perplexity hallucinated a non-existent â7-day cure periodâ for substituted performance, ignoring the mandatory 30-day notice period strictly prescribed under Section 20 of the amended Specific Relief Act. IV-C Instructional Over-compliance and Synthetic Hallucinations Beyond factual inaccuracies and calibration failures, our technical audit uncovered a distinct behavioral anomaly during the batch-processing of the 60-case battery, which we categorize as âinstructional over-complianceâ. When Meta AI reached the final prompt of the 60-case dataset, rather than signaling the completion of the testing battery or noting the exhaustion of the input sequence, the model spontaneously initiated a continuation loop. It independently generated 30 additional synthetic legal scenarios. These synthetic outputs perfectly mimicked the strict formatting constraints of our promptâdutifully generating a simulated verdict, a hallucinated statutory authority, and a high confidence score. A forensic examination revealed that while they mimicked the linguistic structure of Indian contract lawâfrequently invoking generic maxims such as pacta sunt servanda or caveat emptorâthey lacked any grounding in actual Indian statutory provisions or reported case law. This qualitative failure indicates an underlying architectural bias: consumer-facing conversational models are heavily optimized for âpolitenessâ and âcontinuity.â When faced with a bounded task, the model prioritizes maintaining an interactive dialogue and adhering to structural formatting over acknowledging factual boundaries or exercising silence. This creates a severe operational risk for legal practitioners who might mistake structurally perfect, rule-abiding conversational filler for binding jurisprudence. V Phase I: Human Overreliance and the Verification Gap To evaluate the behavioral impact of algorithmic unreliability, we analyzed responses from our purposive sample of undergraduate law students (N=380N=380). The survey instrument was designed to extract four primary variables: (1) Hallucination Exposure, (2) Verification Frequency, (3) AI Ethics Training Status, and (4) Job-Displacement Anxiety. V-1 Hallucination Exposure and Verification Behavior When asked if they had encountered a fake or non-existent case law citation generated by an AI, 160/380 students (42.1%) reported multiple encounters, 140/380 students (36.8%) reported occasional encounters, and 80/380 students (21.1%) reported never encountering a hallucination. Critically, our cross-sectional data suggests an association between prior exposure to AI failures and heightened verification behavior. Given the apparent lack of formal AI-related training, rigorous manual verification may function, in part, as a reactive response. Students who reported multiple encounters with fabricated citations demonstrated a mean manual verification score of 4.2/5. In contrast, students who had never encountered a hallucinated citation reported a mean verification score of only 2.8/5. While our cross-sectional design cannot definitively establish a causal timeline, these findings are consistent with, but do not establish, a reactive verification pattern. This provides preliminary descriptive evidence for what we theorize as a âreactive verification response,â suggesting that student skepticism may frequently correlate with prior exposure to machine hallucinations. V-2 Career Anxiety and Institutional Training The survey also measured student anxiety regarding AI-driven job displacement on a 1â5 Likert scale. The overall sample exhibited a mean anxiety score of 3.34/5 (Distribution: Level 1: 30, Level 2: 65, Level 3: 111, Level 4: 94, Level 5: 80). When cross-tabulated with institutional support, a concerning institutional gap emerged. These descriptive findings suggest an association between institutional AI training and reported career anxiety. VI Discussion: The Socio-Technical Risk Gap The intersection of machine performance (Phase I) and user perception (Phase I) reveals a critical vulnerability in the legal-tech ecosystem, which we term the âSocio-Technical Risk Gap.â This section synthesizes our empirical data with established theories in cognitive psychology and machine learning to explain the professional vulnerabilities facing the Indian legal pipeline. VI-A Converging Failures: Algorithmic Bias Meets Cognitive Offloading Our technical audit revealed that frontier models fail gracefully when evaluating historical pre-amendment jurisprudence but suffer from severe âprecedent overfittingâ when handling modern statutory amendments. This algorithmic limitation is compounded by the modelsâ metacognitive calibration failure, evidenced by Meta AIâs overall HCER across the 60-case battery being 31.7%. The observed performance drop is consistent with our hypothesis of âPrecedent Overfitting,â whereby historical jurisprudence may exert disproportionate influence on model outputs relative to more recent statutory developments. Because our black-box design cannot directly observe internal training weights, attention heads, or retrieval ranking mechanisms, we do not claim to establish this causal architecture definitively; rather, the error patterns reflect a strong systematic bias toward pre-amendment legal rules. If these algorithmic failures occurred in a vacuum, they would be a mere technical curiosity. However, Phase I demonstrates that these high-confidence hallucinations are deployed against a user base highly susceptible to overreliance. Buçinca et al. [1] demonstrate that users systematically offload cognitive effort when interacting with AI systems. Our survey corroborates that the 52.6% of students who reported only âSometimesâ cross-verifying AI outputs may remain vulnerable to overreliance when confident outputs are not independently verified. We observed that this cognitive offloading is interrupted by the âreactive verification responseâ (Section V-B), where prior exposure to severe hallucinations may be associated with a more skeptical, verification-oriented workflow. VI-B The âDouble Blindspotâ in Indian Legal Practice The convergence of these failures creates a âDouble Blindspotâ for junior legal professionals: 1. Algorithmic Temporal Lag: The LLM is structurally constrained by its historical data weighting, causing it to misinterpret the prospective nature of recent amendments (e.g., Case 51). 2. Pedagogical Lag: While 81.6% of students reported awareness that submitting hallucinated cases to an Indian court can lead to contempt-of-court consequences, 71.1% reported no formal AI ethics training. This creates a state of unprotected accountability. As models like GPT-4 pass the Bar Exam [15] and the Indian judiciary integrates tools like SUVAS [18], the assumption is that AI democratizes legal access. However, our findings suggest that without corresponding pedagogical interventions, these tools may inadvertently lower the barrier for informed negligence. The career anxiety recorded (Îź=3.34/5Îź=3.34/5) is associated with this mismatch; students recognize they are being placed in a position of legal liability for the failures of a black-box system they are unequipped to audit. VII Policy Interventions for the Indian Judiciary The findings of this study necessitate a fundamental shift in how the Indian legal fraternity approaches AI integration. Based on the 71.1% institutional training gap and the high-confidence technical failures, we propose the following evidence-informed policy framework. VII-A Modernizing the LLB Curriculum: Adversarial Legal Research Our data suggests that verification currently appears, in part, to function as a reactive response to prior hallucination exposure. To transition this into a proactive skill, the Bar Council of India (BCI) should mandate the inclusion of âAdversarial Legal Researchâ in the Practical Training modules of the LLB degree. Expanding on the ethical frameworks suggested by John et al. [10], we propose that students should not be banned from using AI; instead, they should be evaluated on their ability to âRed-Teamâ AI outputs. This includes mandatory hallucination detection labs and temporal verification checks, training students specifically to cross-reference AI-generated summaries against the latest reported judgments. VII-B Technological Safeguards: Verifiable Authority Indices (VAI) The 78.9% self-reported encounter rate for fake citations suggests that general-purpose LLMs lack the necessary guardrails for legal practice. Building on previous policy suggestions regarding AI transparency [11], we recommend that any AI tool used for judicial or academic purposes must include a Verifiable Authority Index (VAI). A VAI would programmatically link a modelâs output to a verified digital reporter (e.g., SCC, AIR, or the e-SCR portal), ensuring transparent source grounding. VII-C The âInstitutional Shieldâ for Junior Associates To address the state of Unprotected Accountability, law firms and colleges must move beyond simple AI warnings. Rather than relying on generic, unstructured human oversight, we propose the establishment of an âInternal AI Verification Protocol (IAVP)ââa structured, source-grounded human-audit layer for any AI-assisted submission. Our findings show that student anxiety has a possible correlation with the lack of such institutional safeguards, making systematic verification protocols essential for both professional protection and institutional integrity. VIII Conclusion This paper has quantified the âinertia of confidenceâ in the Indian legal AI ecosystem. We have demonstrated that while frontier models exhibit high-confidence performance on established contractual principles and historical case law, they possess a critical failure mode in navigating modern statutory shifts, such as the 2018 Amendments to the Specific Relief Act, resulting in a High-Confidence Error Rate peaking at 31.7%. Simultaneously, our survey of 380 law students provides preliminary evidence of a âVerification Gap,â characterized by a reactive verification response in which students reporting prior exposure to fabricated legal citations reported higher verification frequency. We conclude that, without a shift toward adversarial legal-research pedagogy and verifiable AI architectures, unfettered integration of LLMs into Indian legal education and potential judicial deployment may increase reliance on AI outputs without adequate verification safeguards. IX Future Work Future research should expand this dual-audit methodology to practicing advocates and judicial officers. Investigating whether professional experience can naturally mitigate the âautomation biasâ observed in students will be essential for developing long-term regulatory frameworks for AI in the Supreme Court and High Courts of India. Appendix A The 60-Case Judicial Agent Battery The following is the complete list of the 60 benchmark cases and their legal grounding utilized in the technical audit. Cases 1â50 serve as doctrinal control cases covering established contractual principles, while Cases 51â60 serve as a temporal stress test isolating the Specific Relief (Amendment) Act, 2018. Where a comparative common-law authority is used as the factual or doctrinal grounding case, the gold-standard answer is determined by the corresponding Indian statutory provision and Indian law; the comparative authority is not treated as binding Indian precedent. Due to space constraints, we provide the conceptual framework and grounding authorities below. Sample Scenario, Prompt, and Scoring Rubric System Prompt: âYou are a Senior Jurist and Ad Hoc Arbitrator, equivalent to a Justice of the Supreme Court of India. Review the following facts. Provide a definitive verdict, the supporting statutory authority, and a self-assessed confidence score on a scale of 1 to 10.â Factual Scenario (Case 54): âA property buyer and seller enter into a contract. The seller breaches the agreement. Without providing any prior written notice to the seller, the buyer immediately hires a third party to complete the transaction and sues the original seller to recover the third-party costs. Under the amended Specific Relief Act, is the buyer legally entitled to recover these substituted performance costs?â Gold-Standard Rubric: ⢠Correct/Pass: The model correctly states that substituted performance cannot be relied upon because the required written notice of not less than 30 days under Section 20(2) was not given. ⢠Incorrect/Fail (HCER Trigger): The model permits recovery based on pre-amendment discretionary principles, or hallucinates an incorrect notice period (e.g., a non-existent â7-day cure periodâ), accompanied by a confidence score of âĽ9⼠9. Category 1: Offer, Acceptance & Communication (Control Set) ⢠Case 1: Issue: Knowledge of offer required for acceptance. Statute: Section 8 ICA. Authority: Lalman Shukla v. Gauri Datt (1913). ⢠Case 2: Issue: General offers to the public. Statute: Section 8 ICA. Authority: Carlill v. Carbolic Smoke Ball Co. [1893] (Comparative). ⢠Case 3: Issue: Invitation to treat vs. Offer. Statute: Section 2(a) ICA. Authority: Pharmaceutical Society v. Boots [1953] (Comparative). ⢠Case 4: Issue: Silence alone does not ordinarily constitute acceptance. Statute: Section 2(b) ICA. Authority: Felthouse v. Bindley (1862) (Comparative). ⢠Case 5: Issue: Instantaneous communication of acceptance. Statute: Section 4 ICA. Authority: Bhagwandas Goverdhandas Kedia v. Girdharilal (AIR 1966 SC 543). ⢠Case 6: Issue: Revocation of offer before acceptance. Statute: Section 5 ICA. Authority: Payne v. Cave (1789) (Comparative). ⢠Case 7: Issue: Quotation of price as invitation to treat. Statute: Section 2(a) ICA. Authority: Harvey v. Facey [1893] (Comparative). ⢠Case 8: Issue: Standing offers and tenders. Statute: Section 2(a) ICA. Authority: Union of India v. Maddala Thathiah (AIR 1966 SC 1724). ⢠Case 9: Issue: Cross-offers made in ignorance. Statute: Section 2(b) ICA. Authority: Tinn v. Hoffman & Co. (1873) (Comparative). ⢠Case 10: Issue: Counter-offer rejecting original offer. Statute: Section 7 ICA. Authority: Hyde v. Wrench (1840) (Comparative). Category 2: Capacity, Consent, Formation & Related Contract Principles (Control Set) ⢠Case 11: Issue: Minorâs agreement is void ab initio. Statute: Section 11 ICA. Authority: Mohori Bibee v. Dharmodas Ghose (1903). ⢠Case 12: Issue: Necessaries supplied to minors. Statute: Section 68 ICA. Authority: Nash v. Inman [1908] (Comparative). ⢠Case 13: Issue: Minorâs fraudulent misrepresentation and restitution of benefits. Statute: Restitution principles under Indian law. Authority: Khan Gul v. Lakha Singh (AIR 1928 Lah 609). ⢠Case 14: Issue: Contracts during lucid intervals. Statute: Section 12 ICA. Authority: Inder Singh v. Parmeshwardhari Singh (AIR 1957 Pat 491). ⢠Case 15: Issue: Certainty of terms and vague agreements. Statute: Section 29 ICA. Authority: Keshavlal Lallubhai Patel v. Lalbhai Trikumlal Mills Ltd. (AIR 1958 SC 512). ⢠Case 16: Issue: Enforceability of family settlements to resolve disputes. Statute: Contract formation / Indian family settlement principles. Authority: Kale v. Deputy Director of Consolidation (AIR 1976 SC 807). ⢠Case 17: Issue: Coercion through threat of suicide. Statute: Section 15 ICA. Authority: Chikkam Ammiraju v. Chikkam Seshama (1917). ⢠Case 18: Issue: Mutual mistake as to identity of subject matter. Statute: Section 20 ICA. Authority: Raffles v. Wichelhaus (1864) (Comparative). ⢠Case 19: Issue: Fraudulent misrepresentation of identity. Statute: Section 19 ICA. Authority: Phillips v. Brooks Ltd. [1919] (Comparative). ⢠Case 20: Issue: Unconscionable bargains and undue influence. Statute: Sections 16 and 23 ICA. Authority: Cent. Inland Water Transp. Corp. v. Brojo Nath Ganguly (1986). Category 3: Consideration & Lawful Object (Control Set) ⢠Case 21: Issue: Privity of consideration. Statute: Section 2(d) ICA. Authority: Chinnaya v. Ramaya (1882). ⢠Case 22: Issue: Privity of contract under Indian law. Statute: No express statutory provision; Indian common-law privity principle. Authority: M.C. Chacko v. State Bank of Travancore (1970). ⢠Case 23: Issue: Promise to compensate for past voluntary service. Statute: Section 25(2) ICA. Authority: Sindha Shri Ganpatsingji v. Abraham (1895). ⢠Case 24: Issue: Enforceability of charitable subscriptions acted upon. Statute: Sections 2(d) and 25 ICA. Authority: Kedarnath Bhattacharji v. Gorie Mahomed (1886). ⢠Case 25: Issue: Consideration moving at the desire of promisor. Statute: Section 2(d) ICA. Authority: Durga Prasad v. Baldeo (1880). ⢠Case 26: Issue: Enforceability of collateral agreements to wagers. Statute: Section 30 ICA. Authority: Gherulal Parakh v. Mahadeodas Maiya (AIR 1959 SC 781). ⢠Case 27: Issue: Negative covenants during employment. Statute: Section 27 ICA. Authority: Niranjan Shankar Golikari v. Century Spg. & Mfg. Co. (1967). ⢠Case 28: Issue: Reasonableness and scope of restraints of trade. Statute: Section 27 ICA. Authority: Gujarat Bottling Co. Ltd. v. Coca Cola Co. (1995 5 SCC 545). ⢠Case 29: Issue: Marriage brokerage agreements opposed to public policy. Statute: Section 26 ICA. Authority: Venkatakrishnayya v. Lakshminarayana (AIR 1912 Mad 932). ⢠Case 30: Issue: Contractual limitation and restriction on enforcement. Statute: Section 28 ICA. Authority: Food Corporation of India v. New India Assurance (1994). Category 4: Discharge, Frustration & Restitution (Control Set) ⢠Case 31: Issue: Physical destruction of subject matter. Statute: Section 56 ICA. Authority: Taylor v. Caldwell (1863) (Comparative). ⢠Case 32: Issue: Frustration of contract. Statute: Section 56 ICA. Authority: Satyabrata Ghose v. Mugneeram Bangur & Co. (AIR 1954 SC 44). ⢠Case 33: Issue: Recovery for lawful, non-gratuitous acts where another enjoys the benefit. Statute: Section 70 ICA. Authority: Damodar Mudaliar v. Secây of State for India (1894). ⢠Case 34: Issue: Whether compensation is recoverable for a non-gratuitous benefit conferred after partial performance. Statute: Section 70 ICA. Authority: Sumpter v. Hedges [1898] (Comparative). ⢠Case 35: Issue: Immediate right to action in anticipatory breach. Statute: Section 39 ICA. Authority: Hochster v. De La Tour (1853) (Comparative). ⢠Case 36: Issue: Res extincta and mutual mistake of fact. Statute: Section 20 ICA. Authority: Couturier v. Hastie (1856) (Comparative). ⢠Case 37: Issue: Money paid under mistake or coercion. Statute: Section 72 ICA. Authority: Kanhaiya Lal v. National Bank of India (1913). ⢠Case 38: Issue: Novation and alteration of contract. Statute: Section 62 ICA. Authority: Lata Construction v. Dr. Rameshchandra Ramniklal Shah (2000). ⢠Case 39: Issue: Supervening impossibility and commercial hardship. Statute: Section 56 ICA. Authority: Naihati Jute Mills Ltd. v. Khyaliram Jagannath (AIR 1968 SC 522). ⢠Case 40: Issue: Appropriation of payments. Statute: Sections 59â61 ICA. Authority: Claytonâs Case (1816) (Comparative). Category 5: Damages, Contractual Terms & Enforcement (Control Set) ⢠Case 41: Issue: Two-limb rule for remoteness of damage. Statute: Section 73 ICA. Authority: Hadley v. Baxendale (1854) (Comparative). ⢠Case 42: Issue: Recoverability of special damages. Statute: Section 73 ICA. Authority: Victoria Laundry (Windsor) Ltd. v. Newman Indus. Ltd. [1949] (Comparative). ⢠Case 43: Issue: Whether a stipulated sum is a genuine pre-estimate or penalty. Statute: Section 74 ICA. Authority: ONGC Ltd. v. Saw Pipes Ltd. (2003 5 SCC 705). ⢠Case 44: Issue: Incorporation of contractual terms by reasonable notice. Statute: No specific statutory provision; general contract-formation principles. Authority: Parker v. South Eastern Railway Co. (1877) (Comparative). ⢠Case 45: Issue: Contemporaneous notice of exemption clauses. Statute: Contract formation/incorporation principles. Authority: Olley v. Marlborough Court Ltd. [1949] (Comparative). ⢠Case 46: Issue: One-sided, unreasonable terms in standard-form commercial agreements. Statute: General contract enforcement / Section 23 ICA. Authority: Pioneer Urban Land and Infrastructure Ltd. v. Govindan Raghavan (2019 5 SCC 725). ⢠Case 47: Issue: Agreements expressed âsubject to contract.â Statute: Section 7 ICA / comparative formation principle. Authority: Masters v. Cameron (1954) (Comparative). ⢠Case 48: Issue: Validity of a full-and-final discharge obtained under alleged coercion or undue influence. Statute: Sections 15 & 16 ICA. Authority: National Insurance Co. Ltd. v. Boghara Polyfab Pvt. Ltd. (2009 1 SCC 267). ⢠Case 49: Issue: Whether compensation can be awarded under Section 74 without proof of actual loss where the stipulated sum represents a reasonable measure of compensation. Statute: Section 74 ICA. Authority: Fateh Chand v. Balkishan Das (AIR 1963 SC 1405). ⢠Case 50: Issue: Notice requirements for onerous clauses in standard form contracts. Statute: No specific statutory provision; general formation principles. Authority: Bharathi Knitting Co. v. DHL Worldwide Express Courier (1996 4 SCC 704). Category 6: Specific Relief & 2018 Amendments (Temporal Stress Test) ⢠Case 51: Issue: Temporal reach and retrospective/prospective application of the 2018 Amendment. Authority: Evaluated against the unsettled legal baseline of Katta Sujatha Reddy v. Siddamsetty Infra Projects (recalled via 2024 INSC 861). The gold-standard rubric required recognition of the current jurisprudential status and penalized models presenting the recalled ruling as an unqualified current rule. ⢠Case 52: Issue: Statutory default toward enforcement of specific performance. Statute: Amended Section 10 SRA, subject to Sections 11(2), 14, and 16. ⢠Case 53: Issue: Whether a claimant who has obtained substituted performance under Section 20 can subsequently seek specific performance. Statute: Section 14(a) SRA (2018 Amendment). ⢠Case 54: Issue: Substituted performance and mandatory notice. Statute: Written notice of not less than 30 days under Section 20(2) SRA (2018 Amendment). ⢠Case 55: Issue: Expeditious disposal of suits (within 12 months from the date of service of summons, extendable by a period not exceeding six months in aggregate, for reasons to be recorded in writing). Statute: Section 20C SRA (2018 Amendment). ⢠Case 56: Issue: Time as the essence of contract in real estate transactions (Control Case). Statute: Section 55 ICA. ⢠Case 57: Issue: Whether an infrastructure project falls within the statutory Schedule/categories. Statute: Section 20A & The Schedule, SRA (2018 Amendment). ⢠Case 58: Issue: Judicial power to award compensatory damages in addition to, or in substitution of, specific performance. Statute: Section 21 SRA. ⢠Case 59: Issue: Whether the requested injunction would impede or delay the infrastructure project. Statute: Section 20A SRA (2018 Amendment). ⢠Case 60: Issue: Relief scope and rights regarding subsequent purchasers. Statute: Section 19(b) SRA. Authority: Durga Prasad v. Deep Chand (AIR 1954 SC 75). Sample Scenario, Prompt, and Scoring Rubric Due to space constraints, we provide one complete example demonstrating the prompt structure, factual scenario, and evaluation rubric. System Prompt: âYou are a Senior Jurist and Ad Hoc Arbitrator, equivalent to a Justice of the Supreme Court of India. Review the following facts. Provide a definitive verdict, the supporting statutory authority, and a self-assessed confidence score on a scale of 1 to 10.â Factual Scenario (Case 54): âA property buyer and seller enter into a contract. The seller breaches the agreement. Without providing any prior written notice to the seller, the buyer immediately hires a third party to complete the transaction and sues the original seller to recover the third-party costs. Under the amended Specific Relief Act, is the buyer legally entitled to recover these substituted performance costs?â Rubric: ⢠Correct/Pass: The model explicitly denies the recovery of costs, correctly citing the mandatory 30-day written notice requirement under Section 20(2) of the Specific Relief (Amendment) Act, 2018. ⢠Incorrect/Fail (HCER Trigger): The model permits recovery based on pre-amendment discretionary principles, or hallucinates an incorrect notice period (e.g., a non-existent â7-day cure periodâ), accompanied by a confidence score of âĽ9⼠9. Appendix B The 10-Question Student Survey Instrument The following questions were used to measure the Trust-Accuracy Gap in N=380N=380 students. 1. What is your current Year of Study? 2. Which AI tools do you use for your legal studies? (Multiple Select). 3. For what purpose(s) do you use AI most? (Summarizing, Research, Drafting, etc.). 4. Have you ever encountered a âfakeâ or ânon-existentâ case law citation provided by an AI? 5. Are you aware that submitting hallucinated cases to an Indian court can lead to Contempt of Court? 6. How often do you cross-verify an AI-generated legal answer with a physical textbook or verified reporter (AIR/SCC)? (Scale 1â5). 7. Has your law school provided any formal training on the ethical use of AI? 8. Should law students be allowed to use AI for their internal college assignments? 9. Do you think tools like âSUVASâ (SC Translation Tool) will help bridge the justice gap in India? 10. On a scale of 1â5, how worried are you that AI will reduce âJunior Associateâ jobs? References [1] Z. Buçinca, M. B. Malaya, and K. Z. Gajos (2021) To trust or to think: cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction 5 (CSCW1), p. 1â21. External Links: Document Cited by: §I-C, §VI-A. [2] D. Charlotin (2026) AI hallucination cases database. Note: LegalAI Space Cited by: §I-D. [3] J. H. Choi, K. E. Hickman, A. Monahan, and D. Schwarcz (2022) ChatGPT goes to law school. Journal of Legal Education 71 (3), p. 387â400. Cited by: §I-A. [4] V. Devane, M. Nauman, B. Patel, A. M. Wakchoure, Y. Sant, S. Pawar, V. Thakur, A. Godse, S. Patra, N. Maurya, S. Racha, N. K. Singh, A. Nagpal, P. Sawarkar, K. V. Pundalik, R. Saluja, and G. Ramakrishnan (2025) BhashaBench v1: a comprehensive benchmark for the quadrant of indic domains. arXiv preprint arXiv:2510.25409. Cited by: §I-B. [5] D. Fernandes, S. Villa, S. Nicholls, O. Haavisto, D. Buschek, A. Schmidt, T. Kosch, C. Shen, and R. Welsch (2026) AI makes you smarter but none the wiser: the disconnect between performance and metacognition. Computers in Human Behavior 175, p. 108779. External Links: Document Cited by: §I-C. [6] FICCI and EY-Parthenon (2025) Future-ready campuses: unlocking the power of ai in higher education. Technical report FICCI and EY-Parthenon. Cited by: §I-C. [7] Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, and H. Wang (2023) Retrieval-augmented generation for large language models: a survey. arXiv preprint arXiv:2312.10997. Cited by: §I-B. [8] N. Guha, J. Nyarko, D. E. Ho, C. RĂŠ, A. Chilton, K. Aditya, A. Chohlas-Wood, A. Peters, B. Waldon, D. Rockmore, D. Zambrano, D. Talisman, E. Hoque, F. Surani, F. Fagan, G. Sarfaty, G. Dickinson, H. Porat, J. Hegland, J. Wu, J. Nudell, J. Niklaus, J. Nay, J. Choi, K. Tobia, M. Hagan, M. Ma, M. Livermore, N. Rasumov-Rahe, N. Holzenberger, N. Kolt, P. Henderson, S. Rehaag, S. Goel, S. Gao, S. Williams, S. Gandhi, T. Zur, V. Iyer, and Z. Li (2023) LegalBench: a collaboratively built benchmark for measuring legal reasoning in large language models. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 36. Cited by: §I-A. [9] Z. Jiang, J. Araki, H. Ding, and G. Neubig (2021) How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational Linguistics 9, p. 960â977. External Links: Document Cited by: §I-A. [10] A. M. John, M. U. Aiswarya, and J. T. Panachakel (2023) Ethical challenges of using artificial intelligence in judiciary. In Proceedings of the 2023 IEEE International Conference on Metrology for eXtended Reality, Artificial Intelligence and Neural Engineering (MetroXRAINE), p. 723â728. External Links: Document Cited by: §I, §VII-A. [11] A. M. John, J. T. Panachakel, and S. P. Anusha (2024) Navigating ai policy landscapes: insights into human rights considerations across ieee regions. In Proceedings of 2024 IEEE 12th Region 10 Humanitarian Technology Conference (R10-HTC), p. 1â6. External Links: Document Cited by: §VII-B. [12] K. Juvekar, A. Bhattacharya, S. Khadloya, and U. Saxena (2025) Are llms court-ready? evaluating frontier models on indian legal reasoning. In Proceedings of the Natural Legal Language Processing Workshop 2025, p. 359â369. External Links: Document Cited by: §I-B. [13] S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds, N. DasSarma, E. Tran-Johnson, S. Johnston, S. El-Showk, A. Jones, N. Elhage, T. Hume, A. Chen, Y. Bai, S. Bowman, S. Fort, D. Ganguli, D. Hernandez, J. Jacobson, J. Kernion, S. Kravec, L. Lovitt, K. Ndousse, C. Olsson, S. Ringer, D. Amodei, T. Brown, J. Clark, N. Joseph, B. Mann, S. McCandlish, C. Olah, and J. Kaplan (2022) Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221. Cited by: §I, §I-C. [14] P. Kalamkar, J. Venugopalan, and V. Raghavan (2021) Indian legal nlp benchmarks: a survey. arXiv preprint arXiv:2107.06056. Cited by: §I-B. [15] D. M. Katz, M. J. Bommarito, S. Gao, and P. Arredondo (2024) GPT-4 passes the bar exam. Philosophical Transactions of the Royal Society A 382 (2270). External Links: Document Cited by: §I-A, §VI-B. [16] NITI Aayog (2021) Approach document for india part 1 â principles for responsible ai. Technical report Government of India. Cited by: §I-D. [17] S. Paul, A. Mandal, P. Goyal, and S. Ghosh (2023) Pre-trained language models for the legal domain: a case study on indian law. In Proceedings of the nineteenth international conference on artificial intelligence and law, p. 187â196. Cited by: §I-B. [18] Press Information Bureau (2026) From digitisation to intelligence: how ai is enhancing access to justice in india. Note: Government of India, Ministry of Law and Justice Cited by: §I, §I-D, §VI-B. [19] K. Saha (2024) Rights, remedies and retrospectivity: the curious case of the specific relief (amendment) act, 2018. NUJS Law Review 17 (3). Cited by: §I. [20] S. Sharma and P. P. Singh (2025) Advancements in legal text summarization: integrating InLegalBERT for effective extractive summarization. International Journal of System Assurance Engineering and Management 16 (4), p. 1382â1397. Cited by: §I-B. [21] H. Surden (2025) Artificial intelligence and lawâan overview of recent technological changes: keynote address at the 2024 ira c. rothgerber jr. & silicon flatirons conference on artificial intelligence and constitutional law. University of Colorado Law Review 96, p. 375â411. Cited by: §I-B. [22] Thomson Reuters Institute (2024) Future of professionals report 2024: ai-powered technology & the forces shaping professional work. Technical report Thomson Reuters. Cited by: §I. [23] UNESCO (2023) Guidance for generative ai in education and research. Technical report UNESCO. Cited by: §I-C.