Paper deep dive
Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases
Amogh Raina, Ilias Chalkidis, Daniel Hershcovich, Henrik Palmer Olsen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/19/2026, 4:21:03 AM
Summary
This study evaluates the legal reasoning capabilities of OpenAI GPT-5.4 on European Court of Human Rights (ECtHR) cases regarding Article 10 (Freedom of Expression). Using a dataset of 30 cases, the authors compare three prompting strategies: non-curated, expert-curated step-by-step instructions, and guide-based inference. Results indicate that while GPT-5.4 produces structurally complete reasoning, it is substantively shallow compared to human experts. Crucially, the study finds that LLM-as-a-Judge evaluators are reliable but do not align well with human annotators, and improved reasoning quality via expert prompts does not necessarily correlate with higher predictive accuracy. The authors caution against using automated evaluation or task accuracy as proxies for legal reasoning quality.
Entities (8)
Relation Signals (9)
GPT-5.4 â evaluatedon â European Court of Human Rights
confidence 95% ¡ We evaluate OpenAI GPT 5.4... using legal cases from the European Court of Human Rights (ECtHR) as a testbed.
Expert-Curated Prompt â doesnotimprove â Predictive Accuracy
confidence 93% ¡ which does not result in more accurate predictions compared to the other examined settings.
Expert-Curated Prompt â leadsto â Comprehensive Reasoning
confidence 93% ¡ Overall, the expert-curated prompt leads to more comprehensive reasoning
LLM-as-a-judge â usedfor â GPT-5.4
confidence 92% ¡ We present our findings derived from assessing the model's responses with both human and LLM evaluation... LLM-as-a-Judge evaluators are internally consistent yet align only weakly with our trained annotators
European Court of Human Rights â publishesin â HUDOC
confidence 90% ¡ All cases have been officially published in the HUDOC database
GPT-5.4 â specializesin â Article 10
confidence 90% ¡ We focus on a short, targeted study of ECtHR cases related to Article 10 (Freedom of expression)... GPT-5.4... performs across different prompting settings
LLM-as-a-judge â evaluatedby â Claude Opus 4.7
confidence 85% ¡ We instruct three top-tier models (GPT-5.5, DeepSeek V4 Pro, and Claude Opus 4.7) to follow the same evaluation protocol
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Reasoning has become a standard technique and feature for contemporary LLMs; however, its application and quality in the context of demanding legal-oriented tasks, such as legal case forecasting, remain under explored. We investigate how LLMs reason in the context of legal case forecasting, using legal cases from the European Court of Human Rights (ECtHR) as a testbed. We evaluate OpenAI GPT 5.4, a recent top-tier LLM, by exploring alternative prompting strategies that are more or less suggestive of what counts as legally meaningful reasoning in the context of ECtHR jurisprudence. We present our findings derived from assessing the model's responses with both human and LLM evaluation. We find that the examined model scores far from ideal in legal reasoning, the model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet align only weakly with our trained annotators, i.e., reliable but not a valid substitute for human evaluation. Overall, the expert-curated prompt leads to more comprehensive reasoning, which does not result in more accurate predictions compared to the other examined settings. Based on our findings, we urge the community not to rely solely on automated LLM-based evaluation and to avoid using task accuracy as an appropriate proxy for reasoning quality.
Tags
Links
- Source: https://arxiv.org/abs/2608.17168v1
- Canonical: https://arxiv.org/abs/2608.17168v1
Trouble viewing inline? Open PDF directly â
Full Text
110,629 characters extracted from source content.
Expand or collapse full text
Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases Amogh Raina 1 Ilias Chalkidis 1 Daniel Hershcovich 1 Henrik Palmer Olsen 2 Abstract Reasoning has become a standard technique and feature for contemporary LLMs; however, its application and quality in the context of demanding legal-oriented tasks, such as legal case forecasting, remain underexplored. We investigate how LLMs reason in the context of legal case forecasting, using legal cases from the European Court of Human Rights (ECtHR) as a testbed. We evaluate OpenAI GPT 5.4, a recent top-tier LLM, by exploring alternative prompting strategies that are more or less sug- gestive of what counts as legally meaningful reasoning in the context of ECtHR jurisprudence. We present our find- ings derived from assessing the modelâs responses with both human and LLM evaluation. We find that the exam- ined model scores far from ideal in legal reasoningâthe model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet align only weakly with our trained annotators, i.e., reliable but not a valid substi- tute for human evaluation. Overall, the expert-curated prompt leads to more comprehensive reasoning, which does not result in more accurate predictions compared to the other examined settings. Based on our findings, we urge the community not to rely solely on automated LLM-based evaluation and to avoid using task accuracy as an appropriate proxy for reasoning quality. 1. Introduction Large Language Models (LLMs) that employ reasoning- often referred to as Reasoning Language Models (RLMs)- have demonstrated strong performance across a wide range of complex instruction-following and problem-solving tasks, including mathematics and coding (OpenAI, 2024; Guo et al., 2025). LLM reasoning has been a critical research * Equal contribution 1 Department of Computer Science, Uni- versity of Copenhagen, Denmark 2 Faculty of Law, University of Copenhagen, Denmark. Correspondence to: Amogh Raina <amogh.raina@di.ku.dk>. Proceedings of the42 nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025. Copyright 2025 by the author(s). topic and so far has been substantially benchmarked in do- mains like logic, math, and code; nonetheless, its potential application to complex, legal-oriented tasksâsuch as legal case forecasting (judgment)âremains heavily understudied. In such tasks, reasoning is not merely a path to better predic- tive accuracy but a lens into the modelâs decision-making, which is an explainability factor beyond brute-force pattern matching. Recent industry surveys indicate that a substan- tial proportion of legal professionals, estimated between 60% and 70%, now report using generative AI tools at least weekly for drafting, research, or document analysis. Nee- lamegam & Nirmala (2025) reported that LLMs seem to excel at legal text classification and summarisation tasks, but struggle with tasks requiring structured legal reasoning. Legal case forecasting (judgment) is inherently reliant on sophisticated reasoning that differs across fields of law, e.g., criminal law, contract/commercial law, and human rights law, and jurisdictions, e.g., national or supranational, such as ECtHR. Analyzing and interpreting case facts, connect- ing them to the modelâs internal legal knowledge or other input sources, and making predictive decisions all require an advanced reasoning capability. In this project, we investigate how LLMs reason in the con- text of legal case forecasting, using legal cases from the European Court of Human Rights (ECtHR) as a small-scale testbed: as in prior work (Aletras et al., 2016; Chalkidis et al., 2019; T.y.s.s et al., 2022), the LLM assesses the facts and alleged violations and reasons about how the ECtHR would rule for or against the applicant. The general ques- tion we want to answer is: Can LLMs reason in a legally meaningful manner? We approach this general question by addressing the follow- ing research questions: RQ1:How competent is a recent top-tier LLM, such as Ope- nAI GPT 5.4, in legal reasoning on ECtHR cases? RQ2:How do different prompting strategies, i.e., the inclu- sion (or not) of relevant context, affect the modelâs reasoning performance? RQ3:What is the quality of reasoning as assessed by human annotators and models (LLM-as-a-Judge)? And how does it differ among them? 1 arXiv:2608.17168v1 [cs.CL] 17 Aug 2026 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? Figure 1: Cartoonish abstract depiction of the real-life ECtHR case workflow from case lodging to the courtâs judgment and its publishing; and of our LLM-driven legal judgment-forecasting experiment. We focus on a short, targeted study of ECtHR cases related to Article 10 (Freedom of expression) of the European Con- vention of Human Rights (ECHR), which reads as follows: 1. Everyone has the right to freedom of expression. This right shall include freedom to hold opinions and to receive and impart information and ideas without inter- ference by public authority and regardless of frontiers. This Article shall not prevent States from requiring the licensing of broadcasting, television or cinema enter- prises. 2. The exercise of these freedoms, since it carries with it duties and responsibilities, may be subject to such formalities, conditions, restrictions or penalties as are prescribed by law and are necessary in a democratic society, in the interests of national security, territorial integrity or public safety, for the prevention of disor- der or crime, for the protection of health or morals, for the protection of the reputation or rights of others, for preventing the disclosure of information received in con- fidence, or for maintaining the authority and impartiality of the judiciary. Contributions (a) We release a small dataset compris- ing 30 recent European Court of Human Rights (ECtHR) cases meant to assess LLMsâ capabilities in legal judgment forecasting and reasoning (Section 2). All cases have been officially published in the HUDOC database from April 2025 onwards. 1 1 We use a small dataset of recent cases to reduce (though not fully eliminate) overlap with the examined modelâs training data; in early experiments, we found that models memorize cases, especially popular ones that have been heavily discussed. As the modelâs knowledge cut-off (August 31, 2025) postdates part of our corpus, we discuss the residual contamination risk in the Limitations section. (b) We explore how the most recent OpenAI GPT 5.4 model (Knowledge cut-off: Aug 31, 2025) performs across differ- ent prompting settings: (a) out-of-the-box with minor task guidance, or (b) under different settings where the prompt includes the ideal reasoning strategy step-by-step descrip- tion explicitly (verbatim), or implicitly, providing the official multi-page guide related to the applicable article. (c) We evaluate the modelâs generated responses with a human evaluation (by law students) and models (LLM-as-a- Judge), focusing on the quality of the modelâs assessment (reasoning), rather than the modelâs decision (prediction). We also release, as part of our dataset (point a), the results of the human and LLM-based evaluations for future explo- ration or comparison with other LLM-as-a-Judge schemes. 2. Examined Task and Data The ECtHR hears allegations that Council of Europe mem- ber states (the defendant state(s)) have breached the Eu- ropean Convention of Human Rights (ECHR). Applicants lodge an application; if admissible, the court examines the case and rules on a decision, which is published and serves as precedent (Figure 1). We experiment with 30 ECtHR cases from the HUDOC database, related to Article 10 (Freedom of expression). For each, the dataset provides the facts as presented in the decision (the factual events and relevant legal framework); Article 10 was allegedly violated by the defendant state(s), and the court decided whether it was indeed violated or not. As in T.y.s.s et al. (2022), the model receives the case facts and must reason and predict the (non-)violation of the ar- ticle. 2 It generates both an assessment (reasoning) and a 2 This is different from prior work (Chalkidis et al., 2021), since 2 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? decision (Section 3.2). Dataset ProcessingTo facilitate our experiments, we per- formed several pre-processing steps to extract the necessary information from the officially published courtâs judgment. First, we extract the facts of the case, as represented in the courtâs decision, alongside the presented relevant legal framework, which summarizes the state of national lawâ of the defendant states(s)ârelated to the considerations of the examined case; this is used as the factual basis (input) provided to the model. The average length is 4,993 words. Furthermore, we extract the courtâs assessment, the part where the court reasons on how the facts should be inter- preted under the law, covering interference, lawfulness, legit- imacy, and proportionality, which we treat as our reference of expert reasoning. We keep only the Article-10 merits (excluding admissibility and out-of-scope articles) and ex- tract its case-law references (e.g., â(See Axel Springer AG v. Germany [GC], no. 39954/08)â) and factual references (e.g., â(See paragraphs 12-14 above)â). On average, the assessment includes 25 factual and 10â15 case-law refer- ences. 2.1. Task Simplifications The formulation of the task in our experiment does not replicate the courtâs flow in detail by any means. We follow a thin experimental setup, where, given curated facts, the models have to reason and predict the potential violations of the examined ECHR article. Here are the main differences: Input: The input is the curated, summarised facts as written in the courtâs decision, i.e., the facts as the Court sees them after hearing both parties. The court itself has access to broader, unfiltered case files (factual background, partiesâ arguments, evidence); see Medvedeva et al. (2021) for how facts are established. Background: The models do not receive ECHR case law (precedent) verbatim.Rather than retrieve it via a RAG systemâwhich risks cascading failures into the assessmentâwe provide the list of case-law references from the courtâs assessment as an oracle and ask the model to place them appropriately to support its arguments. Output: In our experiments, the model produces the le- gal reasoning relevant to, and a prediction of the potential violation of the examined ECHR article (Article 10). In contrast, the ECtHRâs ruling is more granular. The court considers the admissibility of a case as an initial pre-trial process and rules on potential damages, costs, and expenses that the defendant state must cover as reparations when the applicant (plaintiff) wins the case. Our model output does the model is âawareâ of the ECHR article under discussion, as is the court ruling the cases in real life. not include any of these components. Despite all the aforementioned limitations, our framework is still the closest to the real judicial process followed by the Court, compared to the literature (Aletras et al., 2016; Chalkidis et al., 2019; T.y.s.s et al., 2022), since our input includes the relevant legal framework, while we also re- quest the modelâs assessment (reasoning) as a crucial part of supporting the modelâs decision that we later evaluate. 2.2. Reasoning Desiderata Qualities of legal reasoning: Appropriate legal reasoning in the context of our work is considered one that (a) is grounded on the case facts, i.e., the interpretation of the facts is part of the reasoning, and (b) is relevant to the articles considered, i.e., the reasoning evolves around the application of the examined article(s) to the case facts in a legal-meaningful matter, i.e., following ECtHR standards as set out in its jurisprudence. Primacy of legal reasoning: Appropriate legal reasoning prevails over predictive accuracy. A wrong prediction rela- tive to the courtâs decision can arise from genuine ambiguity, narrative complexity, lack of context (Section 2.1), or subjec- tivity (Xu et al., 2023); appropriate reasoning, by contrast, is the legally meaningful âvirtueâ that legitimises juridical bodies in democratic societies. 3. Experimental Methodology 3.1. Examined Model We experiment with OpenAIâs GPT-5.4 model (OpenAI, 2026), the second-most capable LLM in the GPT series at the time, with a knowledge cutoff of August 31, 2025. 3 We run the generator with medium reasoning effort (exact per-setting values in Appendix A). We focus on ChatGPT since it remains the most widely used model among legal professionals, ahead of other frontier LLM providers, mak- ing it the most representative choice for studying how legal practitioners actually encounter LLM-generated content. 3.2. Prompting Settings We explore 3 alternative prompting settings: 4 (a) Non-Curated, where the model is instructed to assess and decide on a given case based on the facts and to format its output in a specific fashion, i.e., provide its assessment as a list of titled paragraphs alongside its decision (prediction) regarding the alleged violation (yes or no), but we do not 3 GPT-5.5, which we use as an evaluator (LLM-as-a-Judge), was released in the latter stage of our project. 4 In Appendix B, we provide the instructions (prompts) for all settings verbatim. 3 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? Figure 2: Appropriate reasoning strategy to assess ECHR cases related to Article 10 (Freedom of Expression). provide any explicit information related to the appropriate reasoning strategy for the examined article. (b) Curated (Expert), where we provide a detailed step- by-step description of the appropriate reasoning strategy (Figure 2) related to the examined articleâfollowing EC- tHR standards, curated by our legal expert. In this setting, the model is instructed to follow and apply the predefined reasoning strategy, and we expect a top-tier highly capable model to do this effectively. (c) Curated (Guide), where we provide the official guide of the examined article (Article 10), which is part of the series of Case-Law Guides published by the European Court of Human Rights. 5 This is a two-step process: (i) we first instruct the model explicitly to infer the appropriate reason- ing strategy based on the guideâs content; and (i) we use the generated reasoning strategy as part of the follow-up prompt, where we request the model to assess and decide on a given case by following and applying the inferred strategy. Exact API identifiers, version numbers, and reasoning-effort settings for the generator and all evaluators are reported in Appendix A. 3.3. Reasoning Evaluation We evaluate the formally disclosed case assessment (reason- ing) in the modelâs output (evaluator models are run with high reasoning effort). We ignore the âhiddenâ reasoning (or âthinkingâ) that precedes the final answer, treating it as a scratchpad where the model explores the task, analogous to a studentâs notes during a test, or the Courtâs undisclosed deliberations. Human Evaluation As mentioned (Section 2.2), our main goal is to assess the quality of the modelâs assessment (rea- soning), rather than its decision (prediction). To do so, we recruit three senior law students, each with extensive train- ing in ECtHR jurisprudence and working under an expertâs supervision. For each case, we present to every annotator: (a) the original facts and the courtâs assessment as presented in the Courtâs judgment, and (b) the modelâs generated 5 You can find the official guide here. assessment (response) for two out of the three examined set- tings (a-c). The three annotators independently rate the same set of cases, which lets us report inter-annotator agreement and humanâmodel alignment (Section 4.5). Each annotator is asked to assess (annotate) the model re- sponses based on the following criteria per model response: 1.Step occurrence: If each of the 4 steps of the ideal assessment (reasoning) strategy (Figure 2) occursâis presentâin the modelâs generated assessment. This is a binary label (yes or no). This criterion assesses whether the model follows a step, regardless of quality. 2.Step comprehensiveness: The comprehensiveness of each of the occurred steps in each response in a Likert- scale (1-5) compared to the courtâs assessment. If the expected step did not occur in the modelâs response, the score is 0. This criterion captures the overall quality of the reasoning step (Section 2.2). 3.Overall conciseness: If the overall modelâs assessment for each response is concise, i.e., whether it contains no irrelevant information in a Likert-scale (1-5). This criterion captures a supplementary aspect of quality. This is a mixed evaluation setting, where each annotator assesses the quality of each model response given the ground truth (the courtâs assessment), but also implicitly ranks the two generated responses per step for the examined stepâs comprehensiveness, since both are presented in parallel. This allows us to consider individualâper settingâscores, but also pairwise comparisons, i.e., the âwin rateâ of one setting over the other. Sampling and comparability. Sampling and comparabil- ity. Each case is shown as a pair of two of the three settings, fully balanced: each of the three pairs (AâB, AâC, BâC) is used for exactly 10 cases, so every setting is assessed on 20 cases and paired equally often with each other, with display order randomized. In total, the human evaluation covers 22 distinct cases as 30 (case, pair) instances (some cases are reused across pairs to keep the balance), all rated by the three annotators. This balance makes the per-setting 4 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? Prompt Setting Step OccurrenceStep ComprehensivenessOverall Conciseness Fact. Ref. F1Accuracy HumanModelsHumanModelsHumanModels A: Non-Curated0.96Âą 0.190.92Âą 0.283.95Âą 1.083.61Âą 1.443.87Âą 0.753.83Âą 0.940.51Âą 0.190.82 B: Curated (Expert)1.00Âą 0.061.00Âą 0.004.45Âą 0.794.13Âą 0.824.38Âą 0.674.49Âą 0.540.51Âą 0.190.77 C: Curated (Guide)0.99Âą 0.110.99Âą 0.094.22Âą 0.884.00Âą 0.964.10Âą 0.713.92Âą 0.740.53Âą 0.200.77 Majority baseline (always âviolationâ)â0.77 Table 1: Results for each examined setting (A-C) across all metrics, averaged over all steps and the 22 distinct cases for GPT-5.4. For each setting, both Human (three annotators) and Models (LLM-as-a-Judge) ratings are reported. The always-âviolationâ majority baseline attains0.77accuracy; settings B and C match it exactly and A exceeds it by a single case, while reasoning-quality scores differ markedly across settings .i.e., reasoning quality does not track predictive accuracy (agreement with the courtâs decision; see Section 4.5). averages in Table 1 directly comparable; the parallel presen- tation may still induce contrast effects on absolute scores, a further reason we emphasize rankings over absolute values (Section 4.5). LLM-as-a-Judge Evaluation Similarly, we instruct three top-tier models (GPT-5.5, DeepSeek V4 Pro, and Claude Opus 4.7) to follow the same evaluation protocol as the human labeler. The models must respond with all relevant scores for each criterion, as instructed. The models assessed all three setting pairs across the dataset Other Automated Metrics We also compute the factual reference overlap as a secondary metric of how well the model grounds its assessment in the case facts: we compare the factual paragraphs the model references (in the â(see paragraph P)â format) against those the Court references in its own assessment, using the F1-score to account for both recall and precision. Since the Court also cites para- graphs the model was never shown (e.g., from the partyâs submissions), we additionally restrict recall to paragraphs present in the facts as the reachable target set, and report groundedness, the fraction of the modelâs references that point to a real facts paragraph. For completeness, we also report predictive accuracy, whether the modelâs prediction (violation or not) aligns with the courtâs ruling, though we do not treat it as a measure of reasoning quality (Section 2.2). 4. Results & Discussion 4.1. Hidden & Output Reasoning A reasoning modelâs response involves three distinct quan- tities: (i) internal reasoning, a chain of computation the model is billed for; (i) an emitted reasoning summary, a short, provider-generated summary the API may return (for the OpenAI models we use, the raw chain is never exposed and this summary is often empty); and (i) the output as- sessment, the final answer. We evaluate only the output assessment (i); by not assessing the modelâs âhiddenâ rea- soning we mean the emitted summary (i), which is at most a compressed, possibly unfaithful digest of the internal reason- ing, so there is nothing reliable to score. This is a disclosure limitation, not a claim that the model reasons little Indeed, the model reasons a great deal internally. Table 2 shows it spends1.7kâ3.5k billed tokens on internal rea- soning (i), comparable to or larger than the visible output (1.0kâ1.3k), yet almost none surfaces as text, the emitted summary is empty in50/20/7%of cases (A, B, C) despite those calls being billed thousands of reasoning tokens. Set- ting C triggers the most internal reasoning (partly its higher reasoning-effort; Appendix A), while B is the most econom- ical overall (shortest output, fewest total tokens). Prompt SettingInternal reasoningOutputTotal A: Non-Curated1709Âą 5771341Âą 1383050 B: Curated (Expert)2016Âą 644963Âą 1092979 C: Curated (Guide)3550Âą 13711230Âą 1304780 Table 2: Mean billed tokens (from the API usage) for the modelâs internal reasoning (stage (i): an internal chain of computation, not returned as text) and its disclosed output assessment (stage (i)), with the total (completion tokens), per setting over the 22 cases. The internal reasoning is billed but not observable in the output; the model emits at most a brief, often empty, reasoning summary (stage (i)), which we do not evaluate. 4.2. Overall Quantitative Assessment Table 1 reports the overall scores per metric, aggregated over all reasoning steps and cases. Step occurrence As we observe based on the human evalua- tion, we have very similar high scores (0.94-0.98) across all settings, which means that the examined model, GPT 5.4, is performing the relevant steps in most cases. The model evaluation offers similar results (0.92-1.00) with the same ranking and slightly inflated scores for settings B and C. No- tably, even the Non-Curated setting (A), which contains no explicit description of the reasoning strategy, still elicits the four doctrinal steps inâź90â100% of cases. âNon-Curatedâ therefore means that we do not scaffold the reasoning in the prompt, not that the model is unaware of the four-step test: the structure appears frequently because it is standard Arti- cle 10 doctrine and thus already internalised by the model. 5 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? Figure 3: Results for step occurrence and comprehensiveness per step (1-4) across settings (A-C) . The role of the curated prompts (B, C) is to make the model apply the test reliably (removing the residual8â18%of missed steps under A), not to teach it from scratch. Step comprehensiveness When it comes to comprehen- siveness, both the annotators and the model judges iden- tify more meaningful differences across settings. All set- tings score between 3 (average) and 4 (good) in this metric, which means that their reasoning is far from ideal (com- plete) compared to the courtâs assessment, but, nonethe- less, leans towards positive. The best-performing setting is B (expert-curated), followed by setting C (model-inferred guide-based), with setting A (non-curated) in third place, a ranking shared by the annotators and the judges. The two rater types agree on this ranking but not on the absolute level: with three annotators, the human comprehensiveness scores are, if anything, slightly higher than the judgesâ (e.g., 4.45vs.4.13for setting B), and the AâB gap is comparable for both (0.50for humans,0.52for judges). Absolute com- prehensiveness scores are therefore rater-dependent; we rely on the ranking, and quantify rater agreement in Section 4.5. Overall conciseness For the overall conciseness, we have a very similar picture with the same ranking across examined settings (B>C>A), and scores in the range 3 (average) to 4 (good). Here, however, the judges disagree among themselves more than on the other criteria (DeepSeek V4 Pro is lenient while Claude Opus 4.7 is strict; Section 4.5), so no single leniency direction holds across judges relative to the annotators. Factual Reference Overlap The F1 between the modelâs and the Courtâs factual references is close to0.51â0.53 across settings, but this is driven by low precisionâthe model referencesâź21 paragraphs per case versus the Courtâsâź13, i.e., it over-cites rather than by poor recall: when recall is restricted to paragraphs actually present in the facts, it is high (0.81â0.86), and groundedness is1.00 (no hallucinated paragraph numbers across any response). The output-length instruction (âup to 1000 wordsâ) is a soft guideline, not a hard token limit; since the model over-cites, its coverage is not suppressed by length. The model thus grounds its references reliably and recovers most facts the Court relied on, but hedges by over-citing. Predictive Accuracy All settings reach77â82%accu- racy, but these differences are not meaningful: the corpus is class-imbalanced (77%of cases find a violation), and a trivial âalways-violationâ baseline attains0.77-exactly the accuracy of settings B and C, with A exceeding it by a sin- gle case (pairwiseMcNemartests are non-significant,â¤1 discordant case; bootstrap95%CIs overlap fully). Cru- cially, accuracy does not track reasoning quality: the best- reasoning setting (B) sits exactly on the baseline, and at the case level comprehensiveness and prediction correct- ness are essentially uncorrelated (r = 0.08; Section 4.5). The accuracy the model achieves thus reflects the majority class rather than the substantive legal test, so accurate pre- dictions without legally meaningful reasoning are of little value (Section 2.2).settings where reasoning matters. 4.3. Step-wise Quantitative Assessment In Figure 3, we present step-wise results for occurrence and comprehensiveness to better understand which steps are harder to perform. We observe that step 2, i.e., assessing the lawfulness of the interference, is the step most frequently missed: it occurs in onlyâź90% of setting-A cases accord- ing to the annotators (andâź82% according to the model judges), making it the most fundamental gap identified by both. Elsewhere, misses are only occasionalâthe annota- tors record step 2 missing in setting B (0.98) and step 4, i.e., assessing the necessity of the interference in a democratic society, missing in settings A and C (0.95) while step 1 and step 3 occur in essentially all cases across settings. Moving to step comprehensiveness, we first observe that the comprehensiveness of step 2 for setting A is clearly affected (penalized) by the fact that the examined model did not perform this step in roughly10â18%of the cases (per the annotators and the judges, respectively); but poor comprehensiveness is not merely a matter of not performing the step always, but rather how the step was performed, with some interesting variations between model and human evaluation. On the one hand, the model judges score the comprehensiveness of step 3 lower than the rest for settings 6 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? (a) Human inter-annotator agreement (3 annotators) CriterionChance-corr.Agreement Step occurrenceÎş = 0.1497% (exact) Step comprehensiveness Îą = 0.09 87% (within-1) Overall concisenessÎą = 0.11 91% (within-1) (b) Judgeâ human-consensus alignment Compr.ConciseOcc. JudgeĎbias Ďbias%agr Claude Opus 4.70.16 â0.260.24 â0.5497 DeepSeek V4 Pro0.24 â0.050.28 +0.4496 GPT-5.50.33 â0.010.23 +0.0198 Judgeâjudge Îą0.410.08 Table 3: Reliability and validity of the evaluation. (a) Agree- ment among the three annotators: a chance-corrected co- efficient (FleissâÎşfor binary occurrence; Krippendorffâs ordinalÎąfor the1â5criteria) alongside raw agreement (exact for occurrence, within-one-point for the1â5crite- ria). The low chance-corrected values are a prevalence / restricted-range artifactâraw agreement is high (see text). (b) Alignment of each LLM judge with the human consen- sus: SpearmanĎand mean signed bias (judgeâhuman; positive=more lenient) for comprehensiveness and con- ciseness, and % agreement on occurrence. A and B, signifying considerable deviation from the courtâs assessment, while the annotators identify more profound inconsistencies in step 4 across all settings. In other words, the model reasoning is poorer for the latter (harder to assess) steps (3 and 4), overall. 4.4. Cross-judge Comparison In Figure 4, we compare the three model judges. All judges agree on the ranking of settings (B>C>A) for every cri- terion, with only minor scoring differences on occurrence and comprehensiveness. They diverge most on overall con- ciseness, where DeepSeek V4 Pro is consistently the most lenient (4.4â4.9) and Claude Opus 4.7 the strictest (3.2â4.0). The ranking across settings is thus robust across judges even though their absolute scores are not; we quantify how each judge aligns with the human annotators next (Section 4.5). 4.5. Reliability and Validity of the Evaluation Since our central claims rest on the human evaluation, we assess both its reliability (do the three annotators agree?) and the validity of the LLM judges as a proxy for human assessment (do the judges agree with humans?). Do the annotators agree? Almost always. The three annotators disagree on whether a step occurs in only11of 240cases (97%exact agreement), and on the1â5criteria they land within one point of each other87â91%of the time (Table 3(a)). The chance-corrected coefficients (Îş/Îą) are nonetheless low (0.09â0.14). This is a known artifact: because almost every rating sits at the top of the scale (a step almost always occurs; quality clusters at4â5), that much agreement is already expected by chance, so little is left forÎş/Îąto credit. Agreement is thus strong on the judgments themselves and on the ranking of settings (all three annotators rankB > C > A), even if the exact score on the1â5scale is somewhat rater-dependentâas, we now show, it is for the LLM judges too. Can the LLM judges replace the humans? (reliability is not validity) To check, we form a human consensus score per case by averaging the three annotators, and com- pare each judge to it (Table 3(b)). The result is striking: the three judges agree with each other far more than the humans agree among themselves (judgeâjudgeÎą = 0.41vs. humanâ humanÎą = 0.09on comprehensiveness), yet each judge correlates only weakly with the human consensus (Ď = 0.16â 0.33). The judges are consistent, then, but consistent with a shared model pattern rather than with human judgment: reliability is not validity. Among the three, GPT-5.5 tracks the human ranking best (Ď = 0.33), Claude Opus 4.7 is the strictest (it scores below the humans), and DeepSeek V4 Pro the most lenient on conciseness; all match the hu- man occurrence labels almost perfectly (96â98%). This cautions against using an LLM judge as a drop-in substitute for human evaluation. Does better reasoning mean a correct prediction? No. Pooling all60(case, setting) instances that received human scores, comprehensiveness and prediction correctness are essentially uncorrelated (r = 0.08,p = 0.55): correctly and incorrectly predicted cases have almost the same aver- age comprehensiveness (4.27vs.4.20). Reasoning quality thus carries essentially no signal about whether the modelâs prediction matches the courtâour central decoupling claim, now measured across cases rather than over just three per- setting points. 4.6. Qualitative Analysis Although step-comprehensiveness scores vary between 3 and 4 across raters, to understand the limitations of the modelâs reasoning we have to compare the model outputs against the actual court judgments. One example is the assessment of lawfulness (step 2) in the case of Tergek v T Ě urkiye (Case 1 in Appendix C), where the Court observes that âit is not disputed between the parties that the interference was prescribed by law. Accordingly, it accepts that the interference complained of by the applicant had a legal basis under domestic law, namely either section 62 or section 68(3) of Law no. 5275.â. 7 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? Figure 4: Results per Setting (A-C) across the three model judges (Claude Opus 4.7, DeepSeek V4 Pro, and GPT 5.5). The first examined setting (expert-curated) performs a thor- ough analysis of lawfulness, but only provides a vague con- clusion: âThis strongly supports a finding that the interfer- ence was not lawful in the Convention sense.â The Courtâs conclusions are always firm: Either the interference was lawful, or it wasnât. If it was not lawful, the model should have predicted a violation for that reason, but it did not; instead it moved to the next steps. The score is low given the vagueness. Another issue is that the model places em- phasis on the assessment performed during the national trial, where a lower Turkish court found the interference unlawful. This assessment, however, was overturned by the Turkish Constitutional Court. The fact that different court instances in Turkey assessed the case differently may have impacted the modelâs assessment, explaining the vagueness. Still, agreement og disagreement between national courts should not affect the reasoning; the Court must perform its own assessment and present a firm conclusion. Furthermore, the modelâs finding does not mirror the Courtâs finding, but this may be because the Courtâs argument for finding that the lawfulness requirement is fulfilled is based on information that was not available to the model, e.g., the partiesâ sub- missions (i.e. the information that the lawfulness of the interference was not disputed by the applicant). The second examined setting (non-curated) also gets the formal legal basis for the interference right, but gets a low score because it presents only a very short analysis and convolutes the lawfulness and legitimacy assessment (steps 2-3). Both models score low (2) on step 4 (proportionality anal- ysis). It is interesting to note here that the Court in pr. 63 identifies an important difference between printouts or pho- tocopies on the one hand and officially published books or periodicals on the other. Specifically the court accept the Turkish Constitutional Courtâs assessment that printouts and photocopies sent to prisoners present risks to the security and order of the prison environment because of the possible infiltration of external communications within large number of printouts. Since it would be unreasonably burdensome on prison staff to review large amounts of photocopies or print- outs for authenticity, the Turkish Constitutional Court and the ECtHR found that the interference was proportionate. A reading of this difference is that the model fails to capture the tacit knowledge the Court applies here: The Court, it could be said, understand the need for a workable system to ensure security in Turkish Prisons and therefore sees the interference as reasonable - The LLMâs do not see this. Recurring failure modes Across the corpus, the modelâs errors recur in a few forms: vague or non-committal conclu- sions where the Court is firm; over-reliance on (sometimes overturned) domestic-court findings; shallow proportional- ity analysis that misses the Courtâs practical balancing; a majority-class bias toward predicting a violation; and over- citation of factual paragraphs. The first three concern the substantive legal judgment and are the most consequential. 5. Limitations We identify a series of overall limitations in our work that future work should address: Limited Scope We consider ECtHR cases, and specifi- cally, one of the many articles of the Convention (ECHR). Hence, our findings are limited to the restricted scope of our study. Nonetheless, our experiments are carefully curated to offer a meaningful, well-scoped evaluation. Future work shall explore a wider range of ECHR articles to better cover the evaluation of ECtHR reasoning. Small DatasetOur 30 Article-10 cases may not faithfully reflect the modelâs capability on the article broadly. We limit data contamination by using only very recent cases, but can- 8 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? not fully rule it out: the modelâs cut-off (August 31, 2025) postdates part of our corpus (published from April 2025 onward), so roughly half of the cases could in principle fall within the training window. We expect memorisation to be limited, as these are recent, less-discussed judgments, unlike the popular older cases that motivated our recency constraint, but note this as a circular trade-off, since older, safely out-of-distribution cases require older, non-state-of- the-art models. Single Model We consider a single model, namely GPT 5.4. Hence, our study does not capture the capabilities of other LLMs to address the examined task. Nonetheless, we believe that using a top-tier, popular model is a good proxy for the state-of-the-art in the field. LLM-Judge modelsWe note that the generator (GPT-5.4) and one of the three judges (GPT-5.5) share the same de- veloper (OpenAI); to mitigate any resulting self-preference bias, we deliberately include two additional judges from different developers (Anthropicâs Claude Opus 4.7 and DeepSeek V4 Pro) and report per-judge results (Figure 4), which expose single-judge idiosyncrasies rather than relying on any one evaluator. Human Annotators We collect annotations from three senior law students, which lets us report inter-annotator agreement and humanâmodel alignment (Section 4.5). They agree strongly on coarse judgments and on the ranking of settings, but the fine-grained1â5scale is rater-dependent. All three are law students working under expert supervision rather than independent legal experts; the expert layer of our study is the supervision and the qualitative doctrinal analysis (Section 4.6), while independent expert quantitative annotation and a larger annotator pool remains valuable future work. 6. Ethical and Societal Implications Our findings carry practical risks for deploying LLMs in legal settings. First, a modelâs disclosed assessment need not reflect its actual internal computation: for the closed model we study, substantial hidden reasoning is billed but never returned (Section 4.1), so a fluent assessment gives no guarantee that its stated reasons produced the decision, a transparency concern for closed models in high-stakes use. Second, our reliabilityâvalidity gap (Section 4.5) cautions against automating the evaluation itself: LLM judges agree with one another far more than with trained annotators and over-rate the hardest (proportionality) step, so using them as the sole arbiter of âgood legal reasoningâ risks entrenching a shared model bias under a veneer of consensus. Third, the modelâs tendency to default to the majority outcome could, if deployed, systematically misjudge the minority of cases where restrictions on expression are in fact justified. Such systems should therefore support, not replace, human legal judgment, and automated evaluation should be validated against, not substituted for, expert assessment. 7. Related Work Building on a line of work on legal judgment prediction over ECtHR cases (Aletras et al., 2016; Chalkidis et al., 2019; T.y.s.s et al., 2022), recent work has stress-tested LLMs on free-text legal reasoning. Shi et al. (2025) proposes step- wise verification and correction over Hong Kong court cases, framing legal judgment prediction as a process-supervision problem with an expert-designed taxonomy of reasoning errors. Chlapanis et al. (2025) introduces a Greek Bar exam benchmark scored along three dimensions (Facts, Cited Ar- ticles, Analysis), paired with a meta-evaluation benchmark (GBB-JME) showing that LLM-judges align with human experts only when guided by span-based rubrics over expert- annotated ground-truths. Juvekar et al. (2025) report on Indian legal exams that frontier LLMs exceed human top- pers on objective multiple-choice papers but fail consistently on long-form reasoning, with recurring failure modes in au- thority, discipline, and forum-appropriate voice. (Enguehard et al., 2025) improves reference-free LLM-as-Judge in le- gal Q&A by decomposing answers into atomic Legal Data Points, reporting better correlation with human experts. Compared to this line of work, our study is deliberately narrow (one ECHR article, 30 recent cases, one model) but couples an expert-curated reasoning prompt with a parallel human/model evaluation that decouples reasoning quality from predictive accuracy; our step-wise analysis surfaces where that decoupling matters and cautions against task accuracy as a proxy for legal reasoning quality. 8. Conclusion and Future Work We studied LLM legal reasoning on Article-10 ECtHR cases with GPT-5.4 under three prompting strategies (zero to expert-level guidance), evaluated by three annotators and three LLM judges. Our answer to the title question is not yet fully: the model reliably reproduces the doctrinal structure, but its substantive reasoning remains shallow and the predic- tions it gets right largely track the majority outcome rather than sound legal analysis. The expert-curated prompt yields the most comprehensive reasoning yet no gain in accuracy, and the LLM judges prove reliable but only weakly aligned with human annotators. We thus urge the community not to rely solely on automated evaluation, nor to treat task accuracy as a proxy for reasoning quality. In future work, we plan to extend the study to other ECHR articles (starting with Articles 3 and 11), enlarge the dataset, and evaluate several top-tier models, as well as recruit more (and more senior) annotators to further study inter-annotator agreement and expert reasoning quality. 9 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? References Aletras, N., Tsarapatsanis, D., Preot ̧iuc-Pietro, D., and Lam- pos, V. Predicting judicial decisions of the european court of human rights: A natural language processing perspective. PeerJ computer science, 2:e93, 2016. Chalkidis, I., Androutsopoulos, I., and Aletras, N. Neu- ral legal judgment prediction in English. In Korho- nen, A., Traum, D., and M ` arquez, L. (eds.), Proceed- ings of the 57th Annual Meeting of the Association for Computational Linguistics, p. 4317â4323, Florence, Italy, July 2019. Association for Computational Lin- guistics. doi: 10.18653/v1/P19-1424. URLhttps: //aclanthology.org/P19-1424/. Chalkidis, I., Fergadiotis, M., Tsarapatsanis, D., Aletras, N., Androutsopoulos, I., and Malakasiotis, P. Paragraph-level rationale extraction through regularization: A case study on European court of human rights cases. In Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-Tur, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., and Zhou, Y. (eds.), Proceedings of the 2021 Conference of the North American Chapter of the Association for Com- putational Linguistics: Human Language Technologies, p. 226â241, Online, June 2021. Association for Compu- tational Linguistics. doi: 10.18653/v1/2021.naacl-main. 22.URLhttps://aclanthology.org/2021. naacl-main.22/. Chlapanis, O. S., Galanis, D., Aletras, N., and An- droutsopoulos, I.GreekBarBench: A challenging benchmark for free-text legal reasoning and citations. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V. (eds.), Findings of the Associa- tion for Computational Linguistics: EMNLP 2025, p. 25099â25119, Suzhou, China, November 2025. Asso- ciation for Computational Linguistics. ISBN 979-8- 89176-335-7. doi: 10.18653/v1/2025.findings-emnlp. 1368. URLhttps://aclanthology.org/2025. findings-emnlp.1368/. Enguehard, J., Van Ermengem, M., Atkinson, K., Cha, S., Chowdhury, A. G., Ramaswamy, P. K., Roghair, J., Marlowe, H. R., Negreanu, C. S., Boxall, K., and Mincu, D. LeMAJ (legal LLM-as-a-judge): Bridging legal reasoning and LLM evaluation. In Aletras, N., Chalkidis, I., Barrett, L., Goant , Ě a, C., Preot , iuc-Pietro, D., and Spanakis, G. (eds.), Proceedings of the Nat- ural Legal Language Processing Workshop 2025, p. 318â337, Suzhou, China, November 2025. Association for Computational Linguistics. ISBN 979-8-89176-338- 8. doi: 10.18653/v1/2025.nllp-1.23. URLhttps: //aclanthology.org/2025.nllp-1.23/. Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Rein- forcement Learning. arXiv preprint arXiv:2501.12948, 2025. Juvekar, K., Bhattacharya, A., Khadloya, S., and Sax- ena, U.Are LLMs court-ready?evaluating fron- tier models on Indian legal reasoning. In Aletras, N., Chalkidis, I., Barrett, L., Goant , Ě a, C., Preot , iuc-Pietro, D., and Spanakis, G. (eds.), Proceedings of the Nat- ural Legal Language Processing Workshop 2025, p. 359â369, Suzhou, China, November 2025. Association for Computational Linguistics. ISBN 979-8-89176-338- 8. doi: 10.18653/v1/2025.nllp-1.26. URLhttps: //aclanthology.org/2025.nllp-1.26/. Medvedeva, M., Ě Ust Ě un, A., Xu, X., Vols, M., and Wiel- ing, M. Automatic judgement forecasting for pending applications of the european court of human rights. In ASAIL/LegalAIIA@ ICAIL, p. 12â23, 2021. Neelamegam, P. and Nirmala, S. J. Recent trends in legal ai: A comprehensive review. In 2025 Third International Conference on Augmented Intelligence and Sustainable Systems (ICAISS), p. 152â159. IEEE, 2025. OpenAI. Openai o1 system card, 2024. URLhttps: //arxiv.org/abs/2412.16720. OpenAI. Gpt-5.4 thinking system card. Technical re- port, OpenAI, 2026. URLhttps://openai.com/ index/gpt-5-4-thinking-system-card/. Shi, W., Zhu, H., Ji, J., Li, M., Zhang, J., Zhang, R., Zhu, J., Xu, J., Han, S., and Guo, Y. LegalReasoner: Step- wised verification-correction for legal judgment reason- ing. In Che, W., Nabende, J., Shutova, E., and Pilehvar, M. T. (eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 7297â7313, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 979-8-89176-251-0. doi: 10.18653/v1/2025.acl-long. 361. URLhttps://aclanthology.org/2025. acl-long.361/. T.y.s.s, S., Xu, S., Ichim, O., and Grabmair, M. Decon- founding legal judgment prediction for European court of human rights cases towards better alignment with experts. In Goldberg, Y., Kozareva, Z., and Zhang, Y. (eds.), Proceedings of the 2022 Conference on Em- pirical Methods in Natural Language Processing, p. 1120â1138, Abu Dhabi, United Arab Emirates, Decem- ber 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.emnlp-main.74.URLhttps:// aclanthology.org/2022.emnlp-main.74/. Xu, S., T.y.s.s, S., Ichim, O., Risini, I., Plank, B., and Grabmair, M. From dissonance to insights: Dissect- 10 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? ing disagreements in rationale construction for case out- come classification. In Bouamor, H., Pino, J., and Bali, K. (eds.), Proceedings of the 2023 Conference on Em- pirical Methods in Natural Language Processing, p. 9558â9576, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023. emnlp-main.594. URLhttps://aclanthology. org/2023.emnlp-main.594/. 11 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? A. Models and cvz All models were accessed through the OpenRouter API. Table 4 reports the exact model identifiers, their role in our study, and the reasoning-effort setting used for each. GPT-5.4 serves as the generator; GPT-5.5, Claude Opus 4.7, and DeepSeek V4 Pro serve as the LLM-as-a-Judge evaluators. As these are very recently released models, some readers may be unfamiliar with them; we therefore provide their precise identifiers and versions here for reproducibility. RoleName (main text)API identifier (OpenRouter)Reasoning effortRelease date GeneratorGPT-5.4 openai/gpt-5.4mediumMarch 5, 2026 EvaluatorGPT-5.5 openai/gpt-5.5highApril 23, 2026 EvaluatorClaude Opus 4.7 anthropic/claude-opus-4.7highApril 16, 2026 EvaluatorDeepSeek V4 Pro deepseek/deepseek-v4-prohighApril 24, 2026 Table 4: Exact model identifiers and settings. All models were accessed via the OpenRouter API. A, B, and C denote the Non-Curated, Curated (Expert), and Curated (Guide) prompting settings, respectively (Section 3.2). B. Instructions (Prompts) We present the prompts (instructions) used across the different settings, as described in Section 3.2: A. Non-curated Instruction (Prompt) ââ You are a legal assistant tasked with forecasting the outcome of cases brought to the European Court of Human Rights (ECtHR) by assessing the case facts. Task Instruction You are provided with a summary of the case facts related to a new case. In your analysis, you should focus only on information relevant to answering the question of whether there has been a violation of Article 10 of the European Convention of Human Rights (ECHR) regarding freedom of expression or not. You have to predict whether Article 10 of the ECHR has been violated in the examined case. To do so, you have to assess the facts of the case in light of ECtHR standards as set out in its jurisprudence. Detailed Instructions ⢠Structure your case assessment as a list of titled paragraphs, separated step by step. It should be up to 1000 words. ⢠As part of your assessment, provide references to the factual paragraphs as part of the assessment, following the format used in ECHR proceedings, i.e., â(see paragraph P)â, where P is the number of the relevant paragraph as presented and numbered in the case facts, e.g., â(see paragraph 12)â. ⢠As part of your assessment, provide references to ECHR case law, following the format used in ECHR proceedings, i.e., â(see no. D/D)â. You can only use the precedent cases from the list presented below. ⢠Do not address any other ECHR articles, except Article 10. ⢠Your response shall be in JSON format. The first key, named âassessmentâ, is your assessment (analysis), and the value shall be a list with titled paragraphs, one per reasoning step, i.e., âTitle of the paragraphâ: âContent of the paragraphâ. The second one, named âviolationâ, is your prediction (violation [yes] or not [no]) Relevant Case Law 11882/10 2594/07 1172/12 Examined Case Facts Case A V. D 12 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? 1. [...] 2. [...] 3. [...] ââ B. Curated (Step-by-step) Instruction (Prompt) ââ You are a legal assistant tasked with forecasting the outcome of cases brought to the European Court of Human Rights (ECtHR) by assessing the case facts. Task Instruction [...] Assessment Strategy Instruction 1. âExistence of an interferenceâ: Assess the facts of the case to determine whether there has been an interference with the applicantâs freedom of expression. If you find no interference, predict that there has been no violation of the article and skip the rest of the steps; if you find that there has been an interference, proceed to step 2. 2.âLawfulness of the interferenceâ: Assess if the interference is lawful. Assess whether the interference has a basis in the domestic law of the member state. If you find that the interference lacks a basis in the domestic law of the member state, predict that there has been a violation of the article and skip the rest of the steps; otherwise, proceed to step 3. 3.âLegitimate aims of the interferenceâ: Assess if the interference is legitimate. Assess if the interference was supported by a need to protect one of the following legitimate aims: (i) national security, (i) territorial integrity, (i) public safety, (iv) the prevention of disorder or crime, (v) the protection of health or morals, (vi) the protection of the reputation or rights of others, (vii) preventing the disclosure of information received in confidence, or (viii) maintaining the authority and impartiality of the judiciary. If you find that the interference does not support one of the legitimate aims, predict that there has been a violation of the article, and skip the next step; otherwise, proceed to step 4. 4.âNecessary in a democratic societyâ: Assess if the interference is necessary in a democratic society. Assess: (a) whether there is a pressing social need for interference; (b) whether the interference uses the least restrictive measure possible to protect the legitimate interest identified in the third step, and (c) whether there are relevant and sufficient reasons to support the interference convincingly. If you find that the interference can be defended as necessary in a democratic society, then predict that there has been a violation of the article; if you find that the interference cannot be so defended, predict that there has been no violation. Detailed Instructions ⢠Structure your case assessment as a list of 4 titled paragraphs, separated step by step. It should be up to 1000 words. â˘As part of your assessment, provide references to the factual paragraphs as part of the assessment, following the format used in ECHR proceedings, i.e., â(see paragraph P)â, where P is the number of the relevant paragraph as presented and numbered in the case facts, e.g., â(see paragraph 12)â. â˘As part of your assessment, provide references to ECHR case law, following the format used in ECHR proceedings, i.e., â(see no. D/D)â. You can only use the precedent cases from the list presented below. ⢠Do not address any other ECHR articles, except Article 10. â˘Your response shall be in JSON format. The first key, named âassessmentâ, is your assessment (analysis), and the value shall be a list with 4 titled paragraphs, one per reasoning step (1-4), i.e., âExistence of an interferenceâ: â...â. The second one, named âviolationâ, is your prediction (violation [yes] or not [no]) Relevant Case Law [...] Examined Case Facts 13 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? [...] ââ C. Curated (Guide) Instruction (Prompt) C1. Inferring Strategy from Guide Instruction ââ You are a legal assistant tasked with assisting with tasks related to the European Court of Human Rights (ECtHR). Task Instruction You are provided with the official guide of Article 10 of the European Convention of Human Rights (ECHR), part of the series of Case-Law Guides published by the European Court of Human Rights. You have to extract from the presented guide (below) a step-by-step case assessment strategy (methodology) on how the Courtâs judges shall assess the facts of each case in light of ECtHR standards as set out in the guide. Detailed Instructions ⢠Structure your case assessment strategy as a list of paragraphs, separated step by step. The steps shall describe sequential actions where the court has to assess the facts related to different factors. It should be up to 1000 words. â˘Each step should have a title presenting the number of the step, e.g., â1. Title of Step 1â, and the description of the step, i.e., what the judges have to assess in the specific steps and how they should proceed or not in further steps. For example, the description can be formulated like this: âAssess the facts of the case to determine whether [...] If you find that [...] predict that there has been a violation of Article 10, and skip the next step; otherwise, proceed to the next step.â. â˘Your response shall be in JSON format. The key, named âassessmentstrategyâ, is your assessment strategy, and the value shall be a list with titled paragraphs, one per reasoning step (1-4), i.e., â1. Title of Step 1â: âDescription of Step 1â. Guide on Article 10 of the European Convention on Human Right [...] ââ C2. Applying Inferred Strategy from Guide Instruction ââ Task Instruction [...] Assessment Strategy Instructions 1. âIdentify the complaint structure: interference or positive obligationâ: Assess whether there was an interference with freedom of expression, or instead a failure by the State to protect its effective exercise. Interference can take many forms: criminal conviction, damages, publication ban, confiscation, refusal of access, surveillance, source- disclosure order, dismissal from employment, disciplinary action, blocking of websites, or other measures capable of chilling future expression. Do not rely only on domestic labels; examine the practical effect on expression. Where proceedings ended without sanction, ask whether they still operated as a warning or deterrent. If there was no direct interference, assess whether the State had a positive obligation to secure effective freedom of expression, especially in journalism, private-employment settings, access to venues or information, and protection from intimidation or disruption. If you find neither interference nor a breached positive obligation, predict no violation of Article 10 and stop. If you find one of them, define the context of the speech (press, politics, judiciary, whistleblowing, internet, reputation, national security, elections, health or morals, etc.) and proceed. 14 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? 2.âTest the Article 10 § 2 gateways: lawfulness and legitimate aimâ: Assess whether the interference was âprescribed by lawâ: the legal basis must exist in domestic law and have sufficient quality, meaning accessibility, foreseeability and safeguards against arbitrariness or abuse. Examine whether the rule was clear enough for the person concerned, taking account of the context, the field regulated, and the personâs professional status. Vague or contradictory rules, ad hoc interpretations, overbroad discretion, or lack of safeguards against arbitrary enforcement point toward unlawfulness. Then verify whether the interference pursued one or more legitimate aims listed exhaustively in Article 10 § 2, such as protection of reputation or rights of others, national security, prevention of disorder or crime, health or morals, confidentiality, or the authority and impartiality of the judiciary. If either lawfulness or legitimate aim fails, predict a violation of Article 10 and skip the next step. Otherwise proceed to necessity and proportionality. 3.âAssess necessity in a democratic society through contextual proportionality and balancingâ: Assess whether the interference answered a pressing social need and was proportionate to the legitimate aim pursued. Start by fixing the margin of appreciation: it is narrow for political speech, matters of public interest, elections, press watchdog activity, and debate on the functioning of justice; wider for commercial speech, some morality cases, and general media-regulatory choices. Then ask whether domestic courts gave relevant and sufficient reasons and applied Convention standards to the facts. Examine the core contextual factors that fit the case: contribution to a debate of public interest; status and role of the speaker (journalist, politician, judge, lawyer, NGO, whistleblower, academic, ordinary citizen); target of the statement (politician, public official, private person, company, judiciary); distinction between facts and value judgments; existence of a sufficient factual basis; good faith and responsible journalism; method of obtaining the information and its veracity; content, tone, form, medium, audience and likely impact; and the severity and nature of sanctions, including chilling effect, imprisonment, damages, injunctions, dismissal, blocking or surveillance. Where two Convention rights conflict, especially Article 10 with Articles 8 or 6 § 2, conduct a fair-balance analysis using the Courtâs specific criteria; if domestic courts already carried out that balancing in conformity with Strasbourg criteria, depart only for strong reasons. In thematic cases, add the guideâs specialised modules: source protection requires an overriding public interest and strong procedural safeguards; confidential-information and whistleblowing cases require attention to public interest, authenticity, reporting channels, harm caused and retaliation; access-to-information cases require that the request be instrumental to expression, concern information of public interest, come from a qualifying watchdog-type actor or equivalent, and concern ready and available material; national-security, disorder, health, morals and internet cases require close analysis of context, risk, reach and audience. If the reasons are not relevant and sufficient, the measure was not the least restrictive means, or the sanction is excessive, predict a violation; otherwise predict no violation. Detailed Instructions [...] Relevant Case Law [...] Examined Case Facts [...] ââ LLM-as-a-Judge Evaluation (Prompt) ââ You are a legal assistant tasked with evaluating law studentsâ assessments of cases brought to the European Court of Human Rights (ECtHR). Task Instruction You are provided with a summary of the case facts (see âCase Factsâ below) and the courtâs assessment (see âCourtâs Assessmentâ below) for an ECtHR case related to Article 10 of the European Convention of Human Rights (ECHR). You are also provided with the assessment from two law students for the very same case (see âStudent A Responseâ and âStudent B Responseâ below). You have to evaluate the quality of the studentsâ assessments in light of ECtHR standards as set out in its jurisprudence. The studentsâ assessments should follow the following reasoning strategy, step by step: 15 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? [...] The students had to provide their assessment as a sequence of titled paragraphs, not necessarily following the exact same structure (presented above). Evaluation Criteria â˘âStep occurrenceâ: Report if each of the 4 steps of the ideal assessment (reasoning) strategy (presented above) occursâis presentâin the studentâs response. It is totally fine if the student performed the step, even if the step is part of a paragraph including more steps. This is a binary label (yes or no). â˘âStep comprehensivenessâ: Report the comprehensiveness of each of the occurred steps in each studentâs response in a Likert-scale (1-5, where 5 is greater than 1) in light of how the step is performed by the courtâs assessment. If the expected step did not occur in the modelâs response, the score is 0. If for a step, you score comprehensiveness for student A as 1, and for student B as 4, this implies that student Bâs step was more comprehensive. â˘âOverall concisenessâ: Report if the overall studentâs assessment is concise, i.e., if there is no irrelevant (redundant) information in each studentâs response, as a whole, in a Likert-scale (1-5, where 5 is greater than 1). General Instructions ⢠Your response shall be in JSON format. The first key, named âevaluationâ, is your evaluation (assessment), which has a list of JSON records, one for each student (âstudentaâ, âstudentbâ). â˘For each student, there is a list for the assessment of the 4 steps (âstep1â, âstep2â, âstep3â, âstep4â), where each step has a subfield (âoccurrenceâ) where the value is 0 (no) or 1 (yes), and a sub-field (âcomprehensivenessâ), where the value is 1-5 (int), as described above (Evaluation Criteria). â˘For each student, there is an extra field âoverallconcisenessâ, where the values are 1-5 (int), as described above (Evaluation Criteria). Output Example [...] Case Facts [...] Courtâs Assessment [...] Student A Response [...] Student B Response [...] 16 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? C. Case Examples We present 2 representative cases below. For each case we include: (i) the case facts as provided to the model; (i) the Courtâs Merits assessment (ground truth); and (i) model assessments under each of the three prompting strategies (Section 3.2) using GPT-5.4. Prompts D and E show assessments by Mistral Medium 3.5 under Prompts A and B as a supplementary model comparison not discussed in the main paper. Case 1: TERGEK v. T Ě URK Ě IYE (Application no. 39631/20), 29 April 2025 CASE FACTS 2. The applicant was born in 1989 and is currently serving a prison sentence in Kocaeli T-Type Prison. He was represented by Mr S. AltÄąntas ̧, a lawyer practising in Kocaeli. 3. The Government were represented by their Agent at the time, Mr HacÄą Ali Ac ̧Ĺkg Ě ul, former Head of the Department of Human Rights of the Ministry of Justice of the Republic of T Ě urkiye. 4. At the time of the events giving rise to the present application, the applicant was detained following his conviction for membership of an armed terrorist organisation described by the Turkish authorities as the âFetullahist Terror Organisation/Parallel State Structureâ (Fetullahc ̧Ĺ Ter Ě or Ě Org Ě ut Ě u / Paralel Devlet YapÄąlanmasÄą, hereinafter referred to as âthe FET Ě O/PDYâ). withholding of letters sent to the applicant and related proceedings First letter and related proceedings 5. On 22 October 2018 the prison administrationâs letter-reading committee reviewed a letter sent to the applicant by his sister and found its enclosures to be objectionable (sakÄąncalÄą). The letter, which contained thirty-one pages of documents printed from the internet, was subsequently referred to the prisonâs Disciplinary Board for further examination. 6. On the same day, the Disciplinary Board, citing section 68(3) of the Law on the execution of sentences and preventive measures (âLaw no. 5275â â see paragraph 19 below), decided to withhold the letter. That decision was based on the grounds that the letterâs enclosures contained statements which could potentially pose a threat to prison security, that it was unclear who had published the information or for what purpose, and that the information included phrases which could facilitate communication within the FET Ě O/PDY organisation. 7. On 26 October 2018 the applicant lodged an objection with the Kocaeli enforcement judge against the Disciplinary Boardâs decision. In his objection the applicant explained that he had injured his ankle on 1 October 2018 and had required a cast for several weeks. He stated that some of the withheld documents had been sent to him by his sister, a physiotherapist, and had contained information on various physiotherapy exercises he could do to assist with the rehabilitation of his ankle. The applicant further explained that the remaining documents related to a distance-learning course in real-estate management, which he was taking, and that he needed them in order to prepare for examinations. 8. On 19 August 2019 the enforcement judge upheld the applicantâs objection, citing the relevant principles and case-law of both the Court and the Constitutional Court. The judge ruled that withholding the documents solely on the grounds that the information contained in them was unclear from a general inspection, without a specific assessment of their actual content, had been unlawful in the light of the freedom of communication and expression. 9. On 1 October 2019 the applicant was notified of that decision. On 25 October 2019 he received the letter and its enclosures. Second letter and related proceedings 10. On 18 December 2018, the letter-reading committee deemed another letter sent by the applicantâs wife to be objectionable. The letter, containing sixty-one pages of documents printed from the internet, a one-page handwritten note and four pictures, was forwarded to the Disciplinary Board for further examination. 11. On the same day, the Disciplinary Board, citing section 68(3) of Law no. 5275 and referencing a prior decision issued by the Administration and Monitoring Board on 4 November 2016 concerning the potential risks associated with allowing internet printouts to be handed over to prisoners, decided to withhold the documents in question from the applicant. It authorised the remaining items â the one-page handwritten note and the four pictures â to be handed over to him. The decision did not in any way address the content of the withheld documents. 12. On 24 December 2018 the applicant objected to that decision before the enforcement judge. In his objection, he argued that the documents in question contained information on various physiotherapy exercises and real-estate management and were essential for both his rehabilitation and his ongoing education. 13. On 27 August 2018, citing section 62 of Law no. 5275 (see paragraph 17 below), the enforcement judge dismissed the applicantâs objection. The judge noted that the internet printouts could not be classified as âbooksâ or âcorrespondenceâ under the relevant legislation as a result of their unknown origin and their susceptibility to external interference. The judge also noted that prisoners were free to access publications through the prison library. 14. On 17 September 2019 the Kocaeli Assize Court, ruling on an objection lodged by the applicant, endorsed the reasoning provided by 17 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? the enforcement judge. INDIVIDUAL Application LODGED by the applicant WITH THE CONSTITUTIONAL COURT 15. On 30 October 2019 the applicant lodged an individual application with the Constitutional Court, arguing, inter alia, that his right to respect for correspondence had been breached as a result of (i) the delayed delivery of the first letter and its enclosures, and (i) the seizure of the second letter and its enclosures by the prison administration. 16. On 16 June 2020 the Constitutional Court, sitting as a panel of two judges, dismissed those complaints as manifestly ill-founded. It referred to its leading judgment in the case of Diyadin Akdemir (application no. 2015/9562, 4 April 2018). The reasoning for its decision reads as follows: âHaving reviewed the application within the scope of the Constitutional Courtâs authority to examine individual applications, and having considered the documents submitted, it has been concluded that there was no interference with the fundamental rights and freedoms set forth in the Constitution, or that any interference that may have occurred did not constitute a violation of those rights (see, to the same effect, Diyadin Akdemir, application no. 2015/9562, 4 April 2018).â RELEVANT LEGAL FRAMEWORK AND PRACTICE RELEVANT DOMESTIC LEGISLATION 17. Section 62 of Law no. 5275, entitled âThe right to receive periodical and non-periodical publicationsâ, as in force at the material time, read as follows, in so far as relevant to the present case: ⢠â1. The convicted person shall have the right to receive [any] periodical or non-periodical publication on payment of the [retail] price, provided that [such publications are not] prohibited by a court. ⢠2. Newspapers, books and printed publications of public institutions, universities, professional organisations with the status of public institutions, foundations benefiting from a tax exemption [by decision of] the President of the Republic and associations working in the public interest shall be handed over free of charge to convicted persons. The textbooks used by convicted persons undergoing studies or training shall not be subject to inspection. â˘3. No publication endangering the security of the institution or containing obscene articles, writings, photographs, or comments shall be given to the convicted person....â 18. Following an amendment made by Law no. 7242 of 14 April 2020, subsection 3 of section 62 is now worded as follows: â˘â3. No publication that disturbs or endangers discipline, order, or security within the institution, [hinders] the achievement of the purpose of [rehabilitation] of convicted persons, or contains obscene articles, writings, photographs, or comments, shall be given to the convicted person.â 19. Section 68 of Law no. 5275, entitled âThe right to receive and send letters, faxes, and telegramsâ, as in force at the material time, provided as follows, in so far as relevant: ⢠â1. With the exception of the restrictions set forth in this section, convicted prisoners shall have the right, at their own expense, to send and receive letters, faxes, and telegrams. ⢠2. The letters, faxes and telegrams sent or received by convicted prisoners shall be monitored by the reading committee in those prisons which have such a body or, in those which do not, by the highest authority in the prison. â˘3. [If] letters, faxes, and telegrams [to convicted prisoners] pose a threat to order and security in the prison, single out serving officials as targets, permit communication between members of terrorist or ... criminal organisations, contain false or misleading information likely to cause panic in individuals or institutions or contain threats or insults, they shall not be forwarded to [the addressee]. Nor shall [such letters, faxes, and telegrams] written by convicted prisoners be dispatched....â 20. Under section 116(1) of Law no. 5275, the provisions of the above sections may be applied to remand prisoners in so far as those provisions are compatible with the detention status of the prisoners concerned. 21. Regulation 91 of the Regulations of 20 March 2006 on the management of prisons and the execution of sentences and preventive measures (âthe Regulationsâ), published in the Official Gazette of 6 April 2006, as in force at the material time, provided as follows: ⢠â1. Convicted prisoners shall have the right to send and receive letters, faxes, and telegrams at their own expense. ⢠2. The letters, faxes and telegrams sent or received by convicted prisoners shall be monitored by the reading committee in those prisons which have such a body or, in those which do not, by the highest authority in the prison. â˘3. [If] letters, faxes, and telegrams [to convicted prisoners] pose a threat to order and security in the prison, single out serving officials as targets, permit communication for organisational purposes between members of terrorist or ... criminal organisations, contain false or misleading information likely to cause panic in individuals or institutions or contain threats or insults, they shall not be forwarded to [the addressee]. Nor shall [such letters, faxes, and telegrams] written by convicted prisoners be dispatched....â 18 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? 22. For other provisions of domestic law relevant to the present case, se Halit Kara v. T Ě urkiye (no. 60846/19, § 16-18, 12 December 2023), Mehmt C ̧ iftci v. Turkey (no. 53208/19, § 10-15, 16 November 2021) and Osman and Altay v. T Ě urkiye (nos. 23782/20 and 40731/20, § 13-19, 18 July 2023). RELEVANT case-law OF THE CONSTITUTIONAL COURT 23. In its judgment in the case of Diyadin Akdemir (application no. 2015/9562, 4 April 2018), the Constitutional Court examined an individual application challenging the prison authoritiesâ refusal to provide a prisoner with photocopied documents sent to him by post. The court unanimously declared the application inadmissible as being manifestly ill-founded. It reasoned that photocopied documents were not covered by section 62 of Law no. 5275, which refers specifically to âperiodicals and non-periodicalsâ. Therefore, applying the same inspection criteria to photocopied documents as to periodicals and non-periodicals would impose an unreasonable burden on the prison administrations and domestic courts. The Constitutional Court further noted that photocopied documents could potentially raise copyright issues. The relevant parts of the judgment read as follows: â20.... Regarding publications that are not subject to any prohibition orders, the Constitutional Court has stated that any interference with freedom of expression without justification, or without meeting the criteria established by it (Halil BayÄąk [GK], App. No: 2014/20002, 30/11/2017, § 28-43), would constitute a violation of Article 26 of the Constitution. 21. It is clear that the inspection required in accordance with the principles and criteria outlined above cannot apply to photocopied documents, which do not fall within the scope of âperiodicals and non-periodicalsâ as referred to n section 62 of Law no. 2575. Expecting photocopied documents sent to convicted and remand prisoners to be subject to inspection in accordance with the above- mentioned provision, in the light of the principles and criteria approved by the Constitutional Court, would place an unreasonable burden on the prison administrations and the lower courts. 22. However, it should not be overlooked that documents in the form of photocopied books, as in the case at hand, may raise copyright issues. In this context, it cannot be argued that the intervention in question â refusing to provide photocopied books to the applicant, a convicted person, on the grounds that inspection was not feasible â was unnecessary in a democratic society.â COURTâS ASSESSMENT (ARTICLE 10 â MERITS) 53. The Court notes that the case concerns the applicantâs request to receive information, in the form of internet printouts, which his wife sent to him by post and which the prison authorities refused to deliver. In that connection, the Court reiterates that, in general, prisoners continue to enjoy all the fundamental rights and freedoms guaranteed by the Convention, with the exception of the right to liberty. Thus, they continue to enjoy the right to freedom of expression (see Yankov v. Bulgaria, no. 39084/97, § 126-45, ECHR 2003-XII, and Tapkan and Others v. Turkey, no. 66400/01, § 68, 20 September 2007), which includes the right to receive information or ideas (see Mesut Yurtsever and Others v. Turkey, nos. 14946/08 and 11 others, § 101, 20 January 2015; MehmetC ̧iftci, cited above, § 32; and Osman and Altay, cited above, § 40). 54. The Court considers that the refusal of the national authorities to hand the documents in question over to the applicant amounted to an interference with his right to receive information and ideas (see Mehmet C ̧ iftci, § 33, and Osman and Altay, § 41, both cited above). 55. The Court observes that it is not disputed between the parties that the interference was prescribed by law. Accordingly, it accepts that the interference complained of by the applicant had had a legal basis under domestic law, namely either section 62 or section 68(3) of Law no. 5275. 56. The Court further notes that the interference pursued legitimate aims within the meaning of Article 10 § 2 of the Convention, namely the protection of national security, the prevention of disorder and the prevention of crime. 57. As regards the necessity of the interference, the Court reiterates the principles deriving from its case-law on freedom of expression, which are summarised in, inter alia, B Ě edat v. Switzerland ([GC], no. 56925/08, 29 March 2016) and Kula v. Turkey (no. 20233/06, § 45-46, 19 June 2018). 58. In order to determine whether the interference with the applicantâs right to freedom of expression has been convincingly justified in the present case, the Court must assess, in line with its case-law, whether the reasons provided by the national authorities to justify the interference were ârelevant and sufficientâ and whether the measure taken was âproportionate to the legitimate aim pursuedâ. 59. As to the assessment of whether the reasons provided were ârelevant and sufficientâ, the Court notes that its task in exercising its supervisory jurisdiction is not to take the place of the competent domestic courts, but rather to review under Article 10 the decisions they have taken pursuant to their margin of appreciation. The domestic courts, given their constant contact with the realities of the country, are often better placed than an international judge to determine whether a fair balance was struck at a given moment. If the balancing exercise undertaken by the national authorities was carried out in compliance with the criteria established by the Courtâs case-law, serious reasons are required for the Court to substitute its opinion for that of the domestic courts (see Haldimann and Others v. Switzerland, no. 21830/09, § 54 and 55, ECHR 2015, and B Ě edat, cited above, § 54). 60. In determining the proportionality of a general measure, such as the one at issue in the present case â namely, the withholding of printed documents from a prisoner solely on the basis of their format â the Court further reiterates that the quality of the judicial review of the necessity of the measure at the national level is of particular importance, including with regard to the application of the relevant margin of appreciation (see Animal Defenders International v. United Kingdom [GC], no. 48876/08, § 108, ECHR 2013 (extracts)). The Court has already held that a general measure is a more practical means of achieving the legitimate aim pursued than a provision allowing for case-by-case examination, as the latter system is likely to lead to considerable uncertainty, litigation, costs, and delays, 19 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? or to discrimination and arbitrariness. Nevertheless, the way in which a general measure has been applied to the facts of a given case helps to reveal its practical impact and is therefore relevant to the assessment of its proportionality, making it an important factor to take into account (ibid.). It follows that the more persuasive the general justifications put forward in support of the general measure, the less importance the Court attaches to the impact of that measure in the particular case before it (ibid., § 109). 61. Turning to the present case, the Court notes at the outset that the Constitutional Court, in its judgment in the Diyadin Akdemir case, set out the criteria that prison authorities must consider when examining photocopied documents sent to prisoners (see paragraph 23 above). Those criteria were reiterated and elaborated upon in detail in the written submissions by the Government (see paragraphs 46-51 above). 62. The Constitutional Court explicitly stated that section 62 of Law no. 5275 referred specifically to âperiodicals and non-periodicalsâ and that photocopied documents were not covered by that section. It held that applying the same inspection criteria to photocopied documents as to periodicals and non-periodicals would impose an unreasonable burden on prison administrations and the domestic courts. The Government, in their written submissions, also pointed out the significant risk of intra-organisational communication, particularly on account of the large volume of incoming documents relating to prisoners convicted of terrorism-related crimes. 63. The Court notes that, although the receipt of photocopied or printed documents in prison was not explicitly regulated by domestic law, the Constitutional Court carried out a thorough and detailed assessment of the matter in the Diyadin Akdemir case. In its assessment, that court balanced the right of prisoners to access information and ideas with the duties and workload of the prison authorities, as well as the serious risks associated with intra-organisational communication. The Court recognises that reviewing a large volume of printed or photocopied documents, in addition to the regular publications sent to prisoners could indeed overwhelm prison staff, impede their duties, and place an excessive burden on the judiciary, including the Constitutional Court. It also acknowledges the inherent differences between printouts or photocopies and officially published books or periodicals, which typically undergo thorough reviews and regulatory controls prior to release to ensure compliance with legal standards. By contrast, printouts and photocopies sent to prisoners lack such pre-publication scrutiny, thereby presenting specific risks to the security and order of the prison environment, including the heightened risk of infiltration of certain external communications within large number of printouts (compare with Osman and Altay, cited above, § 53, which concerned publications sent to the applicants via postal services in disregard of the legal procedures laid down in that regard). The Court further notes that similar considerations were reflected in the decisions of the prison authorities (see paragraphs 11, 13 and 14 above). 64. The Court has already noted the various means available to prisoners for obtaining publications in accordance with the relevant domestic law, namely obtaining publications upon request, internal access to publications issued by official authorities and certain organisations, books kept in prison libraries and school textbooks, and the possibility of receiving books as gifts on specific occasions (see the full text of the relevant provisions in Osman and Altay, cited above, § 14-17), which were still valid at the time of the events in the present case. 65. The Court therefore finds that, in the circumstances of the case, it cannot be considered that the applicantâs right to freedom to receive information and ideas was disproportionately restricted by his inability to access the printed material sent to him by post. 66. In the light of the above, the Court finds it reasonable for the national authorities to regulate the manner in which prisoners may obtain photocopied or printed documents, a matter which falls within their margin of appreciation, in order to ensure the efficient functioning of all prison services. The Court therefore sees no reason in the present case to substitute its opinion for that of the Constitutional Court, which carried out a detailed and carefully balanced assessment of the competing interests involved and did not exceed its margin of appreciation. 67. Consequently, the Court concludes that the retention of the printed documents that had been sent to the applicant did not constitute a violation of Article 10 of the Convention. Court-referenced fact paragraphs: 11, 13, 14, 23, 46, 47, 48, 49, 50, 51 MODEL ASSESSMENTS A. Non-Curated (GPT-5.4) Applicability of Article 10 and existence of an interference: The measures complained of concern the applicantâs receipt of a letter from his sister and a letter from his wife, together with printed documents relating, according to him, to physiotherapy exercises and a distance-learning real-estate course (see paragraphs 5, 7, 10 and 12). Even in prison, a detainee retains the right to receive information and ideas, and restrictions on incoming written material fall within Article 10 (see nos. 39084/97, 20233/06 and 66400/01). The withholding of the first set of documents for about one year, and the complete refusal to hand over the second set, therefore amounted to an interference with the applicantâs Article 10 rights (see paragraphs 6, 9 and 11-14). Prescribed by law and legitimate aim: The interference appears to have had a basis in domestic law, namely sections 62 and 68 of Law no. 5275 and the relevant prison regulations (see paragraphs 17, 19 and 21). It also pursued at least on its face legitimate aims recognised under Article 10 § 2, notably the prevention of disorder or crime and the protection of prison security, especially since the applicant had been convicted of membership of the FET Ě O/PDY (see paragraphs 4, 6 and 11). The central issue is therefore whether the authorities showed that the restrictions were necessary in a democratic society. 20 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? The first letter: lack of an individualised assessment: As regards the first letter, the Disciplinary Board relied on general assertions that the internet printouts might threaten prison security, that their source and purpose were unclear, and that they might facilitate organisational communication within FET Ě O/PDY (see paragraph 6). However, the applicant gave a concrete and innocent explanation: part of the material concerned ankle rehabilitation after an injury and part concerned study materials needed for examinations (see paragraph 7). Importantly, the enforcement judge later accepted his objection and expressly held that the withholding had been unlawful because the authorities had relied only on a general inspection and had not conducted any specific assessment of the actual content of the documents (see paragraph 8). The material was ultimately delivered only on 25 October 2019, around a year after it had been withheld (see paragraph 9). In the Courtâs case-law, prison restrictions based on abstract risks rather than a concrete examination of the material do not satisfy the requirement of relevant and sufficient reasons (see nos. 48876/08, 56925/08 and 14946/08). The second letter: blanket exclusion of internet printouts: The position is even clearer regarding the second letter. The Disciplinary Board withheld sixty-one pages of internet printouts simply by referring to section 68(3) and to a 2016 administrative decision about general risks associated with internet printouts; it did not address the content of the withheld documents at all (see paragraph 11). The applicant again explained that the material concerned physiotherapy and real-estate management and was important for his rehabilitation and education (see paragraph 12). Nevertheless, the enforcement judge dismissed the objection on purely formal grounds, stating that internet printouts were neither âbooksâ nor âcorrespondenceâ because of their unknown origin and susceptibility to interference, and suggesting that the prison library was available instead (see paragraph 13). The assize court endorsed that reasoning (see paragraph 14). Such reasoning points to a near-automatic prohibition on a category of material, not a balancing exercise based on the specific content, the applicantâs personal needs, or any demonstrated security risk. Under the Courtâs standards, a blanket or formalistic refusal to allow access to written material in prison is difficult to reconcile with Article 10 (see nos. 20233/06, 21830/09 and 66400/01). Insufficient proportionality and weak procedural safeguards: The proportionality assessment was particularly deficient because the impugned materials were said to serve rehabilitative and educational purposes, which are ordinarily compatible with the aims of detention rather than contrary to them (see paragraphs 7 and 12). The domestic authorities did not identify any concrete phrase, passage or image that endangered security, nor did they explain why less restrictive options were unavailable, such as examining the documents individually or removing only problematic pages. The Constitutional Court then dismissed the applicantâs complaints with a brief formula referring to its previous case-law, without engaging with the specific facts of the present case, including the enforcement judgeâs finding that the first withholding had been unlawful and the absence of any content-based analysis for the second letter (see paragraphs 8, 11, 13 and 16). This weak review further supports the conclusion that the interference was not convincingly justified under Article 10 (see nos. 48876/08 and 56925/08). Overall conclusion under Article 10: Taking the two incidents together, the domestic authorities relied on vague security concerns and categorical objections to internet printouts, but failed to carry out the individualised examination required by Article 10. The first set of documents was withheld for a prolonged period despite later being found to have been blocked unlawfully, and the second set was withheld on an essentially blanket basis without any analysis of its content (see paragraphs 8, 9, 11, 13 and 14). In line with the Courtâs prison-expression jurisprudence, the reasons given were not relevant and sufficient, and the interference cannot be regarded as necessary in a democratic society (see nos. 39084/97, 20233/06, 14946/08 and 66400/01). Predicted violation: yes B. Curated Step-by-step (GPT-5.4) Existence of an interference: There was an interference with the applicantâs Article 10 rights. The prison authorities first withheld a letter from his sister, including thirty-one pages of internet printouts, on 22 October 2018, and the applicant did not receive them until 25 October 2019 after a successful objection (see paragraphs 5-9). They also withheld a second set of materials sent by his wife, consisting of sixty-one pages of internet printouts, while delivering only the handwritten note and photographs (see paragraphs 10-14). Article 10 protects not only the imparting but also the receipt of information and ideas, including by prisoners, although imprisonment allows wider restrictions than in ordinary situations (see nos. 66400/01 and 21830/09). The delayed delivery of the first materials and the complete refusal of the second clearly affected the applicantâs ability to receive educational and medical information which he said related to physiotherapy and a distance-learning real-estate course (see paragraphs 7 and 12). Lawfulness of the interference: The interference had some formal basis in domestic law, since the authorities relied on section 68(3) of Law no. 5275 for both decisions and the enforcement judge later referred, for the second letter, to section 62 of the same Law (see paragraphs 6, 11, 13 and 17-19). However, the quality of the legal basis appears deficient in its application to the present case. For the first letter, the Disciplinary Board used very general formulae about prison security and possible organisational communication, without identifying any concrete passage or making a specific content-based assessment; the enforcement judge later held that approach unlawful (see paragraphs 6-8). For the second letter, the Board relied on a prior abstract decision about risks linked to internet printouts and again did not assess the actual content at all (see paragraph 11). The enforcement judge then shifted the reasoning, saying internet printouts were neither âbooksâ nor âcorrespondenceâ because of their unknown origin (see paragraph 13). That combination of generic reasoning, inconsistent legal characterisation, and absence of individual examination points to an arbitrary and insufficiently foreseeable interference, contrary to the Courtâs approach to the âprescribed by lawâ requirement (see nos. 14946/08 and 56925/08). This strongly supports a finding that the interference was not lawful in the Convention sense. 21 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? Legitimate aims of the interference: The Government could plausibly argue that the measures pursued the prevention of disorder or crime and the protection of prison security, especially given the applicantâs conviction for membership of the FET Ě O/PDY and the prison authoritiesâ concern that the materials might facilitate communication within that organisation (see paragraphs 4, 6 and 11). Those are legitimate aims under Article 10 § 2. Still, the legitimacy of the aims does not remove the need for an individualised demonstration that these particular materials created such a risk. Here, the authorities never explained why documents said to concern ankle rehabilitation and course preparation genuinely endangered prison order or security (see paragraphs 7 and 12). Necessary in a democratic society: The interference does not appear necessary in a democratic society. The domestic authorities did not show a pressing social need supported by relevant and sufficient reasons. The first set of documents was eventually delivered in full after the enforcement judge found the original decision unlawful because it had rested only on a general inspection and not on the actual content (see paragraphs 8-9). That later delivery undermines the claim of any real security risk. As to the second set, the authorities adopted a blanket refusal of internet printouts, without examining whether any page actually contained prohibited material and without considering less restrictive options such as checking, redacting, or withholding only problematic passages (see paragraphs 10-14). The applicant gave concrete and facially legitimate reasons for needing the documents for physiotherapy and education (see paragraphs 7 and 12), and section 62 itself shows special regard for textbooks used in studies (see paragraph 17). The enforcement judgeâs reference to access through the prison library was abstract and unsupported by any indication that equivalent materials were available there (see paragraph 13). Finally, the Constitutional Court dismissed the complaint by a formulaic reference to its earlier case-law and did not perform a real balancing exercise (see paragraph 16). In line with the Courtâs case-law requiring individualised and proportionate reasons for restricting prisonersâ access to information (see nos. 20233/06, 48876/08 and 21830/09), the measures were disproportionate. Predicted violation: yes C. Curated Guide (GPT-5.4) 1. Interference with freedom to receive information: The impugned measures disclose a clear interference with the applicantâs Article 10 rights. The prison authorities withheld enclosures sent by the applicantâs sister and wife, consisting of internet printouts which, according to the applicant, concerned physiotherapy exercises for his injured ankle and study materials for a distance-learning real-estate course (see paragraphs 7 and 12). The first set of documents was delivered only about one year later, after successful judicial challenge (see paragraphs 8-9), and the second set was never delivered (see paragraphs 11-14). Article 10 protects not only imparting but also receiving information, and prisoners retain that protection, subject to justified prison restrictions. Measures preventing access to written material, even in the prison context, are therefore interferences requiring justification (see nos. 39084/97, 14946/08, 56925/08). 2. Legal basis and legitimate aim: The authorities relied on sections 68 and 62 of Law no. 5275, which regulate prisonersâ correspondence and access to publications (see paragraphs 17 and 19). The aims invoked or inferable from the decisions were prison order and security, and prevention of communication within a terrorist organisation (see paragraphs 6 and 11), which are in principle legitimate under Article 10 § 2. However, the lawfulness of the interference is doubtful. In respect of the first letter, the enforcement judge expressly found the withholding unlawful because the administration had relied on a general and abstract assessment rather than the actual content of the documents (see paragraph 8). In respect of the second letter, the authorities oscillated between the regimes for correspondence and publications and ultimately treated internet printouts as neither, largely because of their form and supposed unknown origin (see paragraphs 11 and 13). Such formalistic and inconsistent classification weakens foreseeability and safeguards against arbitrariness. Still, even assuming the interference was prescribed by law, the decisive issue is necessity and proportionality. 3. Necessity and proportionality: first letter: As to the first letter, the Disciplinary Board stated that the enclosures might threaten prison security, that their publisher and purpose were unclear, and that they contained phrases that could facilitate communication within FET Ě O/PDY (see paragraph 6). Yet no concrete passage or identifiable risk was cited. The domestic court later accepted that withholding the material merely because its content was unclear upon a general inspection, without a specific assessment of what the documents actually contained, was incompatible with freedom of expression and communication (see paragraph 8). That finding is highly significant. Moreover, the delayed delivery deprived the material of much of its practical value, since the applicant had said the documents were needed both for ankle rehabilitation and exam preparation (see paragraph 7). A delay of nearly a year in giving access to allegedly innocuous educational and medical material cannot be regarded as a proportionate response to vague security concerns. Strasbourg case-law requires relevant and sufficient reasons and an individualised assessment, not speculative assertions (see nos. 66400/01, 20233/06, 56925/08). 4. Necessity and proportionality: second letter: The second withholding appears even harder to justify. The Disciplinary Board withheld sixty-one pages of internet printouts by relying on section 68(3) and on a prior general decision warning of the risks of internet printouts in prisons, while authorising the handwritten note and photographs (see paragraph 11). Crucially, it did not address the content of the withheld material at all (see paragraph 11). The enforcement judge then rejected the applicantâs objection on the basis that the documents were neither âbooksâ nor âcorrespondenceâ, because of their unknown origin and susceptibility to interference, and added that prisoners could use the library (see paragraph 13). That reasoning is overly formalistic and close to a blanket exclusion of internet printouts as a category. Article 10 does not permit such automatic restrictions where the authorities do not show why the specific material creates a genuine security problem. The applicant had directly explained that the documents related to physiotherapy and his studies (see paragraph 12), and there is nothing in the facts suggesting violence, propaganda, coded messages, or any concrete 22 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? institutional risk. The library argument is also weak, because no court established that equivalent material was actually available there or that it would meet the applicantâs immediate medical and educational needs. Blanket or category-based restrictions on access to written material, without content analysis, are generally incompatible with the proportionality test (see nos. 14946/08, 48876/08, 21830/09). 5. Quality of domestic review: The domestic review did not cure the interference. Although the enforcement judge remedied the first withholding, that occurred only after a lengthy period and without compensation for the delayâs effects (see paragraphs 8-9). As to the second withholding, the enforcement judge and Assize Court essentially endorsed a categorical exclusion based on form rather than substance (see paragraphs 13-14). The Constitutional Court then dismissed the complaint in a formulaic manner, stating either that there had been no interference or that any interference had not amounted to a violation, with a bare reference to its earlier case-law (see paragraph 16). That reasoning did not engage with the applicantâs concrete explanations, the actual content of the documents, or the balancing required by Article 10. Where domestic courts fail to apply Convention standards and give relevant and sufficient reasons, the Strasbourg Court is more likely to find a violation (see nos. 39084/97, 20233/06, 56925/08). 6. Overall conclusion: Even allowing the State a relatively wider margin of appreciation in prison-security matters, especially concerning a prisoner convicted of membership of a terrorist organisation (see paragraph 4), the restrictions here appear disproportionate. The decisive features are the absence of a specific content-based assessment, the reliance on speculative or generic security concerns, the formalistic exclusion of internet printouts as a category, the substantial delay in delivery of the first material, and the insufficiency of the domestic courtsâ reasoning (see paragraphs 6-16). In light of Article 10 principles requiring careful, individualised justification for restrictions on receiving information, the Court would likely conclude that the applicantâs freedom of expression was violated (see nos. 66400/01, 14946/08, 56925/08). Predicted violation: yes D. Non-Curated (Mistral Medium 3.5) (supplementary; not in main paper) Establishment of interference: The withholding of two letters and their enclosures addressed to the applicant by the prison authorities constitutes an interference with his right to receive information, which falls within the scope of Article 10 of the Convention (see paragraphs 5-14). Prescribed by law and legitimate aim: The interference was based on sections 68(3) and 62 of Law no. 5275 and pursued legitimate aims under Article 10(2), namely the prevention of disorder and crime in the prison context (see paragraphs 6, 11, 17-19). Lack of specific content assessment â first letter: The Disciplinary Boardâs decision to withhold the first letter was based on general concerns about prison security and unclear publication details, without a specific assessment of the documentsâ content (see paragraph 6). The enforcement judge later ruled this decision unlawful precisely for this reason, noting that withholding based solely on general inspection without content analysis violated freedom of communication and expression (see paragraph 8). The documents contained physiotherapy and educational material, which are particularly protected forms of information (see paragraph 7). Blanket restriction â second letter: The second letterâs enclosures were withheld based on a blanket policy against internet printouts, as evidenced by the reference to a 2016 decision, without any examination of their specific content (see paragraph 11). The enforcement judge and Assize Court upheld this on formal grounds, ignoring the substance of the materials (see paragraphs 13-14). The Constitutional Courtâs dismissal relied on a formalistic distinction between document types, rather than a proportionality assessment (see paragraph 16). Disproportionate interference under ECtHR standards: The ECtHR has established that restrictions on freedom of expression must be necessary in a democratic society, requiring a pressing social need and proportionality (see no. 48876/08). Blanket restrictions without individual content assessment are generally incompatible with Article 10 (see no. 56925/08). Here, neither withholding decision demonstrated a concrete, content-based justification. The legitimate purposes of the materialsâmedical rehabilitation and educationâfurther underscore the disproportionate nature of the restrictions (see paragraphs 7, 12). Failure of domestic review to meet Convention standards: While the domestic enforcement judge eventually rectified the withholding of the first letter, the initial interference lacked a lawful basis under national standards (see paragraph 8). For the second letter, all domestic instances upheld the restriction based on the formal classification of the documents as internet printouts, rather than their actual content. This approach fails to meet the ECtHRâs requirement for a specific, proportional assessment under Article 10 (see no. 21830/09). Predicted violation: yes E. Curated Step-by-step (Mistral Medium 3.5) (supplementary; not in main paper) Existence of an interference: The withholding of the applicantâs incoming letters and their enclosures by the prison authorities clearly constituted an interference with his right to receive information under Article 10 of the Convention. The applicant was prevented from accessing materials sent to him, which affected his freedom of expression (see paragraphs 5, 6, 10, 11). 23 Can LLMs Reason in a Legally Meaningful Manner on ECtHR cases? Lawfulness of the interference: The interference had a legal basis in domestic law. The withholding of the letters was carried out pursuant to section 68(3) of Law no. 5275, which explicitly permits the restriction of correspondence that poses a threat to prison order or security or facilitates communication between members of terrorist or criminal organisations (see paragraphs 6, 11, 19). Additionally, the enforcement judgeâs reasoning for the second letter cited section 62 of the same law (see paragraph 13). Legitimate aims of the interference: The interference pursued legitimate aims under Article 10(2) of the Convention. The Disciplinary Board justified the withholding of the first letter on the grounds that its enclosures could pose a threat to prison security and facilitate communication within the FET Ě O/PDY organisation (see paragraph 6). For the second letter, the Board relied on prior guidance concerning the risks associated with internet printouts being handed over to prisoners (see paragraph 11). These reasons align with the legitimate aims of preventing disorder or crime and protecting national security. Necessary in a democratic society: The interference was not necessary in a democratic society. Regarding the first letter, the domestic enforcement judge ruled that withholding the documents solely on the grounds that their content was unclear from a general inspection, without a specific assessment of their actual content, was unlawful (see paragraph 8). For the second letter, the Disciplinary Boardâs decision did not address the content of the withheld documents at all, instead applying a blanket policy against internet printouts (see paragraph 11). The applicant demonstrated that the materials served legitimate purposes related to his medical rehabilitation and education (see paragraphs 7, 12). The ECtHR has consistently required that restrictions on prisonersâ correspondence be based on a specific, individualized assessment of the content and be proportionate (see no. 48876/08). The lack of such an assessment, combined with the availability of less restrictive measures (e.g., content review or provision of materials through approved channels), renders the interference disproportionate and unjustified. Predicted violation: yes 24