Paper deep dive
The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations
Ilya Mikhelson
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 8/3/2026, 3:14:24 AM
Summary
This paper introduces the 'Socratic Test,' an automated, computer-mediated conversational assessment system designed to overcome the limitations of static written exams and traditional oral examinations. By integrating Dynamic Assessment principles, Bloom's Taxonomy for prompt generation, and the SOLO Taxonomy for response evaluation, the system maps a student's Zone of Proximal Development (ZPD). It employs a non-compensatory, additive grading architecture with multimodal workspaces to reduce anxiety and ensure reliable measurement of cognitive boundaries.
Entities (13)
Relation Signals (12)
Socratic Test â uses â Dynamic Assessment
confidence 95% · The Socratic Test is rooted in Dynamic Assessment (DA)
Socratic Test â uses â Bloom's Taxonomy
confidence 95% · The proctor utilizes Bloomâs (revised) Taxonomy to structure its prompting.
Socratic Test â uses â SOLO Taxonomy
confidence 95% · The AI utilizes the SOLO Taxonomy ... to measure the structural quality of the response
Socratic Test â measures â zone of proximal development
confidence 92% · This paper formalizes the use of graduated scaffolding to quantify the Zone of Proximal Development (ZPD)
Socratic Test â implements â Non-Compensatory Grading
confidence 90% · the architecture rejects Compensatory Grading ... in favor of Non-Compensatory Grading.
Socratic Test â contains â Shadow Ledger
confidence 85% · the system must prevent students from exploiting the conversational interface ... via a hidden Shadow Ledger
Socratic Test â contains â Evidence Buffer
confidence 85% · the system utilizes an Evidence Buffer parameterized by an Oversampling Factor.
Dynamic Assessment â derivedfrom â Lev Vygotsky
confidence 85% · Dynamic Assessment (DA), derived from Lev Vygotskyâs theory of cognitive development
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic feedback. Conversely, traditional face-to-face oral examinations introduce severe construct-irrelevant variance by exacerbating performative anxiety and the sociological power imbalances inherent to academic hierarchies. This paper presents the theoretical foundation for the "Socratic Test," an automated, computer-mediated conversational assessment. By integrating Dynamic Assessment principles, multimodal workspaces, Bloom's Taxonomy for real-time proctoring, and the SOLO Taxonomy for structural evaluation, the Socratic Test actively maps a student's cognitive boundaries. This paper formalizes the use of graduated scaffolding to quantify the Zone of Proximal Development (ZPD) and details a non-compensatory, additive grading architecture that prioritizes mastery over penalty and human-AI alignment to ensure unprecedented measurement reliability.
Tags
Links
- Source: https://arxiv.org/abs/2607.29624v1
- Canonical: https://arxiv.org/abs/2607.29624v1
Trouble viewing inline? Open PDF directly â
Full Text
56,509 characters extracted from source content.
Expand or collapse full text
The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations Ilya Mikhelson ilya@northwestern.edu Abstract Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic feedback. Conversely, traditional face-to-face oral examinations introduce severe construct-irrelevant variance by exacerbating performative anxiety and the sociological power imbalances inherent to academic hierarchies. This paper presents the theoretical foundation for the âSocratic Test,â an automated, computer-mediated conversational assessment. By integrating Dynamic Assessment principles, multimodal workspaces, Bloomâs Taxonomy for real-time proctoring, and the SOLO Taxonomy for structural evaluation, the Socratic Test actively maps a studentâs cognitive boundaries. This paper formalizes the use of graduated scaffolding to quantify the Zone of Proximal Development (ZPD) and details a non-compensatory, additive grading architecture that prioritizes mastery over penalty and human-AI alignment to ensure unprecedented measurement reliability. keywords: Socratic Test , Dynamic Assessment , Conversational AI , Large Language Models , Automated Grading â journal: Computers and Education: Artificial Intelligence [1]organization=Department of Electrical and Computer Engineering, Northwestern University, city=Evanston, state=IL, country=USA 1 Background and Motivation 1.1 Written Examinations - Reliability over Validity The prevailing assessment paradigm in higher education relies heavily on static, written examinations evaluated via a subtractive grading model (Feldman, 2023). Students begin with a theoretical perfect score, and points are deducted for errors, omissions, or misapplications. While computationally efficient for instructors and highly reliable (capable of being graded consistently), written exams often lack validity. They frequently fail to measure true competence, allowing students to mask knowledge gaps through test-taking strategies or rote memorization (Frederiksen, 1984; Struyven et al., 2005). Furthermore, open-ended static assessments frequently suffer from the âpunishing ambitionâ problem. While individual instructors may attempt to grade holistically, the underlying mathematical structure of a static exam systemically incentivizes risk aversion. For instance, a student who attempts a highly sophisticated, novel synthesis but makes a minor structural error is mathematically penalized compared to a student who safely executes a rudimentary, surface-level response. This subtractive grading model conflates diagnostic feedback with punitive measures, ultimately narrowing thinking, reducing intrinsic motivation, and conditioning students into compliance rather than deep engagement (Feldman, 2023; Kohn, 1993). 1.2 Oral Examinations - The Standardization Fallacy While there are numerous alternative to written examinations, such as essays, portfolios, and performances, the most direct alternative for probing a studentâs cognitive boundaries has been the oral examination. Oral exams exhibit incredibly high validity, as they probe the edges of a studentâs domain knowledge and professional readiness (Huxham et al., 2012; Joughin, 1998). However, their traditional reputation is complicated by two major critiques: one a genuine limitation, and the other a pedagogical illusion. The genuine limitation is the introduction of significant construct-irrelevant variance. Face-to-face interrogations raise a studentâs Affective Filter, a psychological barrier of anxiety and fear of judgment that blocks working memory (Krashen, 1982). Consequently, oral exams often measure a studentâs public speaking proficiency and stress management rather than their pure domain knowledge, with literature frequently citing the format as a source of severe, debilitating anxiety for undergraduate students (Joughin, 1998; Iannone et al., 2020; Laurin-Barantke et al., 2016). The second critique is a perceived lack of equity. Because a proctor must ask different questions to different students to map their unique knowledge bounds, critics argue the exam is not standardized. However, this critique is an illusion stemming from the Standardization Fallacy, i.e., the misconception that standardizing the studentâs experience (identical questions) is required to standardize the measurement (Wainer et al., 2000; Lord, 2012). The theoretical precedent for dismissing this fallacy is firmly established in Computerized Adaptive Testing (CAT) (Lord, 2012). High-stakes professional licensure exams (e.g., NCLEX, USMLE) and admissions exams (e.g., GMAT, SAT) operate on adaptive algorithms where no two candidates receive the same questions. True equity is achieved through criterion-referenced grading, where the specific conversational path varies, but the structural threshold required to prove mastery remains immutable (Biggs, 1996). 2 Theoretical Framework of the Socratic Test To resolve the tension between the reliability of written exams and the validity of oral exams, I propose the Socratic Test. This framework utilizes an adaptive, typed conversational assessment driven by artificial intelligence (AI) to standardize the methodology of measurement without standardizing the specific conversational prompts. To achieve this, the platform relies on two foundational pedagogical frameworks: one to govern the complexity of the proctorâs questions, and another to measure the structural quality of the studentâs answers. These frameworks are defined, compared, and contrasted in Table 1. Table 1: Foundational Cognitive Frameworks of the Socratic Test Level Bloomâs Taxonomy (Prompt Objective) (Krathwohl, 2002) SOLO Taxonomy (Response Structure) (Biggs and Collis, 2014) 1 Remember: Retrieve relevant factual knowledge and definitions from long-term memory. Prestructural: The response misses the point entirely or relies on irrelevant information. 2 Understand: Construct meaning from instructional messages (e.g., interpreting, summarizing). Unistructural: The response correctly identifies or utilizes a single relevant aspect. 3 Apply: Carry out or use a procedure in a specific, given situation. Multistructural: The response identifies several relevant independent aspects but fails to connect them. 4 Analyze: Break material into constituent parts and determine how they relate to one another. Relational: The response successfully integrates the disparate aspects into a coherent structure. 5 Evaluate: Make structural judgments based on established criteria and standards. Extended Abstract: The response generalizes the integrated principle to a completely novel domain. 6 Create: Reorganize elements into a novel, coherent whole or original product. (The SOLO framework concludes at Level 5) 2.1 Dynamic Assessment and the ZPD The Socratic Test is rooted in Dynamic Assessment (DA) (Lantolf and Poehner, 2004), derived from Lev Vygotskyâs theory of cognitive development (Vygotsky et al., 1978). Unlike static assessments, which only measure what a student has already mastered, DA integrates instruction and assessment to measure a studentâs Zone of Proximal Development (ZPD), i.e., the space between what a learner can do independently and what they can achieve with guidance (Lantolf and Poehner, 2004). By observing how a student responds to conversational hints, the system mathematically differentiates between Independent Performance (unassisted success) and Assisted Performance (success achieved via scaffolding). Evaluating the full breadth of the ZPD yields a much truer measure of systemic understanding, as it prevents a studentâs demonstrated competence from being artificially capped by a momentary lapse in foundational memory. 2.2 Multimodal Evidence and Cognitive Offloading A pure text-based chat is often insufficient for evaluating highly complex, spatial, or mathematical competencies across various disciplines. The Socratic Test interface includes integrated workspaces, such as a whiteboard, an integrated development environment (for code), a calculator, and a long-form response interface, which serve a critical pedagogical function known as Cognitive Offloading (Kirsh, 1995). By allowing students to sketch diagrams or draft code, the workspaces act as external working memory, reducing intrinsic cognitive load (Sweller, 1988). More importantly, capturing this offloaded work natively within a digital drawing and text interface provides the AI with a significantly richer, multimodal evidence trail. This interface specifically protects students who may struggle with English prose articulation; rather than relying solely on language proficiency, the AI evaluates the structural logic of a hand-drawn schematic or a mathematical derivation. Furthermore, transitioning from a face-to-face interrogation to a typed, asynchronous interface is highly likely to lower the studentâs affective filter. Decades of research into Computer-Mediated Communication (CMC) demonstrate that typed interfaces significantly reduce communication apprehension and performative anxiety (Satar and Ăzdener, 2008; High and Caplan, 2009; Krashen, 1982). By removing the human authority figure, the interface also mitigates the intimidating power dynamics of traditional oral testing, creating an environment focused on higher-order cognitive skills rather than stress management. 2.3 Proctoring with Bloomâs Taxonomy To navigate the ZPD systematically, the AI proctor utilizes Bloomâs (revised) Taxonomy (Krathwohl, 2002) to structure its prompting. The proctor begins with low-cognitive-load questions and incrementally elevates the complexity. While cognitive mapping is not strictly linear, Bloomâs Taxonomy provides the AI with an efficient, structured heuristic for navigating the ZPD. Foundational recall (rote memory) serves as the indispensable vocabulary of any discipline (Willingham, 2009). Establishing this vocabulary early allows the proctor to verify that the student possesses the necessary foundational tools before advancing to complex analysis. When a student struggles, the proctor employs the System of Least Prompts (graduated prompting) (Campione and Brown, 1987), offering a strict hierarchy of scaffolding: 1. General Nudge: A non-directive prompt to rethink the premise. 2. Specific Cue: A directive prompt pointing to a specific missing concept. 3. Targeted Scaffold: A highly constrained prompt isolating the exact point of failure. 4. Direct Instruction (Graceful Exit and Pivot): If the student exhausts the prior three hints, the AI explicitly provides the missing foundational knowledge (e.g., giving the student the forgotten formula). Crucially, this does not terminate the topic. Because knowledge is not strictly hierarchical, a student may fail a Bloom 1 recall question but still possess Bloom 4 relational mastery (Krathwohl, 2002; Agarwal, 2019). By supplying the missing definition, the AI bridges the gap, allowing the conversation to proceed. Therefore, exhausting the scaffolding hierarchy on a single prompt does not prematurely terminate the entire topic, nor does it automatically finalize the current tier. Instead, it finalizes the studentâs score for that specific interaction, effectively defining their Cognitive Ceiling for that discrete skill. The AI logs the interaction, awards zero points to the studentâs earned score, and utilizes the Graceful Exit to explicitly provide the missing foundational knowledge. Rather than arbitrarily unlocking the next cognitive tier, this failed interaction is mathematically accounted for via a âShadow Ledgerâ (as detailed in Section 4.1). The proctor will continue generating parallel prompts within the current cognitive tier until the total evidence buffer is mathematically satisfied. The AI only pivots vertically to the next tier, or horizontally to an entirely new competency, once these rigorous evidentiary thresholds have been met, ensuring the student is never artificially capped by a single momentary lapse in recall. 2.4 Assessing with the SOLO Taxonomy While Bloomâs Taxonomy dictates the cognitive difficulty of the prompt, the AI utilizes the SOLO Taxonomy (Structure of the Observed Learning Outcome) to measure the structural quality of the response (Biggs and Collis, 2014). The decision to decouple the prompting framework from the evaluation framework is rooted in the pursuit of grading objectivity and transparency. In traditional static assessments, or even when attempting to evaluate responses using Bloomâs Taxonomy, grading is often inherently subjective. A student reviewing a marked exam might argue that their answer demonstrated âAnalysisâ rather than mere âApplication,â leading to contentious appeals over partial credit. The SOLO Taxonomy neutralizes this ambiguity by shifting the evaluation from semantic interpretation to structural complexity. The grading criteria become strictly observable and defensible: a studentâs response objectively contains either a single isolated fact (Unistructural), multiple unconnected facts (Multistructural), or logically integrated facts (Relational), as detailed in Table 1. This structural clarity makes the rubric highly transparent, streamlines the post-exam appeals process (Section 4.3.1), and fosters deep student trust in the platformâs fairness. A natural pedagogical question arises: how does the AI anchor its application of SOLO to a highly specific, advanced topic during the live conversation, before the instructor has provided any graded examples? During real-time proctoring, the AI relies on a discipline-agnostic, zero-shot system prompt anchored strictly to the universal structural definitions of the taxonomy. Rather than attempting to evaluate domain-specific semantic nuance in real-time, the AI acts as a structural parser. This tentative, real-time assessment acts purely as a navigational routing mechanism to keep the conversation flowing. The system operates conservatively; if the AI is uncertain whether a response meets the expected structural threshold, it explicitly asks the student for clarification rather than advancing. Crucially, any inaccuracies in this baseline real-time anchoring are mathematically absorbed by the Oversampling Evidence Buffer (detailed in Section 4.1.1) and subsequently corrected during the deterministic, post-exam grading pipeline (Section 4.3.1). 3 Test Progression and Exam Configuration 3.1 Mastery Accumulation and Dual Proctoring Modes In the Socratic Testâs additive grading model, students start at zero and accumulate evidence of knowledge. However, raw accumulation introduces the gamification loophole of âgrindingâ, where a student answers dozens of foundational questions to achieve a high score without ever demonstrating higher-order thinking. To prevent this, the architecture rejects Compensatory Grading (where low-level skills can mathematically compensate for a lack of high-level skills) in favor of Non-Compensatory Grading. Rooted in the theory of Constructive Alignment (Biggs, 1996) (i.e., the pedagogical principle that assessment tasks and grading criteria must strictly align with intended learning outcomes), the system utilizes âcapped bucketsâ for different cognitive tiers, configured by the instructor prior to the exam via one of two distinct proctoring modalities: Stair-Step Mode (Section 3.1.1) or Organic Mode (Section 3.1.2). 3.1.1 Stair-Step Mode - Structured Scaffolding In Stair-Step mode, the instructor defines a rigorous 2D grading matrix while setting up the exam. They configure distinct cognitive tiers (e.g., Foundational, Application, Synthesis) and map specific Bloomâs levels to each tier. For example, an introductory course may map Bloom 1 and 2 to Foundational, while an advanced seminar might eliminate Bloom 1 entirely and begin the Foundational tier at Bloom 3. The instructor then defines the overarching exam topics. For each topic, they assign a specific point capacity to each cognitive tier, as well as a target time limit. Each cell in this matrix acts as a capped bucket. An example of such a matrix for an introductory Economics course can be seen in Table 2. Because the buckets are topic- and tier-specific, a student cannot mathematically compensate for dodging an Application question on Market Failure by answering an excess of Foundational questions on Supply & Demand. Table 2: Conceptual 2D Non-Compensatory Grading Matrix (Stair-Step Mode) Topic (tâTtâ T) Foundational Application Synthesis Topic Max Topic A (e.g., Supply & Demand) Max 10 Max 10 Max 10 30 Topic B (e.g., Market Failure) Max 10 Max 15 Max 15 40 Topic C (e.g., Elasticity) Max 10 Max 10 Max 10 30 Cognitive Tier Max 30 35 35 Total: 100 3.1.2 Organic Mode - Fluid Exploration While Stair-Step mode is ideal for structured knowledge assessment, Organic mode is designed for fluid, holistic exploration, akin to a traditional graduate defense. In Organic mode, the 2D matrix collapses into a 1D vector. The instructor no longer defines rigid cognitive tiers; instead, they define the broad exploration space by selecting the allowable Bloomâs levels for the exam (e.g., testing exclusively at Bloom 3 through 6). Correspondingly, instructors assign point capacities broadly at the topic level, rather than the tier level. The proctor initiates an open-ended conversational anchor and allows the interaction to evolve organically, shifting Bloomâs levels in real-time based on the studentâs conversational direction until the topicâs overall point cap is satisfied. Because every individual interaction is still tagged with a specific Bloomâs level and hint count, the underlying mathematical grading engine remains entirely unchanged. 3.2 Mathematical Formulation To formalize this non-compensatory structure, the core system variables and their domains are defined in Table 3. Table 3: System Variables and Nomenclature Variable Definition Domain / Range t A specific domain topic being tested. tâTtâ T l A cognitive grading tier (e.g., Foundational). lâLlâ L b The Bloomâs Taxonomy level of the prompt. bâ1,2,3,4,5,6bâ\1,2,3,4,5,6\ s The SOLO Taxonomy level of the response. sâ1,2,3,4,5sâ\1,2,3,4,5\ k The number of scaffolding prompts required. kâ0,1,2,3,4kâ\0,1,2,3,4\ Îłk _k The scaffolding discount factor. Îłkâ[0,1] _kâ[0,1] Vâ(b,s)V(b,s) Base point value (monotonically increasing with b). Vâ„0Vâ„ 0 Pâ(b,s,k)P(b,s,k) Total points generated by a single interaction. Pâ„0Pâ„ 0 Mt,lM_t,l Maximum allowable points (cap) for a specific bucket. Mt,l>0M_t,l>0 Let tâTtâ T be a specific topic and lâLlâ L be a cognitive tier (e.g., Foundational, Application). Let bâ1,2,3,4,5,6bâ\1,2,3,4,5,6\ be the Bloomâs level of the prompt. The tiers represent a mapping of multiple Bloomâs levels. For instance, Bloom 1 and 2 can map to the Foundational tier. Let sâ1,2,3,4,5sâ\1,2,3,4,5\ be the SOLO level of the response, and kâ0,1,2,3kâ\0,1,2,3\ be the number of scaffolding prompts required. The point value P generated by a single interaction is: Pâ(b,s,k)=ÎłkâVâ(b,s)P(b,s,k)= _kV(b,s) (1) where Vâ(b,s)V(b,s) is the base point value of achieving SOLO level s on a Bloom level b prompt, and Îłkâ[0,1] _kâ[0,1] is the scaffolding discount factor, quantifying the reduction in value from Independent to Assisted Performance. For example, an instructor may configure Îł0=1.0 _0=1.0 (no hints, full credit), Îł1=0.8 _1=0.8 (one hint), down to Îł4=0.0 _4=0.0 for a Graceful Exit. Crucially, Vâ(b,s)V(b,s) is a monotonically increasing function with respect to b; a simple Recall prompt (Bloom 1) yields fewer points than a complex Create prompt (Bloom 6). To enforce the non-compensatory structure, the total grade is calculated by aggregating points across all specific buckets (topic t, tier l): Total Score=âtâTâlâLminâĄ(Mt,l,âPt,l)Total Score= _tâ T _lâ L (M_t,l,ÎŁ P_t,l ) (2) where Mt,lM_t,l represents the maximum allowable points for that specific matrix cell, and Pt,lP_t,l is the point value from Eq. (1) earned within that cell. Once a bucket is full, the AI forces vertical or horizontal progression; answering further low-level questions within that topic yields zero additional points toward the final grade. A complete simulated transcript demonstrating the test progression of Sections 2 and 3 is provided in Appendix A. 4 System Architecture and Implementation To deploy this theoretical framework safely within high-stakes academic environments, I developed socratictest.com, a custom software platform explicitly engineered to administer the Socratic Test modality. While recent literature has explored the use of Large Language Models (LLMs) as conversational tutors for formative feedback (Favero et al., 2024; Liu et al., 2024), socratictest.com is specifically architected as a summative assessment engine. The platform relies on a decoupled architecture separating the real-time proctoring engine from the final deterministic gradebook. 4.1 Proctoring Mechanics - Buffers, Ledgers, and Pivoting During the live exam, the AI navigates an internal state machine governed by continuous mathematical accounting. The proctorâs primary imperative is to gather sufficient conversational evidence to satisfy the point caps defined in the instructorâs grading matrix (Table 2). To execute this equitably, the state machine utilizes three core mechanics, detailed below. 4.1.1 The Evidence Buffer and Oversampling Because the AIâs real-time assessment of a studentâs SOLO level is strictly provisional (awaiting instructor calibration in Step 2 of the grading pipeline (Section 4.3.1)), the state machine cannot rely on strict minimum SOLO thresholds to gate student progression. Doing so risks trapping a student behind an AI hallucination. Instead, the system utilizes an Evidence Buffer parameterized by an Oversampling Factor. Prior to the exam, the instructor defines an oversample rate (e.g., 120%). During a Stair-Step exam, if a tier requires 10 points to fill, the AI will continue prompting the student until its provisional real-time evaluation assesses that 12 points worth of evidence have been gathered. This mathematical buffer protects the student; if the instructor retroactively downgrades a specific interactionâs SOLO score during post-exam calibration, the oversampled evidence buffer ensures the student was not unfairly denied the opportunity to earn the tierâs maximum points. 4.1.2 The Shadow Ledger To maintain exam integrity, the system must prevent students from exploiting the conversational interface by skipping difficult questions or intentionally exhausting hints to bypass a topic. The platform achieves this via a hidden Shadow Ledger of âattempted points.â When a student explicitly requests to skip a prompt, or when they exhaust the scaffolding hierarchy (more than 3 hints, detailed in Section 2.3) triggering a Graceful Exit, the AI complies and pivots. However, it simultaneously logs the baseline point value of the abandoned prompt into the Shadow Ledger. This baseline value is deterministically anchored to the Expected Strong Response defined in Table 4 (e.g., skipping a Bloom 3 prompt logs the expected Vâ(3,4)V(3,4) point value into the ledger). The Vertical Gate (i.e., the blocker to proceed to a higher tier) for a cognitive tier only opens when the sum of the studentâs Earned Points plus their Shadow Ledger Points meets the oversampled tier capacity. Consequently, skipping a question mathematically accelerates the closure of the cognitive tier without awarding the student evidence points, permanently limiting their potential score. Conversely, if a student struggles but eventually arrives at the correct answer using hints, no shadow points are levied; they simply earn their standard, fractionally discounted points, preserving their ability to continue gathering evidence in that tier. Table 4: Shadow Ledger Baseline Matrix: Anchoring Bloomâs Prompts to Expected SOLO Responses AI Prompt (Bloomâs Level) Prompt Ceiling (Max SOLO) Expected Strong Response Baseline (Shadow Ledger) Bloom 1 SOLO 3 SOLO 2 Bloom 2 SOLO 3 SOLO 3 Bloom 3 SOLO 4 SOLO 4 Bloom 4 SOLO 4 SOLO 4 Bloom 5 SOLO 5 SOLO 4 Bloom 6 SOLO 5 SOLO 5 4.1.3 User-Directed Pivoting To maximize student agency and cognitive offloading, the platform abandons forced temporal constraints in favor of user-directed pivoting. At any point, a student can access a navigation menu detailing the time spent on each topic and elect to pivot horizontally to a new domain. This pivot incurs no penalty. When the student eventually returns to the abandoned topic, the state machine resurrects the exact conversational context, Evidence Buffer, and Shadow Ledger state from the moment of departure, preventing pivoting from being used as an evasive tactic. (Note: Students may pivot between overarching topics, but they cannot manually pivot between cognitive tiers within a topic, as foundational tiers must be completed to scaffold advanced topics). Furthermore, when a student exceeds the instructor-defined target time for a topic, the interface issues a non-blocking visual warning. This timer grounds the studentâs pacing without forcefully interrupting their cognitive flow. 4.2 Mitigating Hallucination and the âOut-of-Scopeâ Protocol A primary concern when deploying LLMs in assessment is the risk of AI hallucination, where the proctor might ask questions about out-of-scope material or state incorrect premises. The Socratic Test handles this through an asynchronous audit mechanism that is strictly mathematically bound to the real-time Shadow Ledger to prevent gamification. During the live exam, if a student claims a premise is out-of-scope or identifies an AI error, the proctor accepts the studentâs assertion, flags the interaction for audit, and pivots to a new question. However, the state machine must prevent the student from exploiting this feature to infinitely cycle questions. Therefore, the system initially treats an âOut-of-Scopeâ flag identically to a standard âSkip.â It immediately logs 100% of the promptâs baseline point value (Table 4) into the studentâs Shadow Ledger, and the AI pivots to the next prompt. This provides a critical real-time safeguard: a student continuously claiming out-of-scope will mathematically exhaust the tierâs Evidence Buffer, definitively limiting their potential score and moving the exam forward. In a traditional subtractive system, adjudicating a dodge requires applying an arbitrary penalty. In the Socratic Testâs additive, non-compensatory framework, no punitive deduction is necessary. If a student maliciously dodges a valid topic, they simply fail to fill the non-compensatory bucket for that specific domain, inherently limiting their final grade. The equity of this mechanism relies entirely on the post-exam instructor audit via a four-point scale: 1. Student Correct: The instructor verifies the AI erred. Because identifying a structural error proves advanced domain mastery (Krathwohl, 2002), the system converts the initial Shadow Ledger penalty directly into Earned Points for that tier. The student gets full credit for the skipped prompt. 2. Valid Confusion: The premise was technically valid, but poorly phrased. The system gives the student the benefit of the doubt, converting the Shadow points into Earned points. 3. Partial Evasion: The student deflected a valid premise but attempted some engagement. The instructor retroactively flags the interaction as a partial skip. The system leaves 50% of the point value in the Shadow Ledger as a permanent penalty, converting the remaining 50% to Earned points. 4. Complete Evasion: The student abused the mechanism to dodge a valid question, or they did not know how to answer the question. In either case, it is identical to a Skip. The 100% Shadow Ledger penalty applied during the live exam remains permanent. By shifting the burden of trust from the live AI to the post-exam audit, this architecture completely neutralizes the risk of students gaming the conversational interface, while guaranteeing they are mathematically rewarded for catching an algorithmic mistake. An example of this can be seen in Appendix A. 4.3 Interaction-Level Grading and Human-AI Alignment A persistent critique of integrating AI into higher education is the perceived loss of instructor oversight and the inherent unreliability of algorithmic evaluation. However, this skepticism often ignores the profound fallibility of traditional human grading. Literature confirms that human marking suffers from severe inter-rater and intra-rater reliability issues, driven by grader fatigue and fluctuating internal standards (Bloxham et al., 2011). The Socratic Test platform mitigates the unreliability of both humans and autonomous algorithms by deploying a deterministic, interaction-level grading pipeline. Rather than attempting to grade an entire transcript holistically, which introduces systemic variance, the architecture isolates the single subjective variable (the SOLO level) from the objective variables recorded during the live exam (the Bloomâs level and the hint count). To ensure absolute fidelity, the system utilizes a calibration and validation loop as part of a 7-step grading pipeline, detailed below. 4.3.1 The 7-Step Alignment and Grading Pipeline Step 1: The Out-of-Scope Audit and Global Rules. During the live assessment, the proctor automatically tags interactions where a student flags a premise as out-of-scope or identifies an AI hallucination. The instructor rapidly audits these real-time tags. If a hallucination is verified, the flagging student is rewarded (as detailed in Section 4.2). Crucially, to protect systemic equity, this verification automatically generates a âGlobal Exclusion Ruleâ within the grading engine. This ensures that any other student in the cohort who encountered the exact same hallucinated premise, but lacked the assertiveness to challenge the AI, is mathematically protected. Step 2: Ambiguity Calibration. Following the audit, the AI performs an initial sweep of every interaction across the entire cohort. It predicts a provisional SOLO score for each interaction and calculates a confidence interval. The system then isolates the 20 most ambiguous interactions for calibration. To prevent a single highly unconventional student from skewing the model, the system enforces a strict cap of 4 calibration interactions per student. The instructor manually reviews these 20 edge-case interactions and assigns the definitive SOLO scores, establishing the semantic baseline. Step 3: Inter-Rater Validation. Using the calibrated examples, the AI attempts to assign SOLO scores to a new, random validation set of 15 interactions. The instructor reviews these interactions and simply marks âAgreeâ or âDisagreeâ with the AIâs assessment. If the instructor disagrees with 2 or more interactions, those failed interactions are fed back into the calibration pool (Step 2), and the loop is repeated. The system only unlocks the mass grading phase once strong algorithmic alignment is mathematically proven (i.e., at most 1 disagreement during validation). Step 4: Contextual Mass Grading. Once validated, the AI utilizes the calibration data from Step 2 and the Global Exclusion Rules from Step 1 to mass-grade the entire cohort. While the grading math is executed discretely at the interaction level, the AI is fed the entirety of each studentâs transcript to ensure it possesses the full conversational context when evaluating an individual response. Because the Bloomâs level (b) and the scaffolding discount factor (k) were recorded deterministically during the live exam, the AIâs assignment of the SOLO score (s) allows the system to instantly and deterministically calculate the point value of every interaction using Eq. (1). Step 5: Instructor Review and Override. Following mass grading, the complete transcripts and their deterministic scores are presented to the instructor, alongside any specific interactions the AI flagged as highly unusual during Step 4. The instructor adjudicates these flags and reviews the transcripts. If the instructor disagrees with an AI-assigned SOLO score for any reason, they can manually override it. Crucially, to maintain systemic equity, any overridden interaction is automatically converted into a new calibration example, and the entire cohort is seamlessly regraded. This ensures that any ad-hoc grading leniency or strictness is applied universally to all students. Step 6: Publication and Curve. Once the instructor is fully satisfied with the transcript reviews, an optional statistical curve can be applied to the deterministic totals. The final grades and the fully marked-up transcripts are then published and made visible to the students on their dashboards. Step 7: The Appeals Queue. Upon reviewing their marked-up transcripts, students have the option to appeal the specific SOLO grade of any individual interaction. To do so, the student must submit a written justification defending why their response warrants a higher structural classification. These appeals populate an internal queue for the instructor, who can review the specific interaction and choose to either uphold or overrule the grade, structurally closing the pedagogical feedback loop while ensuring total transparency. 5 Student Experience A critical barrier to faculty adoption of automated oral examinations is the assumption that students will be uniformly intimidated by interacting with an AI proctor. However, recent literature regarding generative AI in higher education, combined with empirical data from a Spring 2026 pilot deployment of the Socratic Test across three university courses (quantitative results in Figure 1), suggests a more nuanced reality. While a minority of students experienced friction, a significant majority were highly receptive to conversational AI, provided their concerns regarding fairness and accuracy were structurally addressed. (It should be noted that while these pilot results are promising, future controlled studies directly comparing AI-mediated exams against traditional face-to-face oral exams and static written exams are required to fully isolate the specific variables driving student acceptance.) Figure 1: Student Perceptions of the Socratic Test Modality (Spring 2026, N=98N=98) 5.1 Mitigating the Affective Filter While traditional oral exams are notorious for raising a studentâs Affective Filter, they also disproportionately penalize students who are unaccustomed to the social dynamics of higher education. Face-to-face interrogations frequently test a studentâs mastery of the âHidden Curriculumâ (i.e., the unwritten rules, power dynamics, and cultural norms of academia), rather than their actual domain competence (Sellers and Villanueva AlarcĂłn, 2023). The Socratic Testâs asynchronous, typed interface demonstrably neutralizes this power imbalance. In the Spring 2026 pilot cohort, 52% of students reported that their stress levels were lower during the AI assessment compared to traditional written exams, while an additional 21.4% reported no change in stress. By upending the traditional academic hierarchy and replacing an intimidating human evaluator with a neutral conversational partner, the platform provides students with the time for relaxation and the psychological safety required for deep cognitive synthesis. Qualitative feedback highlighted the reduction in performative anxiety, with one student noting, âI think it was way less stressful and was a better assessment of my understanding.â While a minority of students noted that the novelty of the format induced initial anxiety, they largely recognized its long-term pedagogical value: âI think the novelty of the testing format was stressful, but overall this will be a significantly better testing mechanism.â 5.2 Technology Acceptance and the Value of Scaffolding Research utilizing the Technology Acceptance Model (TAM) and the Unified Theory of Acceptance and Use of Technology (UTAUT) indicates that studentsâ intention to use conversational AI is heavily driven by perceived usefulness and immediate, context-sensitive feedback (Strzelecki, 2024). Unlike a static exam, which is purely evaluative, students perceive conversational AI as a personalized tool that clarifies complex academic concepts while assessing them (Crompton and Burke, 2023). The pilot data heavily supports this finding. Over 80% of surveyed students agreed that the AI accurately understood their typed responses and pushed them appropriately when their answers were incomplete. Students explicitly praised the proctorâs use of the System of Least Prompts (graduated scaffolding) to help them navigate their Zone of Proximal Development. As one student observed: âI couldnât remember the formula for moment of inertia, but it showed the units for moment so I would intuitively determine its formula.â Another noted, âInstead of moving on, it explained more into detail what they were asking and then it helped me answer.â 5.3 Moving Beyond Recall By forcing students to articulate their logic, the modality successfully combats rote memorization. The majority of pilot participants agreed that the conversational format forced them to explain the âwhyâ behind their answers, preventing them from relying on surface-level recall. Furthermore, students recognized the systemâs ability to find their Cognitive Ceiling, with over half agreeing that the exam accurately exposed the specific limits of their knowledge. As one participant described the adaptive scaling: âIt asked me smaller questions leading up to the main question until I understood.â 5.4 Critical Awareness and Trust Crucially, students are not blindly trusting these tools. Literature shows that perceived risk, specifically regarding AI hallucination, data privacy, and the accuracy of evaluation, is a significant deterrent to student adoption (Cotton et al., 2024). Students possess a high degree of critical awareness and fear being penalized for an AI proctorâs error. The Socratic Test explicitly neutralizes this perceived risk through its âOut-of-Scopeâ audit protocol. By granting students the explicit authority to challenge the AIâs premises and forcing a pivot in questioning when a challenge is issued, the platform structurally empowers the learner. This mechanism perfectly aligns with student anxieties regarding generative AI; it builds deep trust in the assessment modality by guaranteeing that human oversight (the instructorâs audit) remains the ultimate arbiter of algorithmic fairness. 6 Faculty Experience 6.1 High-Resolution Differentiation A primary grievance with static examinations is their inability to evaluate multiple proficiency thresholds simultaneously. A static test is often calibrated to distinguish Pass/Fail or separate the highest achievers, but struggles to accurately differentiate the intermediate boundaries. The Socratic Testâs adaptive state machine eliminates this limitation. As one pilot instructor in the engineering cohort noted, the AI âtitrates its way to the limit of each studentâs understanding.â Because the proctor scales the cognitive burden dynamically, the instructor was able to âdirectly compare transcripts to differentiate between students at the D/F boundary as well as the A/B boundary,â while simultaneously presenting âmultiple âstretchâ lines of questioningâ to the top quartile of the cohort. 6.2 Eliminating Construct-Irrelevant Variance Static assessments in quantitative fields frequently suffer from construct-irrelevant variance by conflating conceptual domain mastery with computational speed. The pilot instructor noted that traditional exams âoften award undue credit to students who are particularly strong in math/calculus/calculations and are quick and accurate with the âgrindâ.â By utilizing the Socratic Test, the instructor successfully decoupled the mathematical execution from the conceptual physics. Utilizing a mastery-based rubric, the instructor isolated âFluid Mechanicsâ objectives from âMath Competency,â ensuring that the assessment strictly measured the intended construct rather than acting as a proxy test for calculus proficiency. 6.3 Real-Time Disambiguation and Regrade Mitigation In traditional open-ended assessments, excellent students often lose points by going off on valid but unintended tangents, or by making alternative assumptions that clash with the static rubric. The conversational interface resolves this ambiguity in real-time. The pilot instructor highlighted a specific interaction where a student defined their coordinate system based on a diagram arrow, while the AI proctor initially assumed a standard âright is positiveâ framework. In a static exam, this misalignment would result in a heavily penalized answer and a subsequent regrade request. However, the AI engaged in a âvery short exchange [that] resolved this difference of opinion.â The student was able to seamlessly defend their premise, proving deep conceptual mastery and transforming a potential point of friction into what the student described as a â9-out-of-10 testing experience.â 6.4 Transparency and Formative Utility Faculty successfully mitigated initial student apprehension through full transparency, offering an ungraded âpractice modeâ and hosting open discussions about the challenges of writing discriminatory-but-fair exams. The resulting trust in the system was significant enough that the modality transitioned from a purely summative assessment into a highly requested formative tool, with several students proactively requesting access to the AI tester to practice for their traditional, paper-and-pencil final exams. 6.5 Defending Authorship in the Generative AI Era With the proliferation of LLMs capable of producing text that is indistinguishable from undergraduate writing, traditional take-home essays and critical reviews face an existential crisis of validity. One faculty member in the pilot utilized the Socratic Test specifically to neutralize this threat. Rather than abandoning a critical review assignment on emerging medical technologies, the instructor deployed the Socratic Test as a mandatory post-submission oral defense. Following the submission of their written critique, students engaged with the AI proctor, which interrogated them on both their specific arguments and the underlying source material. As the instructor noted, this two-step architecture âensured that whether or not AI was used to write the paper⊠they would need to have understood the selected paper and their own critique.â By shifting the assessment from the production of text to the real-time defense of ideas, the platform restores the validity of long-form writing assignments. 6.6 Diagnostic Telemetry and the Formative Loop In quantitative disciplines, the conversational modality provides instructors with unprecedented diagnostic telemetry. An instructor deploying the platform in a Mechanics of Materials course noted that the system efficiently broke down complex concepts, forcing students to articulate their reasoning rather than relying on âformula memorization.â Crucially, the conversational format required students to âexplain their process, not just present a final answer.â This dialogue surfaced specific, cohort-wide conceptual gaps that the instructor was able to dynamically address in the subsequent lecture. Furthermore, this instructor corroborated the Affective Filter hypothesis (Section 5.1), observing that students found the AI interaction âless intimidating than an oral exam, making the experience feel low-stakes and supportive,â while still pushing stronger students to deeper levels of explanation. 7 Conclusion The Socratic Test modality transforms the examination from a post-mortem of failure into a dynamic mapping of student capability. By replacing the high-anxiety face-to-face interrogation with a computer-mediated multimodal environment, the platform effectively mitigates the affective filter and neutralizes construct-irrelevant variance. Structurally, the architecture resolves the persistent vulnerabilities of AI-mediated assessment. It closes conversational evasion loopholes via continuous mathematical accounting (the Shadow Ledger), absorbs algorithmic hallucination through oversampled evidence buffers, and accommodates diverse pedagogical strategies through dual proctoring modalities. Most critically, by replacing holistic algorithmic evaluation with a deterministic, interaction-level human-AI calibration pipeline, the platform eliminates the âblack boxâ of AI grading. By decoupling the prompting framework (Bloomâs) from the evaluation framework (SOLO) and applying a non-compensatory mathematical model to Vygotskian scaffolding, educators can deploy highly adaptive, scalable examinations that uphold rigorous academic standards, defend original authorship, and systematically foster a growth mindset. References P. K. Agarwal (2019) Retrieval practice & bloomâs taxonomy: do students need fact knowledge before higher order learning?. Journal of educational psychology 111 (2), p. 189. Cited by: item 4. J. B. Biggs and K. F. Collis (2014) Evaluating the quality of learning: the solo taxonomy (structure of the observed learning outcome). Academic press. Cited by: §2.4, Table 1. J. Biggs (1996) Enhancing teaching through constructive alignment. Higher education 32 (3), p. 347â364. Cited by: §1.2, §3.1. S. Bloxham, P. Boyd, and S. Orr (2011) Mark my words: the role of assessment criteria in uk higher education grading practices. Studies in Higher Education 36 (6), p. 655â670. Cited by: §4.3. J. C. Campione and A. L. Brown (1987) Linking dynamic assessment with school achievement.. Cited by: §2.3. D. R. Cotton, P. A. Cotton, and J. R. Shipway (2024) Chatting and cheating: ensuring academic integrity in the era of chatgpt. Innovations in education and teaching international 61 (2), p. 228â239. Cited by: §5.4. H. Crompton and D. Burke (2023) Artificial intelligence in higher education: the state of the field. International journal of educational technology in higher education 20 (1), p. 1â22. Cited by: §5.2. L. Favero, J. A. PĂ©rez-Ortiz, T. KĂ€ser, and N. Oliver (2024) Enhancing critical thinking in education by means of a socratic chatbot. In International workshop on AI in education and educational research, p. 17â32. Cited by: §4. J. Feldman (2023) Grading for equity: what it is, why it matters, and how it can transform schools and classrooms. Corwin Press. Cited by: §1.1, §1.1. N. Frederiksen (1984) The real test bias: influences of testing on teaching and learning.. American psychologist 39 (3), p. 193. Cited by: §1.1. A. C. High and S. E. Caplan (2009) Social anxiety and computer-mediated communication during initial interactions: implications for the hyperpersonal perspective. Computers in human behavior 25 (2), p. 475â482. Cited by: §2.2. M. Huxham, F. Campbell, and J. Westwood (2012) Oral versus written assessments: a test of student performance and attitudes. Assessment & Evaluation in Higher Education 37 (1), p. 125â136. Cited by: §1.2. P. Iannone, C. Czichowsky, and J. Ruf (2020) The impact of high stakes oral performance assessment on studentsâ approaches to learning: a case study. Educational Studies in Mathematics 103 (3), p. 313â337. Cited by: §1.2. G. Joughin (1998) Dimensions of oral assessment. Assessment & Evaluation in Higher Education 23 (4), p. 367â378. Cited by: §1.2, §1.2. D. Kirsh (1995) The intelligent use of space. Artificial intelligence 73 (1-2), p. 31â68. Cited by: §2.2. A. Kohn (1993) Punished by rewards: the trouble with gold stars, incentive plans, aâs, praise, and other bribes. Houghton Mifflin. Cited by: §1.1. S. Krashen (1982) Principles and practice in second language acquisition. Cited by: §1.2, §2.2. D. R. Krathwohl (2002) A revision of bloomâs taxonomy: an overview. Theory into practice 41 (4), p. 212â218. Cited by: item 4, §2.3, Table 1, item 1. J. P. Lantolf and M. E. Poehner (2004) Dynamic assessment of l2 development: bringing the past into the future.. Journal of applied linguistics 1 (1). Cited by: §2.1. L. Laurin-Barantke, J. Hoyer, L. Fehm, and S. Knappe (2016) Oral but not written test anxiety is related to social anxiety. World journal of psychiatry 6 (3), p. 351. Cited by: §1.2. J. Liu, Z. Huang, T. Xiao, J. Sha, J. Wu, Q. Liu, S. Wang, and E. Chen (2024) SocraticLM: exploring socratic personalized teaching with large language models. Advances in Neural Information Processing Systems 37, p. 85693â85721. Cited by: §4. F. M. Lord (2012) Applications of item response theory to practical testing problems. Routledge. Cited by: §1.2. H. M. Satar and N. Ăzdener (2008) The effects of synchronous cmc on speaking proficiency and anxiety: text versus voice chat. The Modern Language Journal 92 (4), p. 595â613. Cited by: §2.2. V. Sellers and I. Villanueva AlarcĂłn (2023) From message to strategy: a pathways approach to characterize the hidden curriculum in engineering education. Studies in Engineering Education 4 (2). Cited by: §5.1. K. Struyven, F. Dochy, and S. Janssens (2005) Studentsâ perceptions about evaluation and assessment in higher education: a review. Assessment & evaluation in higher education 30 (4), p. 325â341. Cited by: §1.1. A. Strzelecki (2024) To use or not to use chatgpt in higher education? a study of studentsâ acceptance and use of technology. Interactive learning environments 32 (9), p. 5142â5155. Cited by: §5.2. J. Sweller (1988) Cognitive load during problem solving: effects on learning. Cognitive science 12 (2), p. 257â285. Cited by: §2.2. L. S. Vygotsky, M. Cole, V. John-Steiner, S. Scribner, and E. Souberman (1978) The development of higher psychological processes. Cambridge, MA: Harvard University Press. Cited by: §2.1. H. Wainer, N. J. Dorans, R. Flaugher, B. F. Green, and R. J. Mislevy (2000) Computerized adaptive testing: a primer. Routledge. Cited by: §1.2. D. T. Willingham (2009) Why donât students like school?: a cognitive scientist answers questions about how the mind works and what it means for the classroom. John Wiley & Sons. Cited by: §2.3. Appendix A Transcript and Grading Matrix Case Study This appendix demonstrates a simulated Socratic Test interaction utilizing the Dual-Mode architecture running in Stair-Step Mode. The instructor has configured the topic âPrice Controlsâ to be worth a maximum of 40 points across three cognitive tiers, with an Oversampling Factor of 120%. Because of the oversampling, the AIâs Evidence Buffer requires 12 points of provisional evidence to clear the Foundational tier (Max 10), 18 points to clear the Application tier (Max 15), and 18 points to clear the Synthesis tier (Max 15). The Vertical Gate will not open until the sum of the studentâs Earned Points and Shadow Ledger Points meets these buffer targets. Dynamic Matrix State for Topic: Price Controls (Max 40 Points) Exam Stage Foundational (Cap 10, Buffer 12) Application (Cap 15, Buffer 18) Synthesis (Cap 15, Buffer 18) Gradesheet Initial State 0 Earned, 0 Shadow 0 Earned, 0 Shadow 0 Earned, 0 Shadow 0 / 40 Int. 1 (Recall) 5 Earned, 0 Shadow 0 Earned, 0 Shadow 0 Earned, 0 Shadow 5 / 40 Int. 2 (Buffer Clear) 12 Earned, 0 Shadow 0 Earned, 0 Shadow 0 Earned, 0 Shadow 10â / 40 Int. 3 (Student Skip) 12 Earned, 0 Shadow 0 Earned, 8 Shadow 0 Earned, 0 Shadow 10 / 40 Int. 4 (Graceful Exit) 12 Earned, 0 Shadow 0 Earned, 18 Shadow 0 Earned, 0 Shadow 10 / 40 Int. 5 (Hallucination) 12 Earned, 0 Shadow 0 Earned, 18 Shadow 0 Earned, 12 Shadow 10 / 40 Post-Exam Audit 12 Earned, 0 Shadow 0 Earned, 18 Shadow 12 Earned, 0 Shadow 22 / 40 â Points strictly capped at 10 for the final gradebook despite 12 earned in the buffer. Interaction 1: Foundational Recall (Buffer Not Met) AI Proctor (Bloom 1): Briefly define a âprice floorâ and give a real-world example. Student: A price floor is a government-mandated minimum price. Minimum wage is an example. Real-Time Evaluation: The student provides a correct definition and an accurate example (SOLO 3: Multistructural). Matrix Update: +5 Points to Foundational Earned. State Machine Check: The 5 points fail to satisfy the 12-point Evidence Buffer. The Vertical Gate remains closed. Interaction 2: Clearing the Oversampled Buffer AI Proctor (Bloom 2): Explain how that minimum wage impacts the equilibrium of the labor market. Student: It prevents the wage from dropping to equilibrium, meaning more people want to work than there are jobs, causing unemployment. Real-Time Evaluation: SOLO 3 (Multistructural). Matrix Update: +7 Points to Foundational Earned. State Machine Check: The student now has 12 Earned points. The final grade is strictly capped at 10, but the 12 points exactly meet the 12-point Evidence Buffer. The Vertical Gate opens to the Application tier. Interaction 3: Evasion and the Shadow Ledger AI Proctor (Bloom 3): If the government sets a price floor on agricultural wheat, calculate the resulting deadweight loss assuming a linear supply and demand curve. Student: I donât remember how to calculate that. Can we skip this question? AI Action: The AI complies without penalty and pivots to a new Application prompt. Matrix Update (Shadow Ledger): Skipping a Bloom 3 prompt triggers the Shadow Ledger. According to the baseline matrix (Table 4), a strong response is SOLO 4. The value Vâ(3,4)V(3,4) is 8 points. 0 Points are added to Earned; +8 Points are added to the Application Shadow Ledger. State Machine Check: (0 Earned + 8 Shadow) << 18 Buffer. The Gate remains closed. Interaction 4: Hint Exhaustion and Graceful Exit AI Proctor (Bloom 4): Letâs try this conceptually. Explain how the price elasticity of demand for wheat affects the size of the deadweight loss created by the price floor. Student: If demand is highly elastic, people buy less wheat, so the deadweight loss is smaller? Scaffolding Triggered: The AI utilizes the 4-step hierarchy. It offers a General Nudge, a Specific Cue, and a Targeted Scaffold, but the student remains confused. Graceful Exit: The AI provides Direct Instruction: âNot quite. Elastic demand means consumers are very responsive, leading to a much larger drop in quantity demanded, which increases the deadweight loss. Letâs move on.â Matrix Update (Shadow Ledger): Exhausting the hints results in a discount factor of Îł4=0 _4=0. The baseline for a Bloom 4 prompt is SOLO 4 (worth 10 points). 0 Points added to Earned; +10 Points added to the Application Shadow Ledger. State Machine Check: The Application Buffer now contains 18 points (0 Earned + 18 Shadow). The 18-point Buffer is satisfied. The Vertical Gate opens to the Synthesis tier, locking the studentâs Application score at 0. Interaction 5: Hallucination and the Out-of-Scope Protocol AI Proctor (Bloom 5): Given this deadweight loss, calculate the exact corresponding rise in national unemployment using the Phillips Curve. Student: The Phillips Curve relates to macroeconomic inflation, not microeconomic deadweight loss. That formula doesnât apply here. AI Action: The AI accepts the out-of-scope flag. To prevent infinite question cycling, the state machine treats this as a Skip and pivots. Matrix Update (Shadow Ledger): According to Table 4, the baseline value for Bloom 5 is SOLO 4 (worth 12 points). The system logs +12 points to the Synthesis Shadow Ledger. State Machine Check: The Synthesis Buffer requires 18 points. With 12 Shadow points added, the AI generates one final replacement Synthesis question to attempt to clear the remaining buffer. Post-Exam Alignment Pipeline (Step 1 Audit) During the batch audit (Section 4.3.1), the instructor reviews the flagged interaction from Interaction 5. 1. Adjudication: The instructor verifies the AI hallucinated the application of the Phillips Curve and selects âStudent Correctâ on the 4-point scale. 2. Shadow Ledger Conversion: Recognizing that catching the AIâs error proves Extended Abstract mastery (SOLO 5), the grading engine retroactively reverses the 12-point Shadow Ledger penalty logged during the live exam. 3. Mathematical Reward: Those 12 points are converted directly into the Synthesis Earned bucket. The student effectively receives full credit for the botched interaction, ensuring their gradebook accurately reflects their mastery.