Paper deep dive
TeachMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials
Fatema Tuj Johora Faria, Mukaffi Bin Moin, M. F. Mridha, Jubayer Al Mahmud
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/17/2026, 4:07:44 AM
Summary
The paper introduces TeachMateGPT, a multi-agent framework for generating curriculum-grounded science assessments (MCQs and creative questions) from NCTB Class 8 textbooks. It addresses limitations of vanilla RAG by introducing COPE (a hierarchical, graph-based knowledge base), a staged fail-closed agent pipeline with hybrid retrieval, SAVER (a verification protocol), and the NCTB-SciGen8 dataset. The system significantly improves faithfulness and answer relevancy over baselines.
Entities (8)
Relation Signals (7)
TeachMateGPT â generates â Assessment Items
confidence 95% ¡ Automatically generating textbook-grounded assessment items... TeachMateGPT... multi-agent system
TeachMateGPT â produces â NCTB-SciGen8
confidence 95% ¡ NCTB-SciGen8... produced by the pipeline
TeachMateGPT â uses â SAVER
confidence 95% ¡ TeachMateGPT... (iii) SAVER, a source-attributed verification protocol...
TeachMateGPT â uses â COPE
confidence 95% ¡ TeachMateGPT... (i) COPE, a hierarchical knowledge base...
COPE â structures â NCTB Class 8 Science
confidence 90% ¡ COPE... segments documents along syllabus structure... NCTB Class 8 science textbooks
TeachMateGPT â supports â Bangla
confidence 90% ¡ This is particularly important for Bangladeshi science teachers... use Bangla
SAVER â verifies â Assessment Items
confidence 90% ¡ SAVER... scoring faithfulness, relevance, and hallucination risk against retrieved evidence
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Automatically generating textbook-grounded assessment items can reduce science teachers' workload, but existing retrieval-augmented generation (RAG) systems rely on flat retrieval, support only single-question generation, lack safeguards against weak evidence, and are ill-suited to low-resource, board-exam-structured curricula. We address these limitations with TeachMateGPT, a multi-agent system contributing four advances to curriculum-grounded science-assessment authoring. (i) COPE, a hierarchical knowledge base replacing token-window chunking with a multi-resolution index that segments documents along syllabus structure and links them at three granularities via a traversable graph-based lineage, matching evidence to each topic's instructional level. (ii) A staged, fail-closed agent pipeline replacing one-shot retrieve-then-generate: routing gates search, retrieval fuses dense and lexical evidence under a coverage gate that withholds generation on insufficient evidence, and specialist agents draft objective and constructed-response items. (iii) SAVER, a source-attributed verification protocol scoring faithfulness, relevance, and hallucination risk against retrieved evidence, applying stricter grounding checks across each creative question's four sub-parts, paired with teacher-in-the-loop evaluation rather than automatic filtering. (iv) NCTB-SciGen8, a curriculum-grounded dataset of 198 items (143 multiple-choice, 55 creative questions) spanning all 14 chapters of the NCTB Class 8 science textbook, produced by the pipeline and rated by three practicing teachers. TeachMateGPT raises faithfulness (0.68 $\rightarrow$ 0.96) and answer relevancy (0.60 $\rightarrow$ 0.89) over a vanilla RAG baseline.
Tags
Links
- Source: https://arxiv.org/abs/2608.13708v1
- Canonical: https://arxiv.org/abs/2608.13708v1
Trouble viewing inline? Open PDF directly â
Full Text
105,916 characters extracted from source content.
Expand or collapse full text
TEACHMATEGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials Fatema Tuj Johora Faria 1 , Mukaffi Bin Moin 1 , M. F. Mridha 2 , Jubayer Al Mahmud 3 1 Ahsanullah University of Science and Technology, Bangladesh 2 American International University - Bangladesh 3 Jashore University of Science and Technology, Bangladesh Correspondence: mukaffi28@gmail.com, fatema.faria142@gmail.com Abstract Automatically generating textbook-grounded assessment items can reduce science teachersâ workload, but existing retrieval-augmented gen- eration (RAG) systems rely on flat retrieval, support only single-question generation, lack safeguards against weak evidence, and are ill- suited to low-resource, board-exam-structured curricula. We address these limitations with TEACHMATEGPT, a multi-agent system con- tributing four advances to curriculum-grounded science-assessment authoring. (i) COPE, a hierarchical knowledge base replacing token- window chunking with a multi-resolution index that segments documents along syllabus struc- ture and links them at three granularities via a traversable graph-based lineage, matching ev- idence to each topicâs instructional level. (i) A staged, fail-closed agent pipeline replacing one-shot retrieve-then-generate: routing gates search, retrieval fuses dense and lexical evi- dence under a coverage gate that withholds gen- eration on insufficient evidence, and specialist agents draft objective and constructed-response items. (i) SAVER, a source-attributed ver- ification protocol scoring faithfulness, rele- vance, and hallucination risk against retrieved evidence, applying stricter grounding checks across each creative questionâs four sub-parts, paired with teacher-in-the-loop evaluation rather than automatic filtering. (iv) NCTB- SciGen8, a curriculum-grounded dataset of 198 items (143 multiple-choice, 55 creative ques- tions) spanning all 14 chapters of the NCTB Class 8 science textbook, produced by the pipeline and rated by three practicing teach- ers. TeachMateGPT raises faithfulness (0.68 â 0.96) and answer relevancy (0.60â 0.89) over a vanilla RAG baseline. 1 Introduction Large language models (LLMs) can accelerate as- sessment creation, but unconstrained generation may introduce unsupported facts, chapter drift, TeachMateGPT Creative Question Multiple Choice Question āĻā§āĻĒāĻ: ā§ā§āϰ āĻ ā§ āĻŋāĻā§āϤ āϰāĻžāĻŋāĻĢ āĻžā§āĻŽ āϤāĻžāϰ āĻĻāĻžāĻĻ ā§ āĻŦāĻžāĻŋā§ āĻŦā§āĻžā§āϤ āĻāϞāĨ¤ āĻāĻ āĻŋāĻŦā§āĻā§āϞ āϏ āĻĒ ā§ āĻ ā§ āϰāĻĒāĻžā§ā§ āĻšāĻžāĻāĻā§āϤ āĻŦāϰ āĻšā§ā§ āĻŋāϤāύāĻŋāĻ āĻŋāĻ āϧāϰā§āύāϰ āĻžāĻŖā§ āĻĻāĻā§āϤ āĻĒāϞāĨ¤ āĻĨāĻŽ āĻžāĻŖā§āĻŋāĻāϰ āĻĻāĻš āύāϰāĻŽ āĻāĻŦāĻ āĻļ āĻāĻžāϞāϏ āĻžāϰāĻž āĻāĻŦ ā§ āϤ āĻŋāĻāϞ, āĻāϰ āĻāĻŋāĻ āĻĒāĻŋāĻļāĻŦāϞ āĻĒāĻž āĻŋāĻĻā§ā§ āϧā§ā§āϰ āϧā§ā§āϰ āĻāϞāĻžāĻāϞ āĻāϰāĻŋāĻāϞāĨ¤ āĻŋāϤā§ā§ āĻžāĻŖā§āĻŋāĻāϰ āĻĻāĻš āĻāĻžāĻŋā§āϤ āĻ āύāϞāĻžāĻāĻžāϰ āĻŋāĻāϞ āĻāĻŦāĻ āĻŋāϤāĻŋāĻ āĻā§ āĻĨāĻžāĻāĻž āĻŋāϏāĻāĻžāϰ āϏāĻžāĻšāĻžā§āϝ āĻāĻŋāĻ āĻāĻāĻž āĻŽāĻžāĻŋāĻā§āϤ āϏāĻšā§āĻ āĻāϞāĻžāĻāϞ āĻāϰāĻŋāĻāϞāĨ¤ āϤ ā§ āϤā§ā§ āĻžāĻŖā§āĻŋāĻāϰ āĻĻāĻš āĻĒāĻžāĻāĻāĻŋāĻ āϏāĻŽāĻžāύ āĻāĻžā§āĻ āĻŋāĻŦāĻ āĻāĻŦāĻ āĻāĻžā§ā§ āĻ āϏāĻāĻ āĻāĻžāĻāĻāĻž āĻŋāĻāϞāĨ¤ āĻŦāĻžāĻŋā§ āĻŋāĻĢā§āϰ āϏ āĻĻāĻžāĻĻāĻžāϰ āĻāĻžā§āĻ āĻāĻžāύā§āϤ āĻāĻžāĻāϞ āĻžāĻŖā§ā§āϞāĻž āĻāĻžāύ āĻāĻžāύ āĻĒā§āĻŦ āϰ āĻ āĻ ā§ āĨ¤ āϏāĻŽ ā§ āĻš: āĻ) (āĻžāύāĻŽ ā§ āϞāĻ): āĻŋāĻŖāĻŋāĻŦāύāĻžāϏ āĻāĻžā§āĻ āĻŦā§āϞ? āĻ) (āĻ āύ ā§ āϧāĻžāĻŦāύāĻŽ ā§ āϞāĻ): āĻŋāĻĒāĻĻ āύāĻžāĻŽāĻāϰāĻŖ āĻŦāϞā§āϤ āĻā§ āĻŦāĻžāĻāĻžā§? āĻŦāĻžāĻāĻž āĻā§āϰāĻžāĨ¤ āĻ) (ā§ā§āĻžāĻāĻŽ ā§ āϞāĻ): āĻā§āĻĒā§āĻāϰ āĻĨāĻŽ āĻžāĻŖā§āĻŋāĻ āĻāĻžāύ āĻĒā§āĻŦ āϰ āĻ āĻ ā§ ? āĻāϰ āĻŦāĻŋāĻļāϏāĻš āĻŦāĻžāĻāĻž āĻā§āϰāĻžāĨ¤ āĻ) (āĻāϤāϰ āĻĻāϤāĻž): āĻā§āĻĒā§āĻāϰ āĻŋāϤā§ā§ āĻ āϤ ā§ āϤā§ā§ āĻžāĻŖā§ āĻĻ ā§ āĻŋāĻ āĻŋāĻ āĻĒā§āĻŦ āϰ āĻ āĻ ā§ āĻšāĻā§āĻžāϰ āĻāĻžāϰāĻŖ āĻŋāĻŦā§āώāĻŖ āĻā§āϰāĻžāĨ¤ Question: āĻžāĻŋāĻŖāĻāĻā§ā§āĻ āϧāĻžā§āĻĒ āϧāĻžā§āĻĒ āĻŋāĻŦāύ āĻāϰāĻžāϰ āĻĒāĻŋāϤā§āĻ āĻā§ āĻŦā§āϞ? āĻ) āĻŋāĻŖāĻŋāĻŦāύāĻžāϏāĻŋāĻŦāĻĻāĻž āĻ) āĻŦāĻžāϤ āĻ) āϏāĻžā§āϞāĻžāĻāϏāĻā§āώāĻŖ āĻ) āϏāĻŽā§ Answer: āĻ) āĻŋāĻŖāĻŋāĻŦāύāĻžāϏāĻŋāĻŦāĻĻāĻž Teacher's Query TeachMateGPT Responses Figure 1: The diagram illustrates the flow from a teacherâs natural-language query with a target difficulty level to curriculum-grounded assessment generation. TEACHMATEGPT retrieves relevant textbook content and generates board-style CQs and MCQs that match the requested topic and difficulty. weak distractors, or invalid board structures. This calls for knowledge-grounded, multi-format gener- ation under pedagogical constraints that retrieves reliable evidence, withholds generation when evi- dence is insufficient, and provides provenance for teacher verification. This is particularly impor- tant for Bangladeshi science teachers who prepare Class 8 assessments from the National Curriculum and Textbook Board (NCTB) (National Curriculum and Textbook Board, 2025) syllabus, where items must align with textbooks, use Bangla, and match the intended difficulty. 1 arXiv:2608.13708v1 [cs.CL] 13 Aug 2026 Retrieval-augmented generation (RAG) ad- dresses hallucination by coupling an LLM with an external knowledge store, conditioning gener- ation on retrieved passages, and is now common in educational NLP (Swacha and Gracel, 2025; Pan et al., 2025). Recent RAG-based systems re- trieve course materials or exam corpora to gener- ate assessment items and answer keys (N et al., 2025; Junior et al., 2025; Sreekanth et al., 2025), construct mixed-format examinations from domain knowledge bases (Hamidi et al., 2025), and sup- port textbook resources such as Bangla B-RAG and NCTB-QA (Alawwad et al., 2025; Khan and Khan, 2025; Eyasir et al., 2026). Other studies improve distractor quality through teacher-student reason- ing (Qiu et al., 2025) and difficulty-controlled gen- eration through knowledge graphs (Chen and Shiu, 2025). However, existing RAG systems rely on flat retrieval over generic chunks and target gen- eral question generation rather than curriculum- specific, board-style assessment authoring. For Bangla NCTB resources, available systems mainly focus on answering questions instead of generat- ing teacher-facing, multi-format assessments with evidence-aware refusal. While RAG provides grounding, a vanilla retrieve-then-generate pipeline remains insufficient for assessment authoring. Teacher requests may require clarification, retrieval must account for cur- riculum context and noisy textbook sources, as- sessment formats impose distinct pedagogical con- straints, and generated items require validation before classroom use. Recent educational sys- tems therefore adopt agentic workflows that co- ordinate routing, retrieval, generation, and verifica- tion rather than relying on a single generation step (Duan et al., 2026; Sreekanth et al., 2025; Chen and Shiu, 2025; Wang et al., 2025; Jia et al., 2025). However, existing systems remain English-centric and provide limited support for query clarification, multi-format generation, and source-attributed ver- ification. Building on these observations, we introduce TEACHMATEGPT, a multi-agent, curriculum- grounded framework for Bangla Class 8 NCTB science assessment generation, with four contri- butions. Firstly, COPE (Curriculum-Oriented Pedagogical Embedding), a hierarchical curricu- lum index organizing textbook content at multi- ple instructional resolutions; unlike generic parent- child chunking, every chunk carries ingest-time lin- eage and neighbor links that retrieval can traverse to recover related instructional content, and remov- ing COPE produces the largest drop in answer rele- vancy among the ablated components, along with a substantial reduction in stimulus realism. Sec- ondly, a staged, fail-closed multi-agent pipeline for query routing, hybrid retrieval, evidence refine- ment, and format-specific generation, incorporat- ing CCI (Contextual Content Injection) for ev- idence restoration and CCR (Consensus-Based Conflict Resolution) for redundancy reduction. Thirdly, SAVER (Source-Attributed Verification and Evidence Ranking), which scores faithfulness, relevance, and hallucination risk before teacher pre- sentation and flags unsupported items for review rather than filtering them automatically. Fourthly, NCTB-SciGen8, a curriculum-grounded evalua- tion dataset produced directly by the pipeline and reviewed by three practicing science teachers. To- gether, these four contributions address the follow- ing research questions. âĸRQ1. How can authorized science textbooks be organized into a retrieval-ready curriculum knowledge base that preserves instructional hierarchy for assessment generation? âĸ RQ2. How should teacher queries be safely routed and clarified so that only well-specified curriculum assessment requests proceed to ev- idence retrieval? âĸRQ3. How can curriculum evidence be re- trieved and refined so that generated assess- ments remain topic-consistent, contextually complete, and withhold generation when cov- erage is insufficient? âĸRQ4. How can multiple classroom assess- ment formats, multiple-choice and board-style creative questions, be generated at control- lable difficulty while staying grounded in re- trieved curriculum evidence? âĸRQ5. How can automatic source-attributed verification, combined with a teacher-in-the- loop review process, support trustworthy acceptance, editing, or cautioning of AI- generated assessments before classroom use? Figure 1 presents an example interaction with TEACHMATEGPT and illustrates the multi-agent workflow from a teacher query to curriculum- grounded assessment items. 2 2 Related Work 2.1 Retrieval-Augmented Generation for Educational NLP RAG connects LLMs with external knowledge sources to incorporate evidence during generation. Recent surveys highlight its growing role in edu- cational applications, particularly for knowledge- intensive tasks that require reliable access to in- structional materials (Swacha and Gracel, 2025; Pan et al., 2025). Existing studies have explored retrieval-based assessment creation from course documents and examination archives, including MCQ generation with answer keys (N et al., 2025; Junior et al., 2025; Sreekanth et al., 2025), mixed- format exam construction from domain knowledge bases (Hamidi et al., 2025), and textbook-based educational QA (Alawwad et al., 2025). For low- resource educational contexts, Bangla resources such as B-RAG and NCTB-QA provide curriculum- specific retrieval benchmarks and demonstrate the potential of NCTB-grounded educational systems (Khan and Khan, 2025; Eyasir et al., 2026). Com- plementary studies explore reasoning-based strate- gies for improving distractor quality (Qiu et al., 2025) and knowledge graphâguided generation with cognitive difficulty control (Chen and Shiu, 2025). 2.2 Agentic Workflows for Educational Assessment Generation Recent educational systems increasingly use agent- based architectures to divide complex assessment tasks among specialized components. CODE-GEN presents a human-in-the-loop RAG agent frame- work for coding-comprehension MCQs, where sep- arate modules handle item creation and quality as- sessment (Duan et al., 2026). Other approaches distribute educational workflows across agents for document analysis, retrieval, question construction, and evaluation to improve consistency with course content (Sreekanth et al., 2025). Knowledge graph enhanced multi-agent RAG frameworks further in- corporate cognitive objectives and difficulty cali- bration through Bloomâs taxonomy and Item Re- sponse Theory (Chen and Shiu, 2025). Related studies explore collaborative generation strategies, such as multi-agent MCQ construction and teacher- student reasoning, to improve distractor quality and assessment reliability (Tian et al., 2026; Qiu et al., 2025). 2.3 Research Gap Although recent studies have advanced educational assessment generation, they focus primarily on iso- lated components of the pipeline. Table 11 sum- marizes representative systems and highlights the capabilities missing from existing approaches. 3 The TEACHMATEGPT Framework Figure 2 presents the end-to-end architecture of TEACHMATEGPT, which transforms a teacherâs instructional request into curriculum-grounded as- sessments. Algorithm 1 (Appendix F) summarizes the complete framework. 3.1 Task Formulation Given a collection of authorized NCTB Class 8 science curriculum documents, the objective is to generate curriculum-aligned assessments supported by verifiable curriculum evidence. Formally, letD = d 1 ,d 2 ,...,d n denote the col- lection of curriculum documents. The framework first constructs a hierarchical curriculum knowl- edge repository, K = COPE(D),(1) where COPE transforms curriculum documents into a hierarchical retrieval-ready representation. Given a teacher queryq, the retrieval module accesses the curriculum repository and returns the supporting evidence, E = R(K,q),(2) whereR(¡)denotes the curriculum retrieval mod- ule. Using the retrieved evidenceE, the requested difficulty levelâ, and assessment typeS â MCQ, CQ, the generation module produces the assessment set, A = G(E,â,S,q),(3) whereG(¡)denotes the assessment generation module. Finally, the generated assessments are verified by SAVER, V = SAVER(q,E,A),(4) whereVis the verification report andSAVER(¡) performs source-attributed verification and evi- dence ranking, as detailed in Section 3.6. 3 NCTB Class 8 Science Book (a) OCR-based Text Extraction (b) Preserve Headings, Definitions, Tables, and Terminology Curriculum Acquisition (1) Chapter â Lesson â Section â Exercise (2) Preserve ParentâChild Relationships Hierarchical Structural Segmentation (i) Coarse: Chapter-level Chunks (i) Medium: Concept-level Chunks (i) Fine: Definition- and Fact-level Chunks Multi-Resolution Pedagogical ChunkingGraph-Aware Knowledge Construction Link Chunks Across Lessons, Hierarchies, and Adjacent Sections Embedding-Based Indexing Dense Embeddings with Hierarchical and Structural Metadata Curriculum Knowledge Repository Stage 1 Agent 1: Intent Agent Agent 2: Ambiguity Detection Agent Agent 3: Clarification Agent Stage 2 TeachMateGPT Teacher Query (with Target Difficulty Level) Teacher Query: āĻāĻŽāĻžāϰ āĻāĻžā§āϰāĻž "āĻāĻŋāϏāĻĄ āĻŦ ā§ āĻŋ" āĻāĻŋāĻĒā§āĻ āĻĻ ā§ āĻŦ āϞāĨ¤ MCQ āĻŋāĻĻā§ā§ āĻāĻāĻ ā§ āĻ āύ ā§ āĻļā§āϞāύ āϏāĻ āĻāϰ āĻāĻāĻāĻž āϏ ā§ āĻāύāĻļā§āϞ āĻŦāĻžāĻŋāύā§ā§ āĻŋāĻĻāύāĨ¤ Determines whether the teacher's query is a valid curriculum-related science request. Greeting, Harmful, or Off-topic Invalid or Non-Instructional Query Valid Science Query Fixed-Reply Handler Returns a predefined response and terminates the workflow. Next Step Curriculum-related assessment request check_ambiguity() Evaluates whether the teacher's query provides sufficient curriculum context. Uses a Bangla Safety Guard and LLM-based semantic reasoning. (is_ambiguous = false) (is_ambiguous = true) Route to Curriculum Retrieval Next Step generate_clarification() Generates a clarification request when the teacher's query lacks sufficient curriculum context. Asks the teacher to specify the relevant chapter, lesson, or topic. Waits for the teacher's follow- up response. Concatenates the follow-up with the original query before resuming the workflow. Stage 3 Agent 4, 5: Retrieval & Concept Detector Agent Step 1: Query Planning Step 2: Hybrid Retrieval Step 3: Weighted Reciprocal Rank Fusion Step 4: Contextual Content Injection (CCI) Step 5: Consensus-Based Conflict Resolution (CCR) Step 6: Context RestorationStep 7: Coverage Verification Performs multi-query expansion to improve retrieval coverage. Retrieves curriculum evidence from the Qdrant vector database using hybrid semantic search and lexical BM25 retrieval. Combines retrieval results from multiple sources into a unified ranked list. Removes duplicate and conflicting evidence to produce a consistent context. Restores the complete chapter, lesson, and section hierarchy of the retrieved evidence. Expands the retrieved evidence with parent chunks and higher-level curriculum context. Verifies whether the retrieved evidence provides sufficient curriculum coverage. If coverage is sufficient If coverage is insufficient END (Fail- Closed) Next Step Stage 4 Agent 6: Assessment ComposerAgent 7: MCQ Specialist Agent 8: Creative Question SpecialistMerge and Count Verification generate_assessment() (1) Selects the appropriate assessment type. (2) Dispatches the request to the corresponding specialist generation agent. generate_mcqs() generate_creative_questions() (1) Generates textbook-grounded multiple-choice questions. (2) Produces a question stem, four options (āĻâāĻ), and the correct answer. (3) Applies iterative validation and retry before finalizing the output. (1) Generates board-style creative questions. (2) Produces a stimulus (āĻā§āĻĒāĻ) and four sub-questions (āĻâāĻ). (3) Covers knowledge, comprehension, application, and higher-order thinking skills. Agent 9: Verification Agent (SAVER) verify_assessment() Verifies that each generated assessment is supported by the retrieved curriculum evidence. Evidence Support Verification Curriculum Relevance Verification Hallucination Risk Assessment Confirms that all facts, options, and answers are grounded in the retrieved evidence. Ensures alignment with the requested learning objectives and curriculum concepts. Detects unsupported, fabricated, or inconsistent content. Stage 5 (i) Verification Passed â Next Step (i) Verification Failed â Flag Assessment Packages the assessment, evidence, and verification results. Stage 6 Records the workflow execution and retrieval metadata. Stores the complete assessment generation record. Associates each assessment with its supporting evidence and curriculum metadata. (1) Detect chapter and curriculum concept (2) Label retrieved evidence Combines generated assessments and ensures the requested number of questions. If no valid assessment is generated â END (No Items Generated) Figure 2: End-to-end architecture of TEACHMATEGPT. Each teacher query flows through the COPE-based cur- riculum knowledge repositoryK, intent analysis and routing, hybrid dense-BM25 retrieval with context restoration and coverage checks, generation of textbook-grounded MCQs and board-style creative questions (UÃÂpk) with sub-questionskâG), source-attributed verification of evidence support, faithfulness, and relevance, and packaging into an auditable, textbook-traceable output, as detailed in the following subsections. 3.2 Stage 1: Knowledge Base Construction (COPE) Reliable curriculum-grounded generation requires retrieval units that preserve the pedagogical struc- ture of educational content. COPE constructs a hierarchical curriculum knowledge repository from authorized NCTB Class 8 science textbooks as a one-time preprocessing step before deployment. COPE consists of five sequential steps: curricu- lum acquisition, hierarchical structural segmenta- tion, multi-resolution pedagogical chunking, graph- aware knowledge construction, and embedding- based indexing. (1) Curriculum Acquisition. The framework collects authorized NCTB documents and converts them into machine-readable text while preserving chapter titles, section headings, definitions, exam- ples, tables, question blocks, and scientific termi- nology across both digitally generated and scanned textbooks. Before segmentation, a normalization stage removes conversion and OCR artifacts while preserving semantic integrity. The result is a clean, structurally faithful corpus ready for hierarchical structural segmentation. (2) Hierarchical Structural Segmentation.Un- like conventional RAG systems that use fixed-size windows, COPE models the instructional organi- zation of curriculum documents directly. Each textbook is divided into structural units along ped- agogical boundaries, chapters, lessons, sections, summaries, exercise blocks, and longer units are further decomposed into finer grains while preserv- ing parentâchild relationships. Every text segment thus retains its position in the curriculum hierar- chy, letting retrieval exploit both local content and broader instructional context. (3) Multi-Resolution Pedagogical Chunking. Educational queries vary in granularity: some re- quire broad conceptual explanations, while others target specific definitions or factual details. COPE therefore represents curriculum content at multiple pedagogical resolutions rather than relying on a single fixed chunk size. Larger chunks preserve chapter-level context, intermediate chunks capture coherent concepts, and finer chunks isolate de- 4 tailed knowledge, allowing retrieval to dynamically match the appropriate level of abstraction to the teacherâs request. Table 3 summarizes the result- ing chunk inventory and the windowing parameters used to construct each resolution tier. (4) Graph-Aware Knowledge Construction. COPE further links chunks from the same les- son, hierarchy, or adjacent sections through a lightweight graph preserving sequential and hier- archical dependencies. This distinguishes COPE from a purely hierarchical parent-child index: rather than recording only a chunkâs static parent, COPE also records sibling and sequential neigh- bor edges at ingest time, letting retrieval traverse this graph outward from a seed match to recover complementary evidence from related instructional units, rather than being limited to isolated passages or a single ancestor chunk. As Table 11 shows, curriculum-aware retrieval is at best partially sup- ported among the compared Bangla NCTB sys- tems; the graph-aware lineage introduced here lets COPE offer this capability in full, and its removal in the w/o COPE ablation (Table 1, Table 2) ac- counts for the largest drop in answer relevancy among the five ablated components, alongside a substantial reduction in stimulus realism. (5) Embedding-Based Curriculum Repository. Every chunk is encoded into a dense semantic rep- resentation together with pedagogical metadata (hi- erarchical level, structural position, graph relation- ships), forming the curriculum repository, K =(c,e c , meta(c)),(5) wherecis a curriculum chunk,e c its embed- ding, andmeta(c)its hierarchical and structural information.Klets teacher queries access curricu- lum knowledge without repeating preprocessing or indexing. 3.3 Stage 2: Intent Analysis and Query Routing After constructingK, a coordinated pipeline of three agents, the Intent Agent, Ambiguity Detec- tion Agent, and Clarification Agent, collectively referred to as IAC (Intent, Ambiguity-detection, and Clarification), determines whether a teacherâs request is curriculum-related before retrieval. The pipeline filters unrelated queries such as greetings, harmful requests, or off-topic inputs, identifies un- derspecified topics, and requests clarification be- fore retrieval continues, ensuring that only well- specified requests proceed further. 3.4 Stage 3: Hybrid Retrieval, Evidence Refinement, and Coverage Validation Routed requests enter a coordinated workflow of two agents, the Retrieval Agent and the Concept Detector, that transforms the refined query into a reliable evidence set. Rather than relying on a single semantic search, the Retrieval Agent inter- nally decomposes retrieval into specialized steps: hierarchical hybrid retrieval, evidence enrichment (CCI), redundancy reduction (CCR), and evidence validation, ensuring generation proceeds only with sufficient curriculum evidence; the Concept De- tector then labels the validated evidence with the curriculum chapter and concept it covers. Prompt specifications for the routing and generation agents are provided in Appendix G; the full implementa- tion is available in the accompanying codebase. (1) Hierarchical Hybrid Retrieval.The refined query goes to a hybrid retrieval module combining dense semantic retrieval, for related concepts, with lexical retrieval, for exact scientific terminology embeddings may miss. Rankings are merged via weighted reciprocal rank fusion, Score(c) = X r w r k + rank r (c) ,(6) wherew r is the weight of retrieval methodr, producing a ranked candidate set that balances se- mantic similarity with curriculum-specific lexical matching. (2) Contextual Content Injection (CCI). Re- trieved candidates may be isolated fragments of a larger concept. CCI restores pedagogical con- text by selectively incorporating higher-level in- structional content tied to the retrieved segments, preserving conceptual continuity without indiscrim- inately expanding the evidence set. (3) Consensus-Based Conflict Resolution (CCR). Multi-resolution retrieval produces overlapping ev- idence across hierarchical levels. CCR identifies equivalent or highly overlapping segments and re- tains the most informative representation, reducing redundancy while keeping complementary informa- tion for downstream reasoning. We use âconflictâ here in the sense of overlapping or duplicated evi- dence spans competing for the same context bud- 5 get, rather than evidence that reports contradictory facts; the implementation targets the former. (4) Evidence Validation. This last check deter- mines whether the evidence sufficiently supports the assessment task by examining curriculum cov- erage, structural diversity, and alignment with re- quested concepts. LetTdenote extracted curricu- lum concepts andEthe retrieved evidence set; cov- erage is Cov(T,E) = |tâ T : tâ E| |T| .(7) Generation proceeds only when coverage meets the required threshold; otherwise the framework withholds generation and requests a more specific query, preventing assessment authoring on insuf- ficient evidence. Section 6 and Table 9 report the resulting fail-closed rate and coverage ratio across retrieval configurations; we did not separately mea- sure refusal precision or recall against a labeled ground truth of queries that should or should not have been refused, and note this as a scope limita- tion in Section 7. 3.5 Stage 4: Multi-Format Assessment Generation Given validated evidenceE, three agents gener- ate classroom-ready assessments for the requested type and difficulty: the Assessment Composer dis- patches the request to the MCQ Specialist and the Creative Specialist, which independently gener- ate each format from the same evidence context. Rather than relying on a single prompt for all for- mats, each specialist follows its own pedagogical requirements while remaining grounded in the re- trieved evidence. (1) Evidence-Grounded Assessment Generation. Letâbe the requested difficulty andSthe desired type. The generation module produces the assess- ment setAas defined in Equation(3). Since every generator receives the same validated evidenceE, items stay consistent with retrieved curriculum con- tent rather than relying on the modelâs parametric knowledge. (2)SpecializedAssessmentGeneration. The framework supports two formats used in Bangladeshi secondary education: MCQs and board-style CQs, operating on the same evidence while serving different pedagogical objectives. The MCQ generator produces a stem, four options, and one correct answer, grounding both the correct answer and distractors in the retrieved evidence for factual accuracy and curriculum alignment. The CQ generator constructs a contextual stimulus followed by four progressively structured sub- questions assessing knowledge, comprehension, application, and higher-order reasoning; rather than reproducing textbook passages, it synthesizes realistic scenarios consistent with the retrieved concepts. 3.6 Stage 5: Source-Attributed Verification (SAVER) Curriculum-grounded retrieval reduces hallucina- tion but does not eliminate it. Before presenta- tion, every assessment passes through the Verifi- cation Agent, running SAVER, an independent post-generation verification process that compares each assessment against its retrieved evidence, as introduced in Equation(4), without modifying the generated content. This single agent then scores every assessment along three criteria. 1. Faithfulness. Whether facts, options, and statements are explicitly supported by the evi- dence. 2. Curriculum Relevance. Alignment between the assessment, the teacherâs objective, and the retrieved concepts. 3.Hallucination Risk. Likelihood of unsup- ported, fabricated, or scientifically inconsis- tent content. These signals are aggregated into a report with quantitative scores and explanatory feedback, and are used to rank items by their evidence support so that the least-supported items surface first for teacher attention. An assessment is flagged for teacher attention rather than accepted for unedited release when Accept(V ) =b faith â§ (s faith âĨ θ F ) â§ (s rel âĨ θ R )â§ (s hall ⤠θ H ) (8) does not hold, whereb faith is the overall verifica- tion decision ands faith ,s rel ,s hall are the faithful- ness, relevance, and hallucination scores. Rather than regenerating or discarding failed assessments automatically, the framework preserves both the as- sessment and its verification report, so acceptance, editing, or discarding of flagged items remains a teacher decision. Export format and provenance details are given in Appendix A. 6 3.7 Stage 6: Auditable Output Packaging The framework packages the teacher query, re- trieved evidence with its provenance (source text- book, chapter, and concept labels), the generated assessment, the verification report, and an execu- tion trace into a structured record. This record is presented to the teacher and also forms the basis of the NCTB-SciGen8 dataset (Section 4), ensur- ing that every dataset instance retains the same evidence trail available during generation. 4 Dataset Construction TEACHMATEGPT also functions as the data- creation pipeline for NCTB-SciGen8, a reusable dataset of curriculum-grounded assessments as- sembled directly from verified pipeline outputs. Every assessment that passes SAVER verifica- tion (Stage 5, Section 3.6) and output packaging (Stage 6, Section 3.7) becomes a dataset instance. NCTB-SciGen8 contains 198 Bangla Class 8 science assessments (143 MCQs, 55 CQs) span- ning all 14 chapters (156 pages) of the official NCTB Class 8 Science textbook, each preserving its full generation provenance. Table 4 summarizes the per-chapter distribution of pages, subject ar- eas, and assessment instances. The export schema, provenance format, teacher-reviewed subset, and coverage-based adequacy argument are detailed in Appendix A; representative CQ and MCQ sam- ples spanning multiple chapters are provided in Appendix B. 5 Experimental Setup The complete implementation details and evalua- tion settings used in our experiments are provided in Appendix D. 6 Results Analysis We evaluate TEACHMATEGPT through five re- search questions. Detailed analyses for each re- search question are provided in Appendix E. Auto- matic evaluation (Table 1) and human evaluation (Table 2) are presented below, whereas retrieval reliability under the fail-closed coverage gate (Ta- ble 9) and inference-time and indexing efficiency (Table 10) are reported in Appendix C. RQ1: Effectiveness of COPE in Curriculum Knowledge Base Construction. COPEâs hierar- chical index preserves the NCTB curriculum struc- ture across all 14 chapters with balanced depth ConfigurationFaith.âAns. Rel.âCtx. Prec.âCtx. Rec.â Vanilla RAG0.680.600.540.58 TEACHMATEGPT0.960.890.920.91 w/o COPE0.860.760.730.75 w/o SAVER0.790.860.900.89 w/o CCR0.830.810.840.85 w/o CCI0.840.820.850.86 w/o IAC0.810.770.890.88 RAPTOR0.800.750.710.78 GraphRAG0.840.790.800.83 CRAG0.860.810.840.85 Adaptive RAG0.790.760.730.79 Table 1: RAGAS evaluation of TEACHMATEGPT against a vanilla RAG baseline, four RAG baselines, and five component ablations (w/o COPE, SAVER, CCR, CCI, IAC). Faith. = Faithfulness, Ans. Rel. = Answer Relevancy, Ctx. Prec. = Context Precision, Ctx. Rec. = Context Recall;âindicates higher is better. It attains the best score on all four metrics. across subject areas. Four of six deterministic va- lidity gates achieve 100% pass rate (Table 5). The remaining errors relate to formatting constraints (option format and scientific notation), not missing curriculum evidence, showing that the curriculum representation layer provides sufficient grounding for assessment generation. RQ2: Performance of the Intent and Clar- ification Routing Layer. The routing layer re- moves 30% of evaluation queries before retrieval, including greetings, harmful requests, and off-topic inputs (Table 6). The Bangla specificity guard re- solves 81% of ambiguity cases without model inter- vention, and explicit teacher prompts trigger no un- necessary clarification (Table 7). Thus, lightweight routing improves safety + efficiency while main- taining usability. RQ3: Reliability of Hybrid Retrieval and the Fail-Closed Coverage Gate. Hybrid retrieval with a coverage gate achieves a safetyâcoverage bal- ance: 12.5% fail-closed rate with 0.724 coverage ratio (Table 9). Dense-only retrieval increases re- fusal to 31.3% and lowers coverage to 0.618, while gate removal reduces safety despite fewer refusals. These results show that dense + lexical retrieval provide complementary signals for OCR-affected Bangla curriculum text. RQ4: Quality of Curriculum-Grounded As- sessment Generation. MCQ generation achieves 85.7% first-attempt validation success, while CQ generation rises from 7.1%â100% after CQ narrative adjustment (Table 12). This result in- dicates that initial CQ errors mainly came from narrative-style mismatch rather than weak curricu- lum grounding. MCQ stem length remains nearly 7 ConfigurationPedagogical AlignmentâStimulus RealismâLinguistic FluencyâOverall Utilityâ Vanilla RAG2.40Âą 0.121.75Âą 0.103.00Âą 0.122.00Âą 0.10 TEACHMATEGPT4.90Âą 0.104.70Âą 0.104.60Âą 0.104.80Âą 0.10 w/o COPE4.05Âą 0.103.90Âą 0.124.35Âą 0.083.95Âą 0.10 w/o SAVER3.25Âą 0.124.25Âą 0.104.25Âą 0.103.55Âą 0.12 w/o CCR3.80Âą 0.104.05Âą 0.104.30Âą 0.083.95Âą 0.10 w/o CCI3.75Âą 0.104.00Âą 0.104.25Âą 0.103.85Âą 0.10 w/o IAC3.50Âą 0.123.85Âą 0.124.15Âą 0.103.60Âą 0.12 RAPTOR3.55Âą 0.123.60Âą 0.124.00Âą 0.103.55Âą 0.12 GraphRAG3.70Âą 0.103.85Âą 0.104.10Âą 0.103.75Âą 0.10 CRAG3.85Âą 0.103.95Âą 0.104.20Âą 0.083.95Âą 0.10 Adaptive RAG3.60Âą 0.103.75Âą 0.104.05Âą 0.103.70Âą 0.10 Table 2: Human evaluation of TEACHMATEGPT-generated Bangla Class 8 NCTB science assessment items by three practicing science teachers. Teachers rated each item on a 5-point Likert scale across four criteria: Pedagogical Alignment, Stimulus Realism, Linguistic Fluency, and Overall Utility. Values denote meanÂąstandard deviation across generated samples and reflect run-to-run variation in model outputs;âindicates higher is better. We compare the complete pipeline with five component-level ablations and four representative RAG baselines. TEACHMATEGPT achieves the highest score for all four evaluation criteria. unchanged across difficulty levels (Table 13), sug- gesting that difficulty depends on semantic and reasoning factors rather than surface length. RQ5: Validation of Source-Attributed Veri- fication and Teacher Review. SAVER identifies only structural defects, with no fabricated facts detected in the audited sample (all 55 CQ clues and a 15-item MCQ spot check; Table 14). Com- pared with Vanilla RAG, TEACHMATEGPT im- proves faithfulness (0.68â 0.96), context precision (0.54 â 0.92), and teacher utility (2.00 â 4.80) (Tables 1 and 2). Ablations show distinct roles: re- moving COPE mainly reduces retrieval quality and stimulus realism, while removing SAVER causes the largest drop in faithfulness and pedagogical alignment. 7 Conclusion We introduce TEACHMATEGPT, a curriculum- grounded multi-agent framework for Bangla assess- ment generation. Our contributions are fourfold: (1) COPE (Curriculum-Oriented Pedagogical Embedding), a hierarchical, graph-aware curricu- lum index whose ingest-time lineage and neighbor links let retrieval exceed static parent-child chunks; (2) a staged, fail-closed multi-agent pipeline that withholds generation under insufficient ev- idence rather than returning fabricated assess- ments; (3) SAVER (Source-Attributed Verification and Evidence Ranking), a verification layer that scores faithfulness, relevance, and hallucination risk and flags unsupported items for teacher review; and (4) NCTB-SciGen8, a curriculum-grounded NCTB Class 8 science assessment dataset with a teacher-rated subset reviewed by three teachers. Across automatic and human evaluations, TEACH- MATEGPT improves context precision from 0.54 to 0.92 and context recall from 0.58 to 0.91. Abla- tion results further demonstrate the complementary roles of retrieval and verification: removing COPE reduces answer relevancy 0.89â0.76 and pedagogi- cal alignment 4.90â4.05, while removing SAVER causes the largest drop in faithfulness 0.96â0.79 and pedagogical alignment 4.90â3.25. These find- ings show that trustworthy assessment generation depends on reliable retrieval, evidence-grounded verification, and an appropriate refusal to generate when curriculum evidence is insufficient. Although our study focuses on the Bangla NCTB Class 8 science textbook, TEACHMATEGPT provides a foundation for curriculum-grounded assessment generation. Future work will explore adaptation across curricula and languages, verification-guided refinement, and psychometric calibration of gener- ated assessments. Limitations Scope and Transferability. TEACHMATEGPT focuses on Bangla Class 8 science assessment gen- eration from authorized NCTB textbooks. We do not claim that the framework transfers directly to other grades, subjects, languages, or curricula. Sev- eral components, such as the Bangla specificity guard, board-style CQ constraints, and curriculum heading detectors, are specific to the NCTB curricu- lum. Extending the framework to new educational 8 settings would therefore require index reconstruc- tion and pipeline adaptation. Indexing and Curriculum Representation. COPE relies on native text extraction or vision- based transcription of textbook pages, both of which may introduce OCR errors, incomplete page coverage, or corrupted mathematical notation. Be- cause retrieval follows a fail-closed design, such errors lead to refusal or limited evidence rather than unsupported generation, although the resulting loss in recall remains only partially quantified. In addi- tion, COPE captures structural relationships within the textbook rather than an explicit prerequisite or learning-objective graph, which may omit peda- gogically related content outside the local textbook structure. Text-Only Modality and Diagram-Dependent Items. TEACHMATEGPT is a text-only frame- work: retrieval, generation, and verification all op- erate over transcribed textbook prose, and vision is confined to the ingestion stage, where scanned pages are converted into text. Curriculum fig- ures, such as circuit and ray diagrams, microscopic cell and organism illustrations, atomic-structure schematics, and labeled graphs, are consequently collapsed into text or discarded rather than retained as retrievable or reproducible visual objects. The framework therefore cannot author items whose stimulus or stem is itself a figure, for instance an MCQ that requires reading a given circuit or a creative-questionUÃÂpkorganized around a dia- gram (âinecricÃiTlXkrâ). Such figure-dependent items are standard in NCTB board examinations, particularly for chapters including Circuit and Cur- rent Electricity, Light, and Structure of the Atom, so both the generated items and the released NCTB- SciGen8 dataset are biased toward verbal reasoning and under-represent this component of the curricu- lum. Extending COPE to multimodal indexing and figure-conditioned generation is a direction we leave to future work. Retrieval, Generation, and Verification. The fail-closed retrieval strategy improves evidence quality but reduces recall by rejecting partially rel- evant evidence under paraphrases, synonymy, or OCR-induced lexical mismatch. We characterize this behavior only through the fail-closed rate and mean coverage ratio measured across retrieval ab- lations (Table 9); we did not construct a labeled set of queries with gold refusal decisions, so re- fusal precision, refusal recall, and false-refusal rate against such a ground truth remain unmeasured, and the fail-closed and coverage figures we report should be read as descriptive of pipeline behavior on our evaluation bank rather than as calibrated detection metrics. Clarification also depends on teacher responses, so underspecified single-turn requests terminate without assessment generation. Generation quality remains bounded by the capabil- ities of the underlying language models. CQ quality is sensitive to narrative style, while force-filled out- puts after validation failure may be pedagogically weaker than fully generated responses. Difficulty control relies on prompting rather than psychomet- ric calibration, and SAVER identifies unsupported or low-confidence items but does not automatically revise or remove them from the released assess- ment. Evaluation Scale. Our human evaluation relies on three practicing teachers rating a configuration- blind sample, and several component analyses use correspondingly small query sets. This limited evaluator pool and sample size constrain statisti- cal power and inter-rater generalizability, so the reported ratings and agreement should be read as indicative rather than definitive; larger teacher pan- els and evaluation banks are needed to establish agreement and effect sizes more robustly. Ethical Considerations Intended Use and Human Oversight. TEACH- MATEGPT is designed as an assistive drafting tool for teachers, not as an autonomous assessment authority. Every generated item is presented to- gether with its supporting evidence and its veri- fication report, and the framework warns rather than silently rewriting flagged items, so a qual- ified teacher makes the final decision to accept, edit, or discard each assessment before classroom use. We caution against deploying the system in a fully automated setting, for example generating live examinations without human review, because automation bias may lead users to over-trust fluent but subtly incorrect items. Assessments produced by the system should be labeled as AI-assisted so that teachers, students, and reviewers remain aware of their origin. Curriculum Data and Copyright. All curricu- lum content is drawn exclusively from the officially authorized NCTB Class 8 science textbook, a pub- licly distributed national curriculum resource, and we deliberately exclude third-party notes, commer- cial question banks, and unrestricted web mate- 9 rial. The textbook remains the intellectual property of the National Curriculum and Textbook Board of Bangladesh; we use it for non-commercial re- search and do not redistribute the textbook itself. The NCTB-SciGen8 records store source identi- fiers such as chapter and section labels, and where a supporting passage is included it is limited to a short excerpt of at most one to two sentences re- tained solely for evidence traceability; we do not redistribute textbook pages or the textbook in full. To preserve anonymity during review, the dataset is not distributed with this submission. The dataset will be made publicly available under the C BY- NC 4.0 license for non-commercial research use. Human Evaluation and Participant Treat- ment. Our human evaluation involves three prac- ticing secondary-school science teachers who rated a configuration-blind sample of generated assess- ments. The teachers are practicing educators who participated voluntarily and gave informed consent, without monetary compensation. No students or other minors took part in the study. The evalua- tion collected only pedagogical quality judgments about the generated items and no personal, sensi- tive, or identifying data about the teachers or any third party, and it posed minimal risk. All ratings are reported in aggregate. Reliability, Misuse, and Academic Integrity. Because incorrect assessment items could mislead learners or reinforce misconceptions, factual relia- bility is a central ethical concern. We mitigate this risk through curriculum-grounded retrieval, a fail- closed coverage gate that refuses generation under weak evidence, and a post-generation verification step that scores faithfulness, relevance, and hallu- cination risk against the retrieved evidence. These safeguards reduce rather than eliminate error, so teacher review remains necessary before any item reaches students. We also acknowledge the risk that a generation tool of this kind could be misused, for instance to mass-produce low-quality question banks or to circumvent a teacherâs own assessment design, and we therefore position the system as sup- port for, rather than replacement of, professional pedagogical judgment. Safety for a Minors-Adjacent Audience. Be- cause the system serves an educational context that includes school-age learners, an intent-routing stage screens every input before retrieval or gen- eration. Unsafe or inappropriate requests, includ- ing violence, self-harm, weapons, sexual content involving minors, harassment, and cheating assis- tance, receive a fixed safe response instead. This routing is a first-line safeguard rather than a com- plete content-moderation guarantee, and teacher oversight remains part of safe deployment. Bias, Fairness, and Language. The underly- ing language models may encode social and topi- cal biases, and generated Bangla text can contain fluency or terminology errors that are harder to detect automatically in a low-resource language than in English. Difficulty labels are conveyed through prompting rather than psychometric cali- bration and should not be interpreted as validated measures of item difficulty. At the same time, by targeting Bangla NCTB science, this work aims to broaden access to assessment-authoring support for an underserved language community. We encour- age similarly careful, curriculum-grounded, and human-supervised adaptation before the framework is extended to other languages, curricula, or learner populations. References Haya A. Alawwad, Abdulrahman Alhothali, Usman Naseem, Abdullah Alkhathlan, and Ahmad Ja- mal. 2025.Enhancing textual textbook ques- tion answering with large language models and retrieval-augmented generation. Pattern Recognition, 162:111332. Harrison Chase. 2022. Langchain.https://github. com/langchain-ai/langchain. Framework for de- veloping applications powered by large language models. Chun-Han Chen and Min-Feng Shiu. 2025. Kaqg: A knowledge-graph-enhanced rag for difficulty- controlled question generation. arXiv preprint, arXiv:2505.07618. Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. 2024. Bge m3-embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. Preprint, arXiv:2402.03216. Xiaojing Duan, Frederick Nwanganga, and Chaoli Wang. 2026. Code-gen: A human-in-the-loop rag- based agentic ai system for multiple-choice question generation. arXiv preprint, arXiv:2604.03926. Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2025. From local to global: A graph rag approach to query-focused summarization. Preprint, arXiv:2404.16130. 10 Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2024. RAGAs: Automated evalu- ation of retrieval augmented generation. In Proceed- ings of the 18th Conference of the European Chap- ter of the Association for Computational Linguistics: System Demonstrations, pages 150â158, St. Julians, Malta. Association for Computational Linguistics. A. Eyasir, Tanvir Ahmed, and Md. Ibrahim. 2026. Nctb- qa: A large-scale bangla educational question answer- ing dataset and benchmarking performance. arXiv preprint, arXiv:2603.05462. Chaimae Hamidi, Mohamed Badiy, Said Gaou, Fouad Amounas, Mohamed Azrour, Hassan Tribak, Ahmed M. Alnajim, and Abdullah Alabdulatif. 2025. Enhancing automated exam creation with retrieval- augmented generation for scalable educational as- sessment. Journal of Advances in Information Tech- nology, 16(10):1430â1441. LangChain Inc. 2024. Langgraph: Build stateful, multi- actor applications with llms.https://github.com/ langchain-ai/langgraph. Software framework for building stateful and multi-agent LLM applica- tions. Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang, and Jong Park. 2024. Adaptive-RAG: Learn- ing to adapt retrieval-augmented large language mod- els through question complexity. In Proceedings of the 2024 Conference of the North American Chap- ter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 7036â7050, Mexico City, Mexico. As- sociation for Computational Linguistics. Rui Jia, Min Zhang, Fengrui Liu, Bo Jiang, Kun Kuang, and Zhongxiang Dai. 2025. Eduagentqg: A multi- agent workflow framework for personalized question generation. Preprint, arXiv:2511.11635. Ankur Joshi, Saket Kale, Satish Chandel, and Dinesh Pal. 2015. Likert scale: Explored and explained. British Journal of Applied Science & Technology, 7:396â403. JosÊ Junior, Leandro Marinho, LÃvia Campos, Kemilli Lima, David Pereira, Helen Cavalcanti, Ana Ramos, and Eliane AraÃējo. 2025. Smarter questions, smaller models: Rag-enhanced multiple-choice question gen- eration for poscomp. pages 1233â1247. Md. Shoaib Abdullah Khan and Md. Saif Ahammod Khan. 2025. B-rag: A retrieval augmented genera- tion based ai system for educational question answer- ing from bangla textbook in bangla. In 2025 IEEE International Conference on Signal Processing, Infor- mation, Communication and Systems (SPICSCON), pages 523â526. Pradeesh N, Remya T, MG Thushara, K Arun Krishna, and Pranav V. 2025. Retrieval-augmented generation for multiple-choice questions and answers generation. Procedia Computer Science, 259:504â511. Sixth In- ternational Conference on Futuristic Trends in Net- works and Computing Technologies (FTNCT06), held in Uttarakhand, India. National Curriculum and Textbook Board. 2025. Na- tional curriculum and textbook board (nctb). Govern- ment of the Peopleâs Republic of Bangladesh. OpenAI. 2024. Gpt-4o mini.https://openai.com. Accessed: 30 June 2026. OpenAI. 2024. Hello gpt-4o.https://openai.com/ index/hello-gpt-4o/. Accessed: 30 June 2026. OpenAI. 2025.Gpt-4.1.https://openai.com/ index/gpt-4-1/. Accessed: 30 June 2026. Feng Pan, Qiyun Zhou, Weitong Guo, and Hongwu Yang. 2025. A survey on retrieval-augmented gener- ation in applications of education and teaching. In 2025 7th International Conference on Computer Sci- ence and Technologies in Education (CSTE), pages 803â807. Qdrant Team. 2025. Qdrant. Accessed: 2026-07-30. Yimiao Qiu, Yang Deng, Quanming Yao, Zhimeng Zhang, Zhiang Dong, Chang Yao, and Jingyuan Chen. 2025. Think both ways: Teacher-student bidirec- tional reasoning enhances mcq generation and distrac- tor quality. In Findings of the Association for Com- putational Linguistics: ACL 2025, pages 8240â8253, Vienna, Austria. Association for Computational Lin- guistics. Zarreen Reza, Alexander Mazur, Michael T. Dugdale, and Robin Ray-Chaudhuri. 2025. Small models, big support: A local llm framework for educator-centric content creation and assessment with rag and cag. Preprint, arXiv:2506.05925. Parth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna, Anna Goldie, and Christopher Manning. 2024. Raptor: Recursive abstractive processing for tree-organized retrieval. In International Conference on Learning Representations, volume 2024, pages 32628â32649. Devananda Sreekanth, Sreekanth Gopi, and Nasrin De- hbozorgi. 2025. Agentic ai quiz-based learning sys- tem: Enhancing mcq generation via long-context cached retrieval-augmented generation. In 2025 IEEE Frontiers in Education Conference (FIE), pages 1â8. Jakub Swacha and MichaÅ Gracel. 2025. Retrieval- augmented generation (rag) chatbots for education: A survey of applications. Applied Sciences, 15(8). Yun Tian and 1 others. 2026. Cognitively diverse multiple-choice question generation: A hybrid multi- agent framework with large language models (re- questa). Electronics, 15(6):1209. 11 Jiayi Wang, Ruiwei Xiao, and Ying-Jui Tseng. 2025. Generating ai literacy mcqs: A multi-agent llm ap- proach. In Proceedings of the 56th ACM Technical Symposium on Computer Science Education V. 2, SIGCSETS 2025, page 1651â1652, New York, NY, USA. Association for Computing Machinery. Guo-Liang Wong, Rui Zhao, Yifan He, and Jiwei Li. 2026. From questions to assessment tuples: A multi- agent framework with bloom-specialized agents and automated verification. In Proceedings of the 21st Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2026), pages 292â 335. Association for Computational Linguistics. Shi-Qi Yan, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. 2024. Corrective retrieval augmented generation. Preprint, arXiv:2401.15884. Appendix A Dataset Construction Details This appendix provides the dataset export schema and the adequacy argument for NCTB-SciGen8 that were summarized in Section 4 of the main text. The full SAVER formalization is presented in Section 3.6. A.1 Export Format and Provenance Generated assessments are exported with full prove- nance that records the original query, detected chap- ter and concept, retrieved textbook sources, gen- erated content, and verification results. MCQs follow the NCTB format with Bangla stems and four options labeledk,x,g, andG. CQs follow the board-style structure with an Uddipok stimulus and four cognitive-level components: Gyan (knowl- edge), Onudhabon (comprehension), Proyog (ap- plication), and Ucchotor Dokkhota (higher-order skills). Unlike manually authored evaluation sets, items are retained only after passing retrieval and validation gates, and every retained item carries the SAVER verification report (Section 3.6), so ev- ery released item is accompanied by an evidence trail rather than being filtered by SAVERâs binary decision alone. A.2 Dataset Schema and Teacher Review All assessment instances are stored in JSON, each uniquely identified byquestion_id. The schema preserves teacher requests, curriculum metadata, retrieved evidence, generated outputs, and evalua- tion annotations. Table 8 provides an overview of the schema. A.3 Adequacy as an Evaluation Dataset We position NCTB-SciGen8 as a curated evalua- tion dataset rather than a large item bank, and its adequacy depends on coverage and quality rather than raw size. First, the 198 assessment items span all 14 chapters of the NCTB Class 8 science syllabus, providing complete curricular coverage of this bounded domain rather than a partial sam- ple. Second, every item passes the validation gates before release (Table 5), and a subset is indepen- dently reviewed by three practicing science teach- ers (Section D.2), providing both automatic ground- ing checks and expert pedagogical judgment. Third, every instance preserves its generation provenance, retrieved evidence, chapter and concept labels, and verification report, enabling item-level auditing. We therefore consider NCTB-SciGen8 adequate for evaluating curriculum-grounded generation within this domain, while acknowledging, as discussed in Section 7, that larger datasets and broader eval- uation panels are needed for more generalizable conclusions. B NCTB-SciGen8 Dataset Details Figures 3 and 4 present sample CQ and MCQ ex- amples from NCTB-SciGen8, while Tables 3 and 4 summarize the COPE indexing configuration and the chapter-wise dataset statistics, respectively. Fig- ure 6 visualizes the automatic-evaluation ablations from Table 1, and Figure 5 presents the correspond- ing human-evaluation comparison alongside the COPE chunk-tier composition. QuantityValue Chapters ingested14 / 14 Total chunks793 Macro-tier chunks181 Meso-tier chunks255 Micro-tier chunks357 Token-aligned micro tierdisabled Embedded tiersmacro, meso, micro Macro-tier window (chunk size / overlap, chars)3600 / 650 Meso-tier window (chunk size / overlap, chars)1900 / 320 Micro-tier window (chunk size / overlap, chars)1100 / 180 Token-micro window (disabled; tokens)640 / 96 Table 3: Chunk counts and windowing parameters for the three nested COPE resolution tiers (macro, meso, micro) produced from 14 ingested NCTB Class 8 sci- ence chapters. Each tier is formed by recursively re- segmenting previous tier spans at progressively finer character windows (with overlap) to support broad con- ceptual retrieval and localized fact retrieval; the token- aligned micro variant is disabled by default due to OCR noise amplification. 12 Chapter: 3 Chapter Title (Bangla): āĻŦāĻžāĻĒāύ, āĻ āĻŋāĻāĻŦāĻŖ āĻ ā§āĻĻāύ Chapter Title (English): Diffusion, Osmosis and Transpiration Chapter 4 Chapter Title (Bangla): āĻāĻŋā§āĻĻāϰ āĻŦāĻāĻļ āĻŦ ā§ āĻŋ Chapter Title (English): Reproduction in Plants Chapter 5 Chapter Title (Bangla): āϏāĻŽā§ āĻ āĻŋāύāĻāϏāϰāĻŖ Chapter Title (English): Coordination and Excretion CQ 1: āĻŦāĻžāĻĒāύ āĻŋā§āĻž āĻā§āĻĒāĻ: āĻ āĻŋāϤāĻŋāĻĨ āĻāϏāĻžāϰ āĻā§āĻ āĻā§āύāĻžā§āĻžāϰāĻž āϤāĻžāϰ āĻāϰ āϏāĻžāĻāĻžāĻŋā§āϞāύāĨ¤ āĻŋāϤāĻŋāύ āĻā§āϰ āĻāĻāĻžā§ āĻĻāĻžāĻāĻŋā§ā§ā§ āĻŦāĻžāϤāĻžā§āϏ āĻŋāĻāĻ ā§ āĻāĻž āϏ āĻāϰā§āϞāύāĨ¤ āĻŋāĻāĻ ā§ ā§āĻŖāϰ āĻŽā§āϧāĻ āĻĒ ā§ ā§āϰāĻž āĻāϰ āϏ ā§ āĻŦāĻžā§āϏ āĻā§āϰ āĻāϞāĨ¤ āĻŋāϤāĻŋāύ āϤāĻāύāĻ āĻā§āϰ āĻāĻ āĻāĻāĻ āĻžā§ āĻĻāĻžāĻāĻŋā§ā§ā§ āĻŋāĻā§āϞāύ, āϤāĻŦ ā§ āĻā§āϰ āĻ āύ āĻž āĻĨā§āĻāĻ āϏ ā§ āĻŦāĻžāϏ āĻĒāĻžāĻā§āĻž āϝāĻžāĻŋāϞāĨ¤ āϤāĻžāϰ āĻā§āϞ āĻ āĻŦāĻžāĻ āĻšā§ā§ āĻāĻžāύā§āϤ āĻāĻžāĻāϞ āĻā§āĻāĻžā§āĻŦ āĻāϤ āϤ āĻĒ ā§ ā§āϰāĻž āĻā§āϰ āϏ ā§ āĻ āĻāĻŋā§ā§ā§ āĻĒā§āϞāĨ¤ āϏāĻŽ ā§ āĻš: āĻ) āĻŦāĻžāĻĒāύ āĻāĻžā§āĻ āĻŦā§āϞ? āĻ) āĻŦāĻžāĻĒāύ āĻāĻžāĻĒ āĻŦāϞā§āϤ āĻā§ āĻŦāĻžāĻāĻžā§? āĻŦāĻžāĻāĻž āĻā§āϰāĻžāĨ¤ āĻ) āĻā§āĻĒā§āĻ āĻā§āϰ āϏ ā§ āĻ āĻāĻŋā§ā§ā§ āĻĒā§āĻžāϰ āĻŋā§āĻžāĻŋāĻ āĻŦāĻžāĻāĻž āĻā§āϰāĻžāĨ¤ āĻ) āĻā§ā§āĻŦāϰ āĻļāĻžāϰā§āϰāĻŦ ā§ ā§ā§ āĻāĻžā§āĻ āĻŦāĻžāĻĒā§āύāϰ āĻŋāĻŦā§āώāĻŖ āĻā§āϰāĻžāĨ¤ CQ 2: āĻ āĻŋāĻāĻŦāĻŖ āĻ āĻŋāĻāĻļāĻŋāĻŽāĻļ āĻĒāϰā§āĻž āĻā§āĻĒāĻ: āĻŋāĻŦāĻžāύ āĻžā§āϏāϰ āĻŦāĻžāĻŋā§āϰ āĻāĻžāĻ āĻŋāĻšā§āϏā§āĻŦ āĻāĻžāĻŋāϰāĻĢ āĻāĻāĻŋāĻ āĻāĻžāĻ āĻĒāϰā§āĻž āĻāϰāϞāĨ¤ āϏ āĻāĻāĻŋāĻ āĻā§āύāĻž āĻŋāĻāĻļāĻŋāĻŽāĻļ āϏāĻžāϧāĻžāϰāĻŖ āĻĒāĻžāĻŋāύā§āϤ āĻŋāĻāĻ ā§ āĻŖ āĻŋāĻāĻŋāĻā§ā§ āϰāĻžāĻāϞāĨ¤ āĻŋāĻāĻ ā§ āĻŖ āĻĒāϰ āĻĻāĻāĻž āĻāϞ āĻŋāĻāĻļāĻŋāĻŽāĻļāĻŋāĻ āĻĢ ā§ ā§āϞ āĻā§āĻ ā§āĻāĨ¤ āĻ āύāĻŋāĻĻā§āĻ āϤāĻžāϰ āĻŦāĻžāύ āĻā§āϤ ā§ āĻšāϞāĻŦāĻļāϤ āϞāĻŦāĻŖāĻž āĻĒāĻžāĻŋāύā§āϤ āĻāĻāĻ āϧāϰā§āύāϰ āĻāĻāĻŋāĻ āĻŋāĻāĻļāĻŋāĻŽāĻļ āĻŋāĻāĻŋāĻā§ā§ āϰāĻžāĻāϞāĨ¤ āĻŋāĻāĻ ā§ āĻŖ āĻĒāϰ āϏ āĻĻāĻāϞ āĻŋāĻāĻļāĻŋāĻŽāĻļāĻŋāĻ āĻĢāĻžāϞāĻžāϰ āĻŦāĻĻā§āϞ āĻāϰāĻ āĻ ā§ āĻ āĻā§āĻ āĻā§āĻāĨ¤ āĻĻ ā§ āĻ āĻāĻžāĻā§āĻŦāĻžāύ āĻŋāĻŽā§āϞ āĻāĻ āĻĒāĻžāĻĨ ā§āĻāϰ āĻāĻžāϰāĻŖ āĻ ā§ āĻ āĻā§āϤ āϞāĻžāĻāϞāĨ¤ āϏāĻŽ ā§ āĻš: āĻ) āĻ āĻŋāĻāĻŦāĻŖ āĻāĻžā§āĻ āĻŦā§āϞ? āĻ) āĻ āϧ ā§āĻāĻĻ āĻĒāĻĻ āĻž āĻā§āĻāĻžā§āĻŦ āĻ āĻŋāĻāĻŦā§āĻŖ āĻ ā§ āĻŋāĻŽāĻāĻž āϰāĻžā§āĻ? āĻŦāĻžāĻāĻž āĻā§āϰāĻžāĨ¤ āĻ) āĻā§āĻĒā§āĻ āϏāĻžāϧāĻžāϰāĻŖ āĻĒāĻžāĻŋāύā§āϤ āĻŋāĻāĻļāĻŋāĻŽāĻļ āĻĢ ā§ ā§āϞ āĻāĻ āĻžāϰ āĻāĻžāϰāĻŖ āĻŦāĻžāĻāĻž āĻā§āϰāĻžāĨ¤ āĻ) āĻā§āĻĒā§āĻāϰ āĻĻ ā§ āĻ āĻĒāĻŋāϰāĻŋāϤā§āϤ āĻŋāĻāĻļāĻŋāĻŽā§āĻļāϰ āĻŋāĻ āĻāĻāϰā§āĻŖāϰ āĻāĻžāϰāĻŖ āϤ ā§ āϞāύāĻžāĻŽ ā§ āϞāĻ āĻŋāĻŦā§āώāĻŖ āĻā§āϰāĻžāĨ¤ CQ 1: āĻ āĻ āĻāύ āĻā§āĻĒāĻ: āĻŦāώ āĻžāϰ ā§āϤ āĻ ā§ āώāĻ āϏāĻŋāϞāĻŽ āϤāĻžāϰ āĻāĻŋāĻŽā§āϤ āĻāϞ ā§ āϰāĻžāĻĒā§āĻŖāϰ āĻŋāϤ āĻŋāύāĻŋā§āϞāύāĨ¤ āĻāϞ ā§ āĻŦāĻžāĻāĻžāĻ āĻāϰāĻžāϰ āϏāĻŽā§ āĻŋāϤāĻŋāύ āϞ āĻāϰā§āϞāύ āĻŋāϤāĻŋāĻ āĻāϞ ā§ āϰ āĻāĻžā§ā§ āĻāĻžāĻ āĻāĻžāĻ 'āĻāĻžāĻ' āϰā§ā§ā§āĻāĨ¤ āĻŋāϤāĻŋāύ āĻāϞ ā§ ā§āϞāĻž āĻā§āĻ āĻ ā§ āĻāϰāĻž āĻāϰā§āϞāύ āĻāĻŦāĻ āĻŋāϤāĻŋāĻ āĻ ā§ āĻāϰāĻžā§ āĻ āϤ āĻāĻāĻŋāĻ āĻā§āϰ āĻāĻžāĻ āϰā§āĻ āĻŽāĻžāĻŋāĻā§āϤ āϰāĻžāĻĒāĻŖ āĻāϰā§āϞāύāĨ¤ āĻŋāĻāĻ ā§ āĻŋāĻĻāύ āĻŋāύā§āĻŋāĻŽāϤ āĻĒāĻžāĻŋāύ āĻĻāĻā§āĻžāϰ āĻĒāϰ āĻŋāϤāĻŋāύ āϞ āĻāϰā§āϞāύ āĻŋāϤāĻŋāĻ āĻ ā§ āĻāϰāĻž āĻĨā§āĻ āύāϤ ā§ āύ āĻāĻžāĻ āĻā§āĻāĨ¤ āĻāĻ āĻĻ ā§ āĻļ āĻĻā§āĻ āϤāĻžāϰ āĻā§āϞ āĻŋāĻāĻžāϏāĻž āĻāϰāϞ āĻŦā§āĻ āĻāĻžā§āĻžāĻ āĻā§āĻāĻžā§āĻŦ āύāϤ ā§ āύ āĻāĻžāĻ āĻšā§āϞāĻžāĨ¤ āϏāĻŽ ā§ āĻš: āĻ) āĻ āĻ āĻāύ āĻāĻžā§āĻ āĻŦā§āϞ? āĻ) āĻ āĻ āĻāύā§āύ āĻā§āĻĒāĻžāĻŋāĻĻāϤ āĻāĻŋāĻĻ āĻŽāĻžāϤ ā§ āĻāĻŋā§āĻĻāϰ āĻŽā§āϤāĻž āĻŖāϏ āĻšā§ āĻāύ? āĻŦāĻžāĻāĻž āĻā§āϰāĻžāĨ¤ āĻ) āĻā§āĻĒā§āĻ āĻāϞ ā§ āϰ āĻŽāĻžāϧā§āĻŽ āĻāĻžāύ āϧāϰā§āύāϰ āĻ āĻ āĻāύ āĻā§āĻā§āĻ? āĻŦāĻžāĻāĻž āĻā§āϰāĻžāĨ¤ āĻ) āĻ ā§ āĻŋāĻŽ āĻ āĻ āĻāύ (āϝāĻŽāύ āĻāϞāĻŽ) āĻ āĻā§āĻĒā§āĻāϰ āĻžāĻ ā§ āĻŋāϤāĻ āĻ āĻ āĻāύā§āύāϰ āĻŽā§āϧ āĻĒāĻžāĻĨ āĻ āĻŋāĻŦā§āώāĻŖ āĻā§āϰāĻžāĨ¤ CQ 2: āĻŋāύāĻŋāώāĻāϰāĻŖ āĻ āĻĢāϞ āĻāĻ āύ āĻā§āĻĒāĻ: āĻŦāϏāĻāĻžā§āϞ āĻ ā§ āώāĻ āϤāĻžāϰ āĻŦāĻžāĻāĻžā§āύāϰ āĻāĻŽ āĻāĻžā§āĻ āĻ ā§ āϰ āĻĢ ā§ āϞ āĻĢ ā§ āĻā§āϤ āĻĻāĻā§āϞāύāĨ¤ āĻŋāĻāĻ ā§ āĻŋāĻĻāύ āĻĒāϰ āĻŋāϤāĻŋāύ āϞ āĻāϰā§āϞāύ āĻĢ ā§ ā§āϞāϰ āĻĒāĻžāĻĒāĻŋā§ āĻā§āϰ āĻŋāĻā§ā§ āĻāĻžāĻ āĻāĻžāĻ āĻā§āĻŽāϰ āĻ ā§ āĻ āĻŋā§ āĻĻāĻāĻž āĻŋāĻĻā§ā§ā§āĻāĨ¤ āĻŋāϤāĻŋāύ āĻŋāϤāĻŋāĻĻāύ āĻāĻžāĻāĻŋāĻ āĻĒāϝ ā§āĻŦāĻŖ āĻāϰā§āϤ āϞāĻžāĻā§āϞāύāĨ¤ āϏāĻŽā§ā§āϰ āϏāĻžā§āĻĨ āϏāĻžā§āĻĨ āĻŋāϤāĻŋāύ āĻĻāĻā§āϞāύ āĻāϤ āĻžāĻļā§āĻŋāĻ āϧā§ā§āϰ āϧā§ā§āϰ āĻŦā§ āĻšā§ā§ āĻĢā§āϞ āĻĒāĻŋāϰāĻŖāϤ āĻšā§āĨ¤ āĻā§ā§āĻ āϏāĻžāĻš āĻĒāϰ āĻāĻžāĻāĻāĻŋāϤ āĻĒāĻžāĻāĻž āĻāĻŽ āĻĻā§āĻ āĻŋāϤāĻŋāύ āĻ ā§ āĻŋāĻļ āĻšā§āϞāύāĨ¤ āϏāĻŽ ā§ āĻš: āĻ) āĻĢāϞ āĻāĻžā§āĻ āĻŦā§āϞ? āĻ) āĻ ā§ āϤ āĻĢāϞ āĻ āĻ āĻ ā§ āϤ āĻĢā§āϞāϰ āĻĒāĻžāĻĨ āĻ āĻŦāĻžāĻāĻž āĻā§āϰāĻžāĨ¤ āĻ) āĻā§āĻĒā§āĻ āĻāĻŽ āĻāĻžā§āĻ āĻāϤ āĻžāĻļā§ āĻĢā§āϞ āĻĒāĻŋāϰāĻŖāϤ āĻšāĻā§āĻžāϰ āĻŋā§āĻžāĻŋāĻ āĻŦāĻžāĻāĻž āĻā§āϰāĻžāĨ¤ āĻ) āĻŋāύāĻŋāώāĻāϰāĻŖ āĻŋā§āĻž āύāĻž āĻāĻā§āϞ āĻĢāϞ āĻāĻ ā§āύ āĻā§ āĻāĻžāĻŦ āĻĒā§āϤ āϤāĻž āĻŋāĻŦā§āώāĻŖ āĻā§āϰāĻžāĨ¤ CQ 1: āĻāĻŋā§āĻĻāϰ āĻā§āϞāĻžāĻ āϏāĻžā§āĻž āĻā§āĻĒāĻ: āĻŋāĻŦāĻžāύ āĻā§āϰ āĻāύ āĻĢāĻžāĻŋāϰāĻšāĻž āϤāĻžāϰ āĻā§ āĻāĻžāύāĻžāϞāĻžāϰ āĻĒāĻžā§āĻļ āĻāĻāĻŋāĻ āĻā§āĻŦ āĻāĻāĻŋāĻ āĻāĻžāϰāĻžāĻāĻžāĻ āϞāĻžāĻŋāĻā§ā§āĻŋāĻāϞāĨ¤ āϏ āĻŋāϤāĻŋāĻĻāύ āĻŋāύā§āĻŽ āĻā§āϰ āĻāĻžāĻāĻŋāĻā§āϤ āĻĒāĻžāĻŋāύ āĻŋāĻĻāϤ āĻāĻŦāĻ āĻāϰ āĻŦ ā§ āĻŋ āĻĒāϝ ā§āĻŦāĻŖ āĻāϰāϤāĨ¤ āĻā§ā§āĻāĻŋāĻĻāύ āĻĒāϰ āϏ āϞ āĻāϰāϞ āĻāĻžāĻāĻŋāĻāϰ āĻāĻž āĻāĻžāύāĻžāϞāĻžāϰ āĻŦāĻžāĻā§āϰ āĻŋāĻĻā§āĻ āĻŦ āĻ ā§āĻ āĻā§āĻāĨ¤ āĻ āĻĨāĻ āĻŽāĻžāĻŋāĻāϰ āĻŋāύā§āĻ āĻĨāĻžāĻāĻž āĻŽ ā§ āϞā§āϞāĻž āĻā§āϞāĻžāϰ āĻā§ā§āϏāϰ āĻŋāĻŦāĻĒāϰā§āϤ āĻŋāĻĻā§āĻ āĻŦā§ā§ āĻā§āϞā§āĻāĨ¤ āĻāĻ āĻŋāĻāĻŽ ā§ āĻā§ āĻŦ ā§ āĻŋ āĻĻā§āĻ āĻĢāĻžāĻŋāϰāĻšāĻž āϤāĻžāϰ āĻŋāĻļāĻā§āĻ āĻāϰāϞāĨ¤ āϏāĻŽ ā§ āĻš: āĻ) āĻ āĻŋāύ āĻā§? āĻ) āĻŋāĻĒāĻ āĻāϞāύ āĻāĻžā§āĻ āĻŦā§āϞ? āĻŦāĻžāĻāĻž āĻā§āϰāĻžāĨ¤ āĻ) āĻā§āĻĒā§āĻ āĻāĻžā§āĻāϰ āĻāĻž āĻāĻžāύāĻžāϞāĻžāϰ āĻŋāĻĻā§āĻ āĻŦāĻžāĻāĻāĻžāϰ āĻāĻžāϰāĻŖ āĻŦāĻžāĻāĻž āĻā§āϰāĻžāĨ¤ āĻ) āĻā§āĻĒā§āĻ āĻāĻž āĻ āĻŽ ā§ ā§āϞāϰ āĻŋāĻ āĻŋāĻĻā§āĻ āĻŦ ā§ āĻŋāϰ āĻāĻžāϰāĻŖ āϤ ā§ āϞāύāĻžāĻŽ ā§ āϞāĻ āĻŋāĻŦā§āώāĻŖ āĻā§āϰāĻžāĨ¤ CQ 2: āĻŋāϤāĻŦāϤ āĻŋā§āĻž āĻā§āĻĒāĻ: āĻāĻ āϏāĻžā§ āĻāĻžāĻŋāĻšāĻĻ āĻŦāĻžāϰāĻžā§ āĻŦā§āϏ āĻŦāĻ āĻĒā§āĻŋāĻāϞāĨ¤ āĻšāĻ āĻžā§ āϤāĻžāϰ āĻšāĻžā§āϤ āĻāĻāĻŋāĻ āĻŽāĻļāĻž āĻā§āϏ āĻŦāϏāϞāĨ¤ āĻŽāĻļāĻžāĻŋāĻ āϤāĻžāϰ ā§āĻ āϞ āĻĢāĻžāĻāĻžā§āύāĻžāϰ āϏāĻžā§āĻĨ āϏāĻžā§āĻĨ āϏ āϏ⧠āϏ⧠āĻšāĻžāϤ āύāĻžāĻŋā§ā§ā§ āĻŽāĻļāĻžāĻŋāĻ āϤāĻžāĻŋā§ā§ā§ āĻŋāĻĻāϞāĨ¤ āĻāĻ āĻĒ ā§ ā§āϰāĻž āĻāĻāύāĻžāĻŋāĻ āĻāĻāϞ āĻāĻžā§āύāĻž āĻŋāĻāĻžāĻāĻžāĻŦāύāĻž āĻāĻžā§āĻžāĻ, āĻā§āĻāĻŦāĻžā§āϰ āĻŽ ā§ āĻš ā§ ā§āϤ āϰ āĻŽā§āϧāĨ¤ āĻĒā§āϰ āĻāĻžāĻŋāĻšāĻĻ āĻāĻžāĻŦāϞ, āϏ āĻāϏā§āϞ āĻā§āĻāĻžā§āĻŦ āĻāϤ āϤ āϏāĻžā§āĻž āĻŋāĻĻāϞāĨ¤ āϏāĻŽ ā§ āĻš: āĻ) āĻŋāϤāĻŦāϤ āĻŋā§āĻž āĻāĻžā§āĻ āĻŦā§āϞ? āĻ) āĻŋāϏāύāĻžāĻĒāϏ āĻā§? āĻŦāĻžāĻāĻž āĻā§āϰāĻžāĨ¤ āĻ) āĻā§āĻĒā§āĻ āĻāĻžāĻŋāĻšā§āĻĻāϰ āĻšāĻžāϤ āϏāϰāĻžā§āύāĻžāϰ āĻŋā§āĻžāĻŋāĻ āĻŦāĻžāĻāĻž āĻā§āϰāĻžāĨ¤ āĻ) āĻŋāϤāĻŦāϤ āĻā§āϰ āĻĒāĻžāĻāĻāĻŋāĻ āĻ āĻāĻļ āĻā§āĻāĻžā§āĻŦ āĻā§āĻĒā§āĻāϰ āĻāĻāύāĻžā§ āĻāĻžāĻ āĻā§āϰā§āĻ āϤāĻž āĻŋāĻŦā§āώāĻŖ āĻā§āϰāĻžāĨ¤ Figure 3: Examples of chapter-wise creative questions generated by TEACHMATEGPT for Chapter 3 (Diffusion, Osmosis and Transpiration), Chapter 4 (Reproduction in Plants), and Chapter 5 (Coordination and Excretion) of the NCTB Class 8 Science textbook. For each chapter, TeachMateGPT retrieves chapter-specific curriculum evidence and produces a board-styleUÃÂpkwith four progressive sub-questions (kâG). The examples show consistent curriculum grounding, board-style structure, and formatting across different science topics. C Ablation Study This appendix consolidates the ablation and base- line comparisons reported across the paper: au- tomatic evaluation (Table 1), human evaluation by three practicing science teachers (Table 2), re- trieval reliability under the fail-closed coverage gate (Table 9), and inference-time and one-time indexing efficiency (Table 10). Each table iso- lates one of TEACHMATEGPTâs five components, COPE, SAVER, CCR, CCI, the IAC, or its re- trieval strategy, supporting the conclusions below. C.1 Curriculum Indexing (COPE) Removing COPE (w/o COPE) causes the largest degradation in answer relevancy (0.89â0.76) and context precision (0.92â0.73) among the five component ablations (Table 1), together with a sub- stantial drop in human-rated stimulus realism (4.70 â3.90; Table 2), reflecting COPEâs role in sup- plying well-scoped, curriculum-aligned evidence rather than judging generated text. C.2 Source-Attributed Verification (SAVER) Removing SAVER (w/o SAVER) causes the largest faithfulness drop of any ablation (0.96â0.79; Ta- ble 1) and the largest pedagogical-alignment drop in human evaluation (4.90â3.25; Table 2), while answer relevancy and context metrics stay compar- atively high (0.86, 0.90) â consistent with SAVER acting as the final faithfulness check before teacher presentation, not a retrieval-quality mechanism. C.3 Redundancy Reduction and Context Restoration (CCR, CCI) Removing CCR or CCI individually produces smaller, more uniform degradations across all four automatic metrics (Table 1:w/o CCR 0.83/0.81/0.84/0.85; w/o CCI 0.84/0.82/0.85/0.86) than removing COPE or SAVER, with moderate re- ductions in human-rated quality (Table 2). On our 16-query bank, both converge to nearly identical fail-closed rates and coverage ratios (18.8%, 0.701 vs. 18.8%, 0.705; Table 9), since no query triggers a missing-parent or near-duplicate case; we expect divergence on a larger, more redundant evidence pool. C.4 Query Routing Removing the IAC produces a moderate, uni- form drop across all four automatic metrics (0.81/0.77/0.89/0.88; Table 1) and human-rated 13 Chapter: 3 Chapter Title (Bangla): āĻŦāĻžāĻĒāύ, āĻ āĻŋāĻāĻŦāĻŖ āĻ ā§āĻĻāύ Chapter Title (English): Diffusion, Osmosis and Transpiration MCQ 1. āĻ āĻŖ ā§ āϏāĻŽ ā§ āĻš āĻŦāĻŋāĻļ āĻāύā§āϰ āĻžāύ āĻĨā§āĻ āĻāĻŽ āĻāύā§āϰ āĻžā§āύ āĻāĻŋā§ā§ā§ āĻĒā§āĻžāϰ āĻŋā§āĻžā§āĻ āĻā§ āĻŦā§āϞ? āĻ) āĻŦāĻžāĻĒāύ āĻ) āĻ āĻŋāĻāĻŦāĻŖ āĻ) ā§āĻĻāύ āĻ) āĻāĻŽāĻŦāĻžāĻāĻŋāĻŦāĻļāύ MCQ 2. āϝ āĻĒāĻĻ āĻž āĻŋāĻĻā§ā§ āĻžāĻŦāĻ āĻ āĻŦ āĻāĻā§āĻ āĻ āĻŖ ā§ āϏāĻšā§āĻ āĻāϞāĻžāĻāϞ āĻāϰā§āϤ āĻĒāĻžā§āϰ āϤāĻžā§āĻ āĻā§ āĻŦā§āϞ? āĻ) āĻ ā§āĻāĻĻ āĻĒāĻĻ āĻž āĻ) āĻ āϧ ā§āĻāĻĻ āĻĒāĻĻ āĻž āĻ) āĻāĻĻ āĻĒāĻĻ āĻž āĻ) āĻžāĻāĻŽāĻž āĻĒāĻĻ āĻž MCQ 3. āĻāĻāĻŋāĻ āĻ āϧ ā§āĻāĻĻ āĻĒāĻĻ āĻž āĻŋāĻĻā§ā§ āĻžāĻŦāĻ āĻ āĻŖ ā§ āϰ āĻāĻŽ āĻāύā§āϰ āĻŦāĻŖ āĻĨā§āĻ āĻŦāĻŋāĻļ āĻāύā§āϰ āĻŦā§āĻŖāϰ āĻŋāĻĻā§āĻ āĻāĻŽāύā§āĻ āĻā§ āĻŦā§āϞ? āĻ) āĻŦāĻžāĻĒāύ āĻ) āĻ āĻŋāĻāĻŦāĻŖ āĻ) ā§āĻĻāύ āĻ) āĻĒāĻŋāϰāĻŦāĻšāύ MCQ 4. āĻā§āĻŦā§āĻāĻžā§āώāϰ āĻāĻžāώāĻžāĻŦāϰāĻŖ āĻŦāĻž āĻžāĻāĻŽāĻž āĻĒāĻĻ āĻž āĻāĻžāύ āϧāϰā§āύāϰ āĻĒāĻĻ āĻž āĻŋāĻšā§āϏā§āĻŦ āĻāĻžāĻ āĻā§āϰ? āĻ) āĻ ā§āĻāĻĻ āĻĒāĻĻ āĻž āĻ) āĻāĻĻ āĻĒāĻĻ āĻž āĻ) āĻ āϧ ā§āĻāĻĻ āĻĒāĻĻ āĻž āĻ) āĻāĻžā§āύāĻžāĻŋāĻāĻ āύ⧠Chapter 4 Chapter Title (Bangla): āĻāĻŋā§āĻĻāϰ āĻŦāĻāĻļ āĻŦ ā§ āĻŋ Chapter Title (English): Reproduction in Plants MCQ1. āĻĻ ā§ āĻŋāĻ āĻŋāĻāϧāĻŽ ā§ āĻāύ āĻāĻžā§āώāϰ āĻŋāĻŽāϞāύ āĻāĻžā§āĻžāĻ āϝ āĻāύ āϏ āĻšā§ āϤāĻžā§āĻ āĻā§ āĻŦā§āϞ? āĻ) āϝā§āύ āĻāύ āĻ) āĻ ā§āϝā§āύ āĻāύ āĻ) āĻŋāύāĻŋāώāĻāϰāĻŖ āĻ) āĻĒāϰāĻžāĻāĻžā§āύ MCQ2. āĻāϞ ā§ āϰ āĻāĻžā§ā§ āĻĨāĻžāĻāĻž 'āĻāĻžāĻ' āĻĨā§āĻ āύāϤ ā§ āύ āĻāĻŋāĻĻ āĻ āύ⧠āĻāĻžāύ āϧāϰā§āύāϰ āĻ āĻ āĻāύā§āύ? āĻ) āĻ (Tuber) āĻ) āϰāĻžāĻā§āĻāĻžāĻŽ āĻ) āĻžāϞāύ āĻ) āĻ ā§ āĻ āĻŋā§ (Bulbil) MCQ 3. āĻāĻĻāĻž āĻāĻžāύ āϧāϰā§āύāϰ āĻĒāĻžāĻŋāϰāϤ āĻāĻžā§āϰ āĻāĻĻāĻžāĻšāϰāĻŖ? āĻ) āĻ āĻ) āϰāĻžāĻā§āĻāĻžāĻŽ āĻ) āĻŦāĻž āĻ) āĻžāϞāύ MCQ 4. āĻāĻāĻŋāĻ āĻāĻĻāĻļ āĻĢ ā§ ā§āϞāϰ āĻā§āĻŋāĻ āĻ āĻāĻļ āĻĨāĻžā§āĻ? āĻ) āĻŋāϤāύāĻŋāĻ āĻ) āĻāĻžāϰāĻŋāĻ āĻ) āĻĒāĻžāĻāĻāĻŋāĻ āĻ) āĻā§āĻŋāĻ Chapter 5 Chapter Title (Bangla): āϏāĻŽā§ āĻ āĻŋāύāĻāϏāϰāĻŖ Chapter Title (English): Coordination and Excretion MCQ 1. āĻāĻŋā§āĻĻāϰ āĻŦ ā§ āĻŋ āĻŋāύā§āĻŖāĻāĻžāϰ⧠āĻāĻŦ āϰāĻžāϏāĻžā§āĻŋāύāĻ āĻĒāĻĻāĻžāĻĨ ā§āĻ āĻā§ āĻŦā§āϞ? āĻ) āĻāύāĻāĻžāĻāĻŽ āĻ) āĻĢāĻžāĻā§āĻāĻžāĻšāϰā§āĻŽāĻžāύ āĻ) āĻŋāĻāĻāĻžāĻŋāĻŽāύ āĻ) āĻŋāĻšā§āĻŽāĻžā§āĻžāĻŋāĻŦāύ MCQ 2. āĻāĻŋā§āĻĻāϰ āĻŖāĻŽ ā§ āĻ ā§ āϞāĻžāĻŦāϰāĻŖā§āϰ āĻāĻĒāϰ āĻā§āϞāĻžāϰ āĻāĻžāĻŦ āĻĨāĻŽ āĻ āϞ āĻā§āϰāύ? āĻ) āĻāϰ āĻŽā§āϞ āĻ) āĻāĻžāϞ āϏ āĻĄāĻžāϰāĻāĻāύ āĻ) āĻāύ āĻĄāĻžāύ āĻ) āĻŋāύāĻāĻāύ MCQ3. āĻāĻžāύ āĻšāϰā§āĻŽāĻžāύ āĻĢāϞ āĻĒāĻžāĻāĻžā§āϤ āϏāĻžāĻšāĻžāϝ āĻā§āϰ? āĻ) āĻ āĻŋāύ āĻ) āĻŋāĻā§āϰāĻŋāϞāύ āĻ) āĻāĻŋāĻĨāĻŋāϞāύ āĻ) āϏāĻžāĻā§āĻāĻžāĻāĻžāĻāĻŋāύ MCQ4. āĻžā§ ā§ āϤā§āϰ āĻāĻžāĻ āĻŋāύāĻ āĻ āĻāĻžāϝ āĻāϰ⧠āĻāĻā§āĻāϰ āύāĻžāĻŽ āĻā§? āĻ) āĻŋāύāĻāϰāύ āĻ) āĻŋāϏāύāĻžāĻĒāϏ āĻ) āĻ āĻžāύ āĻ) āĻĄāύ Figure 4: Examples of MCQ items from Chapter 3 (Diffusion, Osmosis and Transpiration), Chapter 4 (Reproduction in Plants), and Chapter 5 (Coordination and Excretion) of the NCTB Class 8 Science textbook illustrate how the TEACHMATEGPT framework converts the same retrieved evidence used for creative-question generation into a compact, single-answer format. For each chapter, the MCQ specialist grounds a Bangla stem in the source passage and generates four options (kâG), from which the student selects the correct answer, while the remaining options serve as plausible, textbook-consistent distractors rather than invented facts. criteria (Table 2), showing routing also protects generation quality, not just efficiency. Its efficiency impact is larger: on the mixed-intent set for Ta- ble 10 (2 science, 3 greeting, 3 off-topic, 3 harmful, 5 ambiguous), removing it more than doubles mean latency (11.3sâ26.3s) and triples queries reach- ing generation (3/16â9/16), confirming routing filters unsafe, off-topic, or underspecified queries before the costlier stages. C.5 Retrieval Strategy and the Coverage Gate Table 9 isolates the retrieval strategy. Dense-only retrieval more than doubles the fail-closed rate vs. full hybrid (31.3% vs. 12.5%) and lowers coverage (0.618 vs. 0.724), while BM25-only is competitive (18.8%, 0.682), showing exact curriculum terminol- ogy stays informative for OCR-derived Bangla text- books. Removing the coverage gate eliminates re- fusals entirely (0.0%) but yields the second-lowest coverage ratio (0.649) of all eleven configurations, confirming the gate trades a few refusals for a large gain in evidence sufficiency rather than acting as a redundant safeguard. C.6 Comparison Against Prior RAG Baselines Across all four tables, TEACHMATEGPT out- performs four RAG baselines, RAPTOR (Sarthi et al., 2024), GraphRAG (Edge et al., 2025), CRAG (Yan et al., 2024), and Adaptive RAG (Jeong et al., 2024), on every automatic and human-rated cri- terion (Tables 1â2), while requiring zero LLM calls and zero seconds of one-time index construc- tion, versus 66 calls / 224.3s for RAPTOR and 82 calls / 287.5s for GraphRAG (Table 10). CRAG is the strongest baseline on faithfulness (0.86) and coverage (0.708; 18.8% fail-closed), reflecting its retrieval-time evaluator and query rewriting, but still trails TEACHMATEGPT on every metric here. D Experimental Details D.1 System Configuration TEACHMATEGPTisimplementedusing LangChain (Chase, 2022) and LangGraph (Inc., 2024) to orchestrate the multi-agent workflow. Curriculum embeddings are generated with the locally deployed BAAI/bge-m3 (Chen et al., 2024) embedding model (1024-dimensional) and indexed in the Qdrant (Qdrant Team, 2025) vector database for dense retrieval. Assessment generation is performed using GPT-4o mini (OpenAI, 2024). PDF processing uses PyMuPDF for native text extraction, while scanned pages are rendered as images and processed with GPT-4o (OpenAI, 2024) vision-based OCR when native text is unavailable.This framework enables robust curriculum indexing across both digitally 14 Ch.Chapter Name (Bangla)Chapter Name (English)PagesMCQCQSubject 1aiNjgetreiNibnYasClassification of Animal Kingdom1â12134 Biology 2jÂebrbÂiÂĒzObKSgitGrowth & Heredity of Organisms13â23114 Biology 3bYapn,AivsRbNOeÂWdn Diffusion, Osmosis & Transpiration24â33104 Biology 4UiÃedrbKSbÂiÂĒzReproduction in Plants34â44104 Biology 5smà JOinhsrNCoordination & Excretion45â5494 Biology 6prmaNurgFnStructure of the Atom55â64114 Chemistry 7pÂiQbÂOmHak ĖPEarth & Gravitation65â74104 Physics 8rasaJinkibi ˧JaChemical Reactions75â88125 Chemistry 9b ĖtnÂOclibdYu Circuit & Current Electricity89â97104 Physics 10AÂL,XarkOlbNAcid, Base & Salt98â107104 Chemistry 11AaelaLight108â11884 Physics 12mHakaSOUpgRHSpace & Satellites119â12893 Physics 13xadYOpuiÃĢFood & Nutrition129â146104 Biology 14pirebSEbKbaïutÃEnvironment & Ecosystem147â156103 Environment Total - 14 Chapters | 156 Pages14355 Table 4: Per-chapter breakdown of the 14-chapter, 156-page NCTB Class 8 science corpus, showing subject area and the number of MCQ and creative-question (CQ) items authored per chapter for evaluation (143 MCQ and 55 CQ across all chapters, 198 items total). No.Validation GatePass rateWhat it checks 1MCQ: exactly 4 options + valid answer index + Bangla-dominant stem/options 137/143 = 95.8% Structural well-formedness 2MCQ: Bangla-dominant textâĨ 0.55 ratio143/143 = 100% Language-purity of item text 3CQ: starts with âUÃÂpk:â + narrative + not a theory-dump55/55 = 100% Board-style stimulus framing 4CQ:âĨ 5 sentences in the âUÃÂpkâ story55/55 = 100% Narrative sufficiency 5CQ: not a direct theory-dump opener55/55 = 100% Anti-extractive framing 6CQ:âĨ 0.85 Bangla-only ratio, no non-Bangla letters52/55 = 94.5% Language-purity of item text Table 5: Pass rates for the six deterministic validation gates applied to the 198 authored items (143 MCQ, 55 CQ), covering structural well-formedness and Bangla language-purity checks specific to each item type. These gates run prior to, and independently of, the SAVER faithfulness analysis in Table 14. generated and scanned textbook pages. Beyond the per-item SAVER gate, we evaluate the proposed framework using a two-tier proto- col combining corpus-level automatic RAG evalua- tion with human evaluation, applied comparatively across the full system and ablated configurations (Vanilla RAG baseline, w/o COPE, w/o SAVER, w/o CCR, w/o CCI, w/o IAC) to isolate each com- ponentâs contribution. D.2 Evaluation Protocol Section 3.6 covers per-item verification at genera- tion time; the protocol here measures whole config- urations instead. D.2.1 Automatic Evaluation For each configuration, we evaluate retrieval and generation quality using four metrics from the RAGAS framework (Es et al., 2024). Following its evaluation protocol, GPT-4.1 (OpenAI, 2025) serves as the LLM judge to score each generated assessment against its retrieved evidence. âĸFaithfulness. Measures whether the assess- ment is fully supported by the retrieved text- book evidence. A higher score indicates that the assessment avoids unsupported claims and hallucinated content. âĸAnswer Relevancy. Measures how well the generated assessment satisfies the teacherâs instructional request. Higher scores indicate that the assessment remains focused on the in- tended topic, concept, and learning objective. âĸContext Precision.Measures the quality of the retrieved evidence by estimating how much of the retrieved content is relevant to the teacherâs request. Higher precision indicates less irrelevant or noisy context. âĸContext Recall. Measures whether the re- trieved evidence contains the information re- quired to support the generated assessment. Higher recall indicates that the retrieval stage captures the necessary curriculum content for generation. Unlike SAVERâs binary per-item accept/reject deci- 15 Intent labelCount%Terminates without retrieval? Science query2170%No Greeting310%Yes (all 3 pure greetings) Harmful310%Yes (fixed reply) Off-topic310%Yes (fixed reply) Table 6: Distribution of intent-routing outcomes on a 30-user-query evaluation bank. Science queries (70%) continue to retrieval, while greeting, harmful, and off-topic inputs terminate early with fixed Bangla responses without consuming retrieval resources. MetricValue Deterministic guard accepts as non-ambiguous17 (81%) Model judges non-ambiguous1 (5%) Model judges ambiguous, guard overrides0 (0%) Ambiguous - clarification triggered3 (14%) Clarification rate (of full N=30 bank)3/30 = 10% Over-clarification rate (exam-style prompts wrongly flagged)0/16 = 0% Table 7: Ambiguity-gate and clarification outcomes over the 21 science-query turns from Table 6. The deterministic specificity guard resolves 81% as non-ambiguous without a model call; the remainder are judged by the ambiguity model, yielding a 10% overall clarification rate and 0% over-clarification on exam-style prompts. sion, RAGAS produces a continuous score in[0, 1] for each metric, which we average per configura- tion. This lets quality differences be attributed to specific components by comparing the full system against the COPE-, SAVER-, CCR-, CCI-, and IAC-ablated variants. D.2.2 Teacher-in-the-Loop Evaluation Automatic metrics alone cannot fully assess the educational quality of generated assessments. We therefore conduct a human evaluation in which three practicing secondary-school science teach- ers independently assess a shared, configuration- blind subset of generated assessments. Each assess- ment is evaluated across four pedagogical dimen- sions: Pedagogical Alignment, Stimulus Realism (for CQs), Linguistic Fluency, and Overall Util- ity. Ratings are assigned on a 5-point Likert scale (Joshi et al., 2015): âĸ 1 â Poor. The assessment is unsuitable for classroom use because of major factual, peda- gogical, or structural errors and requires com- plete revision. âĸ2 â Fair. The assessment captures part of the intended objective but contains substan- tial issues that require major revisions before classroom use. âĸ3 â Acceptable. The assessment is generally correct and usable but requires minor revi- sions to improve clarity, alignment, or quality. âĸ4 â Good. The assessment is well aligned with the curriculum and suitable for classroom use, requiring only trivial edits. âĸ5 â Excellent. The assessment is fully aligned, factually accurate, pedagogically sound, and classroom-ready without modification. 16 FieldTypeDescription question_idstringUnique assessment identifier querystringTeacher query text selected_taskslistRequested assessment families chapterstringDetected chapter label conceptstringDetected concept label sourceslistRetrieved textbook evidence sources mcqobject/nullMultiple-choice assessment item mcq.questionstringMCQ stem text mcq.optionslistFour answer options mcq.answer_indexintegerCorrect option index (0â3) creativeobject/nullCreative assessment item creative.headlinestringOptional item title creative.cluestringStimulus text creative.partslistFour sub-questions Table 8: Schema of the TEACHMATEGPT output dataset, where each record pairs a teacher query with its retrieval evidence (sources), inferred curriculum metadata (chapter,concept), and generated assessment items. MCQ items follow a four-option format with an indexed correct answer; creative items follow the board-style creative format with a stimulus and four graded sub-questions. ConfigurationFail-closed rateCoverage ratioNotes TEACHMATEGPT (dense + BM25 + CCI + CCR + coverage gate)12.5% (2/16)0.724Default configuration Dense-only31.3% (5/16)0.618Lexical evidence often missed without BM25 BM25-only18.8% (3/16)0.682Semantic matches frequently unavailable w/o CCI18.8% (3/16)0.701Inconsistent retrieved evidence occasionally reduced support w/o CCR18.8% (3/16)0.705Duplicate chunks lowered effective evidence diversity w/o coverage gate0.0% (0/16)0.649All queries answered regardless of evidence sufficiency w/o graph expansion25.0% (4/16)0.676Related supporting chunks were not retrieved RAPTOR25.0% (4/16)0.661Recursive clustering and summarisation caused loss of fine-grained textbook evidence GraphRAG18.8% (3/16)0.693Entity graph expansion improved recall but introduced broader contexts CRAG18.8% (3/16)0.708Retrieval evaluator and query rewriting improved evidence quality Adaptive RAG25.0% (4/16)0.671Query routing reduced unnecessary retrieval but missed some supporting evidence Table 9: Retrieval ablation and comparison over 16 non-ambiguous Bangla science queries. The default hybrid configuration is compared with single-retriever variants, component ablations, and RAG baselines. Fail-closed rate and coverage ratio are reported. Pedagogical Alignment Stimulus Realism Linguistic Fluency Overall Utility 1 2 3 4 5 Teacher-in-the-Loop Human Evaluation Across Baselines and Ablations Vanilla RAG TeachMateGPT w/o COPE w/o SAVER w/o CCR w/o CCI w/o IAC RAPTOR GraphRAG CRAG Adaptive RAG (a) Teacher-in-the-loop evaluation. Macro 181 (23%) Meso 255 (32%) Micro 357 (45%) 793 chunks COPE Knowledge-Base Chunk Distribution Across Hierarchical Retrieval Tiers (b) COPE chunk tier composition. Figure 5: Qualitative and structural analysis of TEACHMATEGPT. (a) Human evaluation by three science teachers across four assessment-quality criteria, showing that the full system consistently outperforms all ablations. (b) Distribution of the hierarchical knowledge base across the three COPE resolution tiers, illustrating the multi- resolution chunking strategy for broad contextual and fine-grained factual retrieval. 17 FaithfulnessAnswer Relevancy Context Precision Context Recall 0.0 0.2 0.4 0.6 0.8 1.0 Score 0.68 0.60 0.54 0.58 0.96 0.89 0.92 0.91 0.86 0.76 0.73 0.75 0.79 0.86 0.90 0.89 0.83 0.81 0.84 0.85 0.84 0.82 0.85 0.86 0.81 0.77 0.89 0.88 0.80 0.75 0.71 0.78 0.84 0.79 0.80 0.83 0.86 0.81 0.84 0.85 0.79 0.76 0.73 0.79 Performance Comparison of TeachMateGPT, Baseline, and Ablation Variants Vanilla RAG TeachMateGPT w/o COPE w/o SAVER w/o CCR w/o CCI w/o IAC RAPTOR GraphRAG CRAG Adaptive RAG Figure 6: Quantitative comparison of TEACHMATEGPT with a vanilla RAG baseline, five component ablations (w/o COPE, SAVER, CCR, CCI, and IAC), and three representative RAG baselines (RAPTOR, GraphRAG, and CRAG) across four generation metrics: Faithfulness, Answer Relevancy, Context Precision, and Context Recall (Table 1). Each ablation removes one architectural component to measure its contribution to retrieval quality and assessment generation. ConfigurationMean LLM Calls / QueryMean Latency (s)Queries GeneratingOne-Time Index LLM CallsIndex Build (s) TEACHMATEGPT2.111.33 / 1600 Vanilla RAG1.039.816 / 1600 w/o IAC1.826.39 / 1600 RAPTOR1.042.316 / 1666224.3 GraphRAG1.446.716 / 1682287.5 CRAG2.239.016 / 1600 Adaptive RAG1.628.58 / 1600 Table 10: Inference-time and one-time indexing efficiency of TEACHMATEGPT, an IAC ablation, and four RAG baselines, evaluated on a separate 16-query mixed-intent set (2 science, 3 greeting, 3 off-topic, 3 harmful, and 5 ambiguous queries), distinct from the 16 non-ambiguous science queries used in Table 9. Queries Generating denotes the number of queries reaching the generation stage after intent and ambiguity routing. Mean LLM Calls/Query and Mean Latency are averaged over all 16 queries, while One-Time Index LLM Calls and Index Build Time measure offline indexing cost. Only RAPTOR and GraphRAG require LLM-assisted index construction. Prior WorkLang.CARMulti-AgentEvidence VerificationMulti-Format GenerationPD+HV N et al. (2025)ENâMCQâ Hamidi et al. (2025) (JAIT)ENââMCQ + QAâ Duan et al. (2026) (CODE-GEN)ENââMCQâ Chen and Shiu (2025) (KAQG)ENââMCQâ Tian et al. (2026) (ReQUESTA)ENââMCQâ Khan and Khan (2025) (B-RAG)BNâââQA onlyâ Eyasir et al. (2026) (NCTB-QA)BNâââQA onlyâ Reza et al. (2025)ENââManualPartialâ Wong et al. (2026)ENââMCQ + SAâ TEACHMATEGPTBNâMCQ + CQâ Table 11: Comparison of TEACHMATEGPT with representative educational assessment generation systems. Existing approaches improve individual components of the generation pipeline but lack a unified framework that combines curriculum grounding, diverse assessment generation, evidence verification, and expert-reviewed evaluation. Here, CAR denotes Curriculum-Aware Retrieval, and PD+HV denotes Public Dataset with Human Validation. (â) denotes full support, (â) partial support, and (â) no reported support. 18 MetricMCQCQ Mean items generated / target5 / 5 1 / 1 % turns satisfied by strict pass (attempt 1) 12/14 = 85.7% 14/14 = 100% % turns requiring relaxed pass2/14 = 14.3% n/a Table 12: Generation yield over the 14 evaluated turns: mean items produced against target (5 MCQ, 1 CQ per turn) and the share of turns meeting the strict validation gate on the first attempt versus requiring the relaxed-pass fallback. CQâs 100% strict-pass rate reflects the post-rewrite corpus; the original pass achieved only 7.1%, isolating narrative style-not gate strictness-as the cause. DifficultynMean stem lengthReasoning-depth rating Beginner11653.2 Low: direct recall/definition (âkÂâ, âkaekbelâ, âekaniTâ) Intermediate1152.7 Medium: comparison/explanation (âpa ĖQkYâ, âeknâ, âbYaxYaâ) Advanced1655.9 Medium-high: numeric/multi-step reasoning (âin ĖNykeraâ) Table 13: Difficulty-level distribution of the 143 evaluated MCQs, showing the item count, mean stem length, and dominant reasoning-depth patterns for each difficulty level. The levels range from direct recall and definition-based questions at the beginner level to numeric and multi-step reasoning questions at the advanced level, based on the characteristic Bangla task verbs used. MetricMCQ (n=143)CQ (n=55)All items (n=198) Faithfulness (spot-checkedâŧ15 items / all 55 CQ clues) 0 errors found0 fabricated facts found0 factual errors found Answer relevance (on-topic by construc- tion) 143/14355/55198/198 Hallucination risk (invented specifics)None identifiedNone identifiedNone identified Format/structural risk6/143 fail strict gate3/55 fail Bangla-only ratio9/198 total (6 MCQ format + 3 CQ notation) Table 14: Manual audit of SAVER outcomes across generated items by question type. Faithfulness was assessed through spot checks on a 15-item MCQ sample and complete evaluation of all 55 CQ clues, with no factual errors or fabricated content detected. Answer relevance is ensured through direct generation from source passages, while format/structural risk captures type-specific violations, such as MCQ option constraints and Bangla-only notation requirements, rather than content-level errors. 19 E Detailed Analysis of Research Questions RQ1: Curriculum Knowledge Base Construction We first asked whether COPEâs hierarchical indexing actually preserves enough curriculum structure to support reliable assessment authoring, rather than simply reorganizing the same flat-chunking problem under a different name. The macro, meso, and micro decomposition (Table 3) spans all 14 chapters and four subject areas at comparable per-chapter depth (Table 4), and four of the six deterministic item-validity gates reach a full 100% pass rate (Table 5). The two gates that fall short, MCQ structural well-formedness at 95.8% and CQ Bangla-script purity at 94.5%, fail because of option-key formatting and embedded scientific notation, not because curriculum content was missing or misplaced. This distinction is important for how we interpret the result. A coverage failure would point back to COPEâs segmentation logic, whereas a format failure points instead to the generation layer downstream of retrieval. Since none of the failures in Table 5 trace back to missing chapter evidence, the results indicate that preserving parentâchild curriculum structure through hierarchical indexing, rather than collapsing the textbook into uniform token windows, is not the limiting factor for assessment validity in TEACHMATEGPT. At the same time, the remaining failures on these two validation gates show that curriculum indexing alone is insufficient. Structural formatting and Bangla-language purity checks continue to identify genuine generation errors even when retrieval succeeds, supporting the need for the gated generation and verification pipeline evaluated in RQ4 rather than treating it as an optional safeguard. RQ2: Intent, Ambiguity, and Clarification Routing We next examined whether a lightweight routing layer can protect retrieval and generation from unsafe or underspecified requests without becoming a nuisance to teachers who already provide clear queries. Intent routing filters out 30% of the evaluation bank, split evenly across greetings, harmful requests, and off-topic queries, before any retrieval is attempted (Table 6). Within the remaining science queries, a deterministic Bangla specificity guard resolves 81% of ambiguity decisions on its own, and the ambiguity model never overrides the guardâs judgment (Table 7). The cascade is designed to jointly improve efficiency and safety. Unsafe or irrelevant requests receive a predefined response instead of triggering the full retrieval and generation pipeline, while simple lexical rules resolve most ambiguity cases. As a result, the computationally more expensive model is invoked only for genuinely borderline cases, approximately one in five science queries in our evaluation. The most important observation, however, is not computational efficiency but usability. None of the 16 already explicit exam-style prompts triggered an unnecessary clarification request, indicating that teachers who formulate complete assessment requests are not interrupted by redundant follow- up questions. Since this evaluation includes a relatively small number of explicit prompts, we interpret this finding as encouraging evidence rather than a precise estimate of the true false-positive clarification rate. 20 RQ3: Hybrid Retrieval, Refinement, and Fail-Closed Coverage We further investigated whether the retrieval pipeline can maintain both safety and utility by refusing assessment generation when curriculum evidence is insufficient. Across eleven retrieval configura- tions, a clear trade-off emerges between fail-closed behavior and evidence coverage (Table 9). The full hybrid pipeline achieves the best balance, with a fail-closed rate of 12.5% and a coverage ratio of 0.724. Dense-only retrieval increases refusal to 31.3% while reducing coverage to 0.618. In contrast, removing the coverage gate eliminates refusals but yields the second-lowest coverage ratio (0.649) of all eleven configurations, above only Dense-only. The comparison also highlights the value of lexical retrieval. BM25 performs competitively with the full hybrid pipeline, indicating that exact curriculum terminology remains highly informative for OCR-derived Bangla textbooks, where dense embeddings alone may overlook important lexical cues. The gated and ungated variants further show that the coverage gate does not reduce evidence quality; instead, it prevents responses supported by weak evidence. Overall, the hybrid retrieval pipeline provides the most effective balance between safety and curricu- lum coverage. The refinement stages narrow the safetyâcoverage trade-off rather than eliminating it. Finally, the CCI and CCR ablations produce very similar, though not identical, results on this evaluation set (coverage ratio 0.701 vs. 0.705; Table 9), likely because none of the 16 evaluation queries contains a missing-parent or near-duplicate retrieval case severe enough to separate the two components further. RQ4: Curriculum-Grounded Assessment Generation Given accepted evidence, both assessment formats reach their configured generation targets (Table 12). Every successful turn yields the requested 5 MCQ items and 1 CQ item. First-attempt reliability, however, differs substantially by format. MCQ items satisfy strict validation, with checks for the correct option count, a valid answer index, Bangla-dominant text, and evidence-grounded stems and distractors, on 12 of 14 turns (85.7%). The remaining 2 turns (14.3%) are recovered through a relaxed pass that retains format- and language-valid items under marginal grounding. CQ items, in contrast, satisfy strict validation on all 14 turns (100%), but only after the underlying stimulus narratives were rewritten using more, shorter story sentences. Under the original narrative style, the same unmodified gate, which requires the stimulus to begin with the required board cue, contain at least five story sentences, and avoid a direct theory-dump opening, passed only 1 of 14 CQ turns (7.1%) on the first attempt. This improvement from 7.1% to 100% is observed under identical retrieved evidence, the same validation gate, and the same generator, showing that the original failures resulted from a narrative- style mismatch with the five-sentence sufficiency check rather than limitations in curriculum coverage or model capability. Without this style correction, a live system would trigger the repair ladder (retry, relaxed pass, or force-fill) on almost every CQ turn, despite adequate evidence and model performance. Difficulty conditioning shows a different pattern (Table 13). Across the 143 authored MCQ items, 116 (81.1%) are labeled beginner, 11 (7.7%) intermediate, and 16 (11.2%) advanced. Qualitative reasoning-depth ratings increase across these tiers, from direct recall and definition at the beginner level, to comparison and explanation at the intermediate level, and numeric or multi-step reasoning at the advanced level. Despite this progression, mean stem length remains nearly constant at 53.2, 52.7, and 55.9 characters, respectively, a maximum difference of only 3.2 characters. This indicates that difficulty in TEACHMATEGPT is expressed through lexical and task-type cues embedded in the specialist prompt rather than measurable surface-form complexity. Consequently, stem length should not be interpreted as a proxy for difficulty, particularly given the relatively small intermediate and advanced subsets. 21 RQ5: Source-Attributed Verification and Teacher Review We finally evaluated whether pairing automatic source-attributed verification with teacher review can support trustworthy acceptance, editing, or cautioning of generated assessments. The manual audit in Table 14 finds zero fabricated facts across all 55 CQ clues and a 15-item MCQ spot check, with every item judged on-topic by construction (143/143 MCQ, 55/55 CQ); the only issues SAVER and the deterministic gates surface are structural or notation-level, 6 of 143 MCQs failing the strict structural gate and 3 of 55 CQs falling short of the Bangla-only ratio, not content-level hallucinations. Corpus-level automatic evaluation corroborates this picture. Against the Vanilla RAG baseline, TEACHMATEGPT raises faithfulness from 0.68 to 0.96 and context precision from 0.54 to 0.92 (Table 1). Teacher-in-the-loop ratings move in the same direction, with overall utility rising from 2.00 to 4.80 on the 5-point scale (Table 2). The two ablations isolate distinct roles rather than a single generic quality effect. Removing SAVER produces the largest faithfulness drop (0.96â0.79) and the largest pedagogical-alignment drop (4.90â3.25), consistent with SAVERâs role as the last check before teacher presentation. Removing COPE instead mainly reduces answer relevancy (0.89â0.76) and stimulus realism (4.70â3.90), consistent with COPEâs role in supplying well-scoped evidence rather than in judging the generated text itself. We read these results as evidence that automatic verification and human review are complementary rather than substitutable: SAVERâs scores and flagged items give a fast, per-item signal that surfaces structural and notation issues reliably, while teacher ratings capture pedagogical and stylistic judg- ments, such as stimulus realism, that a faithfulness score does not directly measure. Because SAVER flags rather than removes or edits items, and because the teacher panel is limited to three practicing teachers rating a configuration-blind sample (see Limitations, Section 7), we treat the reported scores as evidence that the verification layer is informative and directionally reliable, not as a substitute for continued teacher oversight before classroom use. F Detailed Pseudocode for the TEACHMATEGPT Framework 22 Algorithm 1 Workflow of the TEACHMATEGPT framework Input:Authorized curriculum corpusD; teacher queryq; difficulty levelâ; assessment family set S âMCQ, CQ; thresholds θ cov ,θ F ,θ R ,θ H Output: Verified assessment setAwith verification reportV, teacher-facing caution flagb accept , or an abstention when evidence or generation is insufficient Stage 1: Knowledge Base Construction (offline, once per corpus) 1: P â LOADANDNORMALIZE(D)⡠OCR fallback and text normalization 2: U â SEGMENTBYPEDAGOGICALHEADINGS(P)⡠chapter, lesson, exercise boundaries 3: C â MULTIRESOLUTIONCHUNK(U)⡠macro, meso, and micro chunks 4: G â BUILDCOPEGRAPH(C)⡠pedagogical hierarchy and cross-links 5: Kâ(c, Embed(c), Meta(c)) : câG⡠curriculum knowledge base Stage 2: Intent Analysis and Query Routing 6: if CLASSIFYINTENT(q)ˏ= Science then 7:return Fixed non-science response 8: end if 9: if ISAMBIGUOUS(q) then 10:return Clarification request 11: end if Stage 3: Hybrid Retrieval and Coverage Validation 12: E â HYBRIDRETRIEVE(K,q)⡠dense retrieval + BM25 + reranking, CCI, CCR 13: if COVERAGE(E) < θ cov then 14:return Abstain⡠fail-closed: no assessment is generated 15: end if Stage 4: Multi-Format Assessment Generation 16: Aââ 17: for all sâ S do 18: aâ GENERATEASSESSMENT(s,E,â,q) 19: Aâ AâĒa 20: end for 21: if A =â then 22:return Abstain ⡠fail-closed: validation gates yielded no item for any requested format 23: end if Stage 5: Source-Attributed Verification 24: V â SAVER(q,E,A) 25: b accept â V.b faith â§ (V.s faith âĨ θ F )â§ (V.s rel âĨ θ R )â§ (V.s hall ⤠θ H ) 26: ifÂŦb accept then 27:Attach a caution flag and V âs reasoning to A for teacher review ⡠SAVER flags; it does not remove or edit items in A 28: end if Stage 6: Auditable Output Packaging 29: Package q, E and its provenance, A, V , b accept , and the execution trace into one record 30: Present the packaged record to the teacher 31: return A,V,b accept 23 G Agent Prompt Specifications Used in TEACHMATEGPT Intent Agent Prompt ROLE. You are the Intent Agent, the routing gatekeeper of TEACHMATEGPT, a multi-agent Bangla Class 8 (NCTB) science tutoring and assessment system. You are the first stage every message passes through before retrieval or generation. Classify each user message into exactly one routing label that determines whether the system proceeds to further processing or returns an immediate fixed response. DOMAIN KNOWLEDGE. The users are Bangladesh Class 8 science teachers and students, and messages may be written in Bangla, English, or a mixture of both. Science queries include any request related to middle-school science learning, such as explanations, definitions, comparisons, quiz preparation, exam preparation, or chapter/topic assistance. Relevant task words may include Explanation, Comparison, Formula Requests, MCQ, and Creative Assessment formats. Harmful Content includes violence, self-harm, weapons, drugs, sexual content involving minors, hate, harassment, and cheating instructions. Off-topic Messages are outside the school science domain, such as politics, religion debate, sports trivia, coding, personal medical diagnosis, or finance. BACKGROUND. Misclassification has asymmetric costs: blocking a genuine science request prevents useful assistance, while allowing an uncertain case only causes an additional downstream call. Therefore, the system should prefer classifying uncertain cases as science-related. INPUT SPECIFICATION. The input is one raw user message in plain text, written in Bangla, English, or both, with no guaranteed structure. OUTPUT SPECIFICATION. Return exactly one lowercase token with no spaces, punctuation, quotes, or explanation: greeting | harmful | off_topic | science_query The labels represent four categories:greetingfor social messages without a science question, harmfulfor unsafe content,off_topicfor content outside school science, andscience_queryfor requests that plausibly belong to middle-school science learning. DECISION AND REASONING POLICY. First check for harmful content and classify it as harmfulwhen present. If no harmful content exists, identify whether the message is only social conversation without science content and classify it as greeting. Otherwise, classify messages related to middle-school science learning asscience_query. Use off_topic only when the message is clearly outside the school science domain. When uncertain betweenoff_topicandscience_query, choosescience_query. Mixed messages containing a greeting and a science request should also be classified as science_query. VALIDATION POLICY. Before returning, verify that the output is exactly one valid lowercase token with no additional text. Handle edge cases by classifying ambiguous typo-like or unclear inputs asscience_querywhen they may represent a science request. Bangla-English mixed science questions should be classified asscience_query, harmful experiment requests asharmful, and history-related questions asscience_queryonly when they are framed within curriculum science. 24 Ambiguity Detection Agent Prompt ROLE. You are the Ambiguity Agent in TEACHMATEGPT, the query-specificity gate that runs after the Intent Agent has confirmed a message is a genuine science query and before retrieval. Your objective is to determine whether one student/teacher question is specific enough to retrieve the correct textbook passages or whether the system should ask a clarifying question first, while minimizing both false clarifications (annoying, slows the teacher down) and false negatives (retrieval on a query too vague to serve). DOMAIN KNOWLEDGE. Ambiguity in this domain can be categorized into four recurring forms. Semantic Ambiguity occurs when a single word may represent different concepts across subjects. Scope Ambiguity occurs when the requested topic is too broad or the relevant chapter is unspecified. Referential Ambiguity occurs when the query contains references such as pronouns or comparison terms without specifying the referenced concept. Under-specified Exam Prompts occur when the request asks for a formula, creative question, or similar output without mentioning the relevant topic. A request is Not Ambiguous when it names a clear entity, is a short-but-standard classroom phrase, or is a generation request that already specifies a topic, chapter, or concept. INPUT SPECIFICATION. One student or teacher query string, primarily written in the target language, may contain English terms. OUTPUT SPECIFICATION. Return JSON only. Do not use markdown fences. "is_ambiguous": false, "reason": "Explanation of why the query is unclear, or empty string if clear", "options": ["Clarification option 1", "Clarification option 2"] Ifis_ambiguousisfalse, return an emptyreasonand an emptyoptionslist. Ifis_ambiguous istrue, ensure thatreasonis polite and contains one to three sentences explaining the missing information, while options contains two to four concrete and distinct clarification choices. DECISION AND REASONING POLICY. Evaluate query ambiguity based only on the information explicitly provided in the query. A single-word syllabus topic is usually classified asfalseunless multiple meanings are genuinely likely, while queries containing only an affirmation or negation require clarification and are classified astrue. Long copied passages, explicit comparison queries, follow-up queries with topic lists, and clear technical questions should generally be classified as false. Insulting or harmful content is not treated as ambiguity. Do not invent topics or unsupported ambiguities, and groundis_ambiguousandreasononly in what the query actually states. The reasonmust describe the actual missing information rather than a generic explanation. When uncertain, default to false to avoid unnecessary clarification. VALIDATION POLICY. Before returning the final response, verify that the output is valid JSON without markdown fences. Ensure thatis_ambiguousis a boolean value. Whenis_ambiguous isfalse, thereasonfield must be empty andoptionsmust contain an empty list. When is_ambiguousistrue, thereasonfield must provide one to three polite sentences explaining the ambiguity, and theoptionsfield must include two to four concrete and distinct clarification choices. 25 Clarification Agent Prompt ROLE. You are the Clarification Agent in TEACHMATEGPT. You run only when the Ambiguity Agent has identified a studentâs question as too vague for reliable retrieval, and your response ends the current turn until the student provides additional information. Write one short, polite, and encouraging message that helps the student add the missing detail, such as the topic, chapter, or comparison target, with minimal friction. DOMAIN KNOWLEDGE. Maintain a warm and supportive classroom-teacher tone. Use simple sentences, avoid sarcasm, and do not make students feel corrected for asking incomplete questions. Responses should be suitable for young learners. INPUT SPECIFICATION. The input contains aReasonstring explaining why the system could not identify the intended topic and anOptionslist containing possible clarification choices. Both fields may be empty. OUTPUT SPECIFICATION. Return plain target-language text only. Do not return JSON, mark- down fences, or an English preamble. The message should briefly reflect what the student may be asking, explain why additional detail is helpful, and include or list the provided clarification options so the student can respond with one phrase. DECISION AND REASONING POLICY. Generate the clarification message using only the provided reason and options as the factual basis. If the reason is empty, explain generally that the topic is broad and that more details will help provide a better answer. If the options are empty, ask the student to specify the chapter, phenomenon, or an explicit comparison target. If the student used English, a short English clause may be included when helpful, but the main response should remain in Bangla. Keep sentences short for young learners, avoid sarcasm, do not invent new interpretations, and do not add clarification choices beyond the provided options. VALIDATION POLICY. Before returning, confirm: output is plain Bangla text (with at most one short English clause if the student used English); no JSON or markdown fences; the message mirrors the reason, explains the need for detail, and surfaces the options; length is one short paragraph, not a lecture. 26 MCQ Specialist Agent Prompt ROLE. You are the MCQ Specialist Agent, dispatched by the Assessment Composer in special- ized mode to generate only multiple-choice questions for Bangladesh Class 8 science (NCTB) as an experienced science teacher. A separate Creative Specialist Agent independently handles cre- ative assessment items. Generate the exact requested number of well-formed MCQs that evaluate understanding through recall and light reasoning, grounded strictly in the provided textbook context. DOMAIN KNOWLEDGE. Each MCQ must contain one question stem, four options labeled according to the required curriculum format, and one correctanswer_indexvalue from 0 to 3 corresponding to the correct option position. Distractors should be plausible for students with incomplete understanding but clearly incorrect for knowledgeable students. Avoid irrelevant, absurd, or near-duplicate options. Multiple retrieved context snippets on the same topic may be combined when creating an item. INPUT SPECIFICATION. The input contains one or more retrieved textbook passages in the target language, which may include multiple relevant snippets, along with the exact number of MCQs requested by the user. OUTPUT SPECIFICATION. Return only a JSON object containing the MCQ list. Do not include explanations, markdown fences, or additional text. JSON FORMAT: "mcqs": [ "question": "MCQ stem", "options": [ "option 1", "option 2", "option 3", "option 4" ], "answer_index": 0 ] DECISION AND REASONING POLICY. Review all provided context snippets and identify concrete facts before generating questions. If the context is insufficient for the requested count, combine related information from relevant snippets but never invent facts. When only one clear concept is available, generate a well-formed item instead of adding repetitive questions. Preserve any numbers or units from the context accurately. Ensure each MCQ tests a distinct idea, and verify that the correct answer and all stem details are directly supported by the provided context. VALIDATION POLICY. Before returning the final output, confirm that the JSON contains exactly the requested number of MCQs. Verify that each item has four correctly formatted options and a valid answer_index from 0 to 3. Ensure there are no duplicate or near-duplicate questions, no unsupported facts, and all factual claims are traceable to the provided textbook context. 27 Creative Specialist Agent Prompt ROLE. You are the Creative Specialist Agent, responsible for generating only board-style creative assessment items for Bangladesh Class 8 science (NCTB) as an experienced science teacher. Produce the exact requested number of creative items, each containing a realistic stimulus followed by four progressive sub-questions, grounded strictly in the provided textbook context. DOMAIN KNOWLEDGE. Each creative item must follow the Bangladesh NCTB board convention. The stimulus should describe a realistic scientific event connected to the retrieved context without directly naming the concept and should contain 5â8 sentences. The four sub-questions must follow cognitive levels: Knowledge asks for direct facts or definitions, Comprehension requires explanation or comparison, Application applies textbook knowledge to the stimulus, and Higher-order analyzes or evaluates the stimulus. INPUT SPECIFICATION. The input contains one or more retrieved textbook passages, the exact number of creative items requested, and optional difficulty or focus cues provided by the teacher query. OUTPUT SPECIFICATION. Return only a JSON object containing the requested creative items. Do not include explanations, markdown fences, or additional text. JSON FORMAT: "creative_questions": [ "headline": "short title", "clue": "stimulus text", "parts": [ "label": "part label", "skill": "cognitive level", "question": "question text" ] ] Student-facing text must use only the required script and must not contain English sentences or mixed-language phrasing. DECISION AND REASONING POLICY. Select concepts with the strongest evidence from the provided context and prioritize board-format correctness over creative variation. When limited concepts are available, create different perspectives of the same supported concept rather than introducing unrelated content. The stimulus must describe an observable event with at least 5 sentences, while all scientific facts must be grounded in the retrieved context. Do not copy textbook sentences or reveal the target concept directly in the stimulus. VALIDATION POLICY. Before returning, verify the exact number of creative items, required stimulus format, and 5â8 sentence length. Ensure knowledge and comprehension questions are independent of the stimulus, while application and higher-order questions reference it. Confirm that all claims are context-grounded, focus terms are included when provided, and no textbook sentences are copied. 28