Paper deep dive
When AI Meets Early Childhood Education: Large Language Models as Assessment Teammates in Chinese Preschools
Xingming Li, Runke Huang, Yanan Bao, Yuye Jin, Yuru Jiao, Qingyong Hu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/26/2026, 2:17:36 AM
Summary
The paper introduces TEPE-TCI-370h, a large-scale dataset of 370 hours of naturalistic teacher-child interactions in Chinese preschools, and Interaction2Eval, an LLM-based framework for automated quality assessment using ECQRS-EC and SSTEW rubrics. The system achieves 88% agreement with human experts and demonstrates an 18x efficiency gain in assessment workflows.
Entities (5)
Relation Signals (3)
Interaction2Eval â implements â ECQRS-EC
confidence 95% · The Evaluation Agent automates the assessment of teacher-child interactions by detecting behavioral indicators defined in the SSTEW and ECQRS-EC rubrics.
Interaction2Eval â implements â SSTEW
confidence 95% · The Evaluation Agent automates the assessment of teacher-child interactions by detecting behavioral indicators defined in the SSTEW and ECQRS-EC rubrics.
Interaction2Eval â utilizes â TEPE-TCI-370h
confidence 95% · Building on this dataset, we develop Interaction2Eval
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:High-quality teacher-child interaction (TCI) is fundamental to early childhood development, yet traditional expert-based assessment faces a critical scalability challenge. In large systems like China's-serving 36 million children across 250,000+ kindergartens-the cost and time requirements of manual observation make continuous quality monitoring infeasible, relegating assessment to infrequent episodic audits that limit timely intervention and improvement tracking. In this paper, we investigate whether AI can serve as a scalable assessment teammate by extracting structured quality indicators and validating their alignment with human expert judgments. Our contributions include: (1) TEPE-TCI-370h (Tracing Effective Preschool Education), the first large-scale dataset of naturalistic teacher-child interactions in Chinese preschools (370 hours, 105 classrooms) with standardized ECQRS-EC and SSTEW annotations; (2) We develop Interaction2Eval, a specialized LLM-based framework addressing domain-specific challenges-child speech recognition, Mandarin homophone disambiguation, and rubric-based reasoning-achieving up to 88% agreement; (3) Deployment validation across 43 classrooms demonstrating an 18x efficiency gain in the assessment workflow, highlighting its potential for shifting from annual expert audits to monthly AI-assisted monitoring with targeted human oversight. This work not only demonstrates the technical feasibility of scalable, AI-augmented quality assessment but also lays the foundation for a new paradigm in early childhood education-one where continuous, inclusive, AI-assisted evaluation becomes the engine of systemic improvement and equitable growth.
Tags
Links
- Source: https://arxiv.org/abs/2603.24389v1
- Canonical: https://arxiv.org/abs/2603.24389v1
Trouble viewing inline? Open PDF directly â
Full Text
41,217 characters extracted from source content.
Expand or collapse full text
When AI Meets Early Childhood Education: Large Language Models as Assessment Teammates in Chinese Preschools Xingming Li 1 , Runke Huang 2B , Yanan Bao 2 , Yuye Jin 2 , Yuru Jiao 2 , and Qingyong Hu 3 1 National University of Defense Technology, Changsha, China lixingming@nudt.edu.cn 2 The Chinese University of Hong Kong, Shenzhen, China runkehuang@cuhk.edu.cn, yananbao,yuyejin@link.cuhk.edu.cn, yurujiao03@gmail.com 3 University of Oxford, Oxford, United Kingdom huqingyong15@outlook.com Abstract. High-quality teacher-child interaction (TCI) is fundamental to early childhood development, yet traditional expert-based assessment faces a critical scalability challenge. In large systems like Chinaâsâserving 36 million children across 250,000+ kindergartensâthe cost and time requirements of manual observation make continuous quality monitor- ing infeasible, relegating assessment to infrequent episodic audits that limit timely intervention and improvement tracking. In this paper, we investigate whether AI can serve as a scalable assessment teammate by extracting structured quality indicators and validating their alignment with human expert judgments. Our contributions include: (1) TEPE- TCI-370h (Tracing Effective Preschool Education), the first large-scale dataset of naturalistic teacher-child interactions in Chinese preschools (370 hours, 105 classrooms) with standardized ECQRS-EC and SSTEW annotations; (2) We develop Interaction2Eval, a specialized LLM-based framework addressing domain-specific challengesâchild speech recogni- tion, Mandarin homophone disambiguation, and rubric-based reason- ingâachieving up to 88% agreement; (3) Deployment validation across 43 classrooms demonstrating an 18Ă efficiency gain in the assessment workflow, highlighting its potential for shifting from annual expert au- dits to monthly AI-assisted monitoring with targeted human oversight. This work not only demonstrates the technical feasibility of scalable, AI- augmented quality assessment but also lays the foundation for a new paradigm in early childhood educationâone where continuous, inclu- sive, AI-assisted evaluation becomes the engine of systemic improve- ment and equitable growth. Our project page is available at https: //qingyonghu.github.io/Interaction2Eval/. Keywords: Teacher-child interaction assessment, Educational speech processing, Large language models, Scalable evaluation arXiv:2603.24389v1 [cs.CL] 25 Mar 2026 2X. Li et al. 1 Introduction High-quality teacher-child interaction (TCI) is the cornerstone of effective early childhood education (ECE), directly influencing childrenâs cognitive, linguistic, and socio-emotional development [4, 13, 33]. Grounded in Vygotskyâs sociocul- tural theory (1978), learning is understood as a socially mediated process in which language and interaction function as fundamental tools for thinking and development [36]. Especially for preschoolers who are in the early stages of de- veloping abstract thinking and independent learning abilities, interactions with adults and peers remain a primary driver of learning and development [23]. Cur- rent quality evaluation approach of TCI rely on trained observers using standard- ized instruments like the Childhood Quality Rating Scale-Emergent Curriculum (ECQRS-EC) [34] and Sustained Shared Thinking and Emotional Wellbeing (SSTEW) [30]. However, this approach requires hours of in-person observation and scoring per classroom and is fundamentally unscalable and low efficiencyâa major challenge for large systems like Chinaâs, serving nearly 36 million children across over 250,000 kindergartens [21]. Recent advances in automatic speech recognition (ASR) and large language models (LLMs) suggest a promising alternative: can we automatically assess teacher-child interaction quality directly from classroom audio? While this concept seems intuitive, it presents unprecedented technical challenges. Un- like adult speech processing in controlled environments, preschool classrooms feature overlapping child voices, high background noise, non-standard pronunci- ation, and domain-specific educational terminology [10]. Moreover, meaningful quality assessment requires more than accurate transcription: it also depends on identifying subtle pedagogical patterns and mapping them to standardized rubric indicators. Answering this question therefore requires addressing three in- terrelated issues: whether large-scale naturalistic classroom data can be collected and annotated with sufficient reliability, whether LLM-based systems can achieve reliable agreement with trained human experts on language-accessible rubric in- dicators, and whether such a system can be deployed in real preschool settings with clear operational value. To bridge this gap, we introduce TEPE-TCI, the first comprehensive TCI dataset in Chinese preschools, enabling automated teacher-child interac- tion assessment from classroom audio. This represents a new challenge at the intersection of speech processing, natural language processing, and ed- ucational measurement. Unlike prior work focusing on structured settings like junior or high school dialogues with clear speech and minimal noise [16,19], our task addresses the complexities of naturalistic preschool interactions, which are more spontaneous, noisy, and linguistically diverse. Additionally, our approach emphasizes language-mediated interaction aspects, enabling scalable, rubric- grounded assessment using real-world classroom recordings. Building on this dataset, we develop Interaction2Eval, an automated assessment framework that combines large-scale naturalistic data collection, specialized speech and language processing, and real-world deployment validation. This work demonstrates both When AI Meets Early Childhood Education3 Fig. 1: Overview of the TEPE-TCI dataset, comprising over 370 hours of teacherâchild interaction audio from 41 public preschools (105 classrooms). Data were ethically approved, expert-annotated, and analyzed across levels. the technical feasibility and practical scalability of AI-augmented quality assess- ment in early childhood education. We make three primary contributions: â We present TEPE-TCI-370h, the first comprehensive dataset of naturalistic classroom interactions with expert quality annotations in Chinese preschool contexts, comprising 370 hours of audio from 105 classrooms. â We develop Interaction2Eval, a specialized LLM-based framework address- ing domain-specific challenges including child speech recognition, Mandarin homophone disambiguation, and rubric-based reasoning, achieving up to 88% agreement in interaction quality assessment. â We validate our approach through real-world deployment across 43 classrooms, demonstrating an 18Ă efficiency gain in the assessment workflow and the po- tential for shifting from annual expert audits to continuous AI-assisted mon- itoring. 2 Related Works Educational Speech and Dialogue Datasets. While general speech datasets are abundant, educational speech data for early childhood contexts remains scarce, as summarized in Table 1. Most existing childrenâs speech corpora were developed for speech recognition rather than educational assessment, including CSLU Kids [29], CMU Kids [9], PF-STAR [2], and CHILDES [27]. Datasets targeting classroom interactions predominantly focus on K-12 settings. Talk- Moves [32] provides 567 mathematics lesson transcripts with discursive move 4X. Li et al. Table 1: Summary of childrenâs speech and educational interaction datasets. K denotes kindergarten, G denotes grade. Diar. indicates speaker diarization and Quality FW indicates quality assessment framework. CorpusLanguage Age Range # Speakers HoursStyle Diar.Quality FWYear CHIEDE [12]Spanish3-659âŒ8 Conversation Partial-2008 CSLU Kidâs Speech Corpus [29] EnglishK-G101,100-Read+Spont. N-2007 CMU Kids Corpus [9]English6-1176-Read speech N-1997 PF-STAR Childrenâs Speech [2] English4-1415814.5 Read Speech N-2005 MyST Corpus [26]EnglishG3-G51,371393 Conversation Y-2024 TalkMoves [32]EnglishK-12--ClassroomY-2022 NCTE [8]EnglishG4-G53171,660 les. ClassroomYCLASS+MQI2023 SimClass [1]English--391SimulatedN-2025 WSW [31]English3-5171,592 ClassroomY-2025 Playlogue [15]English Preschool-33Play-based YDPICS2024 SingaKids [5]Chinese7-1225575ReadingN-2016 SLT-CSRC C1 [41]Chinese7-1192728.6ReadingN-2021 SLT-CSRC C2 [41]Chinese4-115429.5 Conversation N-2021 ChildMandarin [42]Chinese3-539741.3 Conversation N-2024 TEPE-TCI (Ours)Chinese3-42,550370 Classroom Y ECQRS-EC+SSTEW 2025 annotations, NCTE [8] contains elementary classroom recordings with CLASS observation scores, MyST [26] offers 393 hours of conversational speech from grades 3-5 students in tutoring contexts, and SimClass [1] synthesizes classroom audio with ambient noise. For early childhood contexts, resources are notably limited: WSW [31] captures preschool speech via wearables, Playlogue [15] of- fers 33 hours of adult-child dialogue for naturalistic play, and NCRECE PreK provides pre-kindergarten videos primarily for U.S. research [24, 25]. Notably, these datasets predominantly focus on English-speaking environments and lack standardized quality assessment annotations suitable for teacher-child interac- tion evaluation. Publicly available childrenâs speech corpora for Chinese are even more limited, particularly for preschool educational contexts. The few existing datasets, including ChildMandarin [42], SingaKids [5], and SLT-CSRC [41], were designed for speech recognition tasks and either lack classroom recordings or contain no quality assessment annotations, which restricts their utility for au- tomated interaction assessment. This gap is particularly significant given the unique challenges in Chinese preschool speech processing, including pervasive Mandarin homophone ambiguities, domain-specific educational terminology, and culturally-shaped pedagogical practices. LLMs for Educational Assessment. Large language models have shown substantial progress in educational assessment, with GPT-4 achieving human- comparable grading accuracy [14, 20] and high inter-coder agreement in class- room dialogue analysis [16,18]. For early childhood education, Wang et al. [37] explored LLM-based instructional support evaluation using CLASS protocols [35], while Whitehill et al. [38] developed automated the CLASS scoring with utterance-level feedback. However, these approaches focus on a single domain of CLASS framework (e.g., instructional support), rather than capturing the full range of interaction quality. Recent work has explored LLMs for developmental assessment [40], yet comprehensive teacher-child interaction assessment using multi-dimensional professional scales remains unexplored, particularly in non- English contexts where cultural and linguistic factors create distinct challenges. When AI Meets Early Childhood Education5 To our knowledge, no existing system addresses automated ECQRS-EC or SSTEW assessment, nor audio-based teacher-child interaction evaluation in Chi- nese preschools that spans multiple classroom scenarios. Our work establishes both a new benchmark dataset and initial baselines for this task. 3 Problem Formulation We formalize automated teacher-child interaction assessment in preschool environments as follows: Given a noisy, multi-speaker classroom audio record- ing A of duration T, the goal is to detect the presence or absence of behavioral indicators defined in standardized educational rubrics such as ECQRS-EC and SSTEW [30,34]. These rubrics are widely used in ECE to assess the TCI quality based on the occurrence of developmentally supportive teaching behaviors. Each rubric item comprises multiple indicators describing concrete, observable teacher behaviors associated with varying levels of quality - from inadequate to excellent. During 3-hour classroom observation, observers assess whether such behaviors occur and assign scores accordingly. For each indicator i, the system outputs a binary judgment y i â 0, 1 indicating whether the corresponding behavior was observed in the interaction. Item-level scores are subsequently derived from indicator patterns following official scoring protocols. 3.1 Technical Challenges This task introduces several challenges unique to the preschool-based speech processing and educational assessment: Data Challenges: As reviewed in Section 2, publicly available childrenâs speech corpora for Chinese preschool contexts are extremely limited. Existing datasets either lack naturalistic classroom recordings or contain no standard- ized quality assessment annotations, making it impossible to train or evaluate automated interaction assessment systems for this domain. Speech Processing Challenges: Chinese preschool classrooms present compounded acoustic and linguistic difficulties. Acoustically, recordings feature high background noise, overlapping multi-speaker speech, non-standard child pronunciation, and distant-microphone conditions. Linguistically, Mandarinâs tonal nature causes extensive homophone ambiguity (e.g., chĂ©nfĂș: [sinking/floating] vs. [submission]), while domain-specific educational terminology (e.g., jĂŹnq Ìu: [en- ter learning centers] vs. [go inside]) and teachersâ child-directed speech patterns further challenge standard ASR systems. Assessment Challenges: Moving from transcription to meaningful evalua- tion requires handling rubric complexityâeducational assessment scales contain nuanced criteria requiring deep contextual understanding. Temporal reasoning is essential as quality judgments must consider interaction patterns across extended time periods. Additionally, pedagogical knowledge is also crucial for accurate as- sessment, requiring understanding of early childhood education principles and developmental appropriateness. 6X. Li et al. 3.2 Task Scope and Limitations We focus on language-accessible aspects of teacher-child interactionâthose dimensions that can be reliably evaluated from verbal exchanges alone. While we acknowledge that high-quality interaction encompasses non-verbal elements (ges- tures, spatial arrangement, materials), our approach addresses the substantial subset of assessment criteria that depend on conversational patterns, questioning strategies, and linguistic scaffolding. 4 Dataset Construction 4.1 Data Collection Protocol We collected audio using professional recording equipment (iFLYTEK H1 Pro) to capture naturalistic teacher-child interactions across diverse classroom contexts including group activities, free play, outdoor activities, and daily routines. Scale and Scope: Forty-one preschools participated in this study, span- ning three quality tiers as defined by the local education authority: district-level (N=14), municipal-level (N=12), and provincial-level (N=15). These tiers reflect differences in overall institutional quality and operating conditions. From each preschool, two to three K1 classrooms serving 3 to 4 years old were recruited, with each classroom accommodating approximately 25 students. Eight research assistants conducted data collection over 6 weeks, resulting in 370+ hours of au- dio from 105 classrooms with average session length of approximately 3.5 hours. Ethical Consideration: This study was approved by the University Ethics Committee of The Chinese University of Hong Kong, Shenzhen (Approval No. EF20241026001). Informed consent was obtained from all teachers and from par- ents or legal guardians, who were fully briefed on the studyâs purpose, recording procedures, data handling protocols, and intended use for academic research. Parents or legal guardians retained the right to withdraw their childâs data at any time without penalty. To minimize privacy risks, speaker diarization preserves only role labels (teacher/child), and all recordings, transcripts, and annotations are securely stored with access restricted to authorized researchers. For future data sharing, we plan to release only anonymized transcripts, expert annotations, and supporting documentation for non-commercial academic research. Because classroom audio from young children may contain identifiable voice information, raw audio recordings will not be publicly released; any exceptional access would require additional ethical review and formal data-use agreements. 4.2 Expert Annotation Process Professional experts (the same 8 assessors) evaluated the quality of the teacher- child interaction. Annotation Protocol: Quality assessment followed standardized ECQRS- EC and SSTEW protocols. ECQRS-EC comprises 22 items spanning four do- mains: Literacy, Mathematics, Science & Environment, and Diversity. SSTEW includes 15 items covering areas such as building independence and emotional wellbeing, language support, and critical thinking. Each item is organized into When AI Meets Early Childhood Education7 four performance levels (1=inadequate, 3=minimal, 5=good, 7=excellent), op- erationalized through behavioral indicators at each level (e.g., Level 3: âChildren are allowed to talk among themselvesâ; Level 7: âAdults scaffold childrenâs conver- sationsâ). Our annotation adopted indicator-level binary coding: assessors judged whether each specified behavior was observed (1) or not (0). Item-level scores were subsequently derived from indicator patterns following official scoring pro- tocols: the score was assigned at the highest level for which all indicators were met; if more than 50% of the indicators for the next level were also satisfied, the midpoint between those levels was assigned. Our study prioritized indicators that are most representative of daily class- room practice and feasible for audio-based observation. As a result, our anno- tations covered 17 of the 22 ECQRS-EC items and 14 of the 15 SSTEW items, comprising a total of 112 ECQRS-EC indicators and 94 SSTEW indicators. Quality Assurance: All assessors underwent extensive training on both as- sessment scales, with inter-rater reliability validation requiring Îș > 0.80 agree- ment before independent scoring. During the validation phase, two assessors scored each classroom; after achieving reliability thresholds, single assessors com- pleted evaluations. This rigorous process ensured high-quality ground truth an- notations essential for automated system development. 4.3 Dataset Characteristics and Processing Figure 1 presents comprehensive statistics of the TEPE-TCI-370h dataset. Speech segment analysis reveals teachers contributing 73.85% of segments while students account for 26.15%. Character-level analysis shows even greater teacher dominance (83.12% vs 16.88%), reflecting typical classroom discourse patterns where teachers provide more extended explanations and instructions. Interest- ingly, speaker count analysis shows more balanced participation (45.09% teach- ers vs 54.91% students), indicating active child engagement despite shorter in- dividual contributions. This dataset fills a critical gap as the first large-scale Chinese preschool resource combining naturalistic classroom audio with stan- dardized quality annotations. 5 Interaction2Eval Framework Building on insights from dataset construction, we develop Interaction2Eval (Figure 2), a specialized framework addressing core challenges of automated assessment through three LLM-empowered agents (Transcription, Refinement, and Evaluation Agent) 5.1 Transcription and Refinement Agents The first stage transforms raw audio into assessment-ready transcripts. Initial ASR processing using Paraformer with diarization and punctuation restoration [3,11] revealed systematic domain-specific errors: professional transcribers cate- gorized errors as homophones (51.67%), extra words (20.80%), speaker identifica- tion (13.72%), punctuation/segmentation (7.75%), and omissions (6.06%). The 8X. Li et al. LLM Scores + Feedback Interaction2Eval 1 ASR Transcript Agent 2 Refinement Agent 4 3 Evaluation Agent LLM Scores + Feedback Human Expert Evaluation ECQRS-EC Teacher-Child Interaction > 18x Faster thanhuman expert evaluation Fig. 2: Overview of the Interaction2Eval pipeline. predominance of homophone errorsâsignificantly higher than general speech tasksâmotivated our Refinement Agent design. Given this challenge, the Refinement Agent leverages large language models for context-aware correction. The agent prompt incorporates educational do- main knowledge, explicitly stating that the input is preschool classroom speech likely containing homophone errors common in educational settings. To guide disambiguation, the prompt includes examples of frequent confusion pairsâsuch as jĂŹnq Ìu (enter learning centers) versus jĂŹnqĂč (go inside), and chĂ©nfĂș (sink- ing/floating) versus chĂ©nfĂș (submission)âand instructs the model to correct errors while preserving original meaning and speaker attributions. To accommo- date context length limits, processing proceeds via sliding window: the model receives transcript segments, produces corrected versions, and corrections are realigned with original timestamps and speaker labels. 5.2 Rubric-Based Evaluation Agent The Evaluation Agent automates the assessment of teacher-child interactions by detecting behavioral indicators defined in the SSTEW and ECQRS-EC rubrics. We focus specifically on language-accessible indicatorsâthose that can be reli- ably identified from verbal exchanges, such as whether teachers use open-ended questions or encourage child-initiated dialogue. To construct the agent, we collaborated with expert assessors to translate rubric specifications into structured prompts. Each prompt provides the model with indicator definitions across performance levels, accompanied by contrastive examples that illustrate both high-quality and low-quality interaction patterns. The model follows an evidence-first reasoning process: it first locates relevant utterances in the transcript, then determines whether the target indicator is present, and finally provides justification grounded in specific textual evidence. For each indicator, the agent outputs a binary judgment indicating pres- ence or absence, along with the supporting transcript segment. The agent also generates pedagogical suggestions to support teacher professional development. Through iterative prompt refinement with domain experts, we addressed com- When AI Meets Early Childhood Education9 Table 2: CER results. â / â : absolute/relative reduction. ModelRaw CER (%) After Refinement (%)â (%) â Whisper-large [28]35.123.2-11.9 (â33.4%) FunASR Paraformer [11]9.94.3-5.6 (â56.6%) mon failure modes including hallucinated evidence and misalignment between detected behaviors and rubric definitions. 6 Experiments and Results 6.1 Transcription Quality Evaluation To ensure high-quality transcriptions, we compared two leading ASR models, Whisper-large-v3 [28] and FunASR [11], on a 5-hour test set with 16,168 refer- ence characters. Performance was measured using Character Error Rate (CER). As shown in Table 2, raw transcription errors were significant due to homo- phones, overlapping speech, and domain-specific vocabulary. Whisper-large-v3 exhibited a high raw CER of 35.1%, likely due to suboptimal Mandarin adapta- tion, while FunASR Paraformer achieved a lower initial CER of 9.9%. Our Re- finement Agent, powered by the Qwen3-Max [39], reduced errors by addressing homophone ambiguities, lowering Paraformerâs CER to 4.3% (56.6% relative improvement) and Whisperâs to 23.2%. These results demonstrate the agentâs effectiveness, particularly for Mandarin-specific challenges. 6.2 Error Analysis and Domain Adaptation Figure 3 visualizes the most frequently misrecognized terms in raw ASR out- puts, with term size reflecting error frequency. These errors concentrate heavily in education-specific vocabulary commonly used in teacher-child interactions, including homophonic pairs like (jĂŹn q Ìu: enter learning centers) vs (jĂŹn qĂč: go inside). This analysis confirms that semantic-level language modeling is critical for accurate transcription in specialized educational settings and validates the importance of our domain-aware refinement approach. Fig. 3: Frequently misrecognized terms in raw ASR outputs (term size reflects error frequency). 10X. Li et al. 6.3 Assessment Consistency Analysis To evaluate the robustness of automated scoring across diverse LLM archi- tectures and cultural alignments, we benchmark four state-of-the-art models: two international (GPT-5 [22], Gemini-2.5-pro [7]) and two Chinese-optimized (DeepSeek-v3.1 [17], Qwen3-Max [39]). Results also reveal three key findings: (1) Chinese-adapted LLMs outperform international counterparts, particularly on both ECQRS-EC and SSTEW dimensions. DeepSeek-v3.1 achieves the highest mean agreement on both scales (87.3% for ECQRS-EC, 87.9% for SSTEW), followed by Qwen3-Max (85.7% and 86.6% respectively). This ad- vantage likely stems from their alignment with Mandarin linguistic patterns, pedagogical terminology, and local classroom discourse norms, highlighting the critical role of cultural and linguistic grounding in educational AI systems. (2) Performance on SSTEW consistently exceeds that on ECQRS- EC across almost all LLMs. This may be attributed to the distinct focuses of the two rubrics: ECQRS-EC focuses on content delivery (e.g., language, early math, science, literacy) while SSTEW indicators focus on how teachers think and talk with children. Therefore, ECQRS-EC requires a more intent-sensitive and curriculum-aligned form of inference, while SSTEW can often be scored based on observable language structure, affective cues, and scaffolding patterns â areas where language models naturally excel. (3) Performance is promising but imperfectâleaving room for fu- ture work. While top models reach 87.9% agreement (DeepSeek-v3.1 on SSTEW), Îș values of 0.71â0.74 indicate moderate-to-substantial but not yet expert-level agreement. For reference, inter-rater reliability among trained human experts in our annotation process was Îșâ 0.82â0.85, suggesting that the best-performing LLMs achieve approximately 85â90% of human expert consistency. This gap highlights the challenge of encoding complex rubric logic (e.g., "sustained shared thinking") into prompts, and underscores TEPE-TCIâs role as a benchmark for improvement, not a final solution. These findings validate the feasibility of audio-driven automated assessment, while clearly delineating its current limitsâespecially for context-heavy con- structs like ECQRS-EC. 6.4 Scalability and Efficiency Impact Our framework effectively reduces assessment time, marking an important step toward scalable deployment. Traditional manual assessment requires approxi- mately 380 minutes of assessment workflow per classroom (240 min in-person observation, 20 min indicator coding, 120 min report writing), with expert pres- ence making evaluation inherently sequential. In contrast, Interaction2Eval com- pletes the corresponding workflow in approximately 21 minutes (5 min audio processing, 12 min transcription and refinement, 4 min evaluation and report generation), achieving an 18Ă efficiency gain. This comparison reflects the time required to complete the assessment procedure for a classroom session. This ef- ficiency gain stems from two factors: (1) asynchronous recording eliminates syn- chronous observation requirementsâteachers record naturally without external When AI Meets Early Childhood Education11 Table 3: Indicator-level agreement between LLM predictions and human expert annotations for ECQRS-EC [34] and SSTEW [30] scales, measured by percentage agreement (%Agr.) and Cohenâs Kappa (Îș) which quantifies true agreement beyond chance following [6]. Models GPT-5 [22] Gemini-2.5-pro [7] DeepSeek-v3.1 [17] Qwen3-Max [39] ScaleDimensionÎș %Agr.Îș%Agr.Îș%Agr.Îș%Agr. ECQRS-EC Literacy0.678 0.844 0.7360.8710.7050.8860.7640.882 Mathematics0.631 0.836 0.6590.8450.7210.8760.6560.853 Science0.605 0.823 0.5910.8520.7040.8580.6110.836 Overall Mean0.638 0.834 0.6620.8560.7100.8730.677 0.857 SSTEW Trust & Self-regulation0.695 0.859 0.6680.8480.7820.8950.7630.903 Language & Communication 0.718 0.863 0.7520.8780.8670.9490.8250.913 Learning & Critical Think. 0.704 0.820 0.6510.8040.6480.8280.6270.815 Planning & Assessment0.679 0.837 0.6930.8340.6650.8430.6500.831 Overall Mean0.699 0.845 0.6910.8410.7410.8790.716 0.866 observers, enabling parallel assessment at near-zero marginal cost; and (2) au- tomated processing reduces manual review through high-accuracy LLM-based refinement (CER 4.3%), requiring expert validation only for flagged ambiguities than full proofreading. At system scale, this translates to substantial resource savings: assessing 100 classrooms monthly would require 633 expert-hours un- der traditional protocols (100 Ă 380 min), compared to only 35 hours with our framework (100Ă 21 min), enabling the critical shift from annual audits to con- tinuous monitoring across Chinaâs 250,000+ kindergarten system where manual assessment remains logistically infeasible. 1. Participants & Scale 3 Public Kindergartens, Shenzhen, China 77 Teachers, 3-6 age Children 8:30 AM â 11:30 AM Small Group Instruction 2. Data Collection Workflow Brief M11 Pro No Tech Expertise 3. Assessment Pipeline Web based platform Speaker Diarization ASR Transcription Rubric based evaluation 18 X Speedup 380 min 21 min Traditional Expert Interaction2Eval 4. Results & Feedback âRapid turnaround!â (Same-day results) âTransparency!â (Exact quotes) âRapid turnaround!â (Same-day results) 127 sessions, 96.8% Success Implications Interactive Continuous Assessment Monthly Quality Monitoring â Responsive PD Planning Feedback Feedback Feedback Feedback Evaluation Fig. 4: Overview of the pilot deployment study and key outcomes. 12X. Li et al. 7 From Research to Practice: Deployment Expe- rience To evaluate real-world viability, we conducted a pilot deployment of Interac- tion2Eval in operational preschool settings, examining both system performance and human-AI workflow integration (Figure 4). Participants and Scale. We partnered with 3 public kindergartens in Shen- zhen, China, covering 43 classrooms serving 3â6 year-old children. Each class- room had one head teacher, one assistant teacher and 25â35 students. Partic- ipating teachers (N=77) received brief training on recording equipment usage but required no technical expertise in AI or assessment protocols. Data Collection Workflow. Teachers wore unobtrusive recording devices (iFLYTEK H1 Pro) during regular morning sessions (8:30â11:30 AM), captur- ing naturalistic interactions across activities including circle time, free play, and small group instruction. The wireless design ensured minimal disruption to class- room routines, with teachers reporting âbarely noticingâ the device after initial adaptation (typically 2â3 days). Assessment Pipeline. Upon session completion, teachers uploaded audio files to our web-based platform via one-click transfer. The system automati- cally executed the full pipeline: (1) speaker diarization, (2) ASR transcription, (3) LLM-based refinement, and (4) rubric-based evaluation with indicator-level feedback. The corresponding assessment workflow required approximately 21 minutes for a 3-hour session, compared to 380 minutes for traditional expert observation and scoring. Results. Over 4 weeks of preliminary deployment, the system processed 127 classroom sessions with a 96.8% success rate (4 sessions required manual inter- vention due to recording quality issues). Average assessment workflow time was 21 minutes per 3-hour session, compared to 380 minutes for traditional man- ual assessmentâan 18Ă efficiency gain in the assessment workflow. This speedup stems from: (1) asynchronous recording eliminating real-time observa- tion requirements, and (2) automated transcription and scoring reducing manual coding time from hours to minutes. Qualitative Feedback. Post-deployment interviews with 12 teachers and 3 administrators revealed positive reception. Teachers highlighted rapid turnaround and transparency, noting that results were available same-day rather than weeks later and that exact quotes supporting each score made evaluations understand- able. A veteran teacher with 22 years of experience described the indicator- level feedback as a âdata mirrorâ that enabled moving from vague intuitions to evidence-based reflection on specific strengths and gaps. Administrators noted the system enabled monthly rather than annual quality monitoring, facilitating more responsive professional development planning. Limitations Observed. Users correctly identified system boundaries: ef- fective for language-accessible interaction dimensions (dialogue quality, ques- tioning strategies) but unable to assess physical environment or non-verbal en- gagementâareas requiring multimodal observation. This aligns with our design When AI Meets Early Childhood Education13 scope and suggests clear division of labor: AI-assisted continuous monitoring of conversational interactions, complemented by periodic expert evaluation of comprehensive quality including physical and visual aspects. Implications. These preliminary results validate technical feasibility for scaled deployment. The 18Ă efficiency gain in the assessment workflow is not merely quantitativeâit enables qualitative transformation from episodic exter- nal audits to integrated continuous assessment, supporting the iterative im- provement cycles emphasized in contemporary early childhood education quality frameworks. 8 Conclusion and Future Directions We introduce automated teacher-child interaction assessment from audio as a new research area, contributing TEPE-TCI-370h, the first large-scale preschool interaction dataset in China, alongside Interaction2Eval, a specialized LLM- based framework for this challenging problem. Our work demonstrates that high-quality automated assessment is achievable despite significant technical challenges, opening pathways for scalable ECE quality improvement. Key findings include the critical importance of domain-specific solutions for educational speech challenges (homophones, terminology), the feasibility of expert-level LLM assessment when properly guided by educational rubrics, and the practical scalability of audio-based assessment while maintaining pedagog- ical validity. Future research directions include multimodal integration (audio- visual), real-time formative feedback, cross-linguistic generalization, and longi- tudinal studies on educational quality improvement. Acknowledgments. This work was supported by the National Natural Science Foundation of China under Grant No. 62407037 and 62306331. References 1. Attia, A.A., Liu, J., Espy-Wilson, C.: Simclass: A classroom speech dataset gener- ated via game engine simulation for automatic speech recognition research. arXiv preprint arXiv:2506.09206 (2025) 2. Batliner, A., Blomberg, M., et al.: The pf_star childrenâs speech corpus (2005) 3. Bredin, H.: pyannote. audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe. In: Interspeech. p. 1983â1987 (2023) 4. Burchinal, M., Howes, C., Pianta, R.C., et al.: Predicting child outcomes at the end of kindergarten from the quality of pre-kindergarten teaching, instruction, activities, and caregiver sensitivity. Applied Developmental Science (2005) 5. Chen, N.F., et al.: Singakids-mandarin: Speech corpus of singaporean children speaking mandarin chinese. In: Interspeech. p. 1545â1549 (2016) 6. Cohen, J.: A coefficient of agreement for nominal scales. Educational and psycho- logical measurement (1960) 7. Comanici, G., Bieber, E., Schaekermann, M., et al.: Gemini 2.5: Pushing the fron- tier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261 (2025) 14X. Li et al. 8. Demszky, D., Hill, H.: The ncte transcripts: A dataset of elementary math class- room transcripts. In: BEA (2023) 9. Eskenazi, M., Mostow, J., et al.: The CMU kids corpus. Linguistic Data Consortium (1997), https://catalog.ldc.upenn.edu/LDC97S63, lDC97S63 10. Foulkes, P., Docherty, G.J., et al.: Sound judgements: perception of indexical fea- tures in childrenâs speech. A reader in sociophonetics (2010) 11. Gao, Z., Li, Z., Wang, J., et al.: Funasr: A fundamental end-to-end speech recog- nition toolkit. In: Interspeech (2023) 12. Garrote, M., Moreno Sandoval, A.: Chiede, a spontaneous child language corpus of spanish. In: LABLITA Workshop (2008) 13. Huang, R., Zheng, H., Siraj, I.: Quality in chinese preschool classrooms: Its struc- tural influencing factors and associations with child development. Early Education and Development (2025) 14. Impey, C., Wenger, M., Garuda, N., et al.: Using large language models for auto- mated grading of student writing about science. International Journal of Artificial Intelligence in Education (2025) 15. Kalanadhabhatta, M., et al.: Playlogue: Dataset and benchmarks for analyzing adult-child conversations during play. IMWUT (2024) 16. Li, X., Han, G., Fang, B., He, J.: Advancing the in-class dialogic quality: Developing an artificial intelligence-supported framework for classroom dialogue analysis. The Asia-Pacific Education Researcher (2025) 17. Liu, A., Feng, B., Xue, B., et al.: Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437 (2024) 18. Long, Y., Luo, H., Zhang, Y.: Evaluating large language models in analysing class- room dialogue. npj Science of Learning (2024) 19. Long, Y., Zhang, Y.: Enhanced classroom dialogue sequences analysis with a hybrid ai agent: Merging expert rule-base with large language models. arXiv preprint arXiv:2411.08418 (2024) 20. Mendonça, P.C., Quintal, F., Mendonça, F.: Evaluating llms for automated scoring in formative assessments. Applied Sciences (2025) 21. Ministry of Education of the Peopleâs Republic of China: Statistical bulletin on the development of national education in 2024. http://w.moe.gov.cn/jyb_sjz l/sjzl_fztjgb/202506/t20250611_1193760.html (2025), accessed: 2025-06-11 22. OpenAI: Gpt-5 system card (2025), https://openai.com/index/gpt-5-syste m-card/, accessed: 2025-08-07 23. Piaget, J., Cook, M., et al.: The origins of intelligence in children, vol. 8. Interna- tional universities press New York (1952) 24. Pianta, R., Hamre, B., Downer, J., et al.: Early childhood professional development: Coaching and coursework effects on indicators of childrenâs school readiness. Early Education and Development (2017) 25. Pianta, R.: National center for research on early childhood education teacher pro- fessional development study (2007-2011) (2016) 26. Pradhan, S., Cole, R., Ward, W.: My science tutor (myst)âa large corpus of chil- drenâs conversational speech. In: LREC-COLING. p. 12040â12045 (2024) 27. Pye, C.: The childes project: Tools for analyzing talk (1994) 28. Radford, A., Kim, J.W., Xu, T., et al.: Robust speech recognition via large-scale weak supervision. In: ICML (2023) 29. Shobaki, K., Hosom, J.P., Cole, R.A.: CSLU: Kidsâ speech version 1.1 (2007). https://doi.org/10.35111/q5tn-8096, https://catalog.ldc.upenn.edu/LDC2 007S18, lDC2007S18 When AI Meets Early Childhood Education15 30. Siraj, I., et al.: The Sustained Shared thinking and Emotional Well-being (SSTEW) Scale: Supporting process quality in early childhood (2023) 31. Sun, A., Feng, T., et al.: Who said what wsw 2.0? enhanced automated analysis of preschool classroom speech. arXiv preprint arXiv:2505.09972 (2025) 32. Suresh, A., Jacobs, J., Harty, C., et al.: The talkmoves dataset: K-12 mathemat- ics lesson transcripts annotated for teacher and student discursive moves. arXiv preprint arXiv:2204.09652 (2022) 33. Sylva, K., et al.: The effective provision of pre-school education (EPPE) project: Final Report: A longitudinal study funded by the DfES 1997-2004 (2004) 34. Sylva, K., Siraj, I., Taggart, B., Kingston, D.: Early Childhood Quality Rating ScaleâEmergent Curriculum (ECQRS-EC). Teachers College Press (2025) 35. Teachstone: The classroom assessment scoring system Âź (class). https://teachs tone.com/class/, accessed: 2025-09-18 36. Vygotsky, L.: Mind in society: The development of higher psychological processes. harvard university press, cambridge, ma (1978) 37. Wang, J., Hankour, K., Zhang, Y., et al.: Classroom observation: Evaluating in- structional support automatically in classroom for young children. PRML (2025) 38. Whitehill, J., LoCasale-Crouch, J.: Automated evaluation of classroom instruc- tional support with llms and bows: Connecting global predictions to specific feed- back. arXiv preprint arXiv:2310.01132 (2023) 39. Yang, A., Li, A., Yang, B., et al.: Qwen3 technical report. arXiv preprint arXiv:2505.09388 (2025) 40. Yang, Y., Shen, Y., Sun, T., Xie, Y.: Validating the effectiveness of a large language model-based approach for identifying childrenâs development across various free play settings in kindergarten. arXiv preprint arXiv:2505.03369 (2025) 41. Yu, F., Yao, Z., Wang, X., et al.: The slt 2021 children speech recognition challenge: Open datasets, rules and baselines. In: SLT. p. 1117â1123. IEEE (2021) 42. Zhou, J., et al.: Childmandarin: A comprehensive mandarin speech dataset for young children aged 3-5. In: ACL. p. 12524â12537 (2025)