Paper deep dive
Incorporating Cognitive Load and Knowledge Transfer for Multi-Domain Knowledge Tracing
Haotian Zhang, Shucun Wang, Jinze Wu, Liang Ding, Shuochen Liu, Zhenya Huang, Jing Sha, Shijin Wang, Qi Liu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/26/2026, 5:19:30 AM
Summary
The paper proposes LT-MKT, a novel method for Multi-Domain Knowledge Tracing that explicitly models cognitive load and knowledge transfer. It utilizes Large Language Models (LLMs) to construct a Multi-domain Hierarchical Graph linking questions and concepts across domains. The model captures temporal and knowledge-dimension cognitive loads and models intra- and inter-domain knowledge state propagation to improve student performance prediction.
Entities (7)
Relation Signals (9)
LT-MKT → addresses → Knowledge Tracing
confidence 95% · In this paper, we focus on exploring these factors to improve students’ knowledge state assessment in multi-domain learning scenarios and propose a novel method... (LT-MKT).
LT-MKT → models → Cognitive Load
confidence 95% · LT-MKT explicitly captures cross-domain interactions from both cognitive and knowledge propagation perspectives... models cross-domain dependencies... to characterize students’ cognitive load
LT-MKT → models → Knowledge Transfer
confidence 95% · LT-MKT incorporates a knowledge transfer module to capture the propagation of knowledge states both within and across domains.
LT-MKT → constructs → Multi-domain Hierarchical Graph
confidence 90% · LT-MKT first integrates textual information from questions and their associated concepts to construct a Multi-domain Hierarchical Graph
Cognitive Load → manifestsin → Temporal Dimension
confidence 90% · Cognitive load mainly manifests in two forms: Temporal dimension: Students frequently switch between different domains over time.
Cognitive Load → manifestsin → Knowledge Dimension
confidence 90% · Knowledge dimension: A single question may simultaneously require knowledge from multiple domains.
Knowledge Transfer → occursas → Intra-domain transfer
confidence 90% · Intra-domain transfer: Knowledge states propagate among related concepts within the same domain through prerequisite relationships
Knowledge Transfer → occursas → Inter-domain transfer
confidence 90% · Inter-domain transfer: Knowledge states transfer across different but related domains through semantic or functional correlations
LT-MKT → uses → Large Language Models
confidence 90% · LT-MKT first integrates textual information... to construct a Multi-domain Hierarchical Graph, leveraging the advanced representational capabilities of large language models (LLMs).
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Knowledge Tracing (KT) aims to assess students' dynamic knowledge states from their learning histories. While most existing KT methods focus on single-domain learning with notable success, real-world learning scenarios often involve multiple domains simultaneously, introducing two critical factors: 1) Cognitive load, arising from managing learning across domains in both temporal and knowledge dimensions. 2) Knowledge transfer, where knowledge states in one domain influence related states both within and across domains. In this paper, we focus on exploring these factors to improve students' knowledge state assessment in multi-domain learning scenarios and propose a novel method incorporating cognitive Load and knowledge Transfer for Multi-domain Knowledge Tracing (LT-MKT). Specifically, to bridge isolated domains, LT-MKT first integrates textual information from questions and their associated concepts to construct a Multi-domain Hierarchical Graph, leveraging the advanced representational capabilities of large language models (LLMs). Then, cross-domain features in both the temporal and knowledge dimensions are explicitly modeled to capture the effects of cognitive load. Additionally, a knowledge transfer module is designed to model the propagation of knowledge states within and across domains. By jointly modeling these factors, LT-MKT enables more accurate prediction of students' future performance. Finally, extensive experiments on real-world datasets demonstrate that our method achieves state-of-the-art performance.
Tags
Links
- Source: https://arxiv.org/abs/2608.24005v1
- Canonical: https://arxiv.org/abs/2608.24005v1
Trouble viewing inline? Open PDF directly →
Full Text
72,652 characters extracted from source content.
Expand or collapse full text
Incorporating Cognitive Load and Knowledge Transfer for Multi-Domain Knowledge Tracing Conference: Proceedings of the 35th ACM International Conference on Information and Knowledge Management; November 7–11, 2026; Rome, ItalyProceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM ’26), November 7–11, 2026, Rome, ItalyISBN: 979-8-4007-2539-5/2026/11DOI: 10.1145/3799682.3841120CCS: Applied computing E-learningCCS: Social and professional topics Student assessment Haotian Zhang Affiliation: State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China, Hefei, China email: sosweetzhang@mail.ustc.edu.cn , Shucun Wang Affiliation: State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China, Hefei, China email: shucunwang@mail.ustc.edu.cn , Jinze Wu Affiliation: iFLYTEK AI Research, Hefei, China email: hxwjz@mail.ustc.edu.cn , Liang Ding Affiliation: iFLYTEK AI Research, Hefei, China email: liangding3@iflytek.com , Shuochen Liu Affiliation: State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China, Hefei, China email: shuochenliu@mail.ustc.edu.cn , Zhenya Huang Affiliation: State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China & Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, Hefei, China email: huangzhy@ustc.edu.cn , Jing Sha Affiliation: iFLYTEK AI Research, Hefei, China email: jingsha@iflytek.com , Shijin Wang Affiliation: State Key Laboratory of Cognitive Intelligence & iFLYTEK AI Research, Hefei, China email: sjwang3@iflytek.com and Qi Liu Affiliation: State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China & Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, Hefei, China email: qiliuql@ustc.edu.cn © c Abstract. Knowledge Tracing (KT) aims to assess students’ dynamic knowledge states from their learning histories. While most existing KT methods focus on single-domain learning with notable success, real-world learning scenarios often involve multiple domains simultaneously, introducing two critical factors: 1) Cognitive load, arising from managing learning across domains in both temporal and knowledge dimensions. 2) Knowledge transfer, where knowledge states in one domain influence related states both within and across domains. In this paper, we focus on exploring these factors to improve students’ knowledge state assessment in multi-domain learning scenarios and propose a novel method incorporating cognitive Load and knowledge Transfer for Multi-domain Knowledge Tracing (LT-MKT). Specifically, to bridge isolated domains, LT-MKT first integrates textual information from questions and their associated concepts to construct a Multi-domain Hierarchical Graph, leveraging the advanced representational capabilities of large language models (LLMs). Then, cross-domain features in both the temporal and knowledge dimensions are explicitly modeled to capture the effects of cognitive load. Additionally, a knowledge transfer module is designed to model the propagation of knowledge states within and across domains. By jointly modeling these factors, LT-MKT enables more accurate prediction of students’ future performance. Finally, extensive experiments on real-world datasets demonstrate that our method achieves state-of-the-art performance. Code is available at https://github.com/sosweetzhang/LT-MKT. Keywords: Knowledge Tracing, Cognitive Load, Knowledge Transfer †c-license: by 1. Introduction Figure 1. Illustration of Multi-domain Knowledge Tracing. (a) Hierarchical Question Concept Graph outlines relationships between concepts, as well as between questions and concepts. Questions in one domain (e.g., q3q3) may link to concepts from another (e.g., c2c2), and concepts across domains can also correlate (e.g., c2c2 and c3c3). (b) In multi-domain learning scenarios, students may alternate between domains like math, physics, and chemistry, enhancing learning through knowledge transfer but also increasing cognitive load. Online learning has expanded rapidly, offering unmatched flexibility for both students and educators, enabling learning to occur anywhere, anytime (Abdelrahman et al., 2023). Platforms such as Coursera (VanLEHN, 2011) and ASSISTments (Anderson et al., 2014) have demonstrated effectiveness by delivering intelligent, adaptive educational services (Song et al., 2022a). These systems collect extensive data on student interactions, including exercise responses, enabling analysis of knowledge levels, learning preferences, and other key attributes. A central component of this analysis is Knowledge Tracing (KT), which aims to model and track students’ evolving knowledge states through their interactions with the system (Abdelrahman et al., 2023; Zhang et al., 2024; Zhang et al., 2026). In the literature, extensive work has advanced knowledge tracing (Shen et al., 2024). Early methods relied on hidden markov models (Corbett and Anderson, 1994; Käser et al., 2017) and logistic functions with various factors (Pavlik et al., 2009; Cen et al., 2006; Vie and Kashima, 2019) to estimate students’ knowledge states. The introduction of deep learning marked a significant shift, with Deep Knowledge Tracing (DKT) pioneering the application of neural networks to KT (Piech et al., 2015). Since then, numerous variants have emerged, including memory-enhanced models (Zhang et al., 2017; Abdelrahman and Wang, 2019), graph-based approaches (Nakagawa et al., 2019; Ni et al., 2023), and attention-driven models (Ghosh et al., 2020; Shin et al., 2021), etc. More recently, with the rapid development of large language models (LLMs), LLM-based approaches have been explored, leveraging their strong representation and reasoning capabilities to enhance knowledge state modeling, improve interpretability, and support knowledge-aware reasoning tasks (Fu et al., 2024; Han et al., 2025). Despite their success, most existing methods primarily focus on single-domain learning (e.g., math). However, many real-world learning scenarios, such as standardized exams like the GRE, require students to engage with multiple domains simultaneously. As illustrated in Figure 1, multi-domain learning scenarios introduce additional complexity beyond single-domain settings, as a single question may simultaneously require knowledge from multiple domains. Such scenarios give rise to two critical factors that significantly affect students’ knowledge states: 1) Cognitive load, resulting from learning across domains in both temporal and knowledge dimensions. Cognitive load theory (Plass et al., 2010) posits that cognitive load originates from the complexity of learning materials and interactions, which implicitly influence students’ knowledge acquisition and the evolution of their knowledge states during the learning process. In multi-domain learning scenarios, cognitive load mainly manifests in two forms: Temporal dimension: Students frequently switch between different domains over time. For instance, a student may study physics at time t3t_3, switch to mathematics at t4t_4, and then move to chemistry at t5t_5, leading to additional cognitive switching costs. Knowledge dimension: A single question may simultaneously require knowledge from multiple domains. For instance, solving the question q3q_3 at t3t_3 may require both c3c_3 force analysis from the physics domain and c2c_2 trigonometric functions from the mathematics domain, increasing the complexity of knowledge processing. 2) Knowledge transfer, where knowledge states in one domain influence related knowledge states both within and across domains. Transfer of learning theory (Cormier and Hagman, 2014) emphasizes that knowledge concepts are inherently interconnected, and knowledge transfer occurs when understanding one concept influences the learning or understanding of related concepts. In multi-domain learning scenarios, knowledge transfer mainly occurs in two ways: Intra-domain transfer: Knowledge states propagate among related concepts within the same domain through prerequisite relationships, such as from addition to multiplication. Inter-domain transfer: Knowledge states transfer across different but related domains through semantic or functional correlations, such as between force analysis and trigonometric functions. Therefore, accurately modeling knowledge states in multi-domain learning scenarios requires explicitly considering both cognitive load and knowledge transfer, which are typically overlooked in conventional single-domain KT methods. In this paper, we focus on exploring the above factors to improve students’ knowledge state assessment in multi-domain learning scenarios, and propose a novel method incorporating cognitive Load and knowledge Transfer for Multi-domain Knowledge Tracing (LT-MKT). Unlike conventional single-domain KT methods that model knowledge states independently within each domain, LT-MKT explicitly captures cross-domain interactions from both cognitive and knowledge propagation perspectives. Specifically, to bridge isolated domains and establish semantic relationships across heterogeneous concepts, LT-MKT first integrates textual information from questions and their associated concepts to construct a Multi-domain Hierarchical Graph, leveraging the strong representational capabilities of large language models (LLMs). Based on this graph structure, LT-MKT further models cross-domain dependencies in both temporal and knowledge dimensions to characterize students’ cognitive load during learning processes. In the temporal dimension, the model captures students’ domain-switching behaviors over sequential interactions, while in the knowledge dimension, it models the cognitive burden introduced by questions requiring concepts from multiple domains simultaneously. Furthermore, to model the influence of related concepts on knowledge evolution, LT-MKT incorporates a knowledge transfer module to capture the propagation of knowledge states both within and across domains. The intra-domain transfer mechanism models prerequisite relationships among concepts within the same domain, whereas the inter-domain transfer mechanism captures semantic and functional correlations between concepts from different domains. By jointly modeling cognitive load and knowledge transfer, LT-MKT provides a more comprehensive representation of students’ learning processes and enables more accurate prediction of future performance. Finally, extensive experiments on real-world datasets demonstrate the effectiveness of LT-MKT, consistently achieving state-of-the-art performance compared with existing knowledge tracing methods. Our main contributions are summarized as follows: • We highlight the critical roles of cognitive load and knowledge transfer in multi-domain learning scenarios, which are largely overlooked in existing single-domain KT methods. • We propose a novel method, LT-MKT, for multi-domain knowledge tracing, which explicitly models cognitive load and knowledge transfer. In addition, LT-MKT leverages LLMs to construct a Multi-domain Hierarchical Graph for capturing semantic relationships across domains. • Extensive experiments on four real-world datasets demonstrate that LT-MKT achieves state-of-the-art performance against 11 baselines, validating the effectiveness of modeling cognitive load and knowledge transfer in multi-domain learning scenarios. 2. Related Work 2.1. Knowledge Tracing Knowledge Tracing (KT) is a fundamental task that aims to dynamically monitor and assess students’ evolving knowledge states within online learning scenarios (Shen et al., 2024; Huang et al., 2019; Liu et al., 2019b). Over the years, a significant body of research has been dedicated to solving KT tasks. Early approaches primarily relied on statistical models, such as hidden Markov models (Corbett and Anderson, 1994; Käser et al., 2017), or logistic regression-based functions that incorporated various influencing factors (Pavlik et al., 2009; Cen et al., 2006; Vie and Kashima, 2019), to estimate students’ conceptual mastery levels. With the rise of deep learning, Deep Knowledge Tracing (DKT) marked a pivotal moment by introducing neural networks into KT (Piech et al., 2015), opening the door to more sophisticated, data-driven modeling techniques. Following DKT, numerous variants of deep knowledge tracing have emerged, each introducing novel mechanisms to improve performance. These include memory-augmented methods (Zhang et al., 2017; Abdelrahman and Wang, 2019), which explicitly model students’ past interactions, graph-based approaches (Nakagawa et al., 2019; Tong et al., 2020; Ni et al., 2023; Zhang et al., 2022; Wu et al., 2024) that better capture the relational structure between concepts, and attention-based models (Ghosh et al., 2020; Pandey and Karypis, 2019; Shin et al., 2021), which focus on highlighting key interactions during learning. Additionally, other KT models have sought to enhance student performance prediction by incorporating auxiliary information or designing more specialized network architectures (Liu et al., 2019a; Wang et al., 2022; Song et al., 2022b; Shen et al., 2021; Wang et al., 2021; Lee et al., 2023; Liu et al., 2024b; Ma et al., 2024; Sun et al., 2024; Liu et al., 2024a). Despite these advances, most existing KT methods mainly focus on single-domain learning scenarios, such as mathematics, while overlooking the influence of multi-domain learning behaviors. This domain-specific assumption limits the ability of these methods to exploit cross-domain knowledge dependencies, which may provide valuable contextual information for accurately understanding students’ overall knowledge states. Recently, several studies have begun exploring the use of multi-domain information to enhance KT performance. For instance, adaptive knowledge tracing methods (Cheng et al., 2022; Tang et al., 2024; Wu et al., 2025) leverage knowledge from related domains to alleviate data sparsity and improve performance in target domains. Meanwhile, some studies have started to investigate multi-domain knowledge tracing directly. PromptKT (Liu et al., 2025) introduces a prompt-enhanced paradigm that utilizes student interaction data from multiple domains to improve KT performance jointly across domains. TransKT (Han et al., 2025) further proposes a contrastive cross-course knowledge tracing framework that leverages concept graph-guided knowledge transfer to model relationships among learning behaviors across different courses, thereby improving knowledge state estimation. However, existing methods primarily focus on transferring information across domains while largely overlooking the underlying cognitive mechanisms that arise from multi-domain learning. In particular, they fail to explicitly model how cross-domain interactions contribute to cognitive load and knowledge transfer during the learning process, which implicitly influences students’ knowledge states. Different from previous studies, our work explicitly models the interplay of knowledge states across multiple domains from both cognitive load and knowledge transfer perspectives. By doing so, our method provides a more comprehensive representation of students’ learning processes and better reflects the interconnected nature of real-world multi-domain learning scenarios. 2.2. LLMs for Education Large Language Models (LLMs) have demonstrated remarkable effectiveness across a wide range of educational tasks, consistently achieving strong performance in learning-related applications (Wang et al., 2024; Liu et al., 2026a; Wang et al., 2026a; Sha et al., 2026; Liu et al., 2026b). Recent studies show that LLMs can reach near student-level proficiency on standardized examinations in subjects such as mathematics and physics (Achiam et al., 2023), highlighting their potential in educational scenarios including tutoring, writing assistance, and reading comprehension (Malinka et al., 2023). Moreover, LLMs have shown strong capabilities in enabling personalized learning and supporting automated educational assessment (Li et al., 2024; Liu et al., 2024c; Wang et al., 2026b; Lv et al., 2025). In the domain of Knowledge Tracing (KT), recent research has increasingly explored the use of LLMs to enhance representation learning and semantic understanding from both student interactions and textual educational resources. For example, LLM-SBCL (Ni et al., 2024) utilizes LLMs to analyze the textual content of questions within student-question interaction networks, enabling more accurate identification of underlying knowledge concepts, especially in cold-start scenarios with limited data. Similarly, DCL4KT+LLM (Lee et al., 2023) leverages LLMs to estimate question difficulty from question stems and associated knowledge concepts, effectively alleviating the issue of missing difficulty annotations for unseen questions. In addition, Sonkar and Baraniuk (2023) explores the reasoning capabilities of LLMs to simulate students’ misconceptions and incorrect responses based on detailed knowledge profiles, demonstrating the potential of LLMs for modeling complex learning behaviors. More recently, several studies have begun integrating LLMs into knowledge structure modeling for KT. SINKT (Fu et al., 2024) introduces the first inductive knowledge tracing framework that directly incorporates LLMs to enhance representation learning and improve generalization to unseen questions and students. TransKT (Han et al., 2025) further employs concept graph-guided knowledge transfer to model relationships among learning behaviors across different courses, demonstrating the effectiveness of graph-based knowledge structures in capturing semantic dependencies across domains. These studies collectively demonstrate the strong capability of LLMs in extracting semantic relationships from educational content and constructing meaningful knowledge representations for KT tasks. Different from previous studies that mainly employ LLMs to model relationships among concepts, we leverage LLMs to construct a multi-domain hierarchical graph that jointly captures question-concept and concept-concept relationships across domains, providing the structural foundation for modeling cognitive load and knowledge transfer in multi-domain knowledge tracing. 3. Problem Definition In an Intelligent Tutoring System (ITS), we consider a set of students S, a set of questions Q, and a set of knowledge concepts C distributed across multiple domains D. For a student s∈s , the learning history is represented as: (1) Rs=(q1d1,r1),(q2d2,r2),…,(qTdk,rT),R_s=\(q_1^d_1,r_1),(q_2^d_2,r_2),…,(q_T^d_k,r_T)\, where qtdj∈q_t^d_j denotes the question answered by the student at time step t, which belongs to domain dj∈d_j , and rt∈0,1r_t∈\0,1\ represents the student’s response correctness, where rt=1r_t=1 indicates a correct response and rt=0r_t=0 otherwise. The goal of multi-domain knowledge tracing is to predict the probability that the student correctly answers the next question qT+1dmq_T+1^d_m from domain dm∈d_m , based on the historical interaction sequence RsR_s. Formally, the prediction objective is defined as: (2) p(rT+1=1∣Rs,qT+1dm).p(r_T+1=1 R_s,q_T+1^d_m). Different from traditional single-domain knowledge tracing, where all learning interactions are restricted to a single domain, multi-domain knowledge tracing involves more complex cross-domain learning behaviors. Specifically, students may frequently switch across different domains during learning, while individual questions may simultaneously require knowledge from multiple domains, thereby introducing additional factors that complicate knowledge state modeling and make it more challenging. 4. The LT-MKT Model Figure 2. Overall framework of LT-MKT. The model first constructs a multi-domain hierarchical graph from question texts and knowledge concepts with LLM-guided reasoning. Based on this graph, LT-MKT models cognitive load through question difficulty, domain transition, and domain coverage, and updates students’ knowledge states via GRU-based state evolution. It then performs domain-aware knowledge transfer with intra-domain prerequisite propagation and inter-domain correlation propagation, followed by state fusion and performance prediction for the next interaction. In this section, we present the proposed LT-MKT model in detail. An overview of the overall framework is illustrated in Figure 2. Specifically, we first introduce the Multi-domain Hierarchical Graph, which serves as the relational foundation for modeling semantic dependencies across domains (§ 4.1). Based on this graph structure, we then model cognitive load (§ 4.2) and knowledge transfer (§ 4.3) to refine students’ knowledge state estimation in multi-domain learning scenarios, ultimately enabling more accurate student performance prediction (§ 4.4). 4.1. Hierarchical Graph Construction In multi-domain learning scenarios, knowledge concepts are often interconnected across domains, making it difficult to capture the semantic dependencies required for modeling cognitive load and knowledge transfer. To address this issue, we construct a Multi-domain Hierarchical Graph (MDHG) to provide a unified relational structure for multi-domain knowledge tracing. Formally, let =(,,ℰ)G=(Q,C,E) denote the MDHG, where Q is the question set, C is the knowledge concept (KC) set, and ℰE contains three types of edges: question-to-concept associations, intra-domain concept prerequisites, and inter-domain concept correlations. Association edges indicate the concepts involved in each question. Prerequisite edges describe directed learning dependencies within the same domain, while correlation edges capture semantic or functional connections across domains. Constructing such educational graphs usually requires extensive expert annotation. To improve scalability, we employ LLMs to construct the MDHG through prompt-guided reasoning. Specifically, for each question, the prompt provides the question text, domain information, main concept, candidate concepts, and existing prerequisite relations. Based on chain-of-thought reasoning (Wei et al., 2022), the LLM first identifies the concepts involved in the question, then selects prerequisite concepts from the same domain and correlated concepts from other domains. For a question qjq_j with primary concept cic_i, this process is formulated as: (3) (qj),ℳ(qj),(qj)=LLMθ((qj,dj,ci,,ℰpre)),A(q_j),M(q_j),N(q_j)=LLM_θ(P(q_j,d_j,c_i,C,E_pre)), where (qj)A(q_j) denotes the associated concepts of qjq_j, ℳ(qj)M(q_j) denotes intra-domain prerequisite concepts, (qj)N(q_j) denotes inter-domain correlated concepts, and (⋅)P(·) is the graph construction prompt. After processing all questions, we aggregate the generated relations at the concept level. For each concept cic_i, let iQ_i be the set of questions whose primary concept is cic_i. The prerequisite and correlation sets of cic_i are obtained by: (4) ℳi=⋃qj∈iℳ(qj),i=⋃qj∈i(qj).M_i= _q_j _iM(q_j), _i= _q_j _iN(q_j). We then construct the final edge sets as: (5) ℰqc _qc =(qj,c)∣c∈(qj), =\(q_j,c) c (q_j)\, ℰpre _pre =(cm,ci)∣cm∈ℳi, =\(c_m,c_i) c_m _i\, ℰcor _cor =(ci,cn)∣cn∈i. =\(c_i,c_n) c_n _i\. Thus, the complete edge set is ℰ=ℰqc∪ℰpre∪ℰcorE=E_qc _pre _cor. To ensure educational consistency, prerequisite relations are constrained to satisfy the DAG property: (6) Cycle(,ℰpre)=∅,Cycle(C,E_pre)= , which avoids circular dependencies in learning paths. Since each new concept can be processed with the same prompt and aggregation procedure, the graph can also be incrementally extended when unseen concepts appear in future learning scenarios. Detailed prompts are provided in the code repository. 4.2. Cognitive Load Driven State Evolution Cognitive load theory (Plass et al., 2010) posits that cognitive load stems from the complexity of learning materials and interactions, thereby implicitly influencing students’ knowledge acquisition and the evolution of their knowledge states during the learning process. In multi-domain learning scenarios, cognitive load is further amplified by cross-domain learning behaviors in both temporal and knowledge dimensions. Therefore, accurately modeling students’ knowledge states requires explicitly capturing the dynamics of cognitive load throughout the learning process. Motivated by findings from existing studies (Lee et al., 2023; Wu et al., 2025), we investigate the effects of three key factors related to cognitive load: question difficulty, domain transition, and domain coverage. These factors characterize different aspects of cognitive complexity introduced by multi-domain learning interactions. While acknowledging the potential influence of other factors, we leave their exploration for future work. Question Difficulty Question difficulty reflects the complexity of learning materials to some extent and serves as a critical factor in assessing cognitive load. It has been proven to be a critical factor in accurately estimating students’ knowledge states (Lee et al., 2023). Following some previous works (Shen et al., 2022), we calculate the difficulty of the question qiq_i as follows: (7) qDi=∑i=1|Si|ai==0|Si|×λP,qD_i= _i=1^|S_i|a_i==0|S_i|× _P, where SiS_i denotes the set of students who have attempted question qiq_i, and as,i∈0,1a_s,i∈\0,1\ indicates whether student s answers question qiq_i correctly (1 for correct and 0 for incorrect). λP _P represents the predefined granularity level of question difficulty. Intuitively, a question is considered more difficult if it is answered incorrectly by a larger proportion of students. We further represent the difficulty embedding using an embedding matrix qD∈ℝλP×dqDE_qD _P× d_qD, where dqDd_qD denotes the embedding dimension. Domain Transition Domain transition refers to shifts between domains during learning, reflecting the complexity of interactions that influence cognitive load and subsequently affect knowledge state estimation. Intuitively, given a student’s interaction history ℛs=(q1d1,r1),(q2d2,r2),…,(qTdk,rT)R_s=\(q_1^d_1,r_1),(q_2^d_2,r_2),…,(q_T^d_k,r_T)\, the domain transition factor up to time t can be calculated as: (8) dTt=∑t′=t−wst(dt′≠dt′−1),dT_t= _t =t-ws^t(d_t ≠ d_t -1), where wsws is the size of the time window. Since a student’s learning is more strongly influenced by recent interactions, considering transitions over the entire time span may not be appropriate. In practice, a student who frequently switches between different domains within a short period will have a higher domain transition value, indicating stronger cross-domain cognitive switching effects. Then we represent the transitions embedding with embedding matrix EdT∈ℝλT×dTE_dT _T× d_dT, where λT _T represents the maximum value of domain transitions observed across all sequences, dTd_dT is the dimension. Domain Coverage Domain coverage refers to the number of domains covered by the concepts assessed in a specific question. Intuitively, the more domains a question involves, the higher its complexity and the greater the interaction required for learning. This increased complexity contributes to higher cognitive load, which in turn affects knowledge state estimation. For a question qiq_i, the domain coverage is defined as follows: (9) dCi=|Di|,dC_i=|D_i|, where i=d1,d2,…D_i=\d_1,d_2,…\ represents its relevant concepts coverage couple of domains. Similarly, the domain coverage is represented as an embedding matrix EdC∈ℝλC×dCE_dC _C× d_dC, where λC _C represents the total number of domains and dCd_dC is the dimension. State Evolution After examining the impact of cognitive load through the three key factors discussed above, we define the cognitive load cltcl_t at time t as: (10) clt=qDt⊕dTt⊕dCt.cl_t=qD_t _t _t. Subsequently, we integrate the cognitive load into the interaction embedding xtx_t as follows: (11) xt=qt⊕rt⊕clt,x_t=q_t _t _t, where qtq_t is obtained by encoding the question text using BERT (Devlin et al., 2019), and for the answer rtr_t, i.e., 0 or 1, we expand it to an all-zero or all-one vector rt∈ℝdar_t ^d_a, dad_a is the dimension. Finally, we use Gated Recurrent Unit (GRU) (Chung et al., 2014) to update the knowledge state to model the temporal effect in the learning process: (12) t _t =σ(r[t−1⊕t]+r), =σ(W_r [h_t-1 _t ]+b_r), (13) t _t =σ(z[t−1⊕t]+z), =σ(W_z [h_t-1 _t ]+b_z), (14) ~t h_t =tanh(h~[t⋅t−1⊕t]+h~), = (W_ h [r_t·h_t-1 _t ]+b_ h), (15) t _t =(1−t)⋅~t+t⋅t−1, =(1-z_t)· h_t+z_t·h_t-1, where r,z,h~∈ℝ(dk+dx)×dkW_r,W_z,W_ h ^(d_k+d_x)× d_k and r,z,h~∈ℝdkb_r,b_z,b_ h ^d_k and σ denotes the sigmoid function. t∈ℝdkh_t ^d_k is a row extracted from the knowledge state matrix ℋtH_t, which represents the knowledge state of the primary KC of the current question. 4.3. Domain Aware Knowledge Transfer Transfer of learning theory (Cormier and Hagman, 2014) emphasizes that knowledge concepts are inherently interconnected, and knowledge transfer occurs when understanding one concept influences the learning and understanding of related concepts. Different from traditional single-domain knowledge tracing, where all concepts are confined to a single domain, multi-domain learning requires modeling knowledge propagation both within and across domains. Specifically, knowledge transfer in multi-domain learning can be categorized into two types: vertical transfer and lateral transfer (Cormier and Hagman, 2014). Vertical transfer refers to prerequisite-driven knowledge propagation within the same domain, where mastery of foundational concepts is essential for learning subsequent concepts. For instance, learning addition can facilitate understanding of multiplication within mathematics. In contrast, lateral transfer captures semantic or functional associations across different domains, where knowledge from one domain may indirectly influence related knowledge states in another domain. For instance, knowledge of trigonometric functions in mathematics can support the learning of force decomposition in physics. Furthermore, learning hierarchy theory (Khan and Coomarasamy, 2006) posits that knowledge transfer within and across domains may occur at different stages of learning. In particular, intra-domain knowledge transfer is typically more immediate, as prerequisite dependencies directly affect subsequent learning within the same domain before cross-domain transfer takes effect (Macaulay and Cree, 1999). Motivated by these inspirations, we design two graph attention modules, namely intraGAT and interGAT, to successively model intra-domain and inter-domain knowledge transfer, respectively. Specifically, intraGAT captures prerequisite-based knowledge propagation within the same domain, while interGAT further models cross-domain semantic interactions among related concepts. To be specific, we first decompose the MDHG into two sub-graphs: intra-graph and inter-graph, as follows: (16) =+, G=G^intra+G^inter, where G, intraG^intra and interG^inter are relation matrices. Then we design the intraGAT layer to capture within-domain knowledge transfer in intraG^intra: (17) αijintra=exp(LeakyReLU(aT[hi⊕hj]))∑k∈ℳiexp(LeakyReLU(aT[hi⊕hk])),α^intra_ij= (LeakyReLU (a^T [Wh_i _j ] ) ) _k _i (LeakyReLU (a^T [Wh_i _k ] ) ), where ∈ℝ2dka ^2d_k is a learnable weight vector, and LeakyReLU denotes the activation function with a negative slope coefficient α=0.2α=0.2. ℳiM_i represents the set of intra-domain successor KCs associated with knowledge concept cic_i. Furthermore, to maintain the stability of students’ knowledge states and avoid unnecessary propagation noise, we only update the primary knowledge concept cic_i and its associated concepts in ℳiM_i according to the question-to-concept associations in the MDHG: (18) hi′=ELU(∑j∈ℳiαijintraWhj),h _i=ELU ( _j _iα^intra_ijWh_j ), where ELU is the exponential linear units activation function. Similarly, the interGAT layer is designed to capture knowledge transfer across domains: (19) αijinter=exp(LeakyReLU(aT[hi′⊕hj′]))∑k∈iexp(LeakyReLU(aT[hi′⊕hk′])),α^inter_ij= (LeakyReLU (a^T [Wh _i _j ] ) ) _k _i (LeakyReLU (a^T [Wh _i _k ] ) ), (20) hi′=ELU(∑j∈iαijinterWhj′),h _i=ELU ( _j _iα^inter_ijWh _j ), where iN_i is the set of inter-domain correlated KCs of KC cic_i. After knowledge transfer within and across domains, we finally get the student’s final knowledge state ℋ′tH^ _t. 4.4. Prediction and Objective Function After modeling cognitive load and knowledge transfer across domains, we obtain the student’s updated knowledge state representation ℋ′tH^ _t. In LT-MKT, the prediction of the student’s performance on the next question qt+1q_t+1 is based on three components: the question embedding t+1 q_t+1, the cognitive load representation t+1 cl_t+1, and the fused knowledge state vector ~t h_t. Specifically, the fused knowledge state is computed through a State Fusion Module as: (21) ~t h_t =F(~ti+~tj+⋯), =F( h^i_t+ h^j_t+·s), where F(⋅)F(·) denotes a mean fusion operation, which has been widely adopted as an effective strategy for aggregating multiple representations. ~ti h^i_t represents the knowledge state corresponding to concept cic_i extracted from ℋ′tH^ _t, while ~tj h^j_t denotes the knowledge state corresponding to another related concept cjc_j involved in question qt+1q_t+1. Other concept-specific knowledge states are obtained similarly. The predicted probability of correctly answering question qt+1q_t+1 is then computed as: (22) yt+1 y_t+1 =σ(p[t+1⊕~t⊕t+1]+p), =σ(W_p[ q_t+1 h_t cl_t+1]+b_p), where p∈ℝ(dq+dk+dcl)×dkW_p ^(d_q+d_k+d_cl)× d_k and p∈ℝdkb_p ^d_k are trainable parameters, ⊕ denotes the concatenation operation, and σ(⋅)σ(·) represents the sigmoid activation function. The output yt+1∈(0,1)y_t+1∈(0,1) denotes the probability that the student correctly answers question qt+1q_t+1. To optimize LT-MKT, we adopt the binary cross-entropy loss between the predicted response yty_t and the ground-truth response rtr_t as the training objective: (23) ℒ=−∑t=1T(rtlogyt+(1−rt)log(1−yt)).L=- _t=1^T (r_t y_t+(1-r_t) (1-y_t) ). Minimizing this objective encourages the model to accurately estimate students’ future performance by jointly modeling cognitive load and knowledge transfer in multi-domain learning scenarios. Table 1. Dataset statistics Datasets JuniorH SeniorH PTADiscJP PTADiscDS #Students 1,081 4,869 29,430 12,271 #Concepts 139 198 1,245 634 #Questions 170 264 27,820 18,702 #Interactions 39,230 133,683 11,172,165 1,788,245 #Avg.Inter 36.29 27.45 379.6 145.73 #Avg.Cross 14.26 12.39 8.72 7.81 Table 2. Results of all comparison methods on the student performance prediction task. Existing state-of-the-art results are marked by the underline, and the best results are bold. * indicates p-value <0.05 in the t-test. Methods JuniorH SeniorH PTADiscJP PTADiscDS AUC↑~ ACC↑~ RMSE↓~ AUC↑~ ACC↑~ RMSE↓~ AUC↑~ ACC↑~ RMSE↓~ AUC↑~ ACC↑~ RMSE↓~ DKT 0.8840 0.7991 0.1338 0.8861 0.7926 0.1377 0.7246 0.7843 0.3687 0.6504 0.7602 0.4358 GKT 0.8818 0.7983 0.1324 0.8883 0.7828 0.1391 0.7298 0.7892 0.3664 0.6534 0.7646 0.4335 AKT 0.8871 0.8032 0.1341 0.8864 0.7928 0.1382 0.7307 0.7904 0.3655 0.6543 0.7651 0.4337 HawkesKT 0.7515 0.7202 0.2273 0.7610 0.7124 0.2376 0.7441 0.7903 0.3612 0.6574 0.7641 0.4358 LPKT 0.8844 0.8004 0.1335 0.8850 0.7922 0.1383 0.7274 0.7876 0.3669 0.6575 0.7634 0.4346 DIMKT 0.8895 0.8035 0.1319 0.8865 0.7927 0.1374 0.7346 0.7948 0.3547 0.6573 0.7701 0.4328 AT-DKT 0.8855 0.8018 0.1334 0.8851 0.7898 0.1385 0.7285 0.7884 0.3664 0.6619 0.7628 0.4351 MIKT 0.8845 0.7969 0.1354 0.8941 0.7991 0.1325 0.7382 0.8006 0.3577 0.6646 0.7741 0.4316 SINKT 0.8947 0.8114 0.1298 0.8986 0.8075 0.1306 0.7473 0.8118 0.3532 0.6701 0.7842 0.4286 promptKT 0.8981 0.8142 0.1286 0.9013 0.8115 0.1294 0.7407 0.8141 0.3611 0.6682 0.7871 0.4276 TransKT 0.9143 0.8272 0.1256 0.9174 0.8241 0.1272 0.7514 0.8213 0.3517 0.6802 0.7943 0.4148 LT-MKT 0.9387* 0.8425* 0.1214* 0.9312* 0.8470* 0.1141* 0.7645* 0.8410* 0.3485* 0.6929* 0.8031* 0.4065* 5. Experiments In this section, we first introduce the datasets, followed by a description of the baseline models and training details. Subsequently, we present the results of extensive experiments. 5.1. Datasets We conduct experiments on four real-world multi-domain learning datasets. Two publicly available datasets are derived from the PTADisc dataset (Hu et al., 2023) 11 1 https://github.com/wahr0411/PTADisc, namely PTADiscJP and PTADiscDS. Specifically, PTADiscJP contains learning records from two programming-related domains: Java and Python (Java&Python), while PTADiscDS consists of records from C programming and Data Structure & Algorithm Analysis (C&DS). In addition, we conduct experiments on two proprietary datasets supplied by iFLYTEK Co., Ltd., collected from the intelligent learning machine 22 2 https://xxj.xunfei.cn/: JuniorH and SeniorH. These datasets contain students’ question-answering records across three academic domains: mathematics, physics, and English. Specifically, the JuniorH dataset includes learning records from grades 7–9, while the SeniorH dataset contains records from grades 10–12. Basic statistics of all datasets are summarized in Table 1. It is worth noting that, as shown in Table 1, we calculate the average number of Domain Transitions (#Avg.Cross) for students in real-world learning scenarios (detailed in Section 4.2, where wsws is set equal to t). The results indicate that students frequently engage in cross-domain learning in multi-domain educational settings. 5.2. Baselines To validate the effectiveness of our proposed LT-MKT in multi-domain learning scenarios, we selected eleven representative KT models as baselines. Their details are as follows: • DKT (Piech et al., 2015) utilizes RNN to model the learning sequence, where the hidden state represents the learner’s knowledge state. • GKT (Nakagawa et al., 2019) generates a transition graph from the dataset and employs GNN to encode students’ knowledge states. • AKT (Ghosh et al., 2020) uses a monotonic attention mechanism to capture dependencies in learning sequences. • Hawkes-KT (Wang et al., 2021) models temporal cross-effects from point process, where prior interactions impact skill mastery over time. • LPKT (Shen et al., 2021) models students’ progress by considering interval times to calculate learning gains and forgetting rates. • DIMKT (Shen et al., 2022) is a sequential model incorporating question/skill difficulty levels as inputs. • AT-DKT (Chen et al., 2023) enhances DKT by including two auxiliary tasks: question tagging and predicting students’ prior knowledge. • MIKT (Sun et al., 2024) models students’ knowledge states at both coarse-grained domain and fine-grained concept levels. • SINKT (Fu et al., 2024) is a structure-aware inductive model that harnesses large language models to effectively generalize to new student responses and unseen questions. • promptKT (Liu et al., 2025) introduces a prompt-enhanced paradigm utilizing a pre-trained Transformer backbone and a soft domain prompt module for multi-domain knowledge tracing. • TransKT (Han et al., 2025) leverages concept graph guided knowledge transfer to model the relationships between learning behaviors across different courses. 5.3. Experimental Setup In our experiments, we split the dataset in an 8:1:1 ratio by learners to obtain the training set, validation set, and testing set. The difficulty granularity parameter λP _P in Eq. 7 is set to 30, and the sliding window size wsws in Eq. 8 is set to 20. During graph construction, we employed Qwen-plus 33 3 https://bailian.console.aliyun.com/ (default) as the LLM backbone to balance capability and cost. The sampling temperature is set to 0 to reduce randomness. All models are optimized using the Adam optimizer. The learning rate is selected from 0.001,0.005\0.001,0.005\ based on validation performance, with a learning rate decay applied every 10 epochs. Hyperparameters are tuned on the validation set, and the model achieving the best validation performance is used for final evaluation. Following previous KT studies, we adopt three widely used evaluation metrics to comprehensively assess model performance: Area Under the ROC Curve (AUC), Accuracy (ACC), and Root Mean Squared Error (RMSE). All experiments were conducted on a cluster of Linux servers with Tesla V100 GPUs. Table 3. Results of ablation experiments across four datasets. Methods JuniorH SeniorH PTADiscJP PTADiscDS AUC↑~ ACC↑~ RMSE↓~ AUC↑~ ACC↑~ RMSE↓~ AUC↑~ ACC↑~ RMSE↓~ AUC↑~ ACC↑~ RMSE↓~ w/o CL 0.9005 0.8116 0.1322 0.8935 0.8150 0.1268 0.7471 0.8237 0.3610 0.6673 0.7896 0.4231 w/o TG 0.9227 0.8322 0.1219 0.9160 0.8360 0.1250 0.7594 0.8302 0.3544 0.6799 0.7955 0.4102 w/o SF 0.9320 0.8419 0.1227 0.9252 0.8455 0.1152 0.7603 0.8388 0.3505 0.6865 0.8011 0.4097 LT-MKT 0.9387 0.8425 0.1214 0.9312 0.8470 0.1141 0.7645 0.8410 0.3485 0.6929 0.8031 0.4065 5.4. Overall Performance We compare LT-MKT with eleven representative baselines, and the results are reported in Table 2. Several observations can be drawn from the results. First, LT-MKT consistently achieves the best performance across all datasets and evaluation metrics, demonstrating the effectiveness of explicitly modeling cognitive load and knowledge transfer in multi-domain knowledge tracing. Second, LT-MKT achieves more significant improvements on JuniorH and SeniorH, likely because the denser cross-domain learning behaviors (#Avg.Cross in Table 1) further validate the rationality and superiority of our method. Third, we have noticed that HawkesKT performs noticeably worse than other methods, especially on JuniorH and SeniorH. This may result from its sensitivity to temporal dynamics, while the relatively short interaction sequences in these datasets (Table 1) limit its effectiveness. Finally, LLM-based KT methods (e.g., SINKT, TransKT) generally outperform traditional KT models, indicating that LLMs are effective at capturing richer semantic relationships for knowledge tracing. 5.5. Ablation Study To further investigate the importance of each module in LT-MKT, we design three variations to conduct the ablation study, each of which removes one part from the original method: • LT-MKT w/o CL, which removes cognitive load effects including question difficulty, domain transition, and domain coverage features (section 4.2). • LT-MKT w/o TG, replaces the intraGAT and interGAT layers with a single GAT layer (section 4.3). • LT-MKT w/o SF, which excludes the state fusion module F in the prediction stage (section 4.4). From Table 3, several key findings can be drawn. Firstly, the complete model achieved the best overall performance. Secondly, the cognitive load effect significantly impacts model performance, emphasizing the importance of cross-domain features in multi-domain learning. Thirdly, knowledge transfer mainly affects cross-domain knowledge evolution, and its removal causes a performance drop, while state fusion has a smaller effect. The collaborative synergy of the components leads to optimal results, as the absence of any component results in a decline in performance. 5.6. Effect of the Graph Construction Since LT-MKT relies on LLMs to construct the Multi-domain Hierarchical Graph (MDHG), we perform a human evaluation to assess the quality of graphs generated by different LLM backbones from both educational and structural perspectives. Specifically, we recruit twenty experts for evaluation, including five students with relevant coursework experience, eight researchers in educational data mining, and seven teachers with practical experience in curriculum design and assessment. We compare seven representative LLM backbones with different scales and families, including Qwen2.5-7B, Qwen2.5-32B, Qwen2.5-72B, Llama3-70B, DeepSeek-R1, Claude3 Sonnet, and GPT-5 pro. Each generated graph is rated on two dimensions: (i) Educational Rationality (ER), which measures whether the prerequisite and correlation relations are pedagogically meaningful and consistent with real learning dependencies; and (i) Structural Consistency (SC), which evaluates whether the graph exhibits coherent structure without invalid or circular prerequisite relations. Both metrics are scored on a five-point Likert scale (1–5). We further report Fleiss’ κ (Fleiss, 1971) to measure inter-rater agreement. The results are shown in Table 4. Among all compared models, GPT-5 achieves the best performance across all metrics, including the highest inter-annotator agreement (Fleiss’ κ). Overall, larger and more capable models consistently achieve higher ER and SC scores, indicating stronger semantic understanding and structural reasoning ability in graph construction. These trends further highlight the compatibility and extensibility of our framework, as LT-MKT consistently benefits from stronger LLM backbones. With the release of more powerful models in the future, the quality of graph construction is expected to be further improved, thereby providing an even stronger foundation for multi-domain knowledge tracing. Table 4. Expert evaluation results of graph quality generated by different LLM backbones. LLM Backbone ER ↑ SC ↑ Fleiss’ κ ↑ Qwen2.5-7B 3.72 3.60 0.61 Qwen2.5-32B 4.07 3.95 0.69 Qwen2.5-72B 4.29 4.16 0.74 Llama3-70B 4.11 4.02 0.71 DeepSeek-R1 4.13 3.98 0.68 Claude3 Sonnet 4.45 4.38 0.79 GPT-5 pro 4.66 4.56 0.84 Figure 3. Analysis of cognitive-load representations. 5.7. Analysis of Cognitive Load Representation To further examine whether the cognitive load module learns meaningful load-aware representations, we conduct a representation-level analysis on the test interactions. For each interaction, we construct a cognitive load index (CLI) based on the three factors used in LT-MKT, namely question difficulty, domain transition, and domain coverage (Eq. 10). Specifically, each factor is normalized using the training set statistics, and the final CLI is computed as the sum of the normalized values. We then divide test interactions into low-, medium-, and high-load groups based on CLI tertiles. We compare the full LT-MKT model with its variant without cognitive load modeling, denoted as LT-MKT w/o CL. For both models, we extract the learned student state representations before the prediction layer and visualize them using UMAP (McInnes et al., 2018). As shown in Figure 3, the representations learned by LT-MKT w/o CL are highly mixed across different cognitive-load groups, suggesting that the model does not explicitly organize student states according to the load structure of multi-domain learning. In contrast, LT-MKT produces a clearer and more continuous low-to-high load gradient in the representation space. This pattern indicates that the proposed cognitive load module helps encode cross-domain learning burden into the student state representation, rather than only acting as an additional input feature. Figure 4. Results in cold start scenarios. 5.8. Performance under Cold-Start Scenarios Benefiting from the constructed Multi-domain Hierarchical Graph, LT-MKT is expected to evolve the knowledge states of KCs through their related concepts within and across domains. To more intuitively investigate the effectiveness of such knowledge transfer, we construct a cold-start scenario on the JuniorH dataset and evaluate the model’s generalization ability on previously unseen concepts. Specifically, the dataset is manually divided into training and testing sets, where the testing set contains 33 concepts that do not appear in the training set, accounting for 19% of all concepts. Under this cold-start setting, the training and testing sets account for 84% and 16% of the interaction records, respectively. Figure 4 presents the performance of different methods in terms of AUC and ACC, from which several important observations can be made. First, LT-MKT consistently outperforms all other KT methods, demonstrating that modeling knowledge transfer within and across domains can effectively alleviate the cold-start problem, even when target concepts are absent from the training set. Second, methods considering knowledge structures (e.g., TransKT, SINKT, and GKT) also achieve relatively strong performance, suggesting that explicitly modeling relationships among concepts is beneficial for improving the generalization ability of KT models. Figure 5. A case study of a student’s cross-domain knowledge state evolution under LT-MKT. Figure 6. Parameter sensitivity analysis of λP _P and wsws. 5.9. Parameter Sensitivity Analysis In this section, we analyze the sensitivity of two key hyperparameters, λP _P in Eq. 7 and wsws in Eq. 8. We vary λP _P within 10,30,50,70\10,30,50,70\ while fixing ws=20ws=20, and vary wsws within 5,10,20,40\5,10,20,40\ while fixing λP=30 _P=30. As shown in Figure 6, LT-MKT achieves the best overall performance when λP _P is set to a moderate value. When λP _P is too small, questions with different empirical difficulty levels are compressed into coarse categories, making it difficult for the model to distinguish fine-grained cognitive load caused by question difficulty. Conversely, an overly large λP _P introduces sparse and noisy difficulty levels, where small variations in correctness rates may be over-amplified. This weakens the stability of difficulty embeddings and leads to degraded performance. These results suggest that question difficulty should be modeled with sufficient but not excessive granularity. The results for wsws show a similar pattern. A very small window only captures immediate domain switches and may miss short-term cross-domain learning patterns accumulated over several recent interactions. In contrast, an overly large window incorporates distant historical interactions, which may dilute the influence of recent domain switching and introduce irrelevant temporal noise. The best performance is generally obtained around ws=20ws=20, indicating that cognitive load caused by domain transitions is mainly reflected within a limited recent context. Overall, the sensitivity analysis confirms the necessity of properly balancing difficulty granularity and temporal context when modeling cognitive load in multi-domain learning. 5.10. Case Study To complement the quantitative results, we present a qualitative case study to examine whether LT-MKT captures knowledge transfer and cognitive load in multi-domain learning. As shown in Figure 5, the left part presents a partial concept graph, and the right part shows the evolution of a representative student’s knowledge states from the SeniorH dataset over a sequence of interactions. When the student answers q6q6, which involves concept c3c3, the model assigns an increased state to c3c3 even though the student has limited direct practice on it. This is because the student has previously shown strong mastery of c2c2, which is correlated with c3c3 in the concept graph. In contrast, the student fails to answer q11q11, which involves c2c2, after frequent domain switching. This indicates that cognitive load can weaken immediate learning gains even for previously well-trained concepts. Overall, this case study provides intuitive evidence that LT-MKT can jointly model beneficial knowledge transfer and load-induced learning friction. 6. Conclusion In this paper, we proposed LT-MKT, a multi-domain knowledge tracing framework that jointly models cognitive load and knowledge transfer. By constructing a multi-domain hierarchical graph, LT-MKT captures load effects from cross-domain learning behaviors and propagates knowledge states through intra- and inter-domain relations. Experiments on real-world datasets show that LT-MKT consistently outperforms representative KT baselines, demonstrating its effectiveness in modeling students’ knowledge states in multi-domain learning scenarios. In future work, we will try to incorporate richer cognitive signals to further improve model performance, and investigate the application of LT-MKT in broader personalized learning and intelligent tutoring scenarios. Acknowledgements. This work was supported by grants from the National Key Research and Development Program of China (Grant No. 2024YFC3308200), the National Natural Science Foundation of China (Nos. U25B2072 and 62477044), the Key Technologies R&D Program of Anhui Province (No. 202423k09020039), the Young Elite Scientists Sponsorship Program by CAST (No. 2024QNRC001), and the Fundamental Research Funds for the Central Universities. GenAI Usage Disclosure During the preparation of this manuscript, generative AI tools such as ChatGPT and Grammarly were used strictly for grammar correction and sentence refinement. This paper does not contain any text generated entirely by large language models. All original ideas, experimental designs, and data analyses were conceived and conducted exclusively by the authors. Finally, the authors thoroughly reviewed all edited content and take full responsibility for the final version of the manuscript. References Abdelrahman et al. (2023) G. Abdelrahman, Q. Wang, and B. Nunes Knowledge tracing: a survey. ACM Comput. Surv. 55 (11). Cited by: §1. Abdelrahman and Wang (2019) G. Abdelrahman and Q. Wang Knowledge tracing with sequential key-value memory networks. In Proceedings of the 42nd international ACM SIGIR conference on research and development in information retrieval, p. 175–184. Cited by: §1, §2.1. Achiam et al. (2023) J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774. Cited by: §2.2. Anderson et al. (2014) A. Anderson, D. Huttenlocher, J. Kleinberg, and J. Leskovec Engaging with massive online courses. In Proceedings of the 23rd International Conference on World Wide Web, p. 687–698. Cited by: §1. Cen et al. (2006) H. Cen, K. Koedinger, and B. Junker Learning factors analysis-a general method for cognitive model evaluation and improvement. In Proceedings of Intelligent Tutoring Systems, Berlin, Heidelberg, p. 164–175. External Links: ISBN 978-3-540-35160-3 Cited by: §1, §2.1. Chen et al. (2023) J. Chen, Z. Liu, S. Huang, Q. Liu, and W. Luo Improving interpretability of deep sequential knowledge tracing models with question-centric cognitive representations. In Proceedings of the AAAI conference on artificial intelligence, Vol. 37, p. 14196–14204. Cited by: 7th item. Cheng et al. (2022) S. Cheng, Q. Liu, E. Chen, K. Zhang, Z. Huang, Y. Yin, X. Huang, and Y. Su Adaptkt: a domain adaptable method for knowledge tracing. In Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining, p. 123–131. Cited by: §2.1. Chung et al. (2014) J. Chung, C. Gulcehre, K. Cho, and Y. Bengio Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555. Cited by: §4.2. Corbett and Anderson (1994) A. T. Corbett and J. R. Anderson Knowledge tracing: modeling the acquisition of procedural knowledge. User Modeling and User-adapted Interaction 4 (4), p. 253–278. Cited by: §1, §2.1. Cormier and Hagman (2014) S. M. Cormier and J. D. Hagman Transfer of learning: contemporary research and applications. Academic press. Cited by: §1, §4.3. Devlin et al. (2019) J. Devlin, M. Chang, K. Lee, and K. Toutanova Bert: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), p. 4171–4186. Cited by: §4.2. Fleiss (1971) J. L. Fleiss Measuring nominal scale agreement among many raters.. Psychological bulletin 76 (5), p. 378. Cited by: §5.6. Fu et al. (2024) L. Fu, H. Guan, K. Du, J. Lin, W. Xia, W. Zhang, R. Tang, Y. Wang, and Y. Yu Sinkt: a structure-aware inductive knowledge tracing model with large language model. In Proceedings of the 33rd ACM international conference on information and knowledge management, p. 632–642. Cited by: §1, §2.2, 9th item. Ghosh et al. (2020) A. Ghosh, N. Heffernan, and A. S. Lan Context-aware attentive knowledge tracing. In KDD, p. 2330–2339. Cited by: §1, §2.1, 3rd item. Han et al. (2025) W. Han, W. Lin, L. Hu, Z. Dai, Y. Zhou, M. Li, Z. Liu, C. Yao, and J. Chen Contrastive cross-course knowledge tracing via concept graph guided knowledge transfer. In Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, p. 7401–7409. Cited by: §1, §2.1, §2.2, 11st item. Hu et al. (2023) L. Hu, Z. Dong, J. Chen, G. Wang, Z. Wang, Z. Zhao, and F. Wu PTADisc: a cross-course dataset supporting personalized learning in cold-start scenarios. Advances in Neural Information Processing Systems 36, p. 44976–44996. Cited by: §5.1. Huang et al. (2019) Z. Huang, Q. Liu, C. Zhai, Y. Yin, E. Chen, W. Gao, and G. Hu Exploring multi-objective exercise recommendations in online education systems. In Proceedings of the 28th ACM international conference on information and knowledge management, p. 1261–1270. Cited by: §2.1. Käser et al. (2017) T. Käser, S. Klingler, A. G. Schwing, and M. Gross Dynamic bayesian networks for student modeling. IEEE Transactions on Learning Technologies 10 (4), p. 450–462. Cited by: §1, §2.1. Khan and Coomarasamy (2006) K. S. Khan and A. Coomarasamy A hierarchy of effective teaching and learning to acquire competence in evidenced-based medicine. BMC medical education 6 (1), p. 59. Cited by: §4.3. Lee et al. (2023) U. Lee, S. Yoon, J. S. Yun, K. Park, Y. Jung, D. Stratton, and H. Kim Difficulty-focused contrastive learning for knowledge tracing with a large language model-based difficulty prediction. arXiv preprint arXiv:2312.11890. Cited by: §2.1, §2.2, §4.2, §4.2. Li et al. (2024) Q. Li, W. Xia, K. Du, Q. Zhang, W. Zhang, R. Tang, and Y. Yu Learning structure and knowledge aware representation with large language models for concept recommendation. arXiv preprint arXiv:2405.12442. Cited by: §2.2. Liu et al. (2024a) F. Liu, C. Bu, H. Zhang, L. Wu, K. Yu, and X. Hu FDKT: towards an interpretable deep knowledge tracing via fuzzy reasoning. ACM Transactions on Information Systems 42 (5), p. 1–26. Cited by: §2.1. Liu et al. (2024b) G. Liu, H. Zhan, and J. Kim Question difficulty consistent knowledge tracing. In Proceedings of the ACM on Web Conference 2024, p. 4239–4248. Cited by: §2.1. Liu et al. (2024c) J. Liu, Z. Huang, T. Xiao, J. Sha, J. Wu, Q. Liu, S. Wang, and E. Chen SocraticLM: exploring socratic personalized teaching with large language models. Advances in Neural Information Processing Systems 37, p. 85693–85721. Cited by: §2.2. Liu et al. (2019a) Q. Liu, Z. Huang, Y. Yin, E. Chen, H. Xiong, Y. Su, and G. Hu EKT: exercise-aware knowledge tracing for student performance prediction. IEEE Transactions on Knowledge and Data Engineering (TKDE) 33 (1), p. 100–115. Cited by: §2.1. Liu et al. (2019b) Q. Liu, S. Tong, C. Liu, H. Zhao, E. Chen, H. Ma, and S. Wang Exploiting cognitive structure for adaptive learning. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, p. 627–635. Cited by: §2.1. Liu et al. (2026a) S. Liu, P. Luo, C. Zhang, Y. Chen, H. Zhang, Q. Liu, X. Kou, T. Xu, and E. Chen Look as you think: unifying reasoning and visual evidence attribution for verifiable document rag via reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 32159–32167. Cited by: §2.2. Liu et al. (2026b) S. Liu, J. Zhu, L. Shu, J. Lin, Y. Chen, H. Zhang, C. Zhang, D. Xu, J. Li, B. Tang, et al. Perma: benchmarking personalized memory agents via event-driven preference and realistic task environments. arXiv preprint arXiv:2603.23231. Cited by: §2.2. Liu et al. (2025) Z. Liu, S. Huang, T. Guo, M. Hou, and Q. Liang A prompt-driven framework for multi-domain knowledge tracing. Machine Learning 114 (4), p. 87. Cited by: §2.1, 10th item. Lv et al. (2025) R. Lv, Q. Liu, W. Gao, H. Zhang, J. Lu, and L. Zhu Genal: generative agent for adaptive learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 577–585. Cited by: §2.2. Ma et al. (2024) H. Ma, Y. Yang, C. Qin, X. Yu, S. Yang, X. Zhang, and H. Zhu HD-kt: advancing robust knowledge tracing via anomalous learning interaction detection. In Proceedings of the ACM on Web Conference 2024, p. 4479–4488. Cited by: §2.1. Macaulay and Cree (1999) C. Macaulay and V. E. Cree Transfer of learning: concept and process. Social work education 18 (2), p. 183–194. Cited by: §4.3. Malinka et al. (2023) K. Malinka, M. Peresíni, A. Firc, O. Hujnák, and F. Janus On the educational impact of chatgpt: is artificial intelligence ready to obtain a university degree?. In Proceedings of the 2023 Conference on Innovation and Technology in Computer Science Education V. 1, p. 47–53. Cited by: §2.2. McInnes et al. (2018) L. McInnes, J. Healy, and J. Melville Umap: uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426. Cited by: §5.7. Nakagawa et al. (2019) H. Nakagawa, Y. Iwasawa, and Y. Matsuo Graph-based knowledge tracing: modeling student proficiency using graph neural network. In Proceedings of IEEE/WIC/ACM International Conference on Web Intelligence (WI), p. 156–163. Cited by: §1, §2.1, 2nd item. Ni et al. (2024) L. Ni, S. Wang, Z. Zhang, X. Li, X. Zheng, P. Denny, and J. Liu Enhancing student performance prediction on learnersourced questions with sgnn-llm synergy. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, p. 23232–23240. Cited by: §2.2. Ni et al. (2023) Q. Ni, T. Wei, J. Zhao, L. He, and C. Zheng HHSKT: a learner–question interactions based heterogeneous graph neural network model for knowledge tracing. Expert Systems with Applications 215, p. 119334. Cited by: §1, §2.1. Pandey and Karypis (2019) S. Pandey and G. Karypis A self-attentive model for knowledge tracing. CoRR abs/1907.06837. External Links: 1907.06837 Cited by: §2.1. Pavlik et al. (2009) P. I. Pavlik, H. Cen, and K. R. Koedinger Performance factors analysis-a new alternative to knowledge tracing. In Proceedings of Conference on Artificial Intelligence in Education: Building Learning Systems That Care: From Knowledge Representation to Affective Modelling, NLD, p. 531–538. External Links: ISBN 9781607500285 Cited by: §1, §2.1. Piech et al. (2015) C. Piech, J. Spencer, J. Huang, S. Ganguli, M. Sahami, L. Guibas, and J. Sohl-Dickstein Deep knowledge tracing. In Proceedings of International Conference on Neural Information Processing Systems (NeurIPS), p. 505–513. Cited by: §1, §2.1, 1st item. Plass et al. (2010) J. L. Plass, R. Moreno, and R. Brünken Cognitive load theory. Cited by: §1, §4.2. Sha et al. (2026) J. Sha, H. Zhang, J. Wu, S. Wang, Z. Ling, S. Wang, and S. Wei Process-supervised multi-step reasoning with llms for knowledge tagging. In 2026 International Annual Conference on Complex Systems and Intelligent Science (CSIS-IAC), p. 534–539. Cited by: §2.2. Shen et al. (2022) S. Shen, Z. Huang, Q. Liu, Y. Su, S. Wang, and E. Chen Assessing student’s dynamic knowledge state by exploring the question difficulty effect. In Proceedings of the 45th international ACM SIGIR conference on research and development in information retrieval, p. 427–437. Cited by: §4.2, 6th item. Shen et al. (2021) S. Shen, Q. Liu, E. Chen, Z. Huang, W. Huang, Y. Yin, Y. Su, and S. Wang Learning process-consistent knowledge tracing. In Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining, p. 1452–1460. Cited by: §2.1, 5th item. Shen et al. (2024) S. Shen, Q. Liu, Z. Huang, Y. Zheng, M. Yin, M. Wang, and E. Chen A survey of knowledge tracing: models, variants, and applications. IEEE Transactions on Learning Technologies 17 (), p. 1898–1919. Cited by: §1, §2.1. Shin et al. (2021) D. Shin, Y. Shim, H. Yu, S. Lee, B. Kim, and Y. Choi SAINT+: Integrating temporal features for ednet correctness prediction. In Proceedings of LAK21: International Learning Analytics and Knowledge Conference, LAK21, New York, NY, USA, p. 490–496. External Links: ISBN 9781450389358 Cited by: §1, §2.1. Song et al. (2022a) X. Song, J. Li, T. Cai, S. Yang, T. Yang, and C. Liu A survey on deep learning based knowledge tracing. Knowledge-Based Systems 258, p. 110036. Cited by: §1. Song et al. (2022b) X. Song, J. Li, Q. Lei, W. Zhao, Y. Chen, and A. Mian Bi-clkt: bi-graph contrastive learning based knowledge tracing. Knowledge-Based Systems 241, p. 108274. Cited by: §2.1. Sonkar and Baraniuk (2023) S. Sonkar and R. G. Baraniuk Deduction under perturbed evidence: probing student simulation (knowledge tracing) capabilities of large language models.. In LLM@ AIED, p. 26–33. Cited by: §2.2. Sun et al. (2024) J. Sun, F. Yu, Q. Wan, Q. Li, S. Liu, and X. Shen Interpretable knowledge tracing with multiscale state representation. In Proceedings of the ACM Web Conference 2024, p. 3265–3276. Cited by: §2.1, 8th item. Tang et al. (2024) Y. Tang, W. Yang, Y. Xie, and M. Yang Domain adaptive knowledge tracing. International Journal of Machine Learning and Cybernetics, p. 1–14. Cited by: §2.1. Tong et al. (2020) H. Tong, Z. Wang, Q. Liu, Y. Zhou, and W. Han HGKT: introducing hierarchical exercise graph for knowledge tracing. arXiv preprint arXiv:2006.16915. Cited by: §2.1. VanLEHN (2011) K. VanLEHN The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist 46 (4), p. 197–221. Cited by: §1. Vie and Kashima (2019) J. Vie and H. Kashima Knowledge tracing machines: factorization machines for knowledge tracing. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), Vol. 33, p. 750–757. Cited by: §1, §2.1. Wang et al. (2021) C. Wang, W. Ma, M. Zhang, C. Lv, F. Wan, H. Lin, T. Tang, Y. Liu, and S. Ma Temporal cross-effects in knowledge tracing. In Proceedings of the 14th ACM International Conference on Web Search and Data Mining, p. 517–525. Cited by: §2.1, 4th item. Wang et al. (2022) F. Wang, Q. Liu, E. Chen, Z. Huang, Y. Yin, S. Wang, and Y. Su NeuralCD: a general framework for cognitive diagnosis. IEEE Transactions on Knowledge and Data Engineering 35 (8), p. 8312–8327. Cited by: §2.1. Wang et al. (2024) S. Wang, T. Xu, H. Li, C. Zhang, J. Liang, J. Tang, P. S. Yu, and Q. Wen Large language models for education: a survey and outlook. arXiv preprint arXiv:2403.18105. Cited by: §2.2. Wang et al. (2026a) X. Wang, Z. Du, S. Wu, Z. Liu, H. Zhang, J. Zhang, J. Zhu, S. Wang, and K. Zhang Rec2: embedding table reconstruction for deep recommender systems. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, p. 5081–5090. Cited by: §2.2. Wang et al. (2026b) X. Wang, Z. Du, H. Xu, S. Yin, Y. Han, J. Zhu, K. Zhang, and Q. Liu Personalized visual content generation in conversational systems. Advances in Neural Information Processing Systems 38, p. 34571–34602. Cited by: §2.2. Wei et al. (2022) J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35, p. 24824–24837. Cited by: §4.1. Wu et al. (2024) J. Wu, H. Zhang, Z. Huang, L. Ding, Q. Liu, J. Sha, E. Chen, and S. Wang Graph-based student knowledge profile for online intelligent education.. In SDM, p. 379–387. Cited by: §2.1. Wu et al. (2025) Z. Wu, Y. Liu, J. Cen, Z. Zheng, and G. Xu A cross-domain knowledge tracing model based on graph optimal transport. World Wide Web 28 (1), p. 1–25. Cited by: §2.1, §4.2. Zhang et al. (2022) H. Zhang, C. Bu, F. Liu, S. Liu, Y. Zhang, and X. Hu APGKT: exploiting associative path on skills graph for knowledge tracing. In Pacific Rim International Conference on Artificial Intelligence, p. 353–365. Cited by: §2.1. Zhang et al. (2024) H. Zhang, S. Shen, B. Xu, Z. Huang, J. Wu, J. Sha, and S. Wang Item-difficulty-aware learning path recommendation: from a real walking perspective. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, p. 4167–4178. Cited by: §1. Zhang et al. (2026) H. Zhang, J. Wu, Q. Liu, R. Lv, L. Ding, Z. Huang, J. Sha, and S. Wang CBEGRec: learning path recommendation via concept bundling and exercise generation. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, p. 6452–6463. Cited by: §1. Zhang et al. (2017) J. Zhang, X. Shi, I. King, and D. Y. Yeung Dynamic key-value memory networks for knowledge tracing. In Proceedings of International Conference on World Wide Web (W), p. 765–774. Cited by: §1, §2.1.