Paper deep dive
Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning
Alperen Kantarci, Visvanathan Ramesh, Gemma Roig
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/29/2026, 4:25:16 AM
Summary
This paper presents a multimodal framework for predicting student engagement in online tutoring videos using the CASED dataset. The model integrates spatiotemporal features from video, audio, and screen content via Perceiver IO fusion, combined with explicit behavioral cues (head pose, gaze, facial action units). It employs a hierarchical Bayesian context layer to model student and instructor personalities and uses evidential regression (NIG) and Gaussian process classification (SNGP) heads for uncertainty-aware prediction. The framework achieves competitive performance on the CASED challenge test set, highlighting the difficulty of the task and the importance of calibrated uncertainty metrics.
Entities (10)
Relation Signals (9)
CASED → usedfor → Student Engagement Prediction
confidence 95% · We use ... CASED (Sharma et al., 2026; Li et al., 2025) dataset for fine-tuning.
Perceiver IO → usedin → Multimodal Fusion
confidence 95% · All token streams are fused with Perceiver IO
V-JEPA 2 → usedfor → Video Encoding
confidence 92% · V-JEPA 2 (Assran et al., 2025) for student and instructor video
CLIP → usedfor → Screen Content Encoding
confidence 92% · CLIP (Radford et al., 2021) for the screen content
AudioMAE → usedfor → Audio Encoding
confidence 92% · AudioMAE (Huang et al., 2022) for the spectrogram.
NIG → usedfor → regression
confidence 90% · For regression, we use an evidential Normal-Inverse-Gamma (NIG) head
SNGP → usedfor → Classification
confidence 90% · For classification, we use a spectral-normalized neural Gaussian process (SNGP) head
DaiSEE → usedfor → Pretraining
confidence 90% · We use DaiSEE ... datasets for pretraining
Aff-Wild2 → usedfor →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The prediction of student engagement from the online tutoring videos is difficult because engagement is a multidimensional construct comprising distinct behavioral, emotional, and cognitive states. A reliable prediction requires bringing together different types of behavioral signals as well as expressive cues. Through our analysis of the CASED dataset, it is clear that engagement prediction gets even harder due to the high inter-person variability as well as the subjectivity of the engagement annotation. To tackle these challenges, we develop a multimodal framework that integrates the implicit spatiotemporal features extracted from pretrained video, audio, and image encoders along with structured behavioral modalities like head pose, gaze, facial action units, emotion, and wavelet-based audio features. We integrate these modalities via a Perceiver IO latent bottleneck. Moreover, student and instructor personalities are modeled as variational posteriors over learnable embeddings to enable partial pooling across participants. We employ evidential regression and spectral-normalized Gaussian process classification heads for uncertainty-aware prediction to further improve robustness and calibration. Benchmark on the CASED challenge test set shows that all participating methods converge near random-chance performance, revealing the difficulty of the dataset. In this highly ambiguous regime, our framework achieves competitive performance while uniquely offering well-calibrated uncertainty metrics, demonstrating that reliable risk-quantification is an essential prerequisite for deploying engagement models in real-world educational tools.
Tags
Links
- Source: https://arxiv.org/abs/2608.24340v1
- Canonical: https://arxiv.org/abs/2608.24340v1
Trouble viewing inline? Open PDF directly →
Full Text
28,587 characters extracted from source content.
Expand or collapse full text
Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning Conference: INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION; October 05–09, 2026; Napoli, ItalyINTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION (ICMI ’26), October 05–09, 2026, Napoli, ItalyDOI: 10.1145/3776574.3832485ISBN: 979-8-4007-2318-6/2026/101002CCS: Computing methodologies Multi-task learningCCS: Computing methodologies Computer vision representationsCCS: Applied computing Interactive learning environments Alperen Kantarcı Affiliation: Institute of Computer Science, Goethe University Frankfurt, Frankfurt am Main, Germany email: kantarci@em.uni-frankfurt.de , Visvanathan Ramesh Affiliation: Institute of Computer Science, Goethe University Frankfurt, The Hessian Center for Artificial Intelligence, Frankfurt am Main, Germany email: vramesh@em.uni-frankfurt.de and Gemma Roig Affiliation: Institute of Computer Science, Goethe University Frankfurt, The Hessian Center for Artificial Intelligence, Frankfurt am Main, Germany email: roignoguera@em.uni-frankfurt.de © none Abstract. The prediction of student engagement from the online tutoring videos is difficult because engagement is a multidimensional construct comprising distinct behavioral, emotional, and cognitive states. A reliable prediction requires bringing together different types of behavioral signals as well as expressive cues. Through our analysis of the CASED dataset, it is clear that engagement prediction gets even harder due to the high inter-person variability as well as the subjectivity of the engagement annotation. To tackle these challenges, we develop a multimodal framework that integrates the implicit spatiotemporal features extracted from pretrained video, audio, and image encoders along with structured behavioral modalities like head pose, gaze, facial action units, emotion, and wavelet-based audio features. We integrate these modalities via a Perceiver IO latent bottleneck. Moreover, student and instructor personalities are modeled as variational posteriors over learnable embeddings to enable partial pooling across participants. We employ evidential regression and spectral-normalized Gaussian process classification heads for uncertainty-aware prediction to further improve robustness and calibration. Benchmark on the CASED challenge test set shows that all participating methods converge near random-chance performance, revealing the difficulty of the dataset. In this highly ambiguous regime, our framework achieves competitive performance while uniquely offering well-calibrated uncertainty metrics, demonstrating that reliable risk-quantification is an essential prerequisite for deploying engagement models in real-world educational tools. 1. Introduction and Background Student engagement is widely recognized as a fundamental prerequisite for effective learning and academic success (Bergdahl et al., 2024; Finn and Zimmer, 2012). From the perspective of educational psychology, engagement is a multidimensional construct with behavioral, emotional, cognitive, and social dimensions that collectively reflect a learner’s involvement in the educational process. High levels of engagement have been associated with improved learning outcomes, higher knowledge retention, enhanced motivation, and reduced dropout rates (Fredricks et al., 2004). Therefore, understanding and analyzing student engagement is an important objective for educational researchers. A significant challenge is that engagement indicators are distributed across different modalities. A student’s face and posture can reflect their attention, or an instructor’s behavior can influence responsiveness. The screen content can affect cognitive load and audio can convey speech-based information (Bali et al., 2026). Moreover, these signals are not equally informative in every instance. For example, gaze direction may become unreliable when the face is partially obscured, or emotion predictions may become unstable at low resolutions. This necessitates the development of architectures capable of integrating diverse modalities while remaining robust against varying signal quality. Figure 1. Overview of the proposed multimodal engagement framework. Multi-source inputs (student/instructor video, slides, audio, and behavioral cues) are encoded via specialized models and fused using a Perceiver IO block to form the core feature representation (d=768) This is combined with a personality-driven Bayesian Context embedding via a gating mechanism. The final concatenated representation (d=32) feeds into an Evidential Regression head and a GP Classification head to simultaneously predict engagement levels and explicit epistemic uncertainty. The CASED (Sharma et al., 2026; Li et al., 2025) challenge reflects this difficulty. Predicting student engagement from short clips of online tutoring sessions is complicated by several factors: limited training data, substantial inter-person variability, the inherent subjective nature of engagement annotation (Khan and Safa, 2024; Khan et al., 2022). In these settings, predicting performance can be near-random or weakly above-baseline results. This can indicate not only model limitations but also dataset ambiguity, label noise, or insufficient signal in the available observations. Therefore, methods for this task should be evaluated not only by predictive accuracy, but also by how well they handle uncertainty, exploit multimodal structure, and generalize to unseen students. To this end, we present a multimodal engagement prediction framework that combines behavioral cues using large pretrained encoders, efficient multimodal fusion, participant-level Bayesian context modeling, and uncertainty-aware prediction heads. Particulary, we extract explicit behavioral features including head pose, gaze, facial action units, emotion estimates and more generic representations from student video, instructor video, screen content and audio. These various features are fused using Perceiver IO (Jaegle et al., 2021) which is an modality-agnostic asymmetric attention mechanism. We also introduce a hierarchical Bayesian context layer over student and instructor embeddings with partial pooling toward a shared prior to model person-specific effects without encouraging identity memorization. 1.1. Proposed method Given a video clip V, we jointly predict a continuous engagement score y∈[1,5]y∈[1,5] and a binary label ℓ∈0,1 ∈\0,1\. Each clip is associated with a student identity s, an instructor identity i, and optional personality vectors ts,ti∈ℝ10t_s,t_i ^10. Evaluation uses a student-independent split which means training and testing sets have different students. Each clip is decomposed into student video, instructor video, screen content, audio, and explicit behavioral features. For 1280×12801280× 1280 composite videos, we crop fixed student, instructor, and screen regions, resize them to 224×224224× 224, and uniformly sample T=64T=64 frames for student and instructor streams. Audio is resampled to 16 kHz mono and converted to a Kaldi filterbank (Povey et al., 2011) spectrogram of shape (1024,128)(1024,128). We use pretrained encoders for the main modalities: V-JEPA 2 (Assran et al., 2025) for student and instructor video, CLIP (Radford et al., 2021) for the screen content, and AudioMAE (Huang et al., 2022) for the spectrogram. Overall architecture visualization can be seen in Figure 1. Their outputs are projected to a shared dimension, d=768. We utilize a single Linear Layer for each modality, followed by Layer Normalization to stabilize the inputs before they enter the Perceiver IO bottleneck. Zstudent,Zinstructor∈ℝN×768,Zscreen∈ℝP×768,Zaudio∈ℝM×768.Z_student,Z_instructor ^N× 768, Z_screen ^P× 768, Z_audio ^M× 768. We additionally extract five explicit feature streams at T=64T=64: head pose, CWT head dynamics, facial action units, gaze/head angles, and emotion prediction probabilities. Each feature sequence Fk∈ℝT×dkF_k ^T× d_k is linearly projected as Zk=FkWk∈ℝT×768Z_k=F_kW_k ^T× 768. All token streams are fused with Perceiver IO (Jaegle et al., 2021). Let ℳM be the set of modalities and ecke_c_k the learned embedding for modality k. The input token set is (1) X=concat[Zk+eck∣k∈ℳ]∈ℝTtotal×768.X=concat [Z_k+e_c_k k ] ^T_total× 768. A latent array L0∈ℝQ×768L_0 ^Q× 768 with Q=64Q=64 queries is updated for L=3L=3 layers via cross-attention to X and latent self-attention. The fused representation is obtained by mean pooling: (2) f=mean(LN(L))∈ℝ768.f=mean(LN(L_L)) ^768. Bayesian Participant Context: To model stable person-specific effects, we assign each student s and instructor i a variational embedding: (3) q(zs)=(μs,diag(exp(σs2))),q(zi)=(μi,diag(exp(σi2))).q(z_s)=N( _s,diag( ( _s^2))), q(z_i)=N( _i,diag( ( _i^2))). Embeddings are sampled using the reparameterization trick at the training time. Posterior means are used, and unseen students receive a learned prior mean μprior _prior during the inference. When personality metadata are available, they are added through (4) z~s=zs+mtanh(Wtraitts), z_s=z_s+m\, (W_traitt_s), where m is a binary mask dropped during training and set to zero at test time. Student and instructor context are combined by c=σ(Wgate[z~s;z~i])∈ℝ32c=σ(W_gate[ z_s; z_i]) ^32, and concatenated with the fused representation: f^=[f;c]∈ℝ800. f=[f;c] ^800. A KL penalty regularizes participant posteriors toward a shared prior: (5) ℒKL=1N∑sDKL((μs,σs2)∥(μprior,I)).L_KL= 1N _sD_KL\! (N( _s, _s^2)\,\|\,N( _prior,I) ). Uncertainty-Aware Prediction Heads: For regression, we use an evidential Normal-Inverse-Gamma (NIG) head (Amini et al., 2020), which predicts (γ,ν,α,β)(γ,ν,α,β), where γ is the engagement estimate and the remaining parameters define predictive uncertainty. The corresponding aleatoric and epistemic uncertainties are β/(α−1)β/(α-1) and β/(ν(α−1)).β/(ν(α-1)). For classification, we use a spectral-normalized neural Gaussian process (SNGP) head (Liu et al., 2020), with predictive probability (6) p(ℓ=1∣f^)=σ(w⊤f^1+πκ2/8),p( =1 f)=σ\! ( w f 1+πκ^2/8 ), where κ2κ^2 is the predictive variance. Training Objective: The total loss is (7) ℒ=(λNIGℒNIG+λCℒC)eσr+σr+(λclsℒcls+λordℒord)eσc+σc+ℒKL,L= ( _NIGL_NIG+ _CL_C )e _r+ _r+ ( _clsL_cls+ _ordL_ord )e _c+ _c+L_KL, where σr,σc _r, _c are learnable task-uncertainty weights (Kendall et al., 2018). We use NIG loss and C loss for regression, weighted cross-entropy for classification, and an auxiliary ordinal consistency loss coupling regression and classification outputs. Optimization and Ensemble: We train with AdamW (Loshchilov and Hutter, 2019), weight decay 5×10−25× 10^-2, cosine decay, and 5% warmup. Backbone encoders use learning rate 10−510^-5, while newly initialized layers use 5×10−45× 10^-4. Training is staged: encoders are frozen for epochs 1–20 and unfrozen until the convergance. We additionally use random temporal sampling, metadata dropout, and KL (Kullback-Leibler) annealing. For our final predictions we use four different training checkpoints of the same model as ensemble of networks and do a majority voting on the predictions. 2. Experiments and Dataset Analysis We use DaiSEE (Gupta et al., 2016) and Aff-Wild2 (Kollias, 2022a; Kollias, 2022b; Kollias and Zafeiriou, 2021) datasets for pretraining and CASED (Sharma et al., 2026; Li et al., 2025) dataset for fine-tuning. We evaluate under a strict student-independent protocol. For our trainings and validation experiments, we use 5-fold cross-validation with student-level fold assignment. Final leaderboard performances are reported from CASED test set. We report F1 Macro, F1 Weighted, MCC, RMSE, MSE, MAE, R2R^2, Pearson Correlation and Concordance Correlation Coefficient (C) depending on Classification and regression tasks. 2.1. Dataset analysis The binary label distribution is highly skewed: approximately 69% of clips are labeled engaged (label 0) and 31% not-engaged (label 1). The continuous engagement scores are similarly compressed, with mean 3.7 and median 4.0, producing a long lower tail. A majority-class classifier therefore achieves approximately 40–45% macro F1 without learning any discriminative features. Student Analysis: Individual students differ substantially in their mean engagement level and within-session variability. Several students are labeled engaged in over 90% of their clips while others fall below 40%. This hints that a large fraction of the label variance is explained by between-student differences rather than within-clip behavioral dynamics, which directly limits what any clip-level visual model can learn. Instructor Analysis: Per-instructor engagement rates do not vary systematically. All three instructors have 3.37 mean engagement with similar standard deviation. Around 68% of clips are labeled as engaged for all instructors. Furthermore, all students are paired with exactly one instructor across all sessions, making student and instructor identity perfectly collinear. 3. Results Table 1 and Table 2 summarize the official challenge leaderboard for the classification and regression tracks. Our method performs competitively among participating teams. However, the margin between the top methods is small. Best-performing method obtained a F1 Macro of 0.52, suggesting that all approaches face with similar limitations imposed by the dataset and the task. Other metrics also show very similar performances. In the regression task, our approach achieved a matching or closely matching the best-performing methods. However, all submissions obtained near-zero or negative values for R2R^2, Pearson correlation, and C, indicating that none of the evaluated methods was able to reliably capture the underlying engagement signal. Collectively, the leaderboard suggests that the proposed framework performs competitively while highlighting the intrinsic difficulty of automatic engagement prediction on this benchmark. Table 1. Classification performance of participating methods on the CASED challenge test set. Participant F1 (Macro) F1 (Weighted) Precision Recall MCC saurabhh 0.52 0.60 0.52 0.52 0.04 mohitvu 0.52 0.61 0.52 0.52 0.04 adim66 0.51 0.58 0.51 0.52 0.03 Ours 0.51 0.60 0.51 0.51 0.02 priscalab 0.51 0.60 0.51 0.51 0.01 caymann 0.50 0.58 0.50 0.50 0.01 KvochurHegel 0.42 0.59 0.35 0.50 0.00 Table 2. Regression performances on the CASED challenge test set. Participant RMSE MSE MAE R2R^2 Pearson C saurabhh 0.68 0.46 0.57 -0.01 -0.00 -0.00 mohitvu 0.78 0.61 0.63 -0.33 -0.00 -0.00 Ours 0.68 0.46 0.57 -0.00 -0.02 -0.00 priscalab 0.68 0.47 0.57 -0.02 0.01 0.00 caymann 0.68 0.47 0.57 -0.01 0.03 0.01 KvochurHegel 0.68 0.47 0.57 -0.01 -0.03 -0.00 Multimodal and Component Ablation We evaluate the impact of different modality combinations and architectural components on engagement prediction. As shown in Table 3, the inclusion of all three modalities—Student, Instructor, and Content (S+I+C)—yields the highest performance, achieving a C of 0.018 and an F1-macro score of 0.498. This confirms that student engagement is heavily contextual and benefits from integrating student behavior alongside instructor and screen dynamics. In the component ablation Table 3, the Full Model outperforms or remains highly competitive with alternative configurations. The base transformer baseline achieves lower RMSE and MAE. The integration of the Bayesian Context layer, NIG and SNGP heads optimizes macro-level classification and alignment metrics, yielding the top F1-macro and a strong C scores. Table 3. Ablation study. (a) Modality contribution using the full model (BayesianContext + NIG + SNGP). (b) Component contribution on the three-modality input (S=student video, I=instructor video, C=content). Metrics are averaged over 5 student-grouped cross-validation folds. Model Modalities C ↑ RMSE ↓ MAE ↓ F1-macro ↑ MCC ↑ (a) Modality ablation — Full model Full (S) Student -0.016 ± 0.027 1.048 ± 0.060 0.848 ± 0.046 0.486 ± 0.015 -0.002 ± 0.029 Full (S+I) Student + Instructor -0.021 ± 0.016 1.020 ± 0.119 0.832 ± 0.111 0.484 ± 0.022 -0.011 ± 0.016 Full (S+I+C) Student + Instructor + Content 0.018 ± 0.029 0.943 ± 0.038 0.755 ± 0.040 0.498 ± 0.019 0.002 ± 0.034 (b) Component ablation — Full modality set (S+I+C) Base Transformer S+I+C 0.015 ± 0.028 0.821 ± 0.022 0.671 ± 0.026 0.497 ± 0.013 0.012 ± 0.019 + NIG + SNGP S+I+C 0.008 ± 0.018 0.935 ± 0.098 0.759 ± 0.093 0.491 ± 0.017 0.002 ± 0.016 + Bayesian Context S+I+C 0.019 ± 0.021 0.848 ± 0.047 0.691 ± 0.041 0.492 ± 0.018 0.017 ± 0.007 Full Model (Ours) S+I+C 0.018 ± 0.029 0.943 ± 0.038 0.755 ± 0.040 0.498 ± 0.019 0.002 ± 0.034 Table 4. NIG uncertainty decomposition by engagement level. Samples are grouped into five equal-width bins on the 1–5 continuous engagement scale. Engagement Level N Mean Pred Aleatoric ↓ Epistemic ↓ Total 1.0–1.8 41 3.45 42.0879 3323.3606 3365.4482 1.9–2.6 657 3.37 33.1598 2623.2993 2656.4590 2.7–3.4 1447 3.36 34.5978 2726.2368 2760.8345 3.5–4.2 2398 3.38 34.7966 2802.9478 2837.7444 4.3–5.0 435 3.42 36.4595 3008.4961 3044.9551 Pearson r (engagement vs aleatoric): 0.016; vs epistemic: 0.025. Uncertainty Decomposition Table 4 analyzes the uncertainty captured by the NIG regression head across different ground-truth engagement levels. Total uncertainty is heavily dominated by epistemic uncertainty across all bins. Pearson correlation coefficients show negligible linear relationships between raw engagement scores and both aleatoric and epistemic uncertainties, indicating that model confidence is driven by factors other than the absolute magnitude of the engagement score itself. Table 5. Bayesian context layer posterior statistics. The KL divergence from the population prior and the posterior mean norm are shown. Entity Name KL from Prior ↓ ‖μ‖2\|μ\|_2 N Clips All students (mean) — 14.3253 0.0273 116 All instructors (mean) — 13.1287 0.0167 1659 Top-3 students by KL (most personalized) Student laurencedu 15.7200 0.0257 26 Student pierreyoussef 15.4660 0.0795 50 Student richard 15.2584 0.0311 55 Bottom-3 students by KL (least personalized) Student administrator 13.4648 0.0191 116 Student sophie 13.3617 0.0205 129 Student akhat 12.6692 0.0165 246 Instructors Instructor Nigel Lu 11.7067 0.0149 706 Instructor Pierre 11.4088 0.0122 1454 Instructor Catherine 11.2344 0.0111 2818 Bayesian Context Layer Personalization We evaluate the behavioral shifts in the learned identity embeddings via the KL divergence from the population prior. Table 5 shows student identities experience a higher deviation from the population prior (mean KL=14.32) compared to instructor identities (mean KL=13.12), indicating stronger personalization for individual learners. Moreover, statistical analysis reveals a strong, significant negative correlation between a student’s clip count and their KL divergence from the prior (r=-0.748, p=0.001). This suggests that the model applies aggressive, highly individualized posterior shifts to sparse data. Conversely, students with abundant training clips converge closer to a well-regularized population norm. 4. Limitations and Discussion The proposed method has several limitations due to both data and architectural choices. Most of evaluated configurations on test set yield validation C below 0.02 and F1-macro not more than 0.52, indicating models do not consistently outperform a constant mean predictor. There are several possible reasons for this. First, the not-engaged class drives the dominant failure mode. Clips near the Likert midpoint occupy an annotation boundary where small annotator perturbations flip the binary label. Secondly, most of representations remain identity-discriminative, causing the model to fit student appearance rather than engagement dynamics. The population prior in the Bayesian context layer is too weakly constrained to compensate. Finally, each clip receive a single label while Perceiver IO mean-pools over 64 frames, suppressing within-clip fluctuations. Finer-grained temporal supervision would be helpful for detecting small engagement cues. 5. Conclusion We presented a multimodal engagement prediction framework that jointly addresses regression and classification under uncertainty by combining Perceiver IO (Jaegle et al., 2021) fusion,a hierarchical Bayesian context layer, evidential NIG regression, SNGP classifier. Our ablation study reveals that combination of student, instructor, screen provides the lowest error across all metrics. The Bayesian context layer learns distinguishable per-student posteriors, with a strong negative correlation between clip count and KL divergence from the population prior, confirming that hierarchical regularisation correctly controls personalization as a function of data availability. On the CASED challenge leaderboard the system achieves competitive regression performance while producing calibrated epistemic uncertainty estimates at inference time. As a future work extending the Bayesian context layer with explicit personality-trait conditioning and cross-session identity tracking are promising directions for further improving personalised engagement modelling. Safe and Responsible Innovation Statement This work processes video recordings of students in tutoring sessions, raising inherent privacy concerns. All data used in this study was collected under informed consent as part of the CASED (Sharma et al., 2026; Li et al., 2025) challenge, DaiSEE (Gupta et al., 2016) and Aff-Wild2 (Kollias, 2022a; Kollias, 2022b; Kollias and Zafeiriou, 2021) datasets. The model relies on face analysis and behavioral signal extraction, which may exhibit performance disparities across demographic groups not well-represented in the relatively small training cohort. The system is intended as a research tool for understanding engagement dynamics, not as a surveillance or performance evaluation instrument for deployed educational settings. References Amini et al. (2020) A. Amini, W. Schwarting, A. Soleimany, and D. Rus Deep evidential regression. Advances in neural information processing systems 33, p. 14927–14937. Cited by: §1.1. Assran et al. (2025) M. Assran, A. Bardes, D. Fan, Q. Garrido, R. Howes, M. Muckley, A. Rizvi, C. Roberts, K. Sinha, A. Zholus, et al. V-jepa 2: self-supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985. Cited by: §1.1. Bali et al. (2026) C. Bali, B. Tasdelen, S. Bandi, and A. Zsidó Understanding the cognitive cost of multimedia learning: effects of visual load and language proficiency. Cognitive Research: Principles and Implications 11 (1), p. 2. Cited by: §1. Bergdahl et al. (2024) N. Bergdahl, M. Bond, J. Sjöberg, M. Dougherty, and E. Oxley Unpacking student engagement in higher education learning analytics: a systematic review. International Journal of Educational Technology in Higher Education 21 (1), p. 63. Cited by: §1. Finn and Zimmer (2012) . D. Finn and . S. Zimmer Student engagement: what is it? why does it matter?. In Handbook of Research on Student Engagement, p. 97–131 (English). Note: Publisher Copyright: © Springer Science+Business Media, LLC 2012. All rights reserved. External Links: Document, ISBN 9781461420170 Cited by: §1. Fredricks et al. (2004) J. A. Fredricks, P. C. Blumenfeld, and A. H. Paris School engagement: potential of the concept, state of the evidence. Review of educational research 74 (1), p. 59–109. Cited by: §1. Gupta et al. (2016) A. Gupta, A. D’Cunha, K. Awasthi, and V. Balasubramanian Daisee: towards user engagement recognition in the wild. arXiv preprint arXiv:1609.01885. Cited by: §2, Safe and Responsible Innovation Statement. Huang et al. (2022) P. Huang, H. Xu, J. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, and C. Feichtenhofer Masked autoencoders that listen. Advances in neural information processing systems 35, p. 28708–28720. Cited by: §1.1. Jaegle et al. (2021) A. Jaegle, S. Borgeaud, J. Alayrac, C. Doersch, C. Ionescu, D. Ding, S. Koppula, A. Brock, E. Shelhamer, O. J. H’enaff, M. M. Botvinick, A. Zisserman, O. Vinyals, and J. Carreira Perceiver io: a general architecture for structured inputs & outputs. ArXiv abs/2107.14795. External Links: Link Cited by: §1.1, §1, §5. Kendall et al. (2018) A. Kendall, Y. Gal, and R. Cipolla Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 7482–7491. Cited by: §1.1. Khan et al. (2022) S. S. Khan, A. Abedi, and T. J. F. Colella Inconsistencies in measuring student engagement in virtual learning - a critical review. ArXiv abs/2208.04548. External Links: Link Cited by: §1. Khan and Safa (2024) S. Khan and S. Safa Revisiting annotations in online student engagement. In Proceedings of the 2024 10th International Conference on Computing and Data Engineering, ICCDE ’24, New York, NY, USA, p. 111–117. External Links: ISBN 9798400709319, Link, Document Cited by: §1. Kollias and Zafeiriou (2021) D. Kollias and S. Zafeiriou Analysing affective behavior in the second abaw2 competition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 3652–3660. Cited by: §2, Safe and Responsible Innovation Statement. Kollias (2022a) D. Kollias ABAW: learning from synthetic data & multi-task learning challenges. arXiv preprint arXiv:2207.01138. Cited by: §2, Safe and Responsible Innovation Statement. Kollias (2022b) D. Kollias Abaw: valence-arousal estimation, expression recognition, action unit detection & multi-task learning challenges. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 2328–2336. Cited by: §2, Safe and Responsible Innovation Statement. Li et al. (2025) J. Li, G. Sharma, and H. Salam Personality-aware engagement prediction in online learning. In Proceedings of the 3rd International Workshop on Multimodal and Responsible Affective Computing, MRAC ’25, New York, NY, USA, p. 119–127. External Links: ISBN 9798400720529, Link, Document Cited by: §1, §2, Safe and Responsible Innovation Statement. Liu et al. (2020) J. Liu, Z. Lin, S. Padhy, D. Tran, T. Bedrax Weiss, and B. Lakshminarayanan Simple and principled uncertainty estimation with deterministic deep learning via distance awareness. Advances in neural information processing systems 33, p. 7498–7512. Cited by: §1.1. Loshchilov and Hutter (2019) I. Loshchilov and F. Hutter Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, External Links: Link Cited by: §1.1. Povey et al. (2011) D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, et al. The kaldi speech recognition toolkit. In IEEE 2011 workshop on automatic speech recognition and understanding, Cited by: §1.1. Radford et al. (2021) A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, p. 8748–8763. Cited by: §1.1. Sharma et al. (2026) G. Sharma, J. Li, and H. Salam SMART challenge series: context-aware student engagement detection. Zenodo. External Links: Document, Link Cited by: §1, §2, Safe and Responsible Innovation Statement.