Paper deep dive
CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation
Fengming Yu, Haiwei Pan, Kejia Zhang, Chunling Chen, Jian Guan, Baoying Ma
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/4/2026, 11:24:31 AM
Summary
The paper introduces CoCaRS, a method for heterogeneous knowledge distillation that addresses representation discrepancies by calibrating feature decorrelation. It employs Confusion Evidence Estimation (CEE) to capture reliable semantic relations and Strength Allocation Control (SAC) to preserve discriminative structures during redundancy suppression. Additionally, Adaptive Coefficient Regulation (ACR) dynamically adjusts the loss weight to reduce sensitivity to hyperparameters. Experiments on CIFAR-100 and ImageNet-1K demonstrate improved performance over existing methods like RSD, OFA, and PAT.
Entities (18)
Relation Signals (14)
CoCaRS → evaluatedon → CIFAR-100
confidence 95% · Extensive experiments on CIFAR-100 ... validate the effectiveness of CoCaRS
CoCaRS → evaluatedon → ImageNet-1K
confidence 95% · Extensive experiments on ... ImageNet-1K validate the effectiveness of CoCaRS
CoCaRS → usesmodule → Confusion Evidence Estimation
confidence 95% · CoCaRS calibrates feature decorrelation through Confusion Evidence Estimation (CEE)
CoCaRS → usesmodule → Strength Allocation Control
confidence 95% · CoCaRS calibrates feature decorrelation through ... Strength Allocation Control (SAC)
CoCaRS → usesmodule → Adaptive Coefficient Regulation
confidence 95% · Adaptive Coefficient Regulation (ACR) further regulates the contribution...
Confusion Evidence Estimation → captures → Semantic Relations
confidence 90% · CEE ... capture reliable semantic relations for correlation estimation
CoCaRS → outperforms → RSD
confidence 90% · Table 1 shows CoCaRS achieving higher accuracy than RSD across various teacher-student pairs
CoCaRS → outperforms → OFA
confidence 90% · Table 1 shows CoCaRS achieving higher accuracy than OFA
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. However, differences in architectural inductive biases between the teacher and student models often result in substantial representation discrepancies, limiting the effectiveness of direct knowledge transfer. Recently, redundancy suppression has offered a new perspective on heterogeneous KD by preserving cross-architecture invariance and reducing feature redundancy through decorrelation of teacher-student feature correlations. Nevertheless, this formulation may weaken useful structural information through uniform decorrelation, while a fixed coefficient may make the effective contribution of redundancy suppression sensitive to teacher-student pairs and training stages. To address these problems, Correlation Calibration-based Redundancy Suppression (CoCaRS) is proposed to better retain structural information while suppressing redundancy and reduce sensitivity to coefficient settings across teacher-student pairs and training stages. Specifically, CoCaRS calibrates feature decorrelation through Confusion Evidence Estimation (CEE) and Strength Allocation Control (SAC), which respectively capture reliable semantic relations for correlation estimation and preserve discriminative structure during decorrelation. Adaptive Coefficient Regulation (ACR) further regulates the contribution of the calibrated redundancy suppression objective according to its relative loss scale, reducing sensitivity to coefficient settings. Extensive experiments on CIFAR-100 and ImageNet-1K validate the effectiveness of CoCaRS in improving distillation performance and reducing sensitivity to coefficient settings. Code will be released soon.
Tags
Links
- Source: https://arxiv.org/abs/2607.27054v1
- Canonical: https://arxiv.org/abs/2607.27054v1
Trouble viewing inline? Open PDF directly →
Full Text
48,456 characters extracted from source content.
Expand or collapse full text
CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation Fengming Yu Haiwei Pan Zhang Chunling Chen Jian Guan Baoying Ma Abstract Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. However, differences in architectural inductive biases between the teacher and student models often result in substantial representation discrepancies, limiting the effectiveness of direct knowledge transfer. Recently, redundancy suppression has offered a new perspective on heterogeneous KD by preserving cross-architecture invariance and reducing feature redundancy through decorrelation of teacher-student feature correlations. Nevertheless, this formulation may weaken useful structural information through uniform decorrelation, while a fixed coefficient may make the effective contribution of redundancy suppression sensitive to teacher-student pairs and training stages. To address these problems, Correlation Calibration-based Redundancy Suppression (CoCaRS) is proposed to better retain structural information while suppressing redundancy and reduce sensitivity to coefficient settings across teacher-student pairs and training stages. Specifically, CoCaRS calibrates feature decorrelation through Confusion Evidence Estimation (CEE) and Strength Allocation Control (SAC), which respectively capture reliable semantic relations for correlation estimation and preserve discriminative structure during decorrelation. Adaptive Coefficient Regulation (ACR) further regulates the contribution of the calibrated redundancy suppression objective according to its relative loss scale, reducing sensitivity to coefficient settings. Extensive experiments on CIFAR-100 and ImageNet-1K validate the effectiveness of CoCaRS in improving distillation performance and reducing sensitivity to coefficient settings. Code will be released soon. Introduction Recent progress in visual recognition has been largely driven by advanced neural architectures, such as CNNs (He et al. 2016; Sandler et al. 2018; Liu et al. 2022b), ViTs (Dosovitskiy et al. 2021; Touvron et al. 2021; Liu et al. 2021b) and MLP-based models (Tolstikhin et al. 2021; Touvron et al. 2023). While these models achieve strong performance, their computational and storage costs often hinder deployment in resource-constrained scenarios. Knowledge distillation (KD) (Hinton, Vinyals, and Dean 2015) has therefore become a widely used approach for model compression, where a lightweight student model is trained with guidance from a strong teacher to reduce inference cost with minimal performance degradation. Existing KD methods can be categorized into response-based, feature-based, and relation-based approaches (Gou, Yu, and Maybank 2021; Pan et al. 2026). Response-based methods transfer probability distributions via soft targets (Hinton, Vinyals, and Dean 2015; Yang et al. 2019; Son et al. 2021), while feature-based methods use intermediate representations as additional supervision (Romero et al. 2015; Heo et al. 2019b; Lin et al. 2022). Relation-based methods further distill structural information, such as relations among samples or feature representations (Yim et al. 2017; Park et al. 2019; Tian, Krishnan, and Isola 2020). These distillation methods have demonstrated their effectiveness in homogeneous settings. Nevertheless, restricting distillation to such homogeneous scenarios limits the flexibility of teacher selection, since a high-performing teacher with the same architecture as the student may not always be available in practice. For heterogeneous model pairs, differences in architectural inductive biases can lead to substantial discrepancies between teacher and student representations, reducing the compatibility of transferred knowledge and resulting in mismatched supervision and suboptimal performance (Raghu et al. 2021; Liu et al. 2022a; Hao et al. 2023). Recent heterogeneous KD studies have explored different ways to adapt teacher supervision across architectures. For example, OFA (Hao et al. 2023) projects intermediate representations into the logit space, FBT (Li et al. 2025) fuses heterogeneous features through an auxiliary model, and PAT (Lin et al. 2025) adapts teacher representations through feature prompting and region-aware attention. Different from these adaptation strategies, RSD (Zhang et al. 2025) formulates heterogeneous KD from the perspective of redundancy suppression. It constructs a feature correlation matrix between teacher and student representations, where the diagonal entries encourage invariance across architectures, while the off-diagonal entries are constrained by a feature decorrelation objective to suppress redundancy, with a fixed coefficient controlling the contribution of the resulting RSD term. However, RSD applies a uniform decorrelation constraint to feature correlations, even though some of them may also encode useful structural information. Consequently, such information may be weakened together with redundancy. Moreover, as the RSD term varies in scale relative to the task loss across teacher-student pairs and training stages, a fixed coefficient may yield varying effective contributions and make distillation performance sensitive to coefficient selection. To resolve these problems, Correlation Calibration-based Redundancy Suppression (CoCaRS) is proposed to calibrate feature decorrelation for better preservation of structural information encoded in feature correlations, while adaptively regulating the effective contribution of redundancy suppression to reduce sensitivity to coefficient selection. Specifically, CoCaRS introduces Semantic Correlation Calibration (SCC) to retain cross-architecture invariance while calibrating feature decorrelation through Confusion Evidence Estimation (CEE) and Strength Allocation Control (SAC). CEE derives confusion weights from teacher responses to capture reliable semantic relations for correlation estimation, whereas SAC constructs a semantic strength map from a discriminative subspace induced by the teacher classifier to preserve discriminative structure during decorrelation. Adaptive Coefficient Regulation (ACR) further regulates the contribution of SCC according to its loss scale relative to the task objective, thereby reducing sensitivity to coefficient settings. The main contributions of this work are summarized as follows: • CoCaRS is proposed to refine redundancy suppression for heterogeneous KD through correlation calibration, retaining structural information while adaptively regulating its effective contribution. • In SCC, CEE captures reliable semantic relations for correlation estimation, while SAC preserves discriminative structure during decorrelation. • ACR adaptively regulates the effective contribution of SCC to reduce sensitivity to coefficient settings. • Extensive experiments on CIFAR-100 and ImageNet-1K demonstrate the effectiveness of CoCaRS across diverse heterogeneous teacher-student pairs. Related Work Knowledge Distillation Knowledge distillation transfer knowledge from a teacher model to a smaller student model through soft labels (Hinton, Vinyals, and Dean 2015). Subsequent methods improve student learning with richer teacher supervision (Yang et al. 2019; Son et al. 2021; Lin et al. 2022; Lao et al. 2023; Tian, Krishnan, and Isola 2020; Xu et al. 2022). Knowledge distillation across heterogeneous architectures has also been explored. (Touvron et al. 2021; Ren et al. 2022; Liu et al. 2022a; Zhao, Song, and Liang 2023). These methods usually target specific architecture pairs or fixed transfer directions, limiting their applicability to diverse heterogeneous pairs. Recent studies therefore explore general frameworks for heterogeneous KD. OFA (Hao et al. 2023) projects intermediate representations into logits to alleviate the semantic mismatch. FBT (Li et al. 2025) integrates teacher and student representations through an auxiliary model. PAT (Lin et al. 2025) adapts teacher representations through prompt tuning and aligns student features through region-aware attention. In contrast, RSD (Zhang et al. 2025) formulates this problem from the perspective of redundancy suppression, preserving architecture invariant knowledge while suppressing redundant correlations. Our work follows this direction and further calibrates the decorrelation process. Semantic-Aware Redundancy Suppression Redundancy suppression has been widely studied in representation learning. Barlow Twins (Zbontar et al. 2021) reduces feature redundancy by driving the cross-correlation matrix between two augmented views toward the identity matrix, while VICReg (Bardes, Ponce, and LeCun 2022) regularizes covariance to reduce dependencies among embedding variables. RSD extends this principle to heterogeneous KD through a teacher-student correlation matrix, whose diagonal terms preserve cross-architecture invariance and off-diagonal terms suppress redundancy through decorrelation. Relational structures in teacher representations have also been explored in KD. SPKD (Tung and Mori 2019) preserves pairwise similarities between samples, C (Peng et al. 2019) transfers instance correlations, RKD (Park et al. 2019) distills distance and angle relations, and ICKD (Liu et al. 2021a) matches inter-channel correlations. These methods show that correlations in teacher representations can encode structural information, which is relevant to the decorrelation term in redundancy suppression. Beyond these relational structures, SemCKD (Chen et al. 2021a) addresses teacher-student semantic mismatch through adaptive cross-layer calibration. SimKD (Chen et al. 2022) reuses the teacher classifier to guide student feature learning. Neural Collapse (Papyan, Han, and Donoho 2020) reveals the geometric alignment between classifier weights and class means, and NCKD (Zhang, Song, and He 2025) exploits such class geometry for KD. These works motivate semantic calibration and classifier structure in redundancy suppression. Method Preliminaries Redundancy Suppression Distillation (RSD) introduces a redundancy suppression perspective for heterogeneous distillation. Given teacher representations t∈ℝB×Df^t ^B× D and student representations oris∈ℝB×dsf^s_ori ^B× d_s, an adaptor h(⋅)h(·) maps the student representations into the teacher feature space, yielding s∈ℝB×Df^s ^B× D. Here, B is the batch size, while D and dsd_s are the teacher and student feature dimensions. Based on tf^t and sf^s, RSD constructs a Pearson correlation matrix ∈ℝD×DP ^D× D between teacher and student feature units: ij=∑k=1B(fkit−f¯it)(fkjs−f¯js)∑k=1B(fkit−f¯it)2∑k=1B(fkjs−f¯js)2,P_ij= _k=1^B(f^t_ki- f^t_i)(f^s_kj- f^s_j) _k=1^B(f^t_ki- f^t_i)^2 _k=1^B(f^s_kj- f^s_j)^2, (1) where f¯ f denotes the mini-batch mean. Using DI_D as the target, diagonal entries are driven toward one to preserve cross-architecture invariance, whereas off-diagonal entries are driven toward zero to suppress feature redundancy. The RSD objective can be written as an off-diagonal reweighted mean-square error: ℒRSD=1D2[∑i=1D(1−ii)2+κ∑i≠jij2],L_RSD= 1D^2[ _i=1^D(1-P_i)^2+κ _i≠ jP_ij^2], (2) where κ controls the strength of the decorrelation objective. With β weighting the RSD term, the overall objective is ℒ=ℒCE+βℒRSD.L=L_CE+ _RSD. (3) Overall Framework CoCaRS revisits redundancy suppression from the perspective of correlation calibration, as shown in Fig. 1. Semantic Correlation Calibration (SCC) calibrates feature decorrelation in P while retaining cross-architecture invariance, rather than imposing a uniform constraint. Confusion Evidence Estimation (CEE) derives confusion weights from positive and reciprocal negative evidence to calibrate correlation estimation, while Strength Allocation Control (SAC) constructs a semantic strength map to allocate decorrelation strength. Adaptive Coefficient Regulation (ACR) further regulates the SCC contribution using its loss scale relative to the task objective, reducing sensitivity to coefficient settings. Overall, CoCaRS combines semantic calibration of feature decorrelation with adaptive regulation of the SCC contribution. Figure 1: Overview of the proposed CoCaRS framework. (a) CoCaRS refines redundancy suppression for heterogeneous distillation through CEE and SAC; (b) The resulting confusion weights and strength map jointly calibrate feature decorrelation in SCC; (c) CEE derives confusion weights from positive and reciprocal negative evidence in retrieved knowledge; (d) SAC constructs the strength map from a discriminative subspace induced by the teacher classifier through QR decomposition. Confusion Evidence Estimation In RSD, feature redundancy is suppressed by penalizing the off-diagonal entries of the teacher-student correlation matrix. CEE introduces sample awareness into correlation estimation by estimating semantic confusion evidence for each training sample. Samples with stronger semantic confusion reflected in teacher responses are treated as more informative for correlation estimation. In this way, sample-level importance is incorporated into correlation estimation in SCC through teacher dark knowledge aggregated over retrieved samples, while the original redundancy suppression objective is preserved. Non-target teacher responses are used as semantic cues for confusion evidence, since they encode dark knowledge about semantic relations beyond the ground-truth label (Zhao et al. 2022). Accordingly, given the training set =(xi,yi)i=1ND=\(x_i,y_i)\_i=1^N and a collection of pretrained teachers Tmm=1M\T_m\_m=1^M, a teacher response bank ℬB is constructed before distillation: ℬ=(i,yi,i,1t,…,i,Mt)i=1N,B=\ (k_i,y_i,z^t_i,1,…,z^t_i,M )\_i=1^N, (4) where i=ϕk(xi)k_i= _k(x_i) is the retrieval key extracted by a pretrained key encoder ϕk _k, and i,mt∈ℝCz^t_i,m ^C denotes the logits produced by teacher TmT_m. For each sample xix_i, a query embedding i=ϕk(xi)q_i= _k(x_i) is used to retrieve a knowledge set iK_i from ℬB. Let i+K_i^+ denote the positive knowledge set containing retrieved samples labeled yiy_i. Aggregating teacher logits over i+K_i^+ provides a local estimate of the non-target responses around xix_i. Given the normalized teacher logits ~j,mt z_j,m^t for each retrieved sample xjx_j, the positive evidence is formulated as i+=Ψyi+(1M|i+|∑m=1M∑xj∈i+~j,mt),e^+_i= _y_i^+( 1M|K^+_i| _m=1^M _x_j ^+_i z_j,m^t), (5) where Ψyi+(⋅) _y_i^+(·) masks the ground-truth entry and rescales the remaining non-target entries. The obtained i+∈ℝCe_i^+ ^C indicates which non-target classes are supported as potential semantic competitors to yiy_i in the local neighborhood of xix_i. To assess whether these semantic competitors reflect credible confusion with yiy_i, negative evidence is further estimated from i−K_i^-, which contains retrieved samples with labels different from yiy_i. For each non-target class c, let i,c−=xj∈i−∣yj=cK^-_i,c=\x_j ^-_i y_j=c\. The corresponding negative evidence entry evaluates the credibility of the semantic confusion between yiy_i and c by measuring the teacher responses to yiy_i for samples in i,c−K^-_i,c. Formally, it is defined as e~i,c−=1M|i,c−|∑m=1M∑xj∈i,c−z~j,m,yit,|i,c−|>0,0,|i,c−|=0. e^-_i,c= \ array[]lr 1M|K^-_i,c| _m=1^M _x_j ^-_i,c z_j,m,y_i^t,&|K^-_i,c|>0,\\ 0,&|K^-_i,c|=0. array . (6) The entries are assembled according to their class coordinates and calibrated as i−=Ψ−([e~i,1−,e~i,2−,…,e~i,C−]),e^-_i= ^-([ e^-_i,1, e^-_i,2,…, e^-_i,C]), (7) where Ψ−(⋅) ^-(·) denotes the calibration of the assembled negative evidence vector. The obtained i−∈ℝCe^-_i ^C serves as calibrated reciprocal evidence for assessing the credibility of the semantic confusion indicated by i+e^+_i. Given the positive and negative evidence defined above, the final confusion weight is obtained under an asymmetric evidence model. The positive evidence determines the support of possible non-target confusion, whereas the negative evidence calibrates the credibility of this support. This design reflects the assumption that reciprocal responses should reinforce an existing confusion pattern rather than create an independent supervision signal. Accordingly, the calibrated confusion weight is defined as iconf=i+⊙(1+γi−),w^conf_i=e^+_i (1+ ^-_i), (8) where γ controls the strength of reciprocal calibration. The resulting iconf∈ℝCw^conf_i ^C is a calibrated confusion weight vector that captures reliable semantic relations between yiy_i and its non-target competitors for subsequent correlation estimation. Strength Allocation Control SAC calibrates the off-diagonal decorrelation term through a semantic strength map that allocates decorrelation strength according to associations between feature dimensions. Neural collapse theory relates classifier weight vectors to class-level feature prototypes at convergence (Papyan, Han, and Donoho 2020), and NCKD (Zhang, Song, and He 2025) further exploits this theory for classifier construction in distillation. Motivated by this, SAC uses the teacher classifier weights to induce a discriminative subspace for constructing the semantic strength map. Given the normalized teacher classifier weights ~ W, an orthonormal basis Q for the discriminative subspace of the teacher classifier is obtained from ~⊤ W through reduced QR decomposition. When ~⊤ W has full row rank, rank reduction is applied before QR decomposition. The target rank is estimated from the stable rank, which measures the effective rank of a matrix. The semantic matrix is therefore defined as sem=|⊤|M^sem=|Q |, which captures the projection structure of the induced discriminative subspace over feature dimensions. Since the diagonal entries serve as invariance anchors rather than decorrelation terms, they are excluded to obtain the off-diagonal component off=sem⊙(1−)M^off=M^sem (1-I). The strength map is defined as ijκ=exp(−τκMijoff)1D(D−1)∑a≠bexp(−τκMaboff),i≠j,M^κ_ij= (- _κM^off_ij) 1D(D-1) _a≠ b (- _κM^off_ab),i≠ j, (9) where τκ _κ controls the semantic modulation strength. A larger ijoffM^off_ij indicates a stronger association between the corresponding feature dimensions within the induced discriminative subspace, thereby reducing the decorrelation strength applied to the corresponding feature correlation to preserve discriminative structure. Distillation Formulation Given the confusion weights and the semantic strength map defined above, SCC integrates them into the decorrelation objective while preserving the diagonal invariance term. Let ^t f^t and ^s f^s denote the normalized teacher and student features. The SCC objective is formulated as ℒSCC _SCC =1D2[∑i=1D(1−ii)2 = 1D^2 [ _i=1^D(1-P_i)^2 (10) +κ∑i≠jijκ(∑k=1Bkconf^k,itf^k,js)2]. +κ _i≠ jM^κ_ij ( _k=1^BS_k^conf f^t_k,i f^s_k,j )^2 ]. In this objective, κ controls the overall decorrelation strength, while the strength map κM^κ specifies its relative coefficients. kconfS_k^conf is derived from the confusion weight as kconf=Norm(1+α‖kconf‖1),S_k^conf=Norm (1+α\|w^conf_k\|_1 ), (11) where α controls the effect of confusion evidence on sample weighting in correlation estimation. A basic training objective combines the cross-entropy task loss with the SCC term as ℒ=ℒCE+λℒSCC,L=L_CE+ _SCC, (12) where λ controls the strength of the SCC term. However, the relative scale of the SCC term can vary across training stages and teacher-student pairs, making a static coefficient less suitable for maintaining a balanced optimization process. To address this, Adaptive Coefficient Regulation (ACR) regulates the effective contribution of SCC according to its relative loss scale with respect to the task objective: rt=ℒSCCtℒCEt+ϵ.r_t= L_SCC^tL_CE^t+ε. (13) The SCC coefficient is regulated according to the deviation of rtr_t from the target ratio ρ. To reduce fluctuations in loss magnitudes, the coefficient is updated through EMA as λt=ηλt−1+(1−η)(ρrt+ϵ), _t=η _t-1+(1-η)G ( ρr_t+ε ), (14) where η denotes the EMA decay factor, and (⋅)G(·) is a bounded modulation function. The final objective is defined as ℒ=ℒCE+λtℒSCC,L=L_CE+ _tL_SCC, (15) Table 1: Top-1 accuracy (%) on CIFAR-100. The best and second best results are in bold and underlined. Teacher Student From Scratch Feature-based Response-based Heterogeneous-KD T. S. FitNet C RKD CRD KD DKD DIST OFA PAT RSD CoCaRS CNN-based students Swin-T ResNet18 89.26 74.01 78.87 74.19 74.11 77.63 78.74 80.26 77.75 80.54 81.22 83.92 85.42 ViT-S ResNet18 92.04 74.01 77.71 74.26 73.72 76.60 77.26 78.10 76.49 80.15 80.11 81.50 85.22 Mixer-B/16 ResNet18 87.29 74.01 77.15 74.26 73.75 76.42 77.79 78.67 76.36 79.39 80.07 81.85 83.85 Swin-T MobileNetV2 89.26 73.68 74.28 71.19 69.00 79.80 74.68 71.07 72.89 80.98 78.78 83.68 85.50 ViT-S MobileNetV2 92.04 73.68 73.54 70.67 68.46 78.14 72.77 69.80 72.54 78.45 78.87 81.68 85.62 Mixer-B/16 MobileNetV2 87.29 73.68 73.78 70.73 68.95 78.15 73.33 70.20 73.26 78.78 78.62 81.74 84.46 ViT-based students ConvNeXt-T DeiT-T 88.41 68.00 60.78 68.01 69.79 65.94 72.99 74.60 73.55 75.76 79.59 82.46 84.09 Mixer-B/16 DeiT-T 87.29 68.00 71.05 68.13 69.89 65.35 71.36 73.44 71.67 73.90 74.66 78.50 81.54 ConvNeXt-T Swin-P 88.41 72.63 24.06 72.63 71.73 67.09 76.44 76.80 76.41 78.32 80.74 82.21 85.08 Mixer-B/16 Swin-P 87.29 72.63 75.20 73.32 70.82 67.03 75.93 76.39 75.85 76.65 78.44 81.28 84.05 MLP-based students ConvNeXt-T ResMLP-S12 88.41 66.56 45.47 67.70 65.82 63.35 72.25 73.22 71.93 75.21 83.50 84.21 86.63 Swin-T ResMLP-S12 89.26 66.56 63.12 68.37 64.66 61.72 71.89 72.82 11.05 73.58 80.94 82.67 84.99 Average 88.85 71.45 66.25 71.12 70.06 71.44 74.62 74.61 69.15 77.64 79.63 82.14 84.70 Experiments Experimental Setup Models Teacher-student pairs are constructed from models with different architectures. For CNN-based models, ResNet (He et al. 2016), MobileNetV2 (Sandler et al. 2018), and ConvNeXt (Liu et al. 2022b) are adopted. Transformer-based models cover ViT (Dosovitskiy et al. 2021), DeiT (Touvron et al. 2021), Swin Transformer (Liu et al. 2021b), and its lightweight variants, Swin-Pico and Swin-Nano (Hao et al. 2023). MLP-based models include MLP-Mixer (Tolstikhin et al. 2021) and ResMLP (Touvron et al. 2023). Datasets Experiments are conducted on CIFAR-100 (Krizhevsky, Hinton et al. 2009) and ImageNet-1K (Deng et al. 2009). CIFAR-100 contains 60,000 images from 100 classes, with 50,000 images used for training and 10,000 images used for testing. ImageNet-1K is a large-scale dataset containing approximately 1.28 million training images and 50,000 validation images from 1,000 classes. Baselines Several representative KD methods are selected for comparison. Feature-based methods include FitNet (Romero et al. 2015), C (Peng et al. 2019), RKD (Park et al. 2019), and CRD (Tian, Krishnan, and Isola 2020), which transfer knowledge through intermediate representations or feature relations. Response-based methods include KD (Hinton, Vinyals, and Dean 2015), DKD (Zhao et al. 2022), and DIST (Huang et al. 2022), which align the output responses of teacher and student models. Heterogeneous KD methods, including OFA (Hao et al. 2023), PAT (Lin et al. 2025), and RSD (Zhang et al. 2025), are also compared. Main Results Results on CIFAR-100 Experiments are conducted on CIFAR-100 using heterogeneous teacher-student pairs involving CNNs, ViTs, and MLPs. As reported in Table 1, CoCaRS achieves the best results across all evaluated pairs, with an average accuracy of 84.70%. Feature-based and response-based methods obtain average Top-1 accuracies of 69.72% and 72.79%, respectively, and the former is below the scratch baseline of 71.45%. These results indicate the limited applicability of homogeneous methods to heterogeneous teacher-student pairs, whereas methods designed for heterogeneous KD achieve higher average accuracies. RSD achieves an average accuracy of 82.14% through redundancy suppression, while CoCaRS raises the average accuracy to 84.70% by calibrating feature decorrelation, yielding a 2.56% gain. This result supports the benefit of calibrating feature decorrelation rather than applying a uniform constraint. Results on ImageNet-1K CoCaRS further achieves the highest average Top-1 accuracy of 74.46% on ImageNet-1K, outperforming RSD by 0.43% points, as reported in Table 2. Feature-based and response-based methods obtain lower average accuracies than methods developed for heterogeneous KD, reflecting the difficulty of transferring knowledge across heterogeneous architectures. The performance of RSD demonstrates the effectiveness of redundancy suppression in this setting, while the further improvement achieved by CoCaRS supports the benefit of correlation calibration beyond the original formulation. Notably, CoCaRS yields its largest improvement over RSD on Swin-T → ResMLP-S12, with a gain of 0.85%. These results further validate correlation calibration for large-scale heterogeneous distillation. Table 2: Top-1 accuracy (%) on ImageNet-1K. The best and second best results are in bold and underlined. Teacher Student From Scratch Feature-based Response-based Heterogeneous-KD T. S. FitNet C RKD CRD KD DKD DIST OFA PAT RSD CoCaRS CNN-based students Swin-T ResNet18 81.35 69.75 71.18 70.07 68.89 69.09 71.14 71.10 70.91 71.85 71.54 72.13 72.63 Mixer-B/16 MobileNetV2 76.55 68.87 71.59 70.79 69.86 68.89 71.92 70.93 71.74 72.12 72.22 71.90 72.15 ViT-based students ConvNeXt-T DeiT-T 82.05 72.17 70.45 73.12 71.47 69.18 74.00 73.95 74.07 74.41 74.44 74.46 74.59 MLP-based students Swin-T ResMLP-S12 81.35 76.65 76.48 76.15 75.10 73.40 76.67 76.99 77.25 77.31 77.59 77.61 78.46 Average 80.33 71.86 72.43 72.53 71.33 70.14 73.43 73.24 73.49 73.92 73.95 74.03 74.46 Ablation Study Effect of Core Components Table 3 shows that removing CEE, SAC, or ACR consistently degrades performance. The performance drops caused by removing CEE or SAC support calibrating both correlation estimation and decorrelation strength rather than applying uniform decorrelation. Meanwhile, the degradation caused by removing ACR supports the need to regulate the effective contribution of SCC according to its loss scale relative to the task objective. Together, these results validate semantic calibration in SCC and adaptive regulation of its contribution. Table 3: Ablation study on CIFAR-100: effect of the core components of CoCaRS. w/ CEE w/ SAC w/ ACR Swin-T ResNet18 Mixer-B/16 ResNet18 ViT-S MobileNetV2 ConvNeXt-T DeiT-T Mixer-B/16 Swin-P ConvNeXt-T ResMLP-S12 RSD Baseline 83.92 81.85 81.68 82.46 81.28 84.21 ✓ – – 84.38 82.45 83.72 82.95 83.04 85.20 – ✓ – 84.52 82.60 83.96 82.93 83.35 85.50 – ✓ ✓ 84.81 83.60 85.09 83.54 83.50 85.69 ✓ – ✓ 84.71 83.37 84.94 83.65 83.44 85.53 ✓ ✓ – 84.50 83.62 84.91 83.62 83.65 85.94 ✓ ✓ ✓ 85.42 83.85 85.62 84.09 84.05 86.63 Effect of CEE Formulation The construction of confusion evidence from retrieved knowledge and the reciprocal calibration provided by negative evidence are examined in Table 4. Removing CEE entirely reduces performance, confirming the contribution of confusion evidence to correlation estimation in SCC. Removing negative evidence (w/o NE) also leads to lower performance, supporting its contribution to calibrating the credibility of the semantic confusion indicated by positive evidence. Replacing the retrieved samples with randomly selected samples (Random) results in a clear performance decrease, suggesting that an effective local estimate of the non-target responses relies on retrieved samples associated with the query sample. These results support the use of retrieved knowledge for positive evidence estimation and negative evidence for reciprocal calibration. Table 4: Ablation study on CIFAR-100: effects of retrieved knowledge and negative evidence in CEE. CEE Setting Swin-T ResNet18 Mixer-B/16 ResNet18 ViT-S MobileNetV2 ConvNeXt-T DeiT-T Mixer-B/16 Swin-P ConvNeXt-T ResMLP-S12 w/o CEE 84.81 83.60 85.09 83.54 83.50 85.69 w/o NE 84.50 83.60 85.04 83.35 83.33 85.76 Random 84.60 83.22 84.64 83.20 82.55 86.05 CoCaRS 85.42 83.85 85.62 84.09 84.05 86.63 Effect of SAC Formulation The construction of the semantic strength map from the induced discriminative subspace and the direction of decorrelation strength allocation are examined in Table 5. Removing SAC entirely reduces performance, showing the contribution of the semantic strength map to decorrelation strength allocation. When the allocation direction is inverted (Inverted), larger values of offM^off lead to greater decorrelation strength, while the resulting performance remains close to that without SAC. This contrast supports allocating lower decorrelation strength to stronger associations in the induced discriminative subspace. Replacing the induced discriminative subspace with a random orthogonal subspace (Random) also performs worse than the complete formulation, supporting the use of teacher classifier weights as the structural basis of the semantic strength map. These results further validate the proposed direction of decorrelation strength allocation. Table 5: Ablation study on CIFAR-100: effects of strength allocation and discriminative subspace in SAC. SAC Setting Swin-T ResNet18 Mixer-B/16 ResNet18 ViT-S MobileNetV2 ConvNeXt-T DeiT-T Mixer-B/16 Swin-P ConvNeXt-T ResMLP-S12 w/o SAC 84.71 83.37 84.94 83.65 83.44 85.53 Inverted 84.68 83.31 85.03 83.24 83.58 85.73 Random 84.96 83.57 84.90 83.78 83.33 86.03 CoCaRS 85.42 83.85 85.62 84.09 84.05 86.63 Effect of Diagonal Invariance The role of diagonal invariance is examined in Table 6. Extending the modulation introduced by CEE, SAC, or both to the diagonal term consistently degrades performance. The diagonal term is intended to preserve cross-architecture invariance between heterogeneous representations. Applying confusion weights or the semantic strength map to this term may weaken its role in preserving cross-architecture invariance. These results support preserving diagonal invariance while restricting CEE and SAC to off-diagonal redundancy suppression. Table 6: Ablation study on CIFAR-100: effect of CEE and SAC on the invariance term of SCC. Invariance Term Swin-T Mixer-B/16 ViT-S ConvNeXt-T Mixer-B/16 ConvNeXt-T CEE SAC ResNet18 ResNet18 MobileNetV2 DeiT-T Swin-P ResMLP-S12 ✓ – 84.48 83.46 84.99 83.01 82.95 85.28 – ✓ 85.02 83.51 84.78 83.19 83.41 85.81 ✓ ✓ 84.81 83.11 84.60 82.60 83.40 85.94 – – 85.42 83.85 85.62 84.09 84.05 86.63 Table 7: Top-1 accuracy (%) on ImageNet-1K for homo. settings. The best and second best results are in bold and underlined. Teacher Student From Scratch Homogeneous-KD Heterogeneous-KD T. S. KD OFD CRD RKD CAT SimKD Review DKD SDD DIST OFA RSD CoCaRS ResNet34 ResNet18 73.31 69.75 70.66 70.81 71.17 71.34 71.26 71.59 71.61 71.70 71.14 72.07 72.10 72.18 72.58 ResNet50 MobileNet 80.36 68.58 68.58 71.25 71.37 71.32 72.24 72.25 72.56 72.05 72.24 73.24 73.28 73.08 74.91 Further Analysis Performance in Homogeneous Settings To further evaluate CoCaRS under homogeneous settings, experiments are conducted on two teacher-student pairs on ImageNet-1K. The comparison additionally includes OFD (Heo et al. 2019a), CAT (Guo et al. 2023), SimKD (Chen et al. 2022), Review (Chen et al. 2021b), and SDD (Wei, Luo, and Luo 2024). CoCaRS achieves the best performance and outperforms RSD on both pairs, as reported in Table 7. These results further support the effectiveness of introducing correlation calibration into redundancy suppression. Together with the results on heterogeneous teacher-student pairs, this additional evaluation demonstrates the robustness of CoCaRS across different distillation settings. Intermediate Representation Similarity The similarity between the intermediate representations of a Swin-T teacher and a ResMLP-S12 student is visualized using CKA (Kornblith et al. 2019) in Fig. 2. Without KD, the student exhibits relatively low feature similarity with the teacher, reflecting the representation discrepancy between heterogeneous architectures. Both RSD and CoCaRS improve the similarity between teacher and student intermediate features. Compared with RSD, CoCaRS exhibits higher similarity, particularly in the shallow and deep regions. The reduced representation discrepancy is consistent with the improved distillation performance and further supports the effectiveness of CoCaRS. Figure 2: Intermediate representation similarity between Swin-T and ResMLP-S12 measured by CKA, where brighter regions indicate higher similarity. Computational Cost A comparison of additional trainable parameters and peak memory usage across methods is presented in Fig. 3. CoCaRS requires fewer additional trainable parameters than OFA and PAT while matching that of RSD, since its proposed components introduce no additional learnable parameters. Its peak memory usage is comparable to that of OFA and substantially lower than that of PAT. Compared to RSD, CoCaRS achieves a better overall balance between distillation performance and training memory, despite its moderate increase in memory overhead. Since the additional components are used only during distillation, the original inference cost of the student is preserved. Figure 3: Comparison of additional trainable parameters and peak memory usage on CIFAR-100. ACR for SCC Balance As shown in Fig. 4(a), the relative scale of SCC to CE differs across the three heterogeneous model pairs and varies during training. Consequently, without ACR, the same fixed coefficient may produce different effective contributions of SCC, making it difficult to maintain a consistent balance with the task objective across model pairs and training stages. Figure 4(b) further shows that the SCC coefficient is regulated to different extents for the three model pairs according to their relative loss scales. This regulation maintains adaptive control over the contribution of SCC despite variations in loss scale, thereby reducing sensitivity to coefficient selection. Figure 4: ACR regulates the effective contribution of SCC across heterogeneous model pairs during training. Conclusion In this work, redundancy suppression in heterogeneous knowledge distillation was revisited by considering the limitation of uniform decorrelation, which may weaken useful structural information encoded in feature correlations. CoCaRS addresses this limitation through semantic calibration of feature decorrelation while preserving cross-architecture invariance, together with adaptive regulation of the calibrated objective during optimization. Experimental results across heterogeneous model pairs consistently support the effectiveness of this formulation. Overall, these findings demonstrate the effectiveness of semantic correlation calibration in improving redundancy suppression for heterogeneous knowledge distillation. References Bardes, Ponce, and LeCun (2022) Bardes, A.; Ponce, J.; and LeCun, Y. 2022. VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning. In The Tenth International Conference on Learning Representations, ICLR 2022. OpenReview.net. Chen et al. (2022) Chen, D.; Mei, J.; Zhang, H.; Wang, C.; Feng, Y.; and Chen, C. 2022. Knowledge Distillation with the Reused Teacher Classifier. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, 11923–11932. IEEE. Chen et al. (2021a) Chen, D.; Mei, J.; Zhang, Y.; Wang, C.; Wang, Z.; Feng, Y.; and Chen, C. 2021a. Cross-Layer Distillation with Semantic Calibration. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, 7028–7036. AAAI Press. Chen et al. (2021b) Chen, P.; Liu, S.; Zhao, H.; and Jia, J. 2021b. Distilling Knowledge via Knowledge Review. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, 5008–5017. Computer Vision Foundation / IEEE. Deng et al. (2009) Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, 248–255. Ieee. Dosovitskiy et al. (2021) Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In 9th International Conference on Learning Representations, ICLR 2021. OpenReview.net. Gou, Yu, and Maybank (2021) Gou, J.; Yu, B.; and Maybank, S. J. 2021. Knowledge distillation: A survey. International Journal of Computer Vision, 129(6): 1789–1819. Guo et al. (2023) Guo, Z.; Yan, H.; Li, H.; and Lin, X. 2023. Class Attention Transfer Based Knowledge Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, 11868–11877. IEEE. Hao et al. (2023) Hao, Z.; Guo, J.; Han, K.; Tang, Y.; Hu, H.; Wang, Y.; and Xu, C. 2023. One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023. He et al. (2016) He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, 770–778. IEEE Computer Society. Heo et al. (2019a) Heo, B.; Kim, J.; Yun, S.; Park, H.; Kwak, N.; and Choi, J. Y. 2019a. A Comprehensive Overhaul of Feature Distillation. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, 1921–1930. IEEE. Heo et al. (2019b) Heo, B.; Lee, M.; Yun, S.; and Choi, J. Y. 2019b. Knowledge Transfer via Distillation of Activation Boundaries Formed by Hidden Neurons. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, 3779–3787. AAAI Press. Hinton, Vinyals, and Dean (2015) Hinton, G. E.; Vinyals, O.; and Dean, J. 2015. Distilling the Knowledge in a Neural Network. CoRR, abs/1503.02531. Huang et al. (2022) Huang, T.; You, S.; Wang, F.; Qian, C.; and Xu, C. 2022. Knowledge Distillation from A Stronger Teacher. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022. Kornblith et al. (2019) Kornblith, S.; Norouzi, M.; Lee, H.; and Hinton, G. 2019. Similarity of Neural Network Representations Revisited. In Proceedings of the 36th International Conference on Machine Learning, volume 97, 3519–3529. PMLR. Krizhevsky, Hinton et al. (2009) Krizhevsky, A.; Hinton, G.; et al. 2009. Learning multiple layers of features from tiny images. Lao et al. (2023) Lao, S.; Song, G.; Liu, B.; Liu, Y.; and Yang, Y. 2023. Masked Autoencoders Are Stronger Knowledge Distillers. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, 6361–6370. IEEE. Li et al. (2025) Li, G.; Wang, Q.; Yan, K.; Ding, S.; Gao, Y.; and Xia, G. 2025. Fuse Before Transfer: Knowledge Fusion for Heterogeneous Distillation. In IEEE/CVF International Conference on Computer Vision, ICCV 2025, 3445–3454. IEEE. Lin et al. (2025) Lin, J.; Yao, Y.; Hsu, C.; Xie, H.; Shuai, H.; and Cheng, W. 2025. Perspective-Aware Teaching: Adapting Knowledge for Heterogeneous Distillation. In IEEE/CVF International Conference on Computer Vision, ICCV 2025, 4178–4187. IEEE. Lin et al. (2022) Lin, S.; Xie, H.; Wang, B.; Yu, K.; Chang, X.; Liang, X.; and Wang, G. 2022. Knowledge Distillation via the Target-aware Transformer. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, 10905–10914. IEEE. Liu et al. (2021a) Liu, L.; Huang, Q.; Lin, S.; Xie, H.; Wang, B.; Chang, X.; and Liang, X. 2021a. Exploring Inter-Channel Correlation for Diversity-preserved Knowledge Distillation. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, 8251–8260. IEEE. Liu et al. (2022a) Liu, Y.; Cao, J.; Li, B.; Hu, W.; Ding, J.; and Li, L. 2022a. Cross-Architecture Knowledge Distillation. In Computer Vision - ACCV 2022, volume 13845, 179–195. Springer. Liu et al. (2021b) Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021b. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, 9992–10002. IEEE. Liu et al. (2022b) Liu, Z.; Mao, H.; Wu, C.; Feichtenhofer, C.; Darrell, T.; and Xie, S. 2022b. A ConvNet for the 2020s. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, 11966–11976. IEEE. Pan et al. (2026) Pan, H.; Yu, F.; Zhang, K.; Lan, H.; Meng, Q.; and Li, Z. 2026. Knowledge Distillation in Visual Algorithms: A Survey. Journal of Computer Research and Development, 63(1): 90–122. Papyan, Han, and Donoho (2020) Papyan, V.; Han, X. Y.; and Donoho, D. L. 2020. Prevalence of Neural Collapse during the terminal phase of deep learning training. CoRR, abs/2008.08186. Park et al. (2019) Park, W.; Kim, D.; Lu, Y.; and Cho, M. 2019. Relational Knowledge Distillation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, 3967–3976. Computer Vision Foundation / IEEE. Peng et al. (2019) Peng, B.; Jin, X.; Li, D.; Zhou, S.; Wu, Y.; Liu, J.; Zhang, Z.; and Liu, Y. 2019. Correlation Congruence for Knowledge Distillation. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, 5006–5015. IEEE. Raghu et al. (2021) Raghu, M.; Unterthiner, T.; Kornblith, S.; Zhang, C.; and Dosovitskiy, A. 2021. Do Vision Transformers See Like Convolutional Neural Networks? In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, 12116–12128. Ren et al. (2022) Ren, S.; Gao, Z.; Hua, T.; Xue, Z.; Tian, Y.; He, S.; and Zhao, H. 2022. Co-advise: Cross Inductive Bias Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, 16752–16761. IEEE. Romero et al. (2015) Romero, A.; Ballas, N.; Kahou, S. E.; Chassang, A.; Gatta, C.; and Bengio, Y. 2015. FitNets: Hints for Thin Deep Nets. In 3rd International Conference on Learning Representations, ICLR 2015. Sandler et al. (2018) Sandler, M.; Howard, A. G.; Zhu, M.; Zhmoginov, A.; and Chen, L. 2018. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, 4510–4520. Computer Vision Foundation / IEEE Computer Society. Son et al. (2021) Son, W.; Na, J.; Choi, J.; and Hwang, W. 2021. Densely Guided Knowledge Distillation using Multiple Teacher Assistants. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, 9375–9384. IEEE. Tian, Krishnan, and Isola (2020) Tian, Y.; Krishnan, D.; and Isola, P. 2020. Contrastive Representation Distillation. In 8th International Conference on Learning Representations, ICLR 2020. OpenReview.net. Tolstikhin et al. (2021) Tolstikhin, I. O.; Houlsby, N.; Kolesnikov, A.; Beyer, L.; Zhai, X.; Unterthiner, T.; Yung, J.; Steiner, A.; Keysers, D.; Uszkoreit, J.; Lucic, M.; and Dosovitskiy, A. 2021. MLP-Mixer: An all-MLP Architecture for Vision. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, 24261–24272. Touvron et al. (2023) Touvron, H.; Bojanowski, P.; Caron, M.; Cord, M.; El-Nouby, A.; Grave, E.; Izacard, G.; Joulin, A.; Synnaeve, G.; Verbeek, J.; and Jégou, H. 2023. ResMLP: Feedforward Networks for Image Classification With Data-Efficient Training. IEEE Trans. Pattern Anal. Mach. Intell., 45(4): 5314–5321. Touvron et al. (2021) Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jégou, H. 2021. Training data-efficient image transformers & distillation through attention. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, volume 139, 10347–10357. PMLR. Tung and Mori (2019) Tung, F.; and Mori, G. 2019. Similarity-Preserving Knowledge Distillation. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, 1365–1374. IEEE. Wei, Luo, and Luo (2024) Wei, S.; Luo, C.; and Luo, Y. 2024. Scale Decoupled Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, 15975–15983. IEEE. Xu et al. (2022) Xu, H.; Fang, J.; Zhang, X.; Xie, L.; Wang, X.; Dai, W.; Xiong, H.; and Tian, Q. 2022. Bag of Instances Aggregation Boosts Self-supervised Distillation. In The Tenth International Conference on Learning Representations, ICLR 2022. OpenReview.net. Yang et al. (2019) Yang, C.; Xie, L.; Qiao, S.; and Yuille, A. L. 2019. Training Deep Neural Networks in Generations: A More Tolerant Teacher Educates Better Students. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, 5628–5635. AAAI Press. Yim et al. (2017) Yim, J.; Joo, D.; Bae, J.; and Kim, J. 2017. A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, 7130–7138. IEEE Computer Society. Zbontar et al. (2021) Zbontar, J.; Jing, L.; Misra, I.; LeCun, Y.; and Deny, S. 2021. Barlow Twins: Self-Supervised Learning via Redundancy Reduction. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, volume 139, 12310–12320. PMLR. Zhang, Song, and He (2025) Zhang, S.; Song, Z.; and He, K. 2025. Neural Collapse Inspired Knowledge Distillation. In Thirty-Ninth AAAI Conference on Artificial Intelligence, AAAI 2025, 22542–22550. AAAI Press. Zhang et al. (2025) Zhang, W.; Liu, Y.; Ran, W.; and Ma, C. 2025. Cross-Architecture Distillation Made Simple with Redundancy Suppression. In IEEE/CVF International Conference on Computer Vision, ICCV 2025, Honolulu, HI, USA, October 19-25, 2025, 23256–23266. IEEE. Zhao et al. (2022) Zhao, B.; Cui, Q.; Song, R.; Qiu, Y.; and Liang, J. 2022. Decoupled Knowledge Distillation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, 11943–11952. IEEE. Zhao, Song, and Liang (2023) Zhao, B.; Song, R.; and Liang, J. 2023. Cumulative Spatial Knowledge Distillation for Vision Transformers. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, 6123–6132. IEEE.