Paper deep dive
Test Case Prioritization for DNNs via Neural Collapse Instability
Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, Su-Juan Qin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/23/2026, 3:05:10 AM
Summary
The paper proposes Neural-Collapse-Inspired Prioritization (NCIP), a framework for test case prioritization in Deep Neural Networks (DNNs) that leverages Neural Collapse (NC) geometry. Instead of relying on single-checkpoint confidence, NCIP selects representative training checkpoints based on classifier weight equiangularity and prioritizes test inputs by their prediction variability (Total Variation Distance) across these checkpoints, combined with final model margin uncertainty. This approach effectively surfaces boundary-adjacent and failure-prone samples, achieving significant gains in early fault discovery compared to existing baselines.
Entities (9)
Relation Signals (6)
NCIP → uses → Neural Collapse
confidence 95% · NCIP is a Neural-Collapse-Inspired Prioritization framework
NCIP → measures → Total Variation Distance
confidence 92% · NCIP measures prediction variability... via Total Variation Distance (TVD)
Neural Collapse → characterizedby → Classifier Weight Equiangularity
confidence 90% · NC theory shows that... classifier weight vectors corresponding to different classes form an approximately equiangular configuration.
NCIP → selectscheckpointsusing → Classifier Weight Equiangularity
confidence 90% · selects an NC-guided representative subset of training checkpoints using an equiangularity score of classifier weights
NCIP → achievesimprovementin → RAUC-ALL
confidence 85% · NCIP achieves strong performance... with 1.5 to 16.6 percent RAUC-ALL gains
NCIP → achievesimprovementin → RAUC-500
confidence 85% · NCIP achieves strong performance... with 4.9 to 20.6 percent RAUC-500 gains
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limited testing budgets has become increasingly important. Existing test case prioritization techniques often rely on single-checkpoint confidence signals derived from output probabilities. However, DNNs can be confidently wrong, and the confidence margin between the predicted and competing classes is frequently small, which weakens early fault discovery. To address this limitation, we propose a Neural-Collapse-Inspired Prioritization (NCIP) framework that replaces absolute confidence with cross-checkpoint prediction variability in the terminal training regime, where model geometry becomes highly structured. NCIP introduces two key components. First, it selects an NC-guided representative subset of training checkpoints using an equiangularity score of classifier weights, quantified as the standard deviation of pairwise cosine similarities among class weight vectors. Second, it prioritizes test inputs by their prediction variability across the selected checkpoints, surfacing boundary-adjacent and failure-prone samples that are unstable under checkpoint-induced decision boundary shifts. Extensive experiments across multiple datasets and architectures show that NCIP achieves strong performance in early fault discovery compared with competitive baselines, with 1.5 to 16.6 percent RAUC-ALL gains and 4.9 to 20.6 percent RAUC-500 gains under the same testing budget. NCIP further attains the best average performance across all dataset-model pairs.
Tags
Links
- Source: https://arxiv.org/abs/2607.20046v1
- Canonical: https://arxiv.org/abs/2607.20046v1
Trouble viewing inline? Open PDF directly →
Full Text
79,265 characters extracted from source content.
Expand or collapse full text
Test Case Prioritization for DNNs via Neural Collapse Instability CHUNYU LIU, State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, China and National Engineering Research Center of Disaster Backup and Recovery, Beijing University of Posts and Telecommunications, China MINGYUAN LI, Information Security Center, Beijing University of Posts and Telecommunications, China and State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecom- munications, China YANG LI, State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, China WENMIN LI ∗ , State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, China FEI GAO, State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, China and National Engineering Research Center of Disaster Backup and Recovery, Beijing University of Posts and Telecommunications, China TENGFEI TU, State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, China SU-JUAN QIN, State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, China With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limited testing budgets has become increasingly important. Existing test case prioritization techniques often rely on single-checkpoint confidence signals derived from output probabilities. However, DNNs can be confidently wrong, and the confidence margin between the predicted and competing classes is frequently small, which weakens early fault discovery. To address this limitation, we propose a Neural- Collapse-Inspired Prioritization (NCIP) framework that replaces absolute confidence with cross-checkpoint prediction variability in the terminal training regime, where model geometry becomes highly structured. NCIP ∗ Corresponding author. Authors’ Contact Information: Chunyu Liu, State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China and National Engineering Research Center of Disaster Backup and Recovery, Beijing University of Posts and Telecommunications, Beijing, China, chunyuliu@bupt.edu.cn; Mingyuan Li, Information Security Center, Beijing University of Posts and Telecommunications, Beijing, China and State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China, henryli_ i@bupt.edu.cn; Yang Li, State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China, liyang02@bupt.edu.cn; Wenmin Li, State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China, liwenmin@bupt.edu.cn; Fei Gao, State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China and National Engineering Research Center of Disaster Backup and Recovery, Beijing University of Posts and Telecommunications, Beijing, China, gaof@bupt.edu.cn; Tengfei Tu, State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China, tutengfei.kevin@bupt.edu.cn; Su-Juan Qin, State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China, qsujuan@bupt.edu.cn. This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. © 2026 Copyright held by the owner/author(s). ACM 2994-970X/2026/10-ARTISSTA092 https://doi.org/10.1145/3832183 Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. arXiv:2607.20046v1 [cs.LG] 22 Jul 2026 ISSTA092:2Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, and Su-Juan Qin introduces two key components. First, it selects an NC-guided representative subset of training checkpoints using an equiangularity score of classifier weights, quantified as the standard deviation of pairwise cosine similarities among class weight vectors. Second, it prioritizes test inputs by their prediction variability across the selected checkpoints, surfacing boundary-adjacent and failure-prone samples that are unstable under checkpoint-induced decision boundary shifts. Extensive experiments across multiple datasets and architectures show that NCIP achieves strong performance in early fault discovery compared with competitive baselines, with 1.5 %–16.6 % RAUC-ALL gains and 4.9 %–20.6 % RAUC-500 gains under the same testing budget. NCIP further attains the best average performance across all dataset-model pairs. CCS Concepts:• Software and its engineering→ Software testing and debugging. Additional Key Words and Phrases: Test Case Prioritization, Deep Learning, Neural Collapse ACM Reference Format: Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, and Su-Juan Qin. 2026. Test Case Prioritization for DNNs via Neural Collapse Instability. Proc. ACM Softw. Eng. 3, ISSTA, Article ISSTA092 (October 2026), 24 pages. https://doi.org/10.1145/3832183 1 Introduction Deep neural networks (DNNs) have revolutionized various industries by excelling at complex pattern recognition and decision-making tasks. However, as these technologies are increasingly adopted in safety-critical domains such as autonomous driving and medical diagnostics, their inherent vulnerabilities have become more apparent. The severe consequences of potential failures in these applications, ranging from fatal traffic accidents [48] to life-threatening medical misdiagnoses, highlight the urgent need for robust test frameworks. As testing budgets are typically limited, an effective strategy is test case prioritization, which ranks a candidate test set so that faulty inputs are revealed as early as possible. A large body of work has studied prioritization criteria for DNN testing [21,35,46], including coverage-driven heuristics [36, 57], uncertainty-based scores [4, 9, 52], and learning-to-rank measures [8, 46, 53]. However, most test prioritization methods assume that confidence-based metrics can reliably separate correct from incorrect predictions, with correct predictions typically assigned higher confidence, whereas mispredictions exhibit low confidence. Recent studies [4,7] have examined this assumption, showing that DNNs can be confidently wrong and that the confidence gap between the predicted class and competing classes is often marginal even for correct cases. This confidence and accuracy mismatch affects prioritization: high-confidence errors may be ranked late, while low-confidence yet correct inputs may be prioritized excessively, consuming limited testing budgets. In this work, we revisit DNN test prioritization from the perspective of terminal training geometry. Recent studies show that many classifiers exhibit Neural Collapse (NC) in late training. In this phase, class features and classifier weights become highly symmetric, and the decision rule approaches a nearest class center classifier. This suggests that training checkpoints are not arbitrary. They can share geometric properties that help characterize model behavior. Motivated by this insight, we propose an NC-inspired prioritization framework (NCIP). NCIP first selects an NC-guided representative subset of checkpoints using classifier-weight equiangularity. It then uses prediction variability across selected checkpoints to improve ranking reliability. NCIP replaces single-checkpoint confidence with cross-checkpoint prediction variability. It first selects an NC-guided representative subset of late checkpoints, then measures prediction variability across this subset. Truly easy inputs tend to remain stable across selected checkpoints, while failure- prone inputs are more likely to show prediction changes under checkpoint-induced boundary shifts (see Fig. 1). Thus, cross-checkpoint prediction variability better reflects boundary proximity and ambiguity, reducing over-prioritization of low-confidence but correct samples. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. Test Case Prioritization for DNNs via Neural Collapse InstabilityISSTA092:3 −75−50−250255075 Dimension 1 −80 −60 −40 −20 0 20 40 60 80 Dimension 2 Wrong True (a) t-SNE of ResNet-18 output vectors on CIFAR10 using the final checkpoint. −75−50−250255075 Dimension 1 −80 −60 −40 −20 0 20 40 60 80 Dimension 2 Wrong True (b) t-SNE using output vectors aggregated over NC- guided checkpoints Fig. 1. Motivation for cross-checkpoint prediction variability on CIFAR10. Red points denote misclassified samples and white points denote correctly classified samples. Compared with a single final checkpoint, aggregating predictions over NC-guided checkpoints improves the separability of failure-prone inputs. We evaluate NCIP on multiple datasets and testing settings. Across these scenarios, NCIP achieves strong performance in early fault discovery, achieving higher RAUC than competitive baselines. In summary, this paper makes the following contributions: •We connect DNN test case prioritization with terminal phase training geometry and show how NC can inform reliable ranking criteria. • We propose NCIP, which selects NC-guided checkpoints using classifier weight equiangularity and prioritizes tests by cross-checkpoint prediction variability. •We conduct extensive experiments demonstrating that NCIP improves early fault discovery and remains effective across fault types, while offering stable and interpretable prioritization signals. 2 Preliminaries 2.1 Deep Neural Networks DNNs are artificial neural networks characterized by multiple hidden layers, designed to tackle complex pattern recognition tasks by emulating the structure and function of the human brain. They consist of neurons organized hierarchically into input layers, hidden layers, and output layers. Each neuron receives input from the previous layer, computes a weighted sum, applies an activation function, and produces an output for the next layer. Mathematically, for the푙-th layer, with input푥 (푙−1) , output푥 (푙) , weight matrix푊 (푙) , bias vector 푏 (푙) , and activation function 푔, the forward propagation can be represented by Equation 1: 푥 (푙) =푔(푧 (푙) )=푔(푊 (푙) 푥 (푙−1) +푏 (푙) )(1) where 푧 (푙) is the result of the linear transformation, and 푥 (푙) is the activation value. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. ISSTA092:4Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, and Su-Juan Qin 2.2 Test Case Prioritization A large body of work has investigated the testing of DNN-based systems [13,14,16,36,43,44,58,66]. Since such systems are developed under a data-driven paradigm [3,41], a central challenge is to expose model failures effectively under limited testing cost. Test case prioritization aims to order a candidate test set so that faults are revealed as early as possible within a constrained budget. For a classification model, a fault occurs when the model misclassifies an input. Formally, given a target model푀, a candidate test set푇, and a budget퐵 (퐵 ≪ |푇|), the goal is to rank푇and select the top-퐵subset푇 푆 such that푇 푆 maximizes the number of faults exposed in푀. Test prioritization has been widely studied in DNN testing as an effective means to accelerate fault discovery under limited time and computation [2, 4, 9, 21, 33, 35, 46, 51, 65]. 2.3 Neural Collapse Neural Collapse (NC) refers to a set of geometric regularities that emerge in the terminal phase of training for classifiers, typically after the training error becomes (near) zero while the loss keeps decreasing [42,67]. Consider a퐶-class model with feature extractorℎ(·) ∈ R 푑 and linear classifier 푊 ∈ R 퐶×푑 (rows푤 푐 퐶 푐=1 ). Let 퐷 푐 be the training samples of class 푐 and define 휇 푐 = 1 |퐷 푐 | ∑︁ 푥 푖 ∈퐷 푐 ℎ(푥 푖 ), 휇 퐺 = 1 퐶 퐶 ∑︁ 푐=1 휇 푐 .(2) NC is commonly summarized by four coupled phenomena [42]: • (NC1) Variability collapse. As training progresses, within-class feature variation becomes negligible, and features concentrate around their corresponding class means. •(NC2) Simplex ETF geometry. The centered class means ̃ 휇 푐 = 휇 푐 −휇 퐺 converge to a simplex equiangular tight frame: they have equal norm, form equal angles between any pair, and achieve the maximally pairwise separated configuration under these constraints. •(NC3) Self-duality. The classifier weights and the centered class means converge to each other up to a shared rescaling (and appropriate centering), yielding a highly symmetric decision structure with no systematic preference for confusions between particular class pairs. •(NC4) Nearest class center behavior. The induced decision rule simplifies to choosing the class whose mean is nearest to the input feature in Euclidean distance. Intuitively, NC means that class directions become more evenly spread in late training. Since NC3 links classifier weights to class means, we use classifier-head weights as a practical proxy for this geometry. We then track the spread of pairwise cosine similarities among class-weight vectors. If directions are evenly arranged, cosine values are close and the spread is low, whereas if some directions are unusually close or far apart, cosine values spread out and the spread increases. This gives a simple checkpoint-level indicator of NC regularity. 3 Methodology 3.1 Problem Definition We consider unlabeled test case prioritization for a퐶-class classifier푓 휃 . Given an unlabeled test set 푋=푥 푖 푁 푖=1 , 푥 푖 ∈X,(3) Our goal is to find a ranking function휋:[푁] → [푁], yielding the ordered test sequence ⟨푥 휋(1) , . . .,푥 휋(푁) ⟩. A fault occurs when the model prediction differs from the ground-truth label푦 푖 , i.e., Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. Test Case Prioritization for DNNs via Neural Collapse InstabilityISSTA092:5 Uncertainty measure 3 푥 ,푥 ...푥 Prioritized Test Set 4 Input Test Set 2 푀 푚 ...... Model Checkpoints 푚 푚 푚 푚 푚 ...... 푚 Select Model Checkpoints 1 휎 =푠푡푑 ( 푤 푤 ∥푤 ∥푤 ∥ ) 푥 ...... Test Dataset 푥 푥 Margin TVD(Total Variation Distance) 푥 Rank 1 Rank 2 Rank 3 푀푎푟푔푖푛푥 =푃 푦 푥 −푃 (푦 |푥 ) 푇푉퐷푥 = 1 2|푆| ∥푝 푥 −푝 (푥 )∥ ∈ Fig. 2. Overview of NCIP. (1) NC checkpoint selection via weight equiangularity to obtain NC-guided model states; (2) cross-checkpoint prediction variability measurement to quantify instability for each test input; and (3) boundary uncertainty integration using the final model margin to refine the prioritization score. I 푖 = I [ 푓 휃 (푥 푖 )≠ 푦 푖 ] ,(4) whereI[·]is the indicator function. Since labels are unavailable during prioritization,휋must be computed without access to푦 푖 푁 푖=1 . The effectiveness of휋is evaluated by how quickly faults appear in the ranked list under a limited inspection budget 퐵 ≪ 푁 . 3.2 Overview of NCIP Framework Our approach, Neural-Collapse-Inspired-Prioritization (NCIP), prioritizes test samples by mea- suring their inconsistency with NC geometry in the training convergence phase, as shown in Fig. 2. Our framework follows a three-stage pipeline. First, we perform NC checkpoint selection by extracting training checkpoints and selecting a representative subset based on classifier-weight equiangularity, then use cross-checkpoint prediction variability across this subset for prioritization. Second, we measure cross-checkpoint prediction variability by quantifying, for each test input, the instability of its predicted probability distribution across the selected checkpoints via Total Variation Distance (TVD), which captures sensitivity to checkpoint-induced decision boundary shifts. Third, we apply boundary uncertainty integration by incorporating the final checkpoint’s prediction margin to emphasize samples close to the decision boundary, where small perturbations are more likely to flip the predicted label. The resulting prioritization score ranks inputs that remain unstable across NC-guided checkpoints as higher risk test cases. 3.3 NC Checkpoint Selection via Weight Equiangularity NC theory shows that, in the late training stage, the classifier weight vectors corresponding to different classes form an approximately equiangular configuration. This geometric property provides a robust indicator of whether a model has entered a stable NC regime. Let푊=[푤 1 , ...,푤 퐶 ] denote the classifier weight vectors. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. ISSTA092:6Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, and Su-Juan Qin To evaluate the equiangularity of the classifier weights, we measure the variability of the angles between weight vectors. Specifically, we compute the standard deviation of the pairwise cosine similarities between all distinct pairs of class weight vectors: 휎 푐표푠 = std 푗≠푘 ( 푤 ⊤ 푗 푤 푘 ∥푤 푗 ∥푤 푘 ∥ )(5) A smaller value indicates stronger equiangularity and thus closer adherence to NC geometry. For a set of candidate checkpoints, we compute an NC score for each checkpoint from classifier- head weight geometry using Eq. 5. We first include the final checkpoint in the selected set, then iteratively add the checkpoint whose NC score has the largest minimum distance to the current set. This yields representatives that remain terminally consistent while covering diverse NC states. In our implementation, we use the last 90 % of available checkpoints as the candidate pool and set the target size to퐾=30 (e.g., if training has 100 total epochs, the candidate pool contains 90 checkpoints, epochs 11–100), as empirically supported in Section 5.4. After selection, we build and cache a single model incorporating all selected checkpoint weights. In later runs, we directly load this cache instead of rebuilding it, reducing repeated model-loading overhead. Algorithm 1 summarizes the implemented NC checkpoint selection procedure. Algorithm 1 NC Checkpoint Selection via Weight Equiangularity Input: Checkpoint setM=M 1 , . . .,M 푚 , collapse metric퐶(·), target checkpoint count 퐾 Output: selected checkpoint subsetS 1: Compute 푐 푗 ← 퐶(M 푗 ) for all 푗 ∈ 1, . . .,푚⊲ NC score per checkpoint 2: Initialize푇 ← 푚 3: S ←푇⊲ Keep last checkpoint 4: while|S|< 퐾 do 5:For each푢∉S, compute 푑(푢,S)= min 푠∈S |푐 푢 −푐 푠 | 6: 푢 ∗ ← arg max 푢∉S 푑(푢,S) 7: S ←S∪푢 ∗ 8: end while 9: SortS by epoch index in ascending order 10: returnS 3.4 NC-Guided Prediction Variability Given the selected NC-guided checkpoints, we quantify prediction instability for each test input. Let푝 (푚) (푥 푖 )denote the predicted class probability distribution of sample푥 푖 under checkpoint푚, let 푝 (푇) (푥 푖 )denote the prediction of the final checkpoint, and let푆denote the NC-guided checkpoint subset. We define the Total Variation Distance (TVD) as: TVD(푥 푖 )= 1 2|S| ∑︁ 푚∈S 퐶 ∑︁ 푐=1 푝 (푚) 푐 (푥 푖 )− 푝 (푇) 푐 (푥 푖 ) .(6) For samples consistent with NC geometry, predictions remain stable across NC-guided check- points, resulting in small TVD values. In contrast, misclassified or ambiguous samples tend to lie near regions where class attraction competes, causing prediction distributions to fluctuate even among stable models. Fig. 3 shows that misclassified samples consistently exhibit larger TVD-based variability than correctly classified ones across late training checkpoints on both CIFAR10-ResNet- 18 and MNIST-LeNet-5, supporting TVD as an effective fault-oriented ranking signal. Thus, TVD Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. Test Case Prioritization for DNNs via Neural Collapse InstabilityISSTA092:7 020406080100 Number of Models (Checkpoints) 0.0 0.1 0.2 0.3 0.4 Mean TVD MNIST - LeNet5 Correct Samples (Mean) Incorrect Samples (Mean) (a) MNIST-LeNet-5: mean TVD (and variability) for correct vs. incorrect samples across checkpoints. 050100150200250300 Number of Models (Checkpoints) 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 Mean TVD CIFAR10 - ResNet18 Correct Samples (Mean) Incorrect Samples (Mean) (b) CIFAR10-ResNet-18: mean TVD (and variability) for correct vs. incorrect samples across checkpoints. Fig. 3. Evolution of prediction instability measured by TVD across training checkpoints. In both settings, incorrect samples maintain substantially higher TVD than correct ones as the number of considered check- points increases, suggesting that mispredictions can be effectively surfaced by variability-based prioritization. serves as a direct, sample-level observable of NC instability rather than a heuristic uncertainty measure; we prioritize samples with higher TVD values, which indicate greater instability and a higher risk of misclassification. While TVD captures global prediction instability, it does not explicitly measure the uncertainty of the target model itself (i.e., the final epoch checkpoint). To complement this, we introduce a boundary uncertainty measure based on the final model’s prediction margin. For each test sample푥 푖 , let푝 (1) (푥 푖 )and푝 (2) (푥 푖 )denote the largest and second largest predicted probabilities, respectively. The margin is defined as follows: Margin(푥 푖 )= 푝 (1) (푥 푖 )− 푝 (2) (푥 푖 )(7) A smaller margin implies that the input lies closer to the decision boundary and is therefore more susceptible to prediction flips. This quantity is inexpensive to compute at a single checkpoint and provides a boundary proximity signal complementary to the cross-checkpoint variability captured by TVD. 3.5 Final Error Prioritization Score We define the final prioritization score of NCIP, which prioritizes unlabeled test inputs by combining the global instability signal (TVD) and the local boundary signal (margin), for each test input푥 푖 as NCIP(푥 푖 )= zscore(TVD(푥 푖 ))+ zscore(1− Margin(푥 푖 ))(8) Here, zscore(·) denotes the standardization process. The test samples are then ranked in descending order of 푁퐶퐼푃(푥 푖 ), with higher-ranked samples prioritized for testing. This formulation requires no training dataset. By integrating training dynamics and decision boundary information, it provides an effective and robust solution for unlabeled error prioritization. Algorithm 2 outlines our algorithm. 4 Experiment Design 4.1 Datasets and Models To ensure a fair comparison, we used datasets and models consistent with previous studies [4, 35, 46,52], as detailed in Table 1. We employed seven well-known image and text classification datasets: Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. ISSTA092:8Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, and Su-Juan Qin Table 1. Dataset and DNN models DatasetDescriptionEpochsDNN ModelAcc MNIST28×28 handwritten digits100 LeNet-1 LeNet-5 0.9887 0.9918 Fashion- MNIST 28×28 grayscale images200 LeNet-1 LeNet-5 0.8917 0.9021 CIFAR1032×32 colored images300 ResNet-18 VGG-11 0.9554 0.8936 CIFAR100 32×32 colored images with 100 classes 500 ResNet-50 DenseNet-121 0.7793 0.7888 TinyImageNet High-resolution images with 200 classes 500 ResNet-152 DenseNet-201 0.5484 0.5814 IMDB Movie reviews for sentiment analysis (binary) 100 TextCNN Transformer 0.8030 0.7546 AGNEWSNews topic classification (4 classes)100 TextCNN Transformer 0.9154 0.9030 Algorithm 2 NC-guided Error Prioritization via NCIP Input: Selected checkpointsS, test set푇=푥 푖 푁 푖=1 Output: Ranking indices 휋 , TVD scores u 푡푣푑 ∈ R 푁 , margins m∈ R 푁 1: Initialize u 푡푣푑 ← 0⊲ TVD per sample 2: Initialize m← 0⊲ Margin per sample 3: Choose reference checkpoint 푟 ← max(S)⊲ Use last selected epoch as reference 4: for each sample 푥 푖 ∈ 푇 do 5: 푝 (푟) ← softmax(M 푟 (푥 푖 ))⊲ Reference distribution 6: for each checkpoint 푗 ∈S, 푗≠ 푟 do 7:푝 (푗) ← softmax(M 푗 (푥 푖 )) 8:푑 푖푗 ← 1 2 ∥푝 (푗) − 푝 (푟) ∥ 1 ⊲ Total variation distance 9:u 푡푣푑 푖 ← u 푡푣푑 푖 +푑 푖푗 10: end for 11: (푝 (1) ,푝 (2) ) ← Top2(푝 (푟) )⊲ Largest and second largest probs 12:m 푖 ← 푝 (1) − 푝 (2) ⊲ Prediction margin 13: end for 14: s← zscore(u 푡푣푑 )+ zscore(1− m)⊲ Final risk score 15: 휋 ← argsort(s, descending) 16: return 휋, u 푡푣푑 , m MNIST [31], Fashion-MNIST [56], CIFAR10 [28], CIFAR100 [28], TinyImageNet [30], IMDB [37] and AGNEWS [63]. The experiments involved widely used DNN architectures: LeNet-1 [31] for MNIST, LeNet-5 [31] for Fashion-MNIST, ResNet-18 [15] and VGG-11 [47] for CIFAR10, ResNet-50 [15] and DenseNet-121 [20] for CIFAR100, and ResNet-152 [15] and DenseNet-201 [20] for TinyImageNet. For the text classification datasets, we adopted both CNN-based [27] and Transformer-based [50] models. For LLM generative-task evaluation, we additionally use the WikiText dataset [40] with two fine-tuned causal language models, GPT-2 [45] and OPT-125m [61]. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. Test Case Prioritization for DNNs via Neural Collapse InstabilityISSTA092:9 Table 2. Adversarial attack configurations Attack NormPerturbation BudgetTargetedOther Settings FGSM 퐿 ∞ 휖= 8/255UntargetedSingle-step PGD 퐿 ∞ 휖= 8/255Untargeted 훼= 2/255, steps=10, random start BIM 퐿 ∞ 휖= 8/255Untargeted훼= 2/255, steps=10 CW퐿 2 Optimization-based 퐿 2 objective Untargeted 푐= 1,휅= 0, steps=50, lr=0.01 4.2 Adversarial Data Adversarial data consists of deliberately altered inputs with subtle perturbations that lead deep learning models to make incorrect predictions despite these changes being nearly imperceptible to humans. To assess the effectiveness of our proposed DNN testing method in detecting errors with adversarial samples, we included four common attack methods: FGSM [11], PGD [39], CW [6], and BIM [29], implemented using the torchattack library [23] with default hyperparameters. Table 2 summarizes the attack settings used in adversarial evaluation. All attacks are run with standard default settings and in untargeted mode. For each dataset-model pair, we generate ad- versarial examples for FGSM, PGD, BIM, and CW separately, and then combine them into one adversarial test set. We report RAUC-500 and RAUC-ALL in Section 5.1 on this combined set. 4.3 Baseline Methods We selected several classic methods, including uncertainty-based methods: DeepGini [9], Entropy [5], PCS [62], and MSP [55], as well as SA-based methods: DSA and LSA [24], implemented via a third-party library. To differentiate our method from approaches that estimate disagreement from a single checkpoint, we include Dropout [18] and EffiMAP [54] as baselines. Dropout estimates uncertainty from repeated stochastic forward passes and uses it as a proxy for the distance to the decision boundary. EffiMAP perturbs both model and input, measures the resulting prediction changes, and prioritizes samples with larger changes. In addition, we incorporate several representative baselines for comparison. NNS [21] considers not only the uncertainty of a DNN on a given test input but also the model’s uncertainty over its neighboring samples. TDPR [46] constructs a learning trajectory for each test input to characterize the evolving learning dynamics of DNNs across training, from which informative features are extracted and fed into a learning-to-rank framework to derive a prioritized ordering. SETS [52] focuses on high uncertainty test inputs to first shrink the candidate pool and then applies an efficient greedy strategy to further reduce the number of fitness evaluations. 4.4 Evaluation Metrics We assessed the prioritization performance across four key dimensions: effectiveness, robustness, diversity, and guidance. Effectiveness and Robustness. To assess the prioritization ranking’s effectiveness and robust- ness, we employed the Ratio of the Area Under Curve (RAUC) [46,65] as a performance metric for evaluating test sample selection methods. Here, robustness refers explicitly to the method’s effectiveness when the input samples are replaced by adversarial data. For classification tasks, the RAUC is mathematically defined as follows: Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. ISSTA092:10Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, and Su-Juan Qin 푅퐴푈퐶= Í 푁 푖=1 푛 푖 푁 × 푁 ′ + 푁 ′ −푁 ′2 2 , where 푛 푖 = 푛 푖−1 + 1,푐(푥 푖 ) is incorrect 푛 푖−1 , otherwise (9) Here,푁is the total number of test inputs, and푁 ′ represents the number of misclassified samples. The term푐(푥 푖 )indicates the model’s classification prediction for input푥 푖 . This metric prioritizes test inputs that are more likely to be misclassified, with Í 푁 푖=1 푛 푖 denoting the sum of indices of misclassified samples in the ranked results. Given the high labeling cost of test inputs and resource constraints, it is essential to identify erroneous samples quickly. To address this, we introduce RAUC-n (RAUC for the top푛prioritized test inputs) to measure prioritization effectiveness under limited budget scenarios. Specifically, we evaluate RAUC-500 and RAUC-ALL. Diversity. While the above metrics evaluate the effectiveness of test methods in identifying erroneous samples, we also employ the퐹푎푢푙푡 _푇푦푝푒metric [10,64] to assess the diversity of detected faults. If misclassified samples are ranked highly but correspond to the same fault type, the prioritization method exhibits limited fault diversity and is therefore less effective. If a test sample 푥 is misclassified as another label, 퐹푎푢푙푡 _푇푦푝푒 is defined as follows: 퐹푎푢푙푡 _푇푦푝푒(푥)=(푦 → 푓(푥)), where 푓(푥)≠ 푦(10) Here,푦denotes the actual label of the input푥, and푓(푥)indicates the predicted label resulting from the misclassification. Guidance. Under each budget, we retrain the DNN with corrected labels from the top-ranked uncertain samples and report test-accuracy gain, defined as retrained accuracy minus original accuracy. 4.5 Research Questions We investigate the effectiveness and practicality of our approach through the following five research questions (RQs): RQ1: Effectiveness on Clean and Adversarial Data. How effectively does our method prioritize error-prone inputs on both clean and adversarial samples? RQ2: Fault Diversity Evaluation. Can our method detect a wider variety of error types compared to other approaches? RQ3: Effectiveness on DNN Enhancement. Does our method outperform existing approaches in terms of guiding data selection for DNN enhancement? RQ4: Ablation Study and Parameter Impact. How does varying the number of selected checkpoints affect the performance of our method? RQ5: Efficiency. How efficient is our method in prioritizing test inputs? 5 Experiment Results We present our experimental results and analyze the outcomes. We implemented our method using PyTorch V1.12.1 in Python. The experiments were conducted on an Ubuntu 22.04 system equipped with 8 x NVIDIA GeForce RTX 4090 GPUs and 512GBDDR5 RAM. We repeated the process five times for experiments and took the average. 5.1 Effectiveness on Clean and Adversarial Data (RQ1) Effectiveness on Clean Data. To evaluate the effectiveness of NCIP in identifying varied errors, we compare it against state-of-the-art baselines across all subject datasets and models. We employ Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. Test Case Prioritization for DNNs via Neural Collapse InstabilityISSTA092:11 Table 3. RAUC-ALL on Clean Data Methods MNISTFMNISTCIFAR10CIFAR100 TinyImageNetIMDBAGNEWS LeNet-1 LeNet-5 LeNet-1 LeNet-5 RN-18 VGG-11 RN-50 DN-121 RN-152 DN-201CNN Trans. CNN Trans. RS0.5130.5750.5470.5150.5280.5190.5600.5610.6500.6350.5540.5740.5070.541 LSA0.9000.8510.8450.7740.7870.7710.8640.8600.8650.8390.7200.6070.8040.700 DSA0.9790.9820.8850.8930.9430.9070.8910.8870.8930.8780.7550.7280.8380.830 Entropy0.9830.9850.8960.8980.9410.9100.8840.8800.8970.8740.7850.7250.8590.827 DeepGini0.9830.9850.8970.8980.9420.9090.8880.8850.8980.8740.7850.7250.8620.828 MSP 0.9830.9850.8970.8980.9420.9090.8890.8860.8980.8730.7850.7250.8630.828 PCS 0.9820.9850.8950.8970.9420.9090.8900.8870.8940.8710.7850.7250.8640.828 Dropout0.7850.9410.5810.8480.8550.7200.4870.5100.5780.6930.6900.7050.7020.830 EffiMAP0.8570.7530.8550.8610.8850.8660.8700.8620.8840.870– NNS0.9810.9840.8990.9100.9430.9090.8930.8920.8960.8790.7850.7240.8670.829 TDPR0.9810.8790.8690.9040.9220.8710.8830.8770.8920.8980.7980.7480.8620.862 SETS0.9540.9520.8890.8910.9190.9010.8830.8800.8950.8700.7850.7250.8450.816 NCIP0.9860.9860.9040.9310.9520.9080.9020.9040.9080.9060.8140.8440.8790.898 Table 4. RAUC-500 on Clean Data Methods MNISTFMNISTCIFAR10CIFAR100 TinyImageNetIMDBAGNEWS LeNet-1 LeNet-5 LeNet-1 LeNet-5 RN-18 VGG-11 RN-50 DN-121 RN-152 DN-201CNN Trans. CNN Trans. RS0.0220.0560.1190.0670.0630.0920.2020.1880.4450.4600.1690.2520.0970.074 LSA0.5710.4620.5030.5370.5020.6170.7060.7000.6700.8580.4450.5060.4650.465 DSA0.6770.6850.5610.5580.5400.6040.7500.7080.8530.8340.4920.5110.4650.461 Entropy0.7070.7180.5660.5630.5300.6220.7650.7740.8860.9030.4960.5220.4450.482 DeepGini0.7060.7200.5640.5610.5410.6300.7810.7930.8880.8960.4960.5220.4680.489 MSP0.7050.7200.5720.5570.5450.6260.7890.8010.9010.8830.4960.5220.4810.487 PCS0.7000.7200.5420.5480.5410.6060.7580.7350.9050.8200.4960.5220.4860.484 Dropout0.1750.5140.0650.5330.2060.1960.1350.2800.2800.4340.3260.5250.1580.462 EffiMAP 0.3940.1630.4600.4720.4320.5280.6950.6850.8600.846– NNS0.7130.7090.5740.5850.5360.6070.8010.7960.8780.9020.4970.5110.4930.490 TDPR0.7220.6770.4090.6550.4610.6000.7450.7980.9160.8860.5190.5500.4320.623 SETS0.4970.4440.4770.4540.3670.5030.6120.6280.8190.7870.4730.5230.3550.353 NCIP0.7570.7390.6110.6380.5630.6210.7960.8390.9040.9170.7040.5840.525 0.610 two complementary metrics: RAUC-ALL to assess the overall quality of the test prioritization ranking, and RAUC-500 to measure practical efficiency under a limited labeling budget. Table 3 reports the RAUC-ALL results for all competing methods. As shown, NCIP consistently outperforms the strongest baselines across most dataset-model pairs. Compared with widely adopted uncertainty-based metrics such as DeepGini and Entropy, NCIP achieves superior ranking performance. For example, on CIFAR10–ResNet-18, NCIP attains a RAUC-ALL score of 0.952, outperforming both DeepGini and Entropy. This result suggests that leveraging the geometric properties induced by NC, especially the equiangular structure of classifier weights, provides a Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. ISSTA092:12Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, and Su-Juan Qin Table 5. Wilcoxon Test Results on RAUC-ALL for NCIP vs. Each Baseline (One-sided, Holm-corrected). Baseline푝 raw 푝 Holm Effect SizeBaseline 푝 raw 푝 Holm Effect Size DeepGini 1.22× 10 −4 7.32× 10 −4 0.981NNS1.22× 10 −4 7.32× 10 −4 0.981 Dropout 6.10× 10 −5 7.32× 10 −4 1.000PCS1.22× 10 −4 7.32× 10 −4 0.981 DSA6.10× 10 −5 7.32× 10 −4 1.000 Random 6.10× 10 −5 7.32× 10 −4 1.000 EffiMAP 9.77× 10 −4 9.77× 10 −4 1.000SETS6.10× 10 −5 7.32× 10 −4 1.000 Entropy 1.22× 10 −4 7.32× 10 −4 0.981TDPR 6.10× 10 −5 7.32× 10 −4 1.000 LSA6.10× 10 −5 7.32× 10 −4 1.000MSP1.22× 10 −4 7.32× 10 −4 0.981 CIFAR10-RN-18CIFAR10-VGG-11CIFAR100-RN-50CIFAR100-DN-121 0.84 0.86 0.88 0.90 0.92 0.94 0.96 RAUC Mixup-Based Calibration CIFAR10-RN-18CIFAR10-VGG-11CIFAR100-RN-50CIFAR100-DN-121 0.84 0.86 0.88 0.90 0.92 0.94 0.96 RAUC Smoothing-Based Calibration EntropyDeepGiniMSPPCSNNSNCIP (a) GPT2-WikitextOPT-125m-Wikitext 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 RAUC Test Case Prioritization Performance on GPT2 Generative Tasks EntropyDeepGiniMSPPCSNNSNCIP (b) Fig. 4. (a) Test case prioritization under model calibration. (b) NCIP performance on LLM generative tasks using only fine-tuning checkpoints more robust indicator of error proneness than softmax confidence alone. NCIP also demonstrates advantages over coverage-guided techniques such as DSA and LSA. On FMNIST–LeNet-5, NCIP achieves a RAUC-ALL score of 0.931, substantially exceeding LSA and remaining competitive with DSA. Efficiency Under Limited Budget. In practical testing scenarios, testers typically face a con- strained budget for inspecting and labeling test inputs. Table 4 presents the RAUC-500 results, reflecting the errors identified within the top 500 prioritized samples. Although NCIP does not achieve the absolute best performance on every individual dataset-model pair, it consistently excels at the early discovery of faults. For challenging configurations such as CIFAR100–DenseNet-121, NCIP achieves a RAUC-500 score of 0.839, surpassing the strongest baseline MSP, which attains 0.801. This improvement can be attributed to NCIP’s explicit exploitation of NC geometry to prioritize geometrically distinct samples, thereby prioritizing fault-revealing test cases under a limited budget. Generalization Across Modalities. It is worth noting that NCIP maintains high performance across diverse modalities. On the text classification task AGNEWS-Transformer, NCIP achieves the highest RAUC-ALL of 0.898, demonstrating its generalizability beyond image datasets. This is evident in AGNEWS-TextCNN, where NCIP achieves a RAUC-500 score of 0.525, significantly out- performing DeepGini with a score of 0.468 and Entropy with a score of 0.445, thereby demonstrating its ability to select an earlier set of error-triggering inputs under the same testing budget. To further validate statistical significance, we applied Wilcoxon signed-rank tests [38] at훼=0.05. We compared NCIP with each baseline on paired RAUC-ALL results (mostly푛=14 pairs;푛=10 for EffiMAP). To handle multiple comparisons, we applied Holm correction and judged significance using adjusted푝-values. Exact푝-values and effect sizes (rank-biserial correlation, RBC) are reported in Table 5. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. Test Case Prioritization for DNNs via Neural Collapse InstabilityISSTA092:13 Table 6. RAUC results on adversarial data. Highlighted values indicate the best results, and underlined values indicate the second-best results. Methods RAUC-500RAUC-ALL CIFAR10CIFAR100TinyImageNetAvg.CIFAR10CIFAR100TinyImageNetAvg. RN-18 VGG-11 RN-50 DN-121 RN-152 DN-201RN-18 VGG-11 RN-50 DN-121 RN-152 DN-201 RS0.6860.6660.8160.7860.9330.8830.7950.7480.7470.8340.8440.9150.9140.834 LSA0.9320.9270.9270.9060.8810.9650.9230.8290.8370.9000.9020.9280.9220.886 DSA0.9260.9720.9940.9730.9410.9700.9630.8430.8750.9050.9080.9360.9210.898 Entropy0.9440.9320.9690.9740.9850.9740.9630.8360.8740.9040.9070.9390.9120.895 DeepGini0.9420.9310.9690.9790.9860.9770.9640.8360.8730.9040.9070.9380.9120.895 MSP 0.9360.9270.9730.9850.9880.9760.9640.8360.8740.9040.9070.9370.9110.895 PCS 0.9450.9740.9700.9730.8770.9240.9440.8350.8720.9030.9050.9350.9110.893 Dropout0.7120.7680.3890.5430.8010.8920.6840.8410.8340.7800.7890.9040.9300.846 EffiMAP0.8790.7550.9570.9550.9720.9560.9120.8520.8580.8980.9080.9370.9360.898 NNS0.8870.8920.9490.9560.9600.9630.9350.8220.8010.8870.8890.9280.9050.872 TDPR0.9360.9280.9740.9820.9880.9880.9660.9370.8990.9530.9610.9590.9640.946 SETS 0.9130.8930.9460.9560.9790.9550.9400.8360.8740.9040.9070.9370.9110.895 NCIP0.9440.9390.9650.9630.9730.9880.9620.9260.8990.9390.9450.9540.9580.937 After Holm correction, RAUC-ALL shows significant improvements for all baselines (푝 Holm ≤ 9.77×10 −4 ) with large positive effects (RBC=0.981–1.000). This result provides consistent statistical evidence that NCIP improves overall ranking quality across dataset-model pairs. To examine the impact of model calibration on test prioritization methods, we additionally evaluate two common calibration strategies, mixup [59] and label smoothing [49] (Fig. 4a). NCIP remains competitive after calibration, suggesting that the improvement is not only due to confidence rescaling at a single checkpoint, but also due to the cross-checkpoint instability signal. We further explore extended usage scenarios of NCIP. NCIP requires access to multiple check- points, so its primary deployment setting is model development and model-serving pipelines where training or fine-tuning trajectories are retained. In the LLM scenario, we do not assume access to pretraining checkpoints. Instead, we use only fine-tuning checkpoints, as shown in Fig. 4b, and treat low BLEU-2 outputs as failures. Under this setting, NCIP remains competitive on WikiText, achieving 0.587 with GPT-2 (best: 0.596) and 0.706 with OPT-125m (best). These results indicate that NCIP can be applied to generative tasks when fine-tuning checkpoints are available. Effectiveness on Adversarial Data. To evaluate the robustness of NCIP under adversarial attacks, we conducted experiments on adversarial datasets. The results are summarized in Table 6. NCIP remains strong under adversarial settings, with average scores reaching 0.962 (RAUC-500) and 0.937 (RAUC-ALL). Although TDPR achieves the best overall average in this table (0.966/0.946), NCIP is better than the remaining baselines across most dataset-model pairs. For example, on CIFAR10-ResNet-18, NCIP achieves 0.926 RAUC-ALL, while entropy-based baselines are around 0.836. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. ISSTA092:14Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, and Su-Juan Qin 0.020.040.060.080.10 Ratio 10 20 30 40 50 60 Fault Type CIFAR10 - ResNet-18 0.020.040.060.080.10 Ratio 0 100 200 300 400 500 600 CIFAR100 - ResNet-50 0.020.040.060.080.10 Ratio 0 200 400 600 800 TINY-IMAGENET - ResNet-152 0.020.040.060.080.10 Ratio 10 20 30 40 50 60 70 80 Fault Type CIFAR10 - VGG-11 0.020.040.060.080.10 Ratio 0 100 200 300 400 500 600 CIFAR100 - DenseNet-121 0.020.040.060.080.10 Ratio 0 100 200 300 400 500 600 700 800 TINY-IMAGENET - DenseNet-201 DeepGini Dropout DSA EffiMAP Entropy LSA NNS PCS RS SETS TDPR MSP NCIP Fig. 5. Fault Diversity on Clean Data Answer to RQ1: NCIP consistently outperforms most baselines on both clean and adversarial data. It improves RAUC by 1.5 %–16.6 % and increases RAUC-500 by 4.9 %– 20.6 %, enabling earlier fault discovery. Under adversarial settings, NCIP remains robust with second-best average scores (RAUC-500/RAUC-ALL: 0.962/0.937), and performs well under model calibration and on LLM generative tasks. 5.2 Fault Diversity Evaluation (RQ2) To quantify the diversity of errors detected by different DNN testing methods under varying budgets, we evaluated the number of error types identified within budgets of 1 %to10 % of the total clean test samples. As shown in Fig. 5, the horizontal axis denotes the test case selection budget as a percentage of the total samples, while the vertical axis represents the Fault Diversity score. Fig. 5 illustrates that NCIP achieves greater fault diversity than baseline methods across various dataset scales. Overall, NCIP outperforms competing approaches in identifying a wider range of fault types for most dataset-model combinations. For example, on CIFAR100 with DenseNet-121, under a stringent budget of 5 %, NCIP identifies 374 distinct failures, exceeding the performance of the strongest baseline, DeepGini, which detects 360 failures. This result highlights NCIP’s effectiveness in prioritizing inputs that expose diverse failure modes at an early stage. Answer to RQ2: NCIP enhances fault diversity by 17.5 %–37.6 % over baselines, ensuring comprehensive defect discovery for safety-critical systems. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. Test Case Prioritization for DNNs via Neural Collapse InstabilityISSTA092:15 AGNEWS TextCNN AGNEWS Transformer CIFAR10 ResNet-18 CIFAR10 VGG-11 CIFAR100 DenseNet-121 CIFAR100 ResNet-50 FMNIST LeNet-1 FMNIST LeNet-5 IMDB TextCNN IMDB Transformer MNIST LeNet-1 MNIST LeNet-5 0.00 0.02 0.04 0.06 0.08 0.10 0.12 0.14 Accuracy Improvement DeepGini Dropout DSA EffiMAP Entropy LSA NNS PCS RS TDPR MSP SETS NCIP Fig. 6. Accuracy improvement after retraining with the top 20 % selected test cases (averaged over five runs). 5.3 Effectiveness on DNN Enhancement (RQ3) To evaluate the effectiveness of the selected test cases for model improvement, we conducted retraining experiments. Specifically, we fine-tune the original models using the top 20 % of samples prioritized by each method and measure the resulting accuracy improvement on a held-out test set. The results are summarized in Fig. 6. The results indicate that retraining with samples selected by NCIP consistently leads to accuracy gains compared with other methods. NCIP achieves average accuracy improvements of 1.52 %– 5.94 %. These improvements suggest that samples with high NCIP scores correspond to challenging and informative instances that are insufficiently learned by the original models. Answer to RQ3: NCIP provides informative retraining data, achieving 1.52 %–5.94 % greater accuracy than baselines, enabling effective model improvement under limited budgets. 5.4 Ablation Study and Parameter Impact (RQ4) To quantify the contribution of each component in NCIP, we conducted an ablation study on the TVD score, the NC checkpoint selection via weight equiangularity, and the prediction margin. We compare the full NCIP variant (All) with three reduced versions: TVD only (TVD), TVD with NC checkpoint selection (TVD+NC), and TVD with margin (TVD+Margin). Results on fault diversity under a 5 % budget and RAUC-ALL are reported in Table 7 and Table 8, respectively. Incorporating the margin term is particularly beneficial for improving fault diversity, especially on challenging classification tasks. On CIFAR100 with DenseNet-121, adding the margin (TVD+Margin) increases the number of distinct failures identified at 5 % budget from 260 to 301, indicating that margin information helps prioritize ambiguous samples near the decision boundary. The TVD metric alone delivers reasonable performance across all datasets, lending support to our hypothesis that prediction variability across checkpoints can serve as a useful indicator of input hardness. For example, on CIFAR10 with ResNet-18, TVD achieves a RAUC-ALL score of 0.952. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. ISSTA092:16Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, and Su-Juan Qin Table 7. Ablation study on fault diversity at 5 % budget (clean data). Variants MNISTFMNISTCIFAR10CIFAR100 TinyImageNet LeNet-1 LeNet-5 LeNet-1 LeNet-5 RN-18 VGG-11 RN-50 DN-121 RN-152 DN-201 TVD514243415666234260366400 TVD+NC514146505771321350425443 TVD+Margin 514242455666267301392418 NCIP (Full)524148505768346374446445 Table 8. Ablation study on RAUC-ALL (clean data). Variants MNISTFMNISTCIFAR10CIFAR100IMDBAGNEWS TinyImageNet Avg. LeNet-1 LeNet-5 LeNet-1 LeNet-5 RN-18 VGG-11 RN-50 DN-121 CNN Trans. CNN Trans. RN-152 DN-201 TVD0.9830.9850.8910.9010.9440.9070.8750.8820.784 0.715 0.864 0.8920.8840.8840.885 TVD+NC 0.9830.9860.8800.9290.9420.9040.8900.8940.7980.8510.880 0.8910.9070.9120.903 TVD+Margin0.9830.9850.8960.9010.9440.9090.8870.8900.786 0.720 0.866 0.8890.8930.8810.888 NCIP (Full) 0.9860.9870.9040.9310.9520.9080.9020.9040.814 0.844 0.8790.8980.9080.9060.909 0.20.40.60.81.0 Ratio 40 45 50 55 60 65 70 Fault Type Fault Type vs Ratio (CIFAR10, FMNIST, MNIST) CIFAR10-ResNet-18 CIFAR10-VGG-11 FMNIST-LeNet-1 FMNIST-LeNet-5 MNIST-LeNet-1 MNIST-LeNet-5 0.20.40.60.81.0 Ratio 300 325 350 375 400 425 450 Fault Type Fault Type vs Ratio (CIFAR100, Tiny-ImageNet) CIFAR100-DenseNet-121 CIFAR100-ResNet-50 TINY-IMAGENET-DenseNet-201 TINY-IMAGENET-ResNet-152 20406080100 Number 40 45 50 55 60 65 70 Fault Type Fault Type vs Number (CIFAR10, FMNIST, MNIST) CIFAR10-ResNet-18 CIFAR10-VGG-11 FMNIST-LeNet-1 FMNIST-LeNet-5 MNIST-LeNet-1 MNIST-LeNet-5 20406080100 Number 360 380 400 420 440 Fault Type Fault Type vs Number (CIFAR100, Tiny-ImageNet) CIFAR100-DenseNet-121 CIFAR100-ResNet-50 TINY-IMAGENET-DenseNet-201 TINY-IMAGENET-ResNet-152 Fig. 7. The impact of the checkpoint selection ratio/number on fault diversity The NC component can improve overall ranking quality in certain settings, most notably for transformer-based NLP models. On IMDB with Transformer, TVD+NC boosts RAUC-ALL from Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. Test Case Prioritization for DNNs via Neural Collapse InstabilityISSTA092:17 0.20.40.60.81.0 Ratio 0.70 0.75 0.80 0.85 0.90 0.95 1.00 RAUC RAUC vs Ratio 20406080100 Number 0.800 0.825 0.850 0.875 0.900 0.925 0.950 0.975 RAUC RAUC vs Number AGNEWS-TextCNN AGNEWS-Transformer CIFAR10-ResNet-18 CIFAR10-VGG-11 CIFAR100-DenseNet-121 CIFAR100-ResNet-50 FMNIST-LeNet-1 FMNIST-LeNet-5 IMDB-TextCNN IMDB-Transformer MNIST-LeNet-1 MNIST-LeNet-5 TINY-IMAGENET-DenseNet-201 TINY-IMAGENET-ResNet-152 Fig. 8. The impact of the checkpoint selection ratio/number on RAUC 0.715 to 0.851, suggesting that enforcing NC-guided structure provides a stronger global signal for uncertainty estimation in such architectures. The full NCIP method consistently delivers the most robust performance across datasets and model families. By combining local boundary information with global geometric regularization, NCIP attains the best overall ranking quality; in particular, its average RAUC-ALL across all dataset-model pairs is the highest, reaching 0.909. To examine hyperparameter sensitivity, we evaluate two factors: the checkpoint selection ratio 훼(candidate-pool ratio) and the checkpoint selection number퐾(number of selected checkpoints). We report RAUC-ALL and fault diversity at a 5 % budget, as shown in Fig. 7 and Fig. 8.훼uses the last 훼 fraction of checkpoints as candidates, and 퐾 selects top-퐾 representatives from that pool. For ratio sensitivity, increasing the candidate pool from훼=0.1 to about 0.7∼0.9 generally improves performance. On Tiny-ImageNet with DenseNet-201, diversity increases from 410 (훼=0.1) to 445 (훼=0.9). For number sensitivity, performance improves when퐾is small and then plateaus; in practice,퐾≥30 already gives near-saturated RAUC and diversity on most settings. Larger퐾brings limited gains but increases multi-checkpoint inference cost. Based on these results, we use훼=0.9 and퐾=30 as defaults, balancing effectiveness and efficiency. To assess performance under weak neural collapse, we report RAUC-ALL and fault type di- versity across early training epochs in Fig. 9. Since models from earlier training epochs exhibit weaker neural collapse, evaluating checkpoints from earlier training epochs approximates weaker NC conditions. RAUC-ALL generally increases as neural collapse becomes stronger. For exam- ple, CIFAR10-ResNet-18 improves from 0.8104 (10 checkpoints) to 0.9429 (100 checkpoints), and TinyImageNet-DenseNet-201 from 0.8144 to 0.9104. These results indicate that stronger neural collapse provides a more effective signal for test prioritization ranking. Answer to RQ4: Ablation suggests that NC constraint and margin improve struc- ture and boundary sensitivity, respectively. Hyperparameter analysis shows that increasing the checkpoint selection ratio improves RAUC-ALL and fault diversity up to about 훼= 0.9, while selection-number gains saturate around 퐾= 30. 5.5 Efficiency (RQ5) To evaluate computational efficiency, we report test case selection times in Table 9. For NCIP, the main overhead is in the selection stage, which includes multi-checkpoint loading and computing Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. ISSTA092:18Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, and Su-Juan Qin 10203050100 Training Epochs 0.75 0.80 0.85 0.90 0.95 RAUC RAUC-ALL vs Training Epochs (Weak NC) 10203050100 Training Epochs 0 100 200 300 400 Fault Type Fault Type vs Training Epochs (Weak NC) AGNEWS-TextCNN AGNEWS-Transformer CIFAR10-ResNet-18 CIFAR10-VGG-11 CIFAR100-DenseNet-121 CIFAR100-ResNet-50 FMNIST-LeNet-1 FMNIST-LeNet-5 IMDB-TextCNN IMDB-Transformer MNIST-LeNet-1 MNIST-LeNet-5 TINY-IMAGENET-DenseNet-201 TINY-IMAGENET-ResNet-152 Fig. 9. NCIP Performance under Weak Neural Collapse Table 9. Execution time (in seconds) for each method. For clarity, we group LSA and DSA under SA, and aggregate uncertainty-based baselines (e.g., DeepGini, Entropy) under Uncert. DatasetModelSAUncert Dropout NNS SETS EffiMAPTDPRNCIP TrainPredictTrainPredictSelect Predict MNIST LeNet-1 78.155.466.388.9267.19220.315.1970.704.581.964.95 LeNet-5 84.135.138.077.8665.46231.235.781.889.432.034.94 Fashion-MNIST LeNet-175.345.577.557.7069.23205.066.32185.164.881.984.73 LeNet-5 72.935.707.208.8267.58217.715.12156.134.962.074.92 CIFAR10 RN-1878.776.9453.4610.96 68.12654.026.99517.266.4619.9133.57 VGG-1182.846.3712.8411.06 74.35436.515.82484.175.9218.2011.55 CIFAR100 RN-50161.698.37116.3918.27 87.23912.308.401881.147.4645.7964.94 DN-121140.417.8290.6012.56 78.10889.288.371548.906.7927.7957.62 IMDB CNN155.5412.7915.2333.71 68.08–98.2110.719.5613.95 Trans.165.2811.9431.5230.26 73.60–120.099.4911.0024.94 AGNEWS CNN 84.149.4912.5211.43 59.85–74.309.3530.6510.91 Trans. 71.948.3116.0912.53 56.91–80.547.909.1912.01 TinyImageNet RN-152493.0118.99749.4336.71 88.904371.5617.338003.9415.30148.31374.21 DN-201762.3716.46560.7431.66 90.553806.2514.685670.9713.02306.87311.10 휎 cos for checkpoint filtering. This selection stage is a one-time cost per model. According to Table 9, NCIP-Select ranges from 1.96s to 306.87s across all settings, and remains much lower than training-heavy baselines such as TDPR-Train (up to 8003.94s) and EffiMAP-Train (up to 4371.56s). NCIP-Predict can still be higher on large models because inference uses multiple selected checkpoints (e.g., TinyImageNet), but the method keeps a practical effectiveness-runtime trade-off through controllable checkpoint budgeting. Answer to RQ5: NCIP provides efficient prioritization with controllable overhead: the selection stage (multi-checkpoint loading and휎 cos computation) is required only once per model and is far cheaper than training-based baselines. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. Test Case Prioritization for DNNs via Neural Collapse InstabilityISSTA092:19 6 Threats to Validity Implementation and Evaluation Factors. Our findings may be influenced by implementation and configuration factors. Any inaccuracies in checkpoint indexing, preprocessing consistency, or numerical stability in similarity/margin computations may bias the prioritization scores. In addition, training and evaluation stochasticity (e.g., random initialization, data shuffling, and GPU nondeterminism) can introduce variance. Moreover, adversarial robustness evaluation is sensitive to attack configurations; although we use widely adopted attacks (FGSM/PGD/CW/BIM) with standard library implementations and default settings, alternative hyperparameters may change the absolute RAUC values and, in some cases, relative rankings. To mitigate these threats, we enforce a unified training and evaluation protocol across all methods and use identical test sets and budgets. We also release our code and configuration to support reproducibility. Difference Between Model Disagreement-Based Methods. A construct-validity threat is that different methods define model disagreement differently. Dropout and EffiMAP derive disagreement from stochastic perturbations around one checkpoint, whereas NCIP uses disagreement across NC-guided training checkpoints and combines it with final-margin information. Compared with disagreement-based methods that measure prediction variation through model perturbation, NCIP estimates prediction variability from naturally available training checkpoints. The two families rely on different sources of prediction instability, which may offer complementary perspectives in different deployment contexts. Practicality and Deployment Constraints. NCIP incurs considerable runtime overhead due to loading checkpoints and performing multi-checkpoint inference to estimate prediction variability. This cost can hinder adoption in industrial pipelines where throughput and latency are critical. NCIP is mainly intended for model development and managed serving settings where training or fine-tuning checkpoints are retained, not for third-party deployed models without checkpoint access. Even with this overhead, NCIP is still useful in regression testing, where running every test in every cycle is often infeasible and reordering tests under a limited budget matters. This scope is also reflected in our LLM setting: we use only fine-tuning checkpoints (Fig. 4b) and NCIP remains competitive on WikiText (GPT-2 (0.587), OPT-125m (0.706)). 7 Related Work 7.1 DNN Testing Testing DNN-based systems has been extensively studied [13,14,16,36,43,44,58,66]. Since DNN development is largely data-driven [3, 41], a core challenge is to expose model failures effectively under constrained labeling, time, and computation budgets. 7.2 Test Input Prioritization Test prioritization has become a particularly effective method [1,2,17,19,21,33,34,51]. This involves ranking test inputs based on their likelihood of failure, which allows for the early detection of critical errors within limited time and computational resources by focusing on high-risk samples. Existing DNN test-input prioritization techniques can be mainly divided into five categories: coverage-based methods, surprise adequacy-based methods, confidence-based methods, model- perturbation-based methods and learning-to-rank methods. Coverage-Based. Coverage-based methods adapt traditional code-coverage ideas to DNNs by quantifying internal activation behaviors (e.g., neuron coverage) and prioritizing inputs that maximize such criteria [12,36,57,60]. Coverage-based methods focus on internal activation diversity, while NCIP focuses on prediction stability across the training trajectory. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. ISSTA092:20Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, and Su-Juan Qin Surprise Adequacy-Based. Surprise adequacy (SA) methods prioritize inputs by measuring how novel an input’s activation trace is relative to the training distribution [25,26,32]. Representative instantiations include KDE-based likelihood estimation and distance-based SA (DSA) [24]. SA has been used as a quantitative target to guide systematic boundary exploration [22]. SA-based methods quantify input novelty relative to training distribution references. In contrast, NCIP operates on checkpoint-level prediction consistency without accessing training data at prioritization time. Confidence-Based. These methods estimate error likelihood from output probabilities, e.g., DeepGini [9], entropy [5], PCS [62], and MSP [55]. Recent works improve uncertainty estimation by incorporating neighborhood uncertainty (NNS) [4], adjusting probability vectors via feature selection (FAST) [7], or jointly optimizing uncertainty and diversity (SETS) [52]. Confidence-based methods derive ranking signals from per-sample output probabilities at a single training state. NCIP complements this perspective by measuring prediction variability across multiple training checkpoints, targeting instability that may not be captured by point-estimate confidence alone. Model-Perturbation-Based. These methods prioritize inputs by measuring prediction changes under model-side perturbations. Dropout [18] estimates uncertainty from repeated stochastic forward passes, while EffiMAP [54] perturbs both model and input and ranks samples by the resulting prediction variation. Model-perturbation methods estimate prediction variation through repeated perturbation at a fixed checkpoint. NCIP instead draws variation from naturally occurring training checkpoints, providing a different source of prediction instability. Learning-to-Rank. Learning-to-rank methods learn a scoring function from labeled supervision so that bug-revealing inputs appear early under a budget. Examples include mutation-based labeling to improve prioritization (PRIMA) [53], feature-based ranking for classical models (MLPrior) [8], and approaches exploiting training dynamics to distinguish bug-revealing trajectories (TDPR) [46]. Learning-to-rank methods train a dedicated ranking model from labeled supervision, making effective use of failure information. NCIP is designed for settings where a trained model and its checkpoints are accessible but labeled test data is not, using cross-checkpoint variability as an unsupervised prioritization signal. 8 Conclusion This paper presents NCIP, a label-free error prioritization approach for deep classifiers that leverages geometric stability during training. NCIP selects an NC-guided representative subset using classifier-weight equiangularity scores, and then prioritizes test inputs by combining (i) temporal prediction variability across these checkpoints with (i) the final model prediction margin to capture decision boundary proximity. By tying misclassification risk to cross-checkpoint prediction variability, NCIP offers a principled alternative to confidence-only heuristics while requiring no training data during prioritization. Extensive experiments show that NCIP achieves competitive performance in early fault detection and delivers robust prioritization across datasets, architectures, and evaluation settings. Future work will explore extending NC-guided selection to broader model families and developing adaptive checkpoint selection strategies to further reduce computational overhead. Acknowledgments This work is supported by National Natural Science Foundation of China (Grant Nos. U25B2014, 62371069, 62372048, 62272056) Data Availability Our code is available at https://github.com/lucky1207/NCIP. Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. Test Case Prioritization for DNNs via Neural Collapse InstabilityISSTA092:21 Declaration of Generative AI Use During the preparation of this work, the authors used GPT 5.2 in order to improve the readability and language of the manuscript. References [1]Amin Abbasishahkoo, Mahboubeh Dadkhah, Lionel Briand, and Dayi Lin. 2025. MetaSel: A Test Selection Approach for Fine-tuned DNN Models. arXiv preprint arXiv:2503.17534 (2025). doi:10.1109/TSE.2025.3612253 [2]Hamzah Al-Qadasi, Changshun Wu, Yliès Falcone, and Saddek Bensalem. 2022. DeepAbstraction: 2-level prioritization for unlabeled test inputs in deep neural networks. In 2022 IEEE International Conference On Artificial Intelligence Testing (AITest). IEEE, 64–71. doi:10.1109/aitest55621.2022.00018 [3] Ibrahim M Alabdulmohsin, Behnam Neyshabur, and Xiaohua Zhai. 2022. Revisiting neural scaling laws in language and vision. Advances in Neural Information Processing Systems 35 (2022), 22300–22312. doi:10.52202/068431-1620 [4]Shenglin Bao, Chaofeng Sha, Bihuan Chen, Xin Peng, and Wenyun Zhao. 2023. In defense of simple techniques for neural network test case selection. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis. 501–513. doi:10.1145/3597926.3598073 [5]Taejoon Byun, Vaibhav Sharma, Abhishek Vijayakumar, Sanjai Rayadurgam, and Darren Cofer. 2019. Input prioritization for testing neural networks. In 2019 IEEE International Conference On Artificial Intelligence Testing (AITest). IEEE, 63–70. doi:10.1109/AITest.2019.000-6 [6] Nicholas Carlini and David Wagner. 2017. Towards evaluating the robustness of neural networks. In 2017 ieee symposium on security and privacy (sp). Ieee, 39–57. doi:10.1109/SP.2017.49 [7] Jialuo Chen, Jingyi Wang, Xiyue Zhang, Youcheng Sun, Marta Kwiatkowska, Jiming Chen, and Peng Cheng. 2024. FAST: Boosting Uncertainty-based Test Prioritization Methods for Neural Networks via Feature Selection. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering. 895–906. doi:10.1145/3691620.3695472 [8]Xueqi Dang, Yinghua Li, Mike Papadakis, Jacques Klein, Tegawendé F Bissyandé, and Yves Le Traon. 2024. Test input prioritization for machine learning classifiers. IEEE Transactions on Software Engineering (2024). doi:10.1109/TSE.2024. 3350019 [9]Yang Feng, Qingkai Shi, Xinyu Gao, Jun Wan, Chunrong Fang, and Zhenyu Chen. 2020. Deepgini: prioritizing massive tests to enhance the robustness of deep neural networks. In Proceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Analysis. 177–188. doi:10.1145/3395363.3397357 [10] Xinyu Gao, Yang Feng, Yining Yin, Zixi Liu, Zhenyu Chen, and Baowen Xu. 2022. Adaptive test selection for deep neural networks. In Proceedings of the 44th International Conference on Software Engineering. 73–85. doi:10.1145/ 3510003.3510232 [11]Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572 (2014). doi:10.48550/arXiv.1412.6572 [12]Hongjing Guo, Chuanqi Tao, Zhiqiu Huang, and Weiqin Zou. 2025. White-Box Test Input Generation for Enhancing Deep Neural Network Models through Suspicious Neuron Awareness. ACM Transactions on Software Engineering and Methodology (2025). doi:10.1145/3736305 [13]Jianmin Guo, Yu Jiang, Yue Zhao, Quan Chen, and Jiaguang Sun. 2018. Dlfuzz: Differential fuzzing testing of deep learning systems. In Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 739–743. doi:10.1145/3236024.3264835 [14]Fabrice Harel-Canada, Lingxiao Wang, Muhammad Ali Gulzar, Quanquan Gu, and Miryung Kim. 2020. Is neuron coverage a meaningful measure for testing deep neural networks?. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 851–862. doi:10.1145/3368089.3409754 [15]Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition. 770–778. doi:10.1109/CVPR.2016.90 [16] Qiang Hu, Yuejun Guo, Xiaofei Xie, Maxime Cordy, Lei Ma, Mike Papadakis, and Yves Le Traon. 2024. Test optimization in DNN testing: a survey. ACM Transactions on Software Engineering and Methodology 33, 4 (2024), 1–42. doi:10.1145/ 3643678 [17] Qiang Hu, Yuejun Guo, Xiaofei Xie, Maxime Cordy, Wei Ma, Mike Papadakis, Lei Ma, and Yves Le Traon. 2025. Assessing the Robustness of Test Selection Methods for Deep Neural Networks. ACM Transactions on Software Engineering and Methodology (2025). doi:10.1145/3715693 [18] Qiang Hu, Yuejun Guo, Xiaofei Xie, Maxime Cordy, Mike Papadakis, Lei Ma, and Yves Le Traon. 2023. Aries: Efficient testing of deep neural networks via labeling-free accuracy estimation. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 1776–1787. doi:10.1109/ICSE48619.2023.00152 Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. ISSTA092:22Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, and Su-Juan Qin [19]Dong Huang, Qingwen Bu, Yichao Fu, Yuhao Qing, Xiaofei Xie, Junjie Chen, and Heming Cui. 2024. Neuron Sensitivity- Guided Test Case Selection. ACM Transactions on Software Engineering and Methodology 33, 7 (2024), 1–32. doi:10. 1145/3672454 [20]Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. 2017. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition. 4700–4708. doi:10.1109/ CVPR.2017.243 [21] Hyekyoung Hwang, Il Yong Chun, and Jitae Shin. 2023. Improved Test Input Prioritization Using Verification Monitors with False Prediction Cluster Centroids. Electronics 13, 1 (2023), 21. doi:10.3390/electronics13010021 [22]Sungmin Kang, Robert Feldt, and Shin Yoo. 2024. Deceiving humans and machines alike: Search-based test input generation for dnns using variational autoencoders. ACM Transactions on Software Engineering and Methodology 33, 4 (2024), 1–24. doi:10.1145/3635706 [23] Hoki Kim. 2020. Torchattacks: A pytorch repository for adversarial attacks. arXiv preprint arXiv:2010.01950 (2020). doi:10.48550/arXiv.2010.01950 [24]Jinhan Kim, Robert Feldt, and Shin Yoo. 2019. Guiding deep learning system testing using surprise adequacy. In 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE). IEEE, 1039–1049. doi:10.1109/ICSE.2019.00108 [25]Jinhan Kim, Robert Feldt, and Shin Yoo. 2023. Evaluating surprise adequacy for deep learning system testing. ACM Transactions on Software Engineering and Methodology 32, 2 (2023), 1–29. doi:10.1145/3546947 [26] Jinhan Kim, Jeongil Ju, Robert Feldt, and Shin Yoo. 2020. Reducing dnn labelling cost using surprise adequacy: An industrial case study for autonomous driving. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 1466–1476. doi:10.1145/3368089. 3417065 [27] Yoon Kim. 2014. Convolutional neural networks for sentence classification. arXiv preprint arXiv:1408.5882 (2014). doi:10.3115/v1/D14-1181 [28] Alex Krizhevsky, Geoffrey Hinton, et al. 2009. Learning multiple layers of features from tiny images. (2009). [29]Alexey Kurakin, Ian J Goodfellow, and Samy Bengio. 2018. Adversarial examples in the physical world. In Artificial intelligence safety and security. Chapman and Hall/CRC, 99–112. doi:10.1201/9781351251389-8 [30] Yann Le and Xuan Yang. 2015. Tiny imagenet visual recognition challenge. CS 231N 7, 7 (2015), 3. [31]Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 1998. Gradient-based learning applied to document recognition. Proc. IEEE 86, 11 (1998), 2278–2324. doi:10.1109/5.726791 [32]Maoxi Li, Daobo Ma, and Yingqi Zhang. 2025. Improving database anomaly detection efficiency through sample difficulty estimation. (2025). doi:10.20944/preprints202504.1527.v1 [33]Wei Li, Zhiyi Zhang, Yifan Jian, Chen Liu, and Zhiqiu Huang. 2023. DeepRank: Test Case Prioritization for Deep Neural Networks.. In SEKE. 262–267. doi:10.18293/seke2023-188 [34]Yinghua Li, Xueqi Dang, Jacques Klein, Yves Le Traon, and Tegawendé F Bissyandé. 2025. PriCod: Prioritizing Test Inputs for Compressed Deep Neural Networks. ACM Transactions on Software Engineering and Methodology (2025). doi:10.1145/3730435 [35]Zhong Li, Zhengfeng Xu, Ruihua Ji, Minxue Pan, Tian Zhang, Linzhang Wang, and Xuandong Li. 2024. Distance-Aware Test Input Selection for Deep Neural Networks. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis. 248–260. doi:10.1145/3650212.3652125 [36] Lei Ma, Felix Juefei-Xu, Fuyuan Zhang, Jiyuan Sun, Minhui Xue, Bo Li, Chunyang Chen, Ting Su, Li Li, Yang Liu, et al. 2018. Deepgauge: Multi-granularity testing criteria for deep learning systems. In Proceedings of the 33rd ACM/IEEE international conference on automated software engineering. 120–131. doi:10.1145/3238147.3238202 [37]Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies. 142–150. [38]Thomas W MacFarland and Jan M Yates. 2016. Wilcoxon matched-pairs signed-ranks test. In Introduction to Nonpara- metric statistics for the biological sciences using R. Springer, 133–175. doi:10.4135/9780857020123.n633 [39] Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2017. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083 (2017). doi:10.48550/arXiv.1706.06083 [40]Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843 (2016). doi:10.48550/arXiv.1609.07843 [41] Michael A Nielsen. 2015. Neural networks and deep learning. Vol. 25. Determination press San Francisco, CA, USA. [42]Vardan Papyan, XY Han, and David L Donoho. 2020. Prevalence of neural collapse during the terminal phase of deep learning training. Proceedings of the National Academy of Sciences 117, 40 (2020), 24652–24663. doi:10.1073/pnas. 2015509117 [43]Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Jana. 2017. Deepxplore: Automated whitebox testing of deep learning systems. In proceedings of the 26th Symposium on Operating Systems Principles. 1–18. doi:10.1145/3132747.3132785 Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. Test Case Prioritization for DNNs via Neural Collapse InstabilityISSTA092:23 [44]Xin Qiu and Risto Miikkulainen. 2022. Detecting misclassification errors in neural networks with a gaussian process model. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 8017–8027. doi:10.1609/aaai.v36i7.20773 [45]Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al.2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9. [46]Jian Shen, Zhong Li, Minxue Pan, and Xuandong Li. 2024. Prioritizing Test Inputs for DNNs Using Training Dynamics. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering. 1219–1231. doi:10. 1145/3691620.3695498 [47] Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014). doi:10.48550/arXiv.1409.1556 [48] Jack Stewart. 2018. Tesla’s autopilot was involved in another deadly car crash. Wired 3 (2018), 30. [49] Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2818–2826. doi:10.1109/cvpr.2016.308 [50]Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017). doi:10.48550/arXiv.1706.03762 [51] Huiyan Wang, Jingwei Xu, Chang Xu, Xiaoxing Ma, and Jian Lu. 2020. Dissector: Input validation for deep learning applications by crossing-layer dissection. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering. 727–738. doi:10.1145/3377811.3380379 [52]Jingling Wang, Huayao Wu, Peng Wang, Xintao Niu, and Changhai Nie. 2025. SETS: A Simple yet Effective DNN Test Selection Approach. ACM Transactions on Software Engineering and Methodology (2025). doi:10.1145/3772084 [53] Zan Wang, Hanmo You, Junjie Chen, Yingyi Zhang, Xuyuan Dong, and Wenbin Zhang. 2021. Prioritizing test inputs for deep neural networks via mutation analysis. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE). IEEE, 397–409. doi:10.1109/ICSE43902.2021.00046 [54]Zhengyuan Wei, Haipeng Wang, Imran Ashraf, and WK Chan. 2022. Predictive mutation analysis of test case prioritization for deep neural networks. In 2022 IEEE 22nd International Conference on Software Quality, Reliability and Security (QRS). IEEE, 682–693. doi:10.1109/QRS57517.2022.00074 [55] Michael Weiss and Paolo Tonella. 2022. Simple techniques work surprisingly well for neural network test prioritization and active learning (replicability study). In Proceedings of the 31st ACM SIGSOFT International Symposium on Software Testing and Analysis. 139–150. doi:10.1145/3533767.3534375 [56] Han Xiao, Kashif Rasul, and Roland Vollgraf. 2017. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747 (2017). doi:10.48550/arXiv.1708.07747 [57] Yining Yin, Yang Feng, Shihao Weng, Xinyu Gao, Jia Liu, and Zhihong Zhao. 2025. Lightweight Probabilistic Coverage Metrics for Efficient Testing of Deep Neural Networks. In Proceedings of the 16th International Conference on Internetware. 474–486. doi:10.1145/3755881.3755915 [58]Hanmo You, Zan Wang, Junjie Chen, Shuang Liu, and Shuochuan Li. 2023. Regression fuzzing for deep learning systems. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 82–94. doi:10.1109/ icse48619.2023.00019 [59] Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. 2017. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412 (2017). doi:10.48550/arXiv.1710.09412 [60] Kai Zhang, Yongtai Zhang, Liwei Zhang, Hongyu Gao, Rongjie Yan, and Jun Yan. 2020. Neuron activation frequency based test case prioritization. In 2020 International Symposium on Theoretical Aspects of Software Engineering (TASE). IEEE, 81–88. doi:10.1109/tase49443.2020.00020 [61]Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al.2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068 (2022). doi:10.48550/arXiv.2205.01068 [62]Xiyue Zhang, Xiaofei Xie, Lei Ma, Xiaoning Du, Qiang Hu, Yang Liu, Jianjun Zhao, and Meng Sun. 2020. Towards characterizing adversarial defects of deep learning software from the lens of uncertainty. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering. 739–751. doi:10.1145/3377811.3380368 [63]Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. Advances in neural information processing systems 28 (2015). [64]Zhiyi Zhang, Huanze Meng, Yuchen Ding, Shuxian Chen, and Yongming Yao. 2025. Efficient adaptive test case selection for DNNs robustness enhancement. Journal of Systems and Software (2025), 112451. doi:10.2139/ssrn.4978392 [65] Haibin Zheng, Jinyin Chen, and Haibo Jin. 2023. CertPri: certifiable prioritization for deep neural networks via movement cost in feature space. In 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 1–13. doi:10.1109/ASE56229.2023.00126 Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026. ISSTA092:24Chunyu Liu, Mingyuan Li, Yang Li, Wenmin Li, Fei Gao, Tengfei Tu, and Su-Juan Qin [66]Haibin Zheng, Zhiqing Chen, Tianyu Du, Xuhong Zhang, Yao Cheng, Shouling Ji, Jingyi Wang, Yue Yu, and Jinyin Chen. 2022. Neuronfair: Interpretable white-box fairness testing through biased neuron identification. In Proceedings of the 44th International Conference on Software Engineering. 1519–1531. doi:10.1145/3510003.3510123 [67]Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu. 2021. A geometric analysis of neural collapse with unconstrained features. Advances in Neural Information Processing Systems 34 (2021), 29820–29834. doi:10.48550/arXiv.2105.02375 Received 2026-01-30; accepted 2026-06-25 Proc. ACM Softw. Eng., Vol. 3, No. ISSTA, Article ISSTA092. Publication date: October 2026.