Paper deep dive
Making Conformal Predictors Robust in Healthcare Settings: a Case Study on EEG Classification
Arjun Chatterjee, Sayeed Sajjad Razin, John Wu, Siddhartha Laghuvarapu, Jathurshan Pradeepkumar, Jimeng Sun
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/20/2026, 4:15:48 PM
Summary
This paper evaluates conformal prediction methods for EEG seizure classification, highlighting challenges posed by patient distribution shifts and label uncertainty. It introduces Neighborhood Conformal Prediction (NCP), a personalized calibration strategy that significantly improves empirical coverage compared to non-personalized baselines like Naive CP and Covariate CP, while maintaining compact prediction set sizes. The study utilizes the TUH EEG Corpus (TUEV and TUAB datasets) and implements the solution via the PyHealth framework.
Entities (11)
Relation Signals (6)
Patient Distribution Shift → violates → i.i.d. Assumptions
confidence 96% · patient distribution shifts violate the i.i.d. assumptions underlying standard conformal methods
PyHealth → implements → Conformal Prediction Methods
confidence 95% · Our implementation is available via PyHealth... enabling easy replication in other clinical settings
Neighborhood Conformal Prediction → improves → Empirical Coverage
confidence 95% · NCP substantially improves coverage on the random split... achieves around 34% greater coverage than Naive CP
Naive CP → failsunder → Patient Distribution Shift
confidence 93% · under patient-level distribution shift, conventional conformal prediction techniques fail to provide adequate coverage
ContraWR → usedfor → EEG Classification
confidence 92% · We train ContraWR... on TUAB and TUEV
NCP → maintains → Compact Prediction Sets
confidence 90% · NCP consistently maintains smaller prediction sets than non-personalized methods despite achieving competitive or higher coverage
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Quantifying uncertainty in clinical predictions is critical for high-stakes diagnosis tasks. Conformal prediction offers a principled approach by providing prediction sets with theoretical coverage guarantees. However, in practice, patient distribution shifts violate the i.i.d. assumptions underlying standard conformal methods, leading to poor coverage in healthcare settings. In this work, we evaluate several conformal prediction approaches on EEG seizure classification, a task with known distribution shift challenges and label uncertainty. We demonstrate that personalized calibration strategies can improve coverage by over 20 percentage points while maintaining comparable prediction set sizes. Our implementation is available via PyHealth, an open-source healthcare AI framework: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2602.19483v2
- Canonical: https://arxiv.org/abs/2602.19483v2
Trouble viewing inline? Open PDF directly →
Full Text
10,176 characters extracted from source content.
Expand or collapse full text
11institutetext: University of Illinois Urbana-Champaign, Urbana, IL 61801, USA 11email: arjunc4,johnwu3@illinois.edu 22institutetext: PyHealth 33institutetext: Bangladesh University of Engineering and Technology Making Conformal Predictors Robust in Healthcare Settings: a Case Study on EEG Classification Arjun Chatterjee∗ Sayeed Sajjad Razin John Wu∗ Siddhartha Laghuvarapu Jathurshan Pradeepkumar Jimeng Sun Abstract Quantifying uncertainty in clinical predictions is critical for high-stakes diagnosis tasks. Conformal prediction offers a principled approach by providing prediction sets with theoretical coverage guarantees. However, in practice, patient distribution shifts violate the i.i.d. assumptions underlying standard conformal methods, leading to poor coverage in healthcare settings. In this work, we evaluate several conformal prediction approaches on EEG seizure classification, a task with known distribution shift challenges and label uncertainty. We demonstrate that personalized calibration strategies can improve coverage by over 20 percentage points while maintaining comparable prediction set sizes. Our implementation is available via PyHealth, an open-source healthcare AI framework: https://github.com/sunlabuiuc/PyHealth. 1 Introduction Figure 1: EEG Classification Challenges. EEG tasks often violate key aspects of a typical machine learning pipeline. (a) Their annotation process consistently leaves room for uncertainty, which is then passed onto the models. (b) Training, validation, and test distributions are not i.i.d. due to patient distribution shift [13]. Both issues make EEG classification a very challenging machine learning problem. For many healthcare tasks, uncertainty is intrinsic to the problem itself. This is exemplified by EEG classification, where labels are derived from group voting among expert neurologists [2, 7]. Whether due to measurement noise or patient-to-patient variance, expert disagreements on waveform readings are not uncommon [2, 10], as illustrated in Fig. 1. Single-class predictions can therefore be overconfident or misleading. Conformal prediction [8] directly addresses this by producing prediction sets rather than point predictions for uncertain samples. However, it assumes i.i.d. calibration and test distributions, a condition that EEG classification routinely violates due to patient differences and recording conditions [13]. We demonstrate that under patient-level distribution shift, conventional conformal prediction techniques fail to provide adequate coverage. We explore conformal prediction approaches robust to such shifts using ContraWR [14] on TUAB [6] and TUEV [4], evaluated under both random and patient-level splits. Our contributions are: (1) a neighborhood conformal prediction (NCP) approach that improves coverage by up to 20% over baselines; (2) theoretical results showing NCP yields provably better coverage under covariate shift; (3) prediction set sizes that remain relatively stable; and (4) a modular PyHealth integration enabling easy replication in other clinical settings. 2 Experiments and Results 2.0.1 Datasets. We evaluate on two benchmark tasks under two splitting regimes from the TUH EEG Corpus [7]: TUEV [4], a six-class EEG event classification task (SPSW, GPED, PLED, EYEM, ARTF, BCKG), and TUAB [6], a binary normal/abnormal detection task. In the random split setting, all samples are pooled and divided globally at 60%, 10%, 15%, and 15% for the training, validation, calibration, and test sets, respectively. In the patient split setting, patients are partitioned first, with the train partition further split at 60%, 20%, and 20% for train, validation, and calibration sets respectively, and held-out patients forming the test set, introducing realistic cross-patient distribution shift. 2.0.2 Models and Conformal Approaches. We train ContraWR [14], a ResNet-based 2D CNN over multi-channel spectrograms, from scratch across five random seeds. We compare four conformal prediction (CP) approaches: Naive CP [8] (standard split conformal prediction assuming i.i.d. calibration and test), Covariate CP [11] (likelihood-ratio reweighting via KDE [5]), K-means CP (threshold derived from the nearest cluster’s calibration samples), and NCP [3] (sample-specific calibration via k-nearest neighbors with relevance weighting). The first two are non-personalized; the latter two are personalized. More theoretical results are provided in [1]. Figure 2: Empirical coverage under random (top) and patient (bottom) splits. The dotted line is target coverage 1−α1-α. Under the random split, NCP substantially outperforms non-personalized baselines on TUEV. Under the patient split, all methods fall short of target coverage, reflecting the difficulty of cross-patient distribution shift. Figure 3: Average prediction set sizes under random (top) and patient (bottom) splits. NCP consistently maintains smaller prediction sets than non-personalized methods despite achieving competitive or higher coverage. Figure 4: Empirical coverage vs. α for NCP using varying calibration set sizes k. Shaded regions denote ±1± 1 std. The dashed line is target coverage 1−α1-α. Covariate CP does not improve coverage over Naive CP. As seen in Fig. 2, the two methods track similarly across all α values and both datasets, as likelihood-ratio estimation via KDE is unreliable in high-dimensional EEG feature spaces [5]. NCP substantially improves coverage on the random split. On TUEV under the random split, NCP achieves around 34% greater coverage than Naive CP at α=0.2α=0.2 (0.87 vs. 0.53), with a modest increase in average prediction set size (1.22 vs. 0.90) (Fig. 3). On TUAB, where Naive CP already satisfies coverage due to the easier binary task, the personalization benefit is less pronounced but NCP still produces notably more compact prediction sets. Patient-level splits reveal the limits of all approaches. Ensuring calibration and test patients are disjoint causes all methods to fall well short of target coverage on TUEV, with high variance across seeds, and a smaller but still present gap on TUAB. NCP’s smaller set sizes under this split suggest its personalization yields overly tight prediction sets under strong distribution shift. Empirical coverage remains imperfect. No method consistently achieves 1−α1-α coverage across all settings, and cross-patient distribution shift remains an open challenge for conformal prediction in EEG classification. 3 Discussion 3.0.1 Future directions. Our EEG case study highlights distribution shift challenges common across healthcare [12, 13]. Applying personalized conformal predictors more broadly could reveal which shift types these methods handle well. NCP’s coverage could also improve with stronger EEG foundation model embeddings, such as TFM-Tokenizer [9], which better characterizes signal structure. Finally, all implementations are available via pip install pyhealth to lower barriers and encourage adoption of conformal prediction in safety-critical healthcare AI. credits 3.0.2 There are no competing interests. References [1] A. Chatterjee, S. S. Razin, J. Wu, S. Laghuvarapu, J. Pradeepkumar, and J. Sun (2026) Making conformal predictors robust in healthcare settings: a case study on eeg classification. External Links: 2602.19483, Link Cited by: §2.0.2. [2] W. Ge, J. Jing, S. An, A. Herlopian, M. Ng, A. F. Struck, B. Appavu, E. L. Johnson, G. Osman, H. A. Haider, et al. (2021) Deep active learning for interictal ictal injury continuum eeg patterns. Journal of neuroscience methods 351, p. 108966. Cited by: §1. [3] S. Ghosh, T. Belkhouja, Y. Yan, and J. R. Doppa (2023-Jun.) Improving uncertainty quantification of deep classifiers via neighborhood conformal prediction: novel algorithm and theoretical analysis. Proceedings of the AAAI Conference on Artificial Intelligence 37 (6), p. 7722–7730. Cited by: §2.0.2. [4] A. Harati, M. Golmohammadi, S. Lopez, I. Obeid, and J. Picone (2015) Improved eeg event classification using differential energy. IEEE Signal Processing in Medicine and Biology Symposium 2015. Cited by: §1, §2.0.1. [5] S. Laghuvarapu, Z. Lin, and J. Sun (2023) Codrug: conformal drug property prediction with density estimation under covariate shift. Advances in Neural Information Processing Systems 36, p. 37728–37747. Cited by: §2.0.2, §2.0.2. [6] S. Lopez, G. Suarez, D. Jungreis, I. Obeid, and J. Picone (2015) Automated identification of abnormal adult eegs. IEEE Signal Processing in Medicine and Biology Symposium 2015. Cited by: §1, §2.0.1. [7] I. Obeid and J. Picone (2016) The temple university hospital eeg data corpus. Frontiers in neuroscience 10, p. 196. Cited by: §1, §2.0.1. [8] H. Papadopoulos, V. Vovk, and A. Gammerman (2007) Conformal prediction with neural networks. In 19th IEEE International Conference on Tools with Artificial Intelligence (ICTAI 2007), Vol. 2, p. 388–395. Cited by: §1, §2.0.2. [9] J. Pradeepkumar, X. Piao, Z. Chen, and J. Sun (2026) Tokenizing single-channel eeg with time-frequency motif learning. External Links: 2502.16060, Link Cited by: §3.0.1. [10] V. Shah, E. von Weltin, S. Lopez, J. R. McHugh, L. Veloso, M. Golmohammadi, I. Obeid, and J. Picone (2018) The temple university hospital seizure detection corpus. Frontiers in Neuroinformatics 12. Cited by: §1. [11] R. J. Tibshirani, R. Foygel Barber, E. Candes, and A. Ramdas (2019) Conformal prediction under covariate shift. In Advances in Neural Information Processing Systems, Vol. 32. Cited by: §2.0.2. [12] Z. Wu, H. Yao, D. Liebovitz, and J. Sun (2023) An iterative self-learning framework for medical domain generalization. In Advances in Neural Information Processing Systems, Vol. 36, p. 54833–54854. Cited by: §3.0.1. [13] C. Yang, M. B. Westover, and J. Sun (2023) ManyDG: many-domain generalization for healthcare applications. In The 11th International Conference on Learning Representations, ICLR 2023, Cited by: Figure 1, §1, §3.0.1. [14] C. Yang, D. Xiao, M. B. Westover, and J. Sun (2023) Self-supervised eeg representation learning for automatic sleep staging. JMIR AI. Cited by: §1, §2.0.2.