Paper deep dive
Interpretable Human-Label-Free Deep Learning for Real-Bogus Classification with Uncertainty Quantification
Raphaël Bonnet-Guerrini, Bruno Sanchez, Dominique Fouchez, Benjamin Racine, Maya Guy, Mariam Sabalbal, Manal Yassine, Vincenzo Piuri
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/8/2026, 6:53:51 AM
Summary
The paper introduces a human-label-free deep learning framework for real-bogus classification in time-domain astronomy. It addresses the high cost and noise of manual labeling by using simulated supernova-like transient injections combined with heavily contaminated survey data. A dual-network model trained with asymmetric co-teaching handles class-dependent label noise. The framework also incorporates uncertainty quantification (UQ) via MC dropout, deep ensembles, and a proposed hybrid strategy, alongside latent-space visualization for interpretability. Results demonstrate robust performance under severe class contamination, competitive UQ calibration, and successful recovery of light-curve classes, offering a scalable solution for upcoming surveys like LSST.
Entities (9)
Relation Signals (7)
Asymmetric Co-teaching → usedfor → Real-Bogus Classification
confidence 95% · we introduce Asym-Co-teaching, a variant of co-teaching designed for asymmetric, class-dependent noise.
Supernova-like Transient Injection → usedtotrain → Real-Bogus Classification
confidence 94% · we propose a human-label-free approach based on the following intuition: the unlabeled survey detections are dominated by bogus examples, while realistic transient examples can be generated through source injection.
Asymmetric Co-teaching → handles → class-dependent label noise
confidence 93% · we introduce Asym-Co-teaching, a variant of co-teaching designed for asymmetric, class-dependent noise.
Uncertainty Quantification → includes → MC Dropout
confidence 92% · For uncertainty quantification (UQ), we compare MC dropout and deep ensembles
Uncertainty Quantification → includes → Deep Ensembles
confidence 92% · For uncertainty quantification (UQ), we compare MC dropout and deep ensembles
Difference Image Analysis → produces → Bogus detections
confidence 90% · In practice, most detections on difference images are spurious. Common sources of bogus detections include imperfect astrometric registration...
Latent-space Analysis → usedfor → Uncertainty Quantification
confidence 88% · we probe the model through latent-space analysis using dimensionality-reduction and visualization tools to seek global structure in the learned representation.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Time-domain surveys generate many transient candidates, making Real-Bogus classification a critical step in automated discovery pipelines. Reliable labels are costly, while community labels can be noisy and survey-dependent. We aim to develop a Real-Bogus classification framework that can be trained without human-labeled data using injected transients and bogus-dominated survey data, remains robust under strong class contamination, and provides calibrated uncertainty quantification. We combine simulated transient injections with a contaminated survey class and train a dual-network model using asymmetric co-teaching for classes with different label-noise levels. We evaluate performance on a benchmark subset and analyze the learned representation with latent-space visualization tools. For uncertainty quantification (UQ), we compare MC dropout and deep ensembles and propose a low-cost hybrid strategy that exploits the dual-network setting to improve calibration. We extend the evaluation to the light-curve domain to assess recovery of light-curve classes. The method achieves strong Real-Bogus performance on the labeled subset and remains stable under severe class contamination. It recovers transient light-curve classes with high fidelity, while single-source identification is limited by ambiguity in light-curve-derived labels. Our hybrid UQ approach achieves competitive calibration relative to more expensive ensemble baselines. Latent-space analyses indicate that uncertainty aligns with the decision boundary and reveal subclasses within the bogus population. Our results show that injection-driven, weakly supervised training can enable scalable and consistent Real-Bogus classification without human-labeled training data while providing calibrated uncertainties. The method is suited for transfer to forthcoming surveys by re-running the injection-based training pipeline.
Tags
Links
- Source: https://arxiv.org/abs/2607.05393v1
- Canonical: https://arxiv.org/abs/2607.05393v1
Trouble viewing inline? Open PDF directly →
Full Text
95,842 characters extracted from source content.
Expand or collapse full text
Astronomy& Astrophysics manuscript no. main©ESO 2026 July 7, 2026 Interpretable Human-Label-Free Deep Learning for Real-Bogus Classification with Uncertainty Quantification Raphaël Bonnet-Guerrini 1, 2 , Bruno Sanchez 2 , Dominique Fouchez 2 , Benjamin Racine 2 , Maya Guy 3, 4 , Mariam Sabalbal 5 , Manal Yassine 2 , and Vincenzo Piuri 1 1 Università degli Studi di Milano, Department of Computer Science, Milan 20133, Italy e-mail: raphael.bonnet-guerrini@unimi.it ⋆ 2 Aix Marseille Univ, CNRS/IN2P3, CPPM, Marseille, France 3 Université Côte d’Azur, INRIA, CNRS, Laboratoire J.A.Dieudonné, Maasai team, Nice 06000, France 4 Université Côte d’Azur, Observatoire de la Côte d’Azur, CNRS, Laboratoire Lagrange, Nice 06000 France 5 Université de Liège, STAR Institute, Liège 4000, Belgium ABSTRACT Context. Time-domain surveys generate large numbers of transient candidates, making Real-Bogus classification a critical step in automated discovery pipelines. Obtaining reliable human-labeled training sets is costly, while community-provided labels can be noisy and survey-dependent. Aims. We aim to develop a Real-Bogus classification framework that can be trained without human-labeled data using physically mo- tivated injected transients and bogus-dominated survey data, that remains robust under strong class contamination, and that provides well-calibrated and reliable uncertainty quantification suitable for downstream decision-making. Methods. We combine simulated transient injections with a heavily contaminated survey class and train a dual-network model using an asymmetric co-teaching strategy for classes with different label-noise levels. We evaluate performance on a labeled benchmark subset and analyze the learned representation with latent-space visualization tools. For uncertainty quantification (UQ), we compare standard approaches (MC dropout and deep ensembles) and propose a low-cost hybrid strategy that exploits the dual-network setting to improve calibration. We further extend the evaluation to the light-curve domain to assess recovery of light-curve classes. Results. The method achieves strong Real-Bogus performance on the labeled subset and remains stable under severe class con- tamination. It recovers transient light-curve classes with high fidelity, while single-source identification is fundamentally limited by intrinsic ambiguity in light-curve-derived labels. Our hybrid UQ approach achieves competitive calibration relative to substantially more expensive ensemble baselines. Latent-space analyses indicate that uncertainty aligns with the decision boundary and reveal that the model is able to identify subclasses within the bogus population. Conclusions. Our results show that injection-driven, weakly supervised training can enable scalable and consistent Real-Bogus clas- sification without human-labeled training data while providing calibrated uncertainties. The methodology is well suited for transfer to forthcoming surveys by re-running the injection-based training pipeline. Key words. Astronomical instrumentation, methods and techniques - Methods: data analysis - Methods: statistical 1. Introduction The new Vera C. Rubin Legacy Survey of Space and Time (LSST) (Abell et al. 2009; Ivezi ́ c et al. 2019) features a time-domain component that will detect transients through- out its 10-year run using Difference Image Analysis (DIA) LSST Science Collaboration et al. (2009); Ivezi ́ c et al. (2019); Liu et al. (2024). In DIA, a template image is subtracted from a newly observed image to identify changes associated with new or variable sources (Alard & Lupton 1998; Alard 2000). This method is crucial for detecting transient events in crowded fields or under varying observational conditions. Typically, supernovae occupy the short-duration region of the luminosity-timescale phase space and therefore benefit from high-cadence observations that can identify them before they fade (Kulkarni & Kasliwal 2009). Because DIA depends on ac- curate image alignment and point-spread-function matching, it also produces false positives ("bogus") arising from noise, arti- ⋆ Corresponding author facts, imperfect subtraction, cosmic rays, bad pixels, and atmo- spheric effects. Illustrated examples of spurious and successful DIA are shown in Fig. 1. In this work, we use the term source to denote a single-epoch DIA detection, and the term object to denote the association of sources across epochs at a consistent sky position; the magni- tudes of the sources associated with a given object across time form a light curve. While light curves provide the richest information for tran- sient typing, they are typically only informative after mul- tiple detections have been accumulated. Relying solely on light-curve-based methods typically implies that identification occurs only after, or at best late in, the visible window of the transient. However, many downstream science goals require rapid spectroscopic follow-up (Matheson et al. 2013; Najita et al. 2016). At LSST-scale, alert brokers and downstream transient pipelines require early transient-bogus discrimination to suppress spurious detections and prioritize candidates before applying more computationally intensive light-curve classifiers Article number, page 1 arXiv:2607.05393v1 [astro-ph.IM] 6 Jul 2026 A&A proofs: manuscript no. main (Narayan et al. 2018; Möller et al. 2020; Carrasco-Davis et al. 2021; Vuj ˇ ci ́ c et al. 2025). Historically, distinguishing real transients from bogus detec- tions has relied on a combination of algorithmic quality flags and manual human inspection. With the rate of detection expected from the Vera C. Rubin Observatory, manual inspection is not feasible (Graham et al. 2024). In recent years, machine learning has demonstrated re- markable success in various domains, including image clas- sification. Traditional image classification relied heavily on hand-crafted features and domain expertise. However, with the advent of Deep Learning (DL), convolutional neural net- works (CNNs) (LeCun et al. 1989) have become the stan- dard for image classification tasks (Hinton et al. 2012), in- cluding astronomical photometry applications (Dieleman et al. 2015; Lanusse et al. 2017). CNNs automatically learn hierar- chical features from raw pixel data, significantly improving accuracy and reducing the need for manual feature engineer- ing. Following these advances, DL methods have been devel- oped for real-bogus classification (Sedaghat & Mahabal 2018; Reyes et al. 2018; Carrasco-Davis et al. 2021; Hosenie et al. 2021; Killestein et al. 2021; Goode et al. 2022; Makhlouf et al. 2022; Chen et al. 2023). Despite these advancements, we identify two main chal- lenges for deploying DL-based methods for real-bogus classi- fication. Challenge (i) DL often requires large labeled training sets, which can be expensive and time-consuming to obtain. Real- bogus classification typically suffers from lack of labeled data, most methods in the literature therefore rely on supervised learn- ing with human-labeled datasets (Sedaghat & Mahabal 2018; Reyes et al. 2018; Carrasco-Davis et al. 2021; Hosenie et al. 2021; Goode et al. 2022; Makhlouf et al. 2022; Chen et al. 2023). While transient simulation through artificial source injec- tion is well established (Kessler et al. 2015; Sánchez et al. 2022), bogus artifacts are heterogeneous, survey-specific, and difficult to model realistically. In addition, cross-survey domain shift has been observed for real-bogus classification, implying that meth- ods trained for one telescope do not necessarily transfer directly to another (Cabrera-Vives et al. 2023). To meet the labeling de- mands of DL in the large-survey era, astronomy has increas- ingly relied on collaborative labeling. While successful, these ef- forts still incur significant overhead and produce labels with non- negligible inter-annotator disagreement, requiring careful aggre- gation and bias correction (Lintott et al. 2010). The use of simulated data has been explored in Killestein et al. (2021) for point spread functions (PSFs) of minor planets superimposed on galaxy images. To date, unsupervised approaches have been tested (Mong et al. 2022), as have active-learning methods that reduce the amount of labeled data required to achieve strong performance (Liu et al. 2025). Nevertheless, real-bogus classification remains largely constrained by the need for survey-specific human-labeled training data. To overcome challenge (i), we propose a human-label-free approach based on the following intuition: the unlabeled sur- vey detections are dominated by bogus examples, while realistic transient examples can be generated through source injection. Transient appearances can be simulated realistically at the im- age level through injections. By treating survey detections as a noisy bogus class and injected examples as transients during training, we formulate the task as a weakly supervised learn- ing (WSL) problem with class-dependent label noise. Building on existing WSL methods, we introduce Asym-Co-teaching, a variant of co-teaching designed for asymmetric, class-dependent noise. To evaluate the method, we construct training sets with different contamination levels and compare the performance of competing approaches across these settings. Challenge (i) DL models are often treated as black boxes because their decisions arise from many nonlinear transfor- mations over high-dimensional internal representations (Lipton 2016). Most explainable AI (XAI) methods explain a predic- tion by using a simplified proxy (or attribution) mechanism around a specific input (Ribeiro et al. 2016; Shrikumar et al. 2017; Binder et al. 2016; Lundberg & Lee 2017). Such tech- niques have been explored for real-bogus transient classification (Reyes et al. 2018), but they are inherently local: they provide insight for a single image and struggle to explain the global be- havior of the model. Recently, mechanistic interpretability has emerged as an approach for identifying global structures in the com- putations learned by neural networks (Elhage et al. 2021; Cammarata et al. 2020; Sharkey et al. 2025). Early work often attempted to associate individual neurons, filters, or layers with interpretable behaviors (Zeiler & Fergus 2013; Olah et al. 2017), but the recognition of polysemanticity (Elhage et al. 2022) has increasingly shifted attention toward the geometry of latent rep- resentations. In this paper, we use latent space to denote the in- ternal feature representation learned by an intermediate layer of the model. Beyond interpretability, standard deterministic deep learn- ing models usually return point estimates and do not, by default, provide well-calibrated predictive uncertainty. Yet sci- entific inference requires uncertainty quantification (UQ) that can be propagated to downstream analyses. As highlighted in (DESC et al. 2026), adapting UQ methods to astronomy-specific settings is therefore essential for trustworthy scientific analyses. This is especially critical for cosmology-oriented transients tar- geted in this study (e.g., Type Ia supernovae and other rare events used for population-level inference), where real-bogus classifi- cation occurs early in the pipeline: missed detections and, more importantly, systematic biases in detection or classification effi- ciency as a function of brightness, host-galaxy properties, or red- shift directly affect the survey selection function and can propa- gate to bias cosmological measurements. To address challenge (i), we probe the model through latent-space analysis using dimensionality-reduction and visu- alization tools to seek global structure in the learned representa- tion. Because interpretability analyses are primarily qualitative, we complement them with a systematic evaluation of uncer- tainty quantification (UQ) methods. This comparison is moti- vated by two considerations: first, epistemic uncertainty provides a natural representation of lack of knowledge in scientific infer- ence, especially when ground truth is incomplete; second, un- certainty estimates in deep learning are often fragile, potentially miscalibrated, and sensitive to distribution shift. We therefore as- sess their reliability empirically in our setting and adapt existing UQ methods to the specific structure of Asym-Co-teaching. In this paper, we address these two challenges as follows. Section 2 introduces the dataset, the injection procedure, and the construction of the ground truth. Section 3 then presents the Real-Bogus classification task, the model architecture, and the training and optimization strategy. Next, we describe in Section 4 our experimental methodology for evaluating WSL approaches including the Co-teaching framework and Asym- Article number, page 2 Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ Co-teaching, our adaptation of Co-teaching to class-dependent asymmetric noise. In Section 5, we propose a methodology to as- sess UQ in this setting and adapt existing UQ methods to Asym- Co-teaching. We present the latent space exploration and visu- alization in Section 6. The results are presented in Section 7, including UQ and WSL experiments as well as an extension to light-curve classification and inference. We discuss the implica- tions and limitations of our findings in Section 8, and conclude in Section 9. 2. Data and injection process We use imaging data from the Hyper Suprime-Cam (HSC) on the 8.2 m Subaru Telescope, located on Maunakea, Hawai‘i (Miyazaki et al. 2017). HSC comprises 104 science CCDs of size 2048 × 4096 pixels with a pixel scale of 0.168 arcsec pix −1 (Miyazaki et al. 2017). For this study, we use two HSC-based datasets processed with the LSST Sci- ence Pipelines (Sánchez et al. 2022; Liu et al. 2024), whose DIA implementation follows the image-subtraction framework of Alard & Lupton (1998). To keep end-to-end processing computationally tractable during method development, we use rc2_subset as the main training dataset. This is a subset of the HSC SSP survey, used for regular testing of the LSST Data Re- lease Production, and consists of six central science CCDs in the COSMOS field for eight visits in each of the grizy filters (40 visits in total). For inference and light-curve follow-up, we use a larger HSC-UDEEP COSMOS subset containing 103 sci- ence CCDs over 862 visits. All images are processed through the LSST Science Pipelines, including difference image analysis (DIA), source detection, and measurement. (a) Transient candidate ScienceTemplateDifference (b) Bogus detection Fig. 1: Example of difference imaging analysis (DIA) stamps. Each triplet shows the Science image (left), the Template image (middle), and the Difference (Science−Template=Difference) (right). Top: a real transient produces a compact, PSF-like pos- itive residual in the subtraction. Bottom: a bogus detection ex- hibiting a structured residual (dipole-like), characteristic of sub- traction artifacts. 2.1. Difference Image Analysis The first step of difference-image analysis (DIA) is to construct a deep template (reference) image by coadding multiple expo- sures of the same field; the template represents the static sky for that field. A science image (single-epoch exposure) is a new ob- servation of the same field acquired at a later epoch. The science image is aligned photometrically and astrometrically to the tem- plate, after which PSF matching and subtraction are performed. Ideally, the resulting difference image contains only sources that have changed between the two epochs (Fig. 1). In practice, most detections on difference images are spu- rious. Common sources of bogus detections include imperfect astrometric registration, PSF or photometric mismatches, vari- able atmospheric conditions, detector defects, cosmic rays, and subtraction artifacts. Robust Real-Bogus filtering is therefore es- sential to suppress these artifacts before downstream transient characterization. From the resulting difference images, candidate point-like detections are identified and represented as 30× 30 pixel cutouts. Each cutout is normalized independently by clipping pixel val- ues to the 1st-99th percentile range, applying an arcsinh trans- form, and then performing robust standardization using the cutout median and median absolute deviation. These cutouts form the input to our real-bogus classifier. 2.2. Supernova-Like injection Many injection pipelines generate source catalogs by sampling positions and magnitudes from generic or ad hoc distributions. In this work, we instead adopt an injection scheme tailored to supernova-like point sources. This choice reflects the scope of our analysis: our primary downstream application is the detec- tion of SNe Ia-like transients, and we therefore design the injec- tion population to approximate the region of astrophysical pa- rameter space most relevant to that use case. The resulting injec- tion model combines host-galaxy information with simple priors on source location and brightness. Host galaxies are identified using the LSST pipeline extend- edness criterion, a binary flag (0 = point-like, 1 = extended) based on a threshold on the ratio between PSF flux and model flux. For each selected galaxy we estimate the size and orienta- tion of its projected elliptical profile, and use these quantities to define the injection-position prior. The projected distance of the injection from the galaxy center is sampled from a normal distribution centered on the host posi- tion, with variance set by the semi-major axis of the galaxy’s elliptical shape. Mathematically, the offset d inj is sampled as d inj ∼ N (μ = 0,σ 2 = a 2 ), where a is the semi-major axis of the galaxy. This produces a host-associated population of syn- thetic point sources whose projected offsets follow the size and orientation of the underlying galaxy. We then sample the injected magnitude as m inj ∼ U(m host − 1, m host + 3), which is intended to span a broad range of source-to-host contrast ratios while re- maining tied to the observed host population. In addition, we include a 10% hostless component. For these injections, positions are drawn uniformly over the image foot- print, and magnitudes are sampled independently in each band from empirical distributions fitted to the UDEEP catalog. This component is intended to prevent the training set from being re- stricted to host-associated events only. The resulting injection catalogs are ingested by the LSST pipeline middleware and injected at the image level as part of the standard processing workflow 1 . The injected exposures are then passed through the same DIA, detection, and measurement steps as the real survey data. To obtain an approximately balanced training set at the de- tection level, we tune the number of injected sources so that the number of recovered injected detections is comparable to the 1 lsst.source.injection Article number, page 3 A&A proofs: manuscript no. main number of detections obtained from the corresponding real sur- vey processing. In rc2_subset, the real survey processing con- tains 43 607 detections, while our injection procedure produces N inj = 89 570 artificial point sources, of which 51 550 are recov- ered by the pipeline. The unrecovered injections predominantly populate the faint end of the injected distribution (Appendix A), and the recovered sample remains close to balanced relative to the survey detections. To verify that injections remain sparse at the image level, we estimate the affected pixel fraction. Averaging over the 40× 6 = 240 CCD images gives⟨N src ⟩ ≃ 373.2 injections per CCD. Approximating the footprint of each source by a disk of radius R = 3 FWHM ≃ 12.5 pixels affects a fractional area of about 2%. This is an approximate upper bound, since source footprints may overlap, and indicates that injections remain sparse on the scale of the full detector. We combine the recovered injection-driven detections with detections from the corresponding real survey processing (i.e. detections not associated with injected truth) to form the base- line training set. The baseline training set therefore consists of single-detection cutouts with two training labels: recovered in- jections are labeled Transient, while all survey detections are labeled Bogus, although this latter class is expected to contain genuine transients that the model aims to recover at inference time. 2.3. Constructing the evaluation set Because our method relies on a noisy class in which genuine transients are mislabeled during training, we need an indepen- dent ground truth for evaluation. We therefore construct a manu- ally curated evaluation set from HSC-UDEEP light curves. Nei- ther human labels nor light curve level information is used for model training and inference, filtering and manual inspection is used only at the evaluation stage. Since rc2_subset does not pro- vide sufficiently sampled light curves for this purpose, we de- rive the evaluation set from the more densely sampled UDEEP dataset. We first filter the UDEEP light curves to isolate well- measured SN-like candidates, and then manually inspect the re- maining objects to define the set of real detections used for eval- uation. Light-curve filtering. We apply a sequence of cuts to isolate a compact set of well-sampled transient-like light curves for man- ual inspection. First, we remove individual source measurements with signal-to-noise ratio SNR < 5. Objects with no remaining measurements above this threshold are discarded. Next, we identify the epoch of maximum flux for each object and retain only light curves whose informative measurements lie within a window of [−30, +100] days around maximum light. We then require at least six distinct observing nights with valid detections and reject objects whose mean PSF flux across re- tained epochs is negative, as these are indicative of systemati- cally spurious or over-subtracted signals. This cut may remove genuine astrophysical negative residuals or fading events, such as disappearing sources or long-term variables. These objects are outside the scope of the present high-purity SN-like curated set, which is designed to evaluate the recovery of positive transient excesses rather than to provide a complete census of variable phenomena. 2 The large number of sources discarded at this step is mostly due to single observations, which account for 2, 959, 860 sources in UDEEP. StepObjects discardedObjects remaining SNR < 5157, 1123, 699, 427 Time window568, 5403, 130, 887 6 Minimum nights3, 130, 064 2 823 Negative flux 352471 Point source host95376 Low flux ratio34342 Table 1: Summary of the filtering steps used for light-curve se- lection for human labeling. For each step, we report the number of objects discarded and the number remaining afterward. To further suppress contamination, we incorporate host- galaxy information from the HSC deep coadds. For each object, we identify the epoch of maximum absolute difference flux and match its sky position to the nearest template source within 1 ′ . Objects are retained if they are either hostless or associated with an extended host, and if the transient-to-host flux ratio exceeds 1.4 in the i band. These cuts intentionally favor well-measured SN-like transients over low-contrast or ambiguous cases and therefore define a high-purity evaluation subset rather than a rep- resentative sample of the full alert stream. The numbers of dis- carded objects are reported in Table 1. ClassObjects (N = 306) Sources (N = 4,820) SN-like33 (10.8%)1,211 (25.1%) OT41 (13.4%)1,340 (27.8%) Transient74 (24.2%)2,551 (52.9%) Bogus232 (75.8%)2,269 (47.1%) Total306 (100%)4,820 (100%) Table 2: Detailed class composition of the manually curated evaluation set after manual labeling and removal of objects la- beled as Unknown. OT denotes other transient or variable ob- jects. We report both the number of objects and the number of source detections in each class. Human labeling. The 342 filtered light curves are then in- spected manually using the coadded template image, band- averaged science and difference images, and the multi-band light curve itself (See App. B). We assign each object to one of four categories: SN-like, Other transient or variable, Bogus, or Unknown. Objects labeled Unknown are excluded from the bi- nary evaluation set, leaving 306 objects and 4,820 source detec- tions for quantitative analysis. Although 232 of the 306 objects are labeled Bogus, the eval- uation set is nearly balanced at the source-detection level be- cause transient light curves typically produce detections over more epochs. Bogus-labeled objects can also contribute multi- ple detections, for example when subtraction residuals at a fixed location near a bright star recur across epochs owing to imper- fect masking, PSF matching, or astrometric registration. These residuals often appear as dipole-like artifacts rather than isolated single-epoch events. We emphasize that this set serves as a man- ually curated evaluation set rather than spectroscopic truth. Article number, page 4 Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ x : (30, 30, 2) 30 × 30 ConvBlock 1 64 MaxPool 2× 2 Dropout2D 15 × 15 ConvBlock 2 128 MaxPool 2× 2 Dropout2D 7 × 7 ConvBlock 3 256 MaxPool 2× 2 Dropout2D Flatten 3× 3× 256 FC 1 64 ReLU Dropout 0.25 FC 2 1 Sigmoid ˆy Fig. 2: Schematic of the CNN architecture selected by hyperparameter optimization. 3. Deep Learning classification 3.1. Classes and Confusion Matrix After injection, the training data consist of two labeled groups: – injected detections, labeled as Transient; – survey detections, labeled as Bogus, although this class is contaminated by genuine transients present in the survey data. The objective of training is therefore not simply to separate injected from survey detections, but to learn a decision rule that separates transient-like from bogus-like detections despite label noise in the survey class. At training time, the confusion matrix must be interpreted relative to the observed training labels, not as a ground-truth con- fusion matrix: – True positives (TP) are injected data classified as transient. – True negatives (TN) are survey data classified as bogus. – False positives (FP) are survey data classified as transient. – False negatives (FN) are injected data classified as bogus. In this setting, the FP category is of particular interest, be- cause it contains the survey detections reclassified by the model as transient and may therefore include genuine transients hidden inside the noisy bogus-labeled class. By contrast, the FN cate- gory corresponds to injected detections that the model fails to recover as transients. Since the injected class is treated as the clean reference class during training, understanding and mini- mizing this failure mode is important. 3.2. Network Architecture The network architecture consists of three convolutional blocks (L = 3). Each block comprises a 3× 3 convolutional layer, fol- lowed by batch normalization, a ReLU activation, 2×2 max pool- ing, and spatial dropout regularization. The number of filters in the three convolutional layers is parameterized by a single base width F, and follows a geometric progression (F → 2F → 4F). This coupling enforces a monotonic increase in channel width and avoids degenerate configurations while still allowing the overall capacity of the model to be controlled through F. Each input x∈ R 30×30×2 is a two-channel cutout, constructed by stacking the difference and template images (coadded). We do not include the science image explicitly, in order to reduce computational cost and input dimensionality while preserving both the subtraction residual and the static host-galaxy context. The convolutional feature extractor computes h L = ConvBlock 3 ◦ ConvBlock 2 ◦ ConvBlock 1 (x), where each ConvBlock ℓ consists of convolution, batch nor- malization, ReLU, max pooling, and dropout. After convolu- tional feature extraction, the feature maps are flattened and passed through two fully connected layers, where the hidden di- mensionality depends on the spatial reduction factor and a con- figurable hidden size, as will be discussed next. The classifier computes f (x) = FC 2 Dropout ReLU FC 1 Flatten(h L ) , and the final prediction is obtained via ˆy = σ f (x) , with σ(·) denoting the sigmoid activation function for binary classification (transient vs. bogus detection). 3.3. Hyperparameter Optimization The CNN hyperparameters are tuned via Bayesian optimiza- tion using the Optuna framework (Akiba et al. 2019), with the objective of minimizing the loss. To keep the search statisti- cally efficient under a limited budget of 50 trials, we adopt a structured and low-dimensional hyperparameterization. Ar- chitecturally, only the base convolutional width F (which de- termines the filter configuration (F, 2F, 4F)) and the number of units in the dense layer are tuned. Similarly, a single base dropout rate (B) controls regularization throughout the network: convolutional blocks apply scaled rates ((0.5B, 0.5B, 0.5B)), while the dense layer uses B. This geometric progression en- forces smooth capacity scaling and reduces redundant architec- tural degrees of freedom. The search is performed using Optuna’s Tree-structured Parzen Estimator (TPE) sampler, introduced in (Bergstra et al. Article number, page 5 A&A proofs: manuscript no. main 2011), which models the objective function with non-parametric density estimators in hyperparameter space. To reduce computa- tional cost, the Bayesian optimization is carried out on a ran- domly selected 30% subset of the available dataset, reserved specifically for hyperparameter search. The best configuration found by the search is summarized in Table 3 and represented in Fig. 2. This configuration is used to train all the models we compare. HyperparameterValue Batch size (batch_size)128 Learning rate (learning_rate)1.62× 10 −4 Base filters F (model_params.base_filters)64 Dense units (model_params.units)64 Dense dropout (model_params.base_dropout)0.25 Table 3: Best architectural hyperparameters obtained from the Bayesian optimization. 3.4. Early Stopping and Model selection strategy Early stopping is commonly based on validation loss or vali- dation accuracy. In our setting, however, these criteria are not ideal because the survey-labeled bogus class is contaminated by genuine transients. For model selection, we instead monitor the false negative rate (FNR) on the injected class, that is, the frac- tion of injected detections classified as bogus. This quantity is not used as the training objective itself; rather, it is used as an early-stopping and checkpoint-selection criterion. The intuition is that the injected class has the most trustworthy labels avail- able during training, so preserving high recall on this class helps avoid selecting models that overfit the noisy survey labels. We therefore use injected-class FNR as a proxy criterion for check- point selection, while the network itself is still optimized using the standard training loss. This strategy biases model selection toward sensitivity to transient-like detections without explicitly rewarding memorization of the noisy bogus-labeled survey class. 4. Weakly Supervised Learning for Injection-based Transient–Bogus Classification Supervised classifiers can be sensitive to label noise. While moderate levels of corruption may be tolerated (Rolnick et al. 2017), they can bias the learned decision boundary and de- grade performance (Arpit et al. 2017). As calibration pipelines improve in modern wide-field surveys (e.g., Padmanabhan et al. 2008; Burke et al. 2017; Huang et al. 2022) and DIA implemen- tations continue to reduce subtraction artifacts and false posi- tives (Liu et al. 2024), the bogus fraction among DIA candidates may decrease, and the survey-labeled class will contain a larger fraction of genuine transients, weakening our injection-based label-purity assumptions. We therefore study the robustness of our framework under controlled one-sided label corruption by artificially contaminating the survey-labeled training class with injection-origin samples. The baseline dataset, described in Section 2.2, contains 51 550 injection-origin detections and 43 607 survey-origin de- tections, for a total of 95 157 examples. In this baseline, the in- jection provenance and the training label coincide. For each example, we retain two binary variables: (i) a prove- nance label y (stored as spy_injected), indicating whether the cutout truly originates from an injection (y = 1) or from survey data (y = 0); and (i) a training label ̃y (stored as is_injection), which is the label used for model training and may be intentionally corrupted. In the baseline dataset, y = ̃y for all samples. To construct noisy variants, we apply a one-sided, class- conditional corruption mechanism: we randomly select N mis injection-origin examples (y = 1) and flip their training label from ̃y = 1 to ̃y = 0. This contaminates the survey-data train- ing class with transient-like samples. Because labels are flipped in only one direction, we additionally subsample survey-origin examples so that the two training-label classes remain approxi- mately balanced, N( ̃y = 1)≃ N( ̃y = 0). Using this procedure, we generate three datasets with η ≃ 15%, 25%, and 35% contamination in the survey-data class. Ta- ble 4 summarizes both the training-label distribution ( ̃y) and the underlying provenance distribution (y). 4.1. Co-teaching Co-teaching (Han et al. 2018) is a method used in machine learn- ing to handle datasets with noisy labels. Two models are ini- tialized and trained simultaneously on the same dataset. During each training iteration, each model selects a subset of training samples with the smallest loss (i.e., the samples it is most confi- dent about). Each network then uses the small-loss samples se- lected by the other network to update its parameters. The intuition is that samples with smaller losses are more likely to be correctly labeled, while samples with higher losses are more likely to be noisy. This cross-teaching mechanism en- sures that networks do not reinforce their own biases and are less likely to overfit to noisy labels. Over time, as the networks learn, they become better at identifying clean samples, and the training process becomes more robust to label noise 3 . In our setting, standard Co-teaching has two practical limi- tations. First, it requires the user to specify a forget rate, i.e. the fraction of samples assumed to be corrupted; in our experiments this quantity is set using prior knowledge from the controlled- noise setup (Appendix C). Second, the sample-selection step is performed globally within each mini-batch, without explicitly accounting for class-dependent noise asymmetry. In our applica- tion, however, the noise is concentrated primarily in the survey- labeled bogus class, while the injection-labeled transient class is intended to remain comparatively clean. As a result, the standard selection rule may discard informative transient examples rather than focusing rejection on the noisy class. A formal description of standard Co-teaching is given in Appendix D. 4.2. Asymmetric Co-teaching In our application, the label noise is highly asymmetric: the frac- tion of corrupted labels is much higher in the bogus class than in the transient class. In addition, for our target use case, miss- ing a true transient is typically more costly than passing some additional false detections to downstream filtering stages. Standard Co-teaching uses a single forget rate and selects small-loss samples over the whole mini-batch, independently of their class. As a result, there is no guarantee that the high- loss samples discarded during training are predominantly noisy examples from the noisy majority class (bogus); the procedure may instead discard rare but informative transient examples that happen to incur higher loss. To address this issue, we propose 3 All parameter configurations for each Weakly Supervised Learning and Uncertainty Quantification method are presented in Appendix C. Article number, page 6 Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ Observed training labelsTrue provenance Datasetη [%]Total N( ̃y = 1)N( ̃y = 0)N(y = 1)N(y = 0)N mis (inj.-labeled)(survey-labeled)(injection-origin)(survey-origin) Baseline0.0095 15751 55043 60751 55043 6070 Low-noise15.0968 82134 39834 42339 59329 2285 195 Medium-noise24.9944 66622 34922 31727 92716 7395 578 High-noise35.2330 47515 26215 21320 6219 8545 359 Table 4: Summary of the datasets used to evaluate robustness to controlled asymmetric label noise. For each dataset, we report both the observed training labels ̃y (used for optimization) and the true provenance labels y (used only for analysis). Noise is introduced by flipping a subset of injection-origin samples from ̃y = 1 to ̃y = 0, so that the survey-labeled class contains a controlled fraction of hidden injection-origin samples. The contamination fraction is η = P(y = 1| ̃y = 0) = N mis /N( ̃y = 0), where N mis = N(y = 1, ̃y = 0) is the number of injection-origin samples relabeled as survey-labeled for training. The total number of samples decreases with increasing noise because survey-origin samples are subsampled to keep the two training-label classes approximately balanced. Asymmetric Co-teaching (Asym-Co-teaching), which extends Co-teaching by introducing class-specific forget rates and class- wise selection of small-loss samples. Co-Teaching: Keep the smallest-loss fraction ρ(t) = 1−r(t) over the whole mini-batch. Sorted samples Loss, Asymmetric Co-Teaching: Keep class-specific smallest-loss fractions ρ a (t),ρ b (t) within each class. Sorted samples within class Loss, SurveyInjected Fig. 3: Illustration of sample-selection rules in Co-teaching vari- ants. Top: standard Co-teaching applies a class-independent for- get rate and selects small-loss samples globally within the mini- batch. Bottom: Asym-Co-teaching performs class-wise small- loss selection with class-specific forget rates. Hollow markers denote discarded samples. Let: B t =(x i , ̃y i ) i∈I t denote the mini-batch at epoch t, where I t is the index set of samples in the batch and ̃y i ∈a, b is the observed training label. We define the class-specific index sets I t,a =i∈ I t : ̃y i = a,I t,b =i∈ I t : ̃y i = b. The per-sample losses of network 1 and network 2 on this batch are defined by their Binary Cross Entropy (BCE) as ℓ (1) i (t) = BCE f θ 1 (x i ), ̃y i , ℓ (2) i (t) = BCE f θ 2 (x i ), ̃y i ,i∈ I t . We introduce class-specific forget rates r a (t) = min r a · t T k , r a ! , r b (t) = min r b · t T k , r b ! , (1) Similarly to Co-teaching, the fractions of samples kept in each class at iteration t are therefore 1− r a (t) and 1− r b (t), and we set k t,a = (1− r a (t))|I t,a | ,k t,b = (1− r b (t))|I t,b | . For network 1, we sort the samples within each class accord- ing to their loss: ℓ (1) i t 1,a (t)≤ ℓ (1) i t 2,a (t)≤·≤ ℓ (1) i t |I t,a |,a (t), ℓ (1) i t 1,b (t)≤ ℓ (1) i t 2,b (t)≤·≤ ℓ (1) i t |I t,b |,b (t), where (i t 1,a ,..., i t |I t,a |,a ) is a permutation of I t,a and (i t 1,b ,..., i t |I t,b |,b ) is a permutation of I t,b . The sets of reliable samples selected by network 1 at iteration t are then R 1,t,a =i t 1,a ,..., i t k t,a ,a , R 1,t,b =i t 1,b ,..., i t k t,b ,b . Equivalently, they can be characterized as R 1,t,a = arg min S⊂I t,a |S|=k t,a X i∈S ℓ (1) i (t), R 1,t,b = arg min S⊂I t,b |S|=k t,b X i∈S ℓ (1) i (t). Analogously, using the losses ℓ (2) i (t) of network 2, we define the setsR 2,t,a ,R 2,t,b . For each mini-batch, both networks compute individual losses, rank the samples by loss within each class, and then train on the small-loss samples selected by the peer network. The Asymmetric Co-teaching losses are thus given by L (1) Asym-co (t) = 1 |R 2,t,a | +|R 2,t,b | X i∈R 2,t,a BCE f θ 1 (x i ), ̃y i + X i∈R 2,t,b BCE f θ 1 (x i ), ̃y i ! , L (2) Asym-co (t) = 1 |R 1,t,a | +|R 1,t,b | X i∈R 1,t,a BCE f θ 2 (x i ), ̃y i + X i∈R 1,t,b BCE f θ 2 (x i ), ̃y i ! . (2) Article number, page 7 A&A proofs: manuscript no. main These class-specific forget rates determine the fraction of samples discarded from each class at epoch t, with remember rates ρ c (t) = 1−r c (t) controlling the number of samples retained for training. 5. Uncertainty Quantification for injection-based Real-Bogus Classification Uncertainty quantification is fundamental to scientific infer- ence. Although machine learning can be viewed as an exten- sion of statistical modeling, obtaining uncertainty estimates that are simultaneously well calibrated, predictive, and computation- ally efficient remains challenging, particularly in deep learning (Guo et al. 2017). In this work, we evaluate and compare several uncertainty quantification approaches for injection-based real- bogus classification. 5.1. Evaluation of the uncertainty methods We evaluate uncertainty methods along two complementary axes: (i) the quality of their probabilistic predictions, assessed with standard calibration and scoring metrics; and (i) the extent to which the resulting uncertainty quantification correlates with physically meaningful indicators of classification difficulty. Calibration metrics assess whether predicted probabilities are statistically reliable estimates of true class frequencies (Niculescu-Mizil & Caruana 2005). For probabilistic evaluation, we report the negative log- likelihood (NLL), the Brier score, and the expected calibration error (ECE). NLL and Brier score are proper scoring rules, and therefore reflect both calibration and sharpness of probabilistic predictions, while ECE more directly quantifies the mismatch between predicted confidence and empirical accuracy. For binary classification, the negative log-likelihood is NLL =− 1 N N X n=1 y n log ̄p n + (1− y n ) log(1− ̄p n ) ,(3) where ̄p n denotes the predictive probability assigned to the pos- itive class for sample n, and y n ∈ 0, 1 is the observed la- bel. Lower NLL indicates better probabilistic predictions and strongly penalizes overconfident errors, making it particularly sensitive to miscalibration. The Brier score is BS = 1 N N X n=1 ( ̄p n − y n ) 2 .(4) Lower Brier scores indicate better probabilistic predictions. The expected calibration error (ECE) measures how well a model’s estimated probabilities match the true probabilities. This is done by calculating the weighted average error of the esti- mated probabilities. ECE = M X m=1 |B m | N | acc(B m )− conf(B m ) | ,(5) where B m is the set of predictions falling in bin m, acc(B m ) is the empirical accuracy in that bin, and conf(B m ) = 1 |B m | X i∈B m ̄p i . In addition to calibration, we assess the quality of our uncer- tainty quantification by examining its correlation with two physi- cally meaningful and interpretable quantities: the signal-to-noise ratio (SNR) and the distance from maximum brightness for su- pernova samples. At low SNR, the pixel-level evidence in a single-epoch difference-image cutout may be insufficient to reliably distin- guish a genuine PSF-like astrophysical transient from a chance noise fluctuation with similar morphology, implying an effective information limit for transient-bogus classification. Similarly, sources far from the peak luminosity of a super- nova are expected to be more difficult for the model to classify, because supernovae observed during their rise or decline phases have less distinctive photometric signatures. Note that this met- ric can be computed only for sources that are part of light curves labeled as SN-like (see Section 2.3), excluding sources that are unlabeled or classified as other transient or variable or Bogus. While these two quantities are correlated, it is important to note that SNR is not evenly distributed across the training classes, as shown in Appendix E.1, and is therefore expected to follow a similar tendency. Using two metrics with different class- dependent behavior makes the evaluation more robust. To deter- mine the correlation between the uncertainty estimator method and the physical quantities, we use the Spearman rank correla- tion coefficient ρ computed as: ρ = 1− 6 P n i=1 d 2 i n(n 2 − 1) ,(6) where d i = rank(x i )− rank(y i ) is the difference between the ranks of corresponding values, and n is the number of samples. 5.2. Deep Ensembles Deep Ensembles (or ensembles) provide a simple and ef- fective approach to predictive uncertainty in deep learning (Lakshminarayanan et al. 2017). Unlike explicitly Bayesian neural-network methods (MacKay 1992), they require no mod- ification of the base architecture and often provide strong em- pirical uncertainty quantification. They also differ from meth- ods such as MC dropout because independent initialization and training encourage exploration of different regions of the loss landscape (Fort et al. 2020). For our application, we are independently training CNN models. Each model, denoted by f m (·), shares the same archi- tecture (see Section 3.2) but is initialized with different random weights. All models are trained separately on the same dataset following the standard procedure described previously. At in- ference time, the ensemble prediction is obtained by averaging the individual model probabilities. Beyond improving prediction stability, the ensemble also facilitates uncertainty quantification. We quantify this uncertainty from the dispersion of the ensem- ble predictions. In practice, we use the standard deviation of the member probabilities for a given input x, with larger dispersion indicating greater predictive disagreement. Further details are given in Appendix F. 5.3. Repulsive Ensembles Repulsive Ensembles (D’Angelo & Fortuin 2021) extend stan- dard deep ensembles by introducing an interaction term during training that encourages ensemble members to represent diverse functions while still fitting the data well. In practice, the repul- sion is imposed in function space rather than directly in weight Article number, page 8 Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ space, since different parameter settings can correspond to simi- lar predictive functions. The motivation is to increase functional diversity across ensemble members, thereby improving uncer- tainty estimates relative to ensembles that remain concentrated around a narrow region of the loss landscape. At inference time, predictions are aggregated in the same way as for standard deep ensembles, and uncertainty is quantified from the dispersion across ensemble members. 5.4. MC Dropout Monte Carlo Dropout (MC Dropout) (Gal & Ghahramani 2016) is a widely used uncertainty quantification method in deep learn- ing. It can be interpreted as a particular BNN construction in which Bernoulli dropout masks induce an approximate distri- bution over the weights; running the model multiple times with dropout enabled at inference then amounts to Monte Carlo sam- pling from the resulting predictive posterior. A practical advan- tage of the method is that it requires no architectural change be- yond the use of dropout during training. The same trained net- work can then be sampled multiple times at inference by apply- ing different dropout masks. Further formal details are given in Appendix F. 5.5. Uncertainty for Co-teaching Methods As discussed above, deep ensembles provide a strong baseline for uncertainty quantification by capturing variability across dif- ferent modes of the loss landscape, and they can be further di- versified through repulsive training. However, these models are computationally expensive: an ensemble of N networks requires N full training runs and N forward passes per sample at infer- ence time. In the Co-teaching setting, two networks are already trained from different random initializations and can therefore be viewed as a small coupled ensemble, although they are not independent in the same sense as a standard deep ensemble. Be- cause an ensemble of two networks may provide only limited diversity, we augment the Co-teaching pair with MC Dropout at inference time. This allows multiple stochastic predictions from each of the two trained models and increases the number of pre- dictive samples at low additional training cost. The predictive probability now becomes: ̄p(x ∗ ) = 1 NM N X n=1 M X m=1 p n,m (x ∗ ) = 1 NM N X n=1 M X m=1 σ f n,m (x ∗ ; ˆ θ n ; z n,m ) , (7) where N = 2 is the number of Co-teaching networks, from dif- ferent initializations, composing our pseudo-ensemble, z n,m the Dropout Mask and M is the number of Dropout masks that will be applied to both models. Following the same logic, we can estimate the uncertainty from the second centered moment: Var Co-Ens-Dropout (x ∗ ) = 1 NM N X n=1 M X m=1 p n,m (x ∗ )− ̄p(x ∗ ) 2 .(8) This construction can be viewed as a heuristic mixture ap- proximation in which each trained Co-teaching network defines one component and MC Dropout provides local stochastic sam- ples around that component. In a Bayesian framework as a mixture approximate posterior, the N networks converge to different solutions (modes) ˆ θ n N n=1 due to independent initializations and training dynamics, while MC Dropout defines a local variational distribution around each configuration by keeping dropout active at test time. This moti- vates the approximation q(θ) = 1 N N X n=1 q n (θ).(9) with corresponding predictive approximation p(y ∗ | x ∗ ,D)≈ Z p(y ∗ | x ∗ ,θ) q(θ) dθ(10) = 1 N N X n=1 Z p(y ∗ | x ∗ ,θ) q n (θ) dθ(11) ≈ 1 N N X n=1 E z∼q(z) p(y ∗ | x ∗ , ˆ θ n , z) (12) ≈ 1 NM N X n=1 M X m=1 p y ∗ | x ∗ , ˆ θ n , z n,m ,(13) where z n,m ∼ q(z) denotes a dropout mask (sampled indepen- dently for each network n and Monte Carlo pass m). 6. Latent Space Representation Analysis of learned internal representations can provide useful insight into model behavior. In this work, we study the latent representation learned by the classifier by extracting hidden ac- tivations and visualizing them in two dimensions. This analysis is intended as a global representation-level probe of the model, rather than a full mechanistic interpretation of its internal com- putations (Sharkey et al. 2025). While much recent work focuses on identifying interpretable concepts in latent representations using sparse autoencoders (SAEs) (Wu & Walmsley 2025), we adopt a simpler approach based on direct exploration of latent activations through nonlin- ear dimensionality reduction. 6.1. Feature Extraction The dimensionality of the latent representation depends on the layer under consideration. We focus on the output of the final hidden layer of the network (Fig. 2), where we expect high- level abstractions and task-relevant patterns to be strongly repre- sented. For each input sample, we record the activation vector at this layer. For a dataset of N samples, this gives a feature matrix F∈ R N×d , where d denotes the dimensionality of the selected layer’s output. When N exceeds a predefined maximum sample size n max , we apply stratified random sampling to obtain a representative subset while preserving the class distribution. For ensemble- based and co-teaching models, features are extracted from the first model in the ensemble. 6.2. Dimensionality Reduction with UMAP The high dimensionality of F makes direct interpretation infea- sible. We therefore use Uniform Manifold Approximation and Projection (UMAP) (McInnes et al. 2020) to compute a two- dimensional embedding suitable for visualization. Article number, page 9 A&A proofs: manuscript no. main Fig. 4: Two-dimensional UMAP projection of the embeddings from the penultimate dense layer of the classifier. Each point corre- sponds to one candidate and is colored by outcome relative to the labels: TP (yellow×), TN (red•), FP (orange ⋆), and FN (gray■). Insets (a-e) show representative Coadd (template) and Diff (difference) cutouts sampled from the indicated regions: (a) negative-Diff region, dominated by negative residuals characteristic of subtraction artifacts or over-subtraction; (b) low-SNR region, where candi- dates are difficult to distinguish from background fluctuations; (c) dipole region, showing positive/negative residual astrometric mis- alignment; (d) low-SNR transient region, containing faint point-source-like residuals consistent with marginal transient detections; and (e) high-SNR transient region, containing compact, symmetric positive residuals characteristic of confident transient-like candi- dates. The cutouts are provided as illustrative examples only, they were assembled manually to make the static figure informative and are therefore approximate with respect to exact UMAP locations and local neighborhoods. The interactive versions of the UMAP visualizations (with different overlays) are available in the Zenodo supplementary material (DOI:10.5281/zenodo.18434676). UMAP constructs a weighted nearest-neighbor graph in the original feature space and optimizes a low-dimensional embed- ding that preserves local neighborhood structure. As with other nonlinear embedding methods, the visualization is primarily use- ful for exploring local organization; global distances in the 2D projection should not be over-interpreted. In contrast to linear projections such as PCA (Wold et al. 1987), UMAP can capture nonlinear structure; similarly to t- SNE (van der Maaten & Hinton 2008), it primarily emphasizes preservation of local neighborhood relationships. In practice, UMAP often produces embeddings with more apparent global continuity than t-SNE, although global distances are not guaran- teed to be preserved. We compute a two-dimensional UMAP embedding with n_components=2 and Euclidean distance in the original feature space. We vary n_neighbors and min_dist over a small grid: – n_neighbors: size of the local neighborhood used to approx- imate the manifold, – min_dist: minimum distance between embedded points, controlling cluster compactness, – n_components: target embedding dimensionality (set to 2), – metric: Euclidean distance in the original feature space. Specifically, we perform a grid search over n_neighbors∈5, 10, 15, 30, 50, min_dist∈0.0, 0.1, 0.25, 0.5. and select the configuration using a simple visualization heuris- tic based on the spread of the embedding coordinates. This heuristic is used only to avoid overly collapsed visualizations and should not be interpreted as an objective measure of repre- sentation quality. Using an interactive graph built with the Bokeh library, we visualize the UMAP representation together with the corre- sponding confusion-matrix categories, as well as image SNR, predicted class probability, and model uncertainty when avail- able. This interactive tool facilitates qualitative exploration of Article number, page 10 Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ model behavior and helps identify structured failure modes or mislabeled examples. The resulting embeddings are used as a qualitative analysis tool to inspect class structure, misclassifi- cations, uncertainty patterns, and potential label noise in the learned representation. 7. Results 7.1. Weakly supervised learning Table 5 compares weakly supervised learning (WSL) approaches under increasing label-noise rates on training sets, reporting subtype-specific recall for the positive class (survey transients) and specificity for the negative class (bogus detections), together with aggregate accuracy and receiver operating characteristic (ROC). When trained on the baseline dataset (current HSC injected survey data), all methods achieve strong performance (Acc ≈ 0.91-0.92; ROC ≈ 0.97), with Co-teaching and Asym-Co- teaching providing the strongest overall balance across metrics. As we artificially increase the label noise in the baseline, the Standard model degrades substantially (e.g., at 35% noise, Acc = 0.812 and ROC = 0.903), whereas the co-teaching variants remain markedly more robust. In particular, Co-teaching and Asym-Co-teaching consistently maintain high ROC AUC across noise levels (≈ 0.95) and improve accuracy relative to the Stan- dard approach. At higher noise (25-35%), Asym-Co-teaching offers the best aggregate results (Acc≈ 0.87-0.89; ROC≈ 0.95), indicating that co-teaching, especially with asymmetric handling of label noise, provides the most stable performance as annotation quality dete- riorates. The comparable recall obtained for the SN-like and OT sub- classes supports the hypothesis that, at the single-image level, SN-like injections provide useful training examples for detect- ing other transient classes as well. In addition to the two methods presented in Section 4, we also evaluated Stochastic Co-teaching (de Vos et al. 2023), which we briefly describe in Appendix D. It exhibits pronounced trade-offs across the noise regimes. This instability was observed throughout training, with difficulty in tuning the α and β parameters and in selecting the delay before rejection begins. In most training cases, Stochastic Co-teaching collapsed to rejecting predominantly one class, leading to unsta- ble training. 7.2. Uncertainty Quantification Table 6 summarizes probabilistic performance and calibration for the different uncertainty-quantification strategies. All UQ approaches improve over the standard deterministic baseline, with consistent reductions in NLL, Brier score, and ECE. The standard CNN achieves an NLL of 0.453 and a Brier score of 0.074, illustrating the comparatively poorer probabilistic quality of deterministic predictions. This behavior is easily noticeable in Fig. 5. Our method, Ensemble-MC-Dropout for Co-teaching (Sec- tion 5.5) achieves the best overall performance across the three reported metrics, with the lowest NLL (0.292), the lowest Brier score (0.063), and the lowest ECE (0.0375). Deep Ensembles and Repulsive Ensembles remain competitive, particularly on NLL and Brier score, but are consistently outperformed by our approach, with the clearest advantage observed in ECE. Rela- tive to the standard model, this corresponds to a 35.5% reduc- tion in NLL, a 14.9% reduction in Brier score, and a 40.8% re- NoiseWSL methodSN-likeOTBAccROC B Standard0.8850.9140.9260.9120.966 Co-t0.9190.9280.9150.9200.972 Asym-Co-t0.8980.9050.9440.9220.971 15% Standard0.7980.8150.9340.8660.946 Co-t0.9200.9250.8720.8990.963 Asym-Co-t0.8430.8380.9320.8830.959 25% Standard0.7920.7530.9370.8490.939 Co-t0.8820.8970.8990.8940.959 Asym-Co-t 0.8870.8750.9030.8910.951 35% Standard0.7640.7610.8680.8120.903 Co-t0.8070.7790.8760.8310.921 Asym-Co-t 0.8120.8330.9250.8710.950 Table 5: Evaluation of weak supervision learning methods un- der varying label noise conditions. For the positive class (survey transients), recall is reported separately for Supernova like (SN- like) and other transient or variable types (OT). For the negative class, specificity is shown for bogus detections (B). Overall ac- curacy (Acc) and ROC AUC summarize aggregate performance. duction in ECE, indicating both improved probabilistic accuracy and substantially better calibration. These results indicate that our Ensemble-MC-Dropout inference strategy, combined with dual-network architecture, produces probability estimates that more accurately reflect true class frequencies. The lower NLL, in particular, demonstrates that our method effectively reduces overconfident incorrect predictions, a critical property for down- stream tasks. The simultaneous improvement in three metrics confirms that our uncertainty quantification approach achieves superior calibration without sacrificing robustness to extreme probabilities. MetricNLLBrier ScoreECE Standard0.4530.0740.0633 MC Dropout0.3340.0690.0405 Deep Ensemble 0.3140.0650.0535 Repulsive Ensemble0.3080.0650.0511 Our method0.2920.0630.0375 Table 6: Negative Log-Likelihood (NLL), Brier Score and Expected Calibration Error (ECE) on 11 bins for the stan- dard model, the Deep Ensemble, Repulsive Ensemble and our method, Ensemble-MC-Dropout for Co-teaching In addition to calibration, a useful uncertainty estimator should assign larger uncertainty to samples that are intrinsically more difficult to classify, such as low-SNR detections or super- novae observed far from peak brightness. Table 7 reports Spear- man correlations between uncertainty and these two physically meaningful quantities. All methods show the expected negative correlation with SNR (ρ < 0), indicating that uncertainty increases as signal quality degrades. Our method achieves the strongest correla- tion (ρ = −0.296), marginally outperforming Deep Ensembles (ρ = −0.290) and substantially improving upon MC Dropout (ρ =−0.239). All methods also exhibit the expected positive correlation with distance from maximum brightness. Here, Deep Ensembles achieve the highest correlation (ρ = 0.221), with our method performing nearly identically (ρ = 0.218). Article number, page 11 A&A proofs: manuscript no. main 0.0 0.2 0.4 0.6 0.8 1.0 Fraction of Positives Perfectly calibrated Standard MC Dropout Ensemble Rep Ensemble Asym Co-teaching 0.00.20.40.60.81.0 Mean Predicted Probability 0.50 0.25 0.00 0.25 0.50 Residual Fig. 5: The top panel shows the fraction of positive samples as a function of the mean predicted probability, using 11 equally spaced probability bins. The dashed diagonal indicates perfect calibration. The bottom panel shows the residuals with respect to perfect calibration. Overall, these results show that the proposed method pro- duces uncertainty quantification that are physically meaningful and competitive with substantially more expensive ensemble- based baselines, while requiring only two jointly trained net- works augmented with MC Dropout at inference time, which greatly reduces training cost compared with large ensembles. Physical valueρ SNRρ Max. bright MC Dropout−0.2390.176 Deep Ensemble−0.2900.221 Repulsive Ensemble−0.2710.198 Our method−0.2960.218 Table 7: Spearman rank correlation coefficients (ρ) between pre- dicted uncertainty and signal-to-noise ratio (SNR) and distance from maximum brightness (Max. bright, in days). Negative ρ SNR values indicate that uncertainty appropriately increases for low- SNR detections, while positive ρ Max. bright values indicate higher uncertainty for supernovae observed far from peak luminosity. It is also informative to examine the class-conditional val- ues of ρ SNR , reported in Appendix E.2. The overall correlation between uncertainty and SNR is driven primarily by the true- positive subset. A weaker but still visible trend remains in the true-negative subset, whereas the correlation is much less pro- nounced for misclassified examples. This suggests that SNR is most directly informative for un- certainty within the transient-like detections, where low signal quality more strongly affects classification confidence. This in- terpretation is consistent with the class-dependent SNR distri- butions shown in Appendix E.1, and with the latent-space vi- sualization in App. G (interactive version available on Zen- odo), where low-SNR samples appear more diffusely distributed within the transient region than in the bogus region. 7.3. Latent space representation The projection displayed in Fig. 4 highlights a broad separation between negative and positive manifolds, with most ambiguous cases and misclassifications concentrated near their interface, thereby shaping the decision boundary. One key finding of the UMAP representation of the latent space is that it allows the user to understand which patterns are being identified by the model even within a class. In addition to the class information, UQ and SNR can be displayed (see Ap- pendix G). The visualization is also useful for inspecting misclassified cutouts and identifying structured failure modes or potential la- bel noise. We therefore use it as a qualitative complement to the quantitative analyses, rather than as a standalone measure of model performance. 7.4. Towards object identification from source classifications Although the model classifies DIA sources independently, its predictions can be grouped at the object level to perform light- curve-level transient-bogus discrimination. Using the method described in Section 5.5, we aggregate source-level predictions within each labeled light curve and assign an object-level label by majority vote. The resulting performance is summarized in Table H.1. Figure 6 shows the distribution, for each light curve in the la- beled evaluation set, of the fraction of sources classified as tran- sient; the corresponding cumulative counts are shown in App. H. A first notable result is the strong internal consistency of the source-level predictions within individual light curves: 232/306 light curves (75.8%) have at least 90% of their sources assigned to the correct class. For the transient class specifically, 79.7% of light curves have more than 80% of their sources classified as transient. On this curated evaluation set, aggregation of the source- level predictions yields very strong object-level performance: 304/306 light curves are assigned the correct binary label (99.35% accuracy), with only two errors overall (one false posi- tive and one false negative). At the source level, class separation also remains strong, with similar recall for supernovae and other transients or variables (0.887 and 0.896, respectively) and a bo- gus specificity of 0.951. The uncertainty quantification remains consistent with ob- servational difficulty. For supernova detections, uncertainty is lowest near maximum light (mean uncertainty ≃ 0.074) and increases substantially for observations far from peak (> 50 days; mean uncertainty ≃ 0.180), where the corresponding ac- curacy falls to 0.75. This behavior is consistent with the uncer- tainty analysis presented above and suggests that epoch-level un- certainty could be useful for down-weighting ambiguous, low- information measurements in downstream light-curve analyses. Together, these results indicate that single-epoch, per- detection classification scores (and their associated uncertain- ties) are sufficiently stable to support reliable light-curve iden- tification after grouping, while also providing a principled signal for down-weighting ambiguous, low-information epochs during downstream light-curve-based analyses. Encouraged by the strong performance obtained on the la- beled set (Table H.1), we then deploy the same grouping rule as an inference-time filter on the full UDEEP dataset, which con- tains 3, 856, 539 distinct objects. Among these, 269, 968 objects have more than 50% of their sources classified as transient. How- ever, a more detailed inspection shows that most of these objects Article number, page 12 Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ 020406080100 Predicted transient (%) 0 20 40 60 80 100 120 140 160 180 Count a) All light curves 020406080100 Predicted transient (%) b) Transient light curves 020406080100 Predicted bogus (%) c) Bogus light curves Fig. 6: Light-curve-level classification score distributions. Shown are the distributions of the fraction of source detections classified as transient for all light curves (a) and for true transient light curves (b), together with the distribution of the fraction of source detections classified as bogus for true bogus light curves (c). are observed on only a single night. Requiring at least five dif- ferent nights of observation reduces this set to 1, 770 transient candidates, including 108 with more than 50 distinct observa- tions. 8. Discussion and open questions The aim of this work was to address the two challenges intro- duced in Section 1: learning real-bogus classification without human-labeled training data, and obtaining uncertainty quantifi- cation and representations that remain interpretable and useful for downstream analysis. Our experiments show that the pro- posed framework recovers strong source-level performance and, after grouping detections, very strong object-level performance on the manually curated evaluation set (Section 7.4). Although the single-source metric may appear less reliable, it is important to note that this limitation is not solely algorithmic. Indeed, our labels are derived from observed light curves, and a fraction of events is intrinsically ambiguous given the available observables. This sets an effective upper bound on achievable performance for the single-source task. Inspection of misclassi- fied cutouts with the interactive visualization supports this in- terpretation. Combined with in-depth latent-space visualization, these results help explain the observed behavior. A key strength of the method is its portability across surveys with limited adaptation. In the context of forthcoming datasets such as LSST, it remains an open question whether direct in- ference on survey data, despite the expected domain shift, is sufficient, or whether performance and calibration are best pre- served by re-running the full pipeline with injection, training, and inference using survey-specific simulations. In both cases, the methodology is survey-agnostic, with the main challenges being adaptation to other data-management tools, as we have shown that our method is robust to extreme noise rates. Beyond the present scope, weakly supervised methods ap- pear promising both for our current approach, fully human label- free training, and as a way to mitigate the inherent label noise expected in community-labeled datasets. In parallel, asymmetric co-teaching could potentially be leveraged to estimate dataset noise levels, for example by exploiting the training dynam- ics to infer class-dependent noise rates and, in turn, a robust Real-Bogus ratio. While noise-rate estimation is discussed in de Vos et al. (2023) for a global noise level, extending it to class- conditional noise in survey data is an interesting direction for future work. Regarding uncertainty quantification (UQ), it is noteworthy that our approach matches or outperforms standard baselines, including MC dropout and large ensembles. Given that ensem- bles with N = 50 are expected to explore a broader area of the loss landscape than methods that differ mainly at inference time (Fort et al. 2020), this result suggests that the training coupling induced by Asym-Co-teaching may play a central role. An im- portant question for future work is whether the gain arises pri- marily from Asym-Co-teaching and the interaction between the two models during training, or whether the combination of MC dropout and a small ensemble could serve as a robust alternative to large ensembles. Finally, we find that UQ correlates with physically motivated quantities and improves calibration. Consistent with this, latent- space projections indicate that the UQ captures the expected de- cision boundary, whereas SNR is less structured in representa- tion space, particularly in the bogus region. This supports the hy- pothesis that the model relies on features that are not well sum- marized by SNR alone, motivating further work to identify the dominant image and contextual factors driving both predictions and uncertainties. In addition, it would be valuable to explore concept-based interpretability methods applied to the learned representations (Kim et al. 2018), as well as more advanced methods such as (Wu & Walmsley 2025), to identify human-interpretable con- cepts associated with both astrophysical transients and bogus de- tections. Such concepts may guide targeted diagnostics, improve the realism of injections, and potentially lead to explicit bogus simulation. 9. Conclusion We presented a human-label-free approach to Real-Bogus clas- sification for transient candidates. The method is trained using physically motivated injected transient classes together with a highly contaminated survey class, enabling learning without re- lying on human-labeled training data. This setup allows us to build on weakly supervised learning and to develop a training strategy dedicated to an asymmetric-noise regime. We show that the approach remains robust under noise levels exceeding the transient-to-bogus ratio typical of current pipelines, supporting its applicability to future surveys. Article number, page 13 A&A proofs: manuscript no. main We also presented a comparative study of uncertainty quan- tification methods using both calibration and physically mo- tivated metrics. Building on the dual-network setting of co- teaching, we introduced a hybrid strategy combining elements of deep ensembles and MC dropout. This approach reduces com- putational cost while achieving UQ performance comparable to, or better than, the state-of-the-art baselines considered here. To support interpretation, we developed a visualization framework enabling in-depth analysis of model behavior. Finally, we ex- tended our evaluation to light-curve space, where we obtain near- perfect classification accuracy on the labeled set. Author contributions. R.B.G. contributed to the data processing with the LSST Gen3 pipeline, and carried out the methodological develop- ment, including the physically motivated injection strategy, weakly su- pervised learning framework, and uncertainty quantification methods. He performed the main analysis, implemented the ML4transients li- brary, and wrote the manuscript. B.S. provided expertise on the LSST Gen2 and Gen3 difference-imaging pipeline, including resolving com- patibility issues, supporting database management, and supervising the project. D.F. conceived the project, initiated the data-processing strat- egy and latent-space analysis, and provided overall scientific supervi- sion. B.R. initiated the analysis of light curves and difference-image cutouts and implemented an initial CNN model for real–bogus classi- fication. M.Y. led the initial data-processing effort, running the LSST Gen2 pipeline for HSC and developing an early source-injection frame- work. M.S. further developed and systematically evaluated the CNN, including performance studies and the incorporation of additional in- puts. M.G. improved the CNN through hyperparameter optimization and developed the UMAP-based analysis and visualization tools. V.P. provided guidance on interpretability. Acknowledgements. Part of this work was supported by the European Union’s Horizon Europe research and innovation programme under the Marie Sklodowska-Curie grant agreement No 101168829, Challenging AI with Chal- lenges from Physics: How to solve fundamental problems in Physics by AI and vice versa (AIPHY). References Abell, P. A., Allison, J., Anderson, S. F., et al. 2009, LSST Science Book, Version 2.0, Tech. rep., LSST Science Collaboration, 596 pages. Also available at full resolution at http://w.lsst.org/lsst/scibook Akiba, T., Sano, S., Yanase, T., Ohta, T., & Koyama, M. 2019, Optuna: A Next- generation Hyperparameter Optimization Framework Alard, C. 2000, A&AS, 144, 363 Alard, C. & Lupton, R. H. 1998, The Astrophysical Journal, 503, 325–331 Arpit, D., Jastrz ̨ebski, S., Ballas, N., et al. 2017, in Proceedings of Machine Learning Research, Vol. 70, Proceedings of the 34th International Conference on Machine Learning, ed. D. Precup & Y. W. Teh (PMLR), 233–242 Bergstra, J., Bardenet, R., Bengio, Y., & Kégl, B. 2011, in Advances in Neural Information Processing Systems, ed. J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, & K. Weinberger, Vol. 24 (Curran Associates, Inc.) Binder, A., Montavon, G., Bach, S., Müller, K.-R., & Samek, W. 2016, Layer- wise Relevance Propagation for Neural Networks with Local Renormaliza- tion Layers Burke, D. L., Rykoff, E. S., Allam, S., et al. 2017, The Astronomical Journal, 155, 41 Cabrera-Vives, G., Bolivar, C., Förster, F., et al. 2023, Domain Adaptation via Minimax Entropy for Real/Bogus Classification of Astronomical Alerts Cammarata,N.,Carter,S.,Goh,G.,etal.2020,Distill, https://distill.pub/2020/circuits Carrasco-Davis, R., Reyes, E., Valenzuela, C., et al. 2021, The Astronomical Journal, 162, 231 Chen, Z., Zhou, W., Sun, G., et al. 2023, TransientViT: A novel CNN - Vision Transformer hybrid real/bogus transient classifier for the Kilodegree Auto- matic Transient Survey D’Angelo, F. & Fortuin, V. 2021, in Proceedings of the 35th International Con- ference on Neural Information Processing Systems, NIPS ’21 (Red Hook, NY, USA: Curran Associates Inc.) de Vos, B. D., Jansen, G. E., & Išgum, I. 2023, Scientific Reports, 13, 16875 DESC, L., Aubourg, E., Avestruz, C., et al. 2026, Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration Dieleman, S., Willett, K. W., & Dambre, J. 2015, Monthly Notices of the Royal Astronomical Society, 450, 1441–1459 Elhage, N., Hume, T., Olsson, C., et al. 2022, Toy Models of Superposition Elhage, N., Nanda, N., Olsson, C., et al. 2021, Transformer Circuits Thread, https://transformer-circuits.pub/2021/framework/index.html Fort, S., Hu, H., & Lakshminarayanan, B. 2020, Deep Ensembles: A Loss Land- scape Perspective Gal, Y. & Ghahramani, Z. 2016, Dropout as a Bayesian Approximation: Repre- senting Model Uncertainty in Deep Learning Goode, S., Cooke, J., Zhang, J., et al. 2022, Monthly Notices of the Royal As- tronomical Society, 513, 1742 Graham, M. L., Bellm, E. C., Guy, L. P., et al. 2024, Rubin Observatory Technical Note DMTN-102 Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. 2017, On Calibration of Mod- ern Neural Networks Han, B., Yao, Q., Yu, X., et al. 2018, Co-teaching: Robust Training of Deep Neural Networks with Extremely Noisy Labels Hinton, G., Krizhevsky, A., Sutskever, I., & Rachmad, Y. 2012, Advances in Neural Information Processing Systems, 1097 Hosenie, Z., Bloemen, S., Groot, P., et al. 2021, Experimental Astronomy, 51, 319–344 Huang, B., Xiao, K., & Yuan, H. 2022, Photometric calibration methods for wide-field photometric surveys Ivezi ́ c, Ž., Kahn, S. M., Tyson, J. A., et al. 2019, ApJ, 873, 111 Kessler, R., Marriner, J., Childress, M., et al. 2015, The Astronomical Journal, 150, 172 Killestein, T. L., Lyman, J., Steeghs, D., et al. 2021, Monthly Notices of the Royal Astronomical Society, 503, 4838–4854 Kim, B., Wattenberg, M., Gilmer, J., et al. 2018, Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV) Kulkarni, S. R. & Kasliwal, M. M. 2009, Transients in the Local Universe Lakshminarayanan, B., Pritzel, A., & Blundell, C. 2017, Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles Lanusse, F., Ma, Q., Li, N., et al. 2017, Monthly Notices of the Royal Astronom- ical Society, 473, 3895 LeCun, Y., Boser, B., Denker, J. S., et al. 1989, Neural computation, 1, 541 Lintott, C., Schawinski, K., Bamford, S., et al. 2010, Monthly Notices of the Royal Astronomical Society, 410, 166–178 Lipton, Z. C. 2016, CoRR, abs/1606.03490 [1606.03490] Liu, S., Wood-Vasey, W. M., Armstrong, R., et al. 2024, The Astrophysical Jour- nal, 967, 10 Liu, Y., Fan, L., Hu, L., et al. 2025, Astronomy & Astrophysics, 693, A105 LSST Science Collaboration, Abell, P. A., Allison, J., et al. 2009, arXiv e-prints, arXiv:0912.0201 Lundberg, S. & Lee, S.-I. 2017, A Unified Approach to Interpreting Model Pre- dictions MacKay, D. J. 1992, Neural Computation, 4, 448 Makhlouf, K., Turpin, D., Corre, D., et al. 2022, Astronomy & Astrophysics, 664, A81 Matheson, T., Fan, X., Green, R., et al. 2013, Spectroscopy in the Era of LSST McInnes, L., Healy, J., & Melville, J. 2020, UMAP: Uniform Manifold Approx- imation and Projection for Dimension Reduction Miyazaki, S., Komiyama, Y., Kawanomoto, S., et al. 2017, Publications of the Astronomical Society of Japan, 70, S1 Mong, Y.-L., Ackley, K., Killestein, T. L., et al. 2022, Monthly Notices of the Royal Astronomical Society, 518, 752–762 Möller, A., Peloton, J., Ishida, E. E. O., et al. 2020, Monthly Notices of the Royal Astronomical Society, 501, 3272–3288 Najita, J., Willman, B., Finkbeiner, D. P., et al. 2016, Maximizing Science in the Era of LSST: A Community-Based Study of Needed US Capabilities Narayan, G., Zaidi, T., Soraisam, M. D., et al. 2018, The Astrophysical Journal Supplement Series, 236, 9 Niculescu-Mizil, A. & Caruana, R. 2005, in Proceedings of the 22nd Interna- tional Conference on Machine Learning, ICML ’05 (New York, NY, USA: Association for Computing Machinery), 625–632 Olah,C.,Mordvintsev,A.,&Schubert,L.2017,Distill, https://distill.pub/2017/feature-visualization Padmanabhan, N., Schlegel, D. J., Finkbeiner, D. P., et al. 2008, The Astrophys- ical Journal, 674, 1217–1233 Reyes, E., Estevez, P. A., Reyes, I., et al. 2018, in 2018 International Joint Con- ference on Neural Networks (IJCNN) (IEEE), 1–8 Ribeiro, M. T., Singh, S., & Guestrin, C. 2016, "Why Should I Trust You?": Explaining the Predictions of Any Classifier Rolnick, D., Veit, A., Belongie, S., & Shavit, N. 2017, in Advances in Neural Information Processing Systems Article number, page 14 Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ Sedaghat, N. & Mahabal, A. 2018, Monthly Notices of the Royal Astronomical Society, 476, 5365–5376 Sharkey, L., Chughtai, B., Batson, J., et al. 2025, Open Problems in Mechanistic Interpretability Shrikumar, A., Greenside, P., Shcherbina, A., & Kundaje, A. 2017, Not Just a Black Box: Learning Important Features Through Propagating Activation Differences Sánchez, B. O., Kessler, R., Scolnic, D., et al. 2022, The Astrophysical Journal, 934, 96 van der Maaten, L. & Hinton, G. 2008, Journal of Machine Learning Research, 9, 2579 Vuj ˇ ci ́ c, V., Sre ́ ckovi ́ c, V. A., Babarogi ́ c, S., & Aleksi ́ c, J. 2025, Contributions of the Astronomical Observatory Skalnate Pleso, 55, 95 Wold, S., Esbensen, K., & Geladi, P. 1987, Chemometrics and Intelligent Lab- oratory Systems, 2, 37, proceedings of the Multivariate Statistical Workshop for Geologists and Geochemists Wu, J. F. & Walmsley, M. 2025, Re-envisioning Euclid Galaxy Morphology: Identifying and Interpreting Features with Sparse Autoencoders Zeiler, M. D. & Fergus, R. 2013, Visualizing and Understanding Convolutional Networks Article number, page 15 A&A proofs: manuscript no. main Appendix A: Injection magnitude We compare the magnitude distributions of injected sources, re- covered injections, and real detections with positive flux. As ex- pected, the recovered injections are systematically brighter than the full injected population, reflecting the detection efficiency of the pipeline at faint magnitudes. The positive-flux real detec- tions have a broadly comparable magnitude distribution, and the resulting sample remains approximately balanced between real and recovered injected detections. 0 1000 2000 3000 4000 5000 6000 Number of Sources N = 86,973 = 24.76 = 1.94 (a) Injection Catalog grizy 0 1000 2000 3000 4000 5000 Number of Sources N = 51,520 = 23.91 = 1.44 Recov. = 59.2% (b) Injections After Pipeline grizy 1416182022242628 AB Magnitude 0 500 1000 1500 2000 Number of Sources N = 22,558 = 24.86 = 1.88 (c) Real Sources After Pipeline grizy Fig. A.1: Magnitude distributions for (a) injected sources, (b) recovered injections after the pipeline, and (c) real sources with positive flux. In each panel we report the number of sources N and the best-fitting Gaussian parameters (μ, σ) describing the distribution. Article number, page 16 Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ Appendix B: Light curve labeling tool Examples of the actual labeling tool are available in the Zenodo supplementary material (DOI: 10.5281/zenodo.18434676). Fig. B.1: Representative examples of a bogus candidate (top) and a supernova-like transient (bottom). Left: multi-band light curves in detected bands as a function of MJD. Right: examples of associated cutouts shown. For each object from top to bottom: coadd image, mean of the science for this band and mean of the difference image for the band. Appendix C: Training configurations For reproducibility, we summarize below the training hyperpa- rameters used for each uncertainty-estimation or robust-training method considered in this work. Unless stated otherwise, all set- tings follow the default values described in the main text, and we report only the parameters specific to each method (num- ber of models or forward passes, repulsion strength and type, dropout probabilities, and the co-teaching forget-rate sched- ules). For the forget rates, the reported values correspond to the Baseline, 15%, 25%, 35% sets. – Ensemble: num_models = 50. – Repulsive ensemble: num_models = 50, λ rep = 0.1, repulsion_type = "prediction", T = 1.0. – MC Dropout: fwd_passes = 50, p dense = 0.25, p conv = 0.125. – Co-teaching: forget_rate∈0.06, 0.16, 0.26, 0.36. – Asym. Co-teaching: (forget_rate 0 , forget_rate 1 ) ∈ (0.05, 0.01), (0.15, 0.01), (0.25, 0.01), (0.35, 0.01). We select the forget-rate schedules using prior knowledge es- timated from the label noise inferred by the standard training method. In particular, we use the observed false-positive and false-negative fractions on the training set as approximate indi- cators of mislabeling. This suggests a noise level of about 5% for the survey (baseline) class and about 1% for the injected class; the reported forget rates are chosen to reflect these estimates. Appendix D: WSL formalism D.1. Co-teaching Formally, letB t = (x i , ̃y i ) i∈I t denote the mini-batch at iteration (or epoch) t, where I t is the set of indices of the samples in the batch and|B t | =|I t |. The per-sample losses of network 1 and network 2 on this batch are defined by their Binary Cross Entropy (BCE) as ℓ (1) i (t) = BCE f θ 1 (x i ), ̃y i , ℓ (2) i (t) = BCE f θ 2 (x i ), ̃y i , i∈ I t . The forget rate is defined as r(t) = min r· t T k , r ! ,(D.1) where T k is a predefined epoch after which the forget rate re- mains constant. Given the forget rate, the fraction of samples kept in the batch at iteration t is 1− r(t). We set k t = (1− r(t))|B t | = (1− r(t))|I t | . For network 1, we sort the indices in I t according to the in- creasing order of its loss: ℓ (1) i t 1 (t)≤ ℓ (1) i t 2 (t)≤·≤ ℓ (1) i t |I t | (t), where (i t 1 ,..., i t |I t | ) is a permutation of I t . The set of reliable samples selected by network 1 at iteration t is then defined as R 1,t =i t 1 ,..., i t k t = arg min S⊂I t |S|=k t X i∈S ℓ (1) i (t). The Co-teaching losses are then given by L (1) co-teach (t) = 1 |R 2,t | X i∈R 2,t BCE f θ 1 (x i ), ̃y i , L (2) co-teach (t) = 1 |R 1,t | X i∈R 1,t BCE f θ 2 (x i ), ̃y i , (D.2) where R 1,t and R 2,t denote the sets of reliable samples selected at iteration t by networks 1 and 2, respectively. D.2. Stochastic Co-teaching Standard Co-teaching requires the user to specify a forget rate in advance. Stochastic Co-teaching (de Vos et al. 2023) replaces this fixed rate with a stochastic rejection rule. At iteration t, each network samples a threshold z t ∼ Beta(α,β), where α and β control the expected strictness of sample rejec- tion. For a mini-batchB t =(x i , ̃y i ) i∈I t , network j∈1, 2 outputs a probability ˆp ( j) i (t) = σ f θ j (x i ) . The probability assigned to the observed training label is p ( j) i (t) = ˆp ( j) i (t),if ̃y i = 1, 1− ˆp ( j) i (t), if ̃y i = 0. Article number, page 17 A&A proofs: manuscript no. main Samples with p ( j) i (t) ≤ z t are rejected, and the remaining samples are passed to the peer network for updating, as in stan- dard Co-teaching. This stochastic rule removes the need to fix a single deterministic forget rate, but in our experiments it re- mained difficult to tune and often led to unstable class-wise re- jection behavior. For this reason, we report it only as an auxiliary baseline and do not retain it among the main methods. Appendix E: Signal-to-noise ratio by class E.1. Signal-to-noise ratio distribution The SNR is computed as: SNR = |psfFlux| psfFluxErr , where psfFlux is the PSF-fitted flux measured on the differ- ence image and psfFluxErr is its associated uncertainty. Fig- ure E.1 shows that the SNR distributions differ substantially between the two training classes. In particular, detections with SNR ≤ 5 are nearly absent from the injected class because such faint events are less likely to be recovered by the detec- tion pipeline. The injected class also extends to systematically higher SNR values than the survey class. This class asymmetry is expected from the construction of the dataset, but it should be kept in mind when interpreting the relationship between SNR and uncertainty. Fig. E.1: Signal-to-noise ratio analysis of the rc2_subset dataset by class. Panel (a) shows the SNR distribution of the injected class, panel (b) the SNR distribution of the survey class, panel (c) the corresponding cumulative distribution functions, and panel (d) the distribution split below and above SNR = 5. E.2. Signal-Noise-Ratio correlation with Uncertainty Quantification per class Table E.1 reports the Spearman correlation between predicted uncertainty and SNR on the manually curated evaluation set, stratified by predictive outcome. The strongest negative correla- tion is observed for true positives, indicating that uncertainty for correctly identified transient detections is strongly linked to sig- nal quality. For true negatives, the correlation is weaker, suggest- ing that uncertainty on bogus detections depends less directly on photometric SNR and more on heterogeneous artifact morphol- ogy. Correlations for false positives and false negatives are small and unstable in sign, indicating that misclassifications are not explained by SNR alone. ρ SNRTPFPTNFN MC Dropout−0.3680.089 −0.1570.063 Deep Ensemble −0.434 −0.028 −0.192 −0.105 Repulsive Ensemble−0.4200.084 −0.1620.123 Our method −0.549 −0.038 −0.1660.108 Table E.1: Spearman rank correlation coefficients (ρ) between predicted uncertainty and signal-to-noise ratio (SNR) for each predictive class, True Positive (TP), False Positive (FP), True Negative (TN) and False Negative (FN). Appendix F: Uncertainty Quantification formalism F.1. Deep Ensemble During inference, the ensemble prediction is obtained by aver- aging the individual model probabilities: ̄p(x) = 1 M M X m=1 σ( f (x; ˆ θ m )),(F.1) where σ(·) denotes the sigmoid activation function, and ˆ θ m the trained weights of the m-th model. The final predicted class label is then given by: ˆy ensemble = I[ ̄p(x) > 0.5],(F.2) with I[·] representing the indicator function. Beyond improving prediction stability, the ensemble also fa- cilitates uncertainty quantification. The predictive uncertainty for a given input x is estimated as the variance of the individ- ual model probabilities: Var ens (x) = 1 M M X m=1 σ( f m (x))− ̄p(x) 2 .(F.3) This measure reflects the degree of agreement among ensemble members, with higher variance indicating greater predictive un- certainty. F.2. Repulsive ensemble Operationally, the interaction term of the repulsive ensemble is encoded through an update of the form φ( f t i ) =∇ f t i log p( f t i |D) − R ∇ f t i k( f t i , f t j ) M j=1 ,(F.4) where the first term promotes agreement with the data while the second term introduces repulsion via a kernel k(·,·) measuring similarity between ensemble functions. The operatorR(·) aggre- gates the repulsive contributions across the other ensemble mem- bers, pushing f i away from functions that are already represented in the ensemble. This repulsivity is motivated by the desire of the uncertainty quantification method to cover a larger area of the loss landscape and avoid ensembles that are accurate but overly concentrated around a single mode. At inference time, the procedure remains the same as for Deep Ensembles: predictions are obtained by aggregating the ensemble outputs, and uncertainty can be quan- tified from the dispersion across members. Article number, page 18 Bonnet-Guerrini et al.: Interpretable Human-Label-Free Real-Bogus with UQ F.3. MC Dropout Let ˆ θ be the trained weights. Given N dropout masks, we sam- ple the dropout masks z n ∼ q(z) (Bernoulli masks based on the dropout rate of the training), so that for each input x ∗ we per- form N forward passes with these different masks. For pass n, let f n (x ∗ ; ˆ θ, z n ) denote the pre-sigmoid logit of the network. The corresponding predictive probability is p n (x ∗ ) = p(y ∗ = 1| x ∗ , ˆ θ, z n ) = σ f n (x ∗ ; ˆ θ, z n ) ,(F.5) with σ the sigmoid function of the last layer. The MC dropout predictive probability is then approximated by the first moment: ̄p(x ∗ ) = 1 N N X n=1 p n (x ∗ ) = 1 N N X n=1 σ f n (x ∗ ; ˆ θ, z n ) .(F.6) The (epistemic) predictive variance can be approximated by the second centered moment: Var MCdropout (x ∗ ) = 1 N N X n=1 p n (x ∗ )− ̄p(x ∗ ) 2 .(F.7) Appendix G: UMAP representation The UMAP projection can be overlaid with additional quan- tities such as predictive uncertainty and signal-to-noise ratio. These overlays provide a qualitative view of how uncertainty and SNR are distributed across the learned representation. Fig- ure G.1 shows that high-uncertainty samples concentrate near the interface between the main regions of the embedding. Fig- ure G.2 shows that SNR only partially follows the same struc- ture, with a clearer trend in the transient-like region than in the bogus region. These visual patterns are consistent with the quan- titative results reported in Section 7.2. Fig. G.1: Static UMAP representation of the latent space with predictive uncertainty shown as a color overlay. Symbols rep- resent model predictions. The manually curated evaluation set is shown, and uncertainty is computed using the method introduced in Section 5.5. The displayed uncertainty is the per-sample stan- dard deviation of the predicted probability. Appendix H: Cumulative light-curve counts Fig. G.2: Static UMAP representation of the latent space with signal-to-noise ratio shown as a color overlay. Values are capped at SNR = 15. Symbols represent model predictions. The man- ually curated evaluation set is shown; accordingly, no displayed sources have SNR < 5. Article number, page 19 A&A proofs: manuscript no. main Threshold (%)All LCsSNother transients or variablesTransients (all)Bogus ≤ 517433 (100.00%)41 (100.00%)74 (100.00%)174 ( 75.00%) ≤ 1019033 (100.00%)41 (100.00%)74 (100.00%)190 ( 81.90%) ≤ 1520733 (100.00%)41 (100.00%)74 (100.00%)207 ( 89.22%) ≤ 2021533 (100.00%)41 (100.00%)74 (100.00%)215 ( 92.67%) ≤ 2522233 (100.00%)41 (100.00%)74 (100.00%)222 ( 95.69%) ≤ 3022333 (100.00%)41 (100.00%)74 (100.00%)223 ( 96.12%) ≤ 3522733 (100.00%)41 (100.00%)74 (100.00%)227 ( 97.84%) ≤ 4023133 (100.00%)40 ( 97.56%)73 ( 98.65%)230 ( 99.14%) ≤ 4523233 (100.00%)40 ( 97.56%)73 ( 98.65%)231 ( 99.57%) ≤ 5023233 (100.00%)40 ( 97.56%)73 ( 98.65%)231 ( 99.57%) ≤ 5523233 (100.00%)40 ( 97.56%)73 ( 98.65%)231 ( 99.57%) ≤ 6023333 (100.00%)40 ( 97.56%)73 ( 98.65%)232 (100.00%) ≤ 6523830 ( 90.91%)39 ( 95.12%)69 ( 93.24%)232 (100.00%) ≤ 7024328 ( 84.85%)35 ( 85.37%)63 ( 85.14%)232 (100.00%) ≤ 7524428 ( 84.85%)34 ( 82.93%)62 ( 83.78%)232 (100.00%) ≤ 8024826 ( 78.79%)33 ( 80.49%)59 ( 79.73%)232 (100.00%) ≤ 8525622 ( 66.67%)28 ( 68.29%)50 ( 67.57%)232 (100.00%) ≤ 9026519 ( 57.58%)23 ( 56.10%)42 ( 56.76%)232 (100.00%) ≤ 9527911 ( 33.33%)16 ( 39.02%)27 ( 36.49%)232 (100.00%) ≤ 1003063 ( 9.09%)8 ( 19.51%)11 ( 14.86%)232 (100.00%) Table H.1: Cumulative number of light-curves below a predicted-transient threshold (All LCs, Bogus) and above the bin lower edge for transient subclasses (SN-like, other transients or variables, Transients all). Percentages are cumulative within each class/subclass. Article number, page 20