Paper deep dive
TransitNet: A Compact Attention-Augmented Deep Learning Framework for Low-SNR Transit Blind Searches
Xingchen Yan, Jian Ge, Qingtian Liu, Kevin Willis, Quanquan Hu, Jiapeng Zhu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 6/21/2026, 3:19:35 AM
Summary
TransitNet is a compact, attention-augmented deep learning framework designed for low-SNR (Signal-to-Noise Ratio) transit blind searches, specifically targeting intermediate-to-long-period Earth-size planets. The architecture consists of a Filtering Module (FM) for local feature extraction and a Multi-Head Attention (MHA) module for capturing long-range dependencies. It outperforms traditional methods like BLS and TLS in accuracy and recovery rates for low-SNR signals. Beyond detection, TransitNet provides transit window and midpoint estimates using attention-based localization. The model is highly efficient, with a 1.5 MB footprint and significant speed-ups over CPU-based traditional algorithms.
Entities (8)
Relation Signals (5)
TransitNet β appliedto β Kepler
confidence 100% Β· Applied to real Kepler observations, the model successfully recovers all 34 selected confirmed Kepler planets
TransitNet β contains β Filtering Module
confidence 100% Β· The model comprises three modules... the filtering module (FM)...
TransitNet β contains β Multi-Head Attention Module
confidence 100% Β· The model comprises three modules... the MHA module...
TransitNet β outperforms β TLS
confidence 100% Β· TransitNet attains 95.2 percent accuracy... and outperforms both TLS and BLS
TransitNet β outperforms β BLS
confidence 100% Β· TransitNet attains 95.2 percent accuracy... and outperforms both TLS and BLS
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Motivated by the observational incompleteness of intermediate-to-long-period Earth-size planets, we present TransitNet, a compact attention-augmented deep-learning framework for low-SNR transit blind searches. To enable realistic method development and objective threshold calibration under blind-search conditions, we develop a unified dataset construction, benchmarking, and threshold-selection framework. On recovery benchmarks constructed from unseen Kepler targets, TransitNet attains 95.2 percent accuracy in the challenging SNR range of 6 to 8 and outperforms both TLS and BLS, achieving ROC-AUC and PR-AP values of 0.974 and 0.982, respectively. In an injected Earth-size and sub-Earth-size transit recovery experiment, TransitNet achieves a recovery rate of 93.0 percent, substantially exceeding those of TLS (63.1 percent) and BLS (60.0 percent). In addition to detection, TransitNet provides attention-based estimates of transit windows and midpoints. On an independent evaluation set, 97.4 percent of injected transits are fully covered by the estimated transit window. Applied to real Kepler observations, the model successfully recovers all 34 selected confirmed Kepler planets, with a mean absolute transit midpoint error of 1.24 hours. The model combines a compact footprint of about 1.5 MB with high inference efficiency, yielding speed-ups of about 12 to 25 times relative to CPU-TLS and about 4 to 5 times relative to CPU-BLS. These results demonstrate that TransitNet provides an accurate, scalable, and computationally efficient framework for low-SNR transit blind searches in the tested regime and motivate its extension to longer-period Earth-size planet searches.
Tags
Links
- Source: https://arxiv.org/abs/2606.18932v1
- Canonical: https://arxiv.org/abs/2606.18932v1
Trouble viewing inline? Open PDF directly β
Full Text
116,817 characters extracted from source content.
Expand or collapse full text
MNRAS 000, 1β20 (2026)Preprint 18 June 2026Compiled using MNRAS L A T E X style file v3.2 TransitNet: A Compact Attention-Augmented Deep Learning Framework for Low-SNR Transit Blind Searches Xingchen Yan 1,2β , Jian Ge 1 β , Qingtian Liu 1,2 , Kevin Willis 3 , Quanquan Hu 1,2 and Jiapeng Zhu 1 1 Shanghai Astronomical Observatory, Shanghai 200030, China 2 University of Chinese Academy of Sciences, Yanqi Lake Campus, East Road 1, Huairou, Beijing 101408, China 3 Science Talent Training Center, Gainesville, FL, 32606 USA Last updated 2026 June 16; in original form 2025 May 17 ABSTRACT Motivated by the observational incompleteness of intermediate-to-long-period Earth-size planets, we present TransitNet, a com- pact attention-augmented deep-learning framework for low-SNR transit blind searches. To enable realistic method development and objective threshold calibration under blind-search conditions, we develop a unified dataset construction, benchmarking, and threshold-selection framework. On recovery benchmarks constructed from unseen Kepler targets, TransitNet attains 95.2 per cent accuracy in the challenging SNR = 6β8 regime and outperforms both TLS and BLS, achieving ROC-AUC and PR-AP values of 0.974 and 0.982, respectively. In an injected Earth-size and sub-Earth-size transit recovery experiment, TransitNet achieves a recovery rate of 93.0 per cent, substantially exceeding those of TLS (63.1 per cent) and BLS (60.0 per cent). In addition to detection, TransitNet provides attention-based estimates of transit windows and midpoints. On an independent evaluation set, 97.4 per cent of injected transits are fully covered by the estimated transit window. Applied to real Kepler observations, the model successfully recovers all 34 selected confirmed Kepler planets, with a mean absolute transit midpoint error of 1.24 h. The model combines a compact footprint (1.5 MB) with high inference efficiency, yielding speed-ups of βΌ12β25Γ relative to CPU-TLS andβΌ4β5Γ relative to CPU-BLS. These results demonstrate that TransitNet provides an accurate, scalable, and computationally efficient framework for low-SNR transit blind searches in the tested regime and motivate its extension to longer-period Earth-size planet searches. Key words: planets and satellites: detection β (stars:) planetary systems β techniques: photometric β methods: statistical β surveys 1 INTRODUCTION The transit method has emerged as the most productive technique for exoplanet detection. Leveraging both ground-based surveys such as WASP (Pollacco et al. 2006) and KELT (Pepper et al. 2012), and space-based missions including CoRoT (Auvergne et al. 2009), Ke- pler/K2 (Koch et al. 2010; Howell et al. 2014), and TESS (Ricker et al. 2014), these methods have yielded over 6000 confirmed exo- planets to date 1 . Among these missions, the Kepler space telescope has been particularly transformative (Koch et al. 2010): its 4-year primary mission discovered over 2700 confirmed exoplanets and identified an additional ~1900 candidates awaiting confirmation 1 . Its successor, the Transiting Exoplanet Survey Satellite (TESS, Ricker et al. 2014), surveys an area 400 times larger than Keplerβs field of view, enabling all-sky monitoring of millions of nearby bright stars (Ricker et al. 2014). Upcoming missions such as PLATO (Catala & The PLATO Consortium 2009; Rauer et al. 2014, 2025) and ET (the β Contact e-mail: yanxingchen0@gmail.com β Contact e-mail: jge@shao.ac.cn 1 https://exoplanetarchive.ipac.caltech.edu/docs/counts_ detail.html Earth 2.0 space mission; Ge et al. 2022a,b,c, 2024a,b) will further expand our capability to detect and characterize small exoplanets. Despite these remarkable observational achievements, a discrep- ancy persists between the number of detected candidates and con- firmed planets, especially for Earth-size (βΌ 1ν β ) bodies. Early Ke- pler-based occurrence studies reported a rising planet occurrence rate toward smaller planet radii around Sun-like stars (Howard et al. 2012). Subsequent analyses accounting for false positives and detec- tion completeness confirmed the high occurrence of small planets (Fressin et al. 2013; Petigura et al. 2013). In particular, Fressin et al. (2013) found that 16.5 Β± 3.6 per cent of FGK stars host at least one Earth-size planet (0.8β1.25 ν β ) with orbital periods up to 85 days, while occurrence rates remained high across the 0.8β4 ν β range. When combined with completeness corrections and extended to longer periods in later studies, these results are often interpreted as implying that a large fraction (up toβΌ50 per cent) of Sun-like stars host at least one planet smaller than 4ν β . Using the Q1βQ17 Kepler Data Release 25 (DR25) catalog, Bryson et al. (2020) estimated the occurrence rate of rocky planets in the conservative habitable zone of Solar-like stars to be 0.37β0.60 planets per star, broadly consis- tent with earlier extrapolations of Earth-size planet occurrence from shorter-period Kepler populations (e.g., Petigura et al. 2013). Given these inferred occurrence rates, the observed exoplanet Β© 2026 The Authors arXiv:2606.18932v1 [astro-ph.EP] 17 Jun 2026 2 Xingchen Yan et al. 10 β1 10 0 10 1 10 2 10 3 10 4 Orbital Period [days] 10 0 10 1 10 2 Planetary Radius [Earth Radii] 10 1.5 Transit Solar System Me V E Ma J S U N 0 500 Planetary Radius β Orbital Period 0500 Figure 1. Confirmed exoplanets on the ν p βν plane (log scale), with marginal histograms of ν (top) and ν p (right). Solar System planets are shown as labelled magenta stars; dashed lines mark ν p = 1ν β , ν = 100 days, and ν = 10 1.5 days (β 31.6 days). The transit-detected population becomes sparse beyond ν β 31.6 days, particularly in the regime of Earth-size planets, leaving an unpopulated region around ν p βΌ 1ν β . We therefore adopt this scale as the operational boundary between short-period and intermediate- to-long-period (I2LP) planets (Section 3.1.1). The distributions also show a strong short-period detection bias and a radius gap near 1.5β2ν β (Fulton et al. 2017). population is dominated by larger (ν > 2ν β ) planets on short- period orbits, while true Earth analogs remain rare in current detec- tions, with the small-radius, intermediate-period regime (ν β©½ 1 ν β , ν > 10 1.5 β 31.6 days, Fig. 1) remaining largely devoid of confirmed planets. The Kepler mission was designed to detect Earth-size tran- sits with depths ofβΌ84 ppm, requiring a 6.5-hour CDPP of β²20 ppm for a β³ 4ν detection (Koch et al. 2010). The distribution of 6-hour CDPP values spans approximately βΌ20 to β³100 ppm, with a peak nearβΌ30 ppm for νΎ ν βΌ 12 targets and systematically higher values for fainter stars (Christiansen et al. 2012). Given the achieved photometric precision for a subset of tar- gets, Kepler provides sensitivity to such signals in favorable cases. However, Earth-size planets at intermediate to long periods are still sparsely represented in the observed sample, which may in part reflect incomplete detection at low signal-to-noise (SNR) ratios, including limitations in the sensitivity of transit-search algorithms. Exoplanet detection pipelines typically consist of an initial triage stage followed by more detailed vetting of candidate signals (Jenkins et al. 2012). In the triage stage, transit-like signals are identified us- ing a range of detection algorithms. For example, the Kepler pipeline employs a wavelet-based matched-filter approach to detect periodic transit signatures (Jenkins et al. 2010a,b). In addition, periodogram- based methods are widely used in the community, among which the Box Least Squares (BLS; KovΓ‘cs et al. 2002) algorithm models tran- sits as box-shaped signals, while the Transit Least Squares (TLS; Hippke & Heller 2019) method improves sensitivity by incorpo- rating more realistic transit shapes with limb darkening. Additional approaches include the Transit Comb Filter (TCF; Gondhalekar et al. 2023) combined with ARIMA noise modeling (Melton et al. 2024), and the Quasi-Periodic Automated Transit Search (QATS; Carter & Agol 2013) algorithm for detecting transits with timing variations. Machine learning is increasingly used for transit detection and val- idation. Examples include dimensionality reduction and similarity- based classification (Thompson et al. 2015), Bayesian model selec- tion (Mullally et al. 2016), self-organizing maps for transit shape classification (Armstrong et al. 2016), and time-series outlier re- jection techniques such as TSARDI (Mislis et al. 2018). Gaussian process classifiers have been applied to probabilistic planet validation (Armstrong et al. 2020), while random forests have been used in au- tomated vetting pipelines such as Autovetter (McCauliff et al. 2015). Ensemble and boosting methods, including XGBoost and GBDT, have further improved performance in identifying rare transit signals in large datasets (Mislis et al. 2016; Malik et al. 2021; Pratyush & Gangrade 2021; Panahi et al. 2022; Melton et al. 2024). A landmark study by Shallue & Vanderburg (2018) introduced AstroNet, a CNN model that classifies Kepler Threshold Crossing Events (TCEs) and demonstrated high accuracy in distinguishing planetary signals from false positives. Subsequent studies have con- firmed the strong performance of CNNs relative to traditional ma- chine learning methods (Pearson et al. 2017), and extended AstroNet- like architectures to K2 and TESS data (Dattilo et al. 2019; Yu et al. 2019; Tey et al. 2023). Improvements such as ExoNet incorporate additional domain knowledge, including stellar parameters and cen- troid information, yielding significant gains in recall for low-SNR signals (Ansdell et al. 2018). More recent architectures, including ExoMiner and related models, have achieved high precision and re- call in automated vetting and enabled the validation of new exoplanets (Valizadegan et al. 2022). These models have since been extended to TESS data and full-frame image pipelines, further improving perfor- mance and scalability (Valizadegan et al. 2023, 2025; Martinho et al. 2026). A number of studies have leveraged simulated or synthetic data to construct more diverse training sets. Zucker & Giryes (2018) tested CNNs for detecting transit signals based on simulated data. Osborn et al. (2020) created a raw TESS artificial dataset using the Lilith simulator (Li et al. 2019) to train ExoNet and validate it on real TESS data. Iglesias Γlvarez et al. (2023) also tested their own CNNs for detecting transit signals based on simulated data. These stud- ies further highlight the importance of synthetic data in addressing data scarcity and improving model generalization across realistic observational domains (Zucker & Giryes 2018; Yeh & Jiang 2020; Iglesias Γlvarez et al. 2023). Most recently, the GPU phase fold- ing and CNN (GPFC) system of Wang et al. (2024a,b) combined semi-synthetic Kepler training with GPU-accelerated blind search, discovering ultra-short-period (USP) planet candidates at the short- est periods and smallest radii reported to date for such objects, with transit SNR as low asβΌ 6-7 (Wang et al. 2024a,b). In addition, several studies have represented transit signals from different period windows as stacked two-dimensional inputs for CNN-based detection (Chintarungruangchai & Jiang 2019; CuΓ©llar et al. 2022). More recent work has moved beyond purely convolu- tional classifiers applied to phase-folded light curves. For example, CNNs have been combined with recurrent or attention modules for wide-area TESS searches (PΓ€tzold et al. 2025; Thomas et al. 2025), while Transformer-based models have been applied both to folded light curves (Salinas et al. 2023) and directly to TESS full-frame image pixels without prior folding (Salinas et al. 2025). Other re- lated developments include Vision Transformers applied to image- like encodings of time series (Choudhary et al. 2025), residual net- works with channel attention for interpretable vetting of Kepler and TESS data (G & Kumari 2023; Xie et al. 2025), and generative flow- matching methods coupled to gradient-boosted vetters (Fiscale et al. 2025). Another active direction focuses on low-SNR and sparse-event regimes, including single- or quasi-single-transit detection, where the strict periodicity assumptions of traditional methods are weak- ened. Progress in this area includes U-Net/GAN-based segmentation (Dvash et al. 2022), deep classifiers applied to lightly pre-filtered MNRAS 000, 1β20 (2026) TransitNet 3 TESS light curves (Vivien et al. 2025), and object-detection-based approaches (Cui et al. 2021). Alternative generative models for exo- planet detection also remain under investigation (AydoΔan 2022). Taken together, the current state of exoplanet transit detection can be characterized from three closely related perspectives. (i) Observed population gap and detection sensitivity. Current exoplanet catalogs exhibit a noticeable deficit in the I2LP and small- radius regime. In the context of Kepler light curves, this feature is generally understood to arise primarily from the sensitivity limits of detection algorithms. (i) Limitations of existing detection frameworks. Representa- tive classical transit blind-search methods such as BLS and TLS rely on template-based search statistics and require exhaustive period searches, leading to reduced efficiency in large-scale surveys and limited flexibility in capturing diverse transit morphologies. Mean- while, existing deep learning-based approaches are rarely explicitly designed for low-SNR blind transit search scenarios, with limited consideration of unified search protocols and threshold selection cri- teria. (i) Limited physical interpretability beyond detection. Most classification-based deep learning approaches generally do not pro- vide additional physically informative outputs, such as initial esti- mates of transit parameters that would facilitate subsequent vetting and candidate characterization. In this work, we introduce TransitNet, a compact attention- augmented deep learning framework designed for low-SNR transit detection in phase-folded Kepler light curves. Unlike conventional template-matching approaches, TransitNet learns data-driven rep- resentations of low-SNR transit morphology while maintaining a lightweight architecture suitable for large-scale survey applications. In addition to binary detection, the model uses attention-informed localization to provide an initial estimate of the transit midpoint ( Λν 0 ), thereby offering physically useful information for downstream vet- ting and characterization. We demonstrate that TransitNet achieves higher detection sensitivity than BLS and TLS in the ν β [30, 60] days, low-SNR regime, while providing a scalable path towards fu- ture transit blind-searches for longer-period terrestrial planets. In a blind-search setting, it is coupled with GPU-based phase folding to scan trial periods and epochs efficiently (Wang et al. 2024a,b; Hu et al. 2026). Although the present work focuses on the intermediate- period range of 30β60 days, it represents an important step towards attention-based searches for longer-period, lower-transit-multiplicity Earth analogs. The remainder of this paper is organized as follows. Section 2 describes the TransitNet architecture, including the filtering module (FM), the multi-head attention (MHA) module, and the fully con- volutional (FCN) mapping layer. Section 3 introduces the dataset construction and benchmarking pipeline for constructing reliable non-transit controls and presents controlled ablation studies that quantify the contributions of MHA and FCN to low-SNR transit sensitivity. Section 4 benchmarks TransitNet against BLS and TLS under realistic transit blind-search conditions on unseen Kepler tar- gets, evaluating single-target recovery sensitivity on the Low-SNR Transit Recovery Set, cross-target generalization on the Cross-KIC Recovery Set, Earth-size transit recovery rates, inference efficiency, and transit-window and midpoint (ν 0 ) estimation. Finally, Section 5 synthesizes our findings, discusses their implications, and outlines future research directions. 2 ARCHITECTURE OF TRANSITNET This section outlines the architecture of the proposed model, Transit- Net. The model comprises three modules with complementary roles: the FM suppresses noise and extracts local temporal features from the phase-folded light curve; the MHA module captures long-range dependencies and enhances sensitivity to low-SNR transit morpholo- gies, in which the attention weight matrix is further leveraged to inform subsequent transit midpoint estimation, while the FCN map- ping module fuses and refines the representation while preserving spatial structure to produce a scalar detection score. The input to the model is the signal sequence obtained after GPU- based phase folding and binning (Wang et al. 2024a,b; Hu et al. 2026): ν = [ ν 1 , ν 2 ,..., ν νΏ 0 ], where νΏ 0 = 4096 is a fixed length (as discussed in Section 3.1.2) chosen to accommodate the intended search period and transit duration, and ν ν denotes the flux in the ν-th phase bin. 2.1 Filtering Module The FM acts as a learnable, data-driven front-end that suppresses noise and extracts initial local temporal features from the phase- folded light curve. Low-SNR transit signals pose a fundamental challenge for detec- tion because they typically appear as low-amplitude, short-duration dips that are difficult to distinguish from photometric noise and stellar variability in the raw light curve. In particular, for low-SNR transits, these time-domain features are readily obscured by noise, intrin- sic variability, and instrumental systematics, thereby limiting the effectiveness of template-matching and threshold-based detection methods. Certain systematic noise patterns that are challenging to characterize in the time domain may, however, exhibit more distinct signatures in the frequency domain. For example, correlated (col- ored) noise and instrumental systematics often produce concentrated spectral features that can be more effectively suppressed, whereas white noise has a flat power spectral density across frequencies. The convolution operation corresponds to a local weighted sum whose kernel parameters are learned under supervision via back- propagation. From a signal-processing perspective, these learned kernels act as data-driven filters, analogous in spirit to classical hand-designed filters. For example, a Gaussian-like kernel tends to smooth and suppress high-frequency noise (low-pass behavior), while a Laplacian-like kernel emphasizes rapid changes and high- frequency content (high-pass behavior). The theoretical foundation linking convolution to filtering is provided by the convolution theo- rem: for one-dimensional time series ν(ν‘) and ν(ν‘), withF denoting the Fourier transform and νΉ(ν),νΊ(ν) their frequency-domain coun- terparts, Fν(ν‘)β ν(ν‘) = νΉ(ν)Β· νΊ(ν),(1) which shows that convolution in the time domain is equivalent to pointwise multiplication in the frequency domain. This equivalence implies that convolutional layers can be understood as learnable frequency-selective filters: with random initialization and gradient- based training, the convolutional kernels in our FM learn to dis- tinguish between noise-dominated and signal-dominated frequency components, adaptively suppressing noise patterns (such as corre- lated instrumental systematics) while preserving transit-related fea- tures, thereby extracting salient structure from low-SNR transit sig- nals and motivating the use of CNNs for the FM. The design of the FM is inspired by the convolutional stack used for global-view feature extraction in AstroNet (Shallue & Vander- MNRAS 000, 1β20 (2026) 4 Xingchen Yan et al. burg 2018), where multiple one-dimensional convolutional layers perform hierarchical feature extraction on the time series. As shown in Fig. 2, the FM comprises five one-dimensional convolutional lay- ers with channel dimensions increasing in powers of two (16, 32, 64, 128, 256). Each layer consists of one-dimensional convolution, max pooling, dropout, and ReLU activation. With our default hy- perparameters, the input ν is transformed by the FM into a feature tensor νΉ β R νΏ 1 Γ256 (specifically νΏ 1 = 128, Fig. 4), where νΏ 1 is determined by the kernel size, stride, and pooling configuration of the convolutional layers. Convolutional operators, however, have a limited receptive field; capturing long-range dependencies requires stacking many layers, which can introduce unnecessary parameters and optimization dif- ficulties in low-SNR regimes. We therefore use a relatively shallow convolutional stack for local feature extraction and preliminary fil- tering and delegate global temporal structure to a dedicated module with a larger effective receptive field β the MHA module described below. 2.2 Multi-Head Attention Module The MHA module provides a global receptive field over the folded transit profile and enhances sensitivity to low-SNR, morphologically consistent transit features. MHA was introduced in the Transformer architecture by Vaswani et al. (2017) to model long-range dependencies in sequence data. Unlike recurrent and convolutional architectures, attention enables direct interactions between all positions in a sequence, providing a global receptive field and supporting parallel computation while assigning context-dependent weights to different parts of the input. Each head in MHA is built from self-attention. Let νΉ β R νΏ 1 Γν ν denote the feature matrix from the FM, where νΏ 1 is the downsampled sequence length and ν ν is the feature dimension. Self-attention is computed within this single representation: the same νΉ serves as both query source and key/value source. We first project νΉ onto query (ν), key (νΎ), and value (ν) matrices via learned linear maps: ν = νΉν ν , νΎ = νΉν νΎ , ν = νΉν ν (2) where ν ν ,ν νΎ ,ν ν β R ν ν Γν model are learnable weight matrices, yielding ν,νΎ,ν β R νΏ 1 Γν model . Here ν model is the attention embed- ding dimension. In our implementation, ν ν = 256 and ν model = 16. ν ν maps νΉ to a query space encoding what to look for (e.g., ingress/egress, duration, depth); ν νΎ to a key space encoding the patterns available for matching; and ν ν to a value space encoding the flux information to be aggregated. For each attention head, we split ν, νΎ, and ν along the feature dimension into β subspaces of dimension ν νΎ = ν ν = ν model /β. Denoting the ν-th head by subscript ν, we have ν ν ,νΎ ν ,ν ν β R νΏ 1 Γν νΎ with ν νΎ = ν ν . The scaled dot-product attention for head ν is then Attention(ν ν ,νΎ ν ,ν ν ) = softmax ν ν νΎ β€ ν β ν νΎ ν ν (3) modeling pairwise interaction between phase bins over the folded transit profile. The similarityν ν νΎ β€ ν encourages morphologically con- sistent transit features across phases: bins that are symmetrically dis- tributed around the transit center may receive higher attention weights when they exhibit similar flux dips, while other bins receive lower weights, enabling low-SNR transit signals to be emphasized and un- related fluctuations to be suppressed. The scale factor 1/ β ν νΎ keeps the dot products from growing with ν νΎ and stabilizes the softmax. The resulting attention matrix ν΄ ν = softmax(ν ν νΎ β€ ν / β ν νΎ ) β R νΏ 1 ΓνΏ 1 encodes pairwise relationships between phase bins: each entry ν΄ ν,ν represents the attention weight assigned to phase bin ν (the source) when constructing the output representation for phase bin ν (the target), where both ν and ν index positions along the phase-folded sequence (ν, ν β 1, 2,..., νΏ 1 ). The output ν΄ ν ν ν is the correspond- ing weighted combination of value vectors, where each row of ν΄ ν ν ν aggregates information from all phase bins according to the attention weights in ν΄ ν . A single attention head may lack capacity to encode the full di- versity of transit morphologies (e.g., depth, duration, and shape vari- ations due to orbital geometry and limb darkening) in the presence of strong noise. We therefore adopt multi-head attention, computing self-attention in parallel over β heads and concatenating their outputs before a final linear projection: MultiHead(νΉ) = Concat(head 1 ,..., head β )ν ν ,(4) head ν = Attention(ν ν ,νΎ ν ,ν ν ),(5) whereν ν β R βν ν Γν model . To keep the concatenated dimension equal to the attention embedding dimension, we set βν ν = ν model and therefore ν ν = ν model /β; we take ν νΎ = ν ν . In our implementation (Fig. 2), νΉ β R νΏ 1 Γ256 is first projected to ν,νΎ,ν β R νΏ 1 Γ16 , then processed by MHA with β = 8 heads and per-head dimensions ν νΎ = ν ν = ν model /β = 2. The MHA output therefore has shape νΏ 1 Γν model = νΏ 1 Γ16, and is subsequently mapped back to ν ν = 256 channels by a convolutional layer before the downstream head. The overall computation flow of MHA is illustrated in Fig. 3. In summary, under low-SNR conditions, MHA strengthens mor- phologically consistent transit-related flux and down-weights random fluctuations through content-dependent, global pairwise weighting over phase bins; multiple heads further allow the model to capture diverse aspects (ingress/egress, flat bottom, depth, duration), improv- ing robustness and discriminative power for subsequent classification and regression. 2.3 Fully Convolutional Network Module The FCN module maps the MHA-refined representation to a scalar detection score whilst preserving the spatial and channel structure of the transit profile. In many CNN-based models, the final mapping to scalar outputs is implemented with fully connected (dense) layers. In our experiments, this design did not make full use of the multi-channel features pro- duced by MHA: flattening the feature matrix and passing it through dense layers discards the spatial structure of the folded transit profile and the coupling between feature channels. To preserve this structure, we use a FCN module. The FCN (Long et al. 2015) employs pointwise (1Γ 1) convolu- tions, analogous to the channel-mixing stage of depthwise-separable architectures (Howard et al. 2017; Chollet 2017; Tan & Le 2019). These layers fuse information across channels at each phase bin while preserving the sequence dimension. This design refines and fuses MHA features while preserving transit morphology, yielding more stable and physically interpretable outputs. As in Fig. 2, the FCN stacks several lightweight convolutional blocks with channel widths(128, 64, 32, 16, 8, 1), followed by a pointwise projection and sigmoid activation, producing a single scalar score per input sequence that indicates the presence or absence of a transit signal. 2.4 Overall Architecture and Relation to Classical Methods The complete architecture, consisting of the FM, the MHA module, and the FCN module, is referred to as TransitNet. MNRAS 000, 1β20 (2026) TransitNet 5 MHA FMFCN Q K V 16 256128326416 Score Conv Block (BN) Conv Block MHA Conv Sigmoid Figure 2. Architecture of TransitNet after multiple optimisations. Our network processes a batch of 1D sequences (Batch, νΏ 0 ) through three stages: Filtering, MHA, and FCN. In the FM, five convolution blocks (CB) expand the channel width (16, 32, 64, 128, 256) and produce features of shape (Batch, 256, νΏ 1 ). Two variants of CBs are used: one with Batch Normalization (BN) and one without, both of which consist of Conv1Dβ MaxPoolβ Dropoutβ ReLU; The MHA module linearly projects features to ν, νΎ, and ν and applies 8-head attention, preserving the shape (Batch, 256, νΏ 1 ). The FCN applies a lightweight Conv repeatedΓ6; the channel widths follow (128, 64, 32, 16, 8, 1). A final pointwise projection with a sigmoid yields a scalar score per sequence, giving the output shape (Batch, 1). Numbers on the blocks denote channel dimensions. Q K V F A O Self-Attention F Scaled Dot- Product Attention heads F Multi-Head Attention β Figure 3. Schematic of the MHA computation. The input feature matrix νΉ, encoding flux patterns from the folded transit profile, is linearly projected into queries ν (search templates for transit-like features), keys νΎ (reference catalogue of flux patterns), and values ν (actual flux information) using learnt weights ν ν ,ν νΎ , and ν ν . Attention weights ν΄ are computed via scaled dot-product attention, softmax ννΎ β€ / β ν νΎ , assigning adaptive weights to different phase bins based on their contribution to transit detection, and applied to the values to produce per-head outputs ν = ν΄ν. Outputs from all heads, each potentially specialising in different aspects of transit morphology, are concatenated and passed through a final linear projection to yield the enhanced feature matrix νΉ β² . In terms of both detection principle and computational mechanism, TransitNet differs fundamentally from classical transit search algo- rithms such as BLS (KovΓ‘cs et al. 2002) and TLS (Hippke & Heller 2019). Classical methods construct transit templates from physical models and perform a grid search over period, phase, and duration to maximize a global detection statistic; their performance therefore depends strongly on the adequacy of the assumed template family and on exhaustive exploration of a potentially large parameter space. At low-SNR, transit shapes can deviate from idealized box or analytic models, and random or systematic noise can introduce non-Gaussian, temporally correlated structure, so that fixed-template matched fil- tering may lose sensitivity while computational cost grows quickly with the search range. TransitNet, by contrast, follows a data-driven, end-to-end representation-learning paradigm. The convolutional layers learn hi- erarchical local filters from data, implicitly capturing ingress/egress slopes, depth variations, and small-scale noise patterns without hand- designed templates. The MHA then reweights features across all phase bins in a content-dependent manner (equation 3), allowing the model to focus on phases that jointly support transit detection and capture long-range structure without assuming a specific analytic light-curve shape. Multiple heads provide complementary "views" of the same folded profile, emphasizing different aspects of transit morphology or noise; their outputs are combined into a single rich representation for classification. Once trained, inference in TransitNet consists of a fixed sequence of convolutions and matrix multiplica- tions. For a given folded sequence length, the computational cost per trial period is fixed and can be efficiently parallelized on mod- ern GPUs. TransitNet thus offers both greater flexibility with respect to template mismatch and better scalability for large-scale surveys targeting low-SNR, noise-dominated transit signals. Fig. 4 shows the full pipeline from input to score for one transit and one non-transit example. For the transit example, the convolutional feature maps show a clear progression with depth: as successive filtering layers are applied, stochastic fluctuations are gradually sup- pressed and the transit signature becomes increasingly sharp and spatially coherent in phase. This behaviour matches the intended de- sign of the FM, which is to denoise the folded light curve while pre- serving and amplifying physically meaningful transit morphology. By contrast, for the non-transit example, the convolutional feature maps remain diffuse and lack any stable, localised temporal structure as depth increases, reflecting the absence of an underlying periodic dimming event. In this regime, the network does not produce a con- centrated transit-like response in this example; instead, it learns to treat the input as noise and avoids concentrating power at specific phases. The MHA representations further accentuate this contrast. For genuine transits, the learned query-key-value interactions and the resulting attention maps assign disproportionately high weight to a small set of phase bins that are jointly consistent with a transit pattern, while down-weighting regions dominated by residual noise. For non- transit inputs, the attention is more diffuse and fragmented, without MNRAS 000, 1β20 (2026) 6 Xingchen Yan et al. Transit | SNR=8.8 | Score=1.0 (16, 2048) FM (32, 1024) (64, 512) (128, 256) (256, 128) Q (16, 128) MHA K (16, 128) Attention Matrix (128, 128) V (16, 128) AV (16, 128) (256, 64) FCN (128, 32) (64, 16) (32, 8) (16, 4) (8, 2) 1.000 (1, 1) 30354045505560 Period [day] Detection Score Spectrum (3060 d) True Period = 35.38 d Non-Transit | Score=0.0 (16, 2048) FM (32, 1024) (64, 512) (128, 256) (256, 128) Q (16, 128) MHA K (16, 128) Attention Matrix (128, 128) V (16, 128) AV (16, 128) (256, 64) FCN (128, 32) (64, 16) (32, 8) (16, 4) (8, 2) 0.000 (1, 1) 30354045505560 Period [day] Detection Score Spectrum (3060 d) Figure 4. End-to-end feature-map visualisation of TransitNet for a transit (left) and a non-transit (right) input. From top to bottom: input light curves; FM (rows 2-6); MHA β ν, νΎ, attention matrix ν΄ = softmax(ννΎ β€ / β ν νΎ ), ν, and ν΄ν (rows 7β11); FCN (rows 12β18); detection score spectrum (bottom). MNRAS 000, 1β20 (2026) TransitNet 7 a coherent phase pattern. Taken together, these feature βfingerprintsβ demonstrate that TransitNet learns to couple local convolutional fil- tering with global, content-dependent reweighting, enabling the final FCN module to separate transit from non-transit sequences based on high-level, physically interpretable representations rather than on low-level noise fluctuations. The FCN rows in the figure show this refinement, and the detection score spectrum at the bottom yields a sharp peak at the true period for the transit case and a flat response for the non-transit. 3 EXPERIMENTAL SENSITIVITY ENHANCEMENT WITH THE DATA-TAILORED ARCHITECTURE This section presents a comprehensive evaluation of the proposed model using semi-synthetic Kepler light curves. We first introduce a standardized dataset construction pipeline for deep-learning-based transit detection to construct a benchmark dataset with controlled transit characteristics and noise properties (Section 3.1). Subse- quently, ablation studies are conducted to quantify the contributions of the MHA and FCN modules to low-SNR transit feature extraction, training stability, and classification performance (Section 3.2). 3.1 Benchmark Dataset Generation for Transit Detection We employ synthetic light curve generation to produce a large dataset for training a robust deep learning model as it is challenging to con- struct a dataset solely from confirmed transits: (1) low-SNR transits are intrinsically rare, and (2) the current sample size of confirmed exoplanets remains limited and does not adequately cover the full transit parameter space, thereby restricting the diversity of available training data for deep learning models. Through extensive experimentation, we have established a stan- dardized and principled pipeline for constructing datasets for deep learning-based exoplanet detection models: (i) Define observational settings and transit parameter space: Determine the sampling cadence of the telescope data, compute an appropriate number of phase bins for folding and binning, and specify the values and ranges of key transit parameters to be explored in the exoplanet search, including the orbital period ν, transit depth νΏ, and transit duration ν 14 . (i) Generate synthetic light curves: Simulated transit signals are constructed using parameter combinations sampled from the pre- defined ranges specified in the previous step. We consider three alter- native approaches: (1) fully synthetic: simulated transit signals are inserted into Gaussian Noise Light Curves (GNLC) or other Non- Transiting Light Curves (NTLC) to generate Artificial Transiting Light Curves (ATLC); (2) semi-synthetic: simulated transit signals are inserted into real Transit-Masked Light Curves (TMLC) to gen- erate ATLC; and (3) real light curves (confirmed exoplanet signals). (i) Obtain transit and non-transit signal samples: Transit sig- nal samples are derived from ATLCs obtained in the previous step after folding and binning. Non-transit signal samples are generated by folding and binning NTLC, ensuring that no transit features are present. (iv) Define dataset specifications: Ensure that the dataset con- tains a sufficient number of samples, that the injected transit signals cover the target parameter space as comprehensively as possible, and that the positive and negative samples are maintained in a balanced ratio (typically 1:1) to avoid training bias. 3.1.1 Distributions of parameters To simulate transit signals that resemble the observed exoplanet popu- lation as closely as possible, representative distributions of the orbital period (ν), transit duration (ν 14 ), and transit depth (νΏ) are required. We first define the target parameter space for sample construction by focusing on orbital periods ν β [30, 60] days and SNRβ [6, 15], corresponding to the onset of the I2LP and low-SNR regime consid- ered in this work, with 30 days adopted as the lower-period boundary for computational convenience (Fig. 1). Using the KOI tables avail- able from the NASA Exoplanet Archive (Thompson et al. 2018; Christiansen et al. 2025), we select confirmed Kepler planets (CPs) within this period range as the reference population. from which the distributions of ν β [30, 60] days, ν 14 β [1.99, 7.99] hours, and νΏ β [79, 1262] ppm are derived (Fig. 5). These empirical distributions define the admissible parameter ranges within which transit parameters are uniformly sampled to provide broad coverage of the parameter space for subsequent transit- signal generation. 3.1.2 Artificial Light Curve We incorporate TMLCs extracted from real Kepler observations to generate ATLCs rather than adopt an idealized noise model, thereby preserving the authentic noise characteristics of space-based pho- tometry (semi-synthetic approach). Sources of TMLCs: We randomly selectβΌ700 Kepler Input Cat- alog (KIC; Brown et al. 2011) targets corresponding to CP systems identified from the Kepler Object of Interest (KOI; Thompson et al. 2018) catalog, within specified parameter ranges ofν β [30, 60] days and SNRβ [6, 15], while excluding known false positives, to serve as the source of TMLCs. The corresponding KIC identifiers are then used to retrieve the Pre-search Data Conditioning Simple Aperture Photometry (PDCSAP; PDCSAP_FLUX) light curves at a long ca- dence of 29.4 min from the Mikulski Archive for Space Telescopes (MAST) 2 . We then concatenate all available quarterly segments, ap- ply detrending and normalization procedures, and mask all confirmed transit events. This avoids using light curves that are prone to imper- fect detrending, thereby reducing residual trends that could distort the injected transit signals and contaminate the resulting training samples. ATLC generation: We adopt a simplified trapezoidal transit model with distinct ingress and egress phases (Mandel & Agol 2002), ne- glecting limb-darkening effects. This approximation is adopted as a computationally efficient simplification for our low-SNR regime, where photometric noise dominates and transit durations are much shorter than orbital periods, while remaining computationally effi- cient for large-scale signal generation. For each TMLC, we inject a single transit signal into the light curve to produce the corresponding ATLC, with parameters (ν,ν 14 , and νΏ) uniformly sampled from pre- defined ranges. The transit epoch ν 0 , defined as the reference transit epoch, is drawn fromU(0, ν) and determines the phase offset of the injected transit signal. The randomly drawn parameters do not guarantee that all generated ATLCs fall within the target SNR range. We thus calculated the SNR of each candidate ATLC following the formula in KovΓ‘cs et al. (2002): SNR = νΏ ν βοΈ ν ν 14 ν ,(6) 2 https://archive.stsci.edu/kepler/ MNRAS 000, 1β20 (2026) 8 Xingchen Yan et al. 30354045505560 Orbital Period (days) 0 1 2 3 4 Count (a) Distribution of Orbital Period min: 30.254 max: 59.284 mean: 39.499 median: 38.208 234567 Transit Duration (hours) 0 1 2 3 4 (b) Distribution of Transit Duration min: 1.991 max: 7.987 mean: 4.651 median: 4.722 10 2 10 3 Transit Depth (ppm) 0 1 2 3 4 5 (c) Distribution of Transit Depth min: 79.000 max: 1262.000 mean: 493.585 median: 442.900 Figure 5. Distributions of transit parameters for Kepler CPs with orbital periods ν β [30, 60] days and SNR β [6, 15]. Panels show (left to right): orbital period ν, transit duration ν 14 , and transit depth νΏ. Inset boxes display summary statistics (min, max, mean, median) for each distribution. These empirical distributions serve as the basis for generating synthetic transit signals in our training dataset. where νΏ denotes the transit depth, ν represents the photometric measurement uncertainty per data point, ν is the total number of observations in the light curve, ν 14 is the transit duration, and ν is the orbital period. The quantity νν 14 /ν corresponds to the effective number of in-transit data points. This strategy ensures broad coverage of the parameter space and fills gaps in the distribution of confirmed exoplanet parameters that cannot be achieved using real observational data alone. The injected transit signals are combined with TMLCs to form ATLCs, preserving the intrinsic photometric trends and noise properties, and yielding data that closely resemble real observations while enabling con- trolled training sample generation. Since TMLCs may contain un- detected weak transit signals from additional planetary companions, this hybrid approach also better reflects realistic survey conditions and exposes the model to realistic low-level contamination, while potentially introducing label noise. Preparation of transit and non-transit signal samples: Each ATLC is subsequently phase-folded at the injected orbital period and resampled into ν = 4096 uniform phase bins. The choice of ν = 4096 is motivated by the requirement that even the shortest transit events be adequately resolved. Consider the limiting case of the longest orbital period (ν max = 60 days) and the shortest transit duration (ν 14,ννν β 2 hours). The temporal width of each phase bin is Ξν‘ = ν max ν = 60Γ 24Γ 60 4096 β 21.1 min.(7) Under this configuration, a 2-hour transit spans approximately 6 phase bins, providing sufficient sampling to characterize the transit morphology. This bin count corresponds to a power of two (2 12 ), which is favorable for binary-based computation and memory align- ment, and may offer implementation advantages in neural network architectures. Note that, after phase folding and binning, transit sig- nals may appear at any position within the phase window, including cases where they extend across window boundaries, consistent with realistic exoplanet search conditions. Considering that non-transit signal samples should ideally contain no transit features, the non-transit signal samples are constructed from two complementary sources to ensure both label purity and realistic noise characteristics: (i) For the first component, GNLCs are generated and phase- folded over specified period ranges to ensure an idealized non-transit case. The noise amplitude (standard deviation) of the GNLCs is sampled from the empirical distribution of the selected TMLCs. (i) For the second component, to preserve realistic instrumen- tal and astrophysical noise characteristics, we implement a Periodic Chunk Permutation (PCP) strategy applied to each TMLC. The pro- cedure is inspired by the quarter-shuffling framework of (Telesco et al. 2026). Each light curve is partitioned and reorganized using randomly selected periods uniformly drawn from [30, 60] days, ef- fectively disrupting residual periodic structures that may arise from masked or undetected transit signals (see Appendix A for details). The final dataset for ablation studies comprises βΌ71000 unique transit signal samples derived from ATLCs andβΌ85000 non-transit signal samples from both GNLCs and PCP TMLCs, with all samples binned to a length of ν = 4096. Given the large scale of this dataset and the need to phase fold each light curve at multiple trial periods, this process utilizes GPU-accelerated phase folding (Wang et al. 2024a,b; Hu et al. 2026) for all phase folding operations, which significantly reduces the computational time required for processing large numbers of light curves. 3.2 Ablation Studies To better understand the role of each component in TransitNet, we perform ablation studies, where selected modules are individually removed or replaced, and the resulting performance changes are analyzed. Specifically, we demonstrate that: (1) MHA can achieve higher sensitivity in detecting transit signals from time-series data. As discussed in Section 2.1, CNNs primarily model local temporal patterns, while capturing long-range dependencies requires increasing the receptive field through deeper architectures. In contrast, MHA enables direct modeling of global dependencies through pairwise interactions between all time steps, providing more efficient information aggregation for long sequence modeling without stacking deep layers. This module also allows MHA to automatically identify and prioritize salient transit features, thereby enhancing detection sensitivity. (2) FCNs are better than Fully-Connected Neural Networks (FCNNs) for extracting features from the output of the MHA, owing to their compatibility with the data structure. This improves the modelβs ability to distinguish between transit and non-transit signals. To validate these two claims, we design three additional mod- els alongside TransitNet and present their structures in a modular, compositional format. Table 1 provides annotations and descriptions for each of these components. Each model is constructed from a MNRAS 000, 1β20 (2026) TransitNet 9 Table 1. Description of the modules used in model construction. The design of these modules aims to investigate the capacity of different architectural building blocks to learn data features. The FM is adapted from the global- view branch of AstroNet (Shallue & Vanderburg 2018) and simplified by removing half of the convolutional layers. Module Description FMFiltering module. MHAMulti-Head Attention. FCNFully Convolutional Network. FCNNFully Connected Neural Network. combination of four distinct building blocks: FM, MHA, FCN, and FCNN. 3.2.1 Evaluation Metrics The comparative experiments conducted in this study are divided into two parts: classical metrics and the evaluation of stability and generalization capability. The metrics adopted are as follows: (1) Confusion matrix: there are four main concepts contained in the confusion matrix. True Positive (TP) refers to the number of samples that are correctly predicted as transit when they are actually transit events. False Positive (FP) denotes the number of samples that are incorrectly predicted as transit when they actually belong to the non-transit class. True Negative (TN) represents the number of samples that are correctly predicted as non-transit when they are actually non-transit events. False Negative (FN) indicates the number of samples that are incorrectly predicted as non-transit when they actually belong to the transit class. (2) Accuracy Accuracy = νν+νν νν+νν+ νΉν+ νΉν (8) (3) Precision, Recall, Specificity, and F1-Score P = νν νν+ νΉν , R = νν νν+ νΉν ,(9) νΉ 1 = 2PR P+R .(10) (4) Precision-Recall (PR) and Receiver Operating Character- istic (ROC) Curve: The PR curve plots precision against recall as the classification threshold varies. The Average Precision (PR-AP) is defined as the area under the PR curve; it summarizes the trade-off between precision and recall in a single scalar, with values in [0, 1] and higher values indicating better performance. The ROC curve plots the true positive rate (TPR) against the false positive rate (FPR) across thresholds, where TPR = νν νν+ νΉν ,FPR = νΉν νΉν+νν ,(11) and the corresponding true negative rate (TNR) and false negative rate (FNR) are defined as TNR = νν νν+ νΉν ,FNR = νΉν νΉν+νν .(12) Here, TNR measures the proportion of correctly identified non- transit samples, whereas FNR quantifies the fraction of transit signals missed by the classifier. The ROC curve provides a global measure of how well the model separates transit from non-transit signals regard- less of class prevalence. The Area Under the ROC Curve (ROC-AUC) is bounded in [0, 1]: a random classifier has AUC β 0.5, whilst a perfect classifier has AUC = 1. The ROC curve characterizes an algorithmβs overall discriminative ability across thresholds, whereas the PR curve emphasizes the trade-off between precision and recall. 3.2.2 Training configurations To address the class imbalance between transit and non-transit sig- nals, we adopt weighted random sampling with replacement during training. Each sample is assigned a weight inversely proportional to the number of samples in its class: ν€ ν = 1/ν ν ν where ν ν denotes the ground-truth class of sample ν, and ν ν ν is the total number of sam- ples belonging to class ν ν . The sampling probability of each sample is proportional to ν€ ν , so that minority-class samples are more likely to be selected during training. To mitigate the variability inherent in stochastic deep learning training, we adopt five-fold cross-validation that promotes a fair and objective comparison, in which the training and test sets are mutually exclusive, with a training-to-test ratio of 4:1 in each fold. Each model underwent multiple training sessions with carefully tuned hyperpa- rameters to reach near-optimal settings. To ensure objective five-fold cross-validation and mitigate training variability, comparison and vi- sualization used the median-performing model from a representative fold. To further enhance the generalization ability of the model and al- leviate overfitting, dropout layers (Srivastava et al. 2014) are added after each convolutional layer in all models. Ioffe & Szegedy (2015) introduced internal covariate shift and proposed BN to mitigate its ad- verse effects, which motivates our use of BN to improve training sta- bility. Model optimization is performed using AdamW (Loshchilov & Hutter 2019), which decouples weight decay from the adaptive gradient updates of Adam (Kingma & Ba 2015), improving regular- ization and generalization. A binary cross-entropy loss function is used. After empirical tuning, the final configuration adopts a learning rate of 8Γ 10 β4 and a weight decay of 1Γ 10 β6 . Dropout is applied at rates of 0.08 and 0.05 in both convolutional and attention layers, respectively, to reduce overfitting. The model is trained with a batch size of 3072, and early stopping (patience = 5) is employed to halt training when no further improvement in validation loss is observed. 3.2.3 Experimental Results Table 2 presents the evaluation metrics on the test set for the four model architectures compared in the experiments. Among these, TransitNet achieves the best overall performance, attaining an average F1-score of 99.1% on the test set after five-fold cross-validation. From the table, we observe that: β’ (a) and (d), (b) and (c): Introducing MHA enables the model to capture global temporal dependencies among data points, enhancing sensitivity to low-SNR transit signatures and improving the modeling of both transit events and complex photometric noise. β’ (a) and (b), (c) and (d): Compared to FCNN, using an FCN as the output mapping layer improves the modeling of spatial structure in the feature matrices while substantially reducing the parameter count by approximately 98% (from 17522K to 350K) and 79% (from 1807K to 380K), respectively. This is because FCN preserves spatial topology and exploits weight sharing, whereas FCNN flattens the features and requires significantly more parameters for the mapping. The paired comparisons above support that the combination of MHA and FCN can both improve the performance of low-SNR transit signal detection and substantially reduce model complexity. MNRAS 000, 1β20 (2026) 10 Xingchen Yan et al. Table 2. Evaluation results on transit samples for four model architectures, evaluated on the test sets and averaged over five-fold cross-validation to ensure statistical robustness: (a) FM + FCN, (b) FM + FCNN, (c) FM + MHA + FCNN, and (d) FM + MHA + FCN (TransitNet). All metrics are reported as percentages (%), and parameter counts are given in thousands (K,Γ1000). Operating-point metrics derived from the ROC and PR curves are reported as TPR at FPR = 1% and recall at P = 98%, consistent with the operating points shown in Fig. 6. Best-performing values are highlighted in bold, while second-best values are shown in italics. Comparative analysis indicates that replacing FCNN with FCN improves performance with substantially fewer parameters (a vs. b; c vs. d), incorporating MHA enhances representation learning (b vs. c; a vs. d). Model Accuracy PrecisionRecallF1-Score TPR @ FPR = 1% Recall @ P = 98% Num. of Params. (K) (a)98.0Β± 0.4 98.7Β± 0.9 97.4Β± 0.2 98.0Β± 0.496.9Β± 1.397.9Β± 0.8350 (b)97.8Β± 0.3 99.2Β± 0.2 96.4Β± 0.6 97.8Β± 0.396.8Β± 0.497.5Β± 0.417,522 (c)98.9Β± 0.2 99.5Β± 0.3 98.4Β± 0.3 98.9Β± 0.298.8Β± 0.499.2Β± 0.31,807 (d)99.1Β± 0.1 99.4Β± 0.2 98.8Β± 0.1 99.1Β± 0.199.1Β± 0.199.4Β± 0.1376 0.0000.0150.030 False Positive Rate 0.93 0.96 0.99 True Positive Rate FPR=0.01 ROC Curve 0.9600.9750.990 Recall 0.90 0.95 1.00 Precision P=0.98 PR Curve FM+FCN FM+MHA+FCNN FM+FCNN TransitNet (FM+MHA+FCN) Figure 6. ROC and PR curves on the test set for four model architectures: FM+MHA+FCNN, FM+FCNN, FM+FCN, and FM+MHA+FCN (Transit- Net). Curves are averaged over five-fold cross-validation. Filled circles in- dicate the operating points at a fixed false positive rate of FPR = 1% on the ROC curves and a fixed precision of Precision = 98% on the PR curves (Table 2), as highlighted by vertical dashed lines. 3.2.4 Stability and generalisation Stability and generalization are both relevant when evaluating a model: the former indicates whether comparable performance can be reproduced across training runs under fixed hyperparameters, whilst the latter relates to performance on held-out data and thus to potential applicability in real-world exoplanet searches. Table 2 reports each metric as the meanΒ± standard deviation over five-fold cross-validation. The standard deviations suggest that the joint use of FCN and MHA is associated with improved training stability. In particular, model (d) exhibits the smallest fold-to-fold variance across all evaluation metrics. Generalization performance is evaluated using test-set metrics (ac- curacy, precision, recall, F1) as well as operating-point statistics de- rived from the ROC and PR curves (TPR at FPR = 1% and R at P = 98%), as listed in the table. Meanwhile, TransitNet achieves the highest TPR = 99.1% at FPR = 1% and the highest R = 99.4% at P = 98% among all models. These operating-point results are consistent with the global trends shown in Fig. 6. The synergistic combination of TransitNet achieves the best per- formance, demonstrating that architectures tailored to transit signal characteristics are essential for optimal low-SNR transit detection, where MHA captures global temporal dependencies and enhances sensitivity to low-SNR transit signatures, while the FCN extracts local features and refines classification boundaries. 4 COMPARATIVE EVALUATION OF BLS, TLS AND TRANSITNET IN LOW-SNR EXOPLANET SEARCHES In this section, we conduct a systematic evaluation of TransitNet un- der realistic transit blind-search scenarios using datasets constructed from 60 randomly selected KIC targets that were not used during training, comparing its performance against BLS and TLS in the search for low-SNR transit signals. The experiments are divided into two parts based on two bench- mark datasets specifically constructed for complementary evaluation objectives. Section 4.2 presents a transit blind-search sensitivity anal- ysis using the Low-SNR Transit Recovery Set, which is built from a single KIC light curve into which individual transit signals with SNRs uniformly sampled are injected. This dataset is designed to quantify the ability of each method to recover low-SNR transit signals under controlled noise conditions. Section 4.3 evaluates cross-target gen- eralization using the Cross-KIC Recovery Set, a dataset constructed from multiple KIC targets spanning diverse stellar variability and noise environments. This benchmark is designed to assess the ro- bustness, stability, and generalization capability of each method in large-scale transit blind-search scenarios. Section 4.4 investigates the performance of different transit blind-search algorithms in terms of recovery rates for ATLCs containing injected planets with radii not exceeding that of the Earth. 4.1 Experimental Setup and Transit Blind-search Configuration To emulate realistic transit blind-search scenarios as closely as possi- ble while ensuring a fair comparison, a common search configuration is adopted for all methods. All analyses presented in this section are based on the detection scores extracted from the spectra of ATLCs and PCP TMLCs gen- erated from 60 randomly selected unseen KIC targets, following the semi-synthetic approach described in Section 3.1.2. The injected transit signals span orbital periods of 30β 60 days and SNRs of 6β 15. BLS, TLS, and TransitNet operate on the same detrended and flux-normalized light curves and produce detection-score spectra evaluated on an identical period grid. The period grid follows the optimal sampling scheme adopted by TLS (Hippke & Heller 2019). To ensure a consistent and fair comparison across methods, we adopt different score-selection strategies for positive and negative samples. For ATLCs containing injected transit signals, the detection score is evaluated at the true injected period for all methods. For PCP TMLCs, which are expected to contain no detectable transit signatures, the score assigned to BLS and TLS is defined as the maximum value in the corresponding detection-score spectrum. For MNRAS 000, 1β20 (2026) TransitNet 11 TransitNet, we instead use the most significant peak relative to the local background level. This choice reflects realistic transit blind- search scenarios in which a dominant peak in an otherwise non- transiting light curve may be interpreted as a candidate signal and selected for further inspection. 4.2 Single-KIC Transit Blind-search Sensitivity in the Low-SNR Regime The objective of this experiment is to evaluate the sensitivity of each algorithm to transit signals with SNR = 6β15 within a single target system. From the 60 randomly selected unseen KIC targets described above, we select one KIC light curve exhibiting minimal residual systematics after detrending and use it to construct the Low- SNR Transit Recovery Set, a benchmark dataset comprising 1000 ATLCs and 1000 PCP TMLCs. 4.2.1 Detection Score Distributions Following the score-extraction procedure described in Section 4.1, we obtain detection score spectra for all ATLCs and PCP TMLCs, yielding the score distributions of BLS, TLS, and TransitNet shown in Fig. 7. For each method, the red and blue histograms represent the score distributions of transit-containing ATLCs and non-transit PCP TMLCs, respectively. A smaller overlap between the two distri- butions indicates a stronger ability to distinguish transit signals from non-transit variability. The figure also highlights the operating thresholds corresponding to a fixed FPR of 1% and a fixedP of 95%, together with the associ- ated performance metrics. Among the three methods, TransitNet ex- hibits the clearest separation between the transit and non-transit score distributions, indicating superior discrimination capability. Conse- quently, it achieves higher transit recovery rates while maintaining a lower FPR, demonstrating a more favorable trade-off between sensi- tivity and reliability in realistic transit-search scenarios. 4.2.2 Transit Blind-search Sensitivity Across low-SNR Regimes This experiment evaluates the performance of each method across different SNR regimes. Fig. 8 shows the ROC and PR curves of BLS, TLS, and TransitNet on the Low-SNR Transit Recovery Set. Tran- sitNet consistently achieves superior performance, with the largest advantage observed in the lowest-SNR bins. The performance gap gradually narrows as the SNR increases. Fig. 9 complements this analysis by showing the score distributions of transit and non-transit samples within each SNR bin. The overlap between the two distributions reflects the difficulty of distinguishing low-SNR transit signals from noise. The shaded NT-prone region below the threshold corresponding to a FPR of 1% highlights signals most susceptible to missed detections, as well as noise fluctuations that may be incorrectly identified as transit candidates. Consistent with the ROC and PR results, TransitNet exhibits substantially less overlap between the two distributions, indicating stronger sensitivity to low-SNR transit signals. 4.3 Cross-KIC Generalisation and Performance Evaluation The objective of this experiment is to evaluate the generalization capability and overall transit-search performance of each algorithm across multiple unseen target systems. To this end, we construct 6810121416 SDE 0 200 400 600 Count BLS SDE Distribution Transit Non-Transit Transit Median Score = 9.66 FPR = 1.0%, TPR=93.7% (ROC) R=93.7%, P=99.0% (PR) Threshold = 6.14 P = 95.0%, R=95.9% (PR) FPR=5.0%, TPR=95.9% (ROC) Threshold = 5.95 1020304050 SDE 0 200 400 600 Count TLS SDE Distribution Transit Non-Transit Transit Median Score = 27.97 FPR = 1.0%, TPR=94.0% (ROC) R=94.0%, P=98.9% (PR) Threshold = 10.35 P = 95.0%, R=95.7% (PR) FPR=5.0%, TPR=95.7% (ROC) Threshold = 9.01 0.00.20.40.60.81.0 Score 0 200 400 600 800 Count TransitNet Score Distribution Transit Non-Transit Transit Median Score = 1.00 FPR = 1.0%, TPR=99.4% (ROC) R=99.4%, P=99.8% (PR) Threshold = 0.26 P = 95.0%, R=99.5% (PR) FPR=5.2%, TPR=99.5% (ROC) Threshold = 0.13 Figure 7. Detection score distributions of transit and non-transit light curves on the Low-SNR Transit Recovery Set for BLS, TLS, and TransitNet. Tran- sitNet exhibits the clearest separation between transit and non-transit popu- lations, indicating improved recoverability of low-SNR transit signals while maintaining strong rejection of non-transit backgrounds. the Cross-KIC Recovery Set, another independent benchmark com- prising 1000 ATLCs with SNR β [6, 15] and 1000 PCP TMLCs generated from TMLCs derived from the 60 randomly selected KIC targets described above. This benchmark is entirely disjoint from the Low-SNR Transit Recovery Set used in the previous experiment and is designed to assess algorithm robustness across diverse stellar targets and noise environments. 4.3.1 Overall Detection Performance and Threshold Optimization For each source KIC target, multiple transit signals spanning different SNR levels are independently injected into the corresponding TMLC to generate ATLCs, while multiple PCP TMLCs are constructed from the same source. The resulting transit and non-transit samples yield a detection- score distribution specific to that KIC target for each algorithm. By computing the ROC-AUC and PR-AP from these score distributions, we obtain a pair of performance metrics for each KIC. Repeating this procedure across all 60 KIC targets produces the distributions of ROC-AUC and PR-AP shown in Fig. 10, while the per-KIC detection- score distributions for individual methods are provided in Figs. B2β B4. A stronger concentration of these metrics near unity indicates bet- ter cross-target generalization, greater robustness to the diverse noise characteristics present in different KIC targets, and more stable sen- sitivity to low-SNR transit signals. TransitNet achieves the strongest MNRAS 000, 1β20 (2026) 12 Xingchen Yan et al. 0.00.20.40.60.81.0 False Positive Rate 0.4 0.5 0.6 0.7 0.8 0.9 1.0 True Positive Rate ROC - SNR [6,8) TLS (AUC=0.912) BLS (AUC=0.894) TransitNet (AUC=0.986) 0.00.20.40.60.81.0 False Positive Rate 0.86 0.88 0.90 0.92 0.94 0.96 0.98 1.00 True Positive Rate ROC - SNR [8,10) TLS (AUC=0.999) BLS (AUC=0.993) TransitNet (AUC=1.000) 0.00.20.40.60.81.0 False Positive Rate 0.86 0.88 0.90 0.92 0.94 0.96 0.98 1.00 True Positive Rate ROC - SNR [10,12) TLS (AUC=1.000) BLS (AUC=1.000) TransitNet (AUC=1.000) 0.00.20.40.60.81.0 False Positive Rate 0.86 0.88 0.90 0.92 0.94 0.96 0.98 1.00 True Positive Rate ROC - SNR [12,15] TLS (AUC=1.000) BLS (AUC=1.000) TransitNet (AUC=1.000) 0.00.20.40.60.81.0 Recall 0.4 0.5 0.6 0.7 0.8 0.9 1.0 Precision PR - SNR [6,8) TLS (AP=0.944) BLS (AP=0.935) TransitNet (AP=0.992) 0.00.20.40.60.81.0 Recall 0.86 0.88 0.90 0.92 0.94 0.96 0.98 1.00 Precision PR - SNR [8,10) TLS (AP=0.999) BLS (AP=0.996) TransitNet (AP=1.000) 0.00.20.40.60.81.0 Recall 0.86 0.88 0.90 0.92 0.94 0.96 0.98 1.00 Precision PR - SNR [10,12) TLS (AP=1.000) BLS (AP=1.000) TransitNet (AP=1.000) 0.00.20.40.60.81.0 Recall 0.86 0.88 0.90 0.92 0.94 0.96 0.98 1.00 Precision PR - SNR [12,15] TLS (AP=1.000) BLS (AP=1.000) TransitNet (AP=1.000) Figure 8. ROC and PR curves for BLS, TLS, and TransitNet on the Low-SNR Transit Recovery Set, evaluated in four SNR bins: [6, 8), [8, 10), [10, 12), and [12, 15]. Each bin contains equal numbers of transit and non-transit samples derived from the same KIC target. TransitNet achieves substantially higher detection performance in the low-SNR regime, while the performance differences diminish as SNR increases. overall performance, with mean ROC-AUC and PR-AP values of 0.974 and 0.982, respectively, outperforming both BLS and TLS. To enable a fair and reproducible comparison across methods, we introduce a unified threshold-selection procedure for transit blind- search evaluation. For each algorithm, all detection scores obtained from the Cross-KIC Recovery Set are collected, and a common binary decision rule is applied. Candidate thresholds are evaluated using macro-averaged TPR and FPR computed across KIC targets, from which the Youden statistic (Youden 1950) is derived. ν½(ν) = TPR(ν)βFPR(ν),(13) The threshold ν that maximizes ν½(ν) is adopted as the global oper- ating point for each algorithm. The corresponding definition of the average classification error is given by νΈ(ν) = 1 2 FPR(ν)+FNR(ν) ,(14) where the equivalence between the maximization of the Youden index and the minimization of νΈ(ν) is provided in Appendix B, along with the corresponding pseudocode for the threshold-selection procedure. The optimal operating threshold for each method is selected by maximizing the difference between the TPR and FPR, equiva- lent to maximizing Youdenβs J statistic. Fig. 11 shows the result- ing threshold-sweep curves. The optimal operating thresholds are SDE = 6.76 for BLS, SDE = 9.80 for TLS, and Score = 0.54 for TransitNet. While BLS and TLS exhibit sharp maxima, indicating strong sensitivity to threshold selection, TransitNet maintains near- optimal performance over a substantially broader range of thresholds. This behavior suggests greater robustness to threshold perturbations and more stable detection performance across heterogeneous stellar noise environments. The selected thresholds are subsequently fixed for the remaining experiments. Using the selected operating thresholds, we further visualize the confusion matrices of all methods on the Low-SNR Transit Recovery Set. TransitNet consistently outperforms TLS and BLS, achieving the highest overall accuracy of 98.9% on the complete dataset (Fig. B1). Its advantage is particularly pronounced in the challenging low-SNR regime (SNR= 6β8), where it attains an accuracy of 95.2% (Fig. 12). 4.3.2 Case studies on semi-synthetic low-SNR transits We further examine the detection-score spectra produced by all three methods on semi-synthetic ATLCs with injected transits at different SNR levels, highlighting the practical implications of the threshold- based classification and illustrating the superior sensitivity of Tran- sitNet (Fig. 13). For an ATLC with an injected transit signal of SNR = 8.8 and a true period of 35.38 days (left column), all three algorithms successfully recover the primary signal, with detection scores exceeding their respective thresholds. Notably, TransitNet additionally identifies a secondary peak above the threshold at approximately 53 days, which may correspond to a period alias or apotential secondary signal. More significantly, for an ATLC with a lower SNR of 6.5 and a true period of 39.51 days (right column), TransitNet successfully detected the signal with a score well above its threshold of 0.54, while BLS and TLS failed to detect the signal as their scores at the true period remained below their respective thresholds. This observation directly demonstrates TransitNetβs enhanced sen- sitivity to low-SNR transit signals that is missed by TLS and BLS under the adopted thresholds in this example, highlighting a key ad- vantage for detecting signals near the detection limit in real transit blind searches. 4.4 Earth-size and Sub-Earth-size Transit Recovery Analysis To further assess the sensitivity of different transit blind-search al- gorithms to Earth-size and sub-Earth-size planets, we conducted an MNRAS 000, 1β20 (2026) TransitNet 13 [6, 8)[8, 10)[10, 12)[12, 15]Non-transit SNR Interval 6 8 10 12 14 SDE 60 158 4 217 0 222 0 333 BLS SDE Distribution by SNR Interval Transit Non-Transit NT-prone Region (FPR=1%) Missed Recovered Mean of Dist. [6, 8)[8, 10)[10, 12)[12, 15]Non-transit SNR Interval 10 20 30 40 SDE 58 160 3 218 0 222 0 333 TLS SDE Distribution by SNR Interval Transit Non-Transit NT-prone Region (FPR=1%) Missed Recovered Mean of Dist. [6, 8)[8, 10)[10, 12)[12, 15]Non-transit SNR Interval 0.0 0.2 0.4 0.6 0.8 1.0 Score 7 211 0 221 0 222 0 333 TransitNet Score Distribution by SNR Interval Transit Non-Transit NT-prone Region (FPR=1%) Missed Recovered Mean of Dist. Figure 9. Violin plots of detection scores on the Low-SNR Transit Recovery Set for BLS, TLS, and TransitNet, stratified by SNR bins [6, 8), [8, 10), [10, 12), and[12, 15], with an additional column aggregating all non-transit light curves (rightmost). The light pink shaded band, termed the non-transit- prone (NT-prone) region, is bounded above by the score threshold corre- sponding to a FPR of 1%. TransitNet consistently places a larger fraction of low-SNR transit signals above the NT-prone region, indicating superior recoverability under realistic false-alarm constraints. ROC-AUC Distribution by KIC BLS Mean AUC=0.947 TLS Mean AUC=0.958 TransitNet Mean AUC=0.974 0.60.70.80.91.0 PR-AP Distribution by KIC BLS Mean AP=0.967 TLS Mean AP=0.975 TransitNet Mean AP=0.982 Figure 10. Distribution of per-KIC ROC-AUC and PR-AP values on the Cross-KIC Recovery Set. For each KIC, the metrics are computed from the corresponding transit and non-transit detection scores spectrum. These dis- tributions evaluate the transit recovery performance of different algorithms across previously unseen KICs. Distributions that are more concentrated near 1 indicate better and more stable performance under varying stellar noise conditions. 51015 SDE 0.0 0.5 1.0 6.76 BLS 2040 SDE 9.80 TLS 0.20.40.60.81.0 Score 0.54 TransitNet Youden J (TPRFPR) 1 2 (FPR + FNR) Figure 11. Threshold selection for BLS, TLS, and TransitNet on the Cross- KIC Recovery Set based on maximising Youdenβs statistic, ν½ = TPRβ FPR, and equivalently minimising the mean classification error, 1 2 (FPR+ FNR). Metrics are computed independently for each KIC and then macro-averaged across all KICs. Vertical dashed lines indicate the adopted operating thresh- olds (BLS: 6.76, TLS: 9.80, and TransitNet: 0.54), which can serve as practical global operating thresholds in transit blind-searches. The narrow optima of BLS and TLS indicate strong sensitivity to threshold selection, whereas Tran- sitNet maintains near-optimal performance over a broad range of thresholds, suggesting greater robustness across diverse stellar noise environments. Earth-recovery experiment based on real Kepler photometric obser- vations. Unlike the previous Low-SNR Transit Recovery Set and Cross- KIC Recovery Set, this experiment specifically targets transit signals with planet radii not exceeding that of the Earth (ν ν β©½ 1ν β ). The objective is to quantify the fraction of Earth-size and sub-Earth-size transit signals that can be recovered by each algorithm under realistic stellar and instrumental noise conditions. The experiment was conducted on an independent set of 60 unseen KIC targets excluded from training, using the corresponding TMLCs as background light curves. For each target, the stellar radius ν β from the KOI catalog was used to define an Earth-size reference transit depth, νΏ β = (ν β /ν β ) 2 , serving as an upper bound on injected signal strength. For each ATLC, a qualified background light curve was randomly selected. The orbital period was uniformly sampled from ν β [30, 60] d, and the transit duration was drawn from the empirical distribution of confirmed KOIs within the same period range. A target SNR was then uniformly sampled from[6, 15], and the correspond- ing transit depth was obtained by inverting the standard SNR scaling relation (Eq. 6), using the measured noise level ν, period ν, and transit duration ν 14 . This yields dynamically varying transit depths constrained to sub-Earth-to-Earth-size regimes via νΏ β . The corre- sponding planetary radius was computed as ν ν = ν β β νΏ, and only samples satisfying ν ν β©½ 1ν β were retained. Finally, transit signals were injected into the original light curves using a trapezoidal tran- sit model, producing a total of 1000 Earth-size and sub-Earth-size ATLCs. TransitNet, TLS, and BLS were subsequently applied to all ATLCs (Fig. 14). The recovery rate is defined as ν recovered /ν total , represent- ing the fraction of injected Earth-size and sub-Earth-size transits successfully recovered under realistic Kepler noise conditions. Tran- sitNet achieves a recall of 93.0%, substantially outperforming TLS (63.1%) and BLS (60.0%). The markedly higher recovery rate indi- cates that our deep-learning approach is significantly more sensitive than conventional search algorithms, highlighting its potential for improving the completeness of terrestrial exoplanet detections. No- tably, the performance gain is concentrated primarily in the SNR range of 6β8, consistent with the results obtained in the previous experiments. This finding confirms that TransitNet provides its greatest advan- tage near the practical detection threshold of traditional transit-search methods, where low-SNR transit signals are most likely to be missed. MNRAS 000, 1β20 (2026) 14 Xingchen Yan et al. Non-TransitTransit Predicted Label Non-Transit Transit True Label 2180 114104 SNR [6,8) BLS Accuracy: 73.9% Non-TransitTransit Predicted Label Non-Transit Transit True Label 2126 51167 SNR [6,8) TLS Accuracy: 86.9% Non-TransitTransit Predicted Label Non-Transit Transit True Label 2180 21197 SNR [6,8) TransitNet Accuracy: 95.2% 0 50 100 150 200 0 50 100 150 200 0 50 100 150 200 Figure 12. Confusion matrices showing the classification performance of BLS, TLS, and TransitNet on the Low-SNR Transit Recovery Set with SNR β [6, 8). Classification is performed using the operating thresholds selected from the macro-averaged threshold optimisation analysis shown in Fig. 11. Among the three methods, TransitNet demonstrates the best overall performance, achieving the highest accuracy (95.2%), compared with TLS (86.9%) and BLS (73.9%). The overall confusion matrix evaluated on the complete dataset is shown in Fig. B1. 3540455055 Period [day] 0 10 20 SDE BLS - ATLC | SNR = 8.8 Threshold = 6.76 True Period = 35.38 d 3540455055 Period [day] 2.5 0.0 2.5 5.0 7.5 SDE BLS - ATLC | SNR = 6.5 Threshold = 6.76 True Period = 39.51 d 3540455055 Period [day] 0 10 20 30 SDE TLS - ATLC | SNR = 8.8 Threshold = 9.80 True Period = 35.38 d 3540455055 Period [day] 0 5 10 SDE TLS - ATLC | SNR = 6.5 Threshold = 9.80 True Period = 39.51 d 3540455055 Period [day] 0.0 0.2 0.4 0.6 0.8 1.0 Score TransitNet - ATLC | SNR = 8.8 Threshold = 0.54 True Period = 35.38 d 3540455055 Period [day] 0.0 0.2 0.4 0.6 0.8 1.0 Score TransitNet - ATLC | SNR = 6.5 Threshold = 0.54 True Period = 39.51 d Figure 13. Detection score spectrum for BLS, TLS, and TransitNet on semi-synthetic ATLCs. The left column shows results for an injected transit signal with SNR = 8.8 and true period of 35.38 days, where all three algorithms successfully detect the target transit, with scores above the operating thresholds selected in Fig. 11. The right column shows results for a lower-SNR signal (SNR = 6.5) with true period of 39.51 days, where the target transit is recovered only by TransitNet. This comparison directly illustrates TransitNetβs enhanced sensitivity, enabling detection of low-SNR transit that would be missed by BLS or TLS. 4.5 Speed and Inference Efficiency Beyond detection accuracy, computational efficiency is a critical con- sideration for large-scale transit blind-search surveys. To evaluate the runtime performance of TransitNet, we conducted a benchmark comparison against both CPU and GPU implementations of tradi- tional transit detection methods, including BLS (KovΓ‘cs et al. 2002) and TLS (Hippke & Heller 2019), as well as GPU-accelerated BLS (GPU-BLS; Hoffman 2022). All GPU benchmarks were performed on an NVIDIA RTX 4090 GPU with 24,564 MiB of memory. We randomly selected a real Kepler light curve from a confirmed exoplanet host star. Following the grid-generation scheme proposed by Hippke & Heller (2019), period grids were generated with oversampling factors (OS) ranging from 1 to 10. For each OS, the corresponding set of trial periods MNRAS 000, 1β20 (2026) TransitNet 15 30405060 Period [d] 0.6 0.7 0.8 0.9 1.0 R p [R ] Recall: 93.0% TransitNet 68101214 SNR 0 200 Count 30405060 Period [d] 0.6 0.7 0.8 0.9 1.0 Recall: 63.1% TLS 68101214 SNR 0 200 30405060 Period [d] 0.6 0.7 0.8 0.9 1.0 Recall: 60.0% BLS 68101214 SNR 0 200 Confirmed PlanetsRecoveredMissed Figure 14. Recovery of injected Earth-size and sub-Earth-size transits in Kepler TMLCs. Top: planet radius versus period; blue and gray circles denote injected transits recovered and missed at algorithm-specific operating thresholds (Fig. 11), and pink diamonds mark confirmed Kepler planets. Dashed guides indicate ν p = 1ν β and ν = 30 d. Bottom: stacked histograms of the injected sample as a function of SNR. TransitNet achieves 93.0% recall, substantially exceeding TLS (63.1%) and BLS (60.0%). 102030 Period grid length (N periods , Γ10 3 ) 0 10 20 30 40 50 60 Mean runtime (s) Inference Speed 102030 Period grid length (N periods , Γ10 3 ) 0 5 10 15 20 25 30 Speedup factor ( T CPU-TLS / T ) Relative Speedup CPU-TLSCPU-BLSGPU-BLSTransitNet Figure 15. Runtime scaling of transit-search algorithms as a function of period-grid size. Dashed lines denote CPU-based implementations, while solid lines denote GPU-accelerated implementations. The left panel shows the mean wall-clock inference time; the right panel shows the speed-up factor relative to CPU TLS, ν CPU-TLS /ν. Curves compare CPU-TLS, CPU-BLS, GPU-BLS, and TransitNet. The horizontal axis gives ν periods in units of 10 3 , controlled by sweeping the period-grid oversampling factor from 1 to 10. covering 30β60 days was constructed, and the runtime of each method was evaluated. To ensure a fair comparison, each algorithm was first subjected to one warm-up run, after which the execution time was measured over three consecutive trials. All methods were evaluated using the same input light curve. For TransitNet, the reported runtime includes both the GPU-based folding stage and the subsequent neural-network inference. This protocol helps ensure that TLS, BLS, and TransitNet are compared under identical search ranges and period-grid configurations. TransitNet combines superior low-SNR transit recovery with com- putational efficiency comparable to the fastest GPU-accelerated base- line. Over the tested range of oversampling factors (OS = 1β10), corresponding to ν periods β 3.8Γ10 3 β3.8Γ10 4 , inference completes within a few seconds, yielding speed-ups of βΌ12β25Γ relative to CPU-TLS andβΌ4β5Γ relative to CPU-BLS, while remaining within a factor ofβΌ2 of GPU-BLS in absolute runtime (Fig. 15). These characteristics make TransitNet well suited for future large- scale transit blind-search surveys, where both sensitivity to low-SNR transits and computational efficiency are essential for discovering planetary signals among millions of monitored stars, despite several thousand exoplanets having been confirmed to date. 4.6 Transit Midpoint Estimation from Multi-Head Attention Transit window and midpoint estimation exploit the content-adaptive attention mechanism of the MHA described in Section 2.2, in which phase segments consistent with transit morphology are assigned higher attention weights. The one-head attention matrix ν΄ β R νΏ 1 ΓνΏ 1 (Eq. 3) is computed from query and key representations projected from the downsampled features produced by the FM, rather than directly on the original νΏ 0 - bin phase grid. Here νΏ 0 = 4096 and νΏ 1 = 128 denote the lengths of the input sequence and the downsampled feature sequence after the FM, respectively. Therefore, the MHA attention matrix (βΓνΏ 1 ΓνΏ 1 ) requires head aggregation and upsampling before being aligned with the transit window on the original νΏ 0 -bin input sequence. Attention aggregation: Let the multi-head attention matrix be denoted as Aβ R βΓνΏΓνΏ (where νΏ = νΏ 1 is used in this work). The outputs of the β attention heads are averaged, followed by aggregation along the query dimension. ν€ ν = 1 βνΏ β βοΈ ν νΏ βοΈ ν ν΄ (ν) ν,ν .(15) This yields the attention vector w β R νΏ , which quantifies the overall importance assigned to each source position for transit morphology recognition under the global attention mechanism. Temporal rescaling: Since νΏ 1 βͺ νΏ 0 , directly estimating the tran- sit midpoint from the low-resolution w would introduce estimation errors. Therefore, w is remapped onto the phase grid corresponding to the original input sequence in order to restore temporal resolution. Specifically, w is first interpolated using cubic spline interpolation and resampled onto a uniform phase grid of length νΏ 0 , yielding the in- terpolated attention vector w β² β R νΏ 0 . Subsequently, the interpolated attention vector is smoothed with a five-bin moving-average filter to MNRAS 000, 1β20 (2026) 16 Xingchen Yan et al. reduce possible interpolation artifacts, followed by renormalization to preserve the total attention mass. Transit-window estimation: The phase bin corresponding to the global maximum of the attention distribution is identified as the refer- ence peak. The algorithm then expands in both directions until the at- tention decreases to the full width at half maximum (FWHM), thereby defining the estimated transit window[ Λν ν , Λν ν ]. Due to interpolation- induced shifts after upsampling, the attention peak location Λν peak is used only for transit-window localization rather than as the final transit-midpoint estimate. Transit midpoint estimation: Within the estimated transit win- dow [ Λν ν , Λν ν ], the phase bin with the minimum normalized flux is selected as the estimate of the transit midpoint Λν 0 . After obtaining Λν 0 , the transit window is re-centered on Λν 0 while preserving its width, Ξ Λν = Λν ν β Λν ν , thereby updating the interval [ Λν β² ν , Λν β² ν ]. This strategy constrains the transit midpoint using the minimum of the normal- ized flux within the estimated transit window. Because the attention- derived window is obtained from a downsampled representation, its boundaries may not perfectly coincide with the true transit interval and are subject to interpolation uncertainty when projected back to the original sequence. Rather than relying on peaks in the atten- tion vector w β² , the proposed method identifies the midpoint directly from the original photometric measurements within the estimated window. This design represents a practical compromise between lo- calization precision and computational efficiency, combining coarse attention-based localization with refinement from higher-resolution flux measurements. For applications requiring higher midpoint lo- calization precision, this step may be further replaced by refined tem- plate matching within the estimated transit window, thereby yielding a more accurate estimate of the transit midpoint. To quantitatively evaluate the accuracy of transit-window estima- tion, we introduce a transit-window coverage criterion (Fig. 16). Let the estimated transit window be denoted by [ Λν ν , Λν ν ] and the true transit window by [ν 1 ,ν 4 ]. The covered duration of the true transit interval is defined as ν· = max 0, min( Λν ν ,ν 4 )β max( Λν ν ,ν 1 ) ,(16) where the total true transit duration is given by ν 14 = ν 4 β ν 1 . The transit-window coverage score is then defined as ν = ν· ν 14 ,(17) which represents the fraction of the true transit interval covered by the estimated window and ranges from 0 to 1. Specifically, ν = 1 when the true transit interval is completely covered by the estimated window, 0 < ν < 1 in cases of partial coverage, and ν = 0 when no part of the true transit interval is covered by the estimated window. In addition, an independent evaluation dataset containing βΌ7000 transit samples was constructed to assess the transit window and midpoint estimation capabilities of TransitNet. Table 3 summarises the proportions of the three coverage categories together with the mean overlap score achieved by TransitNet on this dataset. Fig. 17 further illustrates the variation of the coverage score ν as a function of SNR, ν, ν 14 , and νΏ. Across the investigated parameter space, the coverage score remains consistently high, with most bins satisfying ν > 0.97. The coverage score depends more strongly on SNR and νΏ, approaching nearly complete coverage for stronger transit signals, whereas only minor fluctuations are observed with orbital period and transit duration. Fig. 18 illustrates the estimated Λν 0 for two genuine Kepler targets using TransitNet, along with their visualization of the processed attention weight sequence and the refined transit window. A total of 34 confirmed Kepler exoplanet samples satisfying ν β Figure 16. Schematic illustration of the transit-region scoring scheme com- paring an estimated transit window, [ Λν ν , Λν ν ] (light blue), with the true in- transit interval,[ν 1 , ν 4 ] (magenta). Black points show synthetic photometry generated from a Batman model (solid red curve), and the hatched region denotes the covered transit interval. The three panels illustrate partial cover- age (left; 0 < ν < 1), full coverage (centre; [ν 1 , ν 4 ] β [ Λν ν , Λν ν ], ν = 1), and no coverage (right; ν = 0). For partial coverage, the overlap duration is ν· = max ( 0, min( Λν ν , ν 4 ) β max( Λν ν , ν 1 ) ) , with score ν = ν·/ν 14 and ν 14 = ν 4 β ν 1 . Table 3. Transit-window coverage statistics and corresponding midpoint esti- mation errors on the independent evaluation dataset. The coverage score quan- tifies the fraction of the true transit interval covered by the estimated window, while the midpoint estimation error is defined as|ν 0 β Λν 0 |. The large errors in the no-coverage cases arise from complete misalignment with the true transit interval. Across all samples, the mean coverage score is 0.982Β± 0.124. CategoryFraction (%) Mean score Mean |ν 0 β Λν 0 | (h) Full coverage97.41.00.05 Partial coverage1.30.60.34 No coverage1.30.04.57 8101214 SNR 0.92 0.94 0.96 0.98 1.00 3540455055 Period [day] 0.96 0.97 0.98 0.99 1.00 34567 Duration [hour] 0.96 0.97 0.98 0.99 1.00 10 2 2 Γ 10 2 5 Γ 10 2 10 3 Depth 0.96 0.97 0.98 0.99 1.00 Transit Window Coverage Transit Window Coverage vs. Transit Parameters Figure 17. Transit window overlap versus transit and observational param- eters on a separate evaluation set. Points show bin-averaged overlap scores; shaded bands indicate meanΒ± standard error of the mean (SEM). The score measures agreement between the estimated ingress-egress interval [ Λν ν , Λν ν ] and the true transit window. Panels: SNR (top left), orbital period in days (top right), transit duration in hours (bottom left), and transit depth (bottom right, log scale). [30, 60] d and SNRβ [6, 15] were selected for testing (see Appendix Table C1). Following Section 4.1, all samples in this regime were recovered with scores above the threshold of 0.54. The true transit midpoints ν 0 of all samples fall within the estimated transit windows [ Λν ν , Λν ν ], supporting the usefulness of the attention-based method for initial transit-window and epoch estimates. TransitNet jointly performs transit detection and transit midpoint MNRAS 000, 1β20 (2026) TransitNet 17 0.9990 0.9995 1.0000 1.0005 Normalized Flux KIC 2165002 | P=47.333 d | SNR=14.0 | Score=1.0 Transit Region [8.88, 9.98] True 0 =9.42 Predicted 0 =9.42 8.08.59.09.510.010.511.0 Time [day] 0 1 Attention 0.99900 0.99925 0.99950 0.99975 1.00000 1.00025 1.00050 Normalized Flux KIC 8231667 | P=42.877 d | SNR=14.8 | Score=1.0 Transit Region [26.96, 27.96] True 0 =27.44 Predicted 0 =27.45 26.026.527.027.528.028.5 Time [day] 0 1 Attention Figure 18. Analysis of two examples. The top panel in each figure shows the normalised flux over time. The estimated transit midpoint ( Λν 0 ) is indicated by the red dashed line, while the true ν 0 is marked by the blue dashed line. The bottom panel displays the attention weights, with peaks indicating the temporal segments most relevant for transit detection. estimation in a single forward pass. This end-to-end formulation may support upstream and downstream survey tasks, including automated triage, transit-preserving detrending (by masking candidate transit regions prior to baseline removal), and the initialization of MCMC- based light-curve modeling (by providing initial ν 0 estimates), and other large-scale follow-up studies. Further dedicated validation will be required to assess these potential applications. 5 CONCLUSION AND DISCUSSION 5.1 Conclusion Motivated by the observational incompleteness of intermediate-to- long-period (I2LP) Earth-size planets, we develop TransitNet, a com- pact attention-based framework for low-SNR transit blind searches. Beyond the model architecture, we further introduce a data con- struction and benchmarking pipeline, in which the Periodic Chunk Permutation (PCP) strategy is introduced to construct reliable non- transit controls, and recovery-oriented evaluation benchmarks are designed to assess transit detection sensitivity under realistic stellar- noise conditions. Extensive experiments demonstrate that TransitNet consistently outperforms classical transit-search methods (TLS and BLS). On the target I2LP benchmark (ν β [30, 60] days, SNR β [6, 15]) during model training, the proposed architecture achieves strong classifica- tion performance in our benchmarks while requiring only βΌ376 K trainable parameters. Controlled ablation studies further confirm that the combination of learnable denoising, global attention, and lightweight convolutional mapping is particularly effective for recov- ering low-SNR transit signals embedded in realistic stellar variability and instrumental noise. Under realistic transit blind-search conditions, TransitNet main- tains a clear advantage over both TLS and BLS. On the Low-SNR Transit Recovery Set, it achieves an accuracy of 95.2% in the chal- lenging SNR = 6β8 regime, substantially exceeding the performance of the classical methods. On the Cross-KIC Recovery Set, mean ROC- AUC and PR-AP values of 0.974 and 0.982 demonstrate robust gener- alization across heterogeneous stellar noise environments. The model also exhibits strong threshold robustness, maintaining near-optimal performance across a broad operating range and thereby reducing calibration requirements for large-scale survey deployment. In an injected Earth-size and sub-Earth-size transit recovery experiment, TransitNet achieves a recovery rate of 93.0%, substantially exceeding those of TLS (63.1%) and BLS (60.0%). Beyond transit detection, TransitNet supports joint transit local- ization and ν 0 estimation directly from the learned attention patterns, reducing the need for a separate localization stage. On an independent evaluation set, the estimated transit windows achieve a mean cover- age score of 0.982Β± 0.124, while all 34 confirmed Kepler planets within the target parameter space are successfully recovered. Applied to real observations, the model achieves a mean transit midpoint esti- mation error| Λν 0 βν 0 | of 1.24 h, demonstrating that physically useful initial transit parameters can be estimated alongside detection in a single inference pass. TransitNet further combines a compact footprint (1.5 MB) with computational efficiency comparable to the fastest GPU-accelerated classical baseline. Across the tested oversampling range (OS = 1β 10), corresponding to ν periods β 3.8Γ10 3 β3.8Γ10 4 , inference com- pletes within a few seconds, providing speed-ups ofβΌ12β25Γ relative to CPU-TLS andβΌ4β5Γ relative to CPU-BLS. Future transit surveys will process millions of stellar light curves to identify a comparatively small population of transiting exoplan- ets. In this regime, sensitivity, scalability, and interpretability are equally important. By combining robust low-SNR transit recovery, efficient deployment, and physically meaningful transit localization, TransitNet provides a practical framework for large-scale transit blind searches for Earth-size planets in the scientifically important I2LP regime and, more broadly, for weak-signal detection problems in time-domain astronomy. 5.2 Cross-task Transfer and Generalization Compared to Kepler, other survey missions (e.g., TESS) differ in cadence, photometric precision, and target populations. Conse- quently, the cross-instrument generalization and transferability of deep-learning-based transit-search models merit discussion. Depending on the transfer paradigm, generalization can be cate- gorized into two types. The first is direct transfer, in which a model trained on Kepler data is applied to other survey data without addi- tional modification. The second is indirect transfer, in which survey- specific datasets are constructed and the model is retrained. For direct transfer, its effectiveness depends on the following con- ditions: (i) the photometric precision and cadence of the light curves are sufficient to resolve the target exoplanet transit duration; and (i) the period range of the target signals is covered by the period distri- bution of the training data. For example, in the Kepler-based dataset adopted in this work, the cadence is 29.4 min. The transit samples are generated from ATLCs containing injected transit signals with periods of 30β60 days and SNR values between 6 and 15, whereas the non-transit samples are constructed from PCP TMLCs and GNLCs, with the noise standard deviations of the latter sampled from the corresponding source KICs. Under these conditions, direct transfer may be feasible but requires dedicated validation. Indirect transfer typically involves more complex data generation procedures and imposes higher demands on training efficiency and stability. The proposed model, with a lightweight architecture of approximately 1.5 MB and strong training stability (Section 2), as MNRAS 000, 1β20 (2026) 18 Xingchen Yan et al. well as its demonstrated computational efficiency in inference (Sec- tion 4.5), is well suited to survey-specific retraining and controlled direct-transfer tests. 5.3 Future Work Planned extensions include: (1) systematically evaluating cross-survey transferability beyond Kepler-based training by benchmarking both direct transfer (zero or minimal retuning) and indirect transfer (survey-specific retraining) on TESS and future facilities such as PLATO and ET. This program will quantify performance as a function of cadence, photometric precision, period-domain overlap with the training set, and stellar population mismatch, and will compare transfer learning, domain adaptation, and de novo training strategies. (2) extending the period search from the current ν β [30, 60] d regime to longer periods (150β200 d and beyond), which are more relevant to habitable-zone Earth analogs; this extension requires ex- plicit treatment of lower transit multiplicity, reduced folded SNR, and smaller phase coverage, potentially via hierarchical search strategies, long-period-focused data augmentation, and specialized templates, with sensitivity and false-positive control jointly evaluated. ACKNOWLEDGEMENTS Funding for this study is provided by the Strategic Priority Pro- gram on Space Science of the Chinese Academy of Sciences (XDA15020600) and Chinaβs Space Origins Exploration Program (GJ11030405). JPZ is grateful for the support of the National Natu- ral Science Foundation of China, Grant No. 12203087. The authors thank Hui Zhang, Shiyin Shen, Bo Ma, and Jiwei Xie for their valu- able comments and constructive suggestions, which helped improve the manuscript, and Zhenghong Liu for helpful discussions and as- sistance with the detrending procedure used in this work DATA AVAILABILITY The data used in this study are publicly available from the Kepler mission archive https://archive.stsci.edu/kepler/. The KOI and exoplanet parameters used in this work are obtained from the NASA Exoplanet Archive (Christiansen et al. 2025) (https://exoplanetarchive.ipac.caltech.edu/), including the KOI table (doi:10.26133/NEA4). The ATLCs and PCP TMLCs derived from Kepler TMLCs are based on these public data and can be accessed upon request from the corresponding author. REFERENCES Ansdell M., et al., 2018, ApJL, 869, L7 Armstrong D. J., Pollacco D., Santerne A., 2016, MNRAS, 465, 2634 Armstrong D. J., Gamper J., Damoulas T., 2020, MNRAS, 504, 5327 Auvergne M., et al., 2009, A&A, 506, 411 AydoΔan K., 2022, Exoplanet Detection by Machine Learning with Data Aug- mentation (arXiv:2211.15577), https://arxiv.org/abs/2211. 15577 Brown T. M., Latham D. W., Everett M. E., Esquerdo G. A., 2011, AJ, 142, 112 Bryson S., et al., 2020, AJ, 161, 36 Carter J. A., Agol E., 2013, ApJ, 765, 132 Catala C., The PLATO Consortium 2009, Exp Astron, 23, 329 Chintarungruangchai P., Jiang I.-G., 2019, PASP, 131, 064502 Chollet F., 2017, in Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR). IEEE, p 1251β1258, doi:10.1109/CVPR.2017.195 Choudhary A., Bandari S., Kushvah B. S., Swastik C., 2025, The Astronom- ical Journal, 170, 120 Christiansen J. L., et al., 2012, PASP, 124, 1279 Christiansen J. L., et al., 2025, PSJ, 6, 186 Cui K., Liu J., Feng F., Liu J., 2021, AJ, 163, 23 CuΓ©llar S., Granados P., Fabregas E., CurΓ© M., Vargas H., Dormido-Canto S., Farias G., 2022, PLoS ONE, 17, e0268199 Dattilo A., et al., 2019, AJ, 157, 169 Dvash E., Peleg Y., Zucker S., Giryes R., 2022, AJ, 163, 237 Fiscale S., Ferone A., Ciaramella A., Inno L., Giordano Orsini M., Covone G., Rotundi A., 2025, Electronics, 14, 1738 Fressin F., et al., 2013, ApJ, 766, 81 Fulton B. J., et al., 2017, The Astronomical Journal, 154, 109 G P., Kumari A., 2023, Identification and Classification of Exoplanets Using Machine Learning Techniques (arXiv:2305.09596), https: //arxiv.org/abs/2305.09596 Ge J., et al., 2022a, ET White Paper: To Find the First Earth 2.0 (arXiv:2206.06693), https://arxiv.org/abs/2206.06693 Ge J., Zhang H., Deng H., Howell S. B., 2022b, The Innovation, 3, 100271 Ge J., et al., 2022c, in Coyle L. E., Matsuura S., Perrin M. D., eds, Proc. SPIE Vol. 12180, Space Telescopes and Instrumentation 2022: Optical, In- frared, and Millimeter Wave. SPIE, p. 1218015, doi:10.1117/12.2630656, https://doi.org/10.1117/12.2630656 Ge J., et al., 2024a, Chinese Journal of Space Science, 44, 400 Ge J., et al., 2024b, in Coyle L. E., Matsuura S., Perrin M. D., eds, Proc. SPIE Vol. 13092, Space Telescopes and Instrumentation 2024: Optical, In- frared, and Millimeter Wave. SPIE, p. 1309218, doi:10.1117/12.3018669, https://doi.org/10.1117/12.3018669 Gondhalekar Y., Feigelson E. D., Caceres G. A., Montalto M., Saha S., 2023, ApJL, 959, L16 Hippke M., Heller R., 2019, A&A, 623, A39 Hoffman J., 2022, cuvarbase: fast period finding utilities for GPUs, Astro- physics Source Code Library Howard A. W., et al., 2012, ApJS, 201, 15 Howard A. G., Zhu M., Chen B., Kalenichenko D., Wang W., Weyand T., Andreetto M., Adam H., 2017, MobileNets: Efficient Convolutional Neu- ral Networks for Mobile Vision Applications (arXiv:1704.04861), https://arxiv.org/abs/1704.04861 Howell S. B., et al., 2014, PASP, 126, 398 Hu Q., Ge J., Jin L., Willis K., 2026, GTLS: A method for speeding up periodic transit detection using GPU, in preparation Iglesias Γlvarez S., DΓez Alonso E., SΓ‘nchez RodrΓguez M. L., RodrΓguez Ro- drΓguez J., SΓ‘nchez Lasheras F., de Cos Juez F. J., 2023, Axioms, 12, 348 Ioffe S., Szegedy C., 2015, in Bach F., Blei D., eds, Proceedings of Ma- chine Learning Research Vol. 37, Proceedings of the 32nd International Conference on Machine Learning. PMLR, Lille, France, p 448β456, https://proceedings.mlr.press/v37/ioffe15.html Jenkins J. M., et al., 2010a, in Radziwill N. M., Bridger A., eds, SPIE Astro- nomical Telescopes + Instrumentation. San Diego, California, USA, p. 77400D, doi:10.1117/12.856764 Jenkins J. M., et al., 2010b, ApJ, 713, L87 Jenkins J. M., et al., 2012, Proc. IAU, 8, 94 Kingma D. P., Ba J., 2015, in International Conference on Learning Repre- sentations (ICLR). Koch D. G., et al., 2010, ApJ, 713, L79 KovΓ‘cs G., Zucker S., Mazeh T., 2002, A&A, 391, 369 Li J., Tenenbaum P., Twicken J. D., Burke C. J., Jenkins J. M., Quintana E. V., Rowe J. F., Seader S. E., 2019, PASP, 131, 024506 Long J., Shelhamer E., Darrell T., 2015, in Proceedings of the IEEE Con- ference on Computer Vision and Pattern Recognition (CVPR). IEEE, p 3431β3440, doi:10.1109/CVPR.2015.7298965 Loshchilov I., Hutter F., 2019, in International Conference on Learning Rep- resentations (ICLR). Malik A., Moster B. P., Obermeier C., 2021, MNRAS Mandel K., Agol E., 2002, ApJ, 580, L171 MNRAS 000, 1β20 (2026) TransitNet 19 Martinho M. J. S., Valizadegan H., Jenkins J. M., Caldwell D. A., Twicken J. D., Tofflemire B., Jafariyazani M., 2026, ExoMiner++ 2.0: Vetting TESS Full-Frame Image Transit Signals, https://arxiv.org/abs/ 2601.14877 (arXiv:2601.14877) McCauliff S. D., et al., 2015, ApJ, 806, 6 Melton E. J., Feigelson E. D., Montalto M., Caceres G. A., Rosenswie A. W., Abelson C. S., 2024, AJ, 167, 202 Mislis D., Bachelet E., Alsubai K. A., Bramich D. M., Parley N., 2016, MNRAS, 455, 626 Mislis D., Pyrzas S., Alsubai K. A., 2018, MNRAS, 481, 1624 Mullally F., Coughlin J. L., Thompson S. E., Christiansen J., Burke C., Clarke B. D., Haas M. R., 2016, PASP, 128, 074502 Osborn H. P., et al., 2020, A&A, 633, A53 Panahi A., et al., 2022, A&A, 663, A101 PΓ€tzold M., Grziwa S., Hribar R., Schmerling H., 2025, TRANSCEN- DENCE β A TRANSit CaptureENgine for DEtection and Neural net- work Characterization of Exoplanets, EPSCβDPS Joint Meeting 2025, doi:10.5194/epsc-dps2025-1430 Pearson K. A., Palafox L., Griffith C. A., 2017, MNRAS, 474, 478 Pepper J., Kuhn R. B., Siverd R., James D., Stassun K., 2012, PASP, 124, 230 Petigura E. A., Marcy G. W., Howard A. W., 2013, ApJ, 770, 69 Pollacco D., et al., 2006, PASP, 118, 1407 Pratyush P., Gangrade A., 2021, Automation Of Transiting Exoplanet Detec- tion, Identification and Habitability Assessment Using Machine Learning Approaches (arXiv:2112.03298),https://arxiv.org/abs/2112. 03298 Rauer H., et al., 2014, Experimental Astronomy, 38, 249 Rauer H., et al., 2025, Experimental Astronomy, 59 Ricker G. R., et al., 2014, JATIS, 1, 014003 Salinas H., Pichara K., Brahm R., PΓ©rez-Galarce F., Mery D., 2023, MNRAS, 522, 3201 Salinas H., Brahm R., Olmschenk G., Barry R. K., Pichara K., Ishitani Silva S., Araujo V., 2025, MNRAS, 538, 2031 Shallue C. J., Vanderburg A., 2018, AJ, 155, 94 Srivastava N., Hinton G., Krizhevsky A., Sutskever I., Salakhutdinov R., 2014, Journal of Machine Learning Research, 15, 1929 Tan M., Le Q., 2019, in Chaudhuri K., Salakhutdinov R., eds, Proceed- ings of Machine Learning Research Vol. 97, Proceedings of the 36th International Conference on Machine Learning. PMLR, p 6105β6114, https://proceedings.mlr.press/v97/tan19a.html Telesco M., Ge J., Willis K., Dong C., Yang J., Liu B., Jin L., 2026, A scaled analog of the Solar System around a Sun-like star, Submitted to Science Tey E., et al., 2023, AJ, 165, 95 Thomas B., Bhat V. M., Mohammed S. A., Mohammed A. W., Dessalegn A. A., Mittal M., 2025, Identifying Exoplanets with Deep Learning: A CNN and RNN Classifier for Kepler DR25 and Candidate Vetting, https://arxiv.org/abs/2509.04793 (arXiv:2509.04793) Thompson S. E., Mullally F., Coughlin J., Christiansen J. L., Henze C. E., Haas M. R., Burke C. J., 2015, ApJ, 812, 46 Thompson S. E., et al., 2018, ApJS, 235, 38 Valizadegan H., et al., 2022, ApJ, 926, 120 Valizadegan H., Martinho M. J. S., Jenkins J. M., Caldwell D. A., Twicken J. D., Bryson S. T., 2023, AJ, 166, 28 Valizadegan H., et al., 2025, AJ, 170, 287 Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A. N., Kaiser Ε., Polosukhin I., 2017, in Advances in Neural Information Pro- cessing Systems. Vivien H. G., Deleuil M., Jannsen N., De Ridder J., Seynaeve D., Carpine M.-A., Zerah Y., 2025, A&A, 694, A293 Wang K., Ge J., Willis K., Wang K., Zhao Y., 2024a, MNRAS, 528, 4053 Wang K., Ge J., Willis K., Wang K., Zhao Y., Hu Q., 2024b, MNRAS, 534, 1913 Xie D., Wang Y., Liu F., Sun W., 2025, Research in Astronomy and Astro- physics, 25, 104004 Yeh L.-C., Jiang I.-G., 2020, PASP, 133, 014401 Youden W. J., 1950, Cancer, 3, 32 Yu L., et al., 2019, AJ, 158, 25 Zucker S., Giryes R., 2018, AJ, 155, 147 MNRAS 000, 1β20 (2026) 20 Xingchen Yan et al. Lower scoreHigher score Transit Score Transit Score Distribution Before and After PCP Before PCP After PCP BLS TLS TransitNet Figure A1. Kernel-density distributions of transit detection scores at the injected orbital period for BLS, TLS, and TransitNet. All methods show substantially lower scores after PCP, demonstrating effective attenuation of injected periodic transit signals. APPENDIX A: VALIDATION OF PERIODIC CHUNK PERMUTATION To evaluate the effectiveness of PCP in suppressing residual transit signatures, we randomly selected 1000 ATLCs from the Cross-KIC Recovery Set for analysis. BLS, TLS, and TransitNet were applied to both the original ATLCs and their PCP-transformed counterparts, and the resulting detection scores at the true transit periods were compared (Fig. A1). The results show that transit signals at the true periods are substan- tially weakened after applying PCP. This effect is reflected by the clear separation between the score distributions of the original ATLCs and PCP-transformed ATLCs across all three methods. In particular, the PCP-transformed samples are concentrated in the low-score regime, indicating that PCP effectively disrupts the periodic transit structure. Consequently, when the light curves are phase-folded at the true pe- riod, the transit signatures can no longer be coherently aligned and are therefore strongly attenuated. These results support the use of PCP-transformed TMLCs as non-transit samples, as they preserve the underlying noise characteristics of real light curves while sub- stantially mitigating the influence of residual or undiscovered transit signals. APPENDIX B: THRESHOLD SELECTION CRITERIA To obtain a unified operating threshold under realistic transit blind- search conditions, detection-score distributions are constructed from both ATLCs and PCP-transformed TMLCs across multiple target KICs. The ATLCs provide transit samples embedded in realistic stel- lar and instrumental noise environments, while the PCP-transformed TMLCs serve as non-transit controls that preserve the statistical prop- erties of the original observations. Rather than determining a threshold separately for each target, we seek a single operating threshold that generalizes across heteroge- neous stellar environments. LetD denote the set of scored samples produced by a given transit-search algorithm. A candidate threshold setT is constructed from the detection-score distribution, and each threshold ν β T is evaluated independently for every target KIC. The resulting TPRs and FPRs are aggregated using macro-averaging across KICs, from which the corresponding macro-averaged Youden statistic ν½(ν) is computed. The optimal operating threshold ν β is then selected as the threshold yielding the maximum Youden statistic ν½ β (Youden 1950), corresponding to the strongest overall separation between transit and non-transit samples while reducing sensitivity to target-specific noise characteristics. The complete procedure is summarized in Algorithm 1. The equivalence between maximizing the macro-averaged Youden statistic and minimizing the macro-averaged balanced classification error can be shown as follows. Since FNR(ν) = 1βTPR(ν),(B1) we have νΈ(ν) = 1 2 h FPR(ν)+ 1βTPR(ν) i = 1 2 [ 1β ν½(ν) ] .(B2) Therefore, ν½(ν) = 1β 2νΈ(ν),(B3) and hence arg max ν ν½(ν) = arg min ν νΈ(ν). Applying the operating threshold selected via Algorithm 1, the resulting confusion matrices for all transit blind-search algorithms on the full Low-SNR Transit Recovery Set are presented in Fig. B1. Algorithm 1 Macro-Averaged Youden Threshold Selection Require: Scored samplesD =(ν ν , ν¦ ν , ν ν ) ν ν=1 ; candidate thresh- old count ν ν Ensure: Shared operating threshold ν β 1: Construct candidate threshold setT from the score distribution ν ν 2: ν β ββ , ν½ β βββ 3: for each threshold ν β T do 4: for each KIC ν containing both classes do 5:Compute TPR ν (ν) and FPR ν (ν) 6: end for 7: TPR macro (ν) β mean ν TPR ν (ν) 8: FPR macro (ν) β mean ν FPR ν (ν) 9: ν½(ν) β TPR macro (ν)β FPR macro (ν) 10: if ν½(ν) > ν½ β then 11:ν β β ν 12:ν½ β β ν½(ν) 13: end if 14: end for 15: return ν β APPENDIX C: RECOVERED CONFIRMED KEPLER LOW-SNR PLANETS This paper has been typeset from a T E X/L A T E X file prepared by the author. MNRAS 000, 1β20 (2026) TransitNet 21 Non-TransitTransit Predicted Label Non-Transit Transit True Label 10000 139856 Accuracy: 93.0% BLS Non-TransitTransit Predicted Label Non-Transit Transit True Label 98317 53942 Accuracy: 96.5% TLS Non-TransitTransit Predicted Label Non-Transit Transit True Label 10000 22973 Accuracy: 98.9% TransitNet 0 200 400 600 800 1000 0 200 400 600 800 0 200 400 600 800 1000 Figure B1. Overall confusion matrices showing the classification performance of BLS, TLS, and TransitNet on the Low-SNR Transit Recovery Set. Classification is performed using the operating thresholds selected from the macro-averaged threshold optimisation analysis shown in Fig. 11. Among the three methods, TransitNet demonstrates the best overall performance, achieving the highest accuracy (98.9%), compared with TLS (96.5%) and BLS (93.0%). Table C1. Summary of known I2LP low-SNR transit signals recovered by TransitNet. Targets were drawn from the KOI catalog hosted by the NASA Exoplanet Archive (Christiansen et al. 2025), following Thompson et al. (2018). We consider confirmed Kepler planets with 30β©½ νβ©½ 60 d and 6β©½ SNRβ©½ 15. The catalog transit midpoint is defined as ν 0 = ν 0 mod ν, where ν 0 is the KOI transit epoch (koi_time0bk, BKJD). The corresponding TransitNet prediction is denoted by Λν 0 ; the mean absolute timing error is β¨| Λν 0 β ν 0 |β© = 1.24 h (median 0.62 h). The Score column lists the detection-spectrum value at the catalog period, obtained from a search over approximately 3Γ 10 4 trial periods. All recovered targets have scores above the adopted recovery threshold of 0.54. Kepler NamePeriod ν (d)SNRDepth νΏ (ppm)ν 0 (d)Λν 0 (d)| Λν 0 β ν 0 | (h)Score [1] Kepler-1162 c59.28413.1424.430.18630.1560.720.93 [2] Kepler-1178 b31.80610.9141.723.45723.4930.860.97 [3] Kepler-1251 b45.09114.1436.343.37043.3350.841.00 [4] Kepler-1419 b42.52214.7685.68.5628.4871.801.00 [5] Kepler-1440 b39.86013.9155.535.18335.2421.410.87 [6] Kepler-1444 b33.4208.9436.625.45925.4530.141.00 [7] Kepler-1451 b35.6239.6651.134.58434.5830.021.00 [8] Kepler-1453 b47.16113.6723.28.7768.8020.621.00 [9] Kepler-1454 b47.03213.7367.834.37734.3840.171.00 [10] Kepler-1472 b38.13011.6180.46.8666.7442.931.00 [11] Kepler-1610 c44.98512.7617.815.61515.6230.190.96 [12] Kepler-1697 b33.49713.2163.529.77629.7880.291.00 [13] Kepler-1703 c31.8257.779.017.23117.1980.791.00 [14] Kepler-176 e51.16612.2259.218.81219.2189.741.00 [15] Kepler-1760 b38.32615.0392.127.10527.1120.171.00 [16] Kepler-1853 b48.88813.5462.87.1697.1790.241.00 [17] Kepler-1914 b30.82813.9641.319.83419.8660.770.99 [18] Kepler-1916 b31.25414.7438.030.03029.9302.401.00 [19] Kepler-1918 b47.05612.1762.032.79632.8391.031.00 [20] Kepler-1919 b37.88614.4672.219.37319.3730.001.00 [21] Kepler-1920 b30.25412.3576.315.62215.6260.101.00 [22] Kepler-1926 b42.87714.81068.827.43727.4110.620.99 [23] Kepler-1965 b41.86812.5150.721.37221.7579.241.00 [24] Kepler-1980 b33.02612.2507.128.90728.9581.220.95 [25] Kepler-263 c47.33314.0931.59.4179.4120.120.97 [26] Kepler-265 d43.13111.8447.836.86736.8500.410.82 [27] Kepler-276 d48.64812.2641.826.18426.1590.601.00 [28] Kepler-296 e34.14213.3788.033.61033.6130.070.99 [29] Kepler-299 e38.28615.0322.828.44128.5412.400.96 [30] Kepler-324 d34.20611.6183.211.98211.9710.261.00 [31] Kepler-331 d32.13411.81262.013.74013.7640.581.00 [32] Kepler-383 c31.20113.7360.115.26615.3070.981.00 [33] Kepler-395 c34.99011.5498.54.2504.2500.000.95 [34] Kepler-438 b35.23314.5352.829.66629.6460.481.00 MNRAS 000, 1β20 (2026) 22 Xingchen Yan et al. 0.00.20.40.60.81.0 Score 2164169 2857607 3544640 3749365 4138008 4254466 4673628 4847534 5281113 5531694 5688910 5938970 5942949 6206214 6268648 6343170 6587280 6776401 6846911 7098355 7102227 7265298 7287415 7289577 7449541 7595157 7605093 7939330 8025596 8107225 8257205 8379705 8751933 8845205 9021075 9512687 9631995 9661979 9825625 9834040 9845898 9897364 9941066 9956082 10154388 10271806 10397751 10426656 10489345 10843431 11015108 11015323 11187837 11308499 11497977 11507101 11656918 11669239 11874577 12834874 TransitNet - Score Distribution by KIC Figure B2. Per-KIC distribution of TransitNet detection scores on the Cross-KIC Recovery Set generated from 60 unseen KICs used as TMLC sources. ATLC samples (with injected transit signals) and PCP-TMLC samples (without transit signals) are generally well separated in score space, indicating robust discrimination between transit and non-transit light curves. MNRAS 000, 1β20 (2026) TransitNet 23 01020304050 SDE 2164169 2857607 3544640 3749365 4138008 4254466 4673628 4847534 5281113 5531694 5688910 5938970 5942949 6206214 6268648 6343170 6587280 6776401 6846911 7098355 7102227 7265298 7287415 7289577 7449541 7595157 7605093 7939330 8025596 8107225 8257205 8379705 8751933 8845205 9021075 9512687 9631995 9661979 9825625 9834040 9845898 9897364 9941066 9956082 10154388 10271806 10397751 10426656 10489345 10843431 11015108 11015323 11187837 11308499 11497977 11507101 11656918 11669239 11874577 12834874 TLS - Score Distribution by KIC Figure B3. Per-KIC distribution of TLS detection scores on the Cross-KIC Recovery Set generated from 60 unseen KICs used as TMLC sources. Compared with TransitNet, TLS exhibits greater overlap between the score distributions of ATLC and PCP-TMLC samples. These overlapping regions correspond to false positives and missed transit detections, indicating a weaker separation between transit and non-transit light curves. MNRAS 000, 1β20 (2026) 24 Xingchen Yan et al. 46810121416 SDE 2164169 2857607 3544640 3749365 4138008 4254466 4673628 4847534 5281113 5531694 5688910 5938970 5942949 6206214 6268648 6343170 6587280 6776401 6846911 7098355 7102227 7265298 7287415 7289577 7449541 7595157 7605093 7939330 8025596 8107225 8257205 8379705 8751933 8845205 9021075 9512687 9631995 9661979 9825625 9834040 9845898 9897364 9941066 9956082 10154388 10271806 10397751 10426656 10489345 10843431 11015108 11015323 11187837 11308499 11497977 11507101 11656918 11669239 11874577 12834874 BLS - Score Distribution by KIC Figure B4. Per-KIC distribution of BLS detection scores on the Cross-KIC Recovery Set generated from 60 unseen KICs used as TMLC sources. Relative to TLS, BLS produces a broader and more variable false-positive score distribution, resulting in substantially greater overlap between ATLC and PCP-TMLC samples. The enlarged overlap region reflects reduced separability between transit and non-transit signals. MNRAS 000, 1β20 (2026)