Paper deep dive
B-jet Tagging Using a Hybrid Edge Convolution and Transformer Architecture
Diego F. Vasquez Plaza, Vidya Manian
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/26/2026, 2:26:30 AM
Summary
The paper introduces the Edge Convolution Transformer (ECT), a hybrid deep learning architecture for b-jet tagging at the LHC. By integrating EdgeConv blocks for local geometric feature extraction with transformer self-attention for global correlation modeling, the model achieves state-of-the-art performance (0.9333 AUC) on ATLAS simulation datasets, outperforming ParticleNet and pure transformer baselines, particularly in challenging charm-jet rejection tasks.
Entities (5)
Relation Signals (3)
Edge Convolution Transformer โ performs โ b-jet tagging
confidence 100% ยท ECT model for bottom-quark jet tagging
Edge Convolution Transformer โ outperforms โ ParticleNet
confidence 95% ยท ECT achieves 0.9333 AUC for b-jet versus combined charm and light jet discrimination, surpassing ParticleNet (0.8904 AUC)
Edge Convolution Transformer โ trainedon โ ATLAS
confidence 95% ยท The study utilizes the ATLAS simulation dataset
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Jet flavor tagging plays an important role in precise Standard Model measurement enabling the extraction of mass dependence in jet-quark interaction and quark-gluon plasma (QGP) interactions. They also enable inferring the nature of particles produced in high-energy particle collisions that contain heavy quarks. The classification of bottom jets is vital for exploring new Physics scenarios in proton-proton collisions. In this research, we present a hybrid deep learning architecture that integrates edge convolutions with transformer self-attention mechanisms, into one single architecture called the Edge Convolution Transformer (ECT) model for bottom-quark jet tagging. ECT processes track-level features (impact parameters, momentum, and their significances) alongside jet-level observables (vertex information and kinematics) to achieve state-of-the-art performance. The study utilizes the ATLAS simulation dataset. We demonstrate that ECT achieves 0.9333 AUC for b-jet versus combined charm and light jet discrimination, surpassing ParticleNet (0.8904 AUC) and the pure transformer baseline (0.9216 AUC). The model maintains inference latency below 0.060 ms per jet on modern GPUs, meeting the stringent requirements for real-time event selection at the LHC. Our results demonstrate that hybrid architectures combining local and global features offer superior performance for challenging jet classification tasks. The proposed architecture achieves good results in b-jet tagging, particularly excelling in charm jet rejection (the most challenging task), while maintaining competitive light-jet discrimination comparable to pure transformer models.
Tags
Links
- Source: https://arxiv.org/abs/2603.21326v1
- Canonical: https://arxiv.org/abs/2603.21326v1
Trouble viewing inline? Open PDF directly โ
Full Text
54,329 characters extracted from source content.
Expand or collapse full text
Prepared for submission to JINST ํฉ-jet Tagging Using a Hybrid Edge Convolution and Transformer Architecture Diego F. Vasquez Plazaand Vidya Manian Department of Electrical and Computer Engineering, University of Puerto Rico Mayaguez, PR 00681-9000, USA E-mail: diego.vasquez@upr.edu, vidya.manian@upr.edu Abstract: Jet flavor tagging plays an important role in precise Standard Model measurement en- abling the extraction of mass dependence in jet-quark interaction and quark-gluon plasma (QGP) interactions. They also enable inferring the nature of particles produced in high-energy particle collisions that contain heavy quarks. The classification of bottom jets is vital for exploring new Physics scenarios in the โ ํ = 13 TeV proton-proton collisions. In this research, we present a hybrid deep learning architecture that integrates edge convolutions with transformer self-attention mechanisms, into one single architecture called the Edge Convolution Transformer (ECT) model for bottom-quark jet tagging. ECT processes track-level features (impact parameters, momen- tum, and their significances) alongside jet-level observables (vertex information and kinematics) to achieve state-of-the-art performance. The study utilizes the ATLAS simulation dataset [1]. We demonstrate that ECT achieves 0.9333 AUC for ํ-jet versus combined charm and ํํํโํก jet discrim- ination, surpassing ParticleNet (0.8904 AUC) and the pure transformer baseline (0.9216 AUC). The model maintains inference latency below 0.060 ms per jet on modern GPUs, meeting the stringent requirements for real-time event selection at the LHC. Our results demonstrate that hybrid architectures combining local and global features offer superior performance for challenging jet classification tasks. The proposed architecture achieves good results in ํ-jet tagging, particularly excelling in charm jet rejection (the most challenging task), while maintaining competitive ํํํโํก-jet discrimination comparable to pure transformer models. Keywords: Heavy-flavor jets,ํ-jets,ํ-jets,ํํํโํก jets, machine learning, deep learning, particleNet, particle transformer, multihead attention, transformer, jet tagging ArXiv ePrint: 1234.56789 arXiv:2603.21326v1 [hep-ph] 22 Mar 2026 Contents 1 Introduction1 2 Flavor Tagging Literature Review2 3 Bottom-Jet Tagging Methodology4 3.1 Notation4 3.2 Dataset and Preprocessing4 3.2.1 ATLAS Simulation Dataset4 3.3 Feature Engineering5 3.3.1 Track-Level Features5 3.3.2 Jet-Level Features6 3.3.3 Normalization and Preprocessing7 3.3.4 Classification Tasks7 3.3.5 Data Distributions7 3.4 Architectural Overview8 3.5 Training and Optimization9 3.6 Evaluation Metrics9 4 Bottom-Jet Tagging Results and Discussion9 4.1 Comparative Analysis13 4.1.1 ECT vs. ParticleNet15 4.1.2 ECT vs. ParT16 4.1.3 Key Insights17 5 Conclusion19 1 Introduction Collimated showers of hadrons and leptons are called jets which are the dominant signatures of high energy quark and gluon production at the Large Hadron Collider (LHC). Identifying the flavor of the parton that initiated a jet (bottom, charm, or light quarks; gluons) is essential for precision Standard Model (SM) measurements and searches for physics beyond the SM. Heavy-flavor jets, particularly those originating from bottom-quarks (ํ-jets), play a central role in measurements of the Higgs boson (ํป โ ํ ฬ ํ), top quark properties, and searches for supersymmetry. The key discriminant for heavy-flavor tagging is the displaced secondary vertex arising from the finite proper lifetime of ํ- and ํ-hadrons (ํํ ํ โ 460 ํm, ํํ ํ โ 150 ํm). Tracks originating from these decays exhibit large impact parameters relative to the primary interaction point. While distinguishing ํ-jets from ํํํโํก-flavor jets (u, d, s quarks, gluons) is relatively straightforward due โ 1 โ to the absence of secondary vertices. However, separating ํ-jets from ํ-jets remains challenging due to their similar decay topologies. This ํ vs ํ discrimination is the limiting factor in many analyses, motivating continuous improvements in tagging algorithms. Machine learning has transformed jet flavor tagging over the past decade, progressing from boosted decision trees operating on hand-crafted high-level features to deep neural networks pro- cessing low-level track and vertex information [2]. Recent state-of-the-art approaches employ either graph neural networks (e.g., ParticleNet [3]) that aggregate information from spatially nearby particles, or transformer architectures [4] that capture global correlations via self-attention. Each paradigm has complementary strengths: graph convolutions excel at modeling local vertex topol- ogy, while transformers capture jet-wide patterns. In this work, we present the Edge Convolution Transformer (ECT), a hybrid architecture that integrates both mechanisms within a unified model. ECT employs EdgeConv blocks [3] to extract local geometric features from ํพ-Nearest Neighbor (ํพํ) graphs in (ํ, ํ) space, followed by transformer self-attention layers to capture global par- ticle correlations. A learned class token aggregates particle-level information, which is fused with jet-level vertex statistics for final classification. We demonstrate that this hybrid approach achieves superior performance compared to pure graph-based (ParticleNet) and pure attention-based (Parti- cle Transformer) baselines, particularly for the challenging ํ vsํ separation task, while maintaining inference latency suitable for real-time LHC trigger systems. The main contributions of this research are: โข a novel hybrid architecture (ECT) combining EdgeConv and transformer attention for jet flavor tagging; โข comprehensive evaluation on ATLAS simulation data across three binary classification tasks (ํ vs ํ, ํ vs ํํํโํก, ํ vs ํ+ํํํโํก jets); โข performance comparison against ParticleNet and Particle Transformer baselines; โข inference latency analysis confirming deployment feasibility for high-level trigger systems; โข demonstration that EdgeConv blocks are essential for heavy-flavor separation, while trans- former attention excels at capturing jet-wide patterns. The remainder of this paper is organized as follows: Section 2 reviews related work in jet flavor tagging and transformer architectures. Section 3 describes the dataset, model architecture, and training methodology. Section 4 presents experimental results and comparative analysis. Section 5 concludes with limitations and future directions. 2 Flavor Tagging Literature Review ML methods for jet flavor tagging make use of track properties and reconstructed secondary vertices as additional information both in ATLAS and CMS experiments. A Scodellaro proceedings review summarizes Run 2 developments in heavy-flavor tagging such as track-based, vertex, soft-lepton, and boosted topology taggers and outlines preparations for the high-luminosity LHC era [5]. The study in [6] presents the ATLAS Run 2 ํ-tagging algorithms, combining low level track and secondary vertex taggers into multivariate classifiers. It reports ํ-tagging efficiencies using 13 TeV ํ collision data corresponding to 80.5 fb โ1 , showing good agreement between data and simulation across a wide jet ํ ํ range. For real-time inference, the ATLAS collaboration developed a fast neural โ 2 โ network based ํ-tagger, fastDIPS, for deployment in the high-level trigger during LHC Run 3 [7]. This tagger enables early rejection of ํํํโํก jets, reducing input rates for hadronic ํ-jet events by a factor of five with only a โผ2% drop in signal efficiency. Another significant innovation is the Recurrent Neural Network (RNN)-based ํ-tagging algorithm (ATL-PHYS-PUB-2017-003) that sequences charged particle tracks in jets to exploit trackโtrack correlations and improve ํ-jet identification without relying on secondary vertex reconstruction [8]. In the case of CMS experiments, the study in [9] reviews the heavy-flavor jet tagging ad- vancements including the UnifiedParticleTransformer (UParT) architecture and presents validated performance comparisons and scale factors derived from 13.6 TeV collision data collected during 2022โ2023. ParticleNet has been used for ํ-jet tagging in CMS High level trigger in 2022 and 2023 [10]. Another CMS study [11] presents Run 2 heavy-flavor jet identification methods including track-based, secondary vertex, soft lepton taggers, boosted โdouble-bโ taggers, and c-jet tagging, achieving 68% ํ-jet efficiency at 1% ํํํโํก-jet misidentification and reducing uncertainty to a few percent across 30โ1000 GeV jet ํ ํ . The review in [12] provides a comparative summary of state-of-the-art jet substructure and flavor-tagging methods used by ATLAS and CMS including attention-based transformers, adversarial training, and advanced calibration techniques highlighting their strengths and limitations in LHC Run 3 analyses. The review of ML methods for heavy-flavor jet tagging by Mondal and Mastrolorenzo sys- tematically categorises heavy-flavor jet tagging methods at the LHC into three generations ranging from Boosted Decision Trees (BDTs) and dense neural networks to graph and transformer-based architectures highlighting their evolution, detector input representations, and calibration strategies across Runs 1โ3 [2]. These methods have seen an evolution for Run 1, Run 2, and Run 3 CMS and ATLAS experiments at the LHC. JetVLAD, a set-based heavy-flavor jet tagger built on the NetVLAD architecture that aggregates particle-level descriptors into fixed-length vectors for classi- fication is proposed in [13]. It demonstrates strong performance on simulated heavy-ion jet samples, achieving heavy-flavor identification efficiencies above 80% with background rejection factors up to several hundred. The authors in [14] demonstrate that adversarial training significantly enhances the robustness of jet-flavor tagging neural networks, preserving high classification accuracy while reducing vulnerability to input perturbations by probing and smoothing the loss surface geome- try. Deep neural networks trained on raw track and vertex level information perform similarly to traditional ํ taggers by directly leveraging high-dimensional detector data, showing that adding low-level features significantly improves jet-flavor classification [15]. For future colliders, the study in [16] presents a fast simulation framework based on Delphes with modules for tracking, time-of-flight, and cluster counting, and evaluates jet flavor tagging using the ParticleNet graph neural network tagger. The DeepJet model introduced in [17] lever- ages deep neural networks including separate branches for charged, neutral, and secondary vertex inputs to process all jet constituents without pre-selection, significantly improving heavy-flavor and quark/gluon tagging performance over previous approaches. More recently, in [18] the authors introduce Retentive Networks (RetNet) for efficient ํ-jet identification in simulated 13 TeV ํ col- lision data, achieving competitive performance with only 330k parameters and offering an effective alternative to models like DeepJet and Particle Transformer (ParT). Transformer-based models are increasingly popular in ํ-jet tagging. In [19], the authors present the DeepJetTransformer with scaled-dot product attention and heavy-flavor transformer โ 3 โ block developed for the Future Circular Collider (FCC) experiment simulations. In [20], the authors review state-of-the-art attention-based transformer architectures for heavy-flavor jet tagging, demon- strating improved classification performance and model interpretability through physics-informed network modifications and analysis of the decision-making process. The ATLAS Collaboration introduces GN2, a groundbreaking transformer-based jet flavor tagging algorithm that leverages end-to-end low-level track information and physics-informed auxiliary objectives: replacing tradi- tional vertex-based taggers and delivering validated performance improvements in both simulation and โ ํ = 13.6 TeV Run 3 data analyses [21]. 3 Bottom-Jet Tagging Methodology This section describes the dataset, followed by data preprocessing and jet feature construction, com- putation of per-particle feature embedding, edge convolution operation, transformer self-attention, aggregation, followed by classification using a Feed-Forward Network (FFN). Further optimization is applied to training loops and evaluation metrics are presented. 3.1 Notation Throughout this section, we use the following notation: โข ํต: Batch size โข ํ: Maximum number of tracks per jet (40) โข ํ: Embedding dimension (128) โข ํพํ: Number of nearest neighbors for EdgeConv (16) โข โ: Number of attention heads (8) โข HโR ํตรํรํ : Particle embeddings โข Mโ 0, 1 ํตรํ : Track validity mask 3.2 Dataset and Preprocessing 3.2.1 ATLAS Simulation Dataset We utilize the publicly available ATLAS simulation dataset (Zenodo record 4044628) [1] consisting of jets sampled from ํ โ ํก ฬ ํก events at โ ํ = 14 TeV. Events were generated using Pythia8 [22] and processed through Delphes [23] fast simulation framework configured to emulate the ATLAS detector [24]. This dataset was used for secondary vertex finding in [25] using the set-to-graph model. The dataset contains three jet flavor classes: โข ํ-jets: Originating from bottom-quarks (ํํ ํ โ 460 ํm) โข ํ-jets: Originating from charm quarks (ํํ ํ โ 150 ํm) โข ํํํโํก jets: Originating from up, down, strange quarks and gluons The data is partitioned into training (62.2%), validation (18.9%), and test (18.9%) sets, as detailed in Tables 1 and 2. โ 4 โ 3.3 Feature Engineering This section presents the track-level and jet-level features and the preprocessing methods before training the models. 3.3.1 Track-Level Features Each jet contains a variable number of charged particle tracks (up to ํ max = 40), zero-padded for batch processing. Seven features per track are extracted, chosen for their discriminative power in heavy-flavor identification: 1. ํ ํ : Transverse momentum (log-transformed and z-normalized) 2. ํ 0 : Transverse impact parameter 3. ํง 0 : Longitudinal impact parameter 4. ํ 0 /ํ ํ 0 : Transverse impact parameter significance 5. ํง 0 /ํ ํง 0 : Longitudinal impact parameter significance 6. IP 3ํท : 3D impact parameter 7. IP 3ํท /ํ IP 3ํท : 3D impact parameter significance The impact parameters characterize the spatial displacement of charged particle tracks relative to the primary interaction vertex (PV), exploiting the finite lifetime of heavy-flavor hadrons (ํํ ํ โ 450 ํm, ํํ ํ โ 150 ํm) to distinguish their decay products from prompt particles [26, 27]. Following the ATLAS perigee track parameterization [28]: โข Transverse impact parameter (ํ 0 ): Defined as the distance of closest approach of the track helix to the primary vertex in the ํ-ํ plane (transverse to the beam axis). Mathematically, ํ 0 represents the signed perpendicular distance from the PV to the track trajectory at the point of closest approach. Typical values range fromO(10 ํm) for prompt tracks toO(1 m) for displaced tracks from ํ-hadron decays. The transverse resolution achieves ํ ํ 0 โ 10โ20 ํm for high-ํ ํ tracks in the ATLAS Inner Detector. โข Longitudinal impact parameter (ํง 0 ): Defined as the ํง-coordinate difference between the primary vertex and the point of closest approach along the beam axis direction. In ATLAS analyses, the quantity ํง 0 sinํ is often used to account for the track polar angle ํ, providing a more uniform resolution across pseudorapidity. The longitudinal component is particularly sensitive to pile-up contamination from nearby ํ interactions. โข 3D impact parameter (IP 3D ): Constructed as the quadrature combination of both compo- nents: IP 3D = โ๏ธ ํ 2 0 + ํง 2 0 (3.1) While IP 3D is indeed composed of ํ 0 and ํง 0 , and these quantities are correlated, they provide complementary discriminative information. The transverse component (ํ 0 ) is more robust โ 5 โ against pile-up effects, while the longitudinal component (ํง 0 ) captures additional vertex displacement information. The ATLAS IP3D algorithm exploits both parameters using two-dimensional probability density functions (PDFs) that explicitly model their correlation structure [27], achieving superior performance compared to using either component alone. The significance of each impact parameter is defined as the ratio of the measured value to its uncertainty: ํ 0 /ํ ํ 0 , ํง 0 /ํ ํง 0 , and IP 3D /ํ IP 3D , where ํ denotes the track-by-track measurement uncertainty propagated from the covariance matrix of the track fit. The 3D significance is computed as: IP 3D ํ IP 3D = โ๏ธ ํ 2 0 + ํง 2 0 โ๏ธ ํ 2 ํ 0 + ํ 2 ํง 0 (3.2) Tracks fromํ-hadron decays typically exhibit significance values|ํ 0 /ํ ํ 0 | > 3 and|IP 3D /ํ IP 3D | > 5, while prompt tracks from the primary vertex cluster near zero with widths determined by detec- tor resolution. In our dataset, ํ-jets show mean |ํ 0 /ํ ํ 0 | values approximately 2โ3 times larger than ํํํโํก jets, with pronounced high-significance tails extending beyond 10ํ (Figure 2). These discriminative distributions motivate the inclusion of all three impact parameter types and their sig- nificances as independent features, despite their mathematical correlation, as each captures distinct aspects of the track displacement topology relevant for heavy-flavor identification. Impact parameters and their significances are critical for distinguishing decay vertices: heavy- flavor jets exhibit displaced secondary vertices due to non-zero proper lifetimes of ํ- and ํ-hadrons, resulting in tracks with larger impact parameters compared to prompt tracks from ํํํโํก jets. 3.3.2 Jet-Level Features Eight global jet observables complement track-level information: 1. ํ jet ํ : Jet transverse momentum (log-transformed) 2. ํ jet : Jet pseudorapidity 3. ํ jet : Jet azimuthal angle 4. ํ jet : Jet mass (log-transformed) 5. ํ trk : Number of associated tracks 6. ํ vtx : Number of reconstructed secondary vertices 7. ํฟ max 3ํท : Maximum 3D displacement of secondary vertices 8. ํ vtx,max trk : Maximum tracks per secondary vertex The vertex-related features (ํ vtx , ํฟ max 3ํท , ํ vtx,max trk ) encode crucial information about secondary vertex topology, particularly discriminative for ํ vs ํ separation given the different decay lengths. โ 6 โ 3.3.3 Normalization and Preprocessing The following are the normalization and preprocessing steps applied to the features. โข Track features: Z-score normalization per feature across training set โข Jet ํ ํ and mass: Log-transformation followed by z-normalization โข Vertex features: Min-max scaling to [0, 1] range โข Padding: Jets with < ํ max tracks are zero-padded โข Masking: Boolean mask Mโ 0, 1 ํ max indicates valid tracks (1) vs padding (0) Table 1. Total jet feature count DataJet featureTotal quantity Train 62.23%ํ jet ํ , ํ jet , ํ jet , ํ jet , ํ trk , ํ vtx , 1,071,940 Validation 18.88%ํฟ max 3ํท , ํ vtx,max trk 325,289 Test 18.88%325,290 Table 2. Total track feature count DataTrack featureTotal quantity Train7,624,594 Validation ํ trk ํ , ํ0, ํง0, ํ 0 /ํ ํ 0 , ํง 0 /ํ ํง 0 , ํํ3ํท, IP 3ํท /ํ IP 3ํท 2,312,445 Test2,311,995 3.3.4 Classification Tasks We evaluate the ECT model on three binary classification tasks with increasing difficulty: โข ํ vs ํํํโํก: Baseline task, exploiting large topological differences โข ํ vs ํ+ํํํโํก: Realistic scenario for LHC analyses, combined background โข ํ vs ํ: Most challenging, distinguishing similar heavy-flavor jets 3.3.5 Data Distributions Figures 1 and 2 present the distributions of jet-level and track-level features across the three jet flavor classes. Several discriminative patterns are evident in Figure 1: ํ-jets exhibit broader transverse momentum (ํ ํ ) distributions extending to higher values, reflecting the massive ํ-quark (mass โ4.2 GeV); vertex multiplicity (ํ vtx ) is significantly higher for ํ-jets compared to charm and ํํํโํก jets, consistent with the longer ํ-hadron lifetime enabling multiple displaced vertices to be reconstructed; jet invariant mass distributions show clear separation, with ํ-jets having systematically larger masses due to the heavier quark content. Figure 2 reveals the discriminative power of track-level features, particularly the impact param- eter significance (ํ 0 /ํ ํ 0 , ํง 0 /ํ ํง 0 , IP 3ํท /ํ IP 3ํท ). Heavy-flavor jets (ํ and ํ) exhibit pronounced tails โ 7 โ extending to large significance values (>10ํ), corresponding to tracks originating from displaced secondary vertices. In contrast, ํํํโํก jets show distributions tightly peaked near zero, consistent with prompt tracks pointing to the primary interaction vertex. The broader tails for ํ-jets com- pared to ํ-jets reflect the longer decay length (ํํ ํ /ํํ ํ โ 3), though substantial overlap makes pure cut-based separation challenging. These distributions motivate our choice of input features and demonstrate the rich discriminative information available at both track and jet levels for flavor classification. 3.4 Architectural Overview The ECT architecture processes jets through six sequential stages (Figure 3): 1. Input Processing: In figure 3, each jet contains up to ํ= 40 zero-padded tracks with 7 features per track (momentum and impact parameters), 8 jet-level global features (kinematics and vertex statistics), spatial coordinates (ํ, ํ) for ํพํ graph construction, and a boolean mask M distinguishing valid tracks from padding. 2. Feature Embedding: Track features are embedded from 7 to 128 dimensions via a three- layer Multi-Layer Perceptron (MLP) [7 โ 128 โ 512 โ 128] with ReLU activations. Jet-level features are processed through a separate MLP [8 โ 128 โ 128] with dropout (ํ= 0.15) for regularization. 3. Local Feature Extraction: Three EdgeConv blocks [3] aggregate information from ํพํ= 16 in(ํ, ํ) space, capturing local jet substructure. The embedding dimension evolves through the blocks as 128โ 64โ 128โ 256, with residual connections and dimension-matching projections. A final linear layer projects back to ํ= 128. 4. Global Interaction (Transformer): Four self-attention layers with 8 heads each (ํ ํพ ํ= 16 per head) capture long-range dependencies between particles. Each layer applies pre- LayerNorm, multi-head attention with scaled dot-product, and a position-wise FFN with hidden dimension 512 (4ร embed dim). The track mask M ensures padded positions con- tribute zero attention weight. 5. Jet-Level Aggregation: A learned class token attends to all particle representations via two class-attention layers, producing a producing a permutation-invariant jet-level embedding z cls โR 128 . This mechanism allows the model to learn which particles are most discriminative for flavor classification. 6. Classification: The class token representation z cls is fused with the jet-level feature embed- ding g jet via element-wise addition (โ), combining learned particle patterns with explicit vertex/kinematic information. A final linear classifier [128 โ 2] with softmax activation produces class probabilities. Notation: Batch size ํต, maximum tracks per jet ํ= 40, embedding dimension ํ= 128. Total trainable parameters: 1,699,330 (โ1.7M). The complete architecture contains 1.7M trainable parameters and is optimized end-to-end using the Adam optimizer with cross-entropy loss. Automatic Mixed Precision (AMP) training enables efficient utilization of GPU tensor cores. โ 8 โ Design Rationale: The hybrid design combines the complementary strengths of EdgeConv and transformers. EdgeConv captures local geometric relationships critical for vertex identification (e.g., clusters of displaced tracks), while transformer attention captures global patterns such as the overall track multiplicity and momentum flow characteristic of heavy-flavor decays. This combination proves particularly effective for the challenging ํ vs ํ discrimination task (Section 4). 3.5 Training and Optimization The ECT model is trained using the Adam optimizer [29] with learning rate of ํ 0 = 5ร 10 โ4 . We employ a constant learning rate schedule without warmup or decay, as preliminary experiments showed that simple schedules were sufficient for convergence. AMP training via torch.cuda.amp enables efficient utilization of GPU tensor cores, reducing training time by approximately 40% with no loss in final model accuracy. A batch size of 1024 is used for the ํ vs ํ+ํํํโํก task due to its balanced class distribution, while batch size of 512 is employed for ํ vs ํ and ํ vs ํํํโํก to maintain stable training dynamics with more skewed class ratios. All models are trained for a maximum of 100 epochs with early stopping based on Area Under Curve (AUC). The hyperparameters for the ECT models were optimized empirically to balance training efficiency and performance. Table 3 outlines the key hyperparameters used in the experiments for ECT along with those for ParticleNet and ParT models for comparison. The ECT architecture employs an embedding dimension of 128 with 4 multi-head attention heads. Four particle-level attention blocks and two class-level attention blocks process track and jet features. EdgeConv layers use ํพํ= 16, with progressively deeper MLPs of sizes (64,64,64), (128, 128, 128), and (256, 256, 256). ReLU activations are used in the MLP embedding. The workflow for the ECT architecture is given in Algorithm 1. 3.6 Evaluation Metrics The evaluation metrics on the validation and testing set are classification accuracy, F1-score, Receiver Operating Characteristic (ROC), and AUC. If the AUC does not improve for 25 epochs, the training is stopped early, and the best model is restored. This model is used for the evaluation of the test set. The test set is evaluated by loading the best model checkpoint and computing the metrics of: accuracy, F1-score, and AUC. The ROC and confusion matrices are printed to visualize the performance of the ECT model. 4 Bottom-Jet Tagging Results and Discussion The ECT model alternates local feature aggregation (via edge convolutions) and global context integration (via Transformer attention + class token). The inclusion of distance-based bias in attention and the use of physics-motivated input features (e.g. impact parameters) ensures that the model respects underlying jet substructure physics, consistent with recent advancements in ParT and ParticleNet methodologies. The ECT is trained on both the seven track and eight jet features totaling fifteen features. Figure 4 shows the metrics calculated over the epochs for the ECT model: loss, accuracy, AUC, and F1-score. Table 4 shows the training loss, Accuracy, AUC, and F1-score for ํ-jet tagging using the three models. The ParT architecture gives best performance for b versus ํํํโํก jets. Overall, the ECT architecture outperforms the ParticleNet, and ParT architectures for ํ โ 9 โ 3210123 jet 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Probability density b-jets (training) c-jets (training) light jets (training) b-jets (validation) c-jets (validation) light jets (validation) 01020304050 M jet [GeV] 10 7 10 6 10 5 10 4 10 3 10 2 10 1 Probability density b-jets (training) c-jets (training) light jets (training) b-jets (validation) c-jets (validation) light jets (validation) 3210123 jet [rad] 0.000 0.025 0.050 0.075 0.100 0.125 0.150 0.175 Probability density b-jets (training) c-jets (training) light jets (training) b-jets (validation) c-jets (validation) light jets (validation) 050100150200 p jet T [GeV] 10 7 10 6 10 5 10 4 10 3 10 2 Probability density b-jets (training) c-jets (training) light jets (training) b-jets (validation) c-jets (validation) light jets (validation) b c light Jet flavor 10 5 2 ร 10 5 3 ร 10 5 Number of jets b-jets (training) c-jets (training) light jets (training) b-jets (validation) c-jets (validation) light jets (validation) Figure 1. Distributions of jet-level features for ํ-jets (red), ํ-jets (green), and light jets (blue) in the ATLAS simulation dataset. Top left: Jet pseudorapidity (ํ jet ). Top right: Jet invariant mass (ํ jet ). Middle left: Jet azimuthal angle (ํ jet ). Middle right: Jet transverse momentum (ํ jet T ). Bottom: Jet flavor distribution. Solid lines represent the training set (62.2% of data), while dashed lines show the validation set (18.9%). All distributions are normalized to unit area for comparison. The broader ํ T distribution of ํ-jets reflects the higher mass of bottom quarks (ํ ํ โ 4.2 GeV/ํ 2 ), while vertex multiplicity differences are evident in the flavor distribution. โ 10 โ 0255075100125150175200 p trk T [GeV] 10 7 10 6 10 5 10 4 10 3 10 2 10 1 Probability density b-jets (training) c-jets (training) light jets (training) b-jets (validation) c-jets (validation) light jets (validation) 10+1 Track charge [e] 4.98 ร 10 1 4.99 ร 10 1 5 ร 10 1 5.01 ร 10 1 5.02 ร 10 1 5.03 ร 10 1 Probability density b-jets (training) c-jets (training) light jets (training) b-jets (validation) c-jets (validation) light jets (validation) 42024 d 0 [m] 10 7 10 6 10 5 10 4 10 3 10 2 10 1 10 0 Probability density b-jets (training) c-jets (training) light jets (training) b-jets (validation) c-jets (validation) light jets (validation) 20015010050050100150200 z 0 [m] 10 7 10 6 10 5 10 4 10 3 10 2 10 1 Probability density b-jets (training) c-jets (training) light jets (training) b-jets (validation) c-jets (validation) light jets (validation) /20/2 trk [rad] 0.14 0.15 0.16 0.17 0.18 0.19 0.20 Probability density b-jets (training) c-jets (training) light jets (training) b-jets (validation) c-jets (validation) light jets (validation) 21012 trk 0.16 0.18 0.20 0.22 0.24 Probability density b-jets (training) c-jets (training) light jets (training) b-jets (validation) c-jets (validation) light jets (validation) Figure 2. Distributions of track-level features for charged particles associated with ํ-jets (red), ํ-jets (green), and light jets (blue). Top row: Track transverse momentum (ํ trk T ) showing ํ T > 1 GeV/ํ selection, and track electric charge distribution. Second row: Transverse (ํ 0 ) and longitudinal (ํง 0 ) impact parameters, demonstrating displaced vertices characteristic of heavy-flavor decays. The pronounced tails for ํ- and ํ-jets arise from secondary vertex displacements (ํํ ํ โ 460 ํm, ํํ ํ โ 150 ํm), while light jets peak sharply near zero, consistent with prompt tracks from the primary vertex. Bottom row: Track pseudorapidity (ํ trk ) and azimuthal angle (ํ trk ) showing angular distributions. where smaller uncertainties enable higher-significance discrimination of displaced vertices. Solid lines represent training data (7,624,594 tracks), dashed lines show validation data (2,312,445 tracks). All distributions are normalized to probability density. Heavy-flavor jets exhibit significantly broader impact parameter distributions, providing the primary discriminative power for ํ-jet identification โ 11 โ Padding + feature selection Tracks = [B, P, 7] Jets = [B, 8] Coordinates (ฮท, ฮฆ) = [B, P, 2] Flavors = [B] ROOT files (Tracks & Jets) Input: 7 track features 8 jet features (train, validation, test) TM = MLP Track C = [B, P, 2] MLP Jets [B, P] = [B, P, 7] (ฮท, ฮฆ) = [B, 8] โ [B, P, 128] โ [B, 128] Final FFN + Softmax / Sigmoid โCoor = [B, P, 2] โMLP Track [B, P, 128] โTM = [B, P] ECB = 1 [B, P, 64] 2 [B, P, 128] 3 [B, P, 256] โ[B, P, 128] T = [B, P, 128] โ (4ร self-attn) โ [B, P, 128] โ C = [B, P, 2] (ฮท, ฮฆ) โ ECB = [B, P, 128] โTM = [B, P] โ T = [B, P, 128] CT/GP = [B, P, 128] โ [B, 128] โ FFN = [B, 2] CL = [B] Final prediction 2 = number of classes Class label Dimensions of data in each stage: [ (B) = Batch, (P) = max_tracks), ( ) = features per track ] Flavors = (values โ b, c, light) Track Mask (TM) = [B, P] โ 0, 1 (1=real track, 0=padding) Transformer blocks (T) - Intermediate dimension: (4 self-att blocks with internal FFN 128โ512โ128) Coordinates Track mask Embedding MLP Tracks Edge Conv Blocks Transformer blocks Class Token / Global Poolling Embeding MLP Jets โ CT/GP = [B, 128] + MLP Jets = [B, 128] Element-wise sum = [B, 128] โ [B, 128] = [B, 128] FFN = [B, 128] โ [B, 2] Figure 3. Architecture of the Edge Convolution Transformer (ECT) for ํ-jet tagging, showing complete information flow and tensor dimensions at each stage. โ 12 โ Table 3. Model hyperparameters and architectural configuration. ECT combines EdgeConv blocks (from ParticleNet) with transformer self-attention and class-attention mechanisms. All models trained on NVIDIA RTX A5000 GPU. ParameterECTParticleNetParT Architecture Embedding dimension128128128 Attention heads8โ8 Self-attention blocks4โ8 Class-attention blocks2โ2 EdgeConv blocks33โ EdgeConv ํพํ1616โ EdgeConv channels(64,64,64)(64,64,64)โ (128,128,128)(128,128,128)โ (256,256,256)(256,256,256)โ Global MLP hidden dim128โ128 Global dropout0.15โ0.15 Total parameters1.7M1.7M1.7M Training Configuration Learning rate5ร 10 โ4 1ร 10 โ3 1ร 10 โ3 Batch size1024 / 512 โ 10241024 OptimizerAdamAdamRanger Max epochs100100100 Early stopping patience252525 LR schedulerโCosineAnnealing Activation (embed MLP)ReLUReLUGELU Mixed precision (AMP)YesYesYes Computational GPUNVIDIA RTX A5000 (24GB) FrameworkPyTorch 2.0 + CUDA 11.8 Training time (ํ vs. ํ+ํํํโํก)2.2 hours4.5 hours1.5 hours โ 1024 for ํ vs. ํ+ํํํโํก, 512 for ํ vs. ํ and ํ vs. ํํํโํก tasks vs. ํ, ํ vs. ํํํโํก and ํ vs. ํ and ํํํโํก jet tagging. The plot of efficiency versus misidentification rates for the ECT, ParticleNet, and ParT are given in Figures 5, 6, and 7 respectively. 4.1 Comparative Analysis This section compares the performance of ECT with ParticleNet and ParT for ํ-jet tagging. All models are trained with the same number of samples for each of the cases. The ParT was originally implemented for jet tagging of jet class datasets [30]. As shown in Table 3, ECT has 4 particle attention and 2 class attention blocks, while ParT has 8 class attention and 2 class attention blocks. While the ParticleNet does not have these attention blocks, the ECT has its edge convolution blocks incorporated in its architecture. Table 4 shows that the ParticleNet takes more time for execution compared to the ECT models and the ParT model. Table 4 and Figures 6 and 7 present a comprehensive comparison of ECT against ParticleNet and Particle Transformer (ParT) across three binary classification tasks. The results reveal distinct โ 13 โ Algorithm 1: ECT: Jet Tagging Pipeline Part 1: Data Preprocessing Input: Raw ROOT files with jet and track data foreach jet event in dataset do Extract track featurestrk_pt, trk_d0, trk_z0, trk_d0sig, trk_z0sig, ip3D, ip3D_signal Extract track coordinates(trk_eta, trk_phi)(for ํพํ graph construction) Extract jet features jet_pt, jet_eta, jet_phi, jet_M, n_trks, n_vertex, vertex_L3D, vertex_ntracks Pad tracks to fixed length ํ max = 40, construct mask Assign jet label according to classification mode (e.g., ํ vs ํ) Part 2: Training the Model Require: Preprocessed datasetD, model ํ ํ , loss functionL, optimizer foreach training epoch do foreach mini-batch(X, C, M, y) fromD do Embed track features via MLP: Hโ MLP(X) foreach EdgeConv block do Update Hโ EdgeConv(H, C, M) foreach Transformer block do Hโ SelfAttn(H, C, M) Aggregate with class token or pooling: zโ Aggregate(H) Predict logits: ห yโ FFN(z) Compute loss: โ โL( ห y, y) Update ํ using optimizer and gradient of โ Validate on held-out data; apply early stopping if needed Part 3: Evaluation and Inference Require: Trained model ํ ํ foreach test jet event do Preprocess features and coordinates as in Part 1 Predict jet class: หํฆ= arg max ํ ํ (X, C, M) Collect predictions for metrics (accuracy, AUC, F1, etc.) performance characteristics for each architecture, highlighting the complementary strengths of edge convolutions and transformer attention mechanisms. Figure 5 illustrates the baseline performance of the ECT model. As expected, distinguishing ํ-jets from ํ-jets is the most challenging task, reflected in the higher misidentification rates across the efficiency range. ECT model gives the best results for ํํํโํก-jet identification and the signal efficiency is higher when it approaches unity. ECT results are used as a reference point for subsequent comparisons. โ 14 โ Figure 4. Training (Train) and validation (Val) metrics for the ECT model on the ํ vs ํ+ํํํโํก classification task over 100 epochs. Top left: Cross-entropy loss converges smoothly from 0.37 to 0.31 over the first 50 epochs, then stabilizes, indicating effective optimization without oscillations. Top right: Training (blue) and validation (orange) accuracy (acc) reachโผ87% with minimal gap (< 0.2%), demonstrating that the model generalizes well without overfitting. Bottom left: Area Under Curve (AUC) improves steadily from 0.92 to 0.93, plateauing around epoch 89 where early stopping was triggered (patience = 25). The small train-validation AUC gap confirms robust generalization. Bottom right: F1-score exhibits higher variance during early training due to threshold sensitivity, then stabilizes at 0.81, confirming balanced precision-recall trade-off. Dashed horizontal lines mark the maximum values achieved: training accuracy 87.2%, validation accuracy 87.0%, training AUC 0.930, validation AUC 0.928. These curves demonstrate stable convergence and effective regularization. 4.1.1 ECT vs. ParticleNet ECT demonstrates consistent and substantial improvements over ParticleNet across all classification tasks: โข ํ vs. ํ (charm rejection): ECT achieves an AUC of 0.8853 compared to ParticleNetโs 0.8023, representing a remarkable +8.3% improvement. As shown in Figure 6 (orange curves), ECT maintains significantly lower misidentification rates across the entire efficiency range. At the medium working point. (1% misidentification rate), ECT reaches 65% signal โ 15 โ Table 4. Performance comparison on ATLAS simulation dataset. Accuracy, AUC, and F1-score evaluated on held-out test set (325,290 jets). Latency measured on NVIDIA RTX A5000 GPU with batch size 1024, averaged over 100 iterations after 10 warmup runs using CUDA events. ModelTasklossAccAUCF1-scoreInferenceTotal time (ms)time (h:m:s) ECTํ vs. ํ0.430381.720.88530.81830.0601:17:53 ํ vs. ํํํโํก0.119596.050.98830.95990.0571:43:01 ํ vs. ํ+ํํํโํก0.308087.750.93330.81460.0572:09:04 ParticleNetํ vs. ํ0.540273.470.80230.754512.2314:27:37 ํ vs. ํํํโํก0.272889.080.94510.897317.8824:29:31 ํ vs. ํ+ํํํโํก0.378984.390.89040.887213.0434:41:36 ParTํ vs. ํ0.467078.850.86340.77780.1461:20:39 ํ vs. ํํํโํก0.124395.940.98760.95890.1601:20:52 ํ vs. ํ+ํํํโํก0.332585.920.92160.78140.2221:30:57 โก Inference time measured on NVIDIA RTX A5000 GPU (24GB VRAM) with batch size 1024, averaged over 100 iterations after 10 warmup runs using CUDA events for microsecond-precision timing. Total training time includes data loading, forward/backward passes, and checkpointing. All models comfortably exceed LHC High-Level Trigger requirements (< 1 ms/jet) by factors of 15-20ร, confirming deployment feasibility for real-time event selection. efficiency compared to ParticleNetโs 52%, a gain of 13 percentage points. Working points are defined by the background misidentification rate: Loose (10%), Medium (1%), and Tight (0.1%), following standard ATLAS/CMS conventions for ํ-tagging calibra- tion [26]. โข ํ vs. ํํํโํก (ํํํโํก-jet rejection): ECT achieves 0.9882 AUC versus ParticleNetโs 0.9451 (+4.3% improvement). The green curves in Figure 6 show that ECT maintains 1โ2 orders of magnitude lower misidentification rates at high efficiencies (> 80%), crucial for analyses requiring tight ํ-tagging selections. โข ํ vs. ํ+ํํํโํก (combined background): ECT reaches 0.9274 AUC compared to ParticleNetโs 0.8904 (+3.7% improvement), with consistent gains visible in the blue curves of Figure 6. These improvements demonstrate that adding transformer attention to ParticleNetโs edge con- volution framework provides substantial performance gains, particularly for the challenging charm rejection task where local vertex information must be integrated with global jet context. 4.1.2 ECT vs. ParT The comparison between ECT and ParT reveals more nuanced trade-offs: โข ํ vs. ํ (charm rejection): ECT significantly outperforms ParT with an AUC of 0.8853 versus 0.8634, a +2.2% improvement that translates to better charm discrimination across all working points. Figure 7 (orange curves) shows that ECT consistently achieves lower misidentification rates. At 60% signal efficiency, ECT reaches 5% misidentification rate while ParT is at 7%, a 40% relative reduction in background contamination. โ 16 โ 0.00.20.40.60.81.0 Signal Efficiency 10 3 10 2 10 1 10 0 Misidentification Rate 10% (Loose) 1% (Medium) 0.1% (Tight) Misidentification rate vs. signal efficiency ECT - b vs. c ECT - b vs. light ECT - b vs. c + light Figure 5. Misidentification rate versus signal efficiency for the ECT model across three ํ-jet classification tasks: ํ vs. ํ (orange), ํ vs. ํํํโํก (green), and ํ vs. ํ+ํํํโํก (blue). Horizontal dashed lines indicate standard working points used in ATLAS analyses: Loose (10% mistag rate), Medium (1%), and Tight (0.1%). The ํ vs. ํํํโํก discrimination achieves the best performance, reaching signal efficiencies above 65% even at the Tight working point, reflecting the distinct signatures of ํํํโํก jets compared to heavy-flavor jets. The more challenging ํ vs. ํ task, where both jet types contain displaced vertices from heavy-hadron decays, shows reduced but still competitive performance. ECT achieves inference throughput of approximately 17,380 jets/s on a single NVIDIA RTX A5000 GPU, corresponding to less than 60 ํs per jet, well within LHC High-Level Trigger latency requirements. โข ํ vs. ํํํโํก (ํํํโํก-jet rejection): Both models achieve excellent and statistically equivalent performance, with ECT at 0.9883 AUC and ParT at 0.9876 AUC. Given the test set size (ํ= 325,290 jets), the observed difference of 0.0007 in AUC is within the expected statistical uncertainty and should not be considered significant. The green curves in Figure 7 are virtually superimposed, indicating that both architectures handle this easier task with comparable effectiveness. โข ํ vs. ํ+ํํํโํก (combined background): ECT outperforms ParT (0.9333 vs. 0.9216, +โ 1.3%), maintaining the advantage observed in the charm rejection task. 4.1.3 Key Insights The comparative analysis reveals three critical findings: 1. Edge convolutions are essential for charm rejection: The superior performance of both ECT and ParticleNet over ParT forํ vs. ํ discrimination (despite ParT having deeper attention โ 17 โ 0.00.20.40.60.81.0 Signal Efficiency 10 3 10 2 10 1 10 0 Misidentification Rate 10% (Loose) 1% (Medium) 0.1% (Tight) Misidentification rate vs. signal efficiency ECT - b vs. c ParticleNet - b vs. c ECT - b vs. light ParticleNet - b vs. light ECT - b vs. c + light ParticleNet - b vs. c + light Figure 6. Comparison of misidentification rate versus signal efficiency between ECT (solid lines) and ParticleNet (dotted lines) for ํ-jet tagging. Three classification scenarios are evaluated: ํ vs. ํ (orange), ํ vs. ํํํโํก (green), and ํ vs. ํ+ํํํโํก (blue). ECT consistently outperforms ParticleNet across all tasks and working points. The most significant improvement is observed in ํ vs. ํ discrimination, where ECT achieves an AUC of 0.885 compared to 0.802 for ParticleNet, representing a 10% relative improvement. At the Medium working point (1% mistag rate), ECT provides approximately 8 percentage points higher ํ-tagging efficiency than ParticleNet for the combined ํ vs. ํ+light classification. This improvement stems from ECTโs hybrid architecture, which combines ParticleNetโs edge convolution operations for local geometric feature extraction with transformer self-attention for capturing global track correlations. with 8 particle blocks vs. ECTโs 4) demonstrates that EdgeConvโs local neighborhood aggregation is crucial for capturing the fine-grained vertex displacement differences between ํ- and ํ-jets. The typical decay length difference (ํํ ํ โ 460 ํm, ํํ ํ โ 150 ํm) requires precise modeling of track-level correlations in the (ํ, ํ) space, which EdgeConv provides through its dynamic graph construction. 2. Transformers excel at ํํํโํก-jet rejection: For the ํ vs. ํํํโํก task, both transformer- based models (ECT and ParT) achieve AUC>98.8%, while ParticleNet reaches 94.5%. This indicates that global attention mechanisms effectively capture the topological differences between heavy-flavor and ํํํโํก jets, where the absence of secondary vertices in ํํํโํก jets is a clear global signature rather than a local feature. 3. Hybrid architecture offers best overall performance: ECT combines the strengths of both approaches: local feature extraction via EdgeConv for challenging tasks (ํ vs. ํ) and global context modeling via transformers for easier tasks (ํ vs. ํํํโํก). This is evidenced by ECT achieving the highest AUC across all three tasks, with training times competitive with ParT โ 18 โ 0.00.20.40.60.81.0 Signal Efficiency 10 3 10 2 10 1 10 0 Misidentification Rate 10% (Loose) 1% (Medium) 0.1% (Tight) Misidentification rate vs. signal efficiency ECT - b vs. c ParT - b vs. c ECT - b vs. light ParT - b vs. light ECT - b vs. c + light ParT - b vs. c + light Figure 7. Comparison of misidentification rate versus signal efficiency between ECT (solid lines) and Particle Transformer (ParT, dash-dotted lines) for ํ-jet tagging. Both architectures employ transformer-based attention mechanisms; however, ECT additionally incorporates EdgeConv blocks for explicit local feature extraction in the(ํ, ํ) coordinate space. ECT demonstrates improved charm-jet rejection (AUC: 0.885 vs. 0.863 for ํ vs. ํ) while maintaining comparable ํํํโํก-jet rejection performance (AUC: 0.988 for both models in ํ vs. ํํํโํก). The EdgeConv component enables ECT to better exploit local geometric relationships among tracks originating from displaced vertices, which is particularly beneficial for distinguishing the decay topologies of ํ-hadrons from those of ํ-hadrons. Despite the additional EdgeConv processing, ECT maintains comparable inference latency to ParT, making it suitable for real-time trigger applications. (Table 4). The performance gains of ECT are most pronounced for charm rejection, the limiting factor in many physics analyses involving ํ-jets. At the Tight working point (0.1% misidentification rate), ECT achieves 48% efficiency for ํ vs. ํ, compared to 42% for ParT and 35% for ParticleNet a 15% relative improvement over the best non-hybrid architecture. 5 Conclusion We present the Edge Convolution Transformer (ECT) architecture for tagging of ํ-jet from ATLAS simulation dataset for three binary classification tasks. The ECT architecture combines local relational learning through edge convolutions with global context modeling through transformer self-attention blocks. The addition of edge convolution was demonstrated to be useful for charm jet โ 19 โ rejection, a challenging task in ํ-jet tagging. The ParT model showed superior performance for ํ- jet tagging against ํํํโํก jet background, however it is computationally more intensive than the ECT model which has a lesser number of particle attention blocks. Experimental results on the ATLAS simulation dataset demonstrate that ECT consistently outperforms both baseline architectures: โข Charm rejection: ECT achieves 88.5% AUC forํ vs. ํ discrimination, surpassing ParticleNet by 8.3% and ParT by 2.2%. At the Medium working point (1% misidentification rate), ECT reaches 65% signal efficiency compared to 60% for ParT and 52% for ParticleNet. โข ํฟํํโํก-jet rejection: ECT and ParT both excel at this task with >98.7% AUC, significantly outperforming ParticleNet (94.5%). ECT achieves 94% efficiency at 1% misidentification rate. โข Combined background: ECT maintains superior performance with 93.3% AUC against the combined ํ+ํํํโํก background, demonstrating robustness across realistic experimental scenarios. โข Inference latency below 0.060 ms per jet on modern GPUs (throughput >17K jets/second). The key finding of this work is that edge convolution blocks are essential for charm jet rejection. The superior performance of ECT over ParT despite ParT having twice as many attention layers demonstrates that local neighborhood aggregation in(ํ, ํ) space satisfies stringent requirements for real-time event selection at the LHC critical for resolving the subtle vertex displacement differences between ํ- and ํ-jets. Conversely, transformer attention excels at capturing global jet topology for ํํํโํก-jet suppression, where the absence of secondary vertices is a clear global signature. By combining these complementary mechanisms, ECT achieves the best overall performance while maintaining computational efficiency competitive with pure transformer models (training time: 78โ150 minutes for 100 epochs on a single GPU). The ECT architecture represents a promising direction for heavy-flavor jet tagging at the LHC, offering improved discrimination power for the challenging charm rejection task while maintaining competitive performance and computational efficiency. Acknowledgments This work was supported by US NSF Award 2334265. The authors thank the Artificial Intelligence Imaging Group (AIIG) at the University of Puerto Rico, Mayaguez for the computational facility. The authors thank Dr. Jesse Thaler, Professor at the MIT Department of Physics and Director of the Institute of Artificial Intelligence and Fundamental Interactions (IAIFI) for his useful feedback. References [1] J. Shlomi. Secondary vertex finding in jets dataset. Zenodo, September 2020. Online; accessed Sep. 22, 2025. [2] Spandan Mondal and Luca Mastrolorenzo. Machine learning in high energy physics: A review of heavy-flavor jet tagging at the lhc. Eur. Phys. J. Special Topics, 233:2687โ2711, 2024. โ 20 โ [3] H. Qu and L. Gouskos. ParticleNet: Jet Tagging via Particle Clouds. arXiv preprint, 2019. Based on CMS-ML documentation: ParticleNet inference guide. [4] Huilin Qu, Congqiao Li, and Sitian Qian. Particle Transformer for Jet Tagging. In Proceedings of the 39th International Conference on Machine Learning (ICML), pages 18281โ18292, 2022. Version v3 (arXiv) published 29Jan2024. [5] Luca Scodellaro, CMS, and ATLAS Collaborations. b tagging in atlas and cms. Proceedings of LHCP2017, arXiv:1709.01290, 2017. hep-ex. [6] M. Aaboud et al. Performance of ํ-Jet Identification in the ATLAS Experiment. JINST, 2018. [7] ATLAS collaboration. Fast ํ-tagging at the high-level trigger of the ATLAS experiment in LHC Run 3. JINST, 18(11):P11006, 2023. [8] ATLAS Collaboration. Identification of Jets Containing $b$-Hadrons with Recurrent Neural Networks at the ATLAS Experiment. ATL-PHYS-PUB-2017-003, CERN Document Server, 2017. [9] Uttiya Sarkar and CMS Collaboration. Run 3 performance and advances in heavy-flavor jet tagging in cms. In Proceedings of the 42nd International Conference on High Energy Physics (ICHEP2024), 2024. [10] CMS Collaboration. Run3pnetbtag. https://twiki.cern.ch/twiki/bin/view/CMSPublic/Run3PNetBtag, March 2025. Accessed: 2025-09-01. [11] CMS Collaboration. Identification of heavy-flavour jets with the cms detector in p collisions at 13 tev. JINST, 13(05):P05011, 2018. [12] Andrea Malara. Exploring jets: substructure and flavour tagging in cms and atlas. Proceedings of LHCP2024, arXiv:2410.14330v1, 2024. hep-ex. [13] Jana Bielฤรญkovรก, Raghav Kunnawalkam Elayavalli, Georgy Ponimatkin, Jรถrn H. Putschke, and Josef ล ivic. Identifying Heavy-Flavor Jets Using Vectors of Locally Aggregated Descriptors. JINST, 16:P03017, 2021. [14] Annika Stein. Improving robustness of jet tagging algorithms with adversarial training: exploring the loss surface, 2023. [15] Daniel Guest, Julian Collado, Pierre Baldi, Shih-Chieh Hsu, Gregor Urban, and Daniel Whiteson. Jet flavor classification in high-energy physics with deep neural networks. Phys. Rev. D, 94:112002, 2016. [16] Diogo Buarque Franzosi, Samuel Calvet, Dรฉsirรฉ Damongo, and Michele Selvaggi. Jet Flavour Tagging for Future Colliders with Fast Simulation. JINST, 2022. [17] Emil Bols, Jan Kieseler, Mauro Verzetti, Markus Stoye, and Anna Stakia. Jet flavour classification using deepjet, 2020. [18] Ayse Asu Guvenli and Bora Isildak. B-jet tagging with retentive networks: A novel approach and comparative study, 2024. [19] Freya Blekman, Florencia Canelli, Alexandre De Moor, Kunal Gautam, Armin Ilg, Anna Macchiolo, and Eduardo Ploerer. Tagging more quark jet flavours at fcc-e at 91 gev with a transformer-based neural network, 2025. [20] Ahmed Hammad and Mihoko M. Nojiri. Transformer networks for heavy flavor jet tagging, 2024. [21] ATLAS Collaboration and Georges Aad et al. Transforming jet flavour tagging at atlas. Preprint CERN-EP-2025-103; arXiv:2505.19689 [hep-ex], 2025. โ 21 โ [22] Torbjรถrn Sjรถstrand, Stefan Ask, Jesper R. Christiansen, Richard Corke, Nishita Desai, Philip Ilten, Stephen Mrenna, Stefan Prestel, Christine O. Rasmussen, and Peter Z. Skands. An Introduction to PYTHIA 8.2. Computer Physics Communications, 191:159โ177, Jun. 2015. [23] J. de Favereau, C. Delaere, P. Demin, A. Giammanco, V. Lemaรฎtre, A. Mertens, and M. Selvaggi. DELPHES 3: a modular framework for fast simulation of a generic collider experiment. Journal of High Energy Physics, 2014(2):057, Feb 2014. Published online Feb 2014. [24] G. Aad and ATLAS Collaboration. The ATLAS Experiment at the CERN Large Hadron Collider. JINST, 3(08):S08003, Aug 2008. Published Aug 2008. [25] Jonathan Shlomi, Sanmay Ganguly, Eilam Gross, Kyle Cranmer, Yaron Lipman, Hadar Serviansky, Haggai Maron, and Nimrod Segol. Secondary vertex finding in jets with neural networks. The European Physical Journal C, 81(6), June 2021. [26] ATLAS Collaboration. ATLAS ํ-jet identification performance and efficiency measurement with ํก ฬ ํก events in ํ collisions at โ ํ = 13 TeV. Eur. Phys. J. C, 79:970, 2019. [27] ATLAS Collaboration. Optimisation of the ATLAS ํ-tagging performance for the 2016 LHC Run. Technical report, CERN, Geneva, 2016. ATLAS Public Note. [28] ATLAS Collaboration. ATLAS Inner Detector: Technical Design Report. CERN-LHCC-97-016, ATLAS-TDR-4, 1997. See also ATLAS tracking software documentation: https://atlassoftwaredocs.web.cern.ch/. [29] Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization, 2014. [30] C. Liand S. Qian H. Qu. Jetclass: A large-scale dataset for deep learning in jet physics. Zenodo, June 2022. Online; accessed Aug. 21, 2025. โ 22 โ