Paper deep dive
MDQEC-QAS: Meta-Decoding for Quantum Error Correction with Hardware-Aware VQC Search and Confidence-Gated Recovery
Prashant Kumar Choudhary, Nouhaila Innan, Muhammad Shafique, Rajeev Singh
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/18/2026, 1:10:23 PM
Summary
The paper introduces MDQEC-QAS, a unified meta-decoding framework for quantum error correction that learns syndrome-to-recovery mappings across multiple stabilizer codes (FiveQubit, Steane, Planar3x3, Planar5x5) and noise settings. It compares a classical Meta-MLP baseline with variational quantum circuit (VQC) meta-decoders selected via hardware-aware quantum architecture search. The study highlights that high teacher-label accuracy does not guarantee logical-level reliability, particularly for larger topological codes like Planar5x5. To address this, a confidence-gated hybrid recovery strategy is proposed, which falls back to a teacher decoder when the learned model's confidence is low, significantly reducing logical failure ratios.
Entities (10)
Relation Signals (9)
MDQEC-QAS → evaluateson → Steane
confidence 95% · The benchmark includes FiveQubit, Steane, Planar3x3, and Planar5x5 codes
MDQEC-QAS → evaluateson → Planar3x3
confidence 95% · The benchmark includes FiveQubit, Steane, Planar3x3, and Planar5x5 codes
MDQEC-QAS → evaluateson → Planar5x5
confidence 95% · The benchmark includes FiveQubit, Steane, Planar3x3, and Planar5x5 codes
MDQEC-QAS → evaluateson → FiveQubit
confidence 95% · The benchmark includes FiveQubit, Steane, Planar3x3, and Planar5x5 codes
MDQEC-QAS → uses → Meta-MLP
confidence 95% · We compare a classical Meta-MLP teacher-trained baseline with variational quantum circuit (VQC) meta-decoders
MDQEC-QAS → uses → VQC
confidence 95% · We compare a classical Meta-MLP teacher-trained baseline with variational quantum circuit (VQC) meta-decoders
VQC → selectedby → Quantum Architecture Search
confidence 92% · VQC meta-decoders selected through hardware-aware quantum architecture search
Confidence-Gated Recovery → reduces → Logical Failure
confidence 90% · confidence-gated fallback reduces them to 1.71 and 1.11
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We propose a unified meta-decoding framework for quantum error correction that learns syndrome-to-recovery mappings across multiple stabilizer codes and noise settings, without requiring separate decoders for each configuration. The benchmark includes FiveQubit, Steane, Planar3x3, and Planar5x5 codes, four noise families, and five evaluation regimes: interpolation, unseen-p transfer, unseen-noise transfer, few-shot unseen-code adaptation, and few-shot held-out-size adaptation. We compare a classical Meta-MLP teacher-trained baseline with variational quantum circuit (VQC) meta-decoders selected through hardware-aware quantum architecture search over qubit count, circuit depth, and entangling topology. The Meta-MLP achieves teacher-label accuracies of 0.9993, 0.9118, 0.9342, 0.6304, and 0.7548 across the five regimes, while the hardware-aware VQC achieves 0.9400, 0.8495, 0.8415, 0.5678, and 0.7143. However, logical-level evaluation shows that high teacher-label accuracy alone is insufficient in the most challenging Planar5x5 setting. During interpolation, the raw logical-failure ratios relative to the teacher are 12.08 and 25.91 for the Meta-MLP and VQC, respectively, whereas confidence-gated fallback reduces them to 1.71 and 1.11. These results support confidence-aware selective recovery rather than unconditional teacher replacement.
Tags
Links
- Source: https://arxiv.org/abs/2607.10707v1
- Canonical: https://arxiv.org/abs/2607.10707v1
Trouble viewing inline? Open PDF directly →
Full Text
79,884 characters extracted from source content.
Expand or collapse full text
MDQEC-QASChoudhary et al. MDQEC-QAS: Meta-Decoding for Quantum Error Correction with Hardware-Aware VQC Search and Confidence-Gated Recovery Prashant Kumar Choudhary 1 , Nouhaila Innan 2,3 , Muhammad Shafique 2,3 and Rajeev Singh 1 1 Department of Physics, Indian Institute of Technology (BHU), Varanasi, Uttar Pradesh, India 2 eBRAIN Lab, Division of Engineering, New York University Abu Dhabi (NYUAD), Abu Dhabi, UAE 3 Center for Quantum and Topological Systems (CQTS), NYUAD Research Institute, NYUAD, Abu Dhabi, UAE prashantkchoudhary.rs.phy22@iitbhu.ac.in nouhaila.innan@nyu.edu muhammad.shafique@nyu.edu rajeevs.phy@iitbhu.ac.in Abstract. We propose a unified meta-decoding framework for quantum error correction that learns syndrome-to-recovery mappings across multiple stabilizer codes and noise settings, without separate decoders for each configuration. The benchmark includes FiveQubit, Steane, Planar3×3 and Planar5×5 codes, four noise families, and five regimes: interpolation, unseen-푝transfer, unseen-noise transfer, few-shot unseen-code adaptation, and few-shot held-out-size adaptation. We compare a classical Meta-MLP teacher-trained baseline with variational quantum circuit (VQC) meta-decoders selected through hardware-aware quantum architecture search over qubit count, circuit depth, and entangling topology. The Meta-MLP achieves teacher-label accuracies of 0.9993, 0.9118, 0.9342, 0.6304, and 0.7548 across the five regimes, while the hardware-aware VQC achieves 0.9400, 0.8495, 0.8415, 0.5678, and 0.7143. However, logical-level evaluation shows that high teacher-label accuracy alone is insufficient in the hardest Planar5×5 setting. During interpolation, raw logical-failure ratios relative to the teacher are 12.08 and 25.91 for the Meta-MLP and VQC, respectively, but confidence-gated fallback reduces them to 1.71 and 1.11. These results support confidence-aware selective recovery rather than unconditional teacher replacement. Keywords: Quantum error correction; meta-decoding; variational quantum circuits; quantum architecture search; stabilizer codes; planar codes; transfer learning; hybrid decoding. 1 Introduction Quantum error correction (QEC) is a central requirement for scalable fault-tolerant quantum computation because frag- ile quantum states are continuously exposed to decoherence, control imperfections, and measurement noise. Within this landscape, the stabilizer formalism provides a unifying frame- work for constructing and analyzing a broad class of quantum codes, including both compact algebraic codes and large topo- logical codes [1–5]. In a stabilizer-code setting, noisy physical qubits are probed through stabilizer measurements, producing a classical syndrome that must be translated into a recovery operation. The success of this decoding stage directly deter- mines the resulting logical failure probability and therefore the usefulness of the encoded quantum information [6–8]. For a measured syndrome푠, decoding can be considered abstractly as the inference problem 푟 ★ (푠)= arg max 푟 Pr(푟 | 푠),(1) where푟denotes a candidate recovery operator. The operational metric of interest is the logical failure probability 푝 퐿 (푝)= Pr(logical failure | 푝),(2) where푝is the physical error probability or, more generally, a parameter controlling the noise strength. In other words, a useful decoder is not just one that predicts labels accurately, but one that preserves logical information after recovery. This distinction becomes particularly important for topological codes, where two decoders with similar classification accuracy can still differ substantially in logical performance [9]. Traditional decoders such as lookup-table strategies, heuris- tic code-specific rules, and matching-based methods can per- form strongly within narrowly defined settings, but they are usually specialized to a certain code family or noise model [10,11]. This limitation has motivated machine-learning- based decoders, from early neural approaches to more scalable graph- and transformer-style models [12, 13]. These developments establish two key observations. First, learned decoders can be competitive with hand-engineered approaches, particularly when the noise has correlations or when decoder inference must adapt to experimental data [14]. Second, the most challenging regimes are often larger struc- tured topological codes, where raw classification accuracy alone does not guarantee strong logical performance [9,15]. Thus, modern QEC decoding is no longer only a combinatorial problem; it increasingly also involves representation learning and decoder-system design. Despite this progress, most learned QEC decoders are still trained and evaluated within fixed code/noise regimes. However, a robust decoder should ideally adapt across het- erogeneous physical conditions: different stabilizer structures, different noise families, and different physical error rates. This motivates a meta-decoding perspective, in which one shared model is trained across multiple decoding tasks and receives explicit side information describing the code family, the noise type, and the physical error level. Rather than learning one isolated decoder per setting, the model learns a transferable syndrome-to-recovery representation [16, 17]. In parallel, variational quantum circuits (VQCs) have emerged as a candidate for hybrid quantum-classical learning systems, but their performance depends strongly on ansatz design [18,19]. This makes quantum architecture search (QAS) a relevant tool for QEC decoding as well: instead of fixing a single parametrized quantum circuit by hand, one may search over qubit count, circuit depth, and entangling topology, and then select architectures using both predictive quality and implementation-oriented cost proxies [20–25]. In a decoding pipeline, this leads naturally to a hardware-aware model-selection problem rather than a purely accuracy-driven one. In this work, we propose a unified meta-decoding frame- work for QEC that combines these ideas in a single repro- ducible pipeline. The framework jointly studies two algebraic codes (FiveQubit and Steane), two planar topological codes (Planar3×3 and Planar5×5), four noise families (depolarizing, 푍-biased, correlated푍-burst, and depolarizing with syndrome- measurement flips), and several transfer settings, including 1 arXiv:2607.10707v1 [quant-ph] 12 Jul 2026 MDQEC-QASChoudhary et al. interpolation, unseen-푝transfer, unseen-noise transfer, and few-shot adaptation to an unseen code. We compare a classi- cal meta-MLP against VQC meta-decoders selected through hardware-aware QAS. A key observation arising from our experiments is that the larger planar code, Planar5×5, is substantially more sensitive to raw learned-decoder mistakes at the logical level than the smaller codes. Even rare teacher-label errors can become logically costly, so high supervised accuracy alone is not a sufficient reliability criterion for this regime. To address this, we introduce a confidence-gated hybrid recovery strategy: when the learned decoder is sufficiently confident, its prediction is used directly; otherwise, the pipeline falls back to the teacher decoder. This mechanism reveals a high-precision selective- decoding regime: the learned model can be useful on confident cases even when it is not reliable as a full teacher replacement. The resulting system should therefore be interpreted as a hybrid learned-assisted recovery rather than an independent learned decoding. Confidence-aware and risk-aware mechanisms are standard in broader machine learning and decision systems, but they remain underexplored in QEC decoding pipelines [26]. The novelty of this work does not lie in the isolated use of a neural decoder or a variational quantum classifier. Rather, it lies in the integration of four components within a sin- gle reproducible QEC pipeline: (i) multi-code, multi-noise meta-decoding across algebraic and planar stabilizer codes, (i) transfer evaluation under unseen-푝, unseen-noise, few- shot unseen-code, and held-out-size settings, (i) compact hardware-aware quantum architecture search for VQC decoder design, and (iv) confidence-gated hybrid recovery for the most challenging topological-code regime. This integrated design leads to three system-level findings: confidence can act as a routing signal for selective decoding, structural transfer across code families and code sizes is substantially harder than trans- fer across noise parameters, and cost-aware VQC selection can remain competitive with accuracy-selected circuits within the compact search space. The main contributions of this paper are summarized as follows: •We construct a pooled multi-code, multi-noise QEC meta- decoding benchmark spanning algebraic and planar stabilizer codes, multiple noise families, and several transfer settings. •We compare a classical Meta-MLP with hardware-aware VQC meta-decoders under a shared global recovery-label framework. •We perform compact hardware-aware QAS over qubit count, circuit depth, and entangling topology, and evaluate the selected VQC architectures after full retraining. •We distinguish teacher-label accuracy from logical-level de- coding performance, showing that raw supervised accuracy can be misleading for larger topological codes. •We identify a transfer hierarchy in which unseen-푝and unseen-noise generalization are easier than unseen-code and held-out-size adaptation, indicating that changes in code structure are more challenging than changes in the noise distribution. •We introduce and evaluate confidence-gated hybrid recovery, showing that confidence can serve as a routing signal be- tween learned predictions and teacher fallback in the larger Planar5×5 regime. 2 Background and Related Work Research on quantum error-correction decoders has evolved from analytical and combinatorial constructions toward in- creasingly data-driven and learning-based approaches. This section reviews prior work at the intersection of stabilizer-code decoding, learned decoder transferability, variational quantum models, and hardware-aware decoder design. 2.1 Classical and learned decoding for stabilizer codes Quantum decoding has traditionally been studied through code- specific analytical and combinatorial methods. Within the stabilizer formalism, the decoder receives a syndrome and must infer a recovery that returns the state to the codespace without inducing a logical fault [3,5]. For topological and surface- code families, decoding is closely connected to threshold analysis and fault-tolerant architecture design. In particular, the topological-quantum-memory framework and later surface- code studies established minimum-weight matching and related combinatorial approaches as strong baselines for structured local-noise settings [6,7,9]. These decoders are often highly effective, but they are typically specialized to a particular code structure and noise assumption [27]. This specialization motivates the search for more adaptive decoder families, since modern experimental platforms exhibit changing calibration, asymmetric noise, measurement faults, and correlated errors. As a result, decoder design increasingly requires methods that combine combinatorial optimization, representation learning, and system-level design. This need for adaptability helped motivate machine-learning- based decoders as data-driven alternatives to explicit code- specific decoding rules. Torlai and Melko provided one of the earliest demonstrations that neural networks can decode topological codes by learning syndrome-to-recovery mappings directly from data [28]. Shortly afterward, Varsamopoulos et al. showed that feedforward neural networks can decode small surface codes efficiently while accounting for implementation- oriented decoding latency [29]. These studies established that decoder behavior can be approximated by trainable mod- els rather than relying exclusively on explicit combinatorial procedures. Subsequent work extended learned decoding to more struc- tured regimes. Baireuther et al. studied machine-learning- assisted correction of correlated qubit errors in topological codes, showing that learned approaches can exploit correla- tions that may be difficult to incorporate into conventional decoders [14]. Reinforcement-learning-based approaches fur- ther explored adaptive decoder behavior and control strategies [30]. More recently, error-rate-agnostic decoding studies in- vestigated whether neural decoders can remain robust when the physical error rate differs between training and testing, thereby motivating transfer-oriented evaluations in QEC [16, 17]. These prior works established the feasibility of learned decoding, but most remained centered on a single code family, a single lattice class, or a single noise scenario. In contrast, this study examines a unified meta-decoding framework spanning multiple stabilizer codes, multiple noise families, and explicit transfer settings. 2.2 Structured decoders, transferability, and logical-level reliability Recent learned QEC decoders have increasingly adopted struc- tured architectures, including graph neural networks and transformer-style models. This development is motivated 2 MDQEC-QASChoudhary et al. by the locality and connectivity patterns of many quantum codes, especially topological codes, which are not always captured efficiently by flat multilayer perceptrons. Lange et al. demonstrated graph-neural-network-based decoding of quantum error-correcting codes, showing that structure-aware learned models can provide strong decoding performance in more complex regimes [31]. Similarly, recent large-scale work on processor data showed that high-capacity learned decoders can achieve strong decoding quality under experimental noise, highlighting the importance of expressive architectures and training distributions [12, 13]. These developments are closely related to the motivation of this study, but the present focus differs in two ways. First, rather than optimizing a learned decoder for a single code family, this study considers a multi-code, multi-noise setting in which a shared decoder is conditioned on the code family, noise family, and physical error rate. Second, this study examines the gap between classification performance and logical-level performance, especially for the larger planar code. This distinction is important because raw learned decoders may achieve nontrivial classification accuracy while still producing substantial logical degradation unless a hybrid recovery mechanism is used [15]. Beyond architectural design, transferability has become a central concern for learned decoding. A major open question in QEC machine learning is whether learned decoders can transfer beyond the regime on which they were trained. Recent error- rate-agnostic and near-term surface-code decoding studies have highlighted this issue by showing that decoders trained under one physical-error regime may generalize imperfectly when evaluated under another [16, 17]. This study extends the transfer perspective by evaluating generalization across interpolation, unseen-푝, unseen-noise, few-shot unseen-code, and held-out-size settings. This combi- nation remains relatively uncommon in prior decoder work, particularly when classical and VQC decoders are evaluated within a unified framework. A recurring challenge in learned QEC decoding is that strong label-level prediction does not necessarily imply strong logical-level recovery, particularly for larger structured codes. This distinction is central to the present study: raw learned decoders, including both the classical meta-MLP and the VQC models, can degrade on Planar5×5 even when their supervised accuracy remains nontrivial. This motivates a hybrid recovery formulation in which learned inference is used selectively and low-confidence cases are delegated to a teacher decoder [26]. Although confidence-based hybridization is common in broader machine-learning systems, it has been less explicitly studied as part of QEC decoder pipelines. In this study, confidence-gated hybrid recovery is used to connect learned transferability with logical reliability, especially in the larger planar-code regime. 2.3 Variational quantum models and architecture search Variational quantum circuits have become an important model class in quantum machine learning. Quantum-enhanced feature spaces [18], circuit-centric quantum classifiers [19], data re- uploading architectures [32], and quantum convolutional neural networks [33] demonstrate that trainable quantum circuits can serve as expressive learning models. These models have been studied across a range of applications, including classification, regression, generative modeling, optimization, and domain- specific prediction tasks, often in hybrid quantum-classical pipelines designed for near-term quantum devices [34–37]. At the same time, their performance depends strongly on circuit architecture, including qubit count, circuit depth, entangling pattern, parameter initialization, and trainability [38–40]. This dependence has motivated quantum architecture search. Du et al. studied quantum circuit architecture search for varia- tional quantum algorithms [20], while later work and surveys examined broader directions in automated quantum model design [21,41]. However, most of this literature does not focus specifically on QEC decoding, and relatively little work ad- dresses transfer-oriented multi-code decoder design. This gap motivates hardware-aware QAS for QEC meta-decoding, with circuit architectures evaluated using predictive performance and implementation-oriented cost proxies such as depth and two-qubit gate count. 2.4 Positioning within prior QEC decoding studies Our study builds on these prior directions while differing in scope, evaluation design, and decoder-system formulation. Compared with early neural decoders, the proposed framework moves beyond single-code decoding toward a shared meta- decoder trained across multiple code and noise settings. Com- pared with correlated-noise and error-rate-adaptive learned decoders, it evaluates transfer across both noise distributions and code structures. Compared with recent graph-based and large-scale learned decoders, its focus is not limited to decoder accuracy, but also includes the relationship between supervised prediction quality and logical-level recovery. Finally, relative to general QAS studies, this work applies hardware-aware VQC search specifically to QEC meta-decoding. Accordingly, our work is positioned as an integrated QEC decoding framework that combines multi-code meta-decoding, transfer evaluation, hardware-aware VQC model selection, and confidence-gated hybrid recovery. Table 1 makes this distinction explicit by contrasting the present framework with representative prior directions in learned QEC decoding. 3 Methodology This section describes the proposed learned-assisted meta- decoding pipeline, shown in Fig. 1. The framework consists of six main components: problem formulation, code and noise specification, pooled dataset construction, meta-decoder training, hardware-aware VQC architecture search, and logical- level evaluation with confidence-gated recovery. 3.1 Problem formulation In a stabilizer-code setting, a physical Pauli error푒acting on 푛qubits induces a syndrome through the binary symplectic product 푠= 퐻푒 푇 (mod 2),(3) where퐻is the binary stabilizer-check matrix of the code [3]. Given syndrome푠, the decoder must infer a recovery푟such that the corrected operator 푒 ′ = 푒⊕ 푟,(4) does not induce a logical fault. If 푒 ′ remains nontrivial at the logical level, decoding fails. The corresponding qubit-level mechanism, from physical Pauli errors and stabilizer syndrome extraction to learned recovery prediction and logical decision, is illustrated in Fig. 2. We formulate this task as meta-decoding. Rather than training a separate decoder for each code and noise condition, 3 MDQEC-QASChoudhary et al. Table 1. Positioning of this study relative to prior directions in learned QEC decoding. Prior directionScopeDistinction of this study Early neural decoding [28, 29]Single code and fixed noise model Multi-code, multi-noise meta-decoding across algebraic and planar stabilizer codes Correlated and adaptive learned decoding [14, 16, 30] Topological codes with corre- lated noise or varying error rates Explicit evaluation under unseen-푝, unseen-noise, and few-shot adap- tation settings Graph-based and large-scale learned decoding [12, 13, 31] Larger or more structured topological-code settings Unified evaluation across multiple code families, including both pre- diction accuracy and logical-level performance Quantum-model-based learn- ing [18, 19, 33] VQC, QCNN, and circuit-based learning models Hardware-aware VQC architecture search applied specifically to QEC meta-decoding Our proposed frameworkMulti-code, multi-noise, transfer- oriented, and logical-level QEC decoding Confidence-gated hybrid recovery for larger topological-code regimes, with teacher fallback in low-confidence cases STAGE 1 Meta-dataset construction Transferable syndrome-recovery data Codes: 5-qubit, Steane,Planar 3×3, Planar 5×5 Noise: depolarizing, Z-biased,Z- burst, readout flip Error rates: p = 0.03,0.05, 0.07 Labels: globalsyndrome-recovery classes Splits: interpolation; held-outp, noise, code, size Output: pooled meta-dataset + metadata STAGE 2 Meta-learning and QAS Weighted Meta-MLP and hardware- aware Meta-VQC Inputs: syndrome + code+ noise + p Baseline: weightedMeta-MLP Search: compactMeta-VQC decoders Score: accuracy & hardware cost Retrain: best-accuracyand best- hardware VQCs Export: Selected circuits as QASM STAGE 3 Transfer benchmark Confidence fallback and logical- level evaluation Evaluate transfer splits Code-aware label masking Confidence-gated teacherfallback Compare Meta-MLPand Meta-VQC Metrics: accuracy, groupwise results, logical failure p_L(p) Measured: robustness | fallback impact Figure 1. Overview of the learned-assisted meta-decoding framework. Syndrome-recovery data are generated using code-appropriate teacher decoders, pooled into a global label space, used to train classical and VQC meta-decoders, and evaluated using teacher-label accuracy, confidence-gated fallback, and logical-failure analysis. a single model is trained on tuples (푠,푐,휂, 푝) ↦→ 푦,(5) where푐is the code identifier,휂is the noise-family identifier, 푝is the physical error probability, and푦is a discrete recovery- class label. The learned decoder is therefore represented as 푓 휃 (푠,푐,휂, 푝)= ˆ푦,(6) where휃denotes trainable parameters andˆ푦is mapped back to a recovery operator for logical evaluation. The supervised objective is teacher-label classification: the model is trained to predict recovery labels produced by a code-appropriate teacher decoder. Teacher-label accuracy measures how well the learned model imitates this reference decoder, but it is not identical to the operational QEC ob- jective because physically distinct recovery strings can be logically equivalent. For this reason, the evaluation distin- guishes teacher-label classification accuracy from logical-level performance. Teacher-label accuracy measures supervised imitation and transfer, whereas logical-level performance is estimated through logical failure after applying the predicted or hybrid recovery. 3.2 Meta-dataset construction The meta-dataset is constructed from four QEC settings: the compact algebraic FiveQubit[[5, 1, 3]]stabilizer code, the Steane[[7, 1, 3]]code, and two planar topological-code in- stances, Planar3×3 and Planar5×5. Teacher decoders are selected according to code family. This code-dependent choice reflects the structure of the benchmark. FiveQubit and Steane have compact syndrome spaces, soNaiveDecoderprovides a consistent reference decoder for supervised label genera- tion. Planar3×3 and Planar5×5 are planar topological-code instances, for which matching-based decoding is aligned with the lattice structure. To evaluate transfer across heterogeneous physical condi- tions, samples are generated under four noise families: depolar- izing noise, (Z)-biased Pauli noise, correlated (Z)-burst noise, and depolarizing noise with syndrome-measurement flips. For the (Z)-biased model, the single-qubit Pauli distribution is Pr(퐼)= 1− 푝, Pr(푋)= 0.2푝, Pr(푌)= 0.1푝, Pr(푍)= 0.7푝. (7) The coefficients(0.2, 0.1, 0.7)are not intended as universal hardware-calibrated constants; rather, they define a represen- tative moderately푍-biased Pauli channel in which the total non-identity error probability remains푝, while푍-type errors dominate. The correlated (Z)-burst model augments local Pauli errors with adjacent correlated (Z)-type events, while the measurement-flip variant additionally flips syndrome bits with probability proportional to the physical error rate. Supervised data are generated at 푝 ∈ 0.03, 0.05, 0.07,(8) with logical-failure curves evaluated later on a denser physical- error grid. 4 MDQEC-QASChoudhary et al. Figure 2. Qubit-level mechanism of learned-assisted QEC meta-decoding. A logical qubit|휓 퐿 ⟩ is encoded into physical qubits, where Pauli errors generate stabilizer syndromes according to푠= 퐻푒 푇 (mod 2). The syndrome, code identity, noise identity, and normalized physical error rate are combined into the meta-feature vector푥= [ ̃푠∥onehot(푐)∥onehot(휂)∥ ̃푝]and projected to a compact VQC input. The VQC applies angle embedding, trainable rotation layers, and entangling operations, after which Pauli-푍measurements are mapped by a classical head to a recovery label and recovery operator푟. Decoding success is determined by whether the corrected operator푒 ′ = 푒⊕ 푟is logically trivial or logically nontrivial. For each code, noise family, and physical error rate, raw syndrome-recovery records are first generated as 푠 푖 ,푐 푖 ,휂 푖 , 푝 푖 ,푟 푖 ,(9) where푠 푖 is the syndrome,푐 푖 is the code identifier,휂 푖 is the noise-family identifier,푝 푖 is the physical error probability, and 푟 푖 is the teacher recovery. Because the codes have different numbers of stabilizers and physical qubits, all syndrome vectors and binary symplectic recovery strings are padded to common global dimensions: ̃푠 푖 ∈ 0, 1 푆 max , ̃푟 푖 ∈ 0, 1 2푄 max .(10) All raw records are pooled before the final train/test split. Each distinct code-specific padded recovery string is then assigned a global recovery label, 푦 푖 = LabelMap(푐 푖 , ̃푟 푖 ).(11) Including the code identifier in the label key preserves code- specific recovery semantics while allowing all models to train over one unified supervised label space. For non-held-out settings, the pooled dataset is split using a label-aware protocol. Singleton labels are assigned to training, labels with multiple occurrences retain at least one representa- tive in training, and test labels are restricted to labels already represented in the training split. This prevents interpolation test classes from being absent during training. For held-out transfer settings, unseen test labels are allowed when they arise from the transfer design, since those settings intentionally evaluate out-of-distribution generalization. The resulting benchmark includes one interpolation set- ting and four transfer settings: unseen-푝transfer, unseen- noise transfer, few-shot unseen-code adaptation, and few-shot Table 2. Dataset split summary for the five evaluation settings. The global recovery-label space contains 15,582 code-specific recovery labels. Unseen test labels denote labels present in the test split but absent from the training split. SettingTrain samples Test samples Unseen test labels Interpolation147,76222,7460 Unseen-푝 transfer99,72654,7584,515 Unseen-noise transfer111,26446,9122,920 Few-shot unseen-code adaptation76,50046,50013,529 Few-shot held-out-size adaptation54,00069,00013,505 held-out-size adaptation. In the main unseen-푝experiment, 푝= 0.05is withheld during training. In the unseen-noise experiment, the (Z)-biased noise model is held out. In the few-shot unseen-code experiment, Planar5×5 is held out from full training and only 5% of its records are included as sup- port. In the held-out-size experiment, code size (5) is held out from full training with 5% support. Table 2 summarizes the resulting split sizes and unseen-label counts. The sample counts reported are the resulting train-test counts obtained after applying the dataset-generation and split protocols, rather than independently tuned hyperparameters. 3.3 Meta-decoder models and training protocol Each sample is represented by concatenating the padded syn- drome vector, a one-hot code identifier, a one-hot noise-family identifier, and the normalized physical error rate. The resulting meta-feature vector is 푥= [ ̃푠∥ onehot(푐)∥ onehot(휂)∥ ̃푝](12) where 휂 denotes the noise-family identifier and ̃푝= 푝 0.1 .(13) 5 MDQEC-QASChoudhary et al. This representation allows one shared decoder to condition its prediction on the syndrome, code family, noise family, and physical error rate. The classical baseline is a meta-decoder based on a mul- tilayer perceptron, denoted Meta-MLP. Given input vector푥, the forward map is ℎ 1 = 휙(푊 1 푥+ 푏 1 ),(14) ℎ 2 = 휙(푊 2 ℎ 1 + 푏 2 ),(15) 푧= 푊 3 ℎ 2 + 푏 3 ,(16) where휙(·)is the ReLU activation and푧contains logits over the global recovery-label space. The implementation uses two hidden layers with layer normalization and dropout. To account for the greater difficulty of the larger planar code, the training protocol applies Planar5×5 oversampling and a sample-weighted cross-entropy loss. The weighted objective is L wCE = 1 푁 푁 ∑︁ 푖=1 푤 푖 ℓ CE (푧 푖 , 푦 푖 ),(17) where푤 푖 increases the contribution of harder training regimes. In the final implementation, Planar5×5 samples are oversam- pled by a factor of 3. The sample weight is factorized into code and noise components: the code weights are 1.0, 1.0, 1.2, and 3.0 for FiveQubit, Steane, Planar3×3, and Planar5×5, respectively, while the noise weights are 1.0, 1.0, 1.1, and 1.1 for depolarizing, (Z)-biased, correlated (Z)-burst, and depo- larizing noise with syndrome-measurement flips, respectively. This weighting makes the supervised objective more sensitive to regimes that are harder at the logical level. The quantum counterpart is a hybrid variational quantum classifier. Because the meta-feature vector can exceed the number of available circuit qubits, the input is first projected to the qubit dimension: 푢= 푊 proj 푥+ 푏 proj ,(18) where푢 ∈R 푛 푞 and푛 푞 is the qubit count of the candidate VQC architecture. The projected vector푢is embedded through angle embedding into a parameterized quantum circuit. Each circuit contains퐿trainable layers, with each layer applying single-qubit푅 푌 and푅 푍 rotations on every qubit followed by an entangling pattern selected from the architecture search space. The quantum feature vector is defined by Pauli-(Z) expectation values, 푞(푢)= ⟨푍 1 ⟩,...,⟨푍 푛 푞 ⟩ ,(19) which are passed to a classical linear head: ˆ푧= 푊 head 푞(푢)+ 푏 head .(20) The VQC meta-decoder is trained with the same weighted cross-entropy objective as the Meta-MLP. Because recovery labels are code-specific, prediction is restricted using a code-label mask. For a sample from code푐, logits corresponding to labels associated with other codes are suppressed: 푧 푘 ←−∞ for all 푘∉Y(푐),(21) whereY(푐)is the set of valid labels for code푐. This masking step prevents cross-code recovery-label predictions and ensures that the global label space remains consistent with code-specific recovery semantics. 3.4 Hardware-aware VQC architecture search The VQC architecture is selected through a compact hardware- aware quantum architecture search rather than fixed manually. The search space varies three architectural factors: • qubit count 푛 푞 ∈ 4, 5, 6, • circuit depth 퐿 ∈ 1, 2, • entangling patternE ∈ chain, ring, full. This defines a discrete search space of 18 candidate architec- tures. The search procedure has two stages. In the search stage, each candidate architecture is trained for 10 epochs and eval- uated using validation accuracy. In the retraining stage, the architecture selected by validation accuracy and the architec- ture selected by hardware-aware score are each retrained from scratch for 35 epochs. This protocol keeps the search compu- tationally manageable while ensuring that the final selected VQC models are evaluated after the same full retraining budget. The search is therefore intended as a compact NISQ-oriented architecture-selection study rather than a large combinatorial search. To account for implementation cost, each architecture is also evaluated using hardware-oriented proxies, including raw circuit depth, raw two-qubit gate count, transpiled depth, and transpiled two-qubit gate count. For a candidate architecture 푎, the hardware-aware score is defined as 푆(푎)= Acc val (푎)− 휆 1 퐷 raw (푎)− 휆 2 퐺 raw 2푞 (푎) − 휆 3 퐷 tr (푎)− 휆 4 퐺 tr 2푞 (푎). (22) where퐷denotes circuit depth and퐺 2푞 denotes two-qubit gate count, each measured in raw and transpiled forms. The coefficients휆 1 ,휆 2 ,휆 3 ,휆 4 control the penalty assigned to circuit cost. The best-by-hardware-score VQC is selected according to 푆(푎), whereas the best-by-accuracy VQC is selected using validation accuracy alone. 3.5 Confidence-gated recovery and logical evaluation The larger planar code, Planar5×5, is especially sensitive to raw learned-decoder errors at the logical level. To reduce this risk, we use a confidence-gated fallback mechanism. Given the masked output distribution of a learned decoder, the prediction confidence is defined as 푐(푥)= max 푦 푝 휃 (푦 | 푥).(23) The final recovery rule is 푟(푥)= ( 푟 learned (푥), 푐(푥) ≥ 휏, 푟 teacher (푥), 푐(푥) < 휏, (24) where푟 teacher is the code-appropriate teacher recovery and 휏is a decoder-specific confidence threshold. In this study, fallback is enabled only for Planar5×5. The fixed thresholds are휏= 0.80for the Meta-MLP and휏= 0.75for the VQC decoders. These thresholds are treated as fixed operating points. Logical-level evaluation is performed through Monte Carlo simulation. For each code, noise family, decoder, and physical error rate, we estimate 푝 퐿 (푝)= 푁 fail 푁 run ,(25) 6 MDQEC-QASChoudhary et al. where푁 run is the number of sampled error instances and푁 fail is the number of cases in which the corrected operator remains logically nontrivial. To compare learned decoders against the teacher, we use the logical-failure ratio 푅 logic (푝)= 푝 (learned) 퐿 (푝) 푝 (teacher) 퐿 (푝) .(26) Values near1indicate teacher-level logical behavior. In the final benchmarking, logical-failure curves are evaluated on the grid푝 ∈ 0.01, 0.02,..., 0.10using 20,000 Monte Carlo trials per point. 4 Results and Discussion 4.1 Experimental setup The full pipeline is implemented in Python usingqecsimfor code simulation and teacher decoding,PyTorchfor the classi- cal Meta-MLP,PennyLanefor the VQC meta-decoders, and Qiskitfor hardware-aware circuit costing and export. Table 3 summarizes the experimental settings and implementation details used in this study. 4.2 Teacher-label accuracy and VQC selection We first evaluate the pooled meta-decoding framework at the teacher-label level across five regimes: interpolation, unseen-푝 transfer, unseen-noise transfer, few-shot unseen-code adapta- tion, and few-shot held-out-size adaptation. The evaluated methods are a majority-label baseline, the classical Meta-MLP, the VQC selected by search-stage validation accuracy, and the VQC selected by the hardware-aware score. Table 4 reports the main teacher-label accuracy results. The Meta-MLP is the strongest raw classifier in all five settings. It reaches 0.9993 in interpolation and remains strong under unseen-푝and unseen-noise transfer, with accuracies of 0.9118 and 0.9342, respectively. The largest drops occur in the few- shot adaptation regimes, where the accuracy decreases to 0.6304 for unseen-code adaptation and 0.7548 for held-out- size adaptation. This indicates that shifts in physical error rate or noise family are easier to absorb than changes in code structure or code size. The VQC meta-decoders follow the same difficulty pattern but remain below the Meta-MLP in raw teacher-label accu- racy. Still, both VQC variants stay well above the majority baseline in every regime, showing that they learn nontrivial syndrome-to-recovery structure. The hardware-aware VQC also remains close to the accuracy-selected VQC after retrain- ing and gives higher accuracy in four of the five settings. The only exception is few-shot unseen-code adaptation, where the accuracy-selected VQC is slightly higher. Thus, the VQC results are best interpreted as compact quantum meta-decoders selected under different criteria, rather than as accuracy win- ners over the classical Meta-MLP. The learning curves support this interpretation. The Meta- MLP converges rapidly and stably on the pooled syndrome- plus-metadata representation, as shown in Fig. 3. This behavior is consistent with the near-saturated interpolation result in Table 4. In contrast, the selected VQCs converge more slowly and reach lower final accuracy, as shown in Fig. 4. However, both VQC variants improve during retraining across all five regimes, including the few-shot transfer settings. The VQC architecture-search results show that the candidate circuits are not equivalent. Fig. 5 reports validation and Figure 3. Training and validation accuracy of the Meta-MLP on the pooled multi-code, multi-noise dataset. Figure 4. Retraining curves of the two selected VQC meta-decoders. Panels (a)–(e) correspond to interpolation, unseen-푝transfer, unseen- noise transfer, few-shot unseen-code adaptation, and few-shot held- out-size adaptation, respectively. search-stage test accuracy, while Fig. 6 reports the hardware- aware score. Because the candidates differ in qubit count, circuit depth, and entangling topology, the search produces distinct accuracy-cost tradeoffs. The hardware-aware criterion therefore acts as a model-selection rule during VQC search, not only as a post hoc circuit-cost analysis. The selected circuit architectures in Fig. 7 and Fig. 8 show the effect of the two selection criteria at the circuit level. Com- pared with the accuracy-selected circuit, the hardware-aware selection chooses a lower-cost architecture while retaining competitive teacher-label accuracy. 4.3 Interpolation We begin with the interpolation setting, where all code fam- ilies, noise models, and supervised physical error rates are represented during training. This setting tests whether the pooled meta-decoder can learn a shared syndrome-to-recovery map in the absence of distributional shift. As shown in Table 4, the Meta-MLP achieves a test accuracy of 0.9993, far above the majority baseline of 0.2032. The VQC models also perform well above the baseline, with accuracies of 0.8987 for the accuracy-selected VQC and 0.9400 for the hardware-aware VQC. In this setting, the hardware-aware VQC outperforms the accuracy-selected VQC after retraining, showing that cost-aware selection does not necessarily reduce teacher-label accuracy within the considered VQC search space. The code-wise interpolation results in Table 5 show that the algebraic codes are the easiest cases. FiveQubit is solved by all learned models, and Steane is nearly saturated by the Meta-MLP while remaining high for both VQCs. The planar codes are more difficult for the VQC models. Planar3×3 remains manageable, but Planar5×5 is the main bottleneck: the Meta-MLP reaches 0.999, whereas the accuracy-selected and hardware-aware VQCs reach 0.692 and 0.830, respectively. 7 MDQEC-QASChoudhary et al. Table 3. Experimental setup and implementation details used in this study. Parameter/FunctionValue Physical error rates for supervised data 푝 ∈ 0.03, 0.05, 0.07 Logical-curve grid푝 ∈ 0.01, 0.02,..., 0.10 Logical Monte Carlo trials20,000 per point Global syndrome padding40 stabilizer bits Global recovery padding82 binary symplectic bits Global label-space size15,582 labels Meta featuressyndrome, code one-hot, noise one-hot, normalized 푝/0.1 MLP hidden dimensions512, 256 MLP activation/normalization/dropoutReLU, LayerNorm, dropout 0.15 MLP optimizerAdam, learning rate 10 −3 , weight decay 10 −5 MLP batch size / epochs256 / 35 Planar5×5 oversampling factor3 Code weightsFiveQubit 1.0; Steane 1.0; Planar3×3 1.2; Planar5×5 3.0 Noise weightsDepolarizing 1.0; 푍-biased 1.0; correlated 푍-burst 1.1; depol.+meas. flip 1.1 VQC simulatorPennyLane default.qubit statevector simulator VQC search space푛 푞 ∈ 4, 5, 6, 퐿 ∈ 1, 2, entanglement in chain/ring/full VQC optimizerAdam, learning rate 10 −2 VQC batch size / epochs32; 10 search epochs and 35 final retraining epochs Qiskit backend for cost proxy FakeManilaV2, FakeManila Transpilation optimization level1 Fallback code familyPlanar5×5 only Fallback thresholdsMeta-MLP 휏= 0.80; VQC 휏= 0.75 Table 4. Overall teacher-label classification accuracy across the five evaluation settings. SettingMajority Meta-MLP VQC best-acc VQC best-hw Interpolation0.20320.99930.89870.9400 unseen-푝 transfer0.19660.91180.83970.8495 Unseen-noise transfer0.19750.93420.83840.8415 Few-shot unseen-code adaptation 0.09940.63040.54350.5678 Few-shot held-out-size adaptation 0.05990.75480.67740.7143 Rows are ranked by the mean of validation and test accuracy across panels (a)-(e). Cell values are accuracy (%). Validation accuracy Test accuracy Overall mean No.(a)(b)(c)(d)(e)(a)(b)(c)(d)(e) 1 6q-2L-full 5961589186898078516471.8 2 6q-2L-chain 5961579186898179496571.6 3 6q-1L-full 5959589187897781506271.3 4 6q-2L-ring 6058509287917869506469.9 5 5q-2L-ring 5757568785877580455868.8 6 5q-1L-full 5457539185837575495868.0 7 5q-2L-full 5757579086867664485767.8 8 6q-1L-chain 5253528882817274465765.8 9 5q-2L-chain 5655539085857450475865.3 10 6q-1L-ring 5151518783776972465564.2 11 4q-2L-ring 5150488882776767465463.0 12 5q-1L-chain 5048488782776667465662.7 13 4q-2L-full 4851488882736965465462.4 14 4q-1L-full 4751488981727065465462.2 15 5q-1L-ring 5046488679776265465361.3 16 4q-2L-chain 4851488679726955455360.5 17 4q-1L-chain 4647458378716455435458.8 18 4q-1L-ring 4743468275745964435258.4 50607075 Accuracy (%) Cell intensity: 507090 higher = darker Figure 5. Validation and search-stage test accuracy of candidate VQC architectures explored during quantum architecture search. This code-wise pattern is consistent with the broader group- wise results in Fig. 9(a). Interpolation accuracy is nearly saturated for the Meta-MLP across most code-noise-푝groups, while the VQC errors are concentrated in the planar-code set- tings. Thus, interpolation establishes the main in-distribution result: the pooled decoder can learn the shared teacher-label task, but the larger planar code remains the most difficult case for the compact VQC models. Although interpolation is almost solved at the teacher-label level by the Meta-MLP, label accuracy alone does not fully Ranked matrix view: rows are architectures sorted by mean score; cells show panels (a)-(e); right bars show the cross-panel mean. No.Architecture(a)(b)(c)(d)(e)Mean score 16q-2L-chain 0.580.600.560.900.86 0.700 26q-1L-full 0.580.580.570.900.85 0.696 36q-2L-full 0.560.580.560.890.84 0.688 4 6q-2L-ring 0.590.570.490.910.86 0.685 55q-2L-full 0.550.560.560.890.84 0.680 6 5q-2L-ring 0.560.560.550.860.84 0.677 75q-1L-full 0.540.560.520.900.85 0.672 85q-2L-chain 0.550.540.530.890.84 0.670 96q-1L-chain 0.520.530.520.880.81 0.651 10 6q-1L-ring 0.500.510.510.860.83 0.641 11 4q-2L-ring 0.500.490.470.870.82 0.629 124q-1L-full 0.460.510.480.890.80 0.627 135q-1L-chain 0.500.480.480.860.82 0.627 144q-2L-full 0.470.500.470.870.81 0.625 154q-2L-chain 0.470.500.470.850.78 0.616 16 5q-1L-ring 0.500.460.480.850.79 0.614 174q-1L-chain 0.460.470.450.830.78 0.597 18 4q-1L-ring 0.470.430.450.810.74 0.582 0.550.600.650.70 Score scale 0.400.92 Cell numbers are scores from panels (a)-(e). Figure 6. Hardware-aware score of candidate VQC architectures. The score combines predictive performance with implementation- oriented circuit-cost proxies such as depth and two-qubit gate count. q 0 : R Y (0) R Y (2.272e− 05)R Z (6.283) • R Y (0.5147)R Z (6.283) • q 1 : R Y (0)R Y (1.714e− 05)R Z (3.142) • R Y (3.156)R Z (3.142) • q 2 : R Y (0)R Y (3.142)R Z (−8.761e− 05) • R Y (6.283)R Z (0) • q 3 : R Y (0)R Y (3.142)R Z (1.571) • R Y (−0.2971)R Z (6.283) • q 4 : R Y (0)R Y (3.142)R Z (−1.571) • R Y (0.00618)R Z (3.142) • q 5 : R Y (0)R Y (3.142)R Z (1.571) • R Y (3.155)R Z (−1.394e− 09) • meas : / 6 0 1 2 3 4 5 (a) q 0 : R Y (0) R Y (0.001944)R Z (4.707) • R Y (6.216)R Z (3.142) • q 1 : R Y (0)R Y (3.141)R Z (1.569) • R Y (2.678)R Z (3.142) • q 2 : R Y (0)R Y (4.709)R Z (3.841) • R Y (4.712)R Z (3.142) • q 3 : R Y (0)R Y (4.712)R Z (3.026) • R Y (1.636)R Z (3.142) • q 4 : R Y (0)R Y (6.289)R Z (4.714) • R Y (4.721)R Z (6.283) • q 5 : R Y (0)R Y (1.571)R Z (0.4008)R Y (4.712)R Z (3.142) meas : / 6 0 1 2 3 4 5 (b) q 0 : R Y (0) R Y (3.141)R Z (5.32) • R Y (6.291)R Z (2.462e− 08) • q 1 : R Y (0)R Y (3.141)R Z (2.748) • R Y (0.01228)R Z (3.142) • q 2 : R Y (0)R Y (3.141)R Z (1.561) • R Y (0.05345)R Z (6.283) • q 3 : R Y (0)R Y (6.281)R Z (4.72) • R Y (5.55)R Z (3.142) • q 4 : R Y (0)R Y (4.735)R Z (3.047) • R Y (1.578)R Z (3.142) • q 5 : R Y (0)R Y (4.713)R Z (2.401)R Y (1.452)R Z (1.65e− 09) meas : / 6 0 1 2 3 4 5 (c) q 0 : R Y (0)R Y (1.572)R Z (5.577) • R Y (4.718)R Z (5.733) • q 1 : R Y (0)R Y (6.3)R Z (4.718) • R Y (1.571)R Z (6.158) • q 2 : R Y (0)R Y (0.004706)R Z (4.715) • R Y (4.713)R Z (3.041) • q 3 : R Y (0)R Y (3.08)R Z (1.576) • R Y (−1.575)R Z (3.068) • q 4 : R Y (0)R Y (−0.06068)R Z (4.695) • R Y (1.567)R Z (1.521) • q 5 : R Y (0)R Y (1.571)R Z (4.062) • R Y (4.696)R Z (1.994) • meas : / 6 0 1 2 3 4 5 (d) q 0 : R Y (0)R Y (3.141)R Z (4.725) • R Y (−0.06412)R Z (6.156) • q 1 : R Y (0)R Y (3.145)R Z (1.573) • R Y (0.1509)R Z (3.152) • q 2 : R Y (0)R Y (6.283)R Z (4.715) • R Y (−0.03973)R Z (3.196) • q 3 : R Y (0)R Y (9.765e− 05)R Z (−1.576) • R Y (6.282)R Z (0.1787) • q 4 : R Y (0)R Y (−0.003203)R Z (6.287) • R Y (3.243)R Z (4.959) • q 5 : R Y (0)R Y (−0.001164)R Z (3.14) • R Y (3.145)R Z (3.143) • meas : / 6 0 1 2 3 4 5 (e) Figure 7. Selected VQC architectures obtained using search-stage val- idation accuracy. Panels (a)–(e) correspond to interpolation, unseen-푝 transfer, unseen-noise transfer, few-shot unseen-code adaptation, and few-shot held-out-size adaptation, respectively. 8 MDQEC-QASChoudhary et al. q 0 : R Y (0)R Y (4.712)R Z (2.422) • R Y (4.713)R Z (0) • q 1 : R Y (0)R Y (1.572)R Z (3.154) • R Y (1.645)R Z (3.142) • q 2 : R Y (0)R Y (3.226)R Z (4.744) • R Y (4.693)R Z (0) • q 3 : R Y (0)R Y (7.808e− 05)R Z (4.712) • R Y (1.569)R Z (−1.34e− 09) • q 4 : R Y (0)R Y (1.568)R Z (5.947) • R Y (4.711)R Z (6.283) • q 5 : R Y (0)R Y (4.712)R Z (5.059) • R Y (4.711)R Z (6.283) • meas : / 6 0 1 2 3 4 5 (a) q 0 : R Y (0)R Y (6.295)R Z (7.85) • R Y (4.887)R Z (6.283) • q 1 : R Y (0)R Y (6.339)R Z (4.694) • R Y (4.731)R Z (6.283) • q 2 : R Y (0)R Y (3.183)R Z (1.557) • R Y (4.712)R Z (3.142) • q 3 : R Y (0)R Y (1.572)R Z (3.077) • R Y (1.577)R Z (6.283) • q 4 : R Y (0)R Y (4.712)R Z (0.8241) • R Y (4.711)R Z (2.147) • q 5 : R Y (0)R Y (4.714)R Z (2.597)R Y (1.575)R Z (3.142) meas : / 6 0 1 2 3 4 5 (b) q 0 : R Y (0)R Y (3.142)R Z (3.142) • q 1 : R Y (0)R Y (3.14)R Z (6.283) • q 2 : R Y (0)R Y (3.142)R Z (6.283) • q 3 : R Y (0)R Y (3.142)R Z (3.142) • q 4 : R Y (0)R Y (0)R Z (3.142) • q 5 : R Y (0)R Y (6.276)R Z (6.283) meas : / 6 0 1 2 3 4 5 (c) q 0 : R Y (0)R Y (6.284)R Z (1.571) • R Y (4.712)R Z (3.169) • q 1 : R Y (0)R Y (4.712)R Z (4.023) • R Y (4.713)R Z (5.762) • q 2 : R Y (0)R Y (4.714)R Z (1.365) • R Y (1.566)R Z (6.079) • q 3 : R Y (0)R Y (6.288)R Z (1.57) • R Y (4.713)R Z (0.4079) • q 4 : R Y (0)R Y (3.14)R Z (1.571) • R Y (4.712)R Z (5.669) • q 5 : R Y (0)R Y (1.572)R Z (0.3995) • R Y (4.711)R Z (2.983) • meas : / 6 0 1 2 3 4 5 (d) q 0 : R Y (0)R Y (4.712)R Z (4.078) • R Y (4.694)R Z (5.973) • q 1 : R Y (0)R Y (1.571)R Z (3.243) • R Y (4.726)R Z (2.432) • q 2 : R Y (0)R Y (3.142)R Z (4.712) • R Y (4.71)R Z (6.027) • q 3 : R Y (0)R Y (3.138)R Z (1.572) • R Y (4.712)R Z (3.15) • q 4 : R Y (0)R Y (4.752e− 06)R Z (4.712) • R Y (1.571)R Z (4.061) • q 5 : R Y (0)R Y (1.571)R Z (5.027) • R Y (1.571)R Z (3.131) • meas : / 6 0 1 2 3 4 5 (e) Figure 8. Selected VQC architectures obtained using the hardware- aware score. Panels (a)–(e) correspond to interpolation, unseen-푝 transfer, unseen-noise transfer, few-shot unseen-code adaptation, and few-shot held-out-size adaptation, respectively. Table 5. Average code-wise interpolation accuracy. CodeMeta-MLP VQC best-acc VQC best-hw FiveQubit1.0001.0001.000 Steane1.0000.9470.976 Planar3×3 0.9980.8700.887 Planar5×5 0.9990.6920.830 determine decoder reliability. We therefore return to logical- failure behavior in Sec. 4.8, where Planar5×5 is analyzed under the confidence-gated fallback protocol. 4.4 Transfer to unseen physical error rate The unseen-푝setting tests whether the pooled representation generalizes across physical error magnitude rather than mem- orizing the specific푝-values used in training. As shown in Table 4, all learned models remain clearly above the majority baseline in this regime. The Meta-MLP achieves 0.9118, while the accuracy-selected and hardware-aware VQCs reach 0.8397 and 0.8495, respectively. Relative to interpolation, the performance drop is noticeable but moderate, suggesting that the learned representation captures structure that persists across nearby physical error rates. The code-wise decomposition in Table 6 shows that the degradation is not uniform across codes. FiveQubit and Steane remain near-saturated, and Planar3×3 stays close to 0.9 even for the VQCs. The main loss again occurs in Planar5×5, where the Meta-MLP reaches 0.844, while the accuracy-selected and hardware-aware VQCs reach 0.644 and 0.670, respectively. This same progression is visible in Fig. 9(b): unseen-푝transfer remains robust for most groups, but the larger planar-code setting is where generalization degrades most clearly. Thus, unseen-푝transfer shows that changing the physical error magnitude is manageable at the teacher-label level, but Planar5×5 remains the main bottleneck. 4.5 Transfer to unseen noise family The unseen-noise setting is a stronger generalization test be- cause it changes the structure of the corruption process rather than only its magnitude. As shown in Table 4, the pooled framework remains robust under this shift. The Meta-MLP Table 6. Average code-wise accuracy under unseen-푝 transfer. CodeMeta-MLP VQC best-acc VQC best-hw FiveQubit1.0001.0001.000 Steane1.0000.9740.978 Planar3×3 0.9930.8770.896 Planar5×5 0.8440.6440.670 MLP VQC-acc VQC-hw (a) MLP VQC-acc VQC-hw (b) MLP VQC-acc VQC-hw (c) MLP VQC-acc VQC-hw (d) MLP VQC-acc VQC-hw (e) ●◆▲●◆▲●◆▲●◆▲●◆▲●◆▲●◆▲●◆▲●◆▲●◆▲●◆▲●◆▲●◆▲●◆▲●◆▲●◆▲ Corr. ZD+measDepol.Z-biasedCorr. ZD+measDepol.Z-biasedCorr. ZD+measDepol.Z-biasedCorr. ZD+measDepol.Z-biased Five-qubitPlanar 3x3Planar 5x5Steane p: Noise: Code: 0.0 0.2 0.4 0.6 0.8 1.0 Test accuracy Physical error probability p: ● 0.03 ◆ 0.05 ▲ 0.07 Figure 9. Groupwise teacher-label accuracy across code, noise family, and physical error rate for the final benchmarked decoders. Panels (a)–(e) correspond to interpolation, unseen-푝transfer, unseen-noise transfer, few-shot unseen-code adaptation, and few-shot held-out-size adaptation, respectively. achieves 0.9342, while the accuracy-selected and hardware- aware VQCs reach 0.8384 and 0.8415, respectively. These values are close to the unseen-푝setting, indicating that the syndrome-plus-metadata representation captures structure that transfers across noise families. The code-wise results in Table 7 again show that the difficulty is concentrated mainly in Planar5×5. FiveQubit and Steane remain near-perfect, and Planar3×3 stays strong. Planar5×5 is harder, with accuracies of 0.900 for the Meta-MLP and 0.691 and 0.670 for the two VQCs. Fig. 9(c) shows the same pattern: the algebraic and smaller planar-code settings remain robust, while the larger planar code remains the main transfer bottleneck. Thus, unseen-noise transfer remains feasible at the teacher- label level, but the result also confirms that changing the noise family does not remove the structural difficulty of Planar5×5. 4.6 Few-shot adaptation to an unseen code The few-shot unseen-code setting is the most demanding transfer regime because the model must adapt to a held-out code configuration using only a small support fraction. In the present setup, the held-out code is Planar5×5 with 5% support. As shown in Table 4, this setting gives the strongest overall degradation among all experiments. The Meta-MLP reaches 0.6304, while the accuracy-selected and hardware- aware VQCs reach 0.5435 and 0.5678, respectively. Although these values remain above the majority baseline of 0.0994, they are far below the corresponding interpolation, unseen-푝, and unseen-noise results. Table 8 clarifies the source of this degradation. The non- held-out codes remain almost saturated, but Planar5×5 drops sharply to 0.397 for the Meta-MLP and to 0.270 and 0.308 for the two VQCs. Fig. 9(d) shows that this drop is highly localized: most configurations remain easy, while the truly unseen code structure is substantially harder. This result 9 MDQEC-QASChoudhary et al. Table 7. Average code-wise accuracy under unseen-noise transfer. CodeMeta-MLP VQC best-acc VQC best-hw FiveQubit1.0000.9860.980 Steane1.0000.9670.968 Planar3×3 0.9960.8720.881 Planar5×5 0.9000.6910.670 Table 8. Average code-wise accuracy under few-shot unseen-code adaptation. CodeMeta-MLP VQC best-acc VQC best-hw FiveQubit1.0001.0001.000 Steane1.0000.9940.995 Planar3×3 0.9990.9380.940 Planar5×5 0.3970.2700.308 shows that the pooled decoder transfers well across known settings, but still faces a clear representation gap when the syndrome-recovery map changes in a structurally new way. Thus, few-shot adaptation to an unseen code exposes the clearest limitation of raw teacher-label transfer. The model retains strong performance on known code structures, but direct adaptation to Planar5×5 remains weak under limited support data. 4.7 Few-shot adaptation to held-out code size The held-out-size setting addresses a related but distinct ques- tion: whether the meta-decoder can adapt to a new code scale using only a small support set. As shown in Table 4, this regime is harder than unseen-푝and unseen-noise trans- fer, but easier than few-shot unseen-code adaptation. The Meta-MLP reaches 0.7548, while the accuracy-selected and hardware-aware VQCs reach 0.6774 and 0.7143, respectively. The hardware-aware VQC again exceeds the accuracy-selected VQC in this setting, suggesting that the hardware-aware crite- rion can select compact architectures that remain competitive under transfer. The code-wise breakdown in Table 9 shows that the held-out- size result combines easy and difficult cases. FiveQubit, Steane, and Planar3×3 remain strong, while Planar5×5 remains the main bottleneck. In particular, the Meta-MLP reaches 0.406 on Planar5×5, while the accuracy-selected and hardware-aware VQCs reach 0.266 and 0.321, respectively. This pattern is also visible in Fig. 9(e), where the larger planar-code setting remains the main source of difficulty. Thus, held-out-size adaptation lies between unseen-noise transfer and unseen-code adaptation in difficulty. The results confirm that structural changes in code scale are harder than shifts in푝or noise family, especially for the larger topological- code setting. 4.8 Logical-level reliability and selective fallback Teacher-label accuracy is a useful first measure of meta-decoder performance, but it is not equivalent to logical reliability. In QEC decoding, a small number of incorrect recovery labels can still induce logical failures, particularly in larger topological- code settings. We therefore evaluate the learned decoders at the logical level and test whether confidence-gated fallback can reduce the gap between raw learned recovery and teacher- decoder behavior. The logical-failure curves in Fig. 10 show the resulting de- coder behavior across the five evaluation settings. Planar5×5 Table 9. Average code-wise accuracy under few-shot held-out-size adaptation. CodeMeta-MLP VQC best-acc VQC best-hw FiveQubit1.0000.9761.000 Steane1.0000.9830.989 Planar3×3 1.0000.9070.952 Planar5×5 0.4060.2660.321 0.020.040.060.080.10 Physical error probability, p 10 −3 10 −2 10 −1 Logical failure rate, p L largest regime shift Transfer setting Interpolation Holdout-p Holdout-code Holdout-size low-p gap opens transition / elbow high-p convergence Figure 10. Combined logical-failure curves under the confidence- gated hybrid recovery protocol. is the most demanding regime: raw learned decoding re- mains far from the teacher decoder, while confidence-gated fallback reduces the logical gap by using the learned decoder only on confident cases and routing uncertain cases to the teacher decoder (Additional Planar5×5 curves are reported in Appendix A). Table 10 quantifies the raw and fallback-enabled logical behavior on Planar5×5. In interpolation, the Meta-MLP reaches nearly perfect teacher-label accuracy, yet its raw logical- failure ratio remains 12.08 relative to the teacher decoder. The raw VQC ratios are even larger, at 42.41 for the accuracy- selected VQC and 25.91 for the hardware-aware VQC. After fallback, these ratios drop to 1.71, 1.04, and 1.11, respectively. The same pattern appears across the transfer settings: raw learned decoding is not a reliable teacher replacement, while fallback brings the learned decoders much closer to teacher- level logical behavior. To make the fallback mechanism visible at the operating point used in the experiments, Table 11 reports fixed-threshold learned coverage and fallback rates on Planar5×5. These quantities are computed from saved test-set confidence values, not from new simulations. They measure how often the learned decoder is used at the reported thresholds, while Table 10 reports the corresponding Monte Carlo logical behavior. The coverage results in Table 11 explain the source of this logical improvement. At휏= 0.75, the VQC models are conservative on Planar5×5 and often route most transfer cases to the teacher decoder. For example, in the few-shot regimes, their learned coverage stays near 0.16, while the confident label accuracy remains much higher than the full label accuracy. This indicates that the confidence signal is useful for selective decoding: the learned decoder is not used broadly, but the cases it retains are substantially more reliable. Finally, Fig. 11 summarizes the learned-decoder advantage across code families and noise models. For each code–noise pair, the score is computed as 푆 푐,푛 = max 푚 * log 10 푝 baseline 퐿 푝 푚 퐿 !+ 푡,푝 ,(27) where푝 baseline 퐿 is the logical-failure rate of the code-specific conventional decoder,푝 푚 퐿 is the logical-failure rate of the learned decoder푚, and the average is taken over transfer 10 MDQEC-QASChoudhary et al. Table 10. Mean Planar5×5 logical-failure ratio relative to the teacher decoder across the five settings. Values near 1 indicate teacher-level logical behavior. SettingDecoderRaw ratio Fallback ratio Reduction factor InterpolationMeta-MLP12.081.717.06 InterpolationVQC best-acc42.411.0440.88 InterpolationVQC best-hw25.911.1123.24 unseen-푝Meta-MLP14.611.927.61 unseen-푝VQC best-acc32.621.0929.94 unseen-푝VQC best-hw35.151.0633.03 Unseen-noiseMeta-MLP14.521.778.19 Unseen-noiseVQC best-acc40.391.0438.83 Unseen-noiseVQC best-hw43.273.1713.66 Few-shot unseen-code Meta-MLP32.273.768.59 Few-shot unseen-code VQC best-acc65.801.0960.25 Few-shot unseen-code VQC best-hw50.851.0548.55 Few-shot held-out-size Meta-MLP30.763.209.63 Few-shot held-out-size VQC best-acc54.461.1149.27 Few-shot held-out-size VQC best-hw58.131.0953.23 Correlated Z burst DepolarizingDepol. + meas. flip Z-biased Noise model FiveQubit Steane Planar3x3 Planar5x5 Code family -0.000 near parity MLP +0.030 learned better MLP -0.003 near parity MLP +0.009 near parity MLP +0.019 learned better VQC-hw -0.004 near parity VQC-hw +0.018 learned better VQC-hw +0.018 learned better VQC-hw -0.028 base better MLP -0.022 base better MLP -0.000 near parity MLP -0.036 base better MLP +0.006 near parity VQC-hw +0.009 near parity VQC-acc -0.005 near parity VQC-hw -0.051 base better VQC-acc -0.06 -0.03 0 +0.03 +0.06 S c , n = max m ⟨log 10 ( p base L / p m L )⟩ learned better base better Figure 11. Computed decoder-advantage summary across code families and noise models. Each cell reports the best learned-decoder advantage relative to the corresponding baseline decoder, averaged over transfer settings and physical error probabilities. Positive values indicate lower logical-failure rates for the best learned decoder, whereas negative values indicate that the baseline decoder remains superior. settings푡and physical error probabilities푝. The baseline is the Naive decoder for FiveQubit and Steane, and MWPM for the planar codes. Positive values indicate lower logical- failure rates for the best learned decoder, while negative values indicate that the conventional decoder remains better. The heatmap shows that learned decoding does not provide a universal advantage; instead, its value depends on the code family and noise model. This supports the intended use of the framework: selective learned assistance with trusted fallback, rather than universal replacement of established decoders. 4.9 Discussion The results show that pooled meta-decoding is feasible, but its reliability depends on how the decoder is evaluated. At the teacher-label level, the pooled representation is highly learnable: the Meta-MLP nearly solves interpolation and remains robust under unseen-푝and unseen-noise transfer. Transfer becomes harder when the code structure or code size changes, with Planar5×5 consistently emerging as the main bottleneck. This suggests that shifts in error rate or noise family are easier to absorb than changes in the underlying syndrome-recovery map. The VQC meta-decoders do not surpass the Meta-MLP in raw accuracy, so they should not be interpreted as accuracy winners. Their role is instead to test compact quantum meta- decoders under different selection criteria. The hardware- aware score changes the selected circuit architecture while keeping final accuracy close to the accuracy-selected VQC in most regimes. This supports hardware-aware VQC selection as an architecture-selection tool within the simulated search space. This interpretation is consistent with QAS as a model-selection problem rather than an automatic performance guarantee [20, 21]. The key reliability finding is that teacher-label accuracy and logical reliability are not equivalent. In the hardest Planar5×5 setting, raw learned decoding can remain far from teacher-level logical behavior even when classification accuracy appears acceptable. Confidence-gated fallback changes this behavior by using the learned decoder only on high-confidence cases and routing uncertain cases to the teacher decoder. The resulting system is therefore a selective learned-assisted decoder, not a full replacement for MWPM or other trusted decoders. This supports a QEC-decoder perspective in which logical-level behavior, rather than supervised prediction accuracy alone, guides decoder assessment [9, 15]. These results define the scope of the proposed framework: hardware-aware VQC selection within a compact simulated search space, and reliability-aware learned decoding through confidence-gated fallback. 5 Conclusion This work presented a selective meta-decoding framework for quantum error correction that combines pooled multi- code learning, hardware-aware VQC architecture search, and confidence-gated hybrid recovery. The framework studies transfer across code, noise, and error-rate regimes instead of training a separate learned decoder for each setting. The results show that pooled teacher-label decoding can transfer across several QEC settings. The Meta-MLP provides the strongest raw accuracy, while the VQC meta-decoders learn transferable structure above the majority baseline and support hardware-aware circuit selection within the simulated search space. At the logical level, the main finding is that teacher-label accuracy alone is not sufficient for decoder assessment. In the hardest Planar5×5 regime, confidence-gated fallback re- duces the gap between raw learned decoding and teacher-level logical behavior by using learned recovery only on high- confidence cases. The resulting method is therefore a selective learned-assisted decoder rather than an unconditional replace- ment for established decoders. Future work can extend this framework through calibrated confidence thresholds, broader circuit-search spaces, and implementation-level cost measure- ments. Acknowledgments This work was supported by the HRD Group of the Council of Scientific & Industrial Research (CSIR) provided by the CSIR Research Fellowships, and in part by the NYUAD Center for Quantum and Topological Systems (CQTS), funded by Tamkeen under the NYUAD Research Institute grant CG008. The authors also acknowledge the National Supercomput- ing Mission (NSM) for providing computing resources of “PARAM Shivay” at the Indian Institute of Technology (BHU), Varanasi, which is implemented by C-DAC and supported by the Ministry of Electronics and Information Technology (MeitY) and Department of Science and Technology (DST), Government of India. References [1]Peter W. Shor. Scheme for reducing decoherence in quantum computer memory. Physical Review A, 52(4): 11 MDQEC-QASChoudhary et al. Table 11. Planar5×5 selective-decoding coverage computed from the saved test-set confidence values. Meta-MLP uses휏= 0.80and VQC models use휏= 0.75, matching the reported fallback protocol. Confident error is measured against the teacher recovery label, not against logical equivalence; it is therefore a teacher-label proxy rather than a logical-failure rate. SettingModel휏 Coverage Fallback Label acc. Conf. label acc. Conf. label err. InterpolationMeta-MLP0.800.9989 0.00110.99920.99980.0002 InterpolationVQC best-acc 0.750.3335 0.66650.74630.99940.0006 InterpolationVQC best-hw 0.750.6049 0.39510.88540.99860.0014 Unseen-푝Meta-MLP0.800.6451 0.35490.63450.96730.0327 Unseen-푝VQC best-acc 0.750.2132 0.78680.47340.99410.0059 Unseen-푝VQC best-hw 0.750.3156 0.68440.49260.99230.0077 Unseen-noiseMeta-MLP0.800.7361 0.26390.72570.98170.0183 Unseen-noiseVQC best-acc 0.750.2499 0.75010.54190.99820.0018 Unseen-noiseVQC best-hw 0.750.2431 0.75690.49970.98190.0181 Few-shot unseen-code Meta-MLP0.800.4204 0.57960.39710.91700.0830 Few-shot unseen-code VQC best-acc 0.750.1585 0.84150.26970.96390.0361 Few-shot unseen-code VQC best-hw 0.750.1629 0.83710.30850.98920.0108 Few-shot held-out-size Meta-MLP0.800.4093 0.59070.40640.94630.0537 Few-shot held-out-size VQC best-acc 0.750.1561 0.84390.26600.97370.0263 Few-shot held-out-size VQC best-hw 0.750.1616 0.83840.32080.99200.0080 R2493–R2496, 1995. doi: 10.1103/PhysRevA.52.R249 3. [2] A. M. Steane. Multiple-particle interference and quantum error correction. Proceedings of the Royal Society A, 452 (1954):2551–2577, 1996. doi: 10.1098/rspa.1996.0136. [3]Daniel Gottesman. Stabilizer Codes and Quantum Error Correction. PhD thesis, California Institute of Technol- ogy, 1997. Ph.D. thesis. [4]A. Yu. Kitaev. Fault-tolerant quantum computation by anyons. Annals of Physics, 303(1):2–30, 2003. doi: 10.1016/S0003-4916(02)00018-0. [5]Barbara M. Terhal. Quantum error correction for quantum memories. Reviews of Modern Physics, 87(2):307–346, 2015. doi: 10.1103/RevModPhys.87.307. [6]Eric Dennis, Alexei Kitaev, Andrew Landahl, and John Preskill. Topological quantum memory. Journal of Mathematical Physics, 43(9):4452–4505, 2002. doi: 10.1063/1.1499754. [7]Austin G. Fowler, Matteo Mariantoni, John M. Martinis, and Andrew N. Cleland. Surface codes: Towards practical large-scale quantum computation. Physical Review A, 86 (3):032324, 2012. doi: 10.1103/PhysRevA.86.032324. [8]Earl T. Campbell, Barbara M. Terhal, and Christophe Vuillot. Roads towards fault-tolerant universal quantum computation. Nature, 549(7671):172–179, 2017. doi: 10.1038/nature23460. [9] Antonio de Marti i Olius, Pau Fuentes, Román Orús, Pere M. Crespo, and José Etxezarreta Martinez. Decoding algorithms for surface codes. Quantum, 8:1498, 2024. doi: 10.22331/q-2024-10-10-1498. [10]Nicolas Delfosse and Naomi H. Nickerson. Almost-linear time decoding algorithm for topological codes. Quantum, 5:595, 2021. doi: 10.22331/q-2021-12-02-595. [11] Christopher Chamberland, Guanyu Zhu, Tomas Jochym- O’Connor, and Aleksander Kubica. Machine-learning methods for quantum error correction: A survey. AVS Quantum Science, 5(4):041101, 2023. doi: 10.1116/5.01 66514. [12]Johannes Bausch, A. W. Senior, Francisco J. H. Heras, Tristan Edlich, Alexander Davies, Max Newman, Christo- pher Brej, Hieu Pham, Livia de Oliveira, Charlotte Papa- georgiou, Ioannis Karamcheti, Sagar Joglekar, Simon P. Chao, Mohammad Babaei, Neil Houlsby, Joshua Izaac, Stefano Carrazza, and Pushmeet Kohli. Learning high- accuracy error decoding for quantum processors. Nature, 636:328–334, 2024. doi: 10.1038/s41586-024-08148-8. [13]Ants Remm, Nathan Lacroix, Lukas Bödeker, Elie Genois, Christoph Hellings, Fran çois Swiadek, Graham J. Norris, Christopher Eichler, Alexandre Blais, Markus Müller, Sebastian Krinner, and Andreas Wallraff. Exper- imentally informed decoding of stabilizer codes based on syndrome correlations. Phys. Rev. Res., 8:013044, Jan 2026. doi: 10.1103/z1ng-wg3k. URLhttps: //link.aps.org/doi/10.1103/z1ng-wg3k. [14] Paul Baireuther, Thomas E. O’Brien, Brian Tarasinski, and Carlo W. J. Beenakker. Machine-learning-assisted correction of correlated qubit errors in a topological code. Quantum, 2:48, 2018. doi: 10.22331/q-2018-01-29-48. [15]Alex Fischer and Akimasa Miyake. Hardness results for decoding the surface code with pauli noise. Quantum, 8:1511, 2023. URLhttps://api.semanticschola r.org/CorpusID:262054085. [16] Karl Hammar, Alexei Orekhov, Patrik Wallin Hybelius, Anna Katariina Wisakanto, Basudha Srivastava, An- ton Frisk Kockum, and Mats Granath. Error-rate-agnostic decoding of topological stabilizer codes. Physical Review A, 105(4):042616, 2022. doi: 10.1103/PhysRevA.105.0 42616. 12 MDQEC-QASChoudhary et al. [17]Boris M. Varbanov, Marc Serra-Peralta, David Byfield, and Barbara M. Terhal. Neural network decoder for near-term surface-code experiments. Phys. Rev. Res., 7: 013029, Jan 2025. doi: 10.1103/PhysRevResearch.7.01 3029. URLhttps://link.aps.org/doi/10.1103 /PhysRevResearch.7.013029. [18]Vojtěch Havlíček, Antonio D. Córcoles, Kristan Temme, Aram W. Harrow, Abhinav Kandala, Jerry M. Chow, and Jay M. Gambetta. Supervised learning with quantum- enhanced feature spaces. Nature, 567(7747):209–212, 2019. doi: 10.1038/s41586-019-0980-2. [19]Maria Schuld, Alex Bocharov, Krysta M. Svore, and Nathan Wiebe. Circuit-centric quantum classifiers. Phys- ical Review A, 101(3):032308, 2020. doi: 10.1103/Phys RevA.101.032308. [20] Yuxuan Du, Tao Huang, Shan You, Min-Hsiu Hsieh, and Dacheng Tao. Quantum circuit architecture search for variational quantum algorithms. npj Quantum Informa- tion, 8(1):62, 2022. doi: 10.1038/s41534-022-00590-1. [21] Darya Martyniuk, Johannes Jung, and Adrian Paschke. Quantum architecture search: a survey. In 2024 IEEE International Conference on Quantum Computing and Engineering (QCE), volume 1, pages 1695–1706. IEEE, 2024. [22] Prashant Kumar Choudhary, Nouhaila Innan, Muham- mad Shafique, and Rajeev Singh. Graph-based bayesian optimization for quantum circuit architecture search with uncertainty calibrated surrogates, 2025. URL https://arxiv.org/abs/2512.09586. [23]Alberto Marchisio, Muhammad Kashif, Nouhaila Innan, and Muhammad Shafique. Hybrid quantum-classical neu- ral architecture search. arXiv preprint arXiv:2605.18345, 2026. [24]Muhammad Kashif, Shaf Khalid, Alberto Marchisio, Nouhaila Innan, and Muhammad Shafique. Faqnas: Flops-aware hybrid quantum neural architecture search using genetic algorithm. In 2026 Design, Automation & Test in Europe Conference (DATE), pages 1–7. IEEE, 2026. [25]Siddhant Dutta, Nouhaila Innan, Sadok Ben Yahia, and Muhammad Shafique. Qas-qtns: Curriculum reinforce- ment learning-driven quantum architecture search for quantum tensor networks. In 2025 IEEE International Conference on Quantum Computing and Engineering (QCE), volume 1, pages 1739–1747. IEEE, 2025. [26]Jiahao Pan and Hyeon Kim. Artificial intelligence for quantum error correction: Opportunities and challenges. AVS Quantum Science, 6(3):030801, 2024. doi: 10.111 6/5.0214974. [27]Lucas Berent, Lukas Burgholzer, Peter-Jan H.S. Derks, Jens Eisert, and Robert Wille. Decoding quantum color codes with maxsat. Quantum, 8:1506, October 2024. ISSN 2521-327X. doi: 10.22331/q-2024-10-23-1506. URLhttp://dx.doi.org/10.22331/q-2024-1 0-23-1506. [28]Giacomo Torlai and Roger G. Melko. Neural decoder for topological codes. Physical Review Letters, 119(3): 030501, 2017. doi: 10.1103/PhysRevLett.119.030501. [29] Savvas Varsamopoulos, Ben Criger, and Koen Bertels. Decoding small surface codes with feedforward neu- ral networks. Quantum Science and Technology, 3(1): 015004, 2018. doi: 10.1088/2058-9565/a955a. [30]Philip Andreasson, Joel Johansson, Simon Liljestrand, and Mats Granath. Quantum error correction for the toric code using deep reinforcement learning. Quantum, 3: 183, 2019. doi: 10.22331/q-2019-09-02-183. [31]Moritz Lange, Pontus Havström, Basudha Srivastava, Valdemar Bergentall, Karl Hammar, Olivia Heuts, Evert P. L. van Nieuwenburg, and Mats Granath. Data-driven decoding of quantum error correcting codes using graph neural networks. Physical Review Research, 7(2):023181, 2025. doi: 10.1103/PhysRevResearch.7.023181. [32]Adrián Pérez-Salinas, Alba Cervera-Lierta, Elies Gil- Fuster, and José I. Latorre. Data re-uploading for a universal quantum classifier. Quantum, 4:226, 2020. doi: 10.22331/q-2020-02-06-226. [33]Iris Cong, Soonwon Choi, and Mikhail D. Lukin. Quan- tum convolutional neural networks. Nature Physics, 15 (12):1273–1278, 2019. doi: 10.1038/s41567-019-064 8-8. [34]Nouhaila Innan, Owais Ishtiaq Siddiqui, Shivang Arora, Tamojit Ghosh, Yasemin Poyraz Koçak, Dominic Para- gas, Abdullah Al Omar Galib, Muhammad Al-Zafar Khan, and Mohamed Bennai. Quantum state tomography using quantum machine learning. Quantum Machine Intelligence, 6(1):28, 2024. [35] Nouhaila Innan, Bikash K Behera, Saif Al-Kuwari, and Ahmed Farouk. Qnn-vrcs: A quantum neural network for vehicle road cooperation systems. IEEE Transactions on Intelligent Transportation Systems, 2025. [36]Nouhaila Innan, Alberto Marchisio, Mohamed Bennai, and Muhammad Shafique. Lep-qnn: Loan eligibility prediction using quantum neural networks. In 2025 IEEE International Conference on Quantum Computing and Engineering (QCE), volume 1, pages 1864–1872. IEEE, 2025. [37]Prashant Kumar Choudhary, Nouhaila Innan, Muham- mad Shafique, and Rajeev Singh. Hqnn-fsp: A hybrid classical-quantum neural network for regression-based fi- nancial stock market prediction. Quantum Machine Intel- ligence, 8(55), 2026. doi: 10.1007/s42484-026-00390-9. [38]Andrea Skolik, Jarrod R. McClean, Masoud Mohseni, Patrick van der Smagt, and Martin Leib. Layerwise learning for quantum neural networks. Quantum Machine Intelligence, 3(1):5, 2021. doi: 10.1007/s42484-020-0 0036-4. [39]M. Cerezo, Akira Sone, Tyler Volkoff, Lukasz Cin- cio, and Patrick J. Coles. Cost function dependent barren plateaus in shallow parametrized quantum cir- cuits. Nature Communications, 12(1):1791, 2021. doi: 10.1038/s41467-021-21728-w. 13 MDQEC-QASChoudhary et al. [40]Danil Vyskubov, Kirill Vyskubov, Nouhaila Innan, and Muhammad Shafique. Scaling laws for hybrid quantum neural networks: Depth, width, and quantum-centric diagnostics. arXiv preprint arXiv:2604.06007, 2026. [41]Xiao-Yu Bi, Yi-Ming Yu, Ye-Hong Chen, and Zhi-Rong Zhong. General-purpose quantum architecture search based on deep reinforcement learning. Phys. Rev. A, 112: 052409, Nov 2025. doi: 10.1103/7rc4-p446. URLhttp s://link.aps.org/doi/10.1103/7rc4-p446. 14 MDQEC-QASChoudhary et al. A Planar5×5 Logical-Failure Curves The Planar5×5 curves presented in Figure 12 provide a focused analysis of the hardest code setting and show how confidence-gated fallback changes logical-failure behavior across the five evaluation regimes. 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (a) Correlated Z burstDepolarizingDepolarizing + meas. flipZ-biased 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (b) 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (c) 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (d) 0.020.040.060.080.10 Physical error probability p 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (e) 0.020.040.060.080.10 Physical error probability p 0.020.040.060.080.10 Physical error probability p 0.020.040.060.080.10 Physical error probability p Planar 5 x 5 Planar MWPMHiMeta-MLPHiMeta-VQC (acc)HiMeta-VQC (hw) 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (a) Correlated Z burstDepolarizingDepolarizing + meas. flipZ-biased 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (b) 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (c) 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (d) 0.020.040.060.080.10 Physical error probability p 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (e) 0.020.040.060.080.10 Physical error probability p 0.020.040.060.080.10 Physical error probability p 0.020.040.060.080.10 Physical error probability p Planar 5 x 5 Planar MWPMHiMeta-MLPHiMeta-VQC (acc)HiMeta-VQC (hw) 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (a) Correlated Z burstDepolarizingDepolarizing + meas. flipZ-biased 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (b) 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (c) 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (d) 0.020.040.060.080.10 Physical error probability p 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (e) 0.020.040.060.080.10 Physical error probability p 0.020.040.060.080.10 Physical error probability p 0.020.040.060.080.10 Physical error probability p Planar 5 x 5 Planar MWPMHiMeta-MLPHiMeta-VQC (acc)HiMeta-VQC (hw) 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (a) Correlated Z burstDepolarizingDepolarizing + meas. flipZ-biased 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (b) 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (c) 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (d) 0.020.040.060.080.10 Physical error probability p 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (e) 0.020.040.060.080.10 Physical error probability p 0.020.040.060.080.10 Physical error probability p 0.020.040.060.080.10 Physical error probability p Planar 5 x 5 Planar MWPMHiMeta-MLPHiMeta-VQC (acc)HiMeta-VQC (hw) 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (a) Correlated Z burstDepolarizingDepolarizing + meas. flipZ-biased 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (b) 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (c) 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (d) 0.020.040.060.080.10 Physical error probability p 10 −4 10 −3 10 −2 10 −1 Logical failure rate p L (e) 0.020.040.060.080.10 Physical error probability p 0.020.040.060.080.10 Physical error probability p 0.020.040.060.080.10 Physical error probability p Planar 5 x 5 Planar MWPMHiMeta-MLPHiMeta-VQC (acc)HiMeta-VQC (hw) Figure 12. Additional Planar5×5 logical-failure curves under the confidence-gated hybrid recovery protocol. Panels (a)–(e) correspond to interpolation, unseen-푝 transfer, unseen-noise transfer, few-shot unseen-code adaptation, and few-shot held-out-size adaptation, respectively. 15