Paper deep dive
When Routes Run Out: Adversarial Co-Learning and Explainable Robustness in Quantum Repeater Networks
Brennan Bell, Inti Gabriel Mendoza Estrada, Andreas Trügler, Paul Erker
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 7/13/2026, 6:23:42 AM
Summary
This paper investigates adversarial co-learning for entanglement-based quantum repeater network routing using an adversarial bandit framework. The Exp3 algorithm simulates Alice selecting Ekert-91 (E91) routes and Eve selecting attack surfaces (intercept-resend or memory degradation) across 50 graph topologies. Results demonstrate that learned strategies closely track minimax references (Pearson r=0.99), with bottleneck topologies yielding zero retention and non-bottleneck topologies following a 1-1/N coverage principle. The study further evaluates decision-tree and local LLM-based explainability for the learned strategies, establishing an open-source workflow for interpreting quantum network games.
Entities (10)
Relation Signals (8)
Alice → selectsroutefor → Ekert-91 protocol
confidence 97% · Alice selects an end-to-end repeater route for an Ekert-91 protocol (E91) representing her move
Bottleneck topology → exhibits → Zero retention
confidence 96% · bottleneck families have zero retention, while non-bottleneck families follow a 1−1/N coverage principle
Eve → selectsattacksurface → Intercept-resend attack
confidence 96% · Eve selects an attack surface, either edge intercept–resend or repeater memory degradation
Non-bottleneck topology → follows → 1-1/N coverage principle
confidence 95% · non-bottleneck families follow a 1−1/N coverage principle
Exp3 algorithm → usedfor → Adversarial co-learning
confidence 95% · Performing adversarial co-learning across 50 structured topologies... Both players use Exp3
SeQUeNCe → simulates → Quantum repeater network
confidence 94% · SeQUeNCe is a discrete-event simulator of quantum networks with explicit sources, fibers, memories, and control protocols
decision tree → explains → Learned strategies
confidence 93% · We then fit decision-tree explanation models to graph-, attack-, and route-level topology-corpus targets and report their faithfulness
Ollama → serves →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We study an adversarial bandit problem for entanglement-based quantum-network routing over a modest graph corpus. Alice selects an end-to-end repeater route for an Ekert-91 protocol (E91) representing her move, while Eve selects an attack surface, either edge intercept--resend or repeater memory degradation. Payoffs are drawn from cached SeQUeNCe-simulated E91 transcripts, and Alice accepts a turn when the finite-sample statistic violates the Clauser-Horne-Shimony-Holt (CHSH) bound. Performing adversarial co-learning across 50 structured topologies, we find that learned retention tracks a full-matrix minimax reference closely (Pearson $r=0.99$): under a one-surface Eve action model, bottleneck families have zero retention, while non-bottleneck families follow a $1-1/N$ coverage principle. We then fit decision-tree explanation models to graph-, attack-, and route-level topology-corpus targets and report their faithfulness. Finally, we construct prompt records for local language models to summarize the tree evidence, resulting in an open-source explanation workflow for quantum-repeater network games.
Tags
Links
- Source: https://arxiv.org/abs/2607.09378v1
- Canonical: https://arxiv.org/abs/2607.09378v1
Trouble viewing inline? Open PDF directly →
Full Text
21,462 characters extracted from source content.
Expand or collapse full text
When Routes Run Out: Adversarial Co-Learning and Explainable Robustness in Quantum Repeater Networks Brennan Bell ∗ , Inti Gabriel Mendoza Estrada † , Andreas Tr ̈ ugler ‡ , and Paul Erker § Abstract—We study an adversarial bandit problem for entanglement-based quantum-network routing over a modest graph corpus. Alice selects an end-to-end repeater route for an Ekert-91 protocol (E91) representing her move, while Eve selects an attack surface, either edge intercept–resend or repeater mem- ory degradation. Payoffs are drawn from cached SeQUeNCe- simulated E91 transcripts, and Alice accepts a turn when the finite-sample statistic violates the Clauser-Horne-Shimony-Holt (CHSH) bound. Performing adversarial co-learning across 50 structured topologies, we find that learned retention tracks a full-matrix minimax reference closely (Pearson r = 0.99): under a one-surface Eve action model, bottleneck families have zero retention, while non-bottleneck families follow a 1−1/N coverage principle. We then fit decision-tree explanation models to graph- , attack-, and route-level topology-corpus targets and report their faithfulness. Finally, we construct prompt records for local language models to summarize the tree evidence, resulting in an open-source explanation workflow for quantum-repeater network games. Index Terms—quantum, E91, CHSH, adversarial, Exp3, ban- dits, games, networks, repeaters, XAI, simulations, security I. INTRODUCTION A. Research questions The present works analyses two main questions. Can a standard adversarial bandit, playing from bandit feedback with no topology knowledge, recover the strategic structure of route selection and attack placement in a simulated Ekert-91 (E91) repeater network? Can the learned and reference strategies then be explained—by symbolic representations and small local language models—with measured faithfulness and without security-oriented hallucinations? B. Motivation In E91, Alice and Bob consume entangled pairs and use Bell-test statistics as a security monitor [1], [2], [3]. On a repeater network, route choice also induces a classical interdiction game: Alice’s mixed strategy determines which fibers and memories are exposed, while a limited Eve chooses one component to attack. This route-versus-surface structure is well known in zero-sum network interdiction [4]. In the ∗ RFI-IRFOS and TU Graz, Graz, Austria; bell.brennan.p@gmail.com † openmaindFlexCoandTUGraz,Graz,Austria; inti.mendoza@openmaind.ai ‡ Know Center Research GmbH and University of Graz, Graz, Austria; andreas.truegler@uni-graz.at § Atominstitut,TUWienandIQOQI, ̈ OAW,Vienna,Austria; paul.erker@tuwien.ac.at homogeneous case, a component on every Alice–Bob route exposes a fatal bottleneck to Eve, i.e. with N component- disjoint routes and one attacked surface, uniform routing gives hit probability 1/N , hence a retention of 1− 1/N . The question is therefore not whether an algorithm can rediscover this coverage logic in a binary graph game. It is whether the same strategic structure remains visible when pay- offs are generated by topology-corpus E91 transcripts: CHSH- gated acceptance, shared repeater memories, unattacked fail- ures, and component attacks. Explanations matter because simulator diagnostics are easy to over-interpret as quantum key distribution (QKD) security claims. C. Related work SeQUeNCe and NetSquid make quantum-network hardware and control assumptions explicit [5], [6]; recent trapped- ion network nodes motivate high-efficiency parameter scales [7]. Exponential-weight algorithm for Exploration and Ex- ploitation (Exp3) is the standard adversarial bandit [8]; in zero-sum games, no-regret learning supports time-averaged strategies, while terminal multiplicative-weight iterates may fail to converge [9], [10]. For explanation, we use shallow decision trees with reported target fit [11], [12]. (a) single bottleneck AB (b) disjoint parallel (N= 8) AB Fig. 1. Exemplar topologies. (a) In a single-bottleneck graph every route crosses one node (red): Eve can deny all traffic from one surface. (b) In a disjoint-parallel graph one attack surface covers only one of N routes. I. MODELING AND METHODOLOGY A. Quantum repeater network simulation SeQUeNCe [5] is a discrete-event simulator of quantum networks with explicit sources, fibers, memories, and control arXiv:2607.09378v1 [quant-ph] 10 Jul 2026 protocols; SeQUeNCe’s repeater module is Barrett–Kok entan- glement generation [13] plus entanglement swapping. In our E91 game turn, a budget of 350 entangled pairs is distributed end-to-end over the chosen route; delivered pairs are measured at random E91 settings, a subset is used for the Bell statistics to compute S [3], and matched basis pairs yield sifted key bits. The public outcome of a turn is one ofaccepted, CHSH abort, quantum bit error rate (QBER) abort, delivery failure, with the CHSH check preceding the QBER check. B. Minimax reference and Exp3 For each graph, cached trials define Alice’s payoff matrix M ∈ [0, 1] |Π|×|E| . Let X = ∆(Π) and Y = ∆(E ). The full- matrix reference is v ⋆ = max x∈X min y∈Y x T My, R oracle = clip [0,1] (v ⋆ /v 0 ), (1) where v 0 is the best no-attack value. We measure learned averaged strategies relative to the Nash equilibrium gap g( ̄x, ̄y) = max x∈X x T M ̄y− min y∈Y ̄x T My,(2) which is zero at a saddle point of this empirical zero-sum game. Both players use Exp3 [8]: action probabilities mix nor- malized weights with uniform exploration, played rewards are importance-weighted, and weights are updated multiplica- tively. We use a learning rate of η t = p logK/(K(t + 10 4 )) and an exploration parameter γ t = min0.2,Kη t , where K is the action-space size and t is the game-turn. Since no-regret guarantees concern averaged play, all reported strategies are frequency-of-play averages rather than terminal iterates [9], [10]. I. EXPERIMENTS A. Implementation choices The corpus holds 50 graphs from eight structural fami- lies (single/multi/deep bottleneck, deep/disjoint/layered par- allel, length-variant disjoint, Wheatstone chains), with up to 32 routes of at most 7 hops and edge lengths 400–550 m. Fig. 1 shows the primary structural motifs. Eve’s action set is no attack∪E ir ∪E mem : one intercept–resend action per eligi- ble edge and one degradation action per internal repeater mem- ory. Intercept–resend is the eavesdropper whose intermediate measurements destroy the CHSH violation [2], while memory degradation models repeater node integrity/availability dis- turbances [14]. Practical QKD attacks such as detector tim- ing [15] and Trojan-horse probing [16] motivate component- specific threat modelling, but those detector/source side chan- nels are not implemented in this experiment. To make 5× 10 5 -turn co-learning feasible, Exp3 turns are sampled from caches of pre-simulated SeQUeNCe E91 trials. We store 64 trials per route–hop and per attack–hop profile; during training, we sample the matching pool, while the complete-information matrix uses the mean Alice acceptance over 16 cached draws per cell. Alice’s payoff is CHSH-only: after forming the Bell-test sample, a turn is accepted if and only if |S| > 2. QBER is logged but not used in the payoff, since the CHSH test dominates the statistics in this simulation. Eve payoff is binary, whether or not her attack lies on Alice’s route. We use a memory fidelity of 0.98 and, based on recent multiplexed trapped-ion network results [7], a memory effi- ciency of 0.544; however, we apply them to a SeQUeNCe Barrett–Kok model; this is a deliberate compromise between accurate device modelling and realistic repeater performance. A pre-run health check expects at least 319 of 350 entangled pairs to be delivered, where 4 of 9 combinations feed the CHSH test for Alice and 2 of 9 combinations feed the key bits. Without unattacked failures, the induced game would be zero-sum; CHSH failures occur on roughly a fifth of clean turns, making the game only approximately zero-sum. B. Control and target outcomes As a control, we aggregate the caches by route depth. Clean baselines violate the CHSH bound at every depth: mean |S| falls from 2.75 at one hop to 2.37 at seven hops. Active hits, by contrast, never violate the CHSH bound. QBER adds no independent signal: any attack disruptive enough to corrupt the key already breaks the CHSH violation, and the Bell check occurs first. The target outcome is the strategic structure. Fig. 2 shows that the minimax reference is nearly determined by two features. Non-bottleneck families track the coverage principle 1 − 1/N in the node-disjoint path count N , while every bottleneck family collapses to zero retention: Eve parks on the unavoidable cut. Exp3 recovers this reference value. Final learned retention, measured as mean acceptance over the last 2.5 × 10 4 turns and normalized by v 0 , matches oracle retention at Pearson r = 0.99 across the 50 graphs. Five high- redundancy graphs sit slightly above one, with maximum ratio 1.11, as an artifact of the unclipped finite-sample ratio. 12481632 node-disjoint paths N 0.0 0.2 0.4 0.6 0.8 1.0 oracle retention 1 − 1/N bottleneck families: retention ≈ 0 deep bottleneck deep parallel disjoint parallel layered parallel length variant multi bottleneck single bottleneck wheatstone chain Fig. 2.Oracle retention versus node-disjoint path count N for all 50 topologies (log 2 axis; coincident points jittered). Non-bottleneck families track 1− 1/N ; bottleneck families sit at zero. none edge IR memory disjoint parallel wheatstone chain single bottleneck layered parallel deep parallel multi bottleneck length variant deep bottleneck .00.64.36 .00.23.77 .00.75.24 .00.32.68 .00.52.48 .00.24.76 .00.67.33 .00.34.66 (a) Eve: action type 123456789+ route rank (by learned probability) .28.27.15.06.03.03.03.02.14 .29.26.08.07.03.02.02.02.20 .38.31.12.10.03.02.02.01.00 .15.13.11.09.08.08.06.05.23 .42.42.17.00.00.00.00.00.00 .10.07.07.07.06.06.05.05.48 .16.15.15.15.07.07.06.06.14 .31.22.20.15.04.04.03.02.00 (b) Alice: route rank 0.0 0.2 0.4 0.6 0.8 1.0 mean learned probability mass Fig. 3. Learned (time-averaged) Exp3 strategies aggregated by family. (a) Eve’s probability mass by action type. (b) Alice’s mass by route rank (routes sorted by learned probability per graph; “9+” sums the tail). Cell labels give the mean mass. Strategies in Fig. 3 show Eve rarely plays no-attack, and her surface choice is structural: intercept–resend on the cut edge in single-bottleneck graphs, as in Fig. 1a, and memory degradation in multi-bottleneck, deep-bottleneck, and Wheat- stone families, where one shared repeater memory covers several routes. Alice mirrors this structure: she spreads nearly uniformly over disjoint routes, but learns a long-tailed hedge on the large multi-bottleneck route sets, with 0.48 mass beyond rank 8. C. Connections to theory The Nash gap of the time-averaged Exp3 strategies in Fig. 4 shows the median falls from 2.6 × 10 −2 at 10 5 turns to 1.3 × 10 −2 at 5 × 10 5 turns, consistent with the no-regret view that zero-sum multiplicative-weight dynamics should be evaluated through averaged play rather than terminal iterates [8], [9], [10]. The non-monotonicity is empirical: payoff cells are estimated from discrete cached trials, best responses can switch under small changes in the average mixture, and some unattacked routes fail the finite CHSH test. The final gap grows with joint action-space size K A K E (log–log r = 0.75), but this is not topological vulnerability; high-redundancy non- bottleneck graphs retain the expected 1− 1/N minimax value, while the residual gap measures horizon imbalance in the learned mixture. D. Decision-tree explanation models and LLM interpretation We fit depth-3 decision-tree regressors to topology-corpus targets; Table I reports their faithfulness, which spans the full quality range. The graph-retention tree is nearly exact (R 2 = 0.9984): its active rules use bottleneck presence and disjoint-path count, so graph-level tree summaries are faithful to our topology-corpus target. The Eve-action tree is moderately faithful (R 2 = 0.7048 over 1501 rows), led by target-route coverage and the memory-degradation flag, consistent with the coverage logic of Fig. 3a. The expected- denial and Alice-route trees are weak (R 2 = 0.43 and 0.24), so route-level tree explanations capture the non-uniqueness of the 0100200300400500 turn (×10 3 ) 10 −3 10 −2 10 −1 10 0 Nash gap (averaged strategies) median (50 graphs) 10–90% Fig. 4. Nash-gap decay across all 50 topologies (gray: individual graphs; blue: median and 10–90% band). The median halves between 10 5 and 5×10 5 turns. TABLE I DECISION-TREE FAITHFULNESS TO THE 50 CORPUS TOPOLOGIES. TargetRows Train R 2 Held-out R 2 Held-out MAE Oracle graph retention500.99840.99780.0081 Oracle Eve action probability15010.70480.68860.0257 Expected denial vs. oracle Alice 15010.43440.37810.1305 Oracle Alice route probability7140.24370.17230.0833 oracle route target. Learned-target counterparts fitted for the prompt pack behave analogously (learned retention R 2 = 0.98, learned Eve strategy R 2 = 0.68). These trees are the symbolic object digested by local large language models (LLMs) as part of an otherwise-static prompt requesting a response of at most 300 words and structured into 5 sections: graph interpretation, tree evidence, strategy evidence, caveats, and rubric self-check. The local models were served via Ollama [17]: llama3.1:8b [18], phi4:14b [19], and nemotron-3-super:120b [20]. In Fig. 5, we present scores for the LLM responses with 4 mechanical categories and 1 subjective human category. All LLMs scored near-perfect on the four mechanical checks, i.e. correctly reported the active decision tree leaf and its predicted value. Yet they were unable to inform why the underlying features produced that outcome. No tested model produced responses exceeding a score of 7 for interpretability. IV. CONCLUSION A. Summary of results On a 50-topology corpus of SeQUeNCe E91 transcripts, Exp3 co-learning from pure bandit feedback recovers the minimax strategy structure and equilibrium support: bottle- necks are fatal, retention follows 1− 1/N , learned retention tracks the oracle reference at r = 0.99, and both players’ learned strategies match the coverage logic of the topology. The Nash gap of the averaged strategies decays with horizon, and the residual gap grows with action-space size—finite- horizon learning error, not topological weakness. Decision- tree faithfulness shows graph-level trees are nearly exact, Eve- action trees moderate, and Alice-route trees weak. llama3.1:8b phi4:14b nemotron-3-super:120b-a12b numeric faithfulness scope honesty decision-tree referencing section completeness human interpretability .95.99.92 .961.001.00 .90.921.00 .90.90.80 .51.55.57 Fig. 5. Mean scores for local LLMs (columns). Rubric (rows) evaluation over 50 scored prompt responses per LLM. Rubric criteria: numeric faithfulness (do stated numbers appear verbatim in the prompt), scope honesty (no security or physical-validity over-claims), decision tree referencing (does it refer to the correct tree leaf for its graph), and section completeness (how many of the requested sections are in the response), are automatically verified. The subjective, human interpretability score was evaluated by the authors and cross-referenced against Claude Fable 5 [21] in an independent blind pass. B. Future work and open questions Open directions follow from the topology-corpus design: Alice’s payoff is CHSH acceptance, Eve’s payoff is route collision, and each matrix entry is estimated from cached SeQUeNCe trials. Improvements should distinguish better between denial, acceptance, and information leakage with more-explicit finite-sample and finite-key semantics. The hard- ware fields are controlled simulation parameters, not emulated device physics. By using hardware studies only to motivate the parameter values, we create a non-trivial high-efficiency E91 routing regime. A future improvement should emulate device mappings more precisely. Finally, larger action spaces are less equilibrated over long training runs, and the residual Nash gap exposes horizon imbalance in the learned mixture. A natural next step is to test contextual-bandit variants [22] which exploit route and attack features, rather than treating all actions as uniformly exchangeable under exploration. ACKNOWLEDGMENTS AND CODE AVAILABILITY We acknowledge funding by the Austrian Federal Ministry of Education, Science, and Research via the Austrian Re- search Promotion Agency (Forschungsf ̈ orderungsgesellschaft – FFG) through Quantum Austria project No. 914033 and No. 63956271. This research was co-funded by the European Union (Quantum Flagship project ASPECTS, Grant Agree- ment No. 101080167). The authors acknowledge OpenAI Codex [23] and Claude Fable 5 [21] for repository develop- ment and manuscript editing assistance. The minimal reproducibility artifact for the topology corpus, payoff-cache construction, learning runs, and figure generation is available at the public GitHub repository [24]. The submit- ted experiments correspond to commit ba96b0. REFERENCES [1] J. S. Bell, “On the problem of hidden variables in quantum mechanics,” Rev. Mod. Phys., vol. 38, p. 447–452, Jul 1966. [Online]. Available: https://link.aps.org/doi/10.1103/RevModPhys.38.447 [2] A. K. Ekert, “Quantum cryptography based on Bell’s theorem,” Phys. Rev. Lett., vol. 67, p. 661–663, Aug 1991. [Online]. Available: https://link.aps.org/doi/10.1103/PhysRevLett.67.661 [3] J. F. Clauser, M. A. Horne, A. Shimony, and R. A. Holt, “Proposed experiment to test local hidden-variable theories,” Physical Review Letters, vol. 23, no. 15, p. 880–884, 1969. [4] A. Washburn and K. Wood, “Two-person zero-sum games for network interdiction,” Operations Research, vol. 43, no. 2, p. 243–251, 1995. [5] X. Wu, A. Kolar, J. Chung et al., “SeQUeNCe: A customizable discrete- event simulator of quantum networks,” Quantum Science and Technol- ogy, vol. 6, no. 4, p. 045027, 2021. [6] T. Coopmans, R. Knegjens, A. Dahlberg et al., “NetSquid, a NETwork simulator for QUantum information using discrete events,” Communi- cations Physics, vol. 4, p. 164, 2021. [7] Z.-B. Cui, Z.-Q. Wang, P.-C. Lai et al., “Metropolitan-scale ion-photon entanglement via a quantum network node with hybrid multiplexing en- hancements,” Nature Communications, vol. 17, p. 697, 2026, published online 6 December 2025. [8] P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire, “The non- stochastic multiarmed bandit problem,” SIAM Journal on Computing, vol. 32, no. 1, p. 48–77, 2002. [9] Y. Freund and R. E. Schapire, “Adaptive game playing using multi- plicative weights,” Games and Economic Behavior, vol. 29, no. 1–2, p. 79–103, 1999. [10] J. P. Bailey and G. Piliouras, “Multiplicative weights update in zero-sum games,” in Proceedings of the 2018 ACM Conference on Economics and Computation, ser. EC ’18.New York, NY, USA: Association for Computing Machinery, 2018, p. 321–338. [11] L. Breiman, J. H. Friedman, R. A. Olshen, and C. J. Stone, Classification and Regression Trees.Belmont, CA: Wadsworth International Group, 1984. [12] C. Rudin, “Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead,” Nature Machine Intelligence, vol. 1, no. 5, p. 206–215, 2019. [13] S. D. Barrett and P. Kok, “Efficient high-fidelity quantum computation using matter qubits and linear optics,” Physical Review A, vol. 71, no. 6, p. 060310, 2005. [14] T. Satoh, S. Nagayama, S. Suzuki, T. Matsuo, M. Hajdu ˇ sek, and R. Van Meter, “Attacking the quantum internet,” IEEE Transactions on Quantum Engineering, vol. 2, p. 1–17, 2021, art. no. 4102617. [15] B. Qi, C.-H. F. Fung, H.-K. Lo, and X. Ma, “Time-shift attack in prac- tical quantum cryptosystems,” Quantum Information and Computation, vol. 7, no. 1–2, p. 73–82, 2007. [16] N. Gisin, S. Fasel, B. Kraus, H. Zbinden, and G. Ribordy, “Trojan- horse attacks on quantum-key-distribution systems,” Physical Review A, vol. 73, no. 2, p. 022320, 2006. [17] Ollama, “Ollama,” Software, 2026, accessed Jul. 10, 2026. [Online]. Available: https://github.com/ollama/ollama [18] —, “llama3.1:8b,” Ollama model library, 2024, accessed 2026-07-10. [Online]. Available: https://ollama.com/library/llama3.1:8b [19] —, “phi4:14b,” Ollama model library, 2024, accessed 2026-07-10. [Online]. Available: https://ollama.com/library/phi4:14b [20] —,“nemotron-3-super:120b,”Ollamamodellibrary,2026, accessed 2026-07-10. [Online]. Available: https://ollama.com/library/ nemotron-3-super:120b [21] Anthropic, “Claude fable 5 and claude mythos 5,” Large language model, 2026, accessed 2026-07-10. [Online]. Available: https://w. anthropic.com/news/claude-fable-5-mythos-5 [22] A. Beygelzimer, J. Langford, L. Li, L. Reyzin, and R. E. Schapire, “Contextual bandit algorithms with supervised learning guarantees,” in Proceedings of the Fourteenth International Conference on Artificial In- telligence and Statistics, ser. Proceedings of Machine Learning Research, vol. 15, 2011, p. 19–26. [23] OpenAI, “Codex: Ai coding agents for software engineering,” https: //openai.com/codex/, 2026, accessed 2026-07-10. [24] B. Bell, “Explaining quantum networks,” GitHub repository, 2026, accessed 2026-07-10. [Online]. Available: https://github.com/496crows/ explaining-quantum-networks