Paper deep dive
A four-player potential game for barren-plateau-aware quantum ansatz design
Rubén Darío Guerrero
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 6/21/2026, 9:04:16 AM
Summary
The paper proposes a novel framework for designing parameterized quantum circuits (PQCs) by casting the architecture search as a four-player potential game. The four players optimize for competing objectives: trainability (via the dynamical Lie algebra), non-stabilizerness (via stabilizer Rényi-2 entropy), task performance (e.g., MaxCut or VQE energy), and hardware cost. The state is represented as a directed acyclic graph (DAG). The authors demonstrate that this approach can navigate the Pareto frontier between barren plateaus and classical simulability, specifically on MaxCut K4 problems. The framework is also applied to refine a chemistry-aware LiH ansatz, successfully reducing gate count while maintaining high correlation energy. Benchmarks against simulated annealing on various hardware topologies (heavy-hex, 2x2 grid, Rydberg) show the Nash search achieves higher mean potential, though statistical significance was not reached in the small-sample tests provided.
Entities (7)
Relation Signals (5)
Ruben Dario Guerrero → authored → A four-player potential game for barren-plateau-aware quantum ansatz design
confidence 100% · Ruben Dario Guerrero rudaguerman@gmail.com
LiH → benchmarkedwith → Givens-doubles ansatz
confidence 100% · On LiH/STO-3G, seeding Nash from a 58-gate Givens-doubles ansatz
Nash search → optimizes → Trainability, Non-stabilizerness, Task Performance, Hardware Cost
confidence 100% · whose players encode trainability, non-stabilizerness, task performance, and hardware cost
Parameterized Quantum Circuit → representedas → Directed Acyclic Graph
confidence 100% · whose state is a circuit directed acyclic graph (DAG)
MaxCut K4 → usedforbenchmarking → Pareto frontier navigation
confidence 90% · A single weight sweep on MaxCut K4 traces a Pareto frontier
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We cast the design of parameterized quantum circuits as a four-player potential game whose state is a circuit directed acyclic graph (DAG) and whose players encode trainability, non-stabilizerness, task performance, and hardware cost. Per-player restricted action sets factorize the move space into append, remove, retype, and rewire operations; a block-coordinate $\varepsilon$-Nash residual $\delta_\text{Nash}$ certifies that no single player can improve unilaterally. A single weight sweep on MaxCut $K_4$ traces a Pareto frontier from a Clifford endpoint $(M_2/n,\langle H\rangle)=(0,4.00)$ to a non-Clifford endpoint $(0.48,3.30)$. On three four-qubit hardware topologies (heavy-hex, $2\times 2$ grid, Rydberg all-to-all), Nash search achieves the highest mean potential; on the $2\times 2$ grid Nash reaches the theoretical ceiling $\Phi_\text{max}=4.10$ on two of five seeds while the simulated-annealing baseline does so on one; paired Wilcoxon tests over five seeds cannot reject the null on any single topology ($p\ge 0.22$). On LiH/STO-3G, seeding Nash from a 58-gate Givens-doubles ansatz produces a 48-operation, depth-25 circuit retaining $97.7\%$ of the correlation energy while simultaneously reducing gate count, increasing non-stabilizerness, and controlling trainability. The framework is complementary to energy-only searches such as ADAPT-VQE and k-UpCCGSD, which reach chemical accuracy with fewer operations but do not optimize the other three axes.
Tags
Links
- Source: https://arxiv.org/abs/2604.21955v1
- Canonical: https://arxiv.org/abs/2604.21955v1
Trouble viewing inline? Open PDF directly →
Full Text
25,372 characters extracted from source content.
Expand or collapse full text
A four-player potential game for barren-plateau-aware quantum ansatz design Ruben Dario Guerrero rudaguerman@gmail.com Parametrized-QC-Graphs Project Abstract We cast the design of parameterized quantum circuits as a four-player potential game whose state is a circuit directed acyclic graph (DAG) and whose players encode trainability, non-stabilizerness, task performance, and hardware cost. Per-player restricted action sets factorize the move space into append, remove, retype, and rewire operations; a block-coordinate ε -Nash residual δNash _Nash certifies that no single player can improve unilaterally. A single weight sweep on MaxCut K4K_4 traces a Pareto frontier from a Clifford endpoint (M2/n,⟨H⟩)=(0,4.00)(M_2/n, H )=(0,4.00) to a non-Clifford endpoint (0.48,3.30)(0.48,3.30). On three four-qubit hardware topologies (heavy-hex, 2×22× 2 grid, Rydberg all-to-all), Nash search achieves the highest mean potential; on the 2×22× 2 grid Nash reaches the theoretical ceiling Φmax=4.10 _max=4.10 on two of five seeds while the simulated-annealing baseline does so on one; paired Wilcoxon tests over five seeds cannot reject the null on any single topology (p≥0.22p≥ 0.22). On LiH/STO-3G, seeding Nash from a 58-gate Givens-doubles ansatz produces a 48-operation, depth-25 circuit retaining 97.7%97.7\% of the correlation energy while simultaneously reducing gate count, increasing non-stabilizerness, and controlling trainability. The framework is complementary to energy-only searches such as ADAPT-VQE and k-UpCCGSD, which reach chemical accuracy with fewer operations but do not optimize the other three axes. 1 Introduction Variational quantum algorithms (VQAs) offload the preparation of an entangled state to a parameterized quantum circuit (PQC) trained by a classical outer loop [1, 2]. Two complementary obstacles threaten their scaling. Barren plateaus cause the variance of every loss gradient to decay exponentially in qubit count whenever the dynamical Lie algebra (DLA) is exponentially large [3, 4]; at the other extreme, classical simulation techniques that exploit a small DLA or a low stabilizer Rényi entropy render many barren-plateau-free circuits tractable on classical hardware [5, 13, 12]. Cerezo et al. [5] framed this as a structural tension: provably trainable VQAs in the literature are also provably simulable, so a convincing quantum advantage requires circuits placed in the interior of a trade-off space, not at either extreme. Existing architecture searches attack the problem from a single direction. Hardware-efficient ansätze [6] maximize expressibility at the cost of trainability. ADAPT-VQE [7] grows operator pools adapted to chemistry and routinely reaches chemical accuracy on LiH with O(20)O(20) operations. Differentiable and stochastic architecture searches such as DQAS [9] score candidate structures by task loss alone. None of these methods exposes the trainability/simulability trade-off as a controllable axis; none returns a stopping criterion that certifies balance among competing objectives. In this work we (i) formalize PQC architecture search as a four-player potential game over circuit DAGs in which each player owns a restricted action set on the shared structure, (i) demonstrate that a single weight sweep traces the Pareto frontier of the barren-plateau/simulability tension on MaxCut K4K_4, and (i) show that the framework refines a chemistry-aware LiH ansatz while jointly controlling trainability, non-stabilizerness, and hardware cost. We report all benchmarks with honest uncertainty: our five-seed paired tests against a simulated-annealing baseline cannot reject the null on any single hardware topology despite a positive mean advantage, and we acknowledge that ADAPT-VQE reaches lower gate count than Nash at chemical accuracy on LiH. The contribution is orthogonal: Nash navigates a trade-off that energy-only methods do not see. 2 Framework 2.1 Circuit DAG as a generalized graph-state adjacency A PQC is encoded as a directed acyclic graph D=(Vop∪Vio,Ed)D=(V_op∪ V_io,E_d) whose operation nodes Vop=g1,…,gLV_op=\g_1,…,g_L\ carry a gate type from a native set G and continuous parameters θ, and whose edges carry qubit-wire labels. This DAG is a strict generalization of the graph-state adjacency matrix [10, 11]: a graph state |G⟩|G lifts into a DAG with only CZ operations, while a generic PQC DAG allows heterogeneous gate types, continuous parameters, and multi-layer temporal ordering. The lowering D↦U()D U( θ) is unique up to commuting-gate reordering and is implemented in a single codepath via TensorCircuit [15] with JAX JIT on an NVIDIA RTX 4060 GPU. 2.2 Players, actions, and potential We define four payoffs on the pair (D,)(D, θ): f1 f_1 =deff(D,) =d_eff(D, θ) (trainability) f2 f_2 =M2(|ψD,⟩)/n =M_2(| _D, θ )/n (non-stabilizerness) (1) f3 f_3 =−⟨H⟩D, or +⟨H⟩ =- H _D, θ or + H (task) f4 f_4 =Chw(D) =C_hw(D) (hardware cost) Here deffd_eff is derived from the quantum Fisher information matrix spectrum and controls gradient variance through the DLA bound Var[∂lℒ]∈Θ(1/dim)Var[ _lL]∈ (1/ g) [4]. The stabilizer Rényi-22 entropy M2M_2 [12] is a practical non-stabilizerness proxy whose magnitude lower-bounds the Bravyi–Gosset simulation cost [13]. The task term f3f_3 is the minimized Hamiltonian expectation for VQE tasks and the maximized cut value for MaxCut. The hardware cost ChwC_hw counts native gates under the target connectivity. Crucially, each player i owns a restricted action set iA_i that only player i may exercise on the shared DAG. Player f1f_1 (trainability) can retype a gate into one that enlarges the DLA; player f2f_2 (non-stabilizerness) can retype into non-Clifford primitives; player f3f_3 (task) can rewire two-qubit connections; player f4f_4 (hardware) can remove a gate. Every structural move is restricted to respect the target hardware’s connectivity graph. The potential Φ(D,)=w1f1+w2f2+w3f3−w4f4 (D, θ)=w_1f_1+w_2f_2+w_3f_3-w_4f_4 (2) is a weighted scalarization. We emphasize that the potential form itself is unremarkable: any weighted scalarization of a multi-objective optimization is a potential game in the sense of Monderer and Shapley [14]. The non-trivial content lies in the factorization of moves into per-player restricted action sets, which gives each objective its own gradient-descent direction in circuit-structure space and allows a meaningful residual test. 2.3 Nash certificate For each player i, the restricted best-response gap δNash(i)=maxa∈i[fi(a⋅D,)−fi(D,)]+δ^(i)_Nash= _a _i [f_i(a· D, θ)-f_i(D, θ) ]_+ (3) measures how much player i could gain by a unilateral deviation from (D,)(D, θ). A point is an ε -Nash equilibrium when δNash≡maxiδNash(i)≤ε _Nash≡ _iδ^(i)_Nash≤ . This is the standard coordinate-wise ε -Nash residual adapted to the restricted-action per-player move sets; viewed from the optimization side it is a block-coordinate stationarity condition. It provides a sharper stopping criterion than the single-objective loss plateau used in DQAS and in simulated-annealing baselines. 2.4 Optimizer An outer simulated-annealing loop over the structural moves (append, remove, retype, rewire) alternates with an inner gradient descent on θ. We benchmark against a simulated-annealing baseline over the same structural move set that scores candidates by Φ alone (no per-player residual); this baseline plays the role attributed to SA-DQAS in an earlier version of this manuscript. Figure 1 shows the framework schematic. Figure 1: Four-player potential game on circuit DAGs. The shared state is a PQC DAG whose nodes carry gate types and parameters. Four players own disjoint action sets: f1f_1 (trainability) retypes a gate to enlarge the dynamical Lie algebra, f2f_2 (non-stabilizerness) retypes into non-Clifford primitives, f3f_3 (task) rewires a two-qubit gate, f4f_4 (hardware) removes a gate. For example, starting from a two-gate H2 circuit on heavy-hex connectivity, player f2f_2 may append a T gate on qubit 0, and the resulting DAG is lowered to a TensorCircuit object for gradient descent on the continuous parameters. An outer simulated-annealing loop over structural moves alternates with an inner parameter update. The residual δNash _Nash (bottom) certifies that no player can improve unilaterally. 3 Results 3.1 Pareto navigation on MaxCut K4K_4 A single weight sweep over the (w1,w2)(w_1,w_2) corners traces a frontier in the (non-stabilizerness, energy) plane for MaxCut on K4K_4 (n=4n=4 qubits, cut value 44); Figure 2 shows sixteen corner runs with w3w_3 fixed. The Nash solutions populate a monotone frontier from a Clifford endpoint (M2/n,⟨H⟩)=(0,4.00)(M_2/n, H )=(0,4.00), in which a stabilizer circuit reaches the exact MaxCut optimum, to a non-Clifford endpoint (0.48,3.30)(0.48,3.30) reached when w2w_2 dominates. An interior corner at (w1,w2)=(1,0.3)(w_1,w_2)=(1,0.3) places the circuit at (0.076,3.939)(0.076,3.939), retaining 98.5%98.5\% of the maximum cut while placing the state well inside the non-stabilizer regime. No frontier point is dominated by another in both coordinates, confirming that the restricted-action Nash residual selects for balance rather than for any extremum. The interpretation is direct: the weight sweep navigates the Cerezo–Larocca tension on a problem where both endpoints are accessible. Large w1w_1 drives deffd_eff up, squeezing the circuit toward the simulable Clifford corner; large w2w_2 moves it toward the exponentially simulation-hard regime at the cost of task energy. This demonstrates Pareto navigation on n=4n=4 MaxCut K4K_4; verifying that the frontier persists as n increases is future work, because the combinatorial move space grows rapidly and our n=8n=8 TFIM scaling data (Sec. 3.3) does not yet span a full weight sweep. Figure 2: Pareto frontier for MaxCut on K4K_4. Each point is a Nash-equilibrium circuit at a distinct weight corner (w1,w2)(w_1,w_2); sixteen corners are shown with w3w_3 fixed. The frontier spans from a Clifford endpoint (non-stabilizerness M2/n=0M_2/n=0, ⟨H⟩=4.00 H =4.00) to a non-Clifford endpoint (0.48,3.30)(0.48,3.30). An interior Pareto-knee solution at (0.076,3.939)(0.076,3.939) is reached by (w1,w2)=(1,0.3)(w_1,w_2)=(1,0.3). Results are shown for n=4n=4 qubits only. 3.2 Head-to-head against a simulated-annealing baseline We compare Nash search against a simulated-annealing baseline that searches over the same structural move set but scores candidates by the scalar potential alone, on the four-player objective Φ across three hardware topologies (IBM heavy-hex subset, 2×22× 2 square grid, Rydberg all-to-all). Five independent seeds per condition were run on a single RTX 4060 GPU; Figure 3 shows means with 95% bootstrap confidence intervals, and Table 1 reports paired statistics. Table 1: Nash versus simulated-annealing baseline across three topologies. Five seeds per cell; ΔΦ is the mean paired difference (Nash −- baseline); bootstrap CI covers the mean difference; p is the paired Wilcoxon one-sided (H1H_1: Nash >> baseline); dzd_z is Cohen’s paired effect size; the last column reports the number of seeds reaching Φ≥4.099 ≥ 4.099 (within numerical tolerance of the theoretical ceiling Φmax=4.10 _max=4.10). Topology Nash mean ± sd Baseline mean ± sd ΔΦ 95% CI Wilcoxon p (dzd_z) Ceiling hits heavy-hex 3.82±0.303.82± 0.30 3.73±0.193.73± 0.19 +0.09+0.09 [−0.20,+0.38][-0.20,+0.38] 0.500.50 (0.24)(0.24) 2/52/5 vs. 0/50/5 2×22×2 grid 3.94±0.093.94± 0.09 3.79±0.263.79± 0.26 +0.15+0.15 [−0.05,+0.36][-0.05,+0.36] 0.220.22 (0.60)(0.60) 1/51/5 vs. 1/51/5 Rydberg all-to-all 3.80±0.333.80± 0.33 3.67±0.323.67± 0.32 +0.13+0.13 [−0.28,+0.50][-0.28,+0.50] 0.500.50 (0.25)(0.25) 2/52/5 vs. 0/50/5 Nash attains the highest mean Φ on every topology, but we cannot reject the null hypothesis of no difference on any single topology at five seeds: paired Wilcoxon p-values are 0.500.50 (heavy-hex), 0.220.22 (grid), and 0.500.50 (Rydberg) for the one-sided alternative. Effect sizes are small to moderate (dz∈[0.24,0.60]d_z∈[0.24,0.60]). Bootstrap confidence intervals for the mean paired difference cross zero on all three topologies. The clearest qualitative contrast appears on the 2×22× 2 grid, where Nash reaches the theoretical ceiling Φmax=4.10 _max=4.10 on two of five seeds while the baseline reaches it on one; on heavy-hex and Rydberg Nash also touches the ceiling on two of five seeds while the baseline never does. Larger seed counts and expanded per-seed budgets will be required to establish statistical significance at the single-topology level; we do not claim a significant per-topology advantage on the present data. Figure 3: Head-to-head comparison at matched budget across three hardware topologies. Bars show five-seed means of the scalar potential Φ ; error bars are 95% bootstrap confidence intervals on the mean. Nash achieves the highest mean on every topology with ΔΦ=+0.09 =+0.09 (heavy-hex), +0.15+0.15 (2×22× 2 grid), and +0.13+0.13 (Rydberg). Per-topology paired Wilcoxon tests do not reject the null on five seeds (see Table 1). On the grid, Nash reaches the theoretical ceiling Φmax=4.10 _max=4.10 on 2/52/5 seeds while the baseline reaches it on 1/51/5. 3.3 Scaling on the transverse-field Ising model To probe how Nash search behaves as n grows, we run the framework on the critical one-dimensional transverse-field Ising model for n∈4,6,8n∈\4,6,8\ qubits with five seeds per size and a fifteen-iteration outer budget [Fig. 4(a,b)]. Warm-started from a QAOA p=1p=1 state, Nash reaches relative errors of 5.87%5.87\%, 7.93%7.93\%, and 7.92%7.92\% in the ground-state energy at n=4,6,8n=4,6,8 respectively (five-seed means; 95% bootstrap CI bands shown); the cold start from |+⟩⊗n|+ n plateaus an order of magnitude higher. Wall-clock per iteration grows approximately linearly at ∼4.7 4.7 s/qubit [Fig. 4(c)]. The Nash residual δNash _Nash reported at final iteration is bimodal at larger n: the per-seed values at n=4,6,8n=4,6,8 are 0.10,0.02,0.10,0.20,0.05\0.10,0.02,0.10,0.20,0.05\, 0.20,0.005,0.03,0.03,0.04\0.20,0.005,0.03,0.03,0.04\, and 0.20,0.02,0.01,0.01,0.20\0.20,0.02,0.01,0.01,0.20\, giving medians 0.100,0.033,0.0230.100,0.033,0.023 and means 0.094,0.061,0.0890.094,0.061,0.089, respectively. At n=8n=8 three seeds converge tightly to δNash≲0.02 _Nash 0.02 while two are stuck at the discrete-move resolution 0.200.20; this bimodality is the signature of the structural move set locally exhausting improvements on some seeds while remaining sub-converged on others. The median tightens monotonically with n, consistent with the intuition that larger systems have more co-equal Nash candidates and fewer move directions with large per-player gains, but the mean does not — we report both statistics rather than picking the favorable one. 3.4 Chemistry: H2 sanity check and LiH multi-objective refinement H2/STO-3G as a sanity check. On H2 at bond length 0.74140.7414 Å with heavy-hex connectivity and gate set h,x,y,z,s,t,t†,rx,ry,rz,rzz,CNOT,CZ\h,x,y,z,s,t,t ,r_x,r_y,r_z,r_z,CNOT,CZ\, Nash converges to E=−1.1373E=-1.1373 Ha, reproducing the active-space ground state after the standard parity and spin-symmetry reduction using a two-gate, depth-1 circuit. We include this as a sanity check only: H2/STO-3G after symmetry reduction is a two-qubit problem whose ground state is a product of two single-qubit rotations, and any reasonable search recovers it. A random-init HEA of depth 4 with 14 operations trained with θ only is trapped ≈20≈ 20 mHa above Hartree–Fock, as expected for barren-plateau-prone HEAs on such a shallow problem. LiH active-space VQE. LiH in the frozen-core active space (6 qubits, including Li 2px,y2p_x,y virtuals) has reference EHF=−7.8631E_HF=-7.8631 Ha and active-space ground state EGS=−7.8778E_GS=-7.8778 Ha, with full FCI at EFCI=−7.8828E_FCI=-7.8828 Ha; the correlation gap over Hartree–Fock is 14.6514.65 mHa. A random-init HEA with 24 gates and depth 11, trained with θ only, is trapped at +165+165 mHa above HF — a textbook barren-plateau signature. With the generic hardware-efficient gate set alone, Nash saturates at EHFE_HF with 6 gates and depth 1 robustly across reweightings up to w3=5w_3=5: a gate-set diagnostic that the HE primitives lack particle-number-conserving excitations. Seeding Nash from a chemistry-aware 58-gate Givens-doubles ansatz (Hartree–Fock preparation plus two paired DoubleExcitation gates, verified against PennyLane) dissolves the diagnostic. The structural moves refine the seed ansatz to 48 operations at depth 25 with 10 parameters while retaining 97.7%97.7\% of the correlation energy (E=−7.8775E=-7.8775 Ha, 0.330.33 mHa above exact). The potential climbs monotonically from Φ=29.3 =29.3 to Φ=33.5 =33.5 over 20 outer iterations; δNash _Nash converges to 0.080.08. Pure Adam on the unmodified 58-gate Givens ansatz reaches EGSE_GS exactly in 0.60.6 s on a single GPU, setting the absolute energy ceiling. The honest framing of this result is not a compression win against chemistry baselines. ADAPT-VQE [7] reaches chemical accuracy on LiH with O(20)O(20) operations, and k-UpCCGSD and related unitary coupled-cluster truncations [8] are similarly efficient when the objective is energy alone. Our contribution is orthogonal: Nash refines a chemistry-aware ansatz while simultaneously controlling f1f_1 (trainability), f2f_2 (non-stabilizerness), and f4f_4 (hardware cost), producing a single circuit that is Pareto-balanced along all four axes rather than optimal along one. The 1010-gate reduction at 2.3%2.3\% correlation-energy cost is the trade-off visible in the Nash residual; it is not a claim of superiority over energy-only baselines. Figure 4: Scaling and chemistry summary. (a) TFIM ground-state relative error vs. qubit count n for warm-started (QAOA p=1p=1) and cold-started Nash; five seeds per size; shaded bands are 95% bootstrap confidence intervals on the mean. (b) The same data plotted as a function of Nash outer iteration. (c) Per-iteration wall clock grows approximately linearly at ∼4.7 4.7 s/qubit. Inset: δNash _Nash at final iteration is bimodal at larger n; per-seed values at n=4,6,8n=4,6,8 are shown individually. The median δNash _Nash tightens from 0.1000.100 at n=4n=4 to 0.0230.023 at n=8n=8, but the mean does not because two of five seeds remain at the move-resolution 0.200.20. Annotated: H2 recovery (E=−1.1373E=-1.1373 Ha, two gates, sanity check) and LiH Givens-seeded refinement (97.7%97.7\% correlation, 4848 gates, 4848/5858 gate reduction with 2.3%2.3\% correlation-energy cost). 4 Discussion Four features distinguish the framework. First, the per-player restricted-action residual δNash _Nash is a block-coordinate stationarity certificate that single-objective searches lack. Second, trainability is coupled to the QFIM spectrum by construction through f1f_1, so structural moves that collapse the DLA are penalized by their own player rather than indirectly through the loss. Third, the three hardware benchmarks run the same algorithm, differing only in the connectivity constraint on the structural move set — the framework is topology-generic. Fourth, seeding from a Givens ansatz converts Nash into a chemistry-ansatz refiner without any algorithmic change. Limitations deserve explicit mention. The Pareto frontier is demonstrated on n=4n=4 MaxCut K4K_4 only; extending the weight sweep to n=6,8n=6,8 and mapping the full frontier at scale is future work. Our paired Wilcoxon tests on the three-topology head-to-head do not reject the null on any single topology at five seeds; replication with ≳20 20 seeds and larger per-seed budgets is required to establish per-topology significance. At the chemistry frontier, ADAPT-VQE reaches lower gate counts than Nash at chemical accuracy on LiH; our framework is complementary, optimizing trainability, non-stabilizerness, and hardware cost jointly with energy rather than minimizing energy alone, and we do not claim an energy-wise advantage. The LiH Givens-seeded run does not close the 0.330.33 mHa residual gap to exact correlation because the generic gate set lacks full UCCSD primitives. Finally, JIT compilation of the lowered tc.Circuit dominates wall-clock for deep ansätze, which caps the practical structural-move budget on our single-GPU setup. Three directions extend naturally. Promoting Givens rotations and particle-number-conserving primitives to first-class DAG gates should close the LiH residual. Potential-energy-surface continuation, where a Nash circuit at one bond length warm-starts the next, will test whether equilibria deform smoothly along reaction coordinates. The DAG formulation makes the embedding of graph-state codes [10] into the same optimization pipeline immediate, suggesting a unified search algorithm spanning code synthesis, ansatz design, and compilation. The framework provides an equilibrium notion adapted to the barren-plateau/simulability tension: not the best circuit by any single metric, but a circuit that no player can unilaterally improve subject to hardware and connectivity constraints. Whether this equilibrium notion yields a quantum advantage at scales beyond n=8n=8 remains open and is the central empirical question for the next round of experiments. Data and code availability The code, the per-experiment configuration files, and all JSON results used to produce the figures and tables in this paper are available at the project repository (https://github.com/rdguerrerom/Parametrized-QC-Graphs, Parametrized-QC-Graphs). Raw per-seed files used in this manuscript are results/multi_seed_d1.json (TFIM scaling), results/multi_seed_f6.json (topology head-to-head), results/exp3_weight_sweep.json (Pareto frontier), results/exp_c3_lih_givens_nash.json (LiH), and results/exp2_h2_heavy_hex.json (H2). Author contributions R.D.G. conceived the framework, implemented the codebase, ran the experiments, and wrote the manuscript. Competing interests The author declares no competing interests. Acknowledgments We acknowledge financial support and computational resources provided by NeuroTechNet S.A.S. Core primitives were adapted from the STABILIZER_GAMES project. TensorCircuit and JAX provided the GPU optimization backend. Compute was performed on a single NVIDIA RTX 4060 GPU. References [1] M. Cerezo et al., Variational quantum algorithms, Nat. Rev. Phys. 3, 625 (2021). [2] K. Bharti et al., Noisy intermediate-scale quantum algorithms, Rev. Mod. Phys. 94, 015004 (2022). [3] J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, and H. Neven, Barren plateaus in quantum neural network training landscapes, Nat. Commun. 9, 4812 (2018). [4] M. Ragone et al., A Lie algebraic theory of barren plateaus for deep parameterized quantum circuits, Nat. Commun. 15, 7172 (2024). [5] M. Cerezo et al., Does provable absence of barren plateaus imply classical simulability?, Nat. Commun. 16, 7907 (2025). [6] A. Kandala et al., Hardware-efficient variational quantum eigensolver for small molecules and quantum magnets, Nature 549, 242 (2017). [7] H. R. Grimsley, S. E. Economou, E. Barnes, and N. J. Mayhall, An adaptive variational algorithm for exact molecular simulations on a quantum computer, Nat. Commun. 10, 3007 (2019). [8] J. Lee, W. J. Huggins, M. Head-Gordon, and K. B. Whaley, Generalized unitary coupled cluster wave functions for quantum computation, J. Chem. Theory Comput. 15, 311 (2019). [9] S.-X. Zhang, C.-Y. Hsieh, S. Zhang, and H. Yao, Differentiable quantum architecture search, Quantum Sci. Technol. 7, 045023 (2022). [10] M. Hein, J. Eisert, and H. J. Briegel, Multiparty entanglement in graph states, Phys. Rev. A 69, 062311 (2004). [11] R. Raussendorf and H. J. Briegel, A one-way quantum computer, Phys. Rev. Lett. 86, 5188 (2001). [12] L. Leone, S. F. E. Oliviero, and A. Hamma, Stabilizer Rényi entropy, Phys. Rev. Lett. 128, 050402 (2022). [13] S. Bravyi, G. Smith, and J. A. Smolin, Trading classical and quantum computational resources, Phys. Rev. X 6, 021043 (2016). [14] D. Monderer and L. S. Shapley, Potential games, Games Econ. Behav. 14, 124 (1996). [15] S.-X. Zhang et al., TensorCircuit: a quantum software framework for the NISQ era, Quantum 7, 912 (2023).