Paper deep dive
KQFuzz: Knowledge-Guided Fuzzing for Quantum Libraries via Large Language Models
Fuyuan Xia, Qixin Zhang, Chenhao Ying, Haojin Zhu, Shuai Wang, Yuan Luo, Pingchuan Ma, Yuxuan Du
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As quantum computing continually improves, ensuring the reliability and correctness of quantum libraries has become increasingly critical. To this end, many LLM-based fuzzing approaches towards quantum libraries have been proposed to uncover potential bugs. However, these methods still suffer from limitations such as insufficient flexibility and low efficiency, which hinder the progress of the quantum computing field. To address these challenges, we propose KQFuzz, a novel knowledge-guided fuzzer for quantum libraries. It leverages comprehensive codebase knowledge to ground LLM-based test generation, synergizing this with fitness-guided evaluation and two-level mutations to explore complex execution paths and trigger potential bugs. Firstly, KQFuzz introduces a novel prompting scheme tailored to quantum programs, which strategically incorporates knowledge of the codebase to efficiently generate high-quality quantum seed programs. Moreover, we develop evaluation and mutation strategies to handle the generated seed programs, facilitating efficient fuzzing execution while further enriching the diversity of the resulting test cases. We implement KQFuzz and conduct fuzzing on three popular quantum libraries, including Qiskit, PennyLane, and Cirq. Experimental results demonstrate that our approach significantly outperforms other state-of-the-art methods, with coverage improved by up to 18.44%. During the development of KQFuzz, we discovered 13 bugs, all of which have been confirmed and 12 have already been fixed by the developers.
Tags
Links
- Source: https://arxiv.org/abs/2607.25647v1
- Canonical: https://arxiv.org/abs/2607.25647v1
Trouble viewing inline? Open PDF directly â
Full Text
133,549 characters extracted from source content.
Expand or collapse full text
KQFuzz: Knowledge-Guided Fuzzing for Quantum Libraries via Large Language Models Fuyuan Xia Shanghai Jiao Tong UniversityShanghaiChina fuyuanxia@sjtu.edu.cn , Qixin Zhang Nanyang Technological UniversitySingaporeSingapore qixin.zhang@ntu.edu.sg , Chenhao Ying Shanghai Jiao Tong UniversityShanghaiChina yingchenhao@sjtu.edu.cn , Haojin Zhu Shanghai Jiao Tong UniversityShanghaiChina zhu-hj@sjtu.edu.cn , Shuai Wang Hong Kong University of Science and TechnologyHong KongChina shuaiw@cse.ust.hk , Yuan Luo Shanghai Jiao Tong UniversityShanghaiChina yuanluo@sjtu.edu.cn , Pingchuan Ma Zhejiang University of TechnologyHangzhouChina pma@zjut.edu.cn and Yuxuan Du College of Computing and Data Science, and School of Physical and Mathematical SciencesNanyang Technological UniversitySingaporeSingapore yuxuan.du@ntu.edu.sg (2026) Abstract. As quantum computing continually improves, ensuring the reliability and correctness of quantum libraries has become increasingly critical. To this end, many LLM-based fuzzing approaches towards quantum libraries have been proposed to uncover potential bugs. However, these methods still suffer from limitations such as insufficient flexibility and low efficiency, which hinder the progress of the quantum computing field. To address these challenges, we propose KQFuzz, a novel knowledge-guided fuzzer for quantum libraries. It leverages comprehensive codebase knowledge to ground LLM-based test generation, synergizing this with fitness-guided evaluation and two-level mutations to explore complex execution paths and trigger potential bugs. Firstly, KQFuzz introduces a novel prompting scheme tailored to quantum programs, which strategically incorporates knowledge of the codebase to efficiently generate high-quality quantum seed programs. Moreover, we develop evaluation and mutation strategies to handle the generated seed programs, facilitating efficient fuzzing execution while further enriching the diversity of the resulting test cases. We implement KQFuzz and conduct fuzzing on three popular quantum libraries, including Qiskit, PennyLane, and Cirq. Experimental results demonstrate that our approach significantly outperforms other state-of-the-art methods, with coverage improved by up to 18.44%. During the development of KQFuzz, we discovered 13 bugs, all of which have been confirmed and 12 have already been fixed by the developers. Quantum Library, Fuzzing, Large Language Model. â journalyear: 2026â conference: 41st IEEE/ACM International Conference on Automated Software Engineering; October 12â16, 2026; Munich, Germany 1. Introduction Diverse quantum libraries have been developed by academic institutions and technology companies to accelerate the design and implementation of quantum algorithms, motivated by the promise that quantum computing offers capabilities beyond classical computation across diverse domains (quantum; ying2025tosem). However, an often overlooked issue is that bugs in such quantum libraries may produce unexpected results, leading to flawed scientific or engineering conclusions and obscuring the true potential of quantum computing (ali2022taoyue; fang2024pldi; assolini2024sas; li2026methodological; long2024tosem; upadhyay2026understandingbugsquantumsimulators). Scarce quantum hardware further amplifies the impact of software bugs, leading to significant resource waste. In this regard, much like in classical software engineering, there is an increasing need for systematic quantum library testing to ensure correctness and reliability. However, porting classical testing methodologies to the quantum domain is far from straightforward. Conventional techniques like static analysis (staticanalysis) often fail because they lack the domain-specific semantics required to comprehend complex quantum states, making verification difficult. Furthermore, the rapid evolution of quantum libraries quickly renders manual testing rules obsolete, necessitating automated and adaptive approaches that can systematically explore the complex logic of these systems. To meet these requirements, fuzzing (manes2019tse; jiang2024fuzzing) has emerged as a leading candidate for automatically discovering bugs. By generating and executing a vast number of randomized test cases, fuzzing can effectively explore the complex input space and identify unexpected software behaviors across diverse execution scenarios, which is uniquely suited for quantum libraries. Figure 1. Comparison of Fuzzing Scopes. Limitations of existing fuzzing-based quantum library testing. In this work, quantum-library testing treats the library or platform implementation as the system under test, while quantum programs serve as test inputs. Existing quantum-library fuzzers mainly rely on domain-specific test generation (paltenghi2023morphq; shaking2025oopsla; hu2024issta), using dedicated generators, transformation rules, templates, or constraint specifications to construct structurally valid programs. However, these approaches necessitate extensive domain expertise to manually formulate generation and mutation rules. Moreover, as shown in Figure 1, such manual curation confines their scope to circuit-related APIs, rendering them inflexible and unable to exercise emerging software features or high-level orchestration logic. Consequently, they are incapable of identifying entire classes of bugs in the rapidly evolving quantum computing ecosystem. Alternatively, general-purpose LLM-based fuzzers such as (xia2024fuzz4all) offer a promising path forward, as they can use their domain knowledge to generate seeds invoking diverse APIs without manual rule design. However, instantiating LLM-based fuzzing for quantum programming libraries faces unique Challenges compared to classical software. C1. Low-validity seed generation. Unlike classical software with stable interfaces, the frequent restructuring of quantum APIs hinders LLMs from maintaining up-to-date internal representations of the evolving API surface and its associated semantic constraints. As illustrated in Figure LABEL:fig:fuzzing_eval in § LABEL:sec:challenge, our pilot study reveals that LLMs achieve only 35â46% validity for quantum programming libraries compared to 64â86% for classical programming libraries with disproportionately high import and attribute errors, averaging 118 such errors per 600 programs in each library. C2. Constrained exploration in quantum programs. Unlike classical software, quantum programs are subject not only to hardware-level constraints such as limited qubit counts and restricted connectivity, but also to framework-level execution semantics and hybrid quantumâclassical control flows. Conventional mutation strategies operating at the byte or abstract-syntax-tree level often fail to preserve these structural and semantic constraints, thereby generating invalid or non-executable programs. The core challenge is therefore to design domain-aware mutation mechanisms that systematically explore the input space while respecting quantum-specific structural, semantic, and complicated constraints. Our solution. To address these challenges, we propose KQFuzz, a Knowledge-guided Quantum Fuzzing framework, which incorporates codebase-specific structural and semantic information to guide LLM-driven seed program generation for quantum libraries. Moreover, KQFuzz integrates fitness-guided evaluation and mutation mechanisms to systematically explore complex execution paths and expose latent defects. To address C1, KQFuzz constructs an API corpus capturing four types of knowledge: (1) static metadata, such as signatures and locations; (2) API associations via three dimensions, i.e., proximity, type coupling, and call relationships; (3) semantic models via LLM-based summarization; and (4) evolution metrics tracking API modifications across versions. During generation, KQFuzz employs a probabilistic selection strategy that prioritizes APIs with strong semantic associations and high evolutionary activity, steering the LLM to synthesize valid, version-aligned seed programs that target potentially bug-prone execution paths. In response to C2, KQFuzz employs a two-level mutation strategy, where parameter-level mutation targets numerical corner cases and gate-level structural mutation explores variations through entanglement-aware substitutions. Furthermore, we introduce a fitness function that integrates gate diversity, entangled qubits, API diversity, and call depth. This function acts as a selection criterion to prioritize high-quality seeds, steering the LLM toward generating programs in more complex and potentially bug-prone regions. Seeking a balance between precision and throughput, KQFuzz adopts a tiered synergy: a high-capacity model (i.e., Gemini-3 (team2023gemini)) performs one-time API semantic modeling to ground a cost-effective LLM (e.g., Qwen 2.5-Coder) for iterative fuzzing generation. Finally, we execute the resulting test casesâcomprising both the generated seed programs and their mutated variantsâand apply heuristic filtering to the execution results to identify impactful bugs within the target quantum libraries. We evaluate the performance of KQFuzz on three mainstream quantum libraries: Qiskit111https://github.com/Qiskit/qiskit, PennyLane222https://github.com/PennyLaneAI/pennylane, and Cirq333https://github.com/quantumlib/Cirq. KQFuzz significantly outperforms existing baselines in code validity and bug-finding effectiveness, discovering 13 bugs, all of which have been confirmed and 12 already fixed by the developers. Our contributions are summarized as follows. âą We propose KQFuzz, a novel knowledge-guided quantum fuzzing framework that integrates codebase-specific insights with LLM-based test generation and domain-specific mutation to systematically explore complex execution paths in quantum libraries. âą We construct a comprehensive API corpus encompassing four dimensions of knowledge. Based on this, a seed program generation strategy is designed to steer the LLM toward generating valid, version-aligned, and potentially bug-prone quantum programs. âą We introduce a two-level mutation strategy (parameter and gate-level) and a multi-dimensional fitness function that accounts for gate diversity and entanglement. This approach prioritizes high-quality seeds to effectively navigate the complex state space of quantum programs. âą We perform a comprehensive evaluation of KQFuzz on three mainstream quantum libraries. The results demonstrate its superiority over existing baselines, surpassing the state-of-the-art coverage by a margin of up to 18.44%, successfully identifying 13 new bugs, all of which have been confirmed by developers and 12 of which are already fixed. 2. Background 2.1. Basic Concepts in Quantum Computing Quantum computing leverages principles such as superposition and entanglement to enable new computational capabilities (gill2025quantumc). In practice, these capabilities are commonly formulated within the quantum circuit model (fowler2012surface), in which computation is represented as a sequence of elementary quantum gates, e.g., single-qubit and entangling two-qubit operations, acting on quantum states. After that, a measurement process is applied to extract information from the final quantum state into classical registers, yielding probabilistic outcomes that encode the solution to a given problem (abbas2024nrp; cerezo2021nrp). Refer to Supplementary Material (SM) A for more details. Within the quantum computing ecosystem, Qiskit (qiskit), PennyLane (pennylane), and Cirq (cirq_developers_2025_10.5281/zenodo.4062499) are three major quantum programming libraries maintained by major technology companies or quantum-focused companies. While these libraries share extensive functional commonalities in circuit synthesis, simulation, and hardware integration, they provide alternative development environments that cater to different hardware-specific optimizations and research workflows. Despite their widespread adoption, these libraries remain in a state of perpetual and volatile transformation to keep pace with rapid hardware and algorithmic breakthroughs (paltenghi2023morphq). These rapid shifts often result in hidden defects, leading to persistent logic anomalies and unexpected behaviors that deviate from intended specifications. These recurring bugs represent a significant bottleneck that necessitates rigorous testing and validation to ensure the integrity of the quantum computing ecosystem. 2.2. Fuzzing Fuzzing foundations and LLM-based evolution. Fuzzing is a cornerstone of automated software testing, designed to identify bugs and edge-case behaviors by executing a target system with a vast array of synthesized inputs (fioraldi2020afl; stephens2016driller; bohme2017directed; zhu2022csur; she2024ccs; xie2025ase; bohme2020boosting; wu2022one). Traditional fuzzing taxonomies primarily distinguish between generation-based approaches, which construct inputs from predefined grammars, and mutation-based strategies, which apply stochastic transformations to existing seeds to maximize code coverage (dong2025tdsc; manes2019tse). While effective for low-level protocols, these conventional methods often struggle with high-level software that requires complex, semantically valid structures. Recently, LLMs have redefined this paradigm. By leveraging their inherent semantic comprehension and code synthesis capabilities, LLM-driven fuzzers can autonomously generate syntactically sophisticated and contextually rich test cases (xia2024fuzz4all; deng2024icse). This shift alleviates the burden of manual rule engineering and enables the exploration of deep logic within complex API sequences. Fuzzing in the quantum libraries. The application of fuzzing to quantum libraries has bifurcated into two primary paradigms: circuit-level construction and LLM-driven exploration. Circuit-level fuzzers typically employ domain-specific generators, transformation rules, or templates to generate gate sequences, ensuring high validity but often remaining confined to low-level gate logic (paltenghi2023morphq; shaking2025oopsla). In contrast, LLM-based methods harness extensive semantic knowledge to probe high-level behaviors in quantum programming libraries. A persistent objective in this field is to effectively bridge the gap between the expressive flexibility of LLM-generated test cases and the stringent syntactic requirements of quantum libraries, ensuring that the generated programs remain both diverse and executable. 3. Motivation & Challenges ⏠1# <IMPORT TEMPLATE> 2qc = QuantumCircuit(11, 11, name=âqcâ) 3qc.append(ECRGate(), qargs=[qc.qubits[2], qc.qubits[5]], cargs=[]) 4qc.append(CHGate(), qargs=[qc.qubits[5], qc.qubits[7]], cargs=[]) 5qc.append(ZGate(), qargs=[qc.qubits[5]], cargs=[]) 6qc.append(ZGate(), qargs=[qr[7]], cargs=[]) 7qc.append(RCCXGate(), qargs=[qr[8], qr[2], qr[6]], cargs=[]) 8qc.append(ZGate(), qargs=[qr[0]], cargs=[]) 9qc.append(iSwapGate(), qargs=[qr[1], qr[3]], cargs=[]) 10qc.append(SdgGate(), qargs=[qr[2]], cargs=[]) 11qc.append(CXGate(), qargs=[qr[3], qr[1]], cargs=[]) 12# <RESULT TEMPLATE> lstlisting 13 14 Code Generated by Circuit-Level Method 15 minipage 16 17 minipage8.4cm 18 lstlisting 19# <QUANTUM CODE> 20from qiskit.circuit.library import grover_operator 21oracle = QuantumCircuit(3) 22oracle.h(2) 23oracle.ccx(0, 1, 2) 24oracle.h(2) 25grover_op = grover_operator(oracle) # Involved API 26qc = QuantumCircuit(3) 27qc.h([0, 1, 2]) 28qc.append(grover_op, [0, 1, 2]) 29qc.measure_all() 30# <QUANTUM CODE> lstlisting 31 Realistic Quantum Program 32 minipage 33 34 Comparison between circuit-level fuzzer-generated code and realistic quantum program. fig:example_diff_method 35 figure 36 37 Revisiting Fuzzing for Quantum Libraries 38We summarize two typical families of existing solutions for fuzzing quantum libraries as follows. First, circuit-level fuzzers aim to generate random quantum circuits to trigger bugs in the quantum software stack~ paltenghi2023morphq,shaking2025oopsla. Their primary approaches use domain-specific generators, transformation rules, templates, or constraint specifications to rapidly produce valid quantum programs, prioritizing the efficient discovery of structural defects such as crashes or incorrect circuit outputs. Nevertheless, this design entails an inherent trade-off between efficiency and generality. By enforcing predefined structural rules to ensure efficient and valid circuit generation, such approaches limit their applicability to a broader range of quantum programs. Consequently, these approaches face two key challenges: (1) manually designed rules become increasingly difficult to maintain as quantum programming libraries evolve and introduce new operations; and (2) their highly specialized generation mechanisms primarily target structural circuit anomalies, limiting their ability to expose the diverse defects that arise in complex and real-world quantum libraries. As indicated in Figure~ fig:example_diff_method(a), MorphQ~ paltenghi2023morphq can generate legal circuits containing a large number of random quantum gates. However, in practice ( e.g., the deployment of the Grover algorithm on real quantum hardware), additional APIs beyond circuit gates are required, such as grover\_operator in Figure~ fig:example_diff_method(b). For such APIs that go beyond the circuit level, potential bugs cannot be uncovered by these methods, thus constraining the applicability of circuit-level fuzzers to quantum libraries. 39 40While fuzzers at the circuit level often suffer from limited flexibility, approaches based on LLMs have emerged as a promising alternative~ deng2023issta,xia2024fuzz4all. Attributed to the extensive pretraining on diverse code corpora, LLMs have the potential to effectively capture the complex semantic logic of quantum libraries to generate diverse test cases. However, the syntactic and semantic validity of quantum code generated by LLMs remains a significant bottleneck. As reported in~ xia2024fuzz4all, only 24.9\% of Qiskit code generated by LLMs is valid, a figure that pales in comparison to the 100\% validity rate achieved by circuit-level fuzzers such as MorphQ~ paltenghi2023morphq. This high failure rate severely constrains the practical utility of LLMs in the domain of quantum library testing. 41 42 Motivating Study 43 sec:challenge 44As discussed in ( ~ sec:intro), LLMs can be leveraged to generate fuzzing inputs targeting a given library. To systematically assess the capability of LLMs to generate fuzzing inputs for quantum libraries, we conduct a preliminary empirical comparison between quantum libraries and their well-established classical analogues. 45 46 Experimental configuration. 47We select three prominent quantum libraries, Qiskit, Cirq, and PennyLane, as our primary targets. To provide a comparative baseline, we also include three widely used classical libraries: Numpy https://github.com/numpy/numpy, Pandas https://github.com/pandas-dev/pandas, and PyTorch https://github.com/pytorch/pytorch. We adopt a representative LLM-based fuzzing pipeline similar to~ lin2025ase, utilizing a Chain-of-Thought (CoT)~ cot prompting paradigm to guide the models through library importation, data preparation, and API invocation. For each library, we randomly sample 300 distinct APIs and employ two open-source models, CodeLlama-13B~ codellama and Qwen2.5-Coder-14B~ qwen2.5-coder, to generate one test case per API. We select models at this scale as they strike an optimal balance between robust reasoning capabilities and local reproducibility. This setup yields a total of 600 test cases per library, allowing us to evaluate the modelsâ zero-shot generation capabilities without domain-specific knowledge. 48 49 Quantifying the performance gap. 50As illustrated in Figure~ fig:fuzzing_eval, our results reveal a stark disparity in code validity between domains. While classical libraries achieve a pass rate of 64.50\% to 86.50\%, the validity for quantum libraries plummets to a range of 35.33\% to 46.00\%. Our empirical results in Figure~ fig:fuzzing_eval show that TypeError, AttributeError, and ValueError are the most frequent errors ( e.g., 164 TypeErrors for Qiskit), while ImportError and ModuleNotFoundError are also common. These findings demonstrate that despite their strong general coding capabilities, LLMs exhibit a pronounced lack of domain awareness when encountering quantum-specific syntax and semantics. 51 52 figure[t] 53 54 [width=0.96 ]fuzzing_performance_comparison.pdf 55 Empirical evaluation of the validity and error patterns in LLM-generated fuzzing test cases across quantum and classical libraries. The bars represent absolute error counts (left y-axis), while the dashed line indicates the validity rate (right y-axis). 56 fig:fuzzing_eval 57 figure 58 59 Root cause analysis. 60A representative example illustrated in Figure~ fig:crash_show reveals that such failures are primarily attributed to the volatile evolution and frequent breaking changes inherent in quantum libraries. 61Unlike classical domains where software updates are primarily driven by application logic, these reconfigurations are uniquely induced by the rapid advancement of quantum hardware. 62As new functionalities and physical architectures emerge, library maintainers must frequently reconfigure invocation patterns and structural placements to accommodate hardware-side capabilities. Such instability leads to inconsistent internal representations within LLMs, heightening their proclivity for generating deprecated or hallucinated code. Furthermore, unlike the mature classical computing ecosystem, there is a pronounced scarcity of publicly available codebases aligned with the latest versions of quantum libraries. Consequently, the available resources for knowledge acquisition are largely restricted to the libraries themselves and the LLMâs own (often flawed) outputs. These observations motivate us to investigate the following guiding questions: 182 How can library-specific knowledge be leveraged to guide LLMs in generating valid seed programs, and 63 183 how can these seeds be systematically evolved to expose latent bugs in quantum libraries? 64 65 figure 66 67 [width=3.2in]crash_show.pdf 68 Code example of typical errors in quantum computing libraries generated by LLMs. fig:crash_show 69 figure 70 71 Challenge 72To address the aforementioned guiding questions, a fundamental prerequisite is to understand the key challenges that arise when applying LLM-based fuzzing to quantum libraries. These challenges are inherently progressive: generating valid, library-aligned test cases is the foundation, upon which effective exploration through mutation must be built. 73 74 figure* 75 76 [width=0.99 ]wfl.pdf 77 The workflow of for detecting bugs in quantum libraries. fig:workflow 78 figure* 79 80 Challenges in generating valid quantum seeds. 81LLMs struggle to generate valid test cases for quantum libraries due to the rapid evolution and frequent restructuring of APIs, which lead to outdated or hallucinated invocations. This issue is further exacerbated by the scarcity of up-to-date quantum code, limiting the modelâs ability to learn stable usage patterns and diverse invocation semantics. Furthermore, since fuzzing is a resource-intensive process requiring high-volume iterations, a fundamental challenge lies in balancing computational cost with testing effectiveness. As a result, seed generation becomes a knowledge-grounding problem, requiring alignment with current library implementations ( i.e., source code), while also balancing cost and effectiveness for large-scale fuzzing with open-source models. 82 83 Constraints and inefficiencies in quantum exploration. 84The subsequent exploration phase remains challenging even with valid seeds. Effective exploration of the program space typically relies on an iterative process of mutation, fitness evaluation, and input regeneration. However, directly applying traditional fuzzing strategies in the quantum setting results in severely limited exploration. First, at the mutation level, quantum programs are governed by hardware-imposed structural constraints and framework-defined execution semantics. Generic mutation operators (e.g., byte-level or AST-level transformations) often fail to preserve these constraints, producing invalid programs and significantly reducing the proportion of executable test cases. Second, at the evaluation and selection level, even when syntactically and structurally valid variants are generated, the lack of domain-specific fitness metrics limits the ability to effectively prioritize high-value seeds. Without quantum-aware fitness functions to measure traits like quantum circuit complexity, the fuzzer cannot prioritize high-potential candidates. Consequently, the entire exploration loop becomes inefficient, trapped in generating redundant inputs with limited capability to evolve simple seeds into complex test cases. In this regard, holistic and domain-specific adaptations across the mutation and selection mechanisms are highly demanded to expose latent bugs. 85 86 Approach 87 sec:approach 88KQFuzz addresses a central challenge in quantum-library fuzzing: generating tests that remain executable under evolving APIs while being diverse enough to exercise deep library logic. Existing circuit-level fuzzers provide high validity but limited API coverage, whereas general-purpose LLM fuzzers cover broader APIs but often rely on outdated API knowledge and generate invalid tests. Moreover, generic mutation and selection strategies lack quantum-aware guidance for parameterized gates, entanglement structures, and multi-API interactions. 89 90To bridge this gap, KQFuzz employs a four-stage validity-to-diversity pipeline as illustrated in ~ fig:workflow. First, in Library Knowledge Extraction (Stage 1, ~ sec:know_extra), constructs an API corpus by extracting structural associations, semantic profiles, and evolution metrics from the library code. Subsequently, during the Seed Generation (Stage 2, ~ sec:seed_gen) phase, a Fuzzing LLM takes seeds from the pool and incrementally incorporates target APIs to generate diverse quantum programs. Next, in Seed Selection (Stage 3, ~ sec:seed_selection), the generated programs are executed and assessed via a multi-dimensional fitness function. The generated quantum programs with high semantic complexity are fed back into the adjusted seed pool for the next iteration. Finally, the Mutation (Stage 4, ~ sec:mutation_oracle) phase systematically transforms quantum programs to expand the search space, while a crash oracle and semantic filters are employed to identify potential bugs. 91 92 Knowledge Modeling from Library Code 93 sec:know_extra 94 95As identified in our pilot study ( ~ sec:challenge), LLM-based fuzzers 96frequently generate invalid seed programs due to outdated library knowledge or 97hallucinated API usage. To mitigate these issues, constructs an explicit 98knowledge model directly from the source code of the target quantum library. 99Rather than relying on parametric knowledge embedded in the LLM, extracts 100lightweight, library-specific structural and semantic cues and organizes them 101into an API corpus that guides subsequent seed generation. 102 103 API inventory and static metadata. 104 begins by constructing an inventory of available APIs from the official 105documentation of the target software. Let $I = \1, 2, âŠ, N\$ 106denote the resulting set of APIs. For each API $i â I$, 107collects static metadata by inspecting the corresponding source files, including 108the API signature, its implementation body, and its file and module location. 109This information serves as the foundation for modeling relationships among APIs 110and for providing accurate and version-aligned context to the LLM during seed 111synthesis. 112 113 API association modeling. 114To capture semantic and structural relationships between APIs, models pairwise API associations using three complementary dimensions. These dimensions are designed as lightweight signals derived from source-level inspection. In particular, Proximity, denoted by $S_p(a,b)$, measures the structural locality via file and module boundaries; Type Overlap, denoted by $S_t(a,b)$, quantifies the shared domain-specific input and output types; and Call Relationships, denoted by $S_c(a,b)$, evaluates direct reachability and shared invocation patterns. We aggregate these three dimensions into a joint association score, i.e., 115 equation 116 eq:score_ab 117S(a,b) = w_1 S_p(a,b) + w_2 S_t(a,b) + w_3 S_c(a,b), 118 equation 119where $w_1$, $w_2$, and $w_3$ control the relative importance of these dimensions. Detailed mathematical formulations for each dimension are provided in SM~B. 120 121 API implementation of semantic modeling. introduces an API implementation of a semantic model by prompting an LLM with source code to infer descriptions, constraints, outputs, and fuzzing points (detailed prompt in SM~C). The modeling LLM is limited to this semantic summarization; API inventory construction and the computation of association and evolution metrics remain non-LLM source-analysis tasks. The resulting semantic model is a structured semantic summary grounded in the library source code. Unlike resource-intensive seed generation, semantic extraction is a one-time preprocessing step. We employ Gemini-3~ team2023gemini to synthesize library profiles, using its extended context window to capture long-range dependencies in complex codebases. To clarify the distinct roles of LLMs within our framework, we designate the LLM used for this extraction as the modeling LLM, while the LLM responsible for subsequent seed generation is referred to as the fuzzing LLM. 122 123 API evolution modeling. In addition to structural associations, 124models API evolution across library versions. APIs that undergo frequent 125modifications may expose unstable or under-tested behaviors and thus represent 126promising fuzzing targets. For a given API $iâ I$, denote $B_i $ and $B_i$ 127as the implementation in two successive versions, with $B_i $ 128preceding $B_i$. We quantify the evolution weight $Mâ(i)$ as 129\[ 130Mâ(i) = 1 - |G_n(B_i) â© G_n(B_i )||G_n(B_i) âȘ G_n(B_i )|, 131\] 132where $G_n(·)$ extracts $n$-gram sets from the normalized Abstract Syntax Tree (AST) node sequences of the implementations. To further prioritize unstable APIs, we map $Mâ(i)$ to the degree of modification: 133 equation eq:Mi 134 M(i) = 1 + ÎŽ · (Îł(Mâ(i) - 0.5)). 135 equation 136This non-linear mapping amplifies the testing priority of high-evolution APIs while suppressing that of stable ones, where $Îł$ and $ÎŽ$ are hyper-parameters controlling the sensitivity and influence range, respectively. 137 138 . All extracted metadata, association scores, semantic model, and evolution metrics are integrated into a unified API corpus, which is subsequently used to guide API selection and prompt construction during seed generation for fuzzing. 139 140 LLM-Guided Seed Generation 141 sec:seed_gen 142Building on the API corpus constructed in ( ~ sec:know_extra), 143iteratively generates seed programs for fuzzing by combining probabilistic API 144selection with LLM-based code synthesis. The key idea is to incrementally extend 145existing quantum programs with carefully selected APIs, while using 146library-specific knowledge to constrain and guide the LLM toward valid and 147semantically meaningful code. 148 149 Probabilistic target API selection. 150Given a seed program containing an invocation of API $iâ I$, queries the API 151corpus to retrieve the set of related APIs, i.e., 152\[ 153S_i = \ j â I S(i,j) â 0 \, 154\] 155where $S(i,j)$ is the association score defined in ~( eq:score_ab). 156Instead of selecting a related API uniformly at random, leverages the 157metrics obtained from knowledge modeling to bias selection toward APIs that are 158both strongly associated with the seed API and actively evolving. Specifically, 159it assigns each candidate API $k â S_i$ a selection probability: 160 equation 161 eq:select_pr 162Pr[k;i] = 163 M(k) · S(i,k) 164ÎŁ _j â S_i M(j) · S(i,j) , 165 equation 166where $S(i, k)$ and $M(k)$ are defined in Eqn.~ eq:score_ab and Eqn.~ eq:Mi, respectively. As a result, prioritizes 167semantically meaningful API combinations while avoiding the forced integration 168of unrelated APIs that often lead to invalid or trivial programs. 169 170 Prompt construction and code generation. Given a selected seed program 171and a target API $i_t $ sampled according to the probabilities in 172 ~( eq:select_pr), invokes the fuzzing LLM to synthesize a new seed program 173that incorporates $i_t$ into the seed program. The goal of prompt construction 174is to (i) ground the LLM in the exact library version under test, (i) 175provide sufficient semantic context to avoid hallucinated usage, and (i) steer 176generation toward either introducing $i_t$ or exercising it thoroughly, depending on the stage of fuzzing. 177 178Accordingly, constructs prompts by integrating three key sources of information: (i) the selected seed program, which provides the execution context and structural skeleton; (i) the semantic correlation between existing APIs in the seed and the target API $i_t$, ensuring that the newly introduced API is contextually relevant; and (i) the API implementation semantic model of $i_t$ (detailed prompt in SM~C). Unlike approaches that rely on the LLMâs internal parametric knowledge, provides an explicit knowledge representation synthesized from the source code. This semantic model encapsulates inferred descriptions, argument constraints, expected outputs, and critical fuzzing points. By presenting these structured semantic profiles, explicitly grounds the generation process in the exact library logic, exposing intricate dependencies and usage patterns that are often absent from outdated training data. 179Moreover, by promoting the generation of programs that integrate multiple semantically related APIs, produces seed programs that exercise more complex and diverse execution behaviors, thereby increasing coverage of the target quantum library. 180 181 Comparison with standard RAG. While the integration of an API corpus shares the high-level philosophy of Retrieval-Augmented Generation (RAG)~ lewis2020retrieval, namely augmenting LLMs with external knowledge, our mechanism differs fundamentally from standard RAG pipelines. Specifically, standard RAG relies on vectorizing documents and retrieving text chunks based on semantic similarity to a user query. However, in the highly constrained domain of quantum libraries, semantic similarity alone does not ensure structural compatibility or valid program execution. In addition, instead of treating knowledge as isolated text snippets, our approach constructs a structured and multi-dimensional knowledge model that captures API relationships, semantic properties, and historical evolution. This design enables the fuzzer to preserve structural constraints while prioritizing components with higher defect potential. As a result, our approach provides a more principled and robust basis than conventional text-based retrieval methods. 182 183 Seed Selection 184 sec:seed_selection 185 186 follows an iterative generation-selection loop. At each iteration, a batch 187of new seed programs is generated from the current seed pool, after which 188high-quality programs are selected and promoted as seeds for subsequent 189iterations. Below, we separately describe how the seed pool is initialized and how the generated programs are prioritized. 190 191 Seed pool initialization. 192 initializes the seed pool using officially provided usage examples 193extracted from API docstrings of the target quantum library. The number of 194initial seeds is comparable to the number of APIs, ensuring broad coverage of 195library functionality. Importantly, all seed programs are derived exclusively 196from the target library itself, avoiding reliance on external quantum code that 197may be outdated or incompatible. 198 199 Fitness-guided seed selection. 200 evaluates generated seed programs using a 201fitness function and prioritizes high-quality programs for reseeding. For each 202generated program, first extracts the largest executable fragment using a 203greedy strategy, ensuring that fitness is computed only on runnable code. For example, consider a 15-line program that crashes at line 12. In this case, the fragment consisting of the first 11 lines constitutes the largest executable fragment. It then analyzes the structure of each generated seed program and assigns a fitness score that quantifies its exploration potential. The fitness function is 204designed to favor programs that exercise complex quantum operations and diverse 205API interactions. Specifically, evaluates each program along the following 206dimensions: Circuit Gate Diversity ($G$). The number of 207distinct quantum gate types appearing in the circuit. Higher diversity indicates 208broader coverage of quantum operations, increasing the likelihood of exposing 209bugs triggered by specific gate combinations. Number of 210Entangled Qubits ($Q$). The number of qubits participating in entangled 211operations. Entanglement amplifies circuit complexity and may expose subtle 212compiler or execution errors, although excessively large entangled states may 213incur the hardness of efficient simulation. API Diversity ($ $). The number of 214distinct APIs invoked in the quantum program. This metric reflects semantic coverage 215across different library components, such as circuit construction, simulation, 216and backend configuration. API Calling Depth ($L$). The 217maximum depth of API call chains, capturing the temporal and logical complexity 218of API interactions. Longer call chains stress long execution paths and stateful 219behaviors. Combining these factors, the fitness score of a generated program $C$ 220is defined as 221 equation 222 FitnessScore(C) = G + (Q,Ï) + L + η· . 223 equation 224To prevent the fuzzer from over-optimizing for entanglement at the expense of executability, the term $Q$ is capped at $Ï$. This design penalizes excessively deep entanglement that threatens to crash the target library. Furthermore, $η$ ($η â„ 1$) serves as a tunable weight parameter designed to balance the diversity of API invocations ($ $) against the other evaluation metrics, ensuring that API diversity is appropriately prioritized without overwhelming the fitness score. 225 226In each iteration, selects the top-$Îș$ programs and adds them to the seed 227pool for the next round. This fitness-guided prioritization biases 228the fuzzing process toward semantically rich test 229cases, enabling progressively broader and more complex exploration of the target quantum library over time. 230 231 Mutation and Test Oracle 232 sec:mutation_oracle 233Following LLM-guided generation, expands the search space through structured mutation and utilizes a crash-based test oracle for bug detection. Mutation 234complements generation by systematically exploring semantic variations that are 235unlikely to be produced by prompting alone, while the test oracle provides 236automated bug detection indicators during execution. 237 238 Mutation. applies semantics-aware mutations to generated test cases 239by directly operating on quantum circuits. The mutation strategy consists of two 240levels: parameter-level mutation and gate-level structural mutation. 241 242 -level mutation. Many quantum gates require numerical 243parameters ( e.g., rotational angles in quantum gates $R_x$, $R_y$, and $R_z$). In 244parameter-level mutation, replaces such parameters with values sampled 245from a predefined set of special constants. These constants include values 246historically associated with bugs as well as typical boundary or representative 247values, such as $± 1$, $0$, and $Ï$. This mutation targets numerical corner 248cases while preserving syntactic and semantic validity of the program. 249 250 -level structural mutation. Unlike parameter-level mutation, which changes numerical values while preserving circuit topology, gate-level mutation modifies multi-qubit interaction structures. substitutes circuit gates using three classes of rules: (i) entanglement-preserving, (i) entanglement-expanding, and (i) entanglement-reducing. These mutations generate structurally different quantum inputs that exercise topology-sensitive logic in compilation, decomposition, simulation, and backend-specific processing. Although such variants may execute already-covered code regions, they can expose failures that are difficult to trigger through parameter mutation alone. 251 252 Test oracle. utilizes a crash-based oracle to identify runtime anomalies within the executed test cases. To maintain high detection precision, performs exception filtering based on error messages and applies predefined heuristic rules to eliminate common false-positive patterns. The remaining candidate anomalies are then manually inspected to determine if they trigger genuine bugs in the target library. This workflow ensures that identifies impactful bugs while minimizing noise from trivial execution failures. 253 254 Experimental Setup 255 Research Questions 256 257To enable a comprehensive and rigorous evaluation of the proposed framework, we articulate three key research questions as follows. 258 itemize[leftmargin=*] 259 [] RQ1: How does perform compared to existing state-of-the-art fuzzing techniques? 260 [] RQ2: What drives âs effectiveness, and how robust is it across different LLM backends and model scales? 261 [] RQ3: To what extent can uncover previously unknown bugs in widely used quantum programming libraries? 262 itemize 263 264 Implementation 265We implement a prototype of as a Python-based framework and conduct a multi-faceted evaluation on a high-performance computing platform. All experiments are performed on a Linux server powered by an Intel Xeon Platinum CPU (3.80 GHz, 192 physical cores), 512 GB RAM, and an NVIDIA Tesla H800 GPU, operating under Ubuntu 24.04 LTS. Our evaluation targets three representative quantum computing libraries: Qiskit, PennyLane, and Cirq. The detailed specifications and statistics of these target libraries are summarized in Table~ tab:quantum_libs. Specifically, the Token column denotes the total number of output tokens generated by Gemini-3~ team2023gemini during API semantic modeling. 266For seed program generation, we leverage the Ollama platform to deploy the INT8-quantized versions of CodeLlama~ codellama and Qwen2.5-Coder~ qwen2.5-coder, balancing inference efficiency with model performance. To maximize hardware throughput and fully utilize the multi-core CPU resources, we implement a multi-processing execution engine that facilitates parallel test execution. Furthermore, to prevent potential resource exhaustion or hangs caused by malformed programs, a strict 20-second execution budget is enforced for each test case. 267To evaluate the effectiveness of , we compare it against three state-of-the-art baselines: two quantum-specific fuzzers, MorphQ~ paltenghi2023morphq and FuzzQ~ shaking2025oopsla, and one LLM-based fuzzer, Fuzz4All~ xia2024fuzz4all. A detailed summary of the core hyperparameters, configuration settings, and description of these baselines is provided in SM~C. 268 269 Evaluation 270 sec:results 271 272 table[t] 273 274 275 Overview of mainstream quantum libraries. 276 tab:quantum_libs 277 tabularlccrrr 278 279Library & Version & GitHub Stars & LoC & API Number & Token \\ 280 281Qiskit & 2.3.0 & 7.1K & 45928 & 788 & 0.38M\\ 282PennyLane & 0.44.0 & 3.1K & 69514 & 1206 & 0.61M\\ 283Cirq & 1.6.1 & 4.9K & 31774 & 866 & 0.41M\\ 284 285 tabular 286 table 287 288 table[t] 289 Comparison with baseline fuzzers on three quantum libraries. table:test0 290 291 292 3pt 293 tabular@c l r r r@ 294 295Target & Fuzzer & 296\# Test Cases & 297Coverage & 298\# Unique Crashes \\ 299 300 4*Qiskit 301 & FuzzQ & 42985 & 11636(25.34\%) & 5 \\ 302 & MorphQ & 25714 & 11396(24.81\%) & 13 \\ 303 & Fuzz4All & 23478 & 24344(53.00\%) & 339 \\ 304 & & 35952 & 29078(63.31\%) & 685 \\ 305 306 2*Pennylane 307 & Fuzz4All & 21823 & 31584(45.44\%) & 426 \\ 308 & & 32068 & 40814(58.71\%) & 1027 \\ 309 310 3*Cirq 311 & FuzzQ & 27067 & 10027(31.56\%) & 3 \\ 312 & Fuzz4All & 22547 & 18286(55.35\%) & 391 \\ 313 & & 34224 & 23446(73.79\%) & 706 \\ 314 315 tabular 316 table 317 318 figure[t] 319 320 [width=0.99 ]unique_coverage_heatmaps.pdf 321 Pairwise unique coverage comparison. 322 fig:heatmap 323 figure 324 325 Comparison with Other Fuzzers 326 327To answer RQ1, we compare with each baseline only on targets evaluated by its released implementation: MorphQ on Qiskit, FuzzQ on Qiskit and Cirq, and Fuzz4All on all three libraries. Although the FuzzQ artifact includes a PennyLane adapter, it is presented only as a future-work demonstration rather than an evaluated configuration. We do not port the baselines to additional target libraries, because the resulting performance could be influenced by our implementation choices and might not faithfully represent the original released tools. For every applicable comparison, each fuzzer was executed for a continuous 24-hour period on the same pinned versions (Qiskit 2.3.0, PennyLane 0.44.0, and Cirq 1.6.1), with baseline configurations and hyperparameters set to the defaults recommended in the original papers. Table~ table:test0 summarizes the number 328of generated programs, the resulting code coverage, and the number of unique 329crashes. Here, we measure Python line coverage using coverage.py while 330excluding the librariesâ internal test files, and we regard a generated program 331as valid only if it executes without runtime exceptions in a properly 332configured environment and invokes the target API at least once. To better 333compare how much new behavior each fuzzer reaches, Figure~ fig:heatmap 334further visualizes the pairwise unique 335coverage relationships between and the baselines. Each cell in the heatmap represents the percentage of the target libraryâs total code lines that are covered by one fuzzer but missed by another. 336 337Experimental results show that consistently achieves the strongest 338overall testing effectiveness across all targets. On Qiskit, reaches 33963.31\% coverage, outperforming Fuzz4All (53.00\%) and more than doubling the 340coverage of MorphQ (24.81\%) and FuzzQ (25.34\%). On PennyLane, 341improves coverage from 45.44\% to 58.71\% over Fuzz4All. On Cirq, 342achieves 73.79\% coverage, compared with 55.35\% for Fuzz4All and 31.56\% for 343FuzzQ. This broader exploration also translates into more bug-finding 344opportunities: reports the highest number of unique crashes on every 345framework, including 685 on Qiskit, 1027 on PennyLane, and 706 on Cirq. 346Moreover, it generates the largest number of programs in all three settings, 347indicating that combines higher throughput with stronger exploration. A detailed analysis of the overall time allocation under the default configuration used in RQ1 is provided in SM~ sm:efficiency_overall. 348 349The pairwise heatmaps in Figure~ fig:heatmap confirm that âs gains represent substantial new behaviors rather than redundant paths. On Qiskit, covers an additional 13.4\% of the total library code that Fuzz4All fails to reach. In contrast, misses only 2.2\% of the library code covered by Fuzz4All. The gap is even more pronounced against specialized quantum fuzzers: MorphQ and FuzzQ fail to cover 41.6\% and 41.1\% of the total library code, respectively, that is otherwise successfully explored by . Similarly, on Cirq, misses only 0.5\% of the code reached by Fuzz4All, while covering an additional 16.6\% of the library. On PennyLane, contributes 14.9\% unique library coverage beyond Fuzz4All, while missing only 2.4\%. These results indicate that largely subsumes the coverage achieved by prior works while significantly expanding the test frontier into previously unreachable regions of the codebases. 350 351 rqbox 352Answer to RQ1: achieves the best results on all three libraries: 35363.31\% coverage on Qiskit, 58.71\% on PennyLane, and 73.79\% on Cirq, all above 354the strongest baselines, triggering 2.1$Ă$ more unique crashes and adding 35515.0\% unique coverage on average. 356 rqbox 357 358 Component Contributions and Cross-Model Robustness 359To address RQ2, we first examine how the core components of contribute to its effectiveness and then evaluate whether its advantage over Fuzz4All persists across different LLM backends and model scales. 360 361 Ablation study of core components. 362We form an ablation over Seed Selection and Codebase Knowledge, while the final configuration measures the incremental benefit of adding Mutation to Seed+Repo. Each configuration is evaluated using 6,000 generated test seeds. Table~ tab:ablation_study reports validity, line coverage, and unique crashes for each configuration. 363 364Codebase Knowledge is the primary contributor to effectiveness. Compared with the base configuration, Repo-only improves validity rate by 8.86--16.70 percentage points (p), coverage by 6.74--15.52 p, and unique crashes by 71--262 across the three libraries. For example, on Cirq, it increases coverage from 54.20\% to 69.72\% and validity from 41.37\% to 50.23\%. These results show that library-specific API and constraint information substantially improve the generation of valid and coverage-effective tests. 365 366Seed Selection and Mutation provide complementary benefits. Compared with the base configuration, Seed-only improves validity by 5.58--6.97 p, while producing smaller coverage and crash gains. Adding Mutation to Seed+Repo improves coverage by only 0.28--0.47 p, but discovers 57, 111, and 61 additional unique crashes on Qiskit, PennyLane, and Cirq, respectively. Thus, Seed Selection primarily improves seed quality, whereas Mutation mainly expands crash-triggering exploration. Mutation is evaluated on top of Seed+Repo because it is designed to refine selected, codebase-aware seeds rather than operate as an independent generator. 367 368 figure[t] 369 370 [width=0.82 ]case_coverage_qwen.pdf 371 Line coverage comparison. 372 fig:coverage_qwen 373 figure 374 375 table[t] 376 377 378 Ablation results on three quantum libraries. 379 tab:ablation_study 380 2pt 381 tabular@lcccccc@ 382 383 2*Target & 3cConfiguration & 2*\% valid & 2*Coverage & 2*Unique Crashes \\ 384 (lr)2-4 385 & Seed & Repo & Mut. & & & \\ 386 387 5*Qiskit 388 & $Ă$ & $Ă$ & $Ă$ & 33.72\% & 22379(48.73\%) & 155 \\ 389 & $ $ & $Ă$ & $Ă$ & 40.69\% & 23189(50.49\%) & 165(+10) \\ 390 & $Ă$ & $ $ & $Ă$ & 50.42\% & 27945(60.85\%) & 226(+71) \\ 391 & $ $ & $ $ & $Ă$ & 53.71\% & 28316(61.65\%) & 250(+95) \\ 392 & $ $ & $ $ & $ $ & 53.71\% & 28530(62.12\%) & 307(+152) \\ 393 394 5*PennyLane 395 & $Ă$ & $Ă$ & $Ă$ & 38.48\% & 33227(47.80\%) & 192 \\ 396 & $ $ & $Ă$ & $Ă$ & 44.06\% & 33632(48.38\%) & 206(+14) \\ 397 & $Ă$ & $ $ & $Ă$ & 49.53\% & 37910(54.54\%) & 454(+262) \\ 398 & $ $ & $ $ & $Ă$ & 51.34\% & 38287(54.86\%) & 476(+284) \\ 399 & $ $ & $ $ & $ $ & 51.34\% & 38330(55.14\%) & 587(+395) \\ 400 401 5*Cirq 402 & $Ă$ & $Ă$ & $Ă$ & 41.37\% & 17222(54.20\%) & 165 \\ 403 & $ $ & $Ă$ & $Ă$ & 48.19\% & 17427(54.85\%) & 179(+14) \\ 404 & $Ă$ & $ $ & $Ă$ & 50.23\% & 22154(69.72\%) & 251(+86) \\ 405 & $ $ & $ $ & $Ă$ & 53.08\% & 22540(70.94\%) & 283(+118) \\ 406 & $ $ & $ $ & $ $ & 53.08\% & 22683(71.39\%) & 344(+179) \\ 407 408 tabular 409 table 410 411 table[t] 412 413 414 Comparative analysis of validity and coverage for and Fuzz4All across various LLMs. 415 tab:rq2_table 416 tabular@llcc@ 417 418Model & Method & Val. Rate (\%) $ $ & Total Coverage (Unique) \\ 419 2*Qwen2.5-Coder-3B & Fuzz4All & 21.04 & 57226 (4162) \\ 420 & & 33.16 (+12.12) & 73079 (20015) \\ 421 2*Qwen2.5-Coder-7B & Fuzz4All & 25.51 & 66866 (6004) \\ 422 & & 40.64 (+15.13) & 80203 (19341) \\ 423 2*Qwen2.5-Coder-14B & Fuzz4All & 28.65 & 71741 (7131) \\ 424 & & 52.99 (+24.34) & 82317 (17707) \\ 425 2*CodeLlama-7B & Fuzz4All & 14.42 & 53334 (1968) \\ 426 & & 23.90 (+9.48) & 70893 (19527) \\ 427 2*CodeLlama-13B & Fuzz4All & 9.66 & 57300 (3104) \\ 428 & & 31.04 (+21.38) & 73364 (19168) \\ 429 tabular 430 table 431 432 Robustness across LLM backends and scales. 433To further verify the robustness of our framework, we evaluated across two 434prominent model families (detailed results for CodeLlama are available in SM~D), and compared 435them against the baseline. As illustrated in Figure~ fig:coverage_qwen, consistently 436outperforms the baseline across all tested model scales and library targets. We 437observe a striking performance crossover in Qiskit and Cirq, where 438utilizing the smallest model scale achieves significantly higher code 439coverage than the baseline framework at its largest scale. This massive performance gap indicates that the 440effectiveness of is primarily driven by our specialized, library-aware 441generation strategy rather than the mere scaling of model parameters. 442 443While both methods tend to benefit from increased model capacity, exhibits higher scaling efficiency. Table~ tab:rq2_table shows that 444 maintains a higher validity rate, with absolute improvements 445over the baseline ranging from 9.48 to 24.34 p. Notably, the advantage of 446 often becomes more pronounced as the model size increases, with the 447validity improvement peaking at 24.34\% on the 14B scale. This suggests that our 448framework can better unlock the reasoning potential of larger models. As shown 449by the growth curves in Figure~ fig:coverage_qwen, Fuzz4All coverage is frequently constrained by 450a high rate of execution exceptions. In contrast, maintains a higher 451validity baseline, which allows larger models to focus more on handling complex 452API constraints rather than basic syntax, thereby achieving 453higher final coverage. These results demonstrate 454that is a robust, model-agnostic framework that can effectively leverage 455the latent capabilities of different model families and scales. The corresponding per-sample token consumption and generation time of and Fuzz4All across these model configurations are reported in SM~ sm:efficiency_llm. 456 457 rqbox 458Answer to RQ2: KQFuzzâs effectiveness is driven by its core components, primarily Codebase Knowledge, which boosts coverage by up to 16.1\% ( e.g., on Cirq). Furthermore, KQFuzz exhibits superior scaling efficiency across LLMs, improving validity rates by 9.48\% to 24.34\% and achieving significantly faster coverage convergence. 459 rqbox 460 461 table[t] 462 463 Bugs found by 464 tab:bugs 465 466 tabularlllc l 467 468Target & ID & API Location & Type & Status \\ 469 470 6*Qiskit 471 & \#15550 & synthesis & BV & Fixed \\ 472 & \#4159 & synthesis & SV & Confirmed \\ 473 & \#15657 & circuit & SD & Fixed \\ 474 & \#15665 & synthesis & BV & Fixed \\ 475 & \#15666 & circuit.library & BV & Fixed \\ 476 & \#15780 & synthesis & BV & Fixed \\ 477 478 4*PennyLane 479 & \#8931 & io.qasm\_interpreter & SD & Fixed \\ 480 & \#8726 & decomposition & SV & Fixed \\ 481 & \#9140 & math & SD & Fixed \\ 482 & \#9141 & transforms & SV & Fixed \\ 483 484 3*Cirq 485 & \#7934 & contrib.paulistring & SD & Fixed \\ 486 & \#7939 & transformers & SV & Fixed \\ 487 & \#7941 & neutral\_atoms & SD & Fixed \\ 488 489 tabular 490 table 491 492 Bug Detection 493Regarding RQ3, as of the submission date, has identified 13 bugs, with 13 confirmed by developers and 10 already fixed. Table~ tab:bugs summarizes the identified bugs. We categorize these bugs into three distinct types based on their root causes: (i) Boundary violations (BV). These errors occur when the library fails to handle extreme input dimensions or edge-case configurations. Such bugs typically bypass high-level sanity checks and manifest as low-level system panics, memory exhaustion, or process hangs. (i) State divergence (SD). This category encompasses defects rooted in flawed internal implementations, organizational structures, or data processing logic. This type of bug usually emerges during multi-stage transformations, where internal metadata fails to maintain synchronized states. (i) Semantic violations (SV). These defects refer to implementation behaviors that deviate from documented specifications or logical protocols. This includes incorrect gate decomposition logic, inconsistencies between documentation and implementation, or violations of API invariants. 494 495Listing~ code:qiskit_forloop illustrates a bug case where assigning parameters to a circuit containing ForLoopOp leads to a state divergence, resulting in an inconsistency error. This defect arises from redundant tracking: the loop parameter is 496incorrectly registered in the outer circuitâs parameter table via both the 497control-flow operator slot and the nested circuit body. Such dual-source 498registration creates a synchronization gap between the global and local 499parameter scopes, causing the assignment mechanism to fail. To understand the 500bug, we quote the maintainerâs response to our bug report. 501 quote 502Huh, this seems like it should really have been an error on all Qiskit versions, and the error shouldnât get as far as the internal logic error. 503 quote 504Their response confirms that this issue has silently existed across all Qiskit 505releases and can only be triggered through the specific parameter-reuse pattern 506generated by our fuzzer, representing a scenario unlikely to be covered by 507existing fuzzers. Additional case studies illustrating other bug categories are provided in SM~D. 508 509 lstlisting[style=QuantumPyBordered, caption=Nested Parameter Inconsistency in Qiskit, label=code:qiskit_forloop] 510from qiskit.circuit import QuantumCircuit, Parameter 511from qiskit.circuit.controlflow import ForLoopOp 512 513theta = Parameter(âthetaâ) 514body = QuantumCircuit(1) 515body.rx(theta, 0) 516for_loop_op = ForLoopOp(range(3), theta, body) 517qc = QuantumCircuit(1) 518qc.append(for_loop_op, [0]) 519# This line triggers the error 520qc.assign_parameters(theta: 3.14159 / 2, inplace=True) lstlisting 521 522 rqbox 523Answer to RQ3: successfully identified 13 bugs across major quantum libraries, with 12 already fixed by developers. By exposing critical boundary, state, and semantic violations that elude conventional testing, our tool demonstrates exceptional capability in ensuring the functional reliability of the evolving quantum libraries. 524 rqbox 525 526 Threats to Validity 527There are some threats to the validity of our results and the conclusions drawn from them. First, the nondeterministic nature of LLMs could introduce variability in our metrics. We mitigate this threat to validity by running each fuzzer for 24 hours to reduce measurement variability. Second, our crash oracle relies on heuristic exception filtering and deduplication rules to eliminate false positive patterns. This presents a threat because such mechanisms might inadvertently discard genuine bugs that share error signatures with trivial failures, leading to unreported bugs. Finally, our evaluation focuses on three quantum libraries and specific open source models, and performance might differ on other paradigms. We believe our framework can generalize to other platforms that provide sufficient source code transparency for knowledge extraction. 528 529 Related Work 530 Quantum software engineering. The rapid emergence of quantum computing has driven the need for robust software engineering practices across the entire development lifecycle~ zhao2020quantum,ali2022taoyue,murillo2025quantum,oldfield2025faster,wang2024quantum,muqeet2024machine,mendiluze2025quantum,xia2025quantum,ye2025measurement,guo2025m2qcode,jin2025novaq,long2024equivalence,guo2024repairing,zhao2023bugs4q,arcaini2025introduction,muqeet2024mitigating. Recent advancements have pushed the boundaries of quantum software development, introducing LLM-assisted quantum code synthesis ( e.g.,~ guo2025quanbench), and specialized compiler optimizations 531~ wang2024dac. Alongside these development tools, quality assurance mechanisms are actively being established to ensure platform reliability. Empirical studies~ paltenghi2022bugs have characterized unique bug patterns in quantum software, motivating the creation of static analysis tools like Qchecker~ zhao2023qchecker to detect issues in quantum programs prior to execution. Program-level techniques such as Muskit and QuanFuzz respectively apply syntactically valid mutations and generate test inputs for quantum programs~ muskit2021ase,poster2021icst. In contrast, QDiff, MorphQ, and FuzzQ target quantum platforms or libraries and use quantum programs as test inputs~ wang2021qdiff,paltenghi2023morphq,shaking2025oopsla. QEMI adapts EMI to quantum software-stack testing by removing dead code constructed from quantum control-flow patterns and checking the resulting variants for crash and output-distribution inconsistencies~ luo2026qemi. instead uses source-derived API knowledge to guide program generation toward broad quantum-library API exploration. 532 533 Testing compilers and other developer tools. 534The critical role of complex software infrastructure has motivated extensive research into automated tool testing. Compiler testing frequently employs randomized code generation and equivalence modulo inputs to uncover optimization bugs. Recently, this scope has expanded to lower-level assemblers via error-driven grammar inference~ kim2024asfuzzer. Beyond foundational compilers, specialized techniques target diverse input spaces: WebAssembly engines utilize stack-invariant transformations for semantic-preserving mutation~ zhang2025waltzz, network protocols employ constrained query-response fuzzing to uncover stateful bugs~ zhang2024resolverfuzz, and database management systems leverage dynamic data-dependency analysis for generic query synthesis~ yang2024buzzbee. Furthermore, LLMs are revolutionizing the field by acting as universal, language-agnostic fuzzers~ xia2024fuzz4all and automating the generation of complex API fuzz drivers~ zhang2024effective. 535 536 Conclusion 537This paper presents , a knowledge-guided fuzzing framework for quantum libraries that leverages LLMs to enhance their performance and efficiency. By leveraging a multi-dimensional API knowledge corpus to guide LLM-based seed generation, coupled with fitness-driven seed selection and two-level mutations, successfully overcomes the limitations of insufficient flexibility in traditional circuit-level fuzzers and the low validity rates of generic LLM-based approaches. Experimental evaluations on major quantum libraries, including Qiskit, PennyLane, and Cirq, demonstrate that significantly outperforms state-of-the-art baselines in both code coverage and quantum program validity. Furthermore, the discovery of 13 bugs underscores the frameworkâs practical effectiveness in uncovering critical bugs within the rapidly evolving quantum computing ecosystem. 538 539 Data Availability 540The source code and data involved in our study are publicly available in the Zenodo repository~ lk_qfuzz_repo. 541 542 thebibliography68 543 544%%% ==================================================================== 545%%% NOTE TO THE USER: you can override these defaults by providing 546%%% customized versions of any of these macros before the 547%%% command. Each of them MUST provide its own final punctuation, 548%%% except for , , and . The latter two 549%%% do not use final punctuation, in order to avoid confusing it with 550%%% the Web address. 551%%% 552%%% To suppress output of a particular field, define its macro to expand 553%%% to an empty string, or better, , like this: 554%%% 555%%% [1] % LaTeX syntax 556%%% 557%%% #1 % plain TeX syntax 558%%% 559%%% ==================================================================== 560 561 #1 562 #1#1 563 #1 564 #1 565 #1 566 #1 567 #1#1 568 #1#1 569 570% The following commands are used for tagged output and should be 571% invisible to TeX 572 [2]#2 573 [2]#2 574 [1]#1 575 [2][]arXiv:#2 576 577 [Abbas et~al .(2024)]% 578 abbas2024nrp 579 author personAmira Abbas, personAndris 580 Ambainis, personBrandon Augustino, personAndreas 581 B\"artschi, personHarry Buhrman, personCarleton 582 Coffrin, personGiorgio Cortiana, personVedran 583 Dunjko, personDaniel~J Egger, personBruce~G 584 Elmegreen, et~al . year2024 . 585 Challenges and opportunities in quantum 586 optimization. 587 journal Nature Reviews Physics volume6, 588 number12 ( year2024), pages718--735. 589 590 591 592 [Ali et~al .(2022)]% 593 ali2022taoyue 594 author personShaukat Ali, personTao Yue, 595 and personRui Abreu. year2022 . 596 When software engineering meets quantum computing. 597 journal Commun. ACM volume65, 598 number4 ( year2022), pages84--88. 599 600 601 602 [Anonymous Authors(2026)]% 603 lk_qfuzz_repo 604 author personAnonymous Authors. 605 year2026 . 606 titleReplication package for anonymous submission. 607 608 609 % 610 https://doi.org/10.5281/zenodo.19230459 611 612 613 614 [Arcaini et~al .(2025)]% 615 arcaini2025introduction 616 author personPaolo Arcaini, personAndriy 617 Miranskyy, and personHausi M\"uller. 618 year2025 . 619 titleIntroduction to the Special Section on software 620 engineering for hybrid quantum computing systems. 621 , numpages112362~pages. 622 623 624 625 [Assolini et~al .(2024)]% 626 assolini2024sas 627 author personNicola Assolini, 628 personAlessandra Di~Pierro, and personIsabella 629 Mastroeni. year2024 . 630 Static analysis of quantum programs. In 631 booktitle International Static Analysis Symposium. 632 Springer, pages1--25. 633 634 635 636 [Bergholm et~al .(2018)]% 637 pennylane 638 author personVille Bergholm, personJosh 639 Izaac, personMaria Schuld, personChristian Gogolin, 640 personShahnawaz Ahmed, personVishnu Ajith, 641 personM~Sohaib Alam, personGuillermo Alonso-Linaje, 642 personBharath AkashNarayanan, personAli Asadi, 643 et~al . year2018 . 644 Pennylane: Automatic differentiation of hybrid 645 quantum-classical computations. 646 journal arXiv preprint arXiv:1811.04968 647 ( year2018). 648 649 650 651 [B\"ohme et~al .(2020)]% 652 bohme2020boosting 653 author personMarcel B\"ohme, 654 personValentin~JM Man\âes, and personSang~Kil 655 Cha. year2020 . 656 Boosting fuzzer efficiency: An information 657 theoretic perspective. In booktitle Proceedings of the 28th 658 ACM Joint Meeting on European Software Engineering Conference and Symposium 659 on the Foundations of Software Engineering. pages678--689. 660 661 662 663 [B\"ohme et~al .(2017)]% 664 bohme2017directed 665 author personMarcel B\"ohme, 666 personVan-Thuan Pham, personManh-Dung Nguyen, and 667 personAbhik Roychoudhury. year2017 . 668 Directed greybox fuzzing. In 669 booktitle Proceedings of the 2017 ACM SIGSAC conference on 670 computer and communications security. pages2329--2344. 671 672 673 674 [Cerezo et~al .(2021)]% 675 cerezo2021nrp 676 author personMarco Cerezo, personAndrew 677 Arrasmith, personRyan Babbush, personSimon~C 678 Benjamin, personSuguru Endo, personKeisuke Fujii, 679 personJarrod~R McClean, personKosuke Mitarai, 680 personXiao Yuan, personLukasz Cincio, 681 et~al . year2021 . 682 Variational quantum algorithms. 683 journal Nature Reviews Physics volume3, 684 number9 ( year2021), pages625--644. 685 686 687 688 [Cirq Developers(2025)]% 689 cirq_developers_2025_10.5281/zenodo.4062499 690 author personCirq Developers. 691 year2025 . 692 booktitle Cirq. 693 694 % 695 https://doi.org/10.5281/zenodo.4062499 696 697 698 699 [Deng et~al .(2023)]% 700 deng2023issta 701 author personYinlin Deng, 702 personChunqiu~Steven Xia, personHaoran Peng, 703 personChenyuan Yang, and personLingming Zhang. 704 year2023 . 705 Large language models are zero-shot fuzzers: 706 Fuzzing deep-learning libraries via large language models. In 707 booktitle Proceedings of the 32nd ACM SIGSOFT international 708 symposium on software testing and analysis. pages423--435. 709 710 711 712 [Deng et~al .(2024)]% 713 deng2024icse 714 author personYinlin Deng, 715 personChunqiu~Steven Xia, personChenyuan Yang, 716 personShizhuo~Dylan Zhang, personShujing Yang, and 717 personLingming Zhang. year2024 . 718 Large language models are edge-case generators: 719 Crafting unusual programs for fuzzing deep learning libraries. In 720 booktitle Proceedings of the 46th IEEE/ACM international 721 conference on software engineering. pages1--13. 722 723 724 725 [Dong et~al .(2025)]% 726 dong2025tdsc 727 author personRuiqi Dong, personFanke Tong, 728 personHe Huang, personXiaogang Zhu, 729 personXi Xiao, personShaohua Wang, 730 personSheng Wen, and personYang Xiang. 731 year2025 . 732 One Mutation Fits All: Exploring Universal Library 733 Fuzzing based on Exogenous Mutation. 734 journal IEEE Transactions on Dependable and Secure 735 Computing ( year2025). 736 737 738 739 [Fang and Ying(2024)]% 740 fang2024pldi 741 author personWang Fang and personMingsheng 742 Ying. year2024 . 743 Symbolic execution for quantum error correction 744 programs. 745 journal Proceedings of the ACM on Programming 746 Languages volume8, numberPLDI 747 ( year2024), pages1040--1065. 748 749 750 751 [Fioraldi et~al .(2020)]% 752 fioraldi2020afl 753 author personAndrea Fioraldi, personDominik 754 Maier, personHeiko Ei feldt, and personMarc 755 Heuse. year2020 . 756 $\$AFL++$\$: Combining incremental steps of 757 fuzzing research. In booktitle 14th USENIX workshop on 758 offensive technologies (WOOT 20). 759 760 761 762 [Fowler et~al .(2012)]% 763 fowler2012surface 764 author personAustin~G Fowler, personMatteo 765 Mariantoni, personJohn~M Martinis, and 766 personAndrew~N Cleland. year2012 . 767 Surface codes: Towards practical large-scale 768 quantum computation. 769 journal Physical Review AâAtomic, Molecular, and 770 Optical Physics volume86, number3 771 ( year2012), pages032324. 772 773 774 775 [Gill et~al .(2025)]% 776 gill2025quantumc 777 author personSukhpal~Singh Gill, personOktay 778 Cetinkaya, personStefano Marrone, personDaniel 779 Claudino, personDavid Haunschild, personLeon 780 Schlote, personHuaming Wu, personCarlo Ottaviani, 781 personXiaoyuan Liu, personSree~Pragna Machupalli, 782 et~al . year2025 . 783 Quantum computing: Vision and challenges. 784 In booktitle Quantum computing. 785 publisherElsevier, pages19--42. 786 787 788 789 [Gill et~al .(2022)]% 790 quantum 791 author personSukhpal~Singh Gill, personAdarsh 792 Kumar, personHarvinder Singh, personManmeet Singh, 793 personKamalpreet Kaur, personMuhammad Usman, and 794 personRajkumar Buyya. year2022 . 795 Quantum computing: A taxonomy, systematic review 796 and future directions. 797 journal Software: Practice and Experience 798 volume52, number1 ( year2022), 799 pages66--114. 800 801 802 803 [Guo et~al .(2025a)]% 804 guo2025m2qcode 805 author personXiaoyu Guo, personShinobu 806 Saito, and personJianjun Zhao. 807 year2025 a. 808 M2QCode: A Model-Driven Framework for Generating 809 Multi-Platform Quantum Programs. 810 journal arXiv preprint arXiv:2510.17110 811 ( year2025). 812 813 814 815 [Guo et~al .(2025b)]% 816 guo2025quanbench 817 author personXiaoyu Guo, personMinggu Wang, 818 and personJianjun Zhao. year2025 b. 819 QuanBench: Benchmarking Quantum Code Generation 820 with Large Language Models. 821 journal arXiv preprint arXiv:2510.16779 822 ( year2025). 823 824 825 826 [Guo et~al .(2024)]% 827 guo2024repairing 828 author personXiaoyu Guo, personJianjun Zhao, 829 and personPengzhan Zhao. year2024 . 830 On repairing quantum programs using ChatGPT. In 831 booktitle Proceedings of the 5th ACM/IEEE International 832 Workshop on Quantum Software Engineering. pages9--16. 833 834 835 836 [Hu et~al .(2024)]% 837 hu2024issta 838 author personTianmin Hu, personGuixin Ye, 839 personZhanyong Tang, personShin~Hwei Tan, 840 personHuanting Wang, personMeng Li, and 841 personZheng Wang. year2024 . 842 Upbeat: Test input checks of q\# quantum 843 libraries. In booktitle Proceedings of the 33rd ACM SIGSOFT 844 International Symposium on Software Testing and Analysis. 845 pages186--198. 846 847 848 849 [Hui et~al .(2024)]% 850 qwen2.5-coder 851 author personBinyuan Hui, personJian Yang, 852 personZeyu Cui, personJiaxi Yang, 853 personDayiheng Liu, personLei Zhang, 854 personTianyu Liu, personJiajun Zhang, 855 personBowen Yu, personKeming Lu, et~al . 856 year2024 . 857 Qwen2. 5-coder technical report. 858 journal arXiv preprint arXiv:2409.12186 859 ( year2024). 860 861 862 863 [Javadi-Abhari et~al .(2024)]% 864 qiskit 865 author personAli Javadi-Abhari, personMatthew 866 Treinish, personKevin Krsulich, personChristopher~J 867 Wood, personJake Lishman, personJulien Gacon, 868 personSimon Martiel, personPaul~D Nation, 869 personLev~S Bishop, personAndrew~W Cross, 870 et~al . year2024 . 871 Quantum computing with Qiskit. 872 journal arXiv preprint arXiv:2405.08810 873 ( year2024). 874 875 876 877 [Jiang et~al .(2024)]% 878 jiang2024fuzzing 879 author personYu Jiang, personJie Liang, 880 personFuchen Ma, personYuanliang Chen, 881 personChijin Zhou, personYuheng Shen, 882 personZhiyong Wu, personJingzhou Fu, 883 personMingzhe Wang, personShanshan Li, 884 et~al . year2024 . 885 When fuzzing meets llms: Challenges and 886 opportunities. In booktitle Companion Proceedings of the 887 32nd ACM International Conference on the Foundations of Software 888 Engineering. pages492--496. 889 890 891 892 [Jin et~al .(2025)]% 893 jin2025novaq 894 author personTiancheng Jin, personShangzhou 895 Xia, and personJianjun Zhao. year2025 . 896 NovaQ: Improving Quantum Program Testing through 897 Diversity-Guided Test Case Generation. 898 journal arXiv preprint arXiv:2509.04763 899 ( year2025). 900 901 902 903 [Kim et~al .(2024)]% 904 kim2024asfuzzer 905 author personHyungseok Kim, personSoomin 906 Kim, personJungwoo Lee, and personSang~Kil Cha. 907 year2024 . 908 AsFuzzer: Differential testing of assemblers with 909 error-driven grammar inference. In booktitle Proceedings of 910 the 33rd ACM SIGSOFT International Symposium on Software Testing and 911 Analysis. pages1099--1111. 912 913 914 915 [Klimis et~al .(2025)]% 916 shaking2025oopsla 917 author personVasileios Klimis, personAvner 918 Bensoussan, personElena Chachkarova, personKarine 919 Even-Mendoza, personSophie Fortz, and personConnor 920 Lenihan. year2025 . 921 Shaking Up Quantum Simulators with Fuzzing and 922 Rigour. 923 journal Proceedings of the ACM on Programming 924 Languages volume9, numberOOPSLA2 925 ( year2025), pages1400--1428. 926 927 928 929 [Lewis et~al .(2020)]% 930 lewis2020retrieval 931 author personPatrick Lewis, personEthan 932 Perez, personAleksandra Piktus, personFabio Petroni, 933 personVladimir Karpukhin, personNaman Goyal, 934 personHeinrich K\"uttler, personMike Lewis, 935 personWen-tau Yih, personTim Rockt\"aschel, 936 et~al . year2020 . 937 Retrieval-augmented generation for 938 knowledge-intensive nlp tasks. 939 journal Advances in neural information processing 940 systems volume33 ( year2020), 941 pages9459--9474. 942 943 944 945 [Li et~al .(2026)]% 946 li2026methodological 947 author personYuechen Li, personMinqi Shao, 948 personJianjun Zhao, and personQichen Wang. 949 year2026 . 950 A Methodological Analysis of Empirical Studies in 951 Quantum Software Testing. 952 journal arXiv preprint arXiv:2601.08367 953 ( year2026). 954 955 956 957 [Lin et~al .(2025)]% 958 lin2025ase 959 author personXingshuang Lin, personQinge 960 Xie, personBinbin Zhao, personYuan Tian, 961 personSaman Zonouz, personNa Ruan, 962 personJiliang Li, personRaheem Beyah, and 963 personShouling Ji. year2025 . 964 PROMFUZZ: Leveraging LLM-Driven and Bug-Oriented 965 Composite Analysis for Detecting Functional Bugs in Smart Contracts. 966 journal arXiv preprint arXiv:2503.23718 967 ( year2025). 968 969 970 971 [Long and Zhao(2024a)]% 972 long2024equivalence 973 author personPeixun Long and personJianjun 974 Zhao. year2024 a. 975 Equivalence, identity, and unitarity checking in 976 black-box testing of quantum programs. 977 journal Journal of Systems and Software 978 volume211 ( year2024), pages112000. 979 980 981 982 [Long and Zhao(2024b)]% 983 long2024tosem 984 author personPeixun Long and personJianjun 985 Zhao. year2024 b. 986 Testing multi-subroutine quantum programs: From 987 unit testing to integration testing. 988 journal ACM Transactions on Software Engineering and 989 Methodology volume33, number6 990 ( year2024), pages1--61. 991 992 993 994 [Luo et~al .(2026)]% 995 luo2026qemi 996 author personJunjie Luo, personShangzhou 997 Xia, personFuyuan Zhang, and personJianjun Zhao. 998 year2026 . 999 QEMI: A Quantum Software Stacks Testing Framework 1000 via Equivalence Modulo Inputs. In booktitle International 1001 Conference on Fundamental Approaches to Software Engineering. Springer, 1002 pages149--169. 1003 1004 1005 1006 [Man\âes et~al .(2019)]% 1007 manes2019tse 1008 author personValentin~JM Man\âes, 1009 personHyungSeok Han, personChoongwoo Han, 1010 personSang~Kil Cha, personManuel Egele, 1011 personEdward~J Schwartz, and personMaverick Woo. 1012 year2019 . 1013 The art, science, and engineering of fuzzing: A 1014 survey. 1015 journal IEEE Transactions on Software Engineering 1016 volume47, number11 ( year2019), 1017 pages2312--2331. 1018 1019 1020 1021 [Mendiluze et~al .(2021)]% 1022 muskit2021ase 1023 author personE\~naut Mendiluze, 1024 personShaukat Ali, personPaolo Arcaini, and 1025 personTao Yue. year2021 . 1026 Muskit: A mutation analysis tool for quantum 1027 software testing. In booktitle 2021 36th IEEE/ACM 1028 International Conference on Automated Software Engineering (ASE). IEEE, 1029 pages1266--1270. 1030 1031 1032 1033 [Mendiluze~Usandizaga et~al .(2025)]% 1034 mendiluze2025quantum 1035 author personE\~naut Mendiluze~Usandizaga, 1036 personShaukat Ali, personTao Yue, and 1037 personPaolo Arcaini. year2025 . 1038 Quantum circuit mutants: Empirical analysis and 1039 recommendations. 1040 journal Empirical Software Engineering 1041 volume30, number4 ( year2025), 1042 pages100. 1043 1044 1045 1046 [Muqeet et~al .(2024a)]% 1047 muqeet2024machine 1048 author personAsmar Muqeet, personShaukat 1049 Ali, personTao Yue, and personPaolo Arcaini. 1050 year2024 a. 1051 A machine learning-based error mitigation approach 1052 for reliable software development on IBMâs quantum computers. In 1053 booktitle Companion Proceedings of the 32nd ACM International 1054 Conference on the Foundations of Software Engineering. 1055 pages80--91. 1056 1057 1058 1059 [Muqeet et~al .(2024b)]% 1060 muqeet2024mitigating 1061 author personAsmar Muqeet, personTao Yue, 1062 personShaukat Ali, and personPaolo Arcaini. 1063 year2024 b. 1064 Mitigating noise in quantum software testing using 1065 machine learning. 1066 journal IEEE Transactions on Software Engineering 1067 volume50, number11 ( year2024), 1068 pages2947--2961. 1069 1070 1071 1072 [Murillo et~al .(2025)]% 1073 murillo2025quantum 1074 author personJuan~Manuel Murillo, personJose 1075 Garcia-Alonso, personEnrique Moguel, personJohanna 1076 Barzen, personFrank Leymann, personShaukat Ali, 1077 personTao Yue, personPaolo Arcaini, 1078 personRicardo P\âerez-Castillo, personIgnacio 1079 Garc\â a-Rodr\â guez~de Guzm\âan, et~al . 1080 year2025 . 1081 Quantum software engineering: Roadmap and 1082 challenges ahead. 1083 journal ACM Transactions on Software Engineering and 1084 Methodology volume34, number5 1085 ( year2025), pages1--48. 1086 1087 1088 1089 [Oldfield et~al .(2025)]% 1090 oldfield2025faster 1091 author personNoah~H Oldfield, personChristoph 1092 Laaber, personTao Yue, and personShaukat Ali. 1093 year2025 . 1094 Faster and better quantum software testing through 1095 specification reduction and projective measurements. 1096 journal ACM Transactions on Software Engineering and 1097 Methodology volume34, number7 1098 ( year2025), pages1--39. 1099 1100 1101 1102 [Paltenghi and Pradel(2022)]% 1103 paltenghi2022bugs 1104 author personMatteo Paltenghi and 1105 personMichael Pradel. year2022 . 1106 Bugs in quantum computing platforms: an empirical 1107 study. 1108 journal Proceedings of the ACM on Programming 1109 Languages volume6, numberOOPSLA1 1110 ( year2022), pages1--27. 1111 1112 1113 1114 [Paltenghi and Pradel(2023)]% 1115 paltenghi2023morphq 1116 author personMatteo Paltenghi and 1117 personMichael Pradel. year2023 . 1118 MorphQ: Metamorphic testing of the Qiskit quantum 1119 computing platform. In booktitle 2023 IEEE/ACM 45th 1120 International Conference on Software Engineering (ICSE). IEEE, 1121 pages2413--2424. 1122 1123 1124 1125 [Roziere et~al .(2023)]% 1126 codellama 1127 author personBaptiste Roziere, personJonas 1128 Gehring, personFabian Gloeckle, personSten Sootla, 1129 personItai Gat, personXiaoqing~Ellen Tan, 1130 personYossi Adi, personJingyu Liu, 1131 personRomain Sauvestre, personTal Remez, 1132 et~al . year2023 . 1133 Code llama: Open foundation models for code. 1134 journal arXiv preprint arXiv:2308.12950 1135 ( year2023). 1136 1137 1138 1139 [Shafiuzzaman et~al .(2024)]% 1140 staticanalysis 1141 author personMd Shafiuzzaman, personAchintya 1142 Desai, personLaboni Sarker, and personTevfik 1143 Bultan. year2024 . 1144 STASE: Static analysis guided symbolic execution 1145 for UEFI vulnerability signature generation. In 1146 booktitle Proceedings of the 39th IEEE/ACM International 1147 Conference on Automated Software Engineering. pages1783--1794. 1148 1149 1150 1151 [She et~al .(2024)]% 1152 she2024ccs 1153 author personDongdong She, personAdam 1154 Storek, personYuchong Xie, personSeoyoung Kweon, 1155 personPrashast Srivastava, and personSuman Jana. 1156 year2024 . 1157 Fox: Coverage-guided fuzzing as online stochastic 1158 control. In booktitle Proceedings of the 2024 on ACM SIGSAC 1159 Conference on Computer and Communications Security. 1160 pages765--779. 1161 1162 1163 1164 [Stephens et~al .(2016)]% 1165 stephens2016driller 1166 author personNick Stephens, personJohn 1167 Grosen, personChristopher Salls, personAndrew 1168 Dutcher, personRuoyu Wang, personJacopo Corbetta, 1169 personYan Shoshitaishvili, personChristopher Kruegel, 1170 and personGiovanni Vigna. year2016 . 1171 Driller: Augmenting fuzzing through selective 1172 symbolic execution.. In booktitle NDSS, 1173 Vol.~ volume16. pages1--16. 1174 1175 1176 1177 [Team et~al .(2023)]% 1178 team2023gemini 1179 author personGemini Team, personRohan Anil, 1180 personSebastian Borgeaud, personJean-Baptiste 1181 Alayrac, personJiahui Yu, personRadu Soricut, 1182 personJohan Schalkwyk, personAndrew~M Dai, 1183 personAnja Hauth, personKatie Millican, 1184 et~al . year2023 . 1185 Gemini: a family of highly capable multimodal 1186 models. 1187 journal arXiv preprint arXiv:2312.11805 1188 ( year2023). 1189 1190 1191 1192 [Upadhyay et~al .(2026)]% 1193 upadhyay2026understandingbugsquantumsimulators 1194 author personKrishna Upadhyay, personMoshood 1195 Fakorede, and personUmar Farooq. 1196 year2026 . 1197 titleUnderstanding Bugs in Quantum Simulators: An 1198 Empirical Study. 1199 1200 1201 [arxiv]2603.22789~[quant-ph] 1202 % 1203 https://arxiv.org/abs/2603.22789 1204 % 1205 1206 1207 1208 [Wang et~al .(2024b)]% 1209 wang2024dac 1210 author personHanrui Wang, personDaniel~Bochen 1211 Tan, personPengyu Liu, personYilian Liu, 1212 personJiaqi Gu, personJason Cong, and 1213 personSong Han. year2024 b. 1214 Q-pilot: Field programmable qubit array compilation 1215 with flying ancillas. In booktitle Proceedings of the 61st 1216 ACM/IEEE Design Automation Conference. pages1--6. 1217 1218 1219 1220 [Wang et~al .(2021a)]% 1221 poster2021icst 1222 author personJiyuan Wang, personFucheng Ma, 1223 and personYu Jiang. year2021 a. 1224 Poster: Fuzz testing of quantum program. In 1225 booktitle 2021 14th IEEE Conference on Software Testing, 1226 Verification and Validation (ICST). IEEE, pages466--469. 1227 1228 1229 1230 [Wang et~al .(2021b)]% 1231 wang2021qdiff 1232 author personJiyuan Wang, personQian Zhang, 1233 personGuoqing~Harry Xu, and personMiryung Kim. 1234 year2021 b. 1235 QDiff: Differential testing of quantum software 1236 stacks. In booktitle 2021 36th IEEE/ACM international 1237 conference on automated software engineering (ASE). IEEE, 1238 pages692--704. 1239 1240 1241 1242 [Wang et~al .(2024a)]% 1243 wang2024quantum 1244 author personXinyi Wang, personShaukat Ali, 1245 personTao Yue, and personPaolo Arcaini. 1246 year2024 a. 1247 Quantum approximate optimization algorithm for test 1248 case optimization. 1249 journal IEEE Transactions on Software Engineering 1250 volume50, number12 ( year2024), 1251 pages3249--3264. 1252 1253 1254 1255 [Wei et~al .(2022)]% 1256 cot 1257 author personJason Wei, personXuezhi Wang, 1258 personDale Schuurmans, personMaarten Bosma, 1259 personFei Xia, personEd Chi, personQuoc~V 1260 Le, personDenny Zhou, et~al . 1261 year2022 . 1262 Chain-of-thought prompting elicits reasoning in 1263 large language models. 1264 journal Advances in neural information processing 1265 systems volume35 ( year2022), 1266 pages24824--24837. 1267 1268 1269 1270 [Wu et~al .(2022)]% 1271 wu2022one 1272 author personMingyuan Wu, personLing Jiang, 1273 personJiahong Xiang, personYanwei Huang, 1274 personHeming Cui, personLingming Zhang, and 1275 personYuqun Zhang. year2022 . 1276 One fuzzing strategy to rule them all. In 1277 booktitle Proceedings of the 44th International Conference on 1278 Software Engineering. pages1634--1645. 1279 1280 1281 1282 [Xia et~al .(2024)]% 1283 xia2024fuzz4all 1284 author personChunqiu~Steven Xia, personMatteo 1285 Paltenghi, personJia Le~Tian, personMichael Pradel, 1286 and personLingming Zhang. year2024 . 1287 Fuzz4all: Universal fuzzing with large language 1288 models. In booktitle Proceedings of the IEEE/ACM 46th 1289 International Conference on Software Engineering. pages1--13. 1290 1291 1292 1293 [Xia et~al .(2025)]% 1294 xia2025quantum 1295 author personShangzhou Xia, personJianjun 1296 Zhao, personFuyuan Zhang, and personXiaoyu Guo. 1297 year2025 . 1298 Quantum concolic testing. 1299 journal Proceedings of the ACM on Software 1300 Engineering volume2, numberISSTA 1301 ( year2025), pages1146--1166. 1302 1303 1304 1305 [Xie et~al .(2025)]% 1306 xie2025ase 1307 author personYuchong Xie, personWenhui 1308 Zhang, and personDongdong She. 1309 year2025 . 1310 ZTaint-Havoc: From Havoc mode to zero-execution 1311 fuzzing-driven taint inference. 1312 journal Proceedings of the ACM on Software 1313 Engineering volume2, numberISSTA 1314 ( year2025), pages917--939. 1315 1316 1317 1318 [Yang et~al .(2024)]% 1319 yang2024buzzbee 1320 author personYupeng Yang, personYongheng 1321 Chen, personRui Zhong, personJizhou Chen, and 1322 personWenke Lee. year2024 . 1323 Towards generic database management system 1324 fuzzing. In booktitle 33rd USENIX Security Symposium (USENIX 1325 Security 24). pages901--918. 1326 1327 1328 1329 [Ye et~al .(2025)]% 1330 ye2025measurement 1331 author personJiaming Ye, personXiongfei Wu, 1332 personShangzhou Xia, personFuyuan Zhang, and 1333 personJianjun Zhao. year2025 . 1334 Is Measurement Enough? Rethinking Output Validation 1335 in Quantum Program Testing. 1336 journal arXiv preprint arXiv:2509.16595 1337 ( year2025). 1338 1339 1340 1341 [Ying et~al .(2025)]% 1342 ying2025tosem 1343 author personMingsheng Ying, personLi Zhou, 1344 and personGilles Barthe. year2025 . 1345 Laws of Quantum Programming. 1346 journal ACM Transactions on Software Engineering and 1347 Methodology ( year2025). 1348 1349 1350 1351 [Zhang et~al .(2024b)]% 1352 zhang2024effective 1353 author personCen Zhang, personYaowen Zheng, 1354 personMingqiang Bai, personYeting Li, 1355 personWei Ma, personXiaofei Xie, 1356 personYuekang Li, personLimin Sun, and 1357 personYang Liu. year2024 b. 1358 How effective are they? exploring large language 1359 model based fuzz driver generation. In booktitle Proceedings 1360 of the 33rd ACM SIGSOFT International Symposium on Software Testing and 1361 Analysis. pages1223--1235. 1362 1363 1364 1365 [Zhang et~al .(2025)]% 1366 zhang2025waltzz 1367 author personLingming Zhang, personBinbin 1368 Zhao, personJiacheng Xu, personPeiyu Liu, 1369 personQinge Xie, personYuan Tian, 1370 personJianhai Chen, and personShouling Ji. 1371 year2025 . 1372 Waltzz:$\$WebAssembly$\$ Runtime Fuzzing with 1373 $\$Stack-Invariant$\$ Transformation. In booktitle 34th 1374 USENIX Security Symposium (USENIX Security 25). 1375 pages6159--6178. 1376 1377 1378 1379 [Zhang et~al .(2024a)]% 1380 zhang2024resolverfuzz 1381 author personQifan Zhang, personXuesong Bai, 1382 personXiang Li, personHaixin Duan, 1383 personQi Li, and personZhou Li. 1384 year2024 a. 1385 $\$ResolverFuzz$\$: Automated Discovery of 1386 $\$DNS$\$ Resolver Vulnerabilities with $\$Query-Response$\$ Fuzzing. In 1387 booktitle 33rd USENIX Security Symposium (USENIX Security 1388 24). pages4729--4746. 1389 1390 1391 1392 [Zhao(2020)]% 1393 zhao2020quantum 1394 author personJianjun Zhao. 1395 year2020 . 1396 Quantum software engineering: Landscapes and 1397 horizons. 1398 journal arXiv preprint arXiv:2007.07047 1399 ( year2020). 1400 1401 1402 1403 [Zhao et~al .(2023a)]% 1404 zhao2023bugs4q 1405 author personPengzhan Zhao, personZhongtao 1406 Miao, personShuhan Lan, and personJianjun Zhao. 1407 year2023 a. 1408 Bugs4Q: A benchmark of existing bugs to enable 1409 controlled testing and debugging studies for quantum programs. 1410 journal Journal of Systems and Software 1411 volume205 ( year2023), pages111805. 1412 1413 1414 1415 [Zhao et~al .(2023b)]% 1416 zhao2023qchecker 1417 author personPengzhan Zhao, personXiongfei 1418 Wu, personZhuo Li, and personJianjun Zhao. 1419 year2023 b. 1420 Qchecker: Detecting bugs in quantum programs via 1421 static analysis. In booktitle 2023 IEEE/ACM 4th 1422 International Workshop on Quantum Software Engineering (Q-SE). IEEE, 1423 pages50--57. 1424 1425 1426 1427 [Zhu et~al .(2022)]% 1428 zhu2022csur 1429 author personXiaogang Zhu, personSheng Wen, 1430 personSeyit Camtepe, and personYang Xiang. 1431 year2022 . 1432 Fuzzing: a survey for roadmap. 1433 journal ACM Computing Surveys (CSUR) 1434 volume54, number11s ( year2022), 1435 pages1--36. 1436 1437 1438 1439 thebibliography 1440 1441 1442 1443 1444 1445This supplementary material (SM) provides necessary details omitted from the main text, including quantum computing background (SM~ sm:background), API association modeling definitions (SM~ sm:api_modeling), implementation and experimental setup of (SM~ sm:implementation), additional evaluation results with bug case studies (SM~ sm:results), and efficiency analysis (SM~ sm:efficiency). 1446 1447 Background of quantum computing 1448 sm:background 1449This section provides a more comprehensive overview of the fundamental concepts in quantum computing referenced in the main text. 1450 1451 Qubits and quantum states. 1452The fundamental unit of quantum information is the quantum bit, or qubit. Unlike a classical bit, which must reside in a discrete state of either 0 or 1, a qubit can exist in a linear combination, or superposition, of these states. Mathematically, a single-qubit state is represented as a vector $|Ï $ in a two-dimensional complex Hilbert space $H C^2$. Using Dirac notation, an arbitrary single-qubit state is expressed as:$$|Ï = α|0 + ÎČ|1 ,$$where $|0 $ and $|1 $ form the standard computational basis, and $α, ÎČ â C$ are complex probability amplitudes satisfying the normalization condition $|α|^2 + |ÎČ|^2 = 1$. For an $n$-qubit system, the state space grows exponentially to $2^n$ dimensions, represented by the tensor product of the individual qubit Hilbert spaces ($H n$). Within this expanded space, entanglement manifests as multi-qubit states that cannot be factored into the tensor product of individual qubit states, representing non-local correlations unique to quantum mechanics. 1453 1454 Quantum gates. 1455The evolution of a closed quantum system is governed by unitary transformations. In the quantum circuit model, these transformations are enacted by quantum gates, which are represented by unitary matrices $U$ satisfying $U U = U U = I$. Single-qubit gates manipulate the state of individual qubits. Prominent examples include the Pauli matrices ($X, Y, Z$), which induce generalized rotations around the Bloch sphere, and the Hadamard gate ($H$), which creates an equal superposition from a computational basis state: 1456$$H|0 = 1 2(|0 + |1 ).$$ 1457To generate entanglement and enable universal quantum computation, multi-qubit gates are required. A ubiquitous two-qubit entangling operation is the Controlled-NOT (CNOT) gate, which flips the state of a target qubit if and only if the control qubit is in the $|1 $ state. Table~ tab:quantum_gates summarizes the symbols and mathematical matrix representations of these typical quantum gates. By composing finite sets of single-qubit and entangling two-qubit gates, any arbitrary unitary operation can be approximated to a desired precision. 1458 1459 table[htbp] 1460 1461 1462 Summary of typical quantum gates and their matrix representations. 1463 tab:quantum_gates 1464 1.2 1465 tabular@lcc@ 1466 1467Gate Name & Symbol & Matrix \\ 1468 1469Hadamard & $H$ & $ 1 2 pmatrix 1 & 1 \\ 1 & -1 pmatrix$ \\ 1470Pauli-X (NOT) & $X$ & $ pmatrix 0 & 1 \\ 1 & 0 pmatrix$ \\ 1471Pauli-Y & $Y$ & $ pmatrix 0 & -i \\ i & 0 pmatrix$ \\ 1472Pauli-Z & $Z$ & $ pmatrix 1 & 0 \\ 0 & -1 pmatrix$ \\ 1473Controlled-NOT & CNOT & $ pmatrix 1 & 0 & 0 & 0 \\ 0 & 1 & 0 & 0 \\ 0 & 0 & 0 & 1 \\ 0 & 0 & 1 & 0 pmatrix$ \\ 1474 1475 tabular 1476 table 1477 1478 Measurement. 1479To extract classical information from a quantum system, a measurement operation must be applied. According to the postulates of quantum mechanics, measuring a qubit in the computational basis $\|0 , |1 \$ forces its superposition state to collapse into one of the definite basis states. This outcome is inherently probabilistic; governed by the Born rule, the probability of observing a specific state is given by the squared magnitude of its corresponding amplitude. For instance, measuring $|Ï = α|0 + ÎČ|1 $ yields the outcome $0$ with probability $|α|^2$ and $1$ with probability $|ÎČ|^2$. Following measurement, the quantum state irreversibly collapses into the observed state, destroying any prior superposition or entanglement. 1480 1481 Quantum circuits. A quantum algorithm is practically formulated as a quantum circuitâan ordered, acyclic sequence of operations applied to an initial state, typically initialized to $|0 n$. Read from left to right, a circuit comprises state preparation, a series of unitary quantum gates to execute the logical computation, and terminal measurements to map the final quantum state into classical registers. Due to the probabilistic nature of quantum measurement, a circuit is typically executed repeatedly to sample the output distribution, thereby allowing the estimation of expectation values that encode the solution to the given computational problem. 1482 1483 Details of API Association Modeling 1484 sm:api_modeling 1485This section elaborates on the mathematical formulations for the API association model referenced in ~ sec:know_extra of the main text. To rigorously quantify the semantic and structural relationships between APIs, we formally define the three dimensions used to compute the combined association score. 1486 1487 Proximity. We use file- and module-level locality as a proxy for semantic relatedness. Let $S_p(a,b)$ denote the proximity score between APIs $a$ and $b$, defined as: 1488$$S_p(a, b) = cases α, & if a and b are defined in the same file, \\ ÎČ, & if a and b are in the same module but different files, \\ λ, & otherwise. cases$$ 1489 1490 Type overlap. Quantum software typically introduces numerous domain-specific types. APIs that share input or output types are likely to be used together. Let $Typeset(a)$ denote the set of non-native types appearing in the signature of API $a$. We quantify type-level similarity using the Jaccard index: 1491$$S_t(a, b) = J(Typeset(a), Typeset(b)),$$ 1492where $J(A,B)=|A â© B| / |A âȘ B|$ measures the overlap between two sets. 1493 1494 Call relationships. We approximate API relatedness based on call relationships observed in the source code via reachability and commonality. Reachability is measured using a shortest call distance function $D(a,b)$, where $D(a,b)=1$ indicates a direct call and $D(a,b)=+â$ indicates no call path. Commonality captures shared invocation patterns. Let $Caller(·)$ and $Callee(·)$ denote the sets of APIs that invoke or are invoked by a given API, respectively. The resulting call association score is defined as: 1495\[ 1496 aligned 1497S_c(a,b) = & 12 [ J(Caller(a), Caller(b)) + J(Callee(a), Callee(b)) ] \\ 1498& + 12D(a,b). 1499 aligned 1500\] 1501 1502 Implementation details of 1503 sm:implementation 1504This section elaborates on the implementation details and the experimental setup of omitted in ~ sec:approach. Table~ tab:hyperparams provides a comprehensive summary of the core hyperparameters and configuration settings used throughout the evaluation of . Subsequently, we describe the exact prompt structures designed for and define the primary metrics utilized to evaluate performance. 1505 1506 table[h] 1507 1508 Summary of Core Hyperparameters and Configuration Settings. 1509 tab:hyperparams 1510 1511 tabularllcc 1512 1513Section & Parameter & Symbol & Value \\ 1514 5* ~ sec:know_extra & Intra-file Score & $α$ & 1.0 \\ 1515 & Intra-module Score & $ÎČ$ & 0.6 \\ 1516 & Default Proximity Score & $λ$ & 0 \\ 1517 & Evolution Sensitivity & $Îł$ & 4 \\ 1518 & Evolution Range & $ÎŽ$ & 0.3 \\ 1519 3* ~ sec:know_extra & Locality Weight & $w_1$ & 2 \\ 1520 & Type Overlap Weight & $w_2$ & 1 \\ 1521 & Call Relationship Weight & $w_3$ & 2 \\ 1522 ~ sec:seed_selection & Entanglement Cap & $Ï$ & 6 \\ 1523& API Diversity Weight & $η$ & 3 \\ 1524 2* ~ sec:seed_selection & Selection Top-$Îș$ & $Îș$ & 10 \\ 1525 & Programs per Iteration & - & 30 \\ 1526 tabular 1527 table 1528 1529 figure[t] 1530 1531 [width=0.9 ]prompt_1.pdf 1532 Prompt structure used by for API implementation semantic modeling. 1533 fig:prompt_1 1534 figure 1535 1536 figure[t] 1537 1538 [width=0.94 ]prompt_2.pdf 1539 A structured view of the prompt template used for LLM-guided seed program generation. The template consists of a shared context block (seed program and target API information) and two parallel, specialized variant blocks: coverage-oriented prompt and call-oriented prompt. 1540 fig:prompt 1541 figure 1542 1543 Prompt 1544 1545In this section, we detail the specific prompt structures utilized by , corresponding to the two main phases of our framework: API semantic extraction and seed program generation. 1546 1547 API implementation semantic modeling prompt. 1548As illustrated in Figure~ fig:prompt_1, to isolate the exact logic of the target library, the prompt injects the library name ( <TARGET\_LIB>) and the relevant, filtered source code ( <SRC\_CODE>). To ensure the output is easily parsed by the automated testing pipeline, we constrain the LLM to return a strictly formatted JSON object. This JSON explicitly constructs the semantic model by extracting: 1549 itemize 1550 Input Constraints: rules, types, and boundary conditions for the APIâs arguments; 1551 Output Description: expected return values and behaviors; 1552 Functionality Summary: a concise overview of the APIâs purpose; and 1553 Fuzzing Test Points: specific edge cases, parameter combinations, and structural targets to guide the fuzzing phase. 1554 itemize 1555 1556 LLM-guided seed generation prompt. 1557Figure~ fig:prompt displays the prompt template used by the fuzzing LLM. To maintain contextual awareness and ensure the generation of valid seed programs, the prompt is divided into a Shared Context block and two task-specific variants: 1558 itemize 1559 Shared Context: this foundational block establishes the persona of a professional quantum computing programmer. It grounds the generation process by providing the execution skeleton via the seed program ( <SEED>), alongside the target library ( <TARGET\_LIB>), the target API name ( <API\_NAME>), and its comprehensive semantic model \\( <API\_MODEL>) extracted during the modeling phase; 1560 Variant A (Coverage-Oriented Prompt): activated when the target API ($i_t$) is already present in the seed program. It directs the model to analyze the semantic profile and generate high-coverage calling code, specifically instructing the LLM to leverage the provided fuzzing test points to explore deeper, more diverse execution paths; and 1561 Variant B (Call-Oriented Prompt): activated when $i_t$ is absent from the seed program. It instructs the model to modify the existing Python seed to seamlessly introduce a new invocation of the target API. To guarantee syntactic correctness and a fully runnable output, this variant explicitly injects the necessary import statement ( <IMPORT\_STATEMENT>). 1562 itemize 1563 1564 Metric 1565 1566We adopt three primary metrics in software testing to analyze the collected experimental results. For clarity, their key concepts and functionalities are introduced below. 1567 1568 Code coverage. Code coverage has been widely adopted in software testing and quantum library testing. We measure Python line coverage using the coverage.py tool while excluding the libraryâs internal test files to ensure that the metrics accurately reflect the exploration of core functional logic. Furthermore, we evaluate unique code coverage, defined as the specific code segments triggered exclusively by a particular configuration. This metric allows us to quantify the distinct exploratory effectiveness of each setting and its unique contribution to the overall code space exploration. 1569 1570 Validity. A generated case is defined as valid if it executes without runtime exceptions in a properly configured environment and invokes the target API at least once. We perform deduplication to count only unique instances, which further provides metrics such as the number of valid programs and validity rate. Note that because the validity of mutated samples inherently depends on the specific mutation strategy applied, mutants are excluded from the statistics for this metric. 1571 1572 Bug detection. Following prior work on quantum library fuzzing, we report the number of unique detected bugs. 1573 1574 Baseline 1575To evaluate the effectiveness of , we compared it against three state-of-the-art baselines, including two quantum-specific fuzzers and one LLM-based universal fuzzer: 1576 1577 itemize 1578 MorphQ~ paltenghi2023morphq is a metamorphic testing framework specifically developed for the Qiskit platform, which designs quantum-specific metamorphic relations and a dedicated program generator to expose semantic inconsistencies. 1579 FuzzQ~ shaking2025oopsla encodes QASM semantics in Alloy to generate structurally constrained circuits and employs invariant checking, statistical tests, and cross-simulator unitary consistency as differential oracles. 1580 Fuzz4All~ xia2024fuzz4all is a general-purpose fuzzer that proposes an LLM-based auto-prompting mechanism combined with an iterative fuzzing loop to automatically synthesize diverse and semantically meaningful inputs across programming languages, achieving high coverage and broad applicability beyond quantum systems. 1581 itemize 1582 1583 Experimental results 1584 sm:results 1585 1586This section provides supplementary experimental data and qualitative examples that complement the primary evaluation presented in ~ sec:results of the main text. Specifically, ~ sm:coverage presents additional code coverage results utilizing the CodeLlama model family to reinforce our primary coverage findings. Furthermore, as referenced in the main text, ~ sm:case_study details additional case studies illustrating other bug categories discovered by . 1587 1588 Coverage Results 1589 sm:coverage 1590As illustrated in Figure~ fig:coverage_codellama, the evaluation utilizing the CodeLlama model family (7B and 13B) corroborates the findings presented in the main text. consistently outperforms the Fuzz4All baseline across all three quantum libraries. Most notably, operating at the smaller 7B scale consistently achieves significantly higher line coverage than the baseline framework utilizing the larger 13B model throughout the entire generation process. These supplementary results further validate that the superiority of is fundamentally driven by its meticulously designed, library-aware generation strategy rather than the sheer scaling of model parameters. 1591 1592 figure[t] 1593 1594 [width=0.8 ]case_coverage_codellama.pdf 1595 Line coverage comparison using CodeLlama. 1596 fig:coverage_codellama 1597 figure 1598 1599 Case Study 1600 sm:case_study 1601In the following, we showcase several representative bugs found by to illustrate their impact and root causes. 1602 1603Listing~ code:qiskit_boundary shows a boundary violation where synthesizing a 1-qubit Clifford operator triggers a Rust-level panic. Although the LNN (Linear Nearest Neighbor) synthesis should support any valid circuit dimension, the underlying logic failed to handle the $n=1$ edge cases, assuming multi-qubit connectivity. It is confirmed as an unhandled boundary in the synthesis engine and has since been patched to support minimal qubit configurations. 1604 1605 lstlisting[style=QuantumPyBordered, caption=Single-qubit Clifford Synthesis Panic in Qiskit, label=code:qiskit_boundary] 1606from qiskit.circuit import QuantumCircuit 1607from qiskit.quantum_info import Clifford 1608from qiskit.synthesis.clifford import synth_clifford_depth_lnn 1609 1610qc = QuantumCircuit(1) 1611qc.h(0) 1612clifford_op = Clifford(qc) 1613 1614# This call triggers the Rust panic 1615synthesized_circuit = synth_clifford_depth_lnn(clifford_op) lstlisting 1616 1617Listing~ code:cirq_hash illustrates a state divergence where checking gate compatibility triggers a TypeError. The callee function fails when processing a UniformSuperpositionGate because the gate class is unhashable. This bug stems from a flaw in the relevant \_\_contains\_\_ implementation, which incorrectly assumes that all gate-derived objects are hashable. 1618 1619 lstlisting[style=QuantumPyBordered, caption=Unhashable Gate Membership in Cirq, label=code:cirq_hash] 1620import cirq 1621from cirq.neutral_atoms import is_native_neutral_atom_gate 1622from cirq.ops import UniformSuperpositionGate 1623 1624gate = UniformSuperpositionGate(m_value=3, num_qubits=2) 1625 1626# This call triggers the TypeError 1627print(is_native_neutral_atom_gate(gate)) lstlisting 1628 1629Listing~ code:pennylane_sign_expand illustrates a semantic violation where executing a QNode with the sign\_expand transform results in an error. This defect stems from a packaging oversight where a critical metadata file ( sign\_expand\_data.json) was omitted from the distribution configuration. This missing resource prevents the transformation API from loading its required logic, resulting in a runtime failure that violates the libraryâs functional integrity and documented behavior. 1630 lstlisting[style=QuantumPyBordered, caption=Missing Resource Dependency in PennyLane, label=code:pennylane_sign_expand] 1631import pennylane as qml 1632dev = qml.device("default.qubit", wires=2) 1633obs = qml.Hamiltonian([2.0, -1.5], [qml.PauliZ(0), qml.PauliX(1)]) 1634 1635@qml.transforms.sign_expand 1636@qml.qnode(dev) 1637def circuit(): 1638 qml.RX(0.5, wires=0) 1639 qml.RY(0.5, wires=1) 1640 qml.CNOT(wires=[0, 1]) 1641 return qml.expval(obs) 1642# A FileNotFoundError error occurs during circuit execution. 1643result = circuit() lstlisting 1644 1645 Efficiency Analysis 1646 sm:efficiency 1647 1648This section analyzes the computational efficiency of . First, we examine the one-time modeling overhead and the execution-time distribution of the complete pipeline under the default configuration used in RQ1. Second, we report the per-sample token consumption and generation time across the LLM backends evaluated in RQ2. 1649 1650 Overall Computational Efficiency 1651 sm:efficiency_overall 1652 1653 Cost of the modeling LLM. 1654The modeling LLM is invoked exclusively during the one-time API corpus 1655construction phase, in which the source code of each API is processed once to 1656construct its semantic model. The resulting corpus is reused throughout all 1657subsequent fuzzing iterations without any further invocation of the modeling 1658LLM. 1659As reported in Table~ tab:quantum_libs of the main text, this phase 1660consumes 0.38M, 0.61M, and 0.41M tokens for Qiskit, PennyLane, and Cirq, 1661respectively. Corpus construction takes only 2.3, 3.3, and 1.6 minutes for the 1662three libraries, corresponding to 0.15\%, 0.23\%, and 0.11\% of the fixed 166324-hour budget. Therefore, the one-time preprocessing overhead is negligible 1664in all three cases. 1665 1666 Time breakdown across the pipeline. 1667To examine where computational time is spent, we divide the complete 1668pipeline into four stages: Preparation, Seed Generation, Mutation, and 1669Execution. Preparation corresponds to the one-time API corpus construction. 1670Seed Generation invokes the fuzzing LLM to produce seed programs. Mutation 1671performs source-level transformations on existing programs, whereas Execution 1672runs the generated programs against the target library. 1673 1674 table[t] 1675 1676 Execution-Time Distribution. 1677 tab:pipeline_time 1678 1679 3pt 1680 tabularlrrrr 1681 1682Target & 1683Preparation & 1684Seed Gen. & 1685Mutation & 1686Execution \\ 1687 1688Qiskit & 0.15\% & 57.50\% & 0.13\% & 42.22\% \\ 1689PennyLane & 0.23\% & 50.83\% & 0.15\% & 48.79\% \\ 1690Cirq & 0.11\% & 55.97\% & 0.14\% & 43.78\% \\ 1691 1692 tabular 1693 table 1694 1695As shown in Table~ tab:pipeline_time, the computational cost is almost 1696entirely concentrated in Seed Generation and Execution, which together account 1697for more than 99\% of the total execution time across all three libraries. In 1698contrast, Preparation and Mutation each consume less than 0.23\%. Preparation 1699is negligible because corpus construction is performed only once and reused 1700throughout all subsequent iterations, while Mutation is implemented as a 1701lightweight source-level transformation without LLM invocation. This design 1702allows to efficiently diversify LLM-generated seed programs under the 1703fixed experimental budget. 1704 1705 Per-Sample Generation Cost 1706 sm:efficiency_llm 1707 1708We further analyze the cost of individual LLM generation calls across the 1709model families and scales evaluated in RQ2. For both and Fuzz4All, we 1710report the average input tokens, output tokens, and generation time per sample. 1711Each result is averaged over Qiskit, PennyLane, and Cirq. 1712 1713 table[t] 1714 1715 Average Per-Sample Token Consumption and Generation Time. 1716 tab:per_sample_efficiency 1717 1718 3pt 1719 tabularllrrr 1720 1721Model & 1722Method & 1723Input Tokens & 1724Output Tokens & 1725Time (s) \\ 1726 1727 2*Qwen-3B 1728 & Fuzz4All & 332.15 & 201.43 & 0.92 \\ 1729 & & 367.10 & 300.90 & 1.32 \\ 1730 1731 2*Qwen-7B 1732 & Fuzz4All & 365.40 & 245.56 & 1.44 \\ 1733 & & 329.38 & 257.33 & 1.52 \\ 1734 1735 2*Qwen-14B 1736 & Fuzz4All & 323.67 & 193.33 & 1.89 \\ 1737 & & 385.60 & 344.00 & 3.35 \\ 1738 1739 2*CodeLlama-7B 1740 & Fuzz4All & 518.21 & 318.73 & 1.48 \\ 1741 & & 479.17 & 431.50 & 2.07 \\ 1742 1743 2*CodeLlama-13B 1744 & Fuzz4All & 526.95 & 319.67 & 2.47 \\ 1745 & & 521.84 & 479.50 & 3.85 \\ 1746 1747 tabular 1748 table 1749 1750As shown in Table~ tab:per_sample_efficiency, both methods incur similar input token consumption, while the clear and consistent difference lies in output tokens, where KQFuzz consumes more than Fuzz4All on every model. This is a direct consequence of KQFuzzâs generation paradigm: rather than generating short and standalone programs, KQFuzz iteratively extends existing seed programs by introducing new and semantically related API calls, resulting in longer and structurally richer programs. This design is precisely what enables KQFuzz to achieve significantly higher coverage and validity, and the additional output tokens represent meaningful semantic content rather than redundancy. 1751 1752 document 1753 "