Paper deep dive
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
Daniel Nichols, Konstantinos Parasyris, Caetano Melone, Tal Ben-Nun, Giorgis Georgakoudis, Harshitha Menon
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/14/2026, 2:30:53 AM
Summary
Record-Remix-Replay (R^3) is a hierarchical GPU kernel optimization framework that integrates LLM-driven evolutionary search with Bayesian optimization and record-replay compilation. By decoupling kernel evaluation from full-application execution, R^3 enables efficient, scalable exploration of the optimization space, including source-level implementation, compiler passes, and launch parameters, achieving significant performance improvements in scientific HPC applications.
Entities (5)
Relation Signals (3)
Record-Remix-Replay → utilizes → MAP-Elites
confidence 95% · we use MAP-Elites evolution to search over source-level kernel implementations
Record-Remix-Replay → utilizes → Bayesian Optimization
confidence 95% · we use BO + replay-based auto-tuning to search over lower-level decisions
Record-Remix-Replay → optimizes → QUDA
confidence 90% · R^3 is able to find a 28.4% reduction in compute time in the QUDA scientific code.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As high-performance computing and AI workloads become increasingly dependent on GPUs, maintaining high performance across rapidly evolving hardware generations has become a major challenge. Developers often spend months tuning scientific applications to fully exploit new architectures, navigating a complex optimization space that spans algorithm design, source implementation, compiler flags and pass sequences, and kernel launch parameters. Existing approaches can effectively search parts of this space in isolation, such as launch configurations or compiler settings, but optimizing across the full space still requires substantial human expertise and iterative manual effort. In this paper, we present Record-Remix-Replay (R^3), a hierarchical optimization framework that combines LLM-driven evolutionary search, Bayesian optimization, and record-replay compilation techniques to efficiently explore GPU kernel optimizations from source-level implementation choices down to compiler pass ordering and runtime configuration. By making candidate evaluation fast and scalable, our approach enables practical end-to-end search over optimization dimensions that are typically treated separately. We show that Record-Remix-Replay can optimize full scientific applications better than traditional approaches over kernel parameters and compiler flags, while also being nearly an order of magnitude faster than modern evolutionary search approaches.
Tags
Links
- Source: https://arxiv.org/abs/2604.11109v1
- Canonical: https://arxiv.org/abs/2604.11109v1
Trouble viewing inline? Open PDF directly →
Full Text
75,853 characters extracted from source content.
Expand or collapse full text
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search Daniel Nichols ∗ , Konstantinos Parasyris, Caetano Melone, Tal Ben-Nun, Giorgis Georgakoudis, Harshitha Menon Lawrence Livermore National Laboratory Livermore, USA danielnichols, parasyris1, cmelone, talbn, georgakoudis1, harshitha@llnl.gov ∗ Corresponding author Abstract—As high-performance computing and AI workloads become increasingly dependent on GPUs, maintaining high performance across rapidly evolving hardware generations has become a major challenge. Developers often spend months tuning scientific applications to fully exploit new architectures, navigating a complex optimization space that spans algorithm design, source implementation, compiler flags and pass sequences, and kernel launch parameters. Existing approaches can effec- tively search parts of this space in isolation, such as launch configurations or compiler settings, but optimizing across the full space still requires substantial human expertise and iterative manual effort. In this paper, we present Record-Remix-Replay (R 3 ), a hierarchical optimization framework that combines LLM- driven evolutionary search, Bayesian optimization, and record- replay compilation techniques to efficiently explore GPU kernel optimizations from source-level implementation choices down to compiler pass ordering and runtime configuration. By making candidate evaluation fast and scalable, our approach enables practical end-to-end search over optimization dimensions that are typically treated separately. We show that Record-Remix-Replay can optimize full scientific applications better than traditional approaches over kernel parameters and compiler flags, while also being nearly an order of magnitude faster than modern evolutionary search approaches. I. INTRODUCTION GPUs power much of today’s high-performance computing and large-scale AI workloads, and their efficiency increasingly determines the cost, energy consumption, and time-to-solution of scientific and industrial computing. Small improvements in GPU utilization can translate into substantial savings at cluster scale and enable higher-fidelity simulations or faster model training. Yet finding these small improvements is notoriously difficult: modern GPU software stacks expose a deep hierarchy of optimization opportunities ranging from algorithm choices and implementation to compiler optimizations and launch parameters. Finding ideal kernels within this space is usually an arduous and slow trial-and-error task left to computational scientists and performance engineers. Despite decades of progress in performance engineering and auto-tuning, GPU optimization remains a critical bottleneck in practice because existing methods address only parts of the problem. Domain experts can reason about algorithmic and implementation-level changes, while auto-tuners are ef- fective over structured spaces such as launch parameters, tiling choices, and selected compiler options. However, large performance gains often require improvements that span these layers simultaneously. As GPU systems become more im- portant to scientific computing and AI, this fragmentation becomes increasingly costly: it slows optimization cycles, limits portability across architectures, and leaves substantial performance untapped in expensive production workloads. A method that can systematically optimize across the full vertical stack of GPU kernel decisions would therefore be highly valuable, both for improving application performance and for reducing the manual expertise required to obtain it. Optimizing across the entire range of GPU kernel design decisions is difficult, however. Algorithmic structure, kernel implementation, compiler decisions, and launch parameters all shape the final execution behavior, and their effects are strongly interdependent. Small source-level changes can alter register pressure, memory access patterns, or synchronization structure, which in turn change which compiler transfor- mations are profitable and which launch configurations are effective. This creates a search space that is large, hierarchical, and non-separable: good choices at one layer depend on choices made at the others. Consequently, approaches that optimize only launch parameters, only compiler passes, or only source rewrites leave substantial performance opportunities unexplored. At the same time, exploring this space is expensive be- cause performance feedback is slow to obtain and difficult to stabilize. For large GPU applications, evaluating a candidate optimization often requires rerunning a substantial portion of the application, often several times to account for measure- ment noise. These long feedback cycles make broad search impractical. They are especially problematic for methods that require many evaluations, such as evolutionary optimization, and for methods that generate diverse source variants, such as LLM-based coding agents. Classical auto-tuners mitigate this cost by restricting search to structured parameter spaces, but that restriction also limits the kinds of transformations they can discover. What is missing is a way to make evaluation cheap enough to support rapid search while still covering the full space of kernel optimizations. Our paper introduces Record-Remix-Replay (R 3 ) that com- bines evolutionary optimization using large language models (LLMs) with Bayesian optimization, enabling us to search arXiv:2604.11109v1 [cs.DC] 13 Apr 2026 optimal source code, launch configurations, and compiler passes. To overcome the massive runtime requirement for evaluating applications thousands of times, we utilize record- replay techniques alongside several novel optimizations to accelerate and scale GPU kernel optimization. We find that R 3 is able to accelerate LLM-based evolutionary search by nearly 10× and tune GPU kernels to better speedups than existing approaches. Using R 3 we are able to reduce the time spent in compute in a large, production HPC application by over 28%. Our paper makes the following important contributions: • We introduce a novel approach combining LLM evolu- tionary search and record-replay to rapidly explore an optimization space including launch parameters, compiler passes, and implementations/algorithms. • We introduce novel improvements to existing record- replay techniques to enable rapid replay of general CUDA and HIP kernels. • We introduce several novel optimizations to enable scal- ing evolutionary optimization to hundreds of GPUs. • We release our record-replay contributions in the open- source framework Mneme (https://github.com/Olympus- HPC/Mneme). R 3 is planned to be open-sourced in the coming weeks. • We conduct evaluations of our approach on several multi- kernel scientific applications. R 3 is able to find a 28.4% reduction in compute time in the QUDA scientific code. I. BACKGROUND This section presents background on auto-tuning, record- replay compiler instrumentation techniques, and evolutionary optimization powered by LLMs. A. Auto-tuning Auto-tuning refers to techniques that automatically search for program configurations to optimize a given performance objective, such as execution time, energy consumption, or resource utilization. Rather than relying solely on manual per- formance engineering, auto-tuners explore a space of possible implementations or parameter settings that influence program behavior. These parameters may include algorithmic variants, compiler optimization flags, tiling factors, memory layouts, or runtime launch parameters. Most auto-tuning systems operate iteratively: in each iteration, the tuner generates a candidate configuration, evaluates it by executing the program (or a representative kernel), and uses the measured performance as feedback to guide subsequent choices. The decision pro- cess may be driven by strategies such as random search, evolutionary algorithms, or model-based optimization meth- ods. Through repeated evaluation and feedback, auto-tuning progressively identifies configurations that improve the target performance metric. Many auto-tuning frameworks have been proposed for GPU and HPC applications, including general-purpose systems such as OpenTuner [1], ActiveHarmony [2], Kernel Launcher [3], and CLTune [4]. These frameworks differ in the parameters they optimize and the search strategies they employ, but all rely on iterative evaluation of candidate configurations. Because each candidate typically requires executing the application/k- ernel, the evaluation cost often becomes a dominant limitation. B. Record-Replay Record-and-replay decouples optimization evaluation from full-application execution by recording a replay unit together with the recorded execution context needed to reproduce it, then replaying transformed variants in isolation. A replay- based evaluation loop therefore consists of: (1) selecting a replay unit, (2) capturing sufficient state to reconstruct its execution, and (3) re-executing that unit repeatedly under modified code, compiler settings, or runtime parameters. Its key benefit is lower evaluation cost: replay avoids rerun- ning unrelated application code, enables selective tuning of hotspots, and exposes parallelism across replay instances. Prior work instantiated this model at different granularities. CERE [5]–[7] uses CPU codelets as replay units, extracting hotspot regions at LLVM IR level so they can be modified, recompiled, and replayed independently from the original program. To keep replay representative while reducing cost, CERE captures the working set and cache state and selects representative invocations. In contrast, Parasyris et al. [8] use the GPU kernel invocation itself as the replay unit. Their recorded execution context includes kernel launch arguments, launch configuration, device memory state, relevant globals, and the kernel image; replay reconstructs this state and executes the kernel as a standalone artifact. Compared to codelet replay, kernel replay is a more natural granularity for GPU tuning because it aligns directly with the unit whose compiler transformations and launch parameters determine performance. Because replay instances are self-contained, they can be tuned selectively, migrated to compatible systems, and evaluated in parallel, substantially reducing optimization cost. Our work builds on this replay-based evaluation model, using GPU-kernel replay as the inner evaluation engine for hierarchical search across source implementations, compiler transformations, and launch configurations. C. Evolutionary Optimization with LLMs Evolutionary optimization has been a popular optimization technique for many decades. Consider the task of finding some x ⋆ that minimizes f(x). Evolutionary optimization creates populations of candidate x values, evaluates each candidate using f or a proxy fitness function, and then evolves the pop- ulation to find the next candidate x values. How populations are evolved can take many forms, but is most often based on genetic evolutionary algorithms, where selection, mutation, and/or crossover are used to remove bad candidates, explore changing good candidates, and migrate candidates between populations, respectively. This style of optimization has been very successful at exploring diverse solutions to optimization problems, particularly non-differentiable ones. We refer the reader to [9] for a more complete discussion and background on evolutionary algorithms. Fig. 1. Overview of MAP-Elites evolution as in AlphaEvolve/OpenEvolve. 1 Prompts are constructed from the MAP-Elites population database,2 sent to an LLM randomly selected from an ensemble of LLMs, and 3 evaluated by compiling and running the code. The population database is then updated according to the MAP-Elites algorithm. Multiple parallel controllers execute this loop concurrently to increase throughput. Up until recently, evolutionary algorithms were only useful for search spaces where mutation was clearly defined, that is, where it is clear how to mutate the current population into the next generation. For example, discrete and continuous spaces are generally trivial to mutate; we can increment numbers, negate them, flip bits, choose random subsets of discrete sets, etc. But other search spaces that are expressed as structured text, like algorithms and source code, are much harder to mutate. We could alter, remove, and/or add new characters, but most mutations would lead to invalid code and the evolutionary search would never converge. However, the recent increase in language modeling capabilities via LLMs has made it possible to use evolutionary search over a text space. Recent works like FunSearch [10], AlphaEvolve [11], and ADRS [12] have shown the impressive capabilities of these systems in discovering new algorithms and systems optimizations. Many of these works, like AlphaEvolve [11] and its popular open-source reproduction OpenEvolve [13], are based on the MAP-Elites evolutionary algorithm [14]. Figure 1 presents an overview of the MAP-Elites evolutionary algorithm pow- ered by LLMs for code optimization. In the MAP-Elites algorithm, members of the population are mapped to coor- dinate cells in a cartesian grid based on “feature dimensions”, i.e. each dimension of the grid corresponds to a feature of solutions in the search space. Common feature dimensions are complexity (e.g. lines of code) and diversity (e.g. hamming distance from other samples). When a new solution is sampled, it is evaluated and mapped to a coordinate cell based on its features. If there is already a solution in that cell and it is worse than the new one, we replace it. The solutions in each coordinate cell are called “elites”. We keep mutating existing samples, evaluating the mutations, and adding them back into the grid until we reach a desirable convergence criterion. MAP-Elites was first introduced to increase the diversity of samples explored during evolutionary optimization. Each island maintains its own MAP-Elites database in which multiple parallel controllers concurrently execute the sample, evaluate, and update cycle. When multiple islands are used we employ migration to move data between them; migration occurs regularly, e.g. every N iterations, and moves samples between MAP-Elites databases. One popular migra- tion implementation is to, whenever migration occurs, replace the bottom half of samples in every island, with one of the top samples from the other islands. For GPU kernel optimization, the main limitation of this approach is evaluation cost. Each iteration still requires com- piling and running the candidate to measure correctness and performance. For large GPU applications, where a single eval- uation may require substantial hardware and long execution times, this makes LLM-guided MAP-Elites prohibitively ex- pensive. This evaluation bottleneck is precisely what motivates our use of record-replay as the inner evaluation engine in R 3 . I. RELATED WORK In this section we present related work on auto-tuning, record-replay, and evolutionary based GPU kernel optimiza- tion methods. We focus solely on works used for GPU kernel optimization as that is the target of our work. Prior work in GPU kernel optimization typically targets only one layer of the tuning stack (Table I). Kernel parameter auto-tuners (e.g. CLTune) focus on structured template choices and launch configurations, and therefore achieve low-cost kernel-scope evaluation [3], [4], [15]. DSL auto-schedulers (e.g. TVM) broaden the structured search space, but require adoption of a DSL or specialized IR and still primarily optimize at kernel or graph scope [16]–[20]. In contrast, compiler auto-tuners (e.g. CompilerGym) search compiler flags or phase orderings, while runtime and system tuners (e.g. ActiveHarmony) tune execution or platform parameters; these methods generally operate at application scope and incur higher evaluation cost [1], [2], [21]–[27]. Replay and JIT substrates (e.g. CERE) reduce evaluation cost substantially by isolating codelets or kernels as replay units, but they do not search unstructured source-level trans- formations [6], [8], [28]. Conversely, LLM-based evolutionary systems (e.g. AlphaEvolve) or kernel optimization agents (e.g. opt-r1) can explore unstructured rewrites, but typically rely on full-application compilation and execution and therefore remain expensive for large GPU applications [10], [11], [13], [29], [30]. R 3 combines the strengths of these lines of work: it uses an outer LLM/MAP-Elites loop to search unstructured kernel rewrites, while an inner record-replay loop provides low-cost evaluation and Bayesian optimization over compiler and launch parameters. Quantitatively, our closest baselines are OpenEvolve and Kernel R-style BO. IV. R 3 : RECORD-REMIX-REPLAY OVERVIEW In this section we detail our proposed Record-Remix-Replay (R 3 ) framework. The framework combines LLM-powered TABLE I REPRESENTATIVE RELATED WORK IN GPU KERNEL AUTO-TUNING AND OPTIMIZATION.✓ = FIRST-CLASS SUPPORT; ◦ = TECHNICALLY POSSIBLE BUT INDIRECT / NOT THE PRIMARY DESIGN POINT. EVALUATION COST IS QUALITATIVE AND REFLECTS THE TYPICAL INNER-LOOP TUNING COST: Low = ISOLATED KERNEL / REPLAY / JIT EVALUATION, Med. = OPERATOR OR GRAPH-LEVEL TUNING, High = FULL-APPLICATION. Source-levelCompiler-level RuntimeSystemEval. Work Unstruct. rewrite DSL / template Flags / params Pass order Launch cfg. Env. / resources Required Execution Cost KERNEL PARAMETER AUTO-TUNERS CLTune; Kernel Tuner [4], [15]✓ ◦ KernelLow Kernel Launcher [3]✓KernelLow DSL AUTO-SCHEDULERS TVM / Ansor / MetaSchedule [16]–[18]✓ ◦ ✓Kernel / GraphMed. Halide autoschedulers [19]✓KernelLow Triton [20]✓ ◦ ✓KernelLow COMPILER AUTO-TUNERS OpenTuner (e.g., GCC flags) [1]✓ ◦ App.High MiCOMP; CompilerGym [21], [22]✓App.High RUNTIME/SYSTEM TUNERS ActiveHarmony; Apollo; Artemis [2], [23], [24] ◦ ✓App.High HiPerBOt; GPTune; Periscope [25]–[27]✓App.High REPLAY / JIT SUBSTRATES CERE (codelets) [6] ◦ ✓ ◦ ✓Replay UnitLow Kernel R [8] ◦ ✓Replay UnitLow Proteus (LLVM-IR JIT) [28] ◦ ✓KernelLow LLM EVOLUTIONARY AGENTS FunSearch; AlphaEvolve [10], [11]✓ ◦ App.High Record-Remix-Replay (this work)✓Replay UnitLow MAP-Elites evolution to optimize the source code imple- mentation of GPU kernels with record-replay and Bayesian optimization (BO) to tune the launch configuration and com- piler passes for kernel implementations. Figure 2 provides an overview of the R 3 framework. R 3 decomposes kernel tuning into two levels to most naturally fit the hierarchy of kernel tuning decisions: text-based source level transformations and structured numeric search over compiler passes and launch configurations. At the outer level, we use MAP-Elites evolution to search over source- level kernel implementations. At the inner level, for each source candidate, we use BO + replay-based auto-tuning to search over lower-level decisions such as compiler optimiza- tion pipelines and launch parameters. This structure allows the framework to explore algorithmic and implementation changes without giving up the efficiency of structured auto-tuning where it is most effective. Bayesian optimization has previ- ously demonstrated strong success in searching over structured kernel hyperparameters [8], but is not well suited towards searching new source code implementations. Similarly, MAP- Elites evolution could be used to guide LLMs to generate launch configurations, but we experimentally find that standard BO approaches are much more efficient and effective for this task than LLMs. The input to R 3 is an initial implementation of the target kernel together with a set of recorded executions obtained from representative application runs (see Section I-B for record-replay context). This implementation is what we will be optimizing and becomes the initial MAP-Elites population (step 0 ). At each iteration of the algorithm, the current MAP- Elites population is used to construct prompts that will be used to prompt an LLM to generate optimizations (step 1 ). These prompts will contain a candidate kernel implementation that needs to be optimized, historical good and bad examples from the database, and necessary context about the current task, so that the LLM knows to output an optimized kernel. Rather than forming prompts arbitrarily, R 3 uses a prefix-aware prompt sampler that increases reuse of static prompt prefixes across generations. This reduces generation overhead and improves throughput when many candidate mutations are produced over the course of evolution. Given a formatted prompt, the framework selects an LLM from an ensemble of available models and uses the selected model to propose a new kernel implementation (step 2 ). The model choice is runtime-aware: rather than assigning requests uniformly, R 3 schedules generations based on estimated la- tency and current device availability so that inference resources remain well utilized. Once a model is selected it is used to generate an optimized version of the kernel in its prompt. The generated candidate kernel is then evaluated to deter- Fig. 2. Overview of our Record-Remix-Replay framework.1 First, the prefix-aware prompt sampler is used to construct prompts that optimize prefix cache hits.2 An LLM is selected intelligently based on estimated generation time to maximize GPU utilization and is used to optimize the GPU kernel.3 LLVM IR is generated for the optimized kernel and sent to the replay server, which replays kernels and runs parallel Bayesian optimization to find the best runtime across compiler passes and launch parameters. The population database is then updated according to the MAP-Elites algorithm. mine its correctness and efficiency (step3 ). Evaluation begins by compiling the generated candidate to LLVM IR, which serves as the interface between source-level evolution and replay-based evaluation. We note that this itself is a significant optimization over standard evolution where re-compilation can be a huge bottleneck, particularly if any changed code is in a header that many translation units depend on or if there are many files that need to be linked. By getting the LLVM IR for only the kernel of interest, we can drastically reduce compile and linking time during evaluation. The compiled LLVM IR for the candidate is then sent to the replay server. The replay server reconstructs a previously recorded kernel execution, lowers the candidate IR to an executable kernel, and evaluates it in isolation from the full application (Section V details the record-replay engine). That is, we “replay” the GPU kernel as a standalone binary using the same GPU state as in the full application execution. In R 3 we replay kernels K times (e.g. 10×) to obtain statistically significant performance results. When replaying/evaluating a kernel implementation candidate, the server runs a parallel Bayesian optimization (BO) search over compiler pass config- urations and launch parameters to identify the best-performing realization of that implementation. We discard any compilation pipelines or launch configurations that lead to the kernel being incorrect. By utilizing BO, source mutations are not judged by a single arbitrary compiler setting or launch configuration, but by the best configuration that can be found for that candidate. The replay server returns the best found configuration, its runtime, and the correctness of the kernel (Section V-C has more details on how R 3 checks correctness). Finally, the measured performance and correctness results are returned to the MAP-Elites controller and used to update the population database. The resulting elites therefore repre- sent source implementations paired with their correctness and near best possible performance after tuning. This optimization is run in a loop, repeating 1 ,2 , and3 until some finishing criteria is reached. R 3 supports three different stopping criteria: maximum number of iterations, maximum wall clock time, or after a set number of iterations with no improvement. The above optimization loop runs for a single GPU kernel and many instances of it are run in parallel for a code with multiple GPU kernels. In this sense, R 3 is embarrassingly parallel as we can optimize each GPU kernel in parallel with no synchronization or communication. V. RECORD-REPLAY ENGINE R 3 utilizes record-replay to rapidly increase the rate at which candidate GPU kernels can be evaluated. While there are previous record-replay systems, they lack important ca- pabilities needed for use in tight evolutionary optimization. Relative to prior replay systems, our engine contributes four capabilities that are critical for our setting: 1) a persistent replay server that amortizes replay initialization across many evaluations, 2) direct, in-memory operation on recorded LLVM IR through Proteus and LLVM APIs, enabling low-cost com- piler and runtime transformations, and 3) correctness checking that is conservative by default but configurable when numerical equivalence rather than bitwise equality is desired. 4) the ability to record and replay arbitrary CUDA and HIP kernels. Together, these make replay practical for powering evaluation for large-scale evolutionary optimization. The engine follows the standard replay workflow of instru- mentation, recording, and replay, but specializes each stage for repeated kernel evaluation. Instrumentation prepares the application so GPU kernels can be intercepted and preserved as replay units. Recording selects as Replay Units the ex- ecution context of representative kernel invocations. Replay reconstructs that context, applies transformations, and evalu- ates candidate implementations under alternative compiler and launch settings. A. Instrumentation and Recording Instrumentation is implemented by extending the Proteus JIT infrastructure [28] to identify GPU kernels, extract each kernel together with its transitive dependency closure as LLVM IR, and embed the resulting IR in the application binary. Kernel launches are redirected through the Proteus runtime so they can later be intercepted and replayed. We preserve kernels as LLVM IR because it is the representation that both minimizes re-evaluation cost (by reducing compile time) and enables transformation during replay (via LLVM passes). Recording is performed during representative application runs through a preload library that intercepts kernel launches and device memory operations. For each recorded kernel invocation, the engine stores: 1) the kernel LLVM IR module, 2) launch configuration, 3) kernel arguments, 4) the pre- execution device state, and 5) the post-execution device state. Using the terminology of section I-B, the pre-execution state forms the prologue snapshot used to reconstruct the replay context, while the post-execution state forms the epilogue snapshot used for correctness checking. Because the replay unit is captured independently from the full application, later experiments can be run without rebuilding or re-executing the original code. B. Replay, Transformation, and Auto-Tuning During replay, the engine restores the recorded execution context, recompiles the recorded LLVM IR, and executes the replay unit either with the recorded launch configuration or with alternative settings sampled by the optimizer. Replay operates directly in memory on LLVM IR through Proteus rather than through a file-based recompilation workflow like in other works [8]. This enables low-cost exploration of transformations such as argument specialization by constant propagation, specialization of thread and block dimensions, in- sertion of launch bounds, standard LLVM optimization levels, and custom compiler pipelines. The transformed kernel is then lowered to a device binary and executed on the reconstructed state. Built on top of replay, the engine exposes a programmatic auto-tuning interface over compiler and runtime parameters. A sampled point in the search space corresponds to a replay experiment: compile the candidate with a selected transforma- tion pipeline, execute it under a selected launch configuration, and return its measured runtime if correct. In R 3 this interface is used with Bayesian optimization inside the replay server, but the mechanism itself is agnostic to the search strategy. C. Correctness Checking Correctness is defined relative to the recorded execution context. Given a replay instance, the engine restores the pro- logue snapshot, executes the candidate kernel, and compares the resulting state against the recorded epilogue snapshot. By default, this comparison is bitwise exact: the GPU memory state after replaying the kernel must match the memory state from the full program run exactly. This conservative policy ensures that accepted transformations preserve the observed kernel behavior on the recorded state. Some optimizations, however, are numerically valid while not bitwise identical, especially for floating-point code. A common example we encounter in this work is replacing an expression such as a / sqrtf(b) with a * rsqrtf(b). To support such cases, the engine also provides an optional relaxed mode in which selected output variables may be checked with user-provided numerical predicates, such as absolute or relative tolerances. Such memory addresses are marked in source code with an annotate function. We emphasize that bitwise correctness checking is the default in our engine; checking correctness with relaxed error tolerances requires users to explicitly annotate what variables can have relaxed constraints and what the strictness of those constraints are. Candidates that fail to compile, crash, or violate the correctness criterion are classified as incorrect and rejected. To balance speed and robustness, correctness is checked at two granularities. During search, candidates are typically validated against a single recorded prologue/epilogue pair to keep the inner loop fast. Promising final candidates are then validated against additional recorded instances of the same kernel when available. This preserves low evaluation cost during search while ensuring the candidate kernel is correct for all available recorded code paths. D. Persistent Replay Server A major systems contribution and optimization within our replay engine is a persistent replay server designed for re- peated evaluation. Prior replay systems are effective for of- fline tuning, but repeated kernel evaluation still suffers from overheads such as re-initializing device memory, reloading replay state, materializing intermediate files, and transferring correctness checks back to the host. These costs are a major bottleneck when in the inner loop of evolutionary search. To address this, R 3 uses a replay server that keeps replay state resident across many evaluations and serves replay re- quests over a lightweight interface. On startup, the server loads a database of recorded kernels, maps replay units across available GPU resources, and pre-allocates prologue and epilogue buffers to amortize allocation and data movement costs. The server then exposes replay and tuning endpoints that accept candidate LLVM IR, reconstruct the appropriate replay context, execute the requested experiment, and return runtime and correctness results. Because replay workers persist across requests, initialization overheads are paid once rather than once per candidate. This design is necessary for replay to be suitable for R 3 . It increases throughput enough that candidate generation, rather than replay, becomes the dominant bottleneck in many settings, which in turn motivates the candidate generation optimizations introduced in Section VI. E. Deployment After tuning, the engine exports the best-performing config- uration for each kernel as a JSON specification describing the selected transformations, compiler settings, and launch param- eters. This specification is consumable by the Proteus runtime, which applies the prescribed optimizations when the original application is executed. As a result, tuning can be performed offline on representative replay instances, while deployment remains transparent to the application source code. VI. SCALING EVOLUTIONARY OPTIMIZATION The use of record-replay and our many optimizations in R 3 rapidly increases the rate of evaluation (see Section V) and makes evolution inference-bound. Thus, in this section we present optimizations in R 3 that accelerate inference via prompting strategies and runtime-aware inference scheduling. A. Prefix-cache Aware Prompting Most modern LLM inference engines employ prefix caching, which reuses intermediate computations when suc- cessive generations share a common prompt prefix. Prefix caching is most beneficial when the prefill phase is non- negligible compared to decode time, i.e. for large input prompts and small/moderate outputs. This matches our setting well: evolutionary prompts contain several large code samples from the database, while the model only needs to generate a diff for a single GPU kernel. However, we find that prompting strategies used in open-source AlphaEvolve-style frameworks are not prefix-cache friendly and place dynamic information early in the prompt, leading to frequent cache misses. To address this, we propose a prefix-cache aware prompt sampling algorithm for MAP-Elites-style evolution. First, we place all static prompt content at the beginning and move dynamic content, such as sampled code examples, to the end. This ensures that the static portion of the prompt can be reused across generations. Second, we reorder sampled inspiration codes to maximize reuse of previously cached prompt prefixes. Given a sampled set of inspiration codes E = [e 1 ,...,e k ], let H denote the set of recent prompt orderings. For each H ∈ H, we find the longest prefix H[1 : ℓ] that is an ordered subset of E and select the ordering with the largest such ℓ. If a match is found, we place that shared prefix first and append the remaining sampled codes in their original sampled order; otherwise, we keep E unchanged. The final prompt is then formed as a fixed static prefix followed by this reordered list of inspirations. Because this procedure changes only the order in which sampled codes are presented, and not which codes are sampled, it preserves the underlying MAP-Elites convergence behavior while increasing the likelihood of prefix-cache hits. B. Runtime Aware Scheduling After record-replay reduces candidate evaluation time, the main bottleneck in R 3 shifts from execution to generation. A naive way to increase throughput is to issue more parallel LLM requests, but this also increases staleness: prompts are con- structed from a population snapshot while many generations and evaluations are still in flight. Some staleness is tolerable, but too much reduces sample quality and wastes inference and evaluation resources on candidates proposed from already superseded elites. A straightforward way to control staleness is to increase the number of islands together with the number of parallel controllers. This mitigates contention on any single database, but it is not ideal. With too many islands, each island receives fewer updates and will require more iterations to converge. Previous works [10], [11] find ≈ 5 islands to be ideal and frameworks like OpenEvolve [13] set recommended defaults between 3 and 5 islands. In practice, we would like to keep the number of islands relatively small for good convergence behavior while still scaling candidate throughput across many inference and evaluation devices. To do this, R 3 uses a runtime-aware scheduling policy for heterogeneous LLMs. We alter the LLM sampling method (step 2 in Figure 1) to sample LLMs based on expected generation latency, the current set of in-flight requests, and ex- pected evaluation availability. Intuitively, the scheduler prefers to issue a slower request only when doing so will not leave evaluation resources idle, and otherwise fills short gaps with faster models that can complete generation and replay before the slower request returns. For example, if a slow model (e.g. gpt-5) is generating a new optimization and the GPU it is scheduled to evaluate on is currently sitting idle, then we can try to backfill this bubble by running a quick model (e.g. gpt- 5-mini) and evaluate its output before the slow model is done generating. This is similar to backfilling in HPC schedulers: short jobs are used opportunistically to improve utilization without delaying longer jobs. Concretely, let τ m denote the estimated generation time for model m, and let ˆ t eval denote the current estimate of replay plus autotuning time for a candidate. When a controller is ready to issue a new generation request, the LLM sampler examines the outstanding generations and the expected times at which replay workers will become available. If a fast model can complete generation and have its candidate evaluated within slack time that would otherwise be lost while waiting for a slower model, the fast model is selected. Otherwise, the scheduler may issue a slower, higher-quality model request. This policy increases overlap between inference and replay, keeps GPUs busier, and improves end-to-end throughput. An important design point of this policy is that it increases utilization given a fixed number of parallel controllers and GPUs available. We do not need to increase the number of parallel controllers and, thus, can improve both inference and replay devices utilization while limiting the number of in-flight samples drawn from stale database state. As we show in Section VIII, once replay has made evaluation cheap, this scheduling layer becomes necessary to continue scaling evolution. We also highlight the simplicity of this approach: it can be implemented directly on top of OpenEvolve or AlphaEvolve without major evaluation infrastructure changes. Only the LLM selection mechanism needs to be replaced, which can be trivially done in Python without altering the existing frameworks via monkey patching. VII. EXPERIMENTAL SETUP In this section we detail the applications used for evaluation, baselines for comparison, and the settings used to run R 3 . A. Computing Environment All AMD experiments are executed on a cluster with MI300A GPUs using the ROCm 7.0.2 software stack. NVIDIA experiments are executed on a cluster with H100 GPUs and CUDA 12. When gpt-oss is locally hosted for an experiment, we utilize vLLM [31] to serve the model. B. Applications We run experiments across four different scientific applica- tions to evaluate different approaches on how well they can optimize GPU kernels. Applications are selected to cover a variety of common HPC and scientific computing domains. We compare optimization results on Lulesh [32], miniFE [33], S3D [34], and miniWeather [35] as they all have CUDA and HIP implementations available as well as many GPU kernels in them. Implementations are taken from HeCBench [36], which has full CUDA and HIP implementations of each application. LULESH is a shock hydrodynamics proxy application for ex- plicit Lagrangian methods on unstructured hexahedral meshes. It has 15 GPU kernels. MiniFE is an implicit finite-element proxy application that models matrix generation, assembly, and iterative solution for sparse linear systems. It has 7 GPU kernels. S3D is a turbulent combustion DNS code with detailed chem- istry and molecular transport. It has 54 GPU kernels. miniWeather is a weather mini-app for dry compressible non- hydrostatic flow. It has 9 GPU kernels, but only 7 on the data path we evaluate. C. Baselines and R 3 Configuration We compare R 3 against two baselines: kernel parameter tuning with Bayesian optimization (BO) based on [8] and OpenEvolve [13]. We select these as they represent state-of- the-art in kernel parameter tuning (passes, launch parameters) and implementation. Record-Replay + BO. Many works have previously used variations of BO to optimize GPU kernel launch parameters. We compare with [8], a recent state-of-the-art work that combines BO with record-replay to tune launch parameters and compiler configurations rapidly. Since [8] can only handle OpenMP-offload kernels and not arbitrary GPU code, we re- implement the same BO search in our record-replay engine. We search over the same configuration (launch and compiler parameters) for 200 iterations and keep the best runtime at the end. Runtimes are obtained by running each kernel seven times: the first two as a warmup and the latter five for recording. We note that, while we recreate the search and replay functionality of [8], our implementation will actually be faster due to its many optimizations (see Section V). OpenEvolve. To compare with recent LLM-powered evolu- tionary optimization works we run OpenEvolve [13] (an open- source reproduction of the closed-source AlphaEvolve [11]). OpenEvolve implements a MAP-Elites based evolutionary search process as shown in Figure 1. We run OpenEvolve for 200 iterations with four islands, a population size of 20, and migrations every 20 iterations; these settings are based on findings in [10], [11]. Record-Remix-Replay (R 3 ). We run R 3 with the same MAP- Elites settings as OpenEvolve: 200 iterations, four islands, population of 20, and migrations every 20 iterations. Within the replay server we tune kernels across compiler pass order and launch parameters using parallel Tree Parzen Estima- tion [37] for 30 iterations as the BO search implementation. We utilize the annotate feature in R 3 (see Section V-C) to implement correctness constraints for each application (we have further manually verified the correctness of all final op- timized kernels by checking they yield the expected scientific result in the full application). For both OpenEvolve and R 3 we utilize an LLM ensemble of gpt-oss-120b [38], gpt-5-mini, and gpt-5 [39]. OpenEvolve randomly samples these with probabilities 0.6, 0.3, and 0.1, respectively. We choose these values based on the 90-10 split in [11] where a large, frontier LLM is used 10% of the time (Gemini-Pro) and a small, fast LLM is used 90% of the time (Gemini-flash); this mixture blends fast iteration with high quality optimizations while keeping costs low. We further split the cheap, fast LLM category into a commercial model and local model to increase output diversity and minimize API costs. R 3 begins with the same sampling probabilities for LLM selection, but then implements runtime-aware selection on top. To make fair time-to-solution comparisons we run all methods (record-replay + BO, OpenEvolve, and R 3 ) with 4·K GPUs, where K is the number of kernels in the application. That is, four GPUs are assigned to optimizing each GPU kernel and each optimization loop runs in parallel. Thus, the smallest run is 28 GPUs for MiniFE and the largest is 216 GPUs for S3D. We note that each of these frameworks can be run with less or more GPUs (e.g. R 3 can be run on a single GPU); we choose these settings to facilitate fair comparison between the frameworks as each approach can trivially parallelize search across different GPU kernels. VIII. SCALING EVOLUTIONARY OPTIMIZATION RESULTS In this section we present our scaling results across our optimizations. We first show how record-replay turns sin- gle thread, sequential evolution from evaluation-bound to inference-bound. We then show the inference via prefix-cache aware prompting and runtime-aware LLM selection results. Figure 3 shows the breakdown of times for a single evolution iteration when optimizing the LQCD application QUDA [40] (see Section X). We run the evolution five times and average iteration timings from the 10th to 20th iteration across those five evolution runs; this is to control for the fact that evaluation will vary throughout the course of evolution as the kernels are optimized. Limiting to the 10th to 20th iterations gives us a glimpse across consistent kernel timings, but the trends hold throughout evolution. We see that replacing standard OpenEvolve evaluation with record-replay drastically reduces compile and evaluation time. QUDA compile time is reduced by ≈ 86% and eval time by ≈ 88%. Note that the OpenEvolve baseline uses partial compilation; we re-use existing compiled binaries and only recompile source files with changes. If you wanted to build from scratch, each OpenEvolve evaluation would be signifi- cantly slower, as compiling from scratch took ≈ 30 minutes on our test system. In R 3 we only need to generate IR for the changed code and can avoid compiler lowering costs and linking costs. Most notably we have transitioned the evolution iteration from evaluation-bound in OpenEvolve to inference- bound in R 3 . OpenEvolve R 3 without server R 3 with server R 3 + server + BO tuning Optimizations 0 50 100 150 200 250 300 350 Time (seconds) 1x launch config run 1x launch config run 1x launch config run 30x launch configs run using parallel BO tuning Evolution Iteration Timing Breakdown LLM Compile Eval Fig. 3. The breakdown in times spent in generation, compiling, and evaluation during evolution with each progressive optimization used. Results are shown for 1 sequential evolve process using gpt-oss-120b with high reasoning effort on the QUDA application. We see our proposed optimizations have moved evolutionary optimization from evaluation-bound to LLM inference-bound. gpt-5-mini gpt-oss-120b Model 0 50 100 150 200 Generation Time (seconds) Distribution of Generation Times During Evolution With Prefix-Aware Prompts OpenEvolve prompts R 3 Prefix-Aware prompts Fig. 4. Comparison of time spent generating outputs with and without prefix- aware prompts across a 250 iteration evolution run. We see a reduction in median generation time for both a commercial (gpt-5-mini) and local (gpt- oss-120b) model. Notably, most commercial model providers offer discounts for prefix-cache hits, leading to lower overall costs during evolution. Now that we have successfully reduced the evaluation time and are inference-bound, we focus on optimizing time spent in LLM generation. Figure 4 shows the distribution of time spent generating outputs across LLMs between OpenEvolve and R 3 in a long 250 iteration run (we exclude gpt-5 from this experiment due to cost). We see a reduction in generation time for both models when using R 3 prefix-aware prompts. Notably, many LLM providers provide discounts for tokens that hit the prefill cache leading to reduced inference costs when using R 3 . It is possible for inference times to go down due to the reformatting of the prompts in R 3 leading to LLMs generating shorter reasoning traces or final outputs. To confirm that cache hits are actually what is aiding here, we re-run the experiment, but manually turn off prefix caching in gpt- oss-120b’s vLLM server. The resulting distribution for gpt- oss-120b is the same as OpenEvolve’s, confirming that R 3 is actually benefiting from prefix cache optimizations. Figure 5 further shows how R 3 ’s runtime-aware LLM selection algorithm improves throughput over OpenEvolve’s (4, 8)(4, 16)(8, 16)(8, 32)(16, 32)(16, 64) 0% 20% 40% 60% 80% 100% GPU Utilization (%) 80% 42% 81% 42% 83% 41% 88% 50% 91% 49% 94% 50% GPU Utilization R 3 w/ OpenEvolve LLM SelectionR 3 Runtime-aware LLM Selection (4, 8)(4, 16)(8, 16)(8, 32)(16, 32)(16, 64) 0 250 500 750 1000 1250 Throughput (evals / hour) 219.3 241.3 454.6 476.4 945.8 951.1 236.8 269.3 482.1 536.1 1002 1080 Throughput (No. Parallel Evolution Processes, No. GPUs) Fig. 5.Comparison of OpenEvolve’s LLM selection algorithm to R 3 ’s runtime-aware LLM selection. Whereas increasing the number of parallel evolution processes w.r.t. the number of GPUs can increase utilization and throughput, it comes at the cost of increasing stale MAP-Elites database reads and slowing convergence. R 3 ’s runtime-aware LLM selection increases utilization and throughput over OpenEvolve for the same number of processes. LLM selection algorithm by up to 13.5%. From Figure 5 it is clear to see that increasing the ratio of parallel evolution processes P to number of evaluation GPUs G will increase GPU utilization and throughput. For example, consider the points (4,16) and (8,16) in Figure 5. With the same number of GPUs for evaluation (16) the throughput and utilization are nearly doubled by increasing P from 4 to 8. However, as P increases past the number of islands and/or closer to G we get more stale reads from the MAP-Elites database and potentially worse/slower convergence. If the number of islands is fixed at 4, as in our experiments, changing P from 4 to 8 will nearly double the number of stale database reads and hurt convergence. Thus, it is important to improve utilization and throughput without increasing P . We see that R 3 ’s runtime- aware selection algorithm increases utilization and throughput over OpenEvolve with the same P . In some instances R 3 achieves up to an 11% increase in GPU utilization or 13.5% more evaluations per hour. IX. RESULTS In this section we present results comparing R 3 with record- replay + BO and OpenEvolve on the four applications. Figure 6 shows the distribution of kernel speedups on MI300A from each method over the baseline implementation; LuleshMiniFEs3dminiWeather Application 1 2 3 4 Speedup Kernel Speedup Distribution on MI300A Record-Replay + BOOpenEvolve R 3 (ours) Fig. 6. Comparison of kernel speedups across the approaches and applications on an MI300A. Times are shown for all kernels in an application (Lulesh 15, MiniFE 7, s3d 54, and miniWeather 7) and dots represent outliers. R 3 yields higher median and max speedups over BO and OpenEvolve. results are shown as a distribution over the kernels within an application, e.g. Lulesh has 15 kernels. For all applications R 3 has a higher median and max speedup than OpenEvolve and record-replay + BO. Both R 3 and OpenEvolve tend to perform stronger than record-replay + BO demonstrating the impact of including source code rewrites in kernel tuning. The same results are shown on an H100 GPU in Figure 7 where we see the same trends as with AMD: R 3 yields better optimized kernels than the two baseline approaches. LuleshMiniFEs3dminiWeather Application 1 2 3 4 Speedup Kernel Speedup Distribution on H100 Record-Replay + BOOpenEvolve R 3 (ours) Fig. 7. Comparison of kernel speedups across the approaches and applications on an H100. Times are shown for all kernels in an application and dots represent outliers. R 3 yields higher median and max speedups over baselines. 0.0020.0030.0040.0050.006 Record-Replay + BO Kernel Time (ms) 0.002 0.003 0.004 0.005 0.006 R 3 Kernel Time (ms) Record-Replay + BO faster R 3 (ours) faster 7 kernels Kernel Times for miniWeather Kernel y = x 0.0020.0030.0040.0050.006 OpenEvolve Kernel Time (ms) 0.002 0.003 0.004 0.005 0.006 R 3 Kernel Time (ms) OpenEvolve faster R 3 (ours) faster 7 kernels Kernel Times for miniWeather Kernel y = x Fig. 8.Comparison of absolute times from kernels in the miniWeather application. The y-axis encodes kernel time after R 3 optimization and the x-axis encodes kernel time after optimization from record-replay + BO (left plot) and OpenEvolve (right plot). Points under the line are faster with R 3 . Not only do the distributions of kernel speedups improve from OpenEvolve to R 3 , but so do the individual kernel times. Figure 8 shows the absolute time of the miniWeather kernels as reported by ROCr for all three approaches. We see that for every kernel R 3 yields a faster optimized kernel than record- replay + BO or OpenEvolve. 110100100010000 Time-to-Solution (min) 1.0 1.1 1.2 1.3 1.4 1.5 1.6 1.7 1.8 Speedup over Baseline MinutesHoursDays 8.5 hours faster Application Lulesh MiniFE s3d miniWeather Algorithm Record-Replay + BO OpenEvolve R 3 (ours) End-to-End Speedup vs. Time-to-Solution Fig. 9. Comparison of final achieved speedup versus the time-to-solution. The vertical axis shows speedup over the original application and the horizontal axis shows time-to-solution in minutes (note log scale). While R 3 is slower than BO with record-replay, it produces consistently better tuning results. Furthermore, R 3 is both faster and yields better speedups than OpenEvolve. Finally, Figure 9 shows the speedup of the final application with the optimized kernels from each approach compared with the time-to-solution. We immediately see two things: (1) R 3 achieves the best overall speedups by combining source level and pass/launch parameter tuning, and (2) R 3 ’s record- replay engine drastically speeds up evolutionary LLM code optimization. R 3 is both faster and yields better results than OpenEvolve, while it is slower than record-replay + BO but still yields better speedups. This is expected as introducing LLMs and another axis of search to tuning will increase time- to-solution. However, without its optimizations R 3 would take nearly a day to run for a single application. For MiniFE, R 3 was 8.5 hours faster than OpenEvolve, enabling us to find solutions faster or try more samples in the same allotted time. X. CASE STUDY: LQCD We present a case study using R 3 to optimize an expensive kernel in QUDA [40], a lattice QCD library, achieving a 28.4% reduction in compute time on AMD MI300A GPUs. A. Kernel Description QUDA [40] is a library for lattice quantum chromody- namics (QCD) calculations. It employs iterative solvers with adaptive multigrid preconditioning to solve the large sparse linear systems that dominate these simulations. It is an ideal candidate for stress testing R 3 as it is a full, production HPC application with hundreds of files, tens of thousands of lines of code, long build times, and big workloads. AlphaEvolve style optimization on this repository would take days making it infeasible for practical use. Our optimization target is the CoarseDslash kernel, which applies the coarse-level Dirac operator within QUDA’s multigrid preconditioner. Each thread operates on a single lattice site, computing output spinor values via gauge-covariant nearest-neighbor hopping and local clover contributions. Dur- ing profiling, we find it to be the most expensive kernel in the codebase. The kernel itself, plus all inline device functions it calls, is large spanning several hundred lines of code. B. R 3 Optimization Results We run R 3 with the same setup as the previous experiments to find optimizations for the CoarseDslash kernel. In the final optimized kernel, R 3 reorganizes the kernel inner loops so that each input spinor or halo value is loaded once for a given spin/color-column pair and then reused across the output color rows handled by the thread. The optimized code also hoists repeatedly used invariants, including references to the gauge and clover fields, halo buffers, source spinors, and row- base indices, out of the hottest loops. This reduces redundant field accesses, repeated address arithmetic, and other indexing overhead. These changes are from LLM source rewrites. We also observe several optimizations enabled due to the hierarchical nature of R 3 . Several parts of the kernel are refactored to extract compile-time constants and conditionals, e.g. the dslash/clover selection, warp-fission index calculation, etc.; this enables the record-replay engine’s compiler constant specialization optimizations. Finding such optimizations would not be possible without a hierarchical search as in R 3 . The final kernel speedup on an MI300A is 1.33×. When placed into the whole QUDA application we see a 28.4% re- duction in time spent in compute when run on a representative workload on 64 GPUs (we omit data loading and tear-down time in our profiling). Due to the time and cost associated with it, we do not run OpenEvolve on the same problem, but based on standalone evaluation times we project it would take just under 2 days (42 hours), while R 3 finished in 108 minutes. XI. CONCLUSION We presented Record-Remix-Replay (R 3 ), a hierarchical GPU kernel optimization framework that combines LLM- guided evolution for source-level rewrites with replay-based Bayesian optimization for compiler passes and launch config- urations. By making candidate evaluation fast through record- replay and related systems optimizations, R 3 makes broad search over implementation, compiler, and launch decisions practical. Across four scientific applications, R 3 achieves better kernel and application-level performance than record- replay + BO alone, better final results than OpenEvolve, and substantially lower time-to-solution than evolutionary search. ACKNOWLEDGMENT This work was performed under the auspices of the U.S. De- partment of Energy by Lawrence Livermore National Labora- tory (LLNL) under Contract DE-AC52-07NA27344 (LLNL- CONF-2017736). This work was supported in part by LLNL LDRD projects 25-ERD-058. This material is based upon work supported by the U.S. Department of Energy, Office of Science, Office of Advanced Scientific Computing Research, through solicitation DE-FOA-0003264, “Advancements in Ar- tificial Intelligence for Science,” under Award Number DE- SC0025598 and contract DE-AC52-07NA27344. Computing support for this work came from the Lawrence Livermore National Laboratory (LLNL) Institutional Computing Grand Challenge program. REFERENCES [1] J. Ansel, S. Kamil, K. Veeramachaneni, J. Ragan-Kelley, J. Bosboom, U.-M. O’Reilly, and S. Amarasinghe, “Opentuner: An extensible framework for program autotuning,” in 2014 23rd International Conference on Parallel Architecture and Compilation Techniques (PACT), 2014, p. 303–315. [Online]. Available: https://doi.org/10. 1145/2628071.2628092 [2] C. Tapus, I.-H. Chung, and J. Hollingsworth, “Active harmony: Towards automated performance tuning,” in SC ’02: Proceedings of the 2002 ACM/IEEE Conference on Supercomputing, 2002, p. 44–44. [Online]. Available: https://doi.org/10.1109/SC.2002.10062 [3] S. Heldens and B. van Werkhoven, “Kernel launcher: C++ library for optimal-performance portable cuda applications,” arXiv preprint arXiv:2303.12374, 2023. [Online]. Available: https://doi.org/10.48550/ arXiv.2303.12374 [4] C. Nugteren and V. Codreanu, “Cltune: A generic auto-tuner for opencl kernels,” in 2015 IEEE 9th International Symposium on Embedded Multicore/Many-core Systems-on-Chip (MCSoC).Los Alamitos, CA, USA: IEEE Computer Society, sep 2015, p. 195–202. [Online]. Available: https://doi.org/10.1109/MCSoC.2015.10 [5] M. Popov, C. Akel, W. Jalby, and P. de Oliveira Castro, “Piecewise holistic autotuning of compiler and runtime parameters,” in Euro- Par 2016: Parallel Processing, P.-F. Dutot and D. Trystram, Eds. Cham: Springer International Publishing, 2016, p. 238–250. [Online]. Available: https://doi.org/10.1007/978-3-319-43659-3 18 [6] P. de Oliveira Castro, C. Akel, E. Petit, M. Popov, and W. Jalby, “CERE: llvm-based codelet extractor and replayer for piecewise benchmarking and optimization,” ACM Trans. Archit. Code Optim., vol. 12, no. 1, p. 6:1–6:24, 2015. [Online]. Available: https://doi.org/10.1145/2724717 [7] Z. Xie, J. Liu, J. Li, and D. Li, “Merchandiser: Data placement on heterogeneous memory for task-parallel hpc applications with load- balance awareness,” in Proceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming, ser. PPoPP ’23.New York, NY, USA: Association for Computing Machinery, 2023, p. 204–217. [Online]. Available: https://doi.org/10. 1145/3572848.3577497 [8] K. Parasyris, G. Georgakoudis, E. Rangel, I. Laguna, and J. Doerfert, “Scalable Tuning of (OpenMP) GPU Applications via Kernel Record and Replay,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC ’23. New York, NY, USA: Association for Computing Machinery, Nov. 2023, p. 1–14. [Online]. Available: https://doi.org/10.1145/3581784.3607098 [9] L. Liu, T. Fei, Z. Zhu, K. Wu, and Y. Zhang, “A Survey of Evolutionary Algorithms,” in 2023 4th International Conference on Big Data, Artificial Intelligence and Internet of Things Engineering (ICBAIE), Aug. 2023, p. 22–27. [Online]. Available: https://doi.org/ 10.1109/ICBAIE59714.2023.10281260 [10] B. Romera-Paredes, M. Barekatain, A. Novikov, M. Balog, M. P. Kumar, E. Dupont, F. J. R. Ruiz, J. S. Ellenberg, P. Wang, O. Fawzi, P. Kohli, and A. Fawzi, “Mathematical discoveries from program search with large language models,” Nature, vol. 625, no. 7995, p. 468–475, Jan. 2024. [Online]. Available: https://doi.org/10.1038/s41586-023-06924-6 [11] A. Novikov, N. V ̃ u, M. Eisenberger, E. Dupont, P.-S. Huang, A. Z. Wagner, S. Shirobokov, B. Kozlovskii, F. J. R. Ruiz, A. Mehrabian, M. P. Kumar, A. See, S. Chaudhuri, G. Holland, A. Davies, S. Nowozin, P. Kohli, and M. Balog, “AlphaEvolve: A coding agent for scientific and algorithmic discovery,” Jun. 2025. [Online]. Available: https://doi.org/10.48550/arXiv.2506.13131 [12] A. Cheng, S. Liu, M. Pan, Z. Li, B. Wang, A. Krentsel, T. Xia, M. Cemri, J. Park, S. Yang, J. Chen, L. Agrawal, A. Desai, J. Xing, K. Sen, M. Zaharia, and I. Stoica, “Barbarians at the Gate: How AI is Upending Systems Research,” Oct. 2025. [Online]. Available: https://doi.org/10.48550/arXiv.2510.06189 [13] A.Sharma,“Openevolve:anopen-sourceevolutionarycod- ingagent,”2025.[Online].Available:https://github.com/ algorithmicsuperintelligence/openevolve [14] J.-B. Mouret and J. Clune, “Illuminating search spaces by mapping elites,” Apr. 2015. [Online]. Available: https://doi.org/10.48550/arXiv. 1504.04909 [15] B. van Werkhoven, “Kernel tuner: A search-optimizing gpu code auto- tuner,” Future Generation Computer Systems, vol. 90, p. 347–358, 2019. [Online]. Available: https://doi.org/10.1016/j.future.2018.08.004 [16] T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, M. Cowan, H. Shen, L. Wang, Y. Hu, L. Ceze, C. Guestrin, and A. Krishnamurthy, “TVM: An Automated End-to-End Optimizing Compiler for Deep Learning,” Feb. 2018. [Online]. Available: https://doi.org/10.48550/arXiv.1802.04799 [17] L. Zheng, C. Jia, M. Sun, Z. Wu, C. H. Yu, A. Haj-Ali, Y. Wang, J. Yang, D. Zhuo, K. Sen, J. E. Gonzalez, and I. Stoica, “Ansor: Generating high-performance tensor programs for deep learning,” arXiv, 2020, oSDI 2020 (arXiv version). [Online]. Available: https://doi.org/10.48550/arXiv.2006.06762 [18] J. Shao, X. Zhou, S. Feng, B. Hou, R. Lai, H. Jin, W. Lin, M. Masuda, C. H. Yu, and T. Chen, “Tensor program optimization with probabilistic programs,” arXiv, 2022, accepted to NeurIPS 2022 (arXiv version). [Online]. Available: https://doi.org/10.48550/arXiv.2205.13603 [19] A. Adams, K. Ma, L. Anderson, R. Baghdadi, T.-M. Li, M. Gharbi, B. Steiner, S. Johnson, K. Fatahalian, F. Durand, and J. Ragan-Kelley, “Learning to optimize halide with tree search and random programs,” ACM Transactions on Graphics, vol. 38, no. 4, p. 121:1–121:12, 2019. [Online]. Available: https://doi.org/10.1145/3306346.3322967 [20] P. Tillet, H. T. Kung, and D. Cox, “Triton: An intermediate language and compiler for tiled neural network computations,” in Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages (MAPL 2019), 2019, p. 10–19. [Online]. Available: https://doi.org/10.1145/3315508.3329973 [21] A. H. Ashouri, A. Bignoli, G. Palermo, C. Silvano, S. Kulkarni, and J. Cavazos, “Micomp: Mitigating the compiler phase-ordering problem using optimization sub-sequences and machine learning,” ACM Trans. Archit. Code Optim., vol. 14, no. 3, Sep. 2017. [Online]. Available: https://doi.org/10.1145/3124452 [22] C. Cummins, B. Wasti, J. Guo, B. Cui, J. Ansel, S. Gomez, S. Jain, J. Liu, O. Teytaud, B. Steiner, Y. Tian, and H. Leather, “Compilergym: Robust, performant compiler optimization environments for ai research,” in 2022 IEEE/ACM International Symposium on Code Generation and Optimization (CGO), 2022, p. 92–105. [Online]. Available: https://doi.org/10.1109/CGO53902.2022.9741258 [23] D. Beckingsale, O. Pearce, I. Laguna, and T. Gamblin, “Apollo: Reusable models for fast, dynamic tuning of input-dependent code,” in 2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS).IEEE, 2017, p. 307–316. [Online]. Available: https://doi.org/10.1109/IPDPS.2017.38 [24] C. Wood, G. Georgakoudis, D. Beckingsale, D. Poliakoff, A. Gimenez, K. Huck, A. Malony, and T. Gamblin, “Artemis: Automatic runtime tuning of parallel execution parameters using machine learning,” in High Performance Computing, B. L. Chamberlain, A.-L. Varbanescu, H. Ltaief, and P. Luszczek, Eds.Cham: Springer International Publishing, 2021, p. 453–472. [Online]. Available: https://doi.org/10.1007/978-3-030-78713-4 24 [25] H. Menon, A. Bhatele, and T. Gamblin, “Auto-tuning parameter choices in hpc applications using bayesian optimization,” in 2020 IEEE International Parallel and Distributed Processing Symposium (IPDPS), 2020, p. 831–840. [Online]. Available: https://doi.org/10. 1109/IPDPS47924.2020.00090 [26] Y. Liu, W. M. Sid-Lakhdar, O. Marques, X. Zhu, C. Meng, J. W. Demmel, and X. S. Li, “Gptune: Multitask learning for autotuning exascale applications,” in Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming, ser. PPoPP ’21.New York, NY, USA: Association for Computing Machinery, 2021, p. 234–246. [Online]. Available: https://doi.org/10.1145/3437801.3441621 [27] R. Mijakovi ́ c, M. Firbach, and M. Gerndt, “An architecture for flexible auto-tuning: The periscope tuning framework 2.0,” in 2016 2nd International Conference on Green High Performance Computing (ICGHPC), 2016, p. 1–9. [Online]. Available: https: //doi.org/10.1109/ICGHPC.2016.7508066 [28] G. Georgakoudis, K. Parasyris, and D. Beckingsale, “Proteus: Portable Runtime Optimization of GPU Kernel Execution with Just-in-Time Compilation,” in Proceedings of the 23rd ACM/IEEE International Symposium on Code Generation and Optimization, ser. CGO ’25. New York, NY, USA: Association for Computing Machinery, Mar. 2025, p. 507–522. [Online]. Available: https://doi.org/10.1145/3696443.3708939 [29] D. Nichols, K. Parasyris, C. Jekel, A. Bhatele, and H. Menon, “Integratingperformancetoolsinmodelreasoningforgpu kernel optimization,” 2025, https://doi.org/10.48550/arXiv.2510.17158. [Online]. Available: https://arxiv.org/abs/2510.17158 [30] D. Nichols, P. Polasam, H. Menon, A. Marathe, T. Gamblin, and A. Bhatele, “ Performance-Aligned LLMs for Generating Fast HPC Code ,” IEEE Transactions on Parallel & Distributed Systems, no. 01, p. 1–12, Mar. 5555, https://doi.org/10.1109/TPDS.2026.3675550. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/TPDS. 2026.3675550 [31] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, “Efficient Memory Management for Large Language Model Serving with PagedAttention,” Sep. 2023, arXiv:2309.06180 [cs]. [Online]. Available: https://doi.org/10.48550/ arXiv.2309.06180 [32] I. Karlin, J. Keasler, and J. R. Neely, “Lulesh 2.0 updates and changes,” Lawrence Livermore National Laboratory (LLNL), Tech. Rep., 07 2013. [Online]. Available: https://doi.org/10.2172/1090032 [33] P. T. Lin, M. A. Heroux, R. F. Barrett, and A. B. Williams, “Assessing a mini-application as a performance proxy for a finite element method engineering application,” Concurrency and Computation: Practice and Experience, vol. 27, no. 17, p. 5374–5389, 2015. [Online]. Available: https://doi.org/10.1002/cpe.3587 [34] A. Danalis, G. Marin, C. McCurdy, J. S. Meredith, P. C. Roth, K. Spafford, V. Tipparaju, and J. S. Vetter, “The scalable heterogeneous computing (shoc) benchmark suite,” in Proceedings of the 3rd Workshop on General-Purpose Computation on Graphics Processing Units, ser. GPGPU-3.New York, NY, USA: Association for Computing Machinery, 2010, p. 63–74. [Online]. Available: https://doi.org/10.1145/1735688.1735702 [35] M. R. Norman, “miniweather,” [Computer Software], March 2020. [Online]. Available: https://doi.org/10.11578/dc.20201001.88 [36] Z. Jin and J. S. Vetter, “A Benchmark Suite for Improving Performance Portability of the SYCL Programming Model,” in 2023 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), Apr. 2023, p. 325–327. [Online]. Available: https://doi.org/10.1109/ISPASS57527.2023.00041 [37] J. Bergstra, R. Bardenet, Y. Bengio, and B. K ́ egl, “Algorithms for hyper-parameter optimization,” in Proceedings of the 25th International Conference on Neural Information Processing Systems, ser. NIPS’11. RedHook,NY,USA:CurranAssociatesInc.,Dec.2011, p.2546–2554,https://dl.acm.org/doi/10.5555/2986459.2986743. [Online].Available:https://papers.nips.c/paper/ 4443-algorithms-for-hyper-parameter-optimization [38] OpenAI, S. Agarwal, L. Ahmad, J. Ai, S. Altman, A. Applebaum, E. Arbus, R. K. Arora, Y. Bai, B. Baker, H. Bao, B. Barak, A. Bennett, T. Bertao, N. Brett, E. Brevdo, G. Brockman, S. Bubeck, C. Chang, K. Chen, M. Chen, E. Cheung, A. Clark, D. Cook, M. Dukhan, C. Dvorak, K. Fives, V. Fomenko, T. Garipov, K. Georgiev, M. Glaese, T. Gogineni, A. Goucher, L. Gross, K. G. Guzman, J. Hallman, J. Hehir, J. Heidecke, A. Helyar, H. Hu, R. Huet, J. Huh, S. Jain, Z. Johnson, C. Koch, I. Kofman, D. Kundel, J. Kwon, V. Kyrylov, E. Y. Le, G. Leclerc, J. P. Lennon, S. Lessans, M. Lezcano-Casado, Y. Li, Z. Li, J. Lin, J. Liss, Lily, Liu, J. Liu, K. Lu, C. Lu, Z. Martinovic, L. McCallum, J. McGrath, S. McKinney, A. McLaughlin, S. Mei, S. Mostovoy, T. Mu, G. Myles, A. Neitz, A. Nichol, J. Pachocki, A. Paino, D. Palmie, A. Pantuliano, G. Parascandolo, J. Park, L. Pathak, C. Paz, L. Peran, D. Pimenov, M. Pokrass, E. Proehl, H. Qiu, G. Raila, F. Raso, H. Ren, K. Richardson, D. Robinson, B. Rotsted, H. Salman, S. Sanjeev, M. Schwarzer, D. Sculley, H. Sikchi, K. Simon, K. Singhal, Y. Song, D. Stuckey, Z. Sun, P. Tillet, S. Toizer, F. Tsimpourlas, N. Vyas, E. Wallace, X. Wang, M. Wang, O. Watkins, K. Weil, A. Wendling, K. Whinnery, C. Whitney, H. Wong, L. Yang, Y. Yang, M. Yasunaga, K. Ying, W. Zaremba, W. Zhan, C. Zhang, B. Zhang, E. Zhang, and S. Zhao, “gpt-oss-120b & gpt-oss-20b Model Card,” Aug. 2025, arXiv:2508.10925 [cs]. [Online]. Available: https://doi.org/10.48550/arXiv.2508.10925 [39] A. Singh, A. Fry, A. Perelman, A. Tart, A. Ganesh, A. El-Kishky, A. McLaughlin, A. Low, A. J. Ostrow, A. Ananthram, A. Nathan, A. Luo, A. Helyar, A. Madry, A. Efremov, A. Spyra, A. Baker- Whitcomb, A. Beutel, A. Karpenko, A. Makelov, A. Neitz, A. Wei, A. Barr, A. Kirchmeyer, A. Ivanov, A. Christakis, A. Gillespie, A. Tam, A. Bennett, A. Wan, A. Huang, A. M. Sandjideh, A. Yang, A. Kumar, A. Saraiva, A. Vallone, A. Gheorghe, A. G. Garcia, A. Braunstein, A. Liu, A. Schmidt, A. Mereskin, A. Mishchenko, A. Applebaum, A. Rogerson, A. Rajan, A. Wei, A. Kotha, A. Srivastava, A. Agrawal, A. Vijayvergiya, A. Tyra, A. Nair, A. Nayak, B. Eggers, B. Ji, B. Hoover, B. Chen, B. Chen, B. Barak, B. Minaiev, B. Hao, B. Baker, B. Lightcap, B. McKinzie, B. Wang, B. Quinn, B. Fioca, B. Hsu, B. Yang, B. Yu, B. Zhang, B. Brenner, C. R. Zetino, C. Raymond, C. Lugaresi, C. Paz, C. Hudson, C. Whitney, C. Li, C. Chen, C. Cole, C. Voss, C. Ding, C. Shen, C. Huang, C. Colby, C. Hallacy, C. Koch, C. Lu, C. Kaplan, C. Kim, C. J. Minott-Henriques, C. Frey, C. Yu, C. Czarnecki, C. Reid, C. Wei, C. Decareaux, C. Scheau, C. Zhang, C. Forbes, D. Tang, D. Goldberg, D. Roberts, D. Palmie, D. Kappler, D. Levine, D. Wright, D. Leo, D. Lin, D. Robinson, D. Grabb, D. Chen, D. Lim, D. Salama, D. Bhattacharjee, D. Tsipras, D. Li, D. Yu, D. J. Strouse, D. Williams, D. Hunn, E. Bayes, E. Arbus, E. Akyurek, E. Y. Le, E. Widmann, E. Yani, E. Proehl, E. Sert, E. Cheung, E. Schwartz, E. Han, E. Jiang, E. Mitchell, E. Sigler, E. Wallace, E. Ritter, E. Kavanaugh, E. Mays, E. Nikishin, F. Li, F. P. Such, F. d. A. B. Peres, F. Raso, F. Bekerman, F. Tsimpourlas, F. Chantzis, F. Song, F. Zhang, G. Raila, G. McGrath, G. Briggs, G. Yang, G. Parascandolo, G. Chabot, G. Kim, G. Zhao, G. Valiant, G. Leclerc, H. Salman, H. Wang, H. Sheng, H. Jiang, H. Wang, H. Jin, H. Sikchi, H. Schmidt, H. Aspegren, H. Chen, H. Qiu, H. Lightman, I. Covert, I. Kivlichan, I. Silber, I. Sohl, I. Hammoud, I. Clavera, I. Lan, I. Akkaya, I. Kostrikov, I. Kofman, I. Etinger, I. Singal, J. Hehir, J. Huh, J. Pan, J. Wilczynski, J. Pachocki, J. Lee, J. Quinn, J. Kiros, J. Kalra, J. Samaroo, J. Wang, J. Wolfe, J. Chen, J. Wang, J. Harb, J. Han, J. Wang, J. Zhao, J. Chen, J. Yang, J. Tworek, J. Chand, J. Landon, J. Liang, J. Lin, J. Liu, J. Wang, J. Tang, J. Yin, J. Jang, J. Morris, J. Flynn, J. Ferstad, J. Heidecke, J. Fishbein, J. Hallman, J. Grant, J. Chien, J. Gordon, J. Park, J. Liss, J. Kraaijeveld, J. Guay, J. Mo, J. Lawson, J. McGrath, J. Vendrow, J. Jiao, J. Lee, J. Steele, J. Wang, J. Mao, K. Chen, K. Hayashi, K. Xiao, K. Salahi, K. Wu, K. Sekhri, K. Sharma, K. Singhal, K. Li, K. Nguyen, K. Gu-Lemberg, K. King, K. Liu, K. Stone, K. Yu, K. Ying, K. Georgiev, K. Lim, K. Tirumala, K. Miller, L. Ahmad, L. Lv, L. Clare, L. Fauconnet, L. Itow, L. Yang, L. Romaniuk, L. Anise, L. Byron, L. Pathak, L. Maksin, L. Lo, L. Ho, L. Jing, L. Wu, L. Xiong, L. Mamitsuka, L. Yang, L. McCallum, L. Held, L. Bourgeois, L. Engstrom, L. Kuhn, L. Feuvrier, L. Zhang, L. Switzer, L. Kondraciuk, L. Kaiser, M. Joglekar, M. Singh, M. Shah, M. Stratta, M. Williams, M. Chen, M. Sun, M. Cayton, M. Li, M. Zhang, M. Aljubeh, M. Nichols, M. Haines, M. Schwarzer, M. Gupta, M. Shah, M. Huang, M. Dong, M. Wang, M. Glaese, M. Carroll, M. Lampe, M. Malek, M. Sharman, M. Zhang, M. Wang, M. Pokrass, M. Florian, M. Pavlov, M. Wang, M. Chen, M. Wang, M. Feng, M. Bavarian, M. Lin, M. Abdool, M. Rohaninejad, N. Soto, N. Staudacher, N. LaFontaine, N. Marwell, N. Liu, N. Preston, N. Turley, N. Ansman, N. Blades, N. Pancha, N. Mikhaylin, N. Felix, N. Handa, N. Rai, N. Keskar, N. Brown, O. Nachum, O. Boiko, O. Murk, O. Watkins, O. Gleeson, P. Mishkin, P. Lesiewicz, P. Baltescu, P. Belov, P. Zhokhov, P. Pronin, P. Guo, P. Thacker, Q. Liu, Q. Yuan, Q. Liu, R. Dias, R. Puckett, R. Arora, R. T. Mullapudi, R. Gaon, R. Miyara, R. Song, R. Aggarwal, R. J. Marsan, R. Yemiru, R. Xiong, R. Kshirsagar, R. Nuttall, R. Tsiupa, R. Eldan, R. Wang, R. James, R. Ziv, R. Shu, R. Nigmatullin, S. Jain, S. Talaie, S. Altman, S. Arnesen, S. Toizer, S. Toyer, S. Miserendino, S. Agarwal, S. Yoo, S. Heon, S. Ethersmith, S. Grove, S. Taylor, S. Bubeck, S. Banesiu, S. Amdo, S. Zhao, S. Wu, S. Santurkar, S. Zhao, S. R. Chaudhuri, S. Krishnaswamy, Shuaiqi, Xia, S. Cheng, S. Anadkat, S. P. Fishman, S. Tobin, S. Fu, S. Jain, S. Mei, S. Egoian, S. Kim, S. Golden, S. Q. Mah, S. Lin, S. Imm, S. Sharpe, S. Yadlowsky, S. Choudhry, S. Eum, S. Sanjeev, T. Khan, T. Stramer, T. Wang, T. Xin, T. Gogineni, T. Christianson, T. Sanders, T. Patwardhan, T. Degry, T. Shadwell, T. Fu, T. Gao, T. Garipov, T. Sriskandarajah, T. Sherbakov, T. Kaftan, T. Hiratsuka, T. Wang, T. Song, T. Zhao, T. Peterson, V. Kharitonov, V. Chernova, V. Kosaraju, V. Kuo, V. Pong, V. Verma, V. Petrov, W. Jiang, W. Zhang, W. Zhou, W. Xie, W. Zhan, W. McCabe, W. DePue, W. Ellsworth, W. Bain, W. Thompson, X. Chen, X. Qi, X. Xiang, X. Shi, Y. Dubois, Y. Yu, Y. Khakbaz, Y. Wu, Y. Qian, Y. T. Lee, Y. Chen, Y. Zhang, Y. Xiong, Y. Tian, Y. Cha, Y. Bai, Y. Yang, Y. Yuan, Y. Li, Y. Zhang, Y. Yang, Y. Jin, Y. Jiang, Y. Wang, Y. Wang, Y. Liu, Z. Stubenvoll, Z. Dou, Z. Wu, and Z. Wang, “OpenAI GPT-5 System Card,” Dec. 2025, arXiv:2601.03267 [cs]. [Online]. Available: https://doi.org/10.48550/arXiv.2601.03267 [40] M. Clark, R. Babich, K. Barros, R. Brower, and C. Rebbi, “Solving lattice qcd systems of equations using mixed precision solvers on gpus,” Computer Physics Communications, vol. 181, no. 9, p. 1517–1528, 2010. [Online]. Available: https://doi.org/10.1016/j.cpc.2010.05.002