Paper deep dive
Autonomous discovery of accelerator commissioning algorithms
Thorsten Hellert
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/10/2026, 4:19:35 AM
Summary
This paper introduces an 'autoresearch' framework where a language-model agent autonomously discovers and optimizes accelerator commissioning algorithms through a closed loop of code modification, simulation testing, and validation. Applied to the ALS-U accumulator ring, the system improves RF beam capture procedures and generates a Pareto front of 16 non-dominated algorithms balancing rapid capture against error correction, demonstrating that AI agents can participate directly in discovering physics control strategies rather than just executing human-defined ones.
Entities (8)
Relation Signals (7)
Autoresearch â appliedto â ALS-U Accumulator Ring
confidence 98% · Applied to RF beam capture in the ALS-U accumulator-ring model
Autoresearch â uses â Language Model Agent
confidence 95% · This Letter demonstrates a closed research loop in which a language-model agent writes commissioning code
Autoresearch â generates â 16 non-dominated algorithms
confidence 94% · Extending the same framework to multiple objectives produces 16 non-dominated algorithms
Autoresearch â optimizes â RF Beam Capture
confidence 92% · The task is beam capture in the ALS-U accumulator ring... asking how far an agent can improve on a human-designed baseline.
Autoresearch â employs â Pareto dominance
confidence 90% · The scalar merge predicate is replaced by Pareto dominance
Autoresearch â utilizes â pySC
confidence 88% · implemented in pySC [47, 48]
Language Model Agent â basedon â Anthropic Claude
confidence 85% · We first evaluate this capability across model tiers... Anthropic Claude family
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Simulated commissioning has become essential for de-risking modern light-source design and commissioning, but the procedures being simulated are still designed entirely by human experts. Their labor-intensive redevelopment after lattice changes makes such studies hard to repeat and limits their use during early design iteration. This Letter demonstrates a closed research loop in which a language-model agent writes commissioning code, tests it in simulation, and improves the algorithm from the results. Applied to RF beam capture in the ALS-U accumulator-ring model, the loop substantially improves a working expert procedure and can construct a working one from a minimal starting point, with more capable models succeeding from less initial code. Extending the same framework to multiple objectives produces 16 non-dominated algorithms spanning physically distinct trade-offs between rapid beam capture and correction of seeded machine errors. This reframes commissioning studies from evaluating human-designed procedures toward a mode in which agents participate directly in discovering accelerator algorithms.
Tags
Links
- Source: https://arxiv.org/abs/2608.07138v1
- Canonical: https://arxiv.org/abs/2608.07138v1
Trouble viewing inline? Open PDF directly â
Full Text
44,769 characters extracted from source content.
Expand or collapse full text
Autonomous discovery of accelerator commissioning algorithms Thorsten Hellert Lawrence Berkeley National Laboratory, Berkeley, CA 94720, USA Abstract Simulated commissioning has become essential for de-risking modern light-source design and commissioning, but the procedures being simulated are still designed entirely by human experts. Their labor-intensive redevelopment after lattice changes makes such studies hard to repeat and limits their use during early design iteration. This Letter demonstrates a closed research loop in which a language-model agent writes commissioning code, tests it in simulation, and improves the algorithm from the results. Applied to RF beam capture in the ALS-U accumulator-ring model, the loop substantially improves a working expert procedure and can construct a working one from a minimal starting point, with more capable models succeeding from less initial code. Extending the same framework to multiple objectives produces 16 non-dominated algorithms spanning physically distinct trade-offs between rapid beam capture and correction of seeded machine errors. This reframes commissioning studies from evaluating human-designed procedures toward a mode in which agents participate directly in discovering accelerator algorithms. Figure 1: The autoresearch loop, iterated without a human in the inner loop. A proposer modifies the target algorithm, an independent reviewer screens the diff, and the fixed harness evaluates surviving candidates on the seed ensemble. Only candidates that improve the incumbent ensemble-mean cost are merged. The agent may modify the algorithm and its helper library; the harness, simulator ground truth, and cost model remain protected. I Introduction Fourth-generation multi-bend-achromat (MBA) storage rings [1, 2, 3, 4, 5] combine small dynamic apertures, strong nonlinearities, and tight tolerances, leaving limited room for commissioning errors [6]. Developing the commissioning procedure primarily on the real machine would therefore be both risky and costly in user dark time. Instead, the lattice and its commissioning strategy must be stress-tested in simulation before beam becomes available [3]. Simulated commissioning has consequently become a central design tool, serving both as an error-analysis step for the lattice and as a validation environment for the procedure [7, 8]. Here a proposed correction sequence is run on ensembles of randomly perturbed machines to test whether the design can be commissioned with realistic alignment, calibration, and diagnostics errors. Such studies are now used across MBA projects to assess commissioning feasibility and robustness [9, 10, 11, 12, 13, 14, 15, 16, 17], derive alignment tolerances [18, 19], and develop procedures for first-turn threading and RF capture [20, 21]. This capability, however, depends on having a suitable commissioning procedure to simulate. Such a procedure is itself a substantial expert-developed artifact: it specifies the order of correction steps, the diagnostics used at each stage, the machine variables that may be adjusted, the stopping criteria, and the recovery logic when a simulated seed fails. Because these procedures depend strongly on the lattice, error model, diagnostics, and available controls, they often require substantial redevelopment after design changes. Detailed studies therefore usually begin only after the lattice and hardware layout have matured, limiting their use during earlier comparisons of substantially different concepts. Existing automation reduces the effort required to execute and tune commissioning procedures, but it does not remove the need to design those procedures in the first place. Startup sequences can be encoded in the control system [22], and model-based, machine-learning, Bayesian, and reinforcement-learning methods can optimize orbits, tunes, injection, free-electron-laser performance, and nonlinear-dynamics objectives once the variables and objective have been specified [23, 24, 25, 26, 27, 28, 29]. More recently, language-model agents have begun to interact with accelerator controls and operational tools [30, 31, 32, 33, 34]. These approaches automate execution, optimization, or coordination within a task structure supplied by experts. The objective, accessible variables, evaluation criteria, and overall procedure remain human-defined. This Letter addresses the remaining design problem by moving the search from the machine variables to the commissioning procedure itself. Rather than asking an automated method to execute or tune a fixed algorithm, we ask an agent to modify the algorithm, evaluate the result, and retain only validated improvements. The idea of autoresearch was first introduced by Karpathy [35] as a greedy cycle that edits code, runs a fixed-budget computational experiment, and keeps only improvements. We adapt this pattern to accelerator commissioning: here the object being changed is the commissioning algorithm, and the harness and physics-based evaluation that judge each change are built specifically for this setting. An agent proposes a code change, an independent reviewer checks the diff (the exact set of proposed code changes) for invalid simulator access or unphysical shortcuts, and a fixed harness evaluates each surviving candidate on the same ensemble of error seeds. The change is merged into the retained code only if it improves on the current best algorithm under the predeclared metric. The loop therefore turns a traditionally manual part of simulated commissioning into a closed experimental cycle: propose a modification, test it against the ensemble, and retain only validated improvements. Similar proposeâevaluateâselect loops are emerging in AI-scientist systems, algorithm discovery, self-driving laboratories, and scientific design automation [36, 37, 38, 39, 40, 41, 42, 43, 44]. The demonstration considered here is RF beam capture in the ALS-U accumulator-ring simulated-commissioning model [45, 8]. We choose this step because it is computationally light enough to support repeated autonomous campaigns while remaining a substantive part of the commissioning chain: each attempted injection represents a machine shot, and successful capture has an unambiguous physical definition. The benchmark asks whether an agent can improve a commissioning algorithm through repeated code modification and physics-based evaluation, with no human intervention once a campaign is configured. Although the specific objective is the mean number of injections required for capture, the broader experiment is whether simulation can serve not only to validate commissioning procedures, but also as a laboratory in which those procedures are autonomously developed. I Autonomous search over commissioning algorithms Autonomous procedure design requires a strict separation between the algorithm being developed and the experiment that judges it. In autoresearch [46], the agent may modify the target algorithm and a companion library of reusable routines, but not the lattice construction, seeded errors, simulator state, action costs, capture criterion, evaluation ensemble, or scoring rule. Candidates access the simulated accelerator only through an operator interface exposing control-room actions and measurements: injection, BPM readout, magnet and RF settings, and selected design-model quantities such as response matrices. Simulator ground truth, including seeded alignment and calibration errors, is withheld. Behind it, the fixed harness executes each candidate on randomly perturbed ALS-U accumulator-ring lattices implemented in pySC [47, 48]. Without this separation, an agent could improve the score by exploiting the benchmark rather than the procedure (see Appendix C). Each loop iteration is one computational experiment. As shown in Fig. 1, a proposer develops a candidate modification in an isolated Git worktree, a private copy of the code, then submits a version-controlled diff with its hypothesis. An independent reviewer screens the change, the harness evaluates survivors on the predeclared seed ensemble, and a deterministic merge predicate decides whether the repository is updated. A human configures each campaign once: its task, seed ensemble, action budget, and scoring rule. The proposeâscreenâevaluateâmerge inner loop then runs for a fixed number of experiments with no further human intervention. Propose: The proposer receives three forms of context. First, declarative domain knowledge: short primers on RF capture, sextupole ramping and chromaticity, tune resonances, and regularized orbit feedback; documentation of the operator interface and error model; and a textbook-retrieval subagent over standard accelerator-physics texts [49, 50]. Second, executable procedural knowledge comes from a helper library with routines for first-turn threading, two-turn stitching, sextupole ramping, tune scanning, and RF phase and frequency correction. These are Python ports of the published ALS-U procedures [8], and the agent may call, modify, recombine, or replace them. Third, campaign memory records merged changes, their hypotheses and scores, selected rejected attempts, and the current retained result. A single proposer tends to fall into repetitive approaches, so each experiment assigns one of five rotating personas such as a theorist, an empiricist, or a simplifier. They share the same tools but are prompted toward different strategies, broadening the search. The proposer returns a pull request with a coherent code change, its rationale, and its expected effect on the objective. Screen: A separate reviewer checks that the change respects the benchmark boundary. It rejects candidates that access unavailable quantities, bypass action costs, inspect evaluation state, alter protected components, or otherwise exploit the simulator rather than improve the algorithm. A deterministic pattern scan guards against known forbidden access patterns. Rejected candidates are not executed. Evaluate: The fixed harness evaluates each surviving candidate on the same ensemble of perturbed machines, from identical seeded conditions. A harness-owned objective includes both successful and failed seeds. Section I defines the ensemble, action budget, and scalar objective for beam capture. Merge: In the scalar campaigns, a candidate is merged only if its ensemble-mean score improves on the incumbent. The merge commits it as the new repository head, so subsequent proposals build on the best algorithm so far. The search is therefore greedy rather than globally exhaustive, but every accepted change is version controlled and validated by the unchanged harness. Because campaign memory carries earlier attempts and their outcomes forward, including rejected ones, each new proposal is informed by what has already been tried. In Sec. IV, the scalar rule is replaced by Pareto dominance. I Search for efficient beam-capture algorithms We first apply the loop to a single commissioning task with a well-defined scalar objective, asking how far an agent can improve on a human-designed baseline. The task is beam capture in the ALS-U accumulator ring, the point in the published commissioning chain [8] where a stored, multi-turn beam is first established. Capture requires at least 80 %80\, 37 of 100 tracked particles to survive 500 turns. Each candidate is evaluated on a fixed ensemble of 50 error seeds, with a budget of 500 attempted injections per seed. The objective is the ensemble-mean number of injections required for capture, lower being better. Failed seeds are not discarded. If a seed terminates at phase k of the six-phase baseline procedure (Appendix A), it receives the surrogate cost 500+100â(6âk)500+100(6-k), so an immediate failure costs 1100 injections and a failure at phase 5 costs 600. A captured seed contributes its actual injection count. The rule therefore rewards progress through the procedure even before full capture is achieved. The autoresearch baseline is a faithful pySC port of the published MATLAB algorithm [8]. As noted, beam capture is a tractable but substantive test of whether the loop can improve a working expert procedure. We first evaluate this capability across model tiers. Figure 2 compares three tiers of the Anthropic Claude family (Haiku 4.5, Sonnet 4.6, and Opus 4.6 [51, 52, 53]) under identical task configuration, evaluation ensemble, and 100-experiment campaign budget, with three independent campaigns per tier. The expert baseline scores 207.5 injections. This is a common starting point, not the best performance achievable by a human: the published procedure was designed for robust commissioning across the broader correction chain, not for minimizing injection count in this isolated capture task. Within this benchmark, the loop reduces the objective by a large factor. The best Sonnet and Opus campaigns reach mean scores of 20.3 and 27.5 injections, respectively, while Haiku reaches 53.6. We therefore read this as a capability threshold within the tested tiers: the frontier models enter the tens-of-injections regime, whereas the weaker tier saturates higher. Sonnet and Opus are difficult to separate, suggesting that beam capture distinguishes the weaker tier from the frontier but cannot resolve differences within the frontier group. More complex procedures exercising a broader correction chain may separate them more clearly. Most of the improvement came from streamlining the human procedure rather than from a new capture mechanism: the agent removed cautious or redundant steps and reduced each measurement to the fewest injections that worked. One change stands out as genuinely new: a recovery move that nudges the correctors near the injection point to escape repeated beam loss, rescuing the hardest seeds and driving much of the final gain. We next ask what prior structure the agent needs to build an effective capture procedure, and whether those needs depend on model capability. Figure 3 separates two inputs supplied together in the main capability study: six short documents describing the relevant accelerator physics, and a helper library of reusable commissioning routines derived from the published ALS-U procedure. To expose their contribution, each campaign now begins not from the working baseline but from a minimal stub that does not capture beam, so the agent must construct a functioning procedure from whatever scaffolding is available. With Haiku and Sonnet we test all four combinations: both inputs, the library alone, the documents alone, and neither. This replicated 2Ă22Ă 2 design across two tiers tests not only which input is most valuable, but whether a stronger model can compensate for missing procedural or domain support. Each condition is run in four independent 40-experiment campaigns, compared at the common budget. The helper library is the dominant scaffold. When available, the agent drives the objective into the tens to low hundreds of injections whether or not the documents are present. Without it, the documents alone usually leave campaigns in the hundreds, and the weaker tier essentially never captures from the bare stub. The stronger tier always captures from scratch, but roughly an order of magnitude worse than with it. The two tiers also fail differently when scaffolding is sparse. With neither the helper library nor the knowledge documents, the weaker tier expends far more agent effort per experiment (in reasoning turns and generated code) while still failing to capture, whereas the stronger tier increases effort less and still produces functioning procedures. These differences are quantified in Appendix B. Across all four scaffold conditions, Sonnet also reaches lower final scores than Haiku, with largest gaps in the two without the helper library. The model comparison therefore suggests that stronger models can compensate partly for missing scaffolding, whereas weaker models require more explicit procedural structure. The ablation does not imply that written domain knowledge is unimportant. For this relatively contained coding task, the helper library packages commissioning knowledge in a more directly usable form: executable routines that the agent can inspect, modify, and recombine. Its larger effect is therefore unsurprising. Figure 2: Beam-capture capability across model tiers. Solid traces show the mean best-so-far ensemble score across three campaigns versus experiment number (top) and cumulative compute cost (bottom); shaded bands show the pointwise minimumâmaximum range. The dashed line marks the 207.5-injection expert baseline, evaluated on the same objective but not optimized for injection count. Figure 3: Ablation of domain-knowledge documents and the helper library. Rows show the library present or absent; columns show the documents present or absent. Solid traces give the mean best-so-far injections to capture across four campaigns, and shaded bands show the pointwise minimumâmaximum range. The dotted line marks the 1100-injection non-capture penalty. Because these campaigns start from a non-capturing stub, their scores are not directly comparable to the expert baseline in Fig. 2. IV Search over costâquality trade-offs The scalar beam-capture study asks how efficiently an autonomous agent can minimize one predeclared objective. Real commissioning procedures, however, are often judged by more than the time required to reach beam. A procedure that captures quickly may leave calibration errors, faulty diagnostics, or injection offsets unresolved, while a slower procedure may correct or identify these errors in a way that is useful for subsequent commissioning steps. In manual simulated-commissioning studies, this trade-off is usually compressed into one expert-designed procedure, or at most a small number of hand-written variants, because each additional point in procedure space requires substantial human design, implementation, debugging, and validation. A central advantage of the autonomous loop is that this labor-intensive comparison can be turned into a search problem: rather than producing a single preferred algorithm, it can produce a set of validated algorithms spanning different operational priorities. We test this setting with a harder version of the ALS-U accumulator-ring beam-capture benchmark. The baseline errors are the same as in Sec. I, but this campaign also enables the discrete catastrophic errors in Table 1: reversed corrector and BPM polarities and dead BPMs. These faults add a diagnostic component to the task, because a successful procedure must not only steer and capture the beam but also infer which parts of the control system are miscalibrated or unusable. The additional faults raise the expert-port baseline from 208 injections in the scalar study to 713 injections on this harder ensemble, so the two baseline values should not be compared directly. This campaign evaluates each candidate on two objectives. The first is the same capture cost used above: the ensemble-mean number of injections required to capture beam. The second is a machine-error correction score, ScorrS_corr, defined in Appendix A, that measures how much of the seeded correctable error has been removed. Lower injection cost and higher ScorrS_corr are both desirable. The scalar merge predicate is replaced by Pareto dominance [54]: a candidate is retained if it is not dominated by any existing retained algorithm. The maintained state becomes the non-dominated set rather than a single best repository head, while the rest of the loop is unchanged. This is a small change to the autonomous harness, but a large change in the kind of commissioning study that becomes practical: the campaign is no longer trying to discover the best procedure under one assumed exchange rate between time and calibration quality, but to populate the trade-off surface itself. Figure 4 shows the front from a single 200-experiment Sonnet campaign, whose two ends are physically distinct strategies. The lowest-cost algorithm corrects only the RF phase and frequency, capturing beam in 679 injections. The highest-quality algorithm keeps the capture procedure intact and appends a stored-beam calibration stage. That stage alternates corrector and BPM-gain calibration against the model response matrix to recover reversed polarities, then adds dead-BPM identification from probe kicks and a launch-error fit. It spends 1371 injections but removes roughly two-thirds of the seeded polarity errors, identifies most dead BPMs, and cuts the injection error. The important result is that autonomous algorithm search changes the scale of what can be explored. A single campaign produced 16 validated commissioning procedures spanning physically meaningful choices between rapid capture and more complete machine calibration. Constructing such a set manually would require repeated expert cycles of designing, implementing, debugging, and comparing separate procedures. The output is therefore not necessarily one globally preferred algorithm, but a set of alternatives from which operators can choose according to the priorities of the broader commissioning sequence. Figure 4: Autonomous exploration of beam-capture cost and correction of seeded machine errors. The campaign minimizes injections to capture while maximizing the correction score ScorrS_corr. Orange points show the 16 non-dominated algorithms retained by the harness. One retained algorithm matches approximately the expert-baseline calibration quality using 679 injections, compared with 713 for the baseline; higher-cost algorithms trade additional injections for improved calibration quality. V Discussion The near-term role of autonomous commissioning search is likely to be primarily offline, spanning lattice design, detailed simulated commissioning, and preparation for first beam. Its value is not only that it can reduce the effort required to rewrite procedures after design changes, but that it can make commissionability a quantity explored throughout the design process. Repeated campaigns could compare alternative procedures, expose recurring failure modes, and identify when apparent commissioning difficulty is rooted in the lattice, diagnostics, actuator layout, or error assumptions rather than in the correction algorithm alone. A practical intermediate mode would keep humans in the evaluation loop. Each agent could propose a new change, while a human expert decides whether to approve, reject, or redirect its evaluation. This would provide a gradual path toward greater autonomy as confidence in the agents and the evaluation harness develops. A further extension is co-design of the accelerator and its commissioning strategy. By enlarging the search space beyond procedure code, related loops could evaluate changes to diagnostics, controls, tolerances, or lattice parameters according to both nominal performance and the existence of robust commissioning solutions. The most direct next test is an end-to-end procedure from first injection to a machine state suitable for user operation. This would require substantially longer evaluations and campaigns, but no fundamental change to the framework. The main challenge would instead be to reduce the high-dimensional space of machine-performance objectives to a small set of quantities, or ideally a scalar objective, that can guide the search efficiently without obscuring important operational trade-offs. Acknowledgements.The author thanks Z. Zhang, D. Ratner, N. Steerenberg, and G. Martino for discussions on autonomous agents for accelerator optimization. The author is particularly grateful to M. Venturini for a careful reading of the manuscript and detailed comments that improved it. This work was supported by the Director of the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231. The code, harness, and campaign configurations needed to reproduce this work are openly available [46]. AI-based tools were used for language editing during manuscript preparation; the author reviewed all content and takes full responsibility. Appendix A Capture procedure, machine-error model, and quality objective Baseline capture procedure. The beam-capture baseline is a pySC reimplementation of the published ALS-U commissioning sequence [8], in six phases: (1) first-turn threading; (2) two-turn stitching; (3) sextupole ramping with orbit correction; (4) an initial tune scan (RF off) to set the working point; (5) RF phase and frequency correction; and (6) a final tune scan against the survival criterion. The partial-credit term in Sec. I scores how far a failing seed advances through these phases, but the six-phase structure is a property of the baseline, not a constraint: the agent may reorganize, merge, or replace these stages, and full capture is certified independently by the harness (Appendix C). Both beam-capture studies use the machine-error model summarized in Table 1, implemented in pySC from the published ALS-U accumulator-ring registration-error model. The continuous errors in the upper block are applied in every simulation seed and represent alignment, calibration, RF, and injection uncertainties. The Pareto campaign additionally includes the discrete faults in the lower block: reversed corrector and BPM polarities and inactive BPMs. These faults make the commissioning problem more diagnostic, because the algorithm must identify or compensate for defective instrumentation and actuators rather than merely correct continuous offsets. The machine-error correction score ScorrS_corr used in Sec. IV measures the residual correctable machine error after beam capture. For each active error category, the harness computes the fraction of the initially seeded error that has been removed and clips the result to [â1,1][-1,1]. A score of 11 denotes complete correction, 0 denotes no net improvement, and â1-1 indicates that the error magnitude has increased by at least its initial value. The reported ScorrS_corr is the mean of these category scores for each seed and is then averaged over the evaluation ensemble. Because ScorrS_corr is computed by the harness from the seeded ground-truth errors, which are withheld from the agent (Appendix C), it is optimized only as a returned scalar reward. The scored categories are BPM offsets and signed gains, corrector calibration, the four transverse injection coordinates (x,xâČ,y,yâČ)(x,x ,y,y ), and RF phase and frequency. Treating BPM gain as a signed quantity allows a polarity reversal to be represented by a gain factor of â1-1. Dead-BPM identification is scored separately as (TPâFP)/D(TP-FP)/D, where TPTP and FPFP are the numbers of correctly and incorrectly identified dead BPMs, respectively, and D is the number actually disabled in that seed. Quadrupole and sextupole calibration errors and the global tune are excluded because their correction belongs to subsequent optics-calibration stages, such as LOCO. The score therefore measures only the machine-registration and injection errors that can reasonably be diagnosed or corrected during the beam-capture procedure. Table 1: Seeded machine-error model for the beam-capture campaigns. Continuous sources (upper block) are applied in all campaigns; the discrete catastrophic errors (lower block) are enabled only for the Pareto campaign. Magnet, BPM, and support entries are per-device Gaussian 1âÏ1Ï values; injection and RF entries are fixed per-seed offsets. Error source Magnitude (1âÏ1Ï) Continuous errors Magnet offset (x,yx,y) 50âÎŒ50\, Magnet roll 200âÎŒ200\, Magnet calibration 0.1%0.1\% Corrector calibration 5%5\% BPM offset (x,yx,y) 500âÎŒ500\, BPM gain/calibration 5%5\% BPM roll 44\,mrad BPM noise (CO/TBT) 1/ 10âÎŒ1\,/\,10\, Cavity frequency 100100\,Hz Cavity voltage 55\,kV Cavity sync. phase 0.150.15\,m (â90ââ\!90 ) Girder offset (x,yx,y) / roll 50âÎŒ50\, / 100âÎŒ/\,100\, Section offset (x,yx,y) 100âÎŒ100\, Circumference 11\,ppm Injection offset (x,yx,y) 600, 500âÎŒ600,\,500\, Injection angle (xâČ,yâČx ,y ) 150, 100âÎŒ150,\,100\, Injection energy 10â310^-3 Injection jitter (pos./angle) 50,5/ 5,250,5\,/\,5,2 (ÎŒ ,ÎŒ ) Injection jitter (energy/phase) 10â4/ 0.1â10^-4\,/\,0.1 Discrete errors, Pareto only Corrector polarity reversal 5%5\% of correctors BPM polarity reversal 5%5\% of BPMs Dead BPMs 5%5\% of BPMs Appendix B Computational cost and reproduction An autoresearch experiment has two costs: the agent cycle and the physics evaluation. The agent cycle includes proposing and coding a change, any single-seed checks performed during development, the independent screen, and the merge decision. The evaluation runs the surviving candidate on the fixed seed ensemble. Which part dominates depends on the commissioning task. For the scalar beam-capture study of Sec. I, a commissioning simulation evaluation takes typically ⌠1â3 min13\,min, so the wall-clock time is dominated by the agent cycle. For the harder Pareto campaign of Sec. IV, the larger injection budget and additional diagnostic faults raise the evaluation to tens of minutes, making the physics evaluation comparable to the agent phase. This difference is the practical reason that the model comparison and scaffold ablation were performed on beam capture: they require many complete campaigns, which would be much more expensive on a full lattice-correction objective. The monetary cost of the agent phase was modest in these studies. Median API cost was approximately $0.50.5 per experiment for Haiku and ⌠$1â313 for the frontier-tier models, corresponding to roughly $60â28060280 for a 100-experiment scalar-capture campaign. Dollar cost varied less than reasoning turns or generated tokens because the largely fixed prompt context was served efficiently by prompt caching. For comparing scaffold conditions, turns and output tokens are therefore more informative than monetary cost. In the ablation of Sec. I, removing both the helper library and the domain-knowledge documents caused the weaker tier to spend an order of magnitude more turns and generated tokens without ever capturing, whereas the stronger tier remained much more concise and constructed working routines in every stripped condition. The practical cost of missing scaffolding is therefore not only lower final performance, but also wasted agent effort. The physics evaluation is CPU-bound. In our implementation, the 50 seeds are distributed over worker processes and evaluated in parallel; a commodity 32-core workstation was sufficient for the production campaigns, with a separate 64-core x86 workstation used as an independent platform replicate. Evaluation time scales with the number of available cores until the seed ensemble is saturated, but also depends on the candidate algorithm: seeds that capture early finish quickly, while failed seeds may consume the full injection budget. When multiple experiments are evaluated concurrently, their seed pools contend for cores and the elapsed time increases accordingly. For reproducible timing, multithreaded numerical libraries should be pinned to one thread per worker process to avoid BLAS oversubscription. All production campaigns used a hard per-experiment wall-clock cap to terminate stuck agent cycles. This prevents rare runaway tool-use loops from dominating a campaign, but it also means that extreme upper-tail runtimes should not be interpreted as completed experiment times. For reproduction, the important quantities are the number of experiments, the number of seeds per experiment, the injection budget per seed, the model-token cost, the host core count, and the protected harness used for scoring and merging. The three tiers correspond to the pinned model snapshots claude-haiku-4-5-20251001, claude-sonnet-4-6, and claude-opus-4-6 (accessed via Amazon Bedrock in region us-east-2), and the agent harness is fixed by claude-agent-sdk 0.1.51. Appendix C Benchmark integrity and failure modes The central risk in autonomous commissioning search is reward hacking, a well-known problem in reinforcement learning and automated optimization: the agent optimizes the implemented benchmark rather than the intended scientific objective [55, 56, 57]. Any discrepancy between what the agent is intended to observe or control and what the benchmark actually permits can therefore be exploited by the search. Several such mismatches emerged during development. In an early harness, the nominal inject-and-read operation was priced, but other interface operations that would consume equivalent machine time on a real accelerator were not. The operator interface also exposed simulator quantities unavailable in a real control room, including true alignment errors and analytic lattice information. In another implementation, the candidate could effectively certify its own success, allowing inadequate one-turn beams to count as valid 500-turn captures. These were not sandbox escapes, but valid optimizations of underspecified benchmarks. The final harness addresses these failures structurally. All machine-equivalent actions are priced, simulator ground truth is excluded from the operator interface, and capture is certified only by the harness using a fixed minimum particle count. Protected benchmark components are read-only, and the observed exploits are covered by regression tests. The reviewer and deterministic pattern scan provide additional screening, but cannot substitute for a correctly specified experimental boundary. Autonomous search can also fail in the opposite direction. The scaffold ablation shows that a weak or under-provisioned agent may expend far more reasoning and generated code while making no progress. Agent activity is therefore not evidence of useful search. Both failure modes reinforce the same requirement: progress must be judged by a fixed, harness-owned metric whose connection to the intended scientific task has been independently validated. References Tavares et al. [2014] P. F. Tavares, S. C. Leemann, M. Sjöström, and Ă . Andersson, J. Synchrotron Radiat. 21, 862 (2014). Raimondi et al. [2023] P. Raimondi, C. Benabderrahmane, P. Berkvens, J. C. Biasci, P. Borowiec, J.-F. Bouteille, T. Brochard, N. B. Brookes, N. Carmignani, L. R. Carver, J.-M. Chaize, J. Chavanne, S. Checchia, Y. Chushkin, F. Cianciosi, M. Di Michiel, R. Dimper, A. DâElia, D. Einfeld, F. Ewald, L. Farvacque, L. Goirand, L. Hardy, J. Jacob, L. Jolly, M. Krisch, G. Le Bec, I. Leconte, S. M. Liuzzo, C. Maccarrone, T. Marchial, D. Martin, M. Mezouar, C. Nevo, T. Perron, E. Plouviez, H. Reichert, P. Renaud, J.-L. Revol, B. Roche, K.-B. Scheidt, V. Serriere, F. Sette, J. Susini, L. Torino, R. Versteegen, S. White, and F. Zontone, Commun. Phys. 6, 82 (2023). Sajaev et al. [2025] V. Sajaev, M. Borland, J. Calvey, J. Dooling, L. Emery, K. Harkay, N. Kuklev, O. Mohsen, Y. Sun, N. Arnold, T. Berenc, A. Brill, H. Bui, J. Carwardine, W. Cheng, T. Fors, M. Kelly, R. Lindberg, G. Shen, M. Smith, F. Rafael, H. Shang, R. Soliday, U. Wienands, and B. Yang, in Proc. NAPACâ25 (JACoW Publishing, Geneva, Switzerland, 2025) p. 1â6. Steier et al. [2025] C. Steier, J. Bohon, S. Borra, K. Chow, C. Espino-Devine, E. DiMasi, G. Ganetis, T. Hellert, M. Johansson, J. Joseph, J.-Y. Jung, R. Leftwich-Vann, D. Leitner, H. Lee, A. Lodge, M. Lerche, T. Luo, R. Miller, D. Nett, B. Nicquevert, O. Omolayo, D. Paulic, A. Ratti, D. Robin, C. Sun, C. Swenson, S. Trovati, M. Venturini, W. Waldron, E. Wallen, and D. Wang, J. Phys. Conf. Ser. 3010, 012046 (2025). Schroer et al. [2018] C. G. Schroer, I. Agapov, W. Brefeld, R. Brinkmann, Y.-C. Chae, H.-C. Chao, M. Eriksson, J. Keil, X. Nuel GavaldĂ , R. Röhlsberger, O. H. Seeck, M. Sprung, M. Tischer, R. Wanzenberg, and E. Weckert, J. Synchrotron Radiat. 25, 1277 (2018). Borland et al. [2014] M. Borland, G. Decker, L. Emery, V. Sajaev, Y. Sun, and A. Xiao, J. Synchrotron Radiat. 21, 912 (2014). Sajaev [2019] V. Sajaev, Phys. Rev. Accel. Beams 22, 040102 (2019). Hellert et al. [2019] T. Hellert, P. Amstutz, C. Steier, and M. Venturini, Phys. Rev. Accel. Beams 22, 100702 (2019). Apollonio et al. [2021] M. Apollonio, R. Fielder, H. Ghasem, and I. Martin, in Proc. IPACâ21 (JACoW Publishing, Geneva, Switzerland, 2021) p. 265â268. Amorim et al. [2021] D. Amorim, A. Loulergue, L. S. Nadolski, and R. Nagaoka, in Proc. IPACâ21 (JACoW Publishing, Geneva, Switzerland, 2021) p. 171â174. Blanco-GarcĂa et al. [2022] O. Blanco-GarcĂa, D. Amorim, M. Deniaud, A. Loulergue, L. Nadolski, and R. Nagaoka, in Proc. IPACâ22 (JACoW Publishing, Geneva, Switzerland, 2022) p. 433â436. Hellert et al. [2022] T. Hellert, I. Agapov, S. Antipov, R. Bartolini, R. Brinkmann, Y.-C. Chae, D. Einfeld, M. Jebramcik, and J. Keil, in Proc. IPACâ22 (JACoW Publishing, Geneva, Switzerland, 2022) p. 1442â1444. Habet et al. [2025] S. Habet, A. Loulergue, L. Nadolski, P. Brunelle, and S. Ducourtieux, in Proc. IPACâ25 (JACoW Publishing, Geneva, Switzerland, 2025) p. 698â701. Chen et al. [2024] K. Chen, Z. Wang, G. Wang, T. He, Z. Wang, D. He, M. Hosaka, and W. Xu, in Proc. IPACâ24 (JACoW Publishing, Geneva, Switzerland, 2024) p. 2043â2045. Benedetti et al. [2025] G. Benedetti, F. Perez, M. CarlĂ , O. Blanco-GarcĂa, and Z. MartĂ, in Proc. IPACâ25 (JACoW Publishing, Geneva, Switzerland, 2025) p. 2222â2224. Wu et al. [2024] X. Wu, L. Tan, S. Xuan, S.-Q. Tian, X. Liu, and Y. Gong, in Proc. IPACâ24 (JACoW Publishing, Geneva, Switzerland, 2024) p. 1342â1345. Chao and Martin [2024] H.-C. Chao and I. Martin, in Proc. IPACâ24 (JACoW Publishing, Geneva, Switzerland, 2024) p. 1262â1265. Agapov et al. [2024] I. Agapov, S. Antipov, R. Bartolini, R. Brinkmann, Y. Chae, E. C. Cortes-Garcia, D. Einfeld, and T. Hellert, arXiv preprint (2024), arXiv:2408.07995 . MartĂ et al. [2023] Z. MartĂ, G. Benedetti, M. CarlĂ , and U. Iriso, in Proc. IPACâ23 (JACoW Publishing, Geneva, Switzerland, 2023) p. 3104â3107. Ji et al. [2023] D. Ji, B. Wang, and X. Cui, in Proc. IPACâ23 (JACoW Publishing, Geneva, Switzerland, 2023) p. 492â495. Wang et al. [2021] B. Wang, Z. Duan, D. Ji, Y. Jiao, and Y. Zhao, in Proc. IPACâ21 (JACoW Publishing, Geneva, Switzerland, 2021) p. 1349â1351. Hellert et al. [2023] T. Hellert, C. Steier, and J. Keil, in Proc. IPACâ23 (2023). Huang and Safranek [2015] X. Huang and J. Safranek, Phys. Rev. ST Accel. Beams 18, 084001 (2015), arXiv:1502.07799 . Duris et al. [2020] J. Duris, D. Kennedy, A. Hanuka, J. Shtalenkova, A. Edelen, P. Baxevanis, A. Egger, T. Cope, M. McIntire, S. Ermon, and D. Ratner, Phys. Rev. Lett. 124, 124801 (2020), arXiv:1909.05963 . Roussel et al. [2021] R. Roussel, A. Hanuka, and A. Edelen, Phys. Rev. Accel. Beams 24, 062801 (2021), arXiv:2010.09824 . Xu et al. [2023] C. Xu, T. Boltz, A. Mochihashi, A. S. Garcia, M. Schuh, and A.-S. MĂŒller, Phys. Rev. Accel. Beams 26, 034601 (2023), arXiv:2211.09504 . Kaiser et al. [2024a] J. Kaiser, C. Xu, A. Eichler, and A. S. Garcia, Phys. Rev. Accel. Beams 27, 054601 (2024a), arXiv:2401.05815 . Kaiser et al. [2024b] J. Kaiser, C. Xu, A. Eichler, A. S. Garcia, O. Stein, E. BrĂŒndermann, W. Kuropka, H. Dinter, F. Mayet, T. Vinatier, F. Burkart, and H. Schlarb, Sci. Rep. 14, 15733 (2024b), arXiv:2306.03739 . Roussel et al. [2024] R. Roussel, A. L. Edelen, T. Boltz, D. Kennedy, Z. Zhang, F. Ji, X. Huang, D. Ratner, A. S. Garcia, C. Xu, J. Kaiser, A. F. Pousa, A. Eichler, J. O. LĂŒbsen, N. M. Isenberg, Y. Gao, N. Kuklev, J. Martinez, B. Mustapha, V. Kain, C. Mayes, W. Lin, S. M. Liuzzo, J. S. John, M. J. V. Streeter, R. Lehe, and W. Neiswanger, Phys. Rev. Accel. Beams 27, 084801 (2024), arXiv:2312.05667 . Kaiser et al. [2025] J. Kaiser, A. Lauscher, and A. Eichler, Sci. Adv. 11, eadr4173 (2025), arXiv:2405.08888 . Mayet [2024] F. Mayet, GAIA: A General AI Assistant for Intelligent Accelerator Operations (2024), arXiv:2405.01359 [cs.CL] . Sulc et al. [2024] A. Sulc, T. Hellert, R. Kammering, H. Hoschouer, and J. S. John, in Machine Learning and the Physical Sciences Workshop, NeurIPS 2024 (2024) arXiv:2409.06336 [physics.acc-ph] . Hellert et al. [2026] T. Hellert, D. Bertwistle, S. C. Leemann, A. Sulc, and M. Venturini, Phys. Rev. Research 8, L012017 (2026), arXiv:2509.17255 . Hellert et al. [2025] T. Hellert, J. Montenegro, and A. Sulc, Osprey: Production-Ready Agentic AI for Safety-Critical Control Systems (2025), arXiv:2508.15066 [cs.MA] . Karpathy [2026] A. Karpathy, autoresearch: AI agents running research on single-GPU nanochat training automatically, https://github.com/karpathy/autoresearch (2026), open-source greedy experiment loop for automated ML research. Lu et al. [2024] C. Lu, C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha, The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery (2024), arXiv:2408.06292 [cs.AI] . Boiko et al. [2023] D. A. Boiko, R. MacKnight, B. Kline, and G. Gomes, Nature 624, 570 (2023). Bran et al. [2024] A. M. Bran, S. Cox, O. Schilter, C. Baldassari, A. D. White, and P. Schwaller, Nat. Mach. Intell. 6, 525 (2024), arXiv:2304.05376 . Romera-Paredes et al. [2024] B. Romera-Paredes, M. Barekatain, A. Novikov, M. Balog, M. P. Kumar, E. Dupont, F. J. R. Ruiz, J. S. Ellenberg, P. Wang, O. Fawzi, P. Kohli, and A. Fawzi, Nature 625, 468 (2024). Ma et al. [2024] Y. J. Ma, W. Liang, G. Wang, D.-A. Huang, O. Bastani, D. Jayaraman, Y. Zhu, L. Fan, and A. Anandkumar, in Int. Conf. on Learning Representations (ICLR) (2024) arXiv:2310.12931 [cs.LG] . MacLeod et al. [2020] B. P. MacLeod, F. G. L. Parlane, T. D. Morrissey, F. HĂ€se, L. M. Roch, K. E. Dettelbach, R. Moreira, L. P. E. Yunker, M. B. Rooney, J. R. Deeth, V. Lai, G. J. Ng, H. Situ, R. H. Zhang, M. S. Elliott, T. H. Haley, D. J. Dvorak, A. Aspuru-Guzik, J. E. Hein, and C. P. Berlinguette, Sci. Adv. 6, eaaz8867 (2020). Szymanski et al. [2023] N. J. Szymanski, B. Rendy, Y. Fei, R. E. Kumar, T. He, D. Milsted, et al., Nature 624, 86 (2023). Abolhasani and Kumacheva [2023] M. Abolhasani and E. Kumacheva, Nat. Synth. 2, 483 (2023). Stach et al. [2021] E. Stach, B. DeCost, A. G. Kusne, J. Hattrick-Simpers, K. A. Brown, K. G. Reyes, et al., Matter 4, 2702 (2021). Steier et al. [2019] C. Steier, P. Amstutz, K. Baptiste, P. Bong, E. Buice, P. Casey, K. Chow, R. Donahue, M. Ehrlichman, J. Harkins, T. Hellert, M. Johnson, J.-Y. Jung, S. Leemann, R. Leftwich-Vann, D. Leitner, T. Luo, O. Omolayo, J. Osborn, G. Penn, G. Portmann, D. Robin, F. Sannibale, S. D. Santis, C. Sun, C. Swenson, M. Venturini, S. Virostek, W. Waldron, and E. WallĂ©n, in Proc. IPACâ19 (JACoW Publishing, Geneva, Switzerland, 2019) p. 1639â1642. Hellert [2026] T. Hellert, autoresearch: autonomous discovery of accelerator commissioning algorithms, https://github.com/als-apg/autoresearch-commissioning (2026), research loop, simulated-commissioning harness, and campaign configurations. Malina et al. [2023] L. Malina, I. Agapov, J. Keil, E. S. H. Musa, B. Veglia, N. Carmignani, L. Carver, L. Hoummi, S. M. Liuzzo, T. Perron, S. White, and T. Hellert, in Proc. ICALEPCSâ23 (JACoW Publishing, Geneva, Switzerland, 2023) p. 1637â1642. Liuzzo et al. [2024] S. M. Liuzzo, N. Carmignani, L. R. Carver, L. Hoummi, T. Perron, S. White, I. Agapov, M. Boese, T. Hellert, J. Keil, L. Malina, E. Musa, and B. Veglia, J. Phys. Conf. Ser. 2687, 032001 (2024). Di Mitri [2023] S. Di Mitri, Fundamentals of Particle Accelerator Physics, Graduate Texts in Physics (Springer Cham, 2023). Wiedemann [2015] H. Wiedemann, Particle Accelerator Physics, 4th ed., Graduate Texts in Physics (Springer Cham, 2015). Anthropic [2025] Anthropic, System card: Claude Haiku 4.5, https://assets.anthropic.com/m/99128d009bdcb/original/Claude-Haiku-4-5-System-Card.pdf (2025). Anthropic [2026a] Anthropic, System card: Claude Sonnet 4.6, https://w.anthropic.com/research/claude-sonnet-4-6 (2026a). Anthropic [2026b] Anthropic, System card: Claude Opus 4.6, https://w.anthropic.com/news/claude-opus-4-6 (2026b). Deb [2001] K. Deb, Multi-Objective Optimization Using Evolutionary Algorithms (John Wiley & Sons, Chichester, 2001). Amodei et al. [2016] D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. ManĂ©, arXiv preprint arXiv:1606.06565 (2016), arXiv:1606.06565 [cs.AI] . Leike et al. [2017] J. Leike, M. Martic, V. Krakovna, P. A. Ortega, T. Everitt, A. Lefrancq, L. Orseau, and S. Legg, arXiv preprint arXiv:1711.09883 (2017), arXiv:1711.09883 [cs.LG] . Skalse et al. [2022] J. Skalse, N. H. R. Howe, D. Krasheninnikov, and D. Krueger, Advances in Neural Information Processing Systems 35, 9460 (2022).