Paper deep dive
How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models
Aleksandra Kalisz, Jack Simons, Krisztina Sinkovics, Noam Ghenassia, Shikha Surana, Henry Moss, Paul Duckworth
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Foundation models for protein structure prediction remain unreliable on certain targets. External oracles can flag and correct these failures, but biological oracles are expensive, making oracle budget a critical constraint. Existing guidance methods, such as FK-steering, DPO, and Best K-of-N sampling, differ in how they spend this budget, yet no systematic comparison exists to guide method selection. To bridge this gap, we benchmark these methods alongside the recently proposed Optimisation Over Outputs (O3), which applies off-the-shelf optimisers within a generative model's latent subspace. We extend the usage of O3 to protein structure prediction models. Overall, our work provides the first practical reference for oracle budget-aware guidance. Our evaluation on two protein targets, calmodulin (1CLL) and E. coli aspartate transcarbamoylase (9EEH), reveals that no single method consistently dominates across all budgets and oracles. Specifically, O3 proves most effective at low oracle budgets, while FK-steering and DPO demonstrate improved performance as the budget increases. We distil these findings into actionable recommendations for practitioners operating under real-world oracle-budget constraints.
Tags
Links
- Source: https://arxiv.org/abs/2608.12192v1
- Canonical: https://arxiv.org/abs/2608.12192v1
Trouble viewing inline? Open PDF directly ā
Full Text
53,169 characters extracted from source content.
Expand or collapse full text
How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models Aleksandra Kalisz 1 2 Jack Simons 1 Krisztina Sinkovics 1 Noam Ghenassia 1 Shikha Surana 1 Henry Moss 3 Paul Duckworth 1 Abstract Foundation models for protein structure predic- tion remain unreliable on certain targets. External oracles can flag and correct these failures, but bio- logical oracles are expensive, making oracle bud- get a critical constraint. Existing guidance meth- ods, such as FK-steering, DPO, and Best K-of-N sampling, differ in how they spend this budget, yet no systematic comparison exists to guide method selection. To bridge this gap, we benchmark these methods alongside the recently proposed Optimi- sation Over Outputs (O3), which applies off-the- shelf optimisers within a generative modelās latent subspace. We extend the usage of O3 to protein structure prediction models. Overall, our work provides the first practical reference for oracle budget-aware guidance. Our evaluation on two protein targets, calmodulin (1CLL) and E. coli as- partate transcarbamoylase (9EEH), reveals that no single method consistently dominates across all budgets and oracles. Specifically, O3 proves most effective at low oracle budgets, while FK-steering and DPO demonstrate improved performance as the budget increases. We distil these findings into actionable recommendations for practition- ers operating under real-world oracle-budget con- straints. 1. Introduction Foundation models for structure prediction such as Al- phaFold3 (Abramson et al., 2024), Chai-1 (Chai Discovery et al., 2024), and Boltz-2 (Passaro et al., 2025) can accu- rately predict structures from sequence information alone. Downstream applications, including drug discovery and an- 1 InstaDeepLtd 2 UniversityofOxford 3 Lancaster University.Correspondenceto:PaulDuckworth <p.duckworth@instadeep.com>. Proceedings of the ICML 2026 Workshop on Structured Proba- bilistic Inference & Generative Modeling (SPIGM), Seoul, South Korea. 2026. Copyright 2026 by the author(s). tibody therapeutic design, rely on these predictions being correct. In practice these predictions are not always reliable: a model may collapse onto a single conformation when sev- eral are functionally relevant (Wayment-Steele et al., 2023), produce physically implausible geometries (Buttenschoen et al., 2024), or assemble large biomolecular complexes in- correctly (Yin & Pierce, 2023). A typical procedure to flag unreliable predictions is via the usage of external oracles, which assign a score to a given prediction, e.g. a molecular dynamics simulation of conformational stability or binding energy (Hollingsworth & Dror, 2018). A natural research direction is whether such oracles can be leveraged at train- ing or inference time to return higher scoring structures, a process referred to as guidance. There are two dominant classes of guidance approaches, inference-time steering and fine-tuning; however their relative effectiveness is poorly understood, a gap in the literature we seek to fill. Inference-time methods, such as Feynman-Kac steering (FK- steering) (Singhal et al., 2025), sample multiple interacting particles through the generative process and resample at intermediate steps using oracle scores. Fine-tuning methods, such as Direct Preference Optimisation (DPO) (Rafailov et al., 2023), instead update model parameters using oracle- scored samples. An even simpler approach to guidance with an oracle is Best K-of-N sampling, whereby we drawN samples and return theKwith the highest oracle scores. All three approaches improve outputs given enough oracle evaluations, but biological oracles are usually expensive: a single molecular dynamics simulation can take days of GPU compute (Hollingsworth & Dror, 2018), and wet-lab assays are slower still. As well as benchmarking the methods listed above, we also provide the first application of the newly proposed Optimisa- tion Over Outputs (O3) framework of Willis et al. (2025) to a protein structure prediction model. O3 uses a small set of example model generations to build a low-dimensional sub- space of the latent space over which standard optimisers can be directly applied. More details of the precise methodology are given in Willis et al. (2025). We analyse the relative performances of O3, FK-steering, DPO, and Best K-of-N sampling applied to Boltz-2 (Pas- 1 arXiv:2608.12192v1 [cs.AI] 12 Aug 2026 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models Figure 1. Bayesian Optimisation via O3 in a Boltz-2.d = 3surrogate subspace created from three Calmodulin seed samples (top row), showing the 2DUspace (bottom row), the GP training data (white dots) and mean predictive posterior overU, the max acquisition point (red cross) for two acquisition rounds, and a 25x25 exhaustive ground-truth grid directly evaluating the oracle (bottom right). saro et al., 2025) across a range of budgets, with the goal to improve the quality of the predicted protein structures. We find that our new application of O3 to Boltz-2 dom- inates at low-to-mid budgets that we evaluate on, while FK-steering and DPO perform better with larger budgets. This qualitative trade-off is demonstrated in our experimen- tal study reported in Section 5. The primary goal of this paper is straightforward: provide the first practical reference for which guidance strategies are effective under different oracle evaluation budgets. 2. Related Work Protein structure prediction. Recent foundation mod- els for protein structure prediction, including AlphaFold3 (Abramson et al., 2024), Chai-1 (Chai Discovery et al., 2024), and Boltz-2 (Passaro et al., 2025), share a common ar- chitecture: an attention-based trunk, typically a Pairformer, that processes the input sequence into single and pair repre- sentations, which is then used as conditioning information to a diffusion block over atomic coordinates. Throughout our experiments we use the open-source model Boltz-2. Latent space optimisation. Black-box optimisation de- scribes optimisation of an objective function which is explic- itly unknown a priori but can be queried through function evaluations. These evaluations can be expensive, noisy, and derivative-free. A popular approach to such problems is Bayesian Optimisation (BO) (see e.g. Garnett (2023)). However, standard BO approaches begin to fail in high- dimensional or highly structured spaces. To tackle this shortcoming, several Latent space Bayesian Optimisation (LSBO) approaches have been proposed which instead per- form BO in the latent space associated with a generative model. Examples of latent space Bayesian optimisation in- clude VAE-BO (G Ģ omez-Bombarelli et al., 2018), weighted retraining (Tripp et al., 2020), LOL-BO (Maus et al., 2022), and COWBOYS (Moss et al., 2025). Generative models often have latent spaces that are too large to apply these methods directly, fortunately, recent work from Bodin et al. (2024) and Willis et al. (2025) proposes an alternative ap- proach that extracts low-dimensional subspaces from gener- ative models, over which standard optimisers can be applied. Inference-time guidance.Inference-time guidance biases the sampling trajectory of a diffusion model toward a tar- get distribution without modifying its weights. Classifier guidance (Dhariwal & Nichol, 2021) perturbs the sampling trajectory by adding the gradient of an externally-trained noise-aware classifier to the denoising step, while classifier- free guidance (Ho & Salimans, 2022) trains a single model on both conditional and unconditional inputs and combines their score estimates at inference time. We choose Feynman- Kac steering (Singhal et al., 2025) as a baseline for this work since it does not require propagating gradients from the ora- cle, and does not require an auxiliary network. Reward-based fine-tuning. DDPO (Black et al., 2023) casts sampling as a multi-step Markov decision process and applies policy-gradient updates with the oracle as re- ward; DRaFT (Clark et al., 2023) backpropagates the oracle gradient through the sampling chain when the oracle is dif- ferentiable. DPO (Rafailov et al., 2023) sidesteps explicit reward modelling by fine-tuning on pairs of oracle-scored samples to increase the relative likelihood of the preferred one, and Diffusion-DPO (Wallace et al., 2023) adapts this objective to the diffusion likelihood. We choose DPO as our fine-tuning baseline because it requires only sample pairs and oracle scores. 2 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models Typically, biological oracles are non-differentiable, meaning that gradient information is not available, making many common guidance approaches unsuitable. 3. Guidance Methods We denote our frozen generative modelg : Z ā X, that maps random noisez ā¼ p 0 from within a latent spacez āZ to structuresx ā¼ p 1 in data spacex ā X. An oracler : X ā R scores each generated output. In our experimentsg is the protein structure prediction model Boltz-2 (Passaro et al., 2025) withZ =X = R 3n for ann-atom system and p 0 =N(0,I). The goal of guidance is to use this model to generate a batch ofKcandidate outputs whose scores under rshould be as high as possible, under the constraint that we only makeNoracle queries. We will now introduce three guidance methods, explaining the key modifications and design choices that were required to apply them to protein structure prediction models. 3.1. Bayesian Optimisation via O3 We now describe how the O3 framework of Willis et al. (2025) is used to guide the generative process of Boltz-2. The algorithm relies on two key components, each necessi- tating a number of oracle queries: (i) the construction of a suitable low-dimensional subspaceUwithin which genera- tion is constrained (the firstMqueries), and (i) a resource- efficient optimization procedure to effectively explore this subspace (the remaining N ā M queries). (i) ConstructingU. Given a total budget ofN, we first sam- pleM < Ninitial structures fromgand score them with r. Exploiting the determinism ofg, we collect thedlatents corresponding to the d highest scoring structures 1 as seeds, Z = [z 1 ,...,z d ] ⤠. Hereddetermines the dimension of the resulting optimisation space,dā1. We then follow the guide- lines of O3 (Willis et al., 2025), and constructU ā [0, 1] dā1 via a surrogate chartĻ :U āZthat decodes eachuāU to a corresponding latentzthrough two maps (see Algo- rithm 1). The first, a KnotheāRosenblatt transform (Knothe, 1957; Rosenblatt, 1952)Ļ : U ā S dā1 + , mapsuto a non- negative weight vectorwon the simplex. The second, a LOL projection (Bodin et al., 2024)ā(w,Z) := w ⤠Z, projects these weight vectors back onto the support of our isotropic Gaussian priorp 0 . This is effective because the original seeds themselves were drawn fromp 0 and the weights are on the simplex. Oncedseeds are chosen, the subspace is fixed, and never re-built. Optimisation occurs withinU, where Uhas dimensionalitydā 1and crucially is independent of the dimensionality ofZ, thus opening up opportunities 1 O3 requires a deterministic generative model, and so we con- verted the stochastic Boltz-2 generation process to its equivalent probability-flow formulation (Karras et al., 2022) and disable stochastic SE(3) augmentation. for efficient optimization. In Appendix Fig. A.2 we select two seedsd = 2, and linearly interpolate between them us- ing 5 intermediate weights, displaying their corresponding generations g(Ā·)ā¼ p 1 . Algorithm 1 Mapping fromU to protein structures 1: u = [u 1 ,...,u dā1 ] ⤠// selected by optimiser 2: w ā Ļ(u)// KnotheāRosenblatt transform 3: z ā ā(w,Z) := w ⤠Z// LOL projection 4: xā g(z)// decoded structure, ready for scoring (i) Bayesian optimization overU. Given a subspaceU, we use Bayesian optimisation (Garnett, 2023, BO) ā the gold-standard for low-dimensional black-box optimisation ā to efficiently utilise our remainingN ā Moracle budget. We begin by fitting a Gaussian process (GP) (Williams & Rasmussen, 2006) surrogate model ond + 2initial points, wheredof them are the seed latents projected ontoUand the remaining2are sampled randomly from withinU. Each subsequent round uses Log Expected Improvement (Ament et al., 2023) to drive data acquisition, each time updating the GP model with the newly acquired point. In order to provide additional training data, we also project thedseeds into the Usubspace using the surrogate chartĻ, and include those as additional training data points (as demonstrated in Fig. 1). We use an RBF kernel, constant mean function, single task GP, and Monte Carlo acquisition function sampler from BOTorch (Balandat et al., 2020). 3.2. FK-Steering Feynman-Kac steering (Singhal et al., 2025) is an inference- time method which uses the signal from an oracle to alter the sampling trajectory towards high-reward regions of the target distribution via Sequential Monte Carlo. The proce- dure generatesMinteracting stochastic processes called particles. Particles are resampled proportionally to weights wcalculated for each particleiusing the oracle rewardr evaluated on the denoised predictions fromx i t (for denoising timestep t): w(x i t ) = g(x i t | x i t prev ) g proposal (x i t | x i t prev ) Ī» exp(r(x i t )ā r(x i t prev )) whereĪ»controls the scale of the reward signal, and g proposal (x t | x t prev )is the proposal transition kernel. In our experiments, since most high-quality biological oracles are non-differentiable, we use the original BOLTZ-2 tran- sition kernelg(x i t | x i t prev ) , demonstrating the gradient-free nature of FK-steering. The rewards are calculated only at the resampling steps and there are several ways to combine them across the reward trajectory (difference, max, average); 3 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models in this work we use the reward difference, which favours particles with increasing rewards. To restrict to an oracle budget ofNand return a batch of Kdistinct samples, we requireM ā„ Kparticles, each of which has access toN/Moracle calls. Practically, this ratio determines the number of resampling steps in our algorithm. Boltz-2 employs a particular noising schedule, which stops injecting noise in the last quarter of the trajectory in order to avoid atomic jitter. Since FK-steering relies on the diversity of particles, we only resample at intermediate denoising steps where Boltz-2 adds positive noise. We call the oracle once at the start of the denoising process, once at the end, and spread the remainder of theN/Mā 2budget across the first 3/4 of the denoising trajectory. 3.3. DPO Unlike our other approaches, Direct Preference Optimisa- tion (DPO) (Rafailov et al., 2023) fine-tunes and directly modifies the generative model parameters based on oracle- scored sample pairs, with the aim of increasing the likeli- hood of generating high-scoring structures. DPO does not require a differentiable oracle (no explicit reward modeling is required, rather just pairs of structures and their relative oracle rankings) making it amenable to common biologi- cal oracles. However, we find that a large oracle budget is required to support effective fine-tuning at foundation model-scale. Diffusion-DPO training objective. DPO updates the modelās weightsĪøin a direction that maximises the mar- gin between the implicit rewards of preferred and rejected samples under a Kullback-Leibler regularised policy, as measured by the DPO loss: L DPO (Īø) =āE (x w ,x l ) h logĻ Ī² log p Īø (x w ) p ref (x w ) ā log p Īø (x l ) p ref (x l ) i , whereβ > 0controls the strength of the KL penalty against moving the new modelp Īø too far from the pre-trained refer- ence modelp ref (initialized at the pre-trained Boltz-2 check- point), andĻrepresents the sigmoid function. Preference pairs always contain a preferred structurex w and a rejected structurex l , each sampled from ourNscored structures and lie above and below the median oracle score respec- tively. Obtaining log-likelihoods from diffusion models is prohibitively expensive (Song et al., 2020), we follow Wal- lace et al. (2023), and approximate the log-likelihood ratio in the DPO loss using differences in the individual train- ing losses (L DSM ) of the reference and fine-tuned model as log(p Īø (x)/p ref (x))āL ref DSM (x)āL Īø DSM (x). Online and offline DPO. We evaluate two variants of DPO that differ in how oracle queries are distributed across training. The offline variant draws allNstructures in a single batch, forms preference pairs once, and trains to convergence on the resulting fixed dataset. The online variant interleaves sampling and training, at each of the Eepochs: we generateN/Estructures from the current model, score the structures with the oracle, form preference pairs, update parameters for one epoch; until a total budget ofN = (N/E)ĆEqueries has been utilised. Critically, the reference modelp ref is reset to the current policy at the start of each epoch, thus softening the KL regularisation. How- ever, this resetting weakens the influence of the pre-trained model ā a tradeoff we investigate in Section 5.4. 4. Experimental Setup As our main experiment, we benchmark the four methods by comparing the predicted structures for the calmodulin protein from its amino acid sequence (of length 144). As ground truth, we compare against the 1CLL structure (Chat- topadhyaya et al., 1992) consisting of 1184 atom coordinates from the Protein Data Bank (PDB) (Berman et al., 2000). Additionally, we benchmark on E. coli aspartate transcar- bamoylase (PDB: 9EEH), with results in Appendix B. For all experiments, the base model we use,g, is Boltz-2 (Pas- saro et al., 2025) and for 1CLLZ = R 3552 , and 9EEH Z = R 21,696 . All methods are permittedNoracle calls in total and return a batch ofKcandidate structures. Next, we describe the TM-score oracle used in the main experiment. Figure 3. TM-score on 1CLL. Left: the 1CLL crystal structure (blue). Centre: a vanilla Boltz-2 sample (orange) overlaid on the ground truth (grey) at TM-score0.45. Right: a guided sample at TM-score0.80. Structure pairs with TM-score above0.5are considered to share the same fold. Oracle. The reward,r, is the TM-score (Zhang & Skol- nick, 2004) of a generated structurexagainst the ground- truth 1CLL backbonex GT . To comparexwithx GT , we find the optimal rigid alignment, namely the rotation and transla- tion ofxthat minimise the distances to the corresponding residues of x GT . The TM-score is then r(x) = max 1 n n ali X i=1 1 1 + (d i (x,x GT )/C) 2 , wherenis the number of residues in the ground truth struc- ture (144 for 1CLL);n ali ⤠nis the number of residues fromx GT matched to a residue inxby the optimal align- ment, with residues that have no nearby counterpart ex- cluded from the sum;d i is the distance between thei-th 4 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models N=20N=50N=100N=200N=500N=1000 0.4 0.5 0.6 0.7 0.8 0.9 TM Score (Mean of K) Method FK-steering DPO Best K-of-N O3 Figure 2. Mean-of-KTM-score across oracle budgets. Comparison of FK-steering, DPO (in the online setting), Best K-of-N, and O3 on 1CLL across the six(N,K)configurations of Section 4. Bars are the mean over5random seeds; error bars are±1std. Higher is better. O3 method dominates at budgets evaluated, while FK-steering and DPO improve steadily with N . aligned residue pair 2 ;C = 1.24 3 ā nā 15ā 1.8is a length- dependent normalisation chosen so that random structure pairs yield approximately constant TM-score; and the maxi- mum is over the rigid alignments described above. Scores lie in(0, 1], with values above0.5indicating the same fold and below0.2indicating random pairs; Fig. 3 shows exam- ples on 1CLL. We use TM-score because it is relatively fast to compute, thus enabling our extensive experimentation. However, we emphasise that our analysis exposes trade-offs between the studied methodologies in settings where in- deed the oracle is expensive or non-differentiable, a setting common in biology. For PDB: 9EEH experiments, we use the MolProbity oracle, described in Appendix B.1. Budget settings. We evaluate each method at six(N,K) configurations spanning low to moderate oracle-call bud- gets:(20, 2),(50, 5),(100, 10),(200, 20),(500, 50), and (1000, 100). We are specifically interested inK > 1set- tings throughout because downstream stages of biological design pipelines, such as wet-lab assay screening, often test batches of candidates rather than committing to a single proposal. Evaluation metrics. At each(N,K)configuration, we report two performance metrics: the max-of-Kscore, the best single structure returned, and the mean-of-Kscore, the average score across the returned batch. The former measures single-shot design quality; the latter measures batched-screening quality. Each plot in Section 5 reports one of the two, with the complementary results presented in our appendices. All experiments are repeated over5random seeds unless stated otherwise. 2 it is a distance between the C α atoms of the two residues. 5. Results Section 5.1 discusses the comparison of the four methods on PDB: 1CLL target, across six(N,K)configurations introduced in Section 4. Sections 5.2 to 5.4 elaborate on method-specific findings and ablations for O3, FK-steering, and DPO, respectively. PDB: 9EEH results are provided in Appendix B. 5.1. Model Comparison Across Budgets Fig. 2 reports the mean-of-KTM-score for our four meth- ods (and Appendix Section A.1 presents the max-of-Kre- sults). O3 outperforms all baselines at low-mid budgets (N ⤠1000) with the mean value plateauing atā¼ 0.81. Best K-of-N is roughly flat atā¼ 0.60across all budgets. FK- steering and DPO improve with increasedN, but neither are competitive with simple Best K-of-N at low budgets (N ⤠100), with FK-steering achieving onlyā¼ 0.55at N = 20. FromN = 200toN = 1000, FK-steering per- forms better, reachingā¼ 0.73atN = 1000; DPO improves withNbut does not match FK-steering or O3 within our range of settings, reachingā¼ 0.71atN = 1000. O3 is the only method that meaningfully improves on the Best K-of-N baseline at low oracle budgets (N ⤠100), supporting the case for example-defined latent-subspace optimisation when oracle calls are scarce. Appendix B.2 reports a model com- parison across budgets on a different protein (PDB:9EEH) and MolProbity oracle (Appendix B.1). 5.2. O3 Fig. 4 isolates the optimisation step in O3, comparing Bayesian optimisation in the subspace against uniform ran- dom sampling in the same subspace forN = 100. We create the subspaceUby selecting to bestdseeds fromM. Random sampling inUis shown to improve over the Best 5 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models K-of-N sampling baseline, demonstrating the value of the low-dimensional subspace. Furthermore, the superior per- formance of O3 seen in Fig. 2 comes from a combination of both the optimiser and the subspace construction. 3456789101520253035404550 Number of seed latents (d) 0.5 0.6 0.7 0.8 TM-Score (Mean of K) O3 (Random) O3 (Bayes Opt) Best-of-N (mean=0.60, n=5) Figure 4. Bayesian optimisation versus random sampling in the O3 subspace. Mean-of-KTM-score atN = 100for O3 with Bayesian optimisation and O3 with uniform random sampling in the same subspace, on 1CLL. The Best K-of-N baseline is invariant under our fixedK/Nratio. Bars are means over3random seeds; error bars are±1 std. In the experiment summarised by Fig. 5 we sweep across subspace dimensionsdfor different available budgets. The key takeaway is that best-performingdvaries with the bud- get ā different oracle-call regimes favour different sub- space dimensions. For example, whenN = 20the best- performing dimension isd = 6; however, forN = 200 dimensiond = 10yields better predictions. It is clear there is a trade-off between the richness of the O3 subspace, guar- anteed by increasing the number of seed latentsd, and the resulting difficulty of the optimisation task in that subspace. Additional O3 ablations and results are presented in Appen- dices A.2 and B.3. 345678910 Number of seed latents (d) 0.60 0.65 0.70 0.75 0.80 0.85 TM Score (Max of K) Budget N=20 N=50 N=100 N=200 N=500 Figure 5. Performance for different O3 subspace dimensions. Max-of-KTM-score for O3 with Bayesian optimisation, across subspace dimensionsdfor different budgets, on 1CLL. The best- performingdvaries with the budget. Bars are means over3random seeds; error bars are±1 std. 5.3. FK-steering We ablate three of FK-steeringās hyperparameters: the num- ber of particles,K, the number of resampling stepsN/K, and the reward-scaling coefficientĪ». Fig. 6 shows that, with KandNfixed, increasingĪ»improves performance across all budgets, however, it likely hinders output diversity. A highĪ» = 50amplifies the signal from the oracle, forcing the algorithm to greedily resample the best-performing particles across the trajectory, driving the mean-of-Kand max-of-K (see Fig. A.7) closer. In Fig. A.8 we varyKper fixed bud- gets ofNand find that increasing the number of resampling stepsN/Kat the cost of particle populationKgenerally yields better performance at higher oracle budgets. Addi- tional FK-Steering ablations are available in Appendix A.3. Figure 6. Effect of FK-steeringĪ»hyperparameter across differ- ent budgetsN. Mean-of-KTM-score achieved by FK-steering on 1CLL. Bars are means over10random seeds; error bars are±1 std. HigherĪ»increases the signal from the rewards and leads to better performance across budgets. 5.4. DPO Fig. 7 compares online and offline DPO across oracle bud- gets. (Max-of-Kcounterpart of this result is reported in Appendix A.4). Online DPO improves steadily withN, reaching a mean TM-score of0.708and a max TM-score of0.783atN = 1000. In contrast, offline DPO shows little sensitivity to budget, plateauing at0.55ā0.57across all bud- gets. The gap between the two variants widens with scale: negligible atN = 100, it becomes substantial atN = 500. These results suggest that the gains from DPO arise primar- ily from adaptive on-policy resampling, rather than from additional optimisation steps alone, consistent with prior observations that online preference optimisation methods can significantly outperform offline counterparts under a fixed budget (Tang et al., 2024). We also find that train- ing hyperparameters such as the preference batch size and the number of samples per epoch can meaningfully affect performance at largeN, though a controlled study isolating each factor is left for future work. N=100N=500N=1000 0.5 0.6 0.7 0.8 TM Score (Mean of K) DPO (online) DPO (offline) Figure 7. Online vs. offline DPO across oracle budgets. Mean- of-KTM-score achieved by online and offline DPO on 1CLL, for oracle budgetsN ā 100, 500, 1000. Bars are means over 5 random seeds; error bars show±1std. Online DPO improves consistently withN, while offline DPO remains flat across budgets. 6 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models 6. Conclusions Notably, our work represents the first application of O3 to protein structure prediction. We demonstrate that this novel approach equips practitioners with a highly effec- tive tool, yielding exceptional results particularly in budget- constrained tasks. Using Boltz-2 and two protein targets along with the TM-score and MolProbity oracles, we com- pare O3, FK-steering, DPO, and Best K-of-N sampling under oracle-budget constraints, thus providing a practical reference for budget-aware guidance in biological settings. Across the budgets we evaluated, O3 dominates the per- formance. At⤠100queries, it is the only method that meaningfully improves on the Best K-of-N baseline, while FK-steering and DPO both steadily improve withN. As the only method that updates the model parameters, DPO is the natural candidate at substantially larger budgets than we evaluate here. The conclusion from this work can be summarised in the following piece of practical advice: use O3 at constrained oracle budgets. FK-steering can be competitive at mod- erate budgets, and reach for DPO once budgets are large enough to fine-tune. An additional practical consideration is batch size and required batch diversity. DPO amortises its training budget into the model weights, making sampling large subsequent batches cheap, whereas inference-time and search methods usually discard oracle values afterwards. Fu- ture work will include additional protein targets paired with other, potentially expensive and non-differentiable, oracles. Impact Statement Our work is a methodological comparison of guidance meth- ods for protein-structure foundation models, aimed at help- ing practitioners spend limited biological oracle budgets effectively, including for downstream applications such as drug discovery and therapeutic design. As with all general- purpose tools for protein design, the same techniques could in principle be applied to harmful targets; our experiments use only the publicly deposited calmodulin structure 1CLL and E. coli aspartate transcarbamoylase structure 9EEH and pose no specific dual-use risk in themselves. References Abramson, J., Adler, J., Dunger, J., Evans, R., Green, T., Pritzel, A., Ronneberger, O., Willmore, L., Ballard, A. J., Bambrick, J., Bodenstein, S. W., Evans, D. A., Hung, C.- C., OāNeill, M., Reiman, D., Tunyasuvunakool, K., Wu, Z., Ė Zemgulyt Ģ e, A., Arvaniti, E., Beattie, C., Bertolli, O., Bridgland, A., Cherepanov, A., Congreve, M., Cowen- Rivers, A. I., Cowie, A., Figurnov, M., Fuchs, F. B., Gladman, H., Jain, R., Khan, Y. A., Low, C. M. R., Perlin, K., Potapenko, A., Savy, P., Singh, S., Stecula, A., Thillaisundaram, A., Tong, C., Yakneen, S., Zhong, E. D., Zielinski, M., Ė Z Ģ Ä±dek, A., Bapst, V., Kohli, P., Jader- berg, M., Hassabis, D., and Jumper, J. M. Accurate structure prediction of biomolecular interactions with Al- phaFold 3. Nature, 630(8016):493ā500, May 2024. ISSN 1476-4687. URLhttp://dx.doi.org/10.1038/ s41586-024-07487-w. Ament, S., Daulton, S., Eriksson, D., Balandat, M., and Bakshy, E. Unexpected improvements to expected im- provement for bayesian optimization. Advances in neural information processing systems, 36:20577ā20612, 2023. Balandat, M., Karrer, B., Jiang, D. R., Daulton, S., Letham, B., Wilson, A. G., and Bakshy, E. BoTorch: A Frame- work for Efficient Monte-Carlo Bayesian Optimization. In Advances in Neural Information Processing Systems 33, 2020. URLhttp://arxiv.org/abs/1910. 06403. Berman, H. M., Westbrook, J., Feng, Z., Gilliland, G., Bhat, T. N., Weissig, H., Shindyalov, I. N., and Bourne, P. E. The Protein Data Bank. Nucleic Acids Research, 28(1): 235ā242, January 2000. ISSN 1362-4962. URLhttp: //dx.doi.org/10.1093/nar/28.1.235. Black, K., Janner, M., Du, Y., Kostrikov, I., and Levine, S. Training Diffusion Models with Reinforcement Learn- ing, 2023. URLhttps://arxiv.org/abs/2305. 13301. Bodin, E., Stere, A., Margineantu, D. D., Ek, C. H., and Moss, H. Linear combinations of latents in generative 7 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models models: subspaces and beyond, 2024. URLhttps: //arxiv.org/abs/2408.08558. Buttenschoen, M., Morris, G. M., and Deane, C. M. PoseBusters: AI-based docking methods fail to gen- erate physically valid poses or generalise to novel se- quences. Chemical Science, 15(9):3130ā3139, 2024. ISSN 2041-6539. URLhttp://dx.doi.org/10. 1039/d3sc04185a. Chai Discovery, Boitreaud, J., Dent, J., McPartlon, M., Meier, J., Reis, V., Rogozhnikov, A., and Wu, K. Chai- 1: Decoding the molecular interactions of life. Octo- ber 2024. URLhttp://dx.doi.org/10.1101/ 2024.10.10.615955. Chattopadhyaya, R., Meador, W. E., Means, A. R., and Quiocho, F. A. Calmodulin structure refined at 1.7 Ģ a reso- lution. Journal of Molecular Biology, 228(4):1177ā1192, December 1992. ISSN 0022-2836. URLhttp://dx. doi.org/10.1016/0022-2836(92)90324-D. Chen, V. B., Arendall, W. B., Headd, J. J., Keedy, D. A., Immormino, R. M., Kapral, G. J., Murray, L. W., Richard- son, J. S., and Richardson, D. C. MolProbity: all- atom structure validation for macromolecular crystal- lography. Acta Crystallographica Section D: Biologi- cal Crystallography, 66(1):12ā21, 2010. doi: 10.1107/ S0907444909042073. Clark, K., Vicol, P., Swersky, K., and Fleet, D. J. Di- rectly Fine-Tuning Diffusion Models on Differentiable Rewards, 2023. URLhttps://arxiv.org/abs/ 2309.17400. Dhariwal, P. and Nichol, A. Diffusion Models Beat GANs on Image Synthesis, 2021. URLhttps://arxiv. org/abs/2105.05233. Garnett, R. Bayesian optimization. Cambridge University Press, 2023. G Ģ omez-Bombarelli, R., Wei, J. N., Duvenaud, D., Hern Ģ andez-Lobato, J. M., S Ģ anchez-Lengeling, B., She- berla, D., Aguilera-Iparraguirre, J., Hirzel, T. D., Adams, R. P., and Aspuru-Guzik, A. Automatic Chemical De- sign Using a Data-Driven Continuous Representation of Molecules. ACS Central Science, 4(2):268ā276, January 2018. ISSN 2374-7951. URLhttp://dx.doi.org/ 10.1021/acscentsci.7b00572. Ho, J. and Salimans, T. Classifier-Free Diffusion Guid- ance, 2022.URLhttps://arxiv.org/abs/ 2207.12598. Hollingsworth, S. A. and Dror, R. O. Molecular Dynamics Simulation for All. Neuron, 99(6):1129ā1143, September 2018. ISSN 0896-6273. URLhttp://dx.doi.org/ 10.1016/j.neuron.2018.08.011. Karras, T., Aittala, M., Aila, T., and Laine, S. Elucidating the Design Space of Diffusion-Based Generative Mod- els, 2022. URLhttps://arxiv.org/abs/2206. 00364. Knothe, H. Contributions to the theory of convex bod- ies. Michigan Mathematical Journal, 4(1), January 1957. ISSN 0026-2285. URLhttp://dx.doi.org/10. 1307/mmj/1028990175. Maus, N., Jones, H. T., Moore, J. S., Kusner, M. J., Brad- shaw, J., and Gardner, J. R. Local Latent Space Bayesian Optimization over Structured Inputs, 2022. URLhttps: //arxiv.org/abs/2201.11872. Miller, R. C., Patterson, M. G., Bhatt, N., Pei, X., and Ando, N. Cooperativity in e. coli aspartate transcarbamoylase is tuned by allosteric breathing. Nature Communications, 17(1):4285, mar 2026. ISSN 2041-1723. doi: 10.1038/ s41467-026-70909-y. URLhttps://doi.org/10. 1038/s41467-026-70909-y. Moss, H., Ober, S. W., and Diethe, T. Return of the latent space cowboys: Re-thinking the use of vaes for bayesian optimisation of structured spaces. In International Con- ference on Machine Learning, p. 44956ā44970. PMLR, 2025. Passaro, S., Corso, G., Wohlwend, J., Reveiz, M., Thaler, S., Somnath, V. R., Getz, N., Portnoi, T., Roy, J., Stark, H., Kwabi-Addo, D., Beaini, D., Jaakkola, T., and Barzilay, R. Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction. June 2025. URLhttp://dx.doi.org/ 10.1101/2025.06.14.659707. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. Direct Preference Optimization: Your Language Model is Secretly a Reward Model, 2023. URL https://arxiv.org/abs/2305.18290. Rosenblatt, M. Remarks on a Multivariate Transformation. The Annals of Mathematical Statistics, 23(3):470ā472, September 1952. ISSN 0003-4851. URLhttp://dx. doi.org/10.1214/aoms/1177729394. Singhal, R., Horvitz, Z., Teehan, R., Ren, M., Yu, Z., McK- eown, K., and Ranganath, R. A General Framework for Inference-time Scaling and Steering of Diffusion Mod- els, 2025. URLhttps://arxiv.org/abs/2501. 06848. Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Er- mon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020. 8 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models Tang, Y., Guo, D. Z., Zheng, Z., Calandriello, D., Cao, Y., Tarassov, E., Munos, R., Ģ Avila Pires, B., Valko, M., Cheng, Y., and Dabney, W. Understanding the perfor- mance gap between online and offline alignment algo- rithms. arXiv preprint arXiv:2405.08448, 2024. Tripp, A., Daxberger, E., and Hern Ģ andez-Lobato, J. M. Sample-Efficient Optimization in the Latent Space of Deep Generative Models via Weighted Retraining, 2020. URL https://arxiv.org/abs/2006.09191. Wallace, B., Dang, M., Rafailov, R., Zhou, L., Lou, A., Purushwalkam, S., Ermon, S., Xiong, C., Joty, S., and Naik, N. Diffusion Model Alignment Using Direct Pref- erence Optimization, 2023. URLhttps://arxiv. org/abs/2311.12908. Wayment-Steele, H. K., Ojoawo, A., Otten, R., Apitz, J. M., Pitsawong, W., H Ģ omberger, M., Ovchinnikov, S., Colwell, L., and Kern, D.Predicting multiple conformations via sequence clustering and AlphaFold2. Nature, 625(7996):832ā839, November 2023. ISSN 1476-4687. URLhttp://dx.doi.org/10.1038/ s41586-023-06832-9. Williams, C. J., Headd, J. J., Moriarty, N. W., Prisant, M. G., Videau, L. L., Deis, L. N., Verma, V., Keedy, D. A., Hintze, B. J., Chen, V. B., Jain, S., Lewis, S. M., Aren- dall, W. B., Snoeyink, J., Adams, P. D., Lovell, S. C., Richardson, J. S., and Richardson, D. C. MolProbity: More and better reference data for improved all-atom structure validation. Protein Science, 27(1):293ā315, 2018. doi: 10.1002/pro.3330. Williams, C. K. and Rasmussen, C. E. Gaussian processes for machine learning, volume 2. MIT press Cambridge, MA, 2006. Willis, S., Stere, A. I., Margineantu, D. D., Oldroyd, H. T., Fozard, J. A., Ek, C. H., Moss, H., and Bodin, E. Defining latent spaces by example: optimisation over the outputs of generative models, 2025. URLhttps://arxiv. org/abs/2509.23800. Yin, R. and Pierce, B. G. Evaluation of AlphaFold antibodyā antigen modeling with implications for improving predic- tive accuracy. Protein Science, 33(1), December 2023. ISSN 1469-896X. URLhttp://dx.doi.org/10. 1002/pro.4865. Zhang, Y. and Skolnick, J. Scoring function for automated assessment of protein structure template quality. Proteins: Structure, Function, and Bioinformatics, 57(4):702ā710, October 2004. ISSN 1097-0134. URLhttp://dx. doi.org/10.1002/prot.20264. 9 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models A. Additional Results on PDB: 1CLL A.1. Complementary Model Comparison Across Budgets N=20N=50N=100N=200N=500N=1000 0.4 0.5 0.6 0.7 0.8 0.9 TM Score (Max of K) Method FK-steering DPO Best K-of-N O3 Figure A.1. Max-of-KTM-score across oracle budgets. Comparison of FK-steering, DPO, Best K-of-N, and O3 on 1CLL across the six (N,K) configurations of Section 4, reporting the maximum TM-score across the returned batch of K candidate structures. Bars are the mean over 5 random seeds; error bars are±1 std. Higher is better. A.2. Additional O3 Details and Results on PDB: 1CLL O3 EXPERIMENTAL DETAILS Table A.1. Experimental configurations for O3 protein structure optimisation (Boltz-2 generator, TM-score oracle on 1CLL).N: total Oracle budget;K: top-Kbatch selection;M: initial samples filtered to create subspaceU;d: number of seed latents used to create thedā 1dimensionalU;d + 2: initial points used to fit the Gaussian process inU;n rounds : number of BO acquisition rounds is set to N ā M ā 2for filtered seed-latents (andN ā (d + 2)for random seed latents); Wall-clock time of experiments run using a H100 GPU with 80Gb RAM. TaskN KMd n rounds wall-clock time (mins) n20k220610584.5 (±2.7) n50k5507257236.1 (±0.4) n100k10100105054810.6 (± 1.4) n200k2020010100109820.6 (±8.6) n500k5050072503024848.4588 (±2.2) n1000k100100010040050598153.5 (±1.6) LOL INTERPOLATION Figure A.2. LOL interpolation trajectory in a Boltz-2d = 2surrogate subspace of two Calmodulin examples, using 5 intermediate values of(w, 1ā w)forw ā (0, 1). Structures vary smoothly along the interpolation while remaining within the support of Boltz-2. For more details see (Bodin et al., 2024). 10 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models COMPLEMENTARY O3 METRICS (a) 3456789101520253035404550 Number of seed latents (d) 0.5 0.6 0.7 0.8 0.9 TM-Score (Max of K) O3 (Random) O3 (Bayes Opt) Best-of-N (mean=0.65, n=5) (b) 345678910 Number of seed latents (d) 0.60 0.65 0.70 0.75 0.80 0.85 TM Score (Mean of K) Budget N=20 N=50 N=100 N=200 N=500 Figure A.3. Complementary O3 metrics. (a) Max-of-K counterpart to Fig. 4: BO versus random sampling at N = 100 and K = 10. (b) Mean-of-K counterpart to Fig. 5: subspace dimension sweep across budgets. O3 ABLATE BUDGET ALLOCATION (a) (b) Figure A.4. O3 Ablation: Budget allocation M. (a) Mean-of-KTM Score reported for a fixed budgetN = 100, varying the initial number of samples used to create theUsubspace as per Section 3.1, resulting in varying remaining BO budget,N ā Macquisition rounds. (b) Max-of-Kcounterpart. In all cases we keep the subspace-dimension constantd = 7,K = 10, and drawj = 10uniform randomusamples as initial GP training data. Bars are the mean over 3 random seeds; error bars are±1std. Higher is better. Experiment uses a previous set of hyperparameters. O3 EFFECT OF GP COVARIANCE FUNCTION (a) (b) Figure A.5. O3 Ablation: GP covariance function. (a) Mean-of-KTM Score across different subspace dimensionality for an RBF kernel vs Matern 5/2 kernel as Gaussian process covariance function. (b) Max-of-Kcounterpart to (a). In all cases we keep the budget N = 100, batch sizeK = 10, withj = 10uniform randomusamples, as per Section 3. Bars are the mean over 5 random seeds; error bars are±1 std. Higher is better. Experiment uses a previous set of hyperparameters. 11 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models O3 EFFECTS OF SEED SELECTION STRATEGY (a) 3456789101520253035404550 Subspace dimensionality (d) 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 TM Score (Mean of K) Best d seeds Random d seeds (b) 3456789101520253035404550 Subspace dimensionality (d) 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 TM Score (Max of K) Best d seeds Random d seeds Figure A.6. O3 Ablation: Choice of seed selection strategy. (a) Mean-of-KTM Score on 1CLL for different approaches to construct the Usubspace: either the ābestdseedsā or ārandomdseedsā are used, across a sweep of subspace dimensionalities. Using random seeds means M = 0, so number of BO rounds isN ā (d + 2), whereas bestdseeds has fewerN ā M ā 2. (b) Max-of-Kcounterpart to (a). In all cases we keep the budgetN = 100and batch sizeK = 10. Bars are the mean over 3 random seeds; error bars are±1std. Higher is better. A.3. Additional FK-Steering Results on PDB: 1CLL COMPLEMENTARY FK-STEERING METRICS Figure A.7. Effect of FK-steeringĪ»hyperparameter across different budgetsN. Max-of-K TM-score achieved by FK-steering on 1CLL. HigherĪ»increases the signal from the rewards and leads to better performance across budgets. Bars are means over10random seeds; error bars are±1 std. FK-STEERING K VS N/K TRADEOFF (a) (b) Figure A.8. FK-steeringKvsN/Ktradeoff for a set budget. (a) Mean-of-KTM Score for different approaches to allocateKparticles andN/Kresampling steps per set budget ofNoracle calls. (b) Max-of-Kcounterpart to (a). For a given budgetN, increasing the number of resampling stepsN/Kat the cost of particle populationKgenerally improves performance in the higher budget groups. In all cases Ī» = 50. Bars are the mean over 10 random seeds; error bars are±1 std. Higher is better. 12 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models FK-STEERING REWARD TRAJECTORIES (a)(b) Figure A.9. FK-steering reward trajectory forK = 10particles with different number of resampling stepsN/K. (a)N/K = 10 resampling steps with a budget ofN = 100oracle calls. (b)N/K = 50resampling steps with a budget ofN = 500oracle calls. In both casesĪ» = 50. Colors preserveN/Kcolorscheme from Fig. A.8. Resampling more frequently at higherN/Kpropagates the reward signal better and reduces reward variance in the final resampling steps. A.4. Additional DPO Results on PDB: 1CLL N=100N=500N=1000 0.5 0.6 0.7 0.8 TM Score (Max of K) DPO (online) DPO (offline) Figure A.10. Online vs. offline DPO across oracle budgets. Max-of-KTM-score achieved by online and offline DPO on 1CLL, for oracle budgets N ā100, 500, 1000. Bar are means over 5 random seeds; error bars show±1 std. Online DPO improves consistently with N , while offline DPO remains flat across budgets. 13 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models A.5. Compute resources used in PDB: 1CLL Experiments Table A.2. Compute resources used by the Boltz-2 guidance PDB: 1CLL experiment reported in this paper. āEvals/runā counts oracle (TM-score) evaluations, each of which corresponds to one full generation through Boltz-2. The FK-steering rows aggregate over the 3 values of Ī» swept per budget. ExperimentModel / pipelineHardwareEvals/run# runsPer-run wall-clockTotal wall-clock O3 (Willis et al., 2025)BOLTZ-2 (PF-ODE)H1006 budgets up to N=10005 per budgetā¼ 0.1 h ā¼ 3 h N=2052.7 m 13 m N=5052.8 m 14 m N=10053.1 m 16 m N=20052.9 m 15 m N=50053.4 m 17 m N=1000518 m 1 h 31 m Best K-of-NBOLTZ-2H1006 budgets up to N=10005 per budgetā¼ 0.1 h ā¼ 3 h N=2052.7 m 13 m N=5052.8 m 14 m N=10053.1 m 16 m N=20052.9 m 15 m N=50053.4 m 17 m N=1000518 m 1 h 31 m FK-steering (Singhal et al., 2025)BOLTZ-2H1006 budgets, 3 Ī» values each N=20, 3 Ī» vals103.3 m 1 h 39 m N=50, 3 Ī» vals103.8 m 1 h 54 m N=100, 3 Ī» vals104.1 m 2 h 3 m N=200, 3 Ī» vals105 m 2 h 30 m N=500, 3 Ī» vals106.9 m 3 h 27 m N=1000, 3 Ī» vals1011 m 5 h 30 m Online DPO (Rafailov et al., 2023; Wallace et al., 2023)BOLTZ-2 (fine-tuned)H1006 budgets up to N=1000 N=2044 m 16 m N=5046 m 24 m N=100511 m 55 m N=200317 m 51 m N=500544 m 3 h 40 m N=100041 h 31 m 6 h 04 m Offline DPO (Rafailov et al., 2023; Wallace et al., 2023)BOLTZ-2 (fine-tuned)H1006 budgets up to N=1000 N=2044 m 16 m N=5057 m 35 m N=100418 m 1 h 12 m N=200556 m 4 h 40 m N=50045 h 40 m 22 h 40 m N=1000321 h 57 m 65 h 51 m B. Experimental Setup and Results on PDB: 9EEH B.1. MolProbity Oracle TM-score (Section 4) requires a ground-truth structure, which is usually unavailable in the design settings this paper targets. We therefore also evaluate against a reference-free oracle that scores the physical plausibility of a generated structure, built on MolProbity (Chen et al., 2010; Williams et al., 2018). The oracle runs six checks. Backbone dihedral quality comes from a Ramachandran analysis and Cβdeviation count viammtbx. Steric quality is the clashscore, the number of serious all-atom overlaps (probeābad overlapā contacts with a gap belowā0.4 Ģ A) per 1000 atoms, computed after adding and optimising all hydrogens withreduce. Covalent geometry is scored as RMSZ-scores of backbone bond lengths and bond angles against the Engh-Huber reference values, and peptide planarity contributes counts of cis-nonPro and twistedĻangles. These terms are combined into a single log-weighted penalty following the MolProbity score formulation, which we flip and normalise to a validity score in [0, 1], with higher values indicating fewer physical violations. B.2. PDB: 9EEH Results We repeat the experiment of Section 5.1 on E. coli aspartate transcarbamoylase (PDB: 9EEH) (Miller et al., 2026), scored with the reference-free MolProbity oracle described above. There are three key differences between this setup and the one with the TM-score oracle on the calmodulin protein (PDB: 1CLL). First, 9EEH was deposited to the PDB after the training cutoff of Boltz-2, thus it is likely an out-of-distribution task for the base model, whereas 1CLL is not. Second, 9EEH is a 7,232-atom complex, which is much larger than 1CLL with 1,184 atoms. Third, the MolProbity score measures physical plausibility rather than similarity to a reference structure, and its effective range is a lot narrower, with most structures typically scoring between 0.5 and 0.7 under our MolProbity implementation. 14 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models Fig. B.1 reports the mean-of-KMolProbity score. O3 scores highest at all but the smallest budget, where it matches Best K-of-N. Best K-of-N has a relatively constant performance across budgets, as expected under fixedK/Nratio. DPO improves with N, although it performs significantly worse than Best K-of-N at low budgets. FK-steering performs poorly throughout and does not improve with budget, in contrast to 1CLL where it does better than Best K-of-N for budgets above N = 100. In Fig. B.2 under max-of-K, Best K-of-N is strongest at large budgets and both it and FK-steering improve with N , while O3 is unchanged. N=20N=50N=100N=200N=500N=1000 0.56 0.58 0.60 0.62 0.64 0.66 MP Score (Mean of K) Method FK-steering DPO Best K-of-N O3 Figure B.1. Mean-of-KMolProbity score across oracle budgets. Comparison of FK-steering, DPO (in the online setting), Best K-of-N, and O3 on 9EEH across the six(N,K)configurations of Section 4. Bars are the mean over3random seeds; error bars are±1std. Higher is better. N=20N=50N=100N=200N=500N=1000 0.56 0.58 0.60 0.62 0.64 0.66 MP Score (Max of K) Method FK-steering DPO Best K-of-N O3 Figure B.2. Max-of-KMolProbity score across oracle budgets. Comparison of FK-steering, DPO, Best K-of-N, and O3 on 9EEH across the six(N,K)configurations of Section 4, reporting the maximum MolProbity score across the returned batch ofKcandidate structures. Bars are the mean over 3 random seeds; error bars are±1 std. Higher is better. Under max-of-K, O3 performs worse than Best K-of-N, and worse than all methods at the highest budget. We hypothesise that this is because it is the only method that requires sampling Boltz-2 with the probability-flow ODE while the other three methods use the stochastic Boltz-2 sampler which improves the output diversity. Under mean-of-Kthis yields worse performance overall because the lower-scoring samples drag down the average, however it improves the max-of-K scores, where only the best sample matters and increasingNgives better scores even without any guidance signal (i.e. Best K-of-N baseline). FK-steering is the only method that scores intermediate denoised predictions, and it is the weakest method under the MolProbity oracle. TM-score varies smoothly with structural accuracy: the same global fold already scores above 0.5, and refinement moves the score toward 1, so intermediate predictions receive informative scores. MolProbity instead measures local geometry, in which residual noise dominates until the final denoising steps. Since the quality of the denoised predictions is much worse at the intermediate steps than at the final ones, this leaves noise-sensitive intermediate MolProbity scores nearly uninformative. Moreover, since FK-steering relies on a positive noise scale for particle diversity, in Boltz-2 resampling is concentrated in the first three quarters of the denoising trajectory, which exacerbates the issues with signal quality from the MolProbity oracle at intermediate steps. Finally, the differences between methods on this target are small, and the error bars over three seeds overlap at several budgets. Three of the four methods behave much as they do on 1CLL ā O3 scores highest, DPO improves withN, and Best K-of-N is relatively constant ā with FK-steering the exception, for the reason above. We therefore read these orderings as consistent trends across budgets rather than as individually meaningful differences, and note that guidance is less effective on this target than on 1CLL with TM-score. 15 How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models B.3. Additional O3 Details and Results on PDB: 9EEH O3 EXPERIMENTAL DETAILS Table B.1. Experimental configurations for O3 protein structure optimisation (Boltz-2 generator, MolProbity oracle on 9EEH).N: total Oracle budget;K: top-Kbatch selection;M: initial samples filtered to create subspaceU;d: number of seed latents used to create thedā 1dimensionalU;d + 2: initial points used to fit the Gaussian process inU;n rounds : number of BO acquisition rounds is set to N ā M ā 2for filtered seed-latents (andN ā (d + 2)for random seed latents); Wall-clock time of experiments run using a H100 GPU with 80Gb RAM. TaskN KMd n rounds wall-clock time (mins) n20k220210589.1 (±0.1) n50k55052552319.2 (±0.3) n100k101001050204840.0 (± 0.7) n200 k2020020100209875.9 (±2.0) n500k505005025030248224.2 (±9.0) n1000k100100010040050598368.0 (±7.4) O3 EFFECTS OF SEED SELECTION STRATEGY (a) 35791520253035404550 Subspace dimensionality (d) 0.60 0.61 0.62 0.63 0.64 0.65 MP Score (Mean of K) Best d seeds Random d seeds (b) 35791520253035404550 Subspace dimensionality (d) 0.60 0.61 0.62 0.63 0.64 0.65 MP Score (Max of K) Best d seeds Random d seeds Figure B.3. O3 Ablation: Choice of seed selection strategy. (a) Mean-of-KMolProbity Score on 9EEH for different approaches to construct theUsubspace: either the ābestdseedsā or ārandomdseedsā are used, across a sweep of subspace dimensionalities. Using random seeds means M = 0, so the number of BO rounds isN ā (d + 2), whereas bestdseeds has fewerN ā M ā 2. (b) Max-of-K counterpart to (a). In all cases, we keep the budgetN = 100and batch sizeK = 10. Bars are the mean over 3 random seeds; error bars are±1 std. Higher is better. 16