Paper deep dive
Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design
Gyubok Lee, Kiwoong Yoo, Jimin Seo, Kyunghoon Hur, Edward Choi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/24/2026, 5:12:15 AM
Summary
This paper investigates the use of Large Language Models (LLMs) to generate interpretable, multi-metric ranking policies for post-generation shortlisting of protein binders. Using a shared panel of 17 precomputed proxy scores (from AF2-Multimer, Boltz-2, Protenix, and Rosetta), the study compares fixed heuristics, supervised ML models, and LLM-generated policies (global and target-conditioned) on held-out datasets. Results show that iterative LLM policies (gpt-4o and gpt-5.4) modestly improve Recall@10 over strong single-feature baselines like Protenix ipTM, demonstrating that LLMs can effectively synthesize heterogeneous proxy metrics for binder prioritization.
Entities (19)
Relation Signals (12)
GPT-5.4 â achievesperformanceon â Recall@10
confidence 95% ¡ target-conditioned iterative gpt-5.4 policies reach the strongest LLM performance, with 0.519 Recall@10
GPT-4o â achievesperformanceon â Recall@10
confidence 95% ¡ averaging performance over five sampled global iterative gpt-4o policies reaches 0.589 Recall@10
Nipah â ispartof â 3-target held-out subset
confidence 95% ¡ 3-target held-out subset comprising Nipah, RBX1, and TREM2
TREM2 â ispartof â 3-target held-out subset
confidence 95% ¡ 3-target held-out subset comprising Nipah, RBX1, and TREM2
RBX1 â ispartof â 3-target held-out subset
confidence 95% ¡ 3-target held-out subset comprising Nipah, RBX1, and TREM2
Protenix â providesmetric â ipTM
confidence 95% ¡ Protenix binder ipTM, which reaches 0.571 Recall@10
XGBoost â isbaselinefor â Shortlisting
confidence 90% ¡ supervised ML baselines based on ... XGBoost
Logistic Regression â isbaselinefor â Shortlisting
confidence 90% ¡ supervised ML baselines based on logistic regression
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Modern de novo design workflows generate many candidate protein binders, but wet-lab validation capacity remains limited, making shortlisting a major bottleneck. We study whether LLMs can generate multi-metric ranking policies from precomputed structural-confidence and interface-quality proxy scores. Rather than proposing a new protein binder design pipeline, we focus on post-generation binder shortlisting: selecting the final top-K candidates from already generated binder pools using a shared panel of precomputed proxy scores. On the 10-target held-out split, averaging performance over five sampled global iterative gpt-4o policies reaches 0.589 Recall@10, modestly improving over the strongest single-feature fixed baseline, Protenix binder ipTM, which reaches 0.571 Recall@10. On the 3-target held-out subset comprising Nipah, RBX1, and TREM2, target-conditioned iterative gpt-5.4 policies reach the strongest LLM performance, with 0.519 Recall@10 and 0.583 NDCG@10. These results suggest that LLM-generated ranking policies can act as an interpretable post-generation decision layer for combining heterogeneous proxy metrics to prioritize binders from large candidate pools.
Tags
Links
- Source: https://arxiv.org/abs/2608.20755v1
- Canonical: https://arxiv.org/abs/2608.20755v1
Trouble viewing inline? Open PDF directly â
Full Text
60,130 characters extracted from source content.
Expand or collapse full text
Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design Gyubok Lee 1 Kiwoong Yoo 2 Jimin Seo 3 Kyunghoon Hur 4 Edward Choi 1 Abstract Modern de novo design workflows generate many candidate protein binders, but wet-lab valida- tion capacity remains limited, making shortlist- ing a major bottleneck. We study whether LLMs can generate multi-metric ranking policies from precomputed structural-confidence and interface- quality proxy scores. Rather than proposing a new protein binder design pipeline, we focus on post-generation binder shortlisting: selecting the final top-Kcandidates from already gener- ated binder pools using a shared panel of pre- computed proxy scores. On the 10-target held- out split, averaging performance over five sam- pled global iterativegpt-4opolicies reaches 0.589Recall@10, modestly improving over the strongest single-feature fixed baseline, Protenix binder ipTM, which reaches 0.571Recall@10. On the 3-target held-out subset comprising Ni- pah, RBX1, and TREM2, target-conditioned iter- ativegpt-5.4policies reach the strongest LLM performance, with 0.519Recall@10and 0.583 NDCG@10. These results suggest that LLM- generated ranking policies can act as an inter- pretable post-generation decision layer for com- bining heterogeneous proxy metrics to prioritize binders from large candidate pools. 1. Introduction Recent de novo protein binder pipelines couple genera- tive backbone models such as RFdiffusion (Watson et al., 2023), ProteinMPNN-style sequence designers (Dauparas 1 Kim Jaechul Graduate School of AI, Korea Advanced In- stitute of Science and Technology (KAIST), Daejeon, South Korea 2 LG AI Research, Seoul, South Korea 3 Department of Electrical and Computer Engineering, Seoul National Univer- sity, Seoul, South Korea 4 Korea Electronics Technology Institute (KETI), Seongnam, South Korea. Correspondence to: Gyubok Lee <gyubok.lee@kaist.ac.kr>. Accepted at the 2026 Workshop on Generative and Agentic AI for Biology (ICML 2026) et al., 2022), and structure-prediction filters based on Al- phaFold2 (Jumper et al., 2021; Evans et al., 2021) or Boltz- 2 (Passaro et al., 2025). These advances have made large- scale candidate generation increasingly routine, but exper- imental validation capacity remains limited. As a result, shortlisting has become a central determinant of how ef- ficiently generated binders are converted into validated hits (Bennett et al., 2023; Pacesa et al., 2025; Adaptyv Bio, 2026). A common practice is to prioritize or filter candidates us- ing fixed thresholds, single-score rankings, or hand-tuned combinations of structure-prediction confidence scores and interface-centric proxies (Bennett et al., 2023; Pacesa et al., 2025; Adaptyv Bio, 2026). Representative examples in- clude BindCraftâs fixed AF2/Rosetta filter set, Adaptyvâs Boltz-2 ipSAE-based computational selection, and PXDe- signâs Protenix-based confidence filters (Pacesa et al., 2025; Adaptyv Bio, 2026; Team et al., 2025a;b). These workflow- specific filters are important for candidate curation, but they do not fully resolve the final selection problem: after de- signs have been generated, filtered, or collected from dif- ferent workflows, only a small number can be experimen- tally tested. A recent meta-analysis of 3,766 experimentally tested de novo binders reports that interface-focused confi- dence metrics such as ipSAE and orthogonal physicochemi- cal descriptors can improve binder selection, while predic- tive performance still varies substantially by target (Overath et al., 2025). Together, these observations motivate a post- generation shortlisting setting that combines complementary proxy scores and tests whether the ranking rule should be global or target-conditioned. To study this setting, we separate shortlisting from candidate binder generation and treat it as a post-generation decision problem. For each target protein, we fix the generated can- didate pool and a common 17-feature panel of precomputed proxy scores before shortlisting. The panel combines model- native confidence scores from AF2-Multimer, Boltz-2, and Protenix with interface-level proxy descriptors computed from predicted complexes. We use policy broadly to de- note any deterministic shortlisting rule that maps candidate proxy scores to a ranked or selected subset. We compare fixed heuristics, supervised machine learning (ML) base- 1 arXiv:2608.20755v1 [cs.AI] 21 Aug 2026 Natural-Language-Guided Binder Shortlisting Target protein Candidate pool ipTM, pTM, binder pLDDT, interface PAE (iPAE) ipTM, pTM, binder pLDDT, interface ipSAE pDockQ2, Rosetta interface ÎG, buried SASA, shape complementarity Candidate pool Scores LR / XGBoost Trained on labeled source data + Proxy scores âShared scoring model Ranked candidate List âTop-K shortlist K Fixed Heuristics Feature descriptions + Dev. single-feature performance + Dev. pool statistics + Held-out. pool statistics âTarget-specific policy Binder A Binder B Binder C ... Target ... Candidate 1 Candidate 2 Candidate K Candidate H Candidate I ... Evaluation Recall@K Hit@K NDCG@K Input Data Protenix Boltz-2 AF2-Multimer pair ipTM, complex pTM, chain ipTM, chain pTM, binder pLDDT Interface descriptors Shared proxy-score panel Shared policy AF2 ipTM/iPAE Boltz-2 ipTM/pDockQ2 Protenix Protenix binder pTM/ipTM Supervised ML Models Feature descriptions + Dev. single-feature performance + Dev. pool statistics âShared policy Target-conditioned LLM Policy Global LLM Policy Figure 1. Overview of the post-generation binder shortlisting task. For each target protein, a fixed pool of candidate binders is generated in advance by upstream design workflows. A shared proxy-score panel is then computed for all candidates using AF2-Multimer, Boltz-2, Protenix/PXDesign, and interface descriptors such as Rosetta interfaceâG. Shortlisting methods, including fixed heuristics, supervised ML models, global LLM policies, and target-conditioned LLM policies, rank the candidates and return a top-Kshortlist for experimental testing. Global LLM policies apply a single shared ranking rule to all held-out targets, whereas target-conditioned LLM policies generate a target-specific ranking rule. lines, and large language model (LLM)-based shortlisting methods that generate either a single global policy for all held-out targets or a separate target-conditioned policy for each held-out target. Our contributions are threefold: â˘A post-generation shortlisting task. We formulate binder shortlisting as final candidate selection from fixed generated pools using a common 17-feature proxy panel and a shared top-K recall protocol. â˘A controlled comparison of shortlisting strategies. We compare fixed single-feature ranking heuristics, supervised ML baselines, global LLM policies, and target-conditioned LLM policies on the same held-out candidate pools and top-K metrics. â˘LLM-generated ranking policies provide a competi- tive post-generation decision layer. Across held-out binder pools, iterative LLM policies synthesize inter- pretable feature-weighted combinations of structural- confidence and interface-quality proxy scores. These policies are competitive with strong single-feature and supervised baselines in top-Krecall, and in the best settings modestly improve Recall@10. 2. Related Work Binder pipelines and candidate selection. Modern binder-design workflows often combine RFdiffusion back- bone generation (Watson et al., 2023), ProteinMPNN-style sequence design (Dauparas et al., 2022), and structure-based validation or filtering with predictors such as AF2 (Bennett et al., 2023). Existing systems typically implement can- didate selection through workflow-specific filters or rank- ing rules. BindCraft (Pacesa et al., 2025) uses fixed AF2- confidence, Rosetta, and interface-quality filters; Adaptyvâs Nipah release used Boltz-2 ipSAE ranking together with community voting and expert curation (Adaptyv Bio, 2026); and PXDesign (Team et al., 2025b) uses Protenix/AF2- based filtering and releases PXDesignBench for standard- ized monomer and binder evaluation. These pipelines show that predictor-derived confidence and interface metrics are useful for binder triage, but they leave open how to com- bine heterogeneous proxy scores once a fixed candidate pool must be shortlisted for experimental testing. Consis- tent with this gap, a 3,766-binder meta-analysis reports that ipSAE-based scores outperform common interface- confidence metrics and that Rosetta-derived descriptors pro- vide complementary signal, while predictive performance remains target-dependent (Overath et al., 2025). Our work is complementary: we do not introduce a generator or re- produce a pipeline-specific filtering stack, but study a post- generation, generator-agnostic shortlisting layer over het- erogeneous candidate pools using fixed proxy scores from multiple predictor families. Target-aware binder design. Cao et al. (2022) generate binders from target structure alone. Gainza et al. (2023) learn surface fingerprints to parameterize interaction design. 2 Natural-Language-Guided Binder Shortlisting Table 1. Composition of the development and held-out evaluation datasets after target-level merging and exact sequence de-duplication. Counts report designs, experimentally confirmed binders, and targets for each source. For pMHC targets, the table reports short target names. The full peptideâHLA pairs are SLLMWITQCâHLA-A*02:01 and RVTDESILSYâHLA-A*01:01. SplitSourceCandidate originTarget namesDesigns Binders Targets DevelopmentBoltzGenBoltzGenAMBP, HNMT, IDI2, IL-7Ra, Insulin Receptor, MZB1/PERP1, PDGFR Beta, PD-L1, PHYH, PMVK, RFK 32710311 Held-out eval.BindCraft1 revalidationBindCraftIFNAR2, spCas9, Der f 21, Der f 776314 Merged EGFRMixed participant-submitted methods, in- cluding Adaptyv EGFR competition sub- missions and BindCraft EGFR revalidation EGFR605681 pMHC minibinders RFdiffusion + ProteinMPNN + Al- phaFold2 filtering NY-ESO-1 pMHC; RVTDESILSY pMHC13732 Nipah releaseMixed participant-submitted methodsNipah Virus Glycoprotein G (NiV-G)1,0301031 GEM/AdaptyvMixed participant-submitted methodsRBX132191 BioArena/AdaptyvMixed human/agent submissions with het- erogeneous methods TREM2100371 Held-out subtotal 2,26925110 Overall total (development + held-out) 2,59635421 APPRAISE (Ding et al., 2024) ranks engineered proteins by target-binding propensity through structure modeling. These works focus on generation or pairwise compatibility scoring. Our work differs in task: given a generated pool and precomputed proxy scores, we synthesize a separate decision policy for selecting K candidates. LLMs and agents for protein design. ProtAgents (Gha- farollahi & Buehler, 2024) frames protein discovery as a multi-agent collaboration among LLM-backed roles that can retrieve knowledge, analyze structures, and call physics or machine-learning tools. ProteinCrow (Ponnapati et al., 2025) similarly builds an agentic protein-design assistant around curated tools, structural inputs, literature, and bio- chemical context. More broadly, Lee et al. (2025) review language-model use in protein design, including sequence modeling, context-conditioned design, and structure inte- gration. Our use of LLMs is complementary: we apply them at the post-generation decision layer, where they pro- pose ranking policies for constructing final shortlists from already-generated candidate pools. 3. Problem Setup Shortlisting as policy synthesis.For a target proteint, we are given a generated candidate binder poolC t ofn t designs. Each candidatecâC t has a fixed proxy-score vectorx t,c â R m computed before shortlisting, wheremis the number of common features available for every candidate. The task is to choose a subsetS t âC t of sizeKthat maximizes recall of experimentally verified binders under the validation budget. A method therefore outputs a ranking policyĎ t â Î , and a deterministic executor scores every candidate byĎ t and returns the topK. In this formulation, ordinary metric- based ranking is a special case: sorting by one score, such as Boltz-2 ipSAE, is a one-feature policy. Policy synthesis generalizes this by choosing which proxy scores to combine and how strongly to weight them for the target pool. Policy space. We use the term policy for the structured triple(F,w,g)that specifies a ranking function. HereF â f 1 ,...,f m selects a subset of features,w â1, 2, 3 |F| assigns integer weights, and each selected feature has a pre- defined higher-is-better or lower-is-better direction. In the main LLM policy space,gis a weighted normalized sum: the executor normalizes every selected feature within the target pool, computes the aggregate score, sorts candidates in descending order, and returns the top K. 4. Datasets 4.1. Sources and split We collected labeled binder-design data from eight public sources or workflow releases. A source is included only when it reports candidate-level experimental outcomes for tested designs, so that each target defines a retrospective shortlisting episode: the candidate pool is fixed, and every candidate has a binder/non-binder label. Table 1 summarizes the target-disjoint development and held-out splits. The development split contains 11 tar- gets from BoltzGen, a de novo binder-generation work- flow and validation dataset (Stark et al., 2025). The 10- target held-out split combines independent workflow out- puts and public validation releases, including BindCraft revalidation (Pacesa et al., 2025), pMHC minibinders, Nipah (Adaptyv Bio, 2026), RBX1 (GEM Workshop & Adaptyv Bio, 2026), TREM2 (bioArena & Adaptyv Bio, 2026), and merged EGFR challenge/revalidation pools; its experimental labels were released after the docu- mentedgpt-4o-2024-11-20cutoff date used for the knowledge-leakage audit (October 1, 2023). We also re- port a 3-target held-out subset consisting of Nipah, RBX1, and TREM2, whose labels were released after the docu- 3 Natural-Language-Guided Binder Shortlisting Table 2. Candidate-level proxy score panel (17 active features). Feature definitions, monotonic directions, and extraction details are in Appendix C. FamilyFeaturesDescription AF2-MultimeripTM,pTM,binder pLDDT, interface PAE Model-nativecomplexconfi- dence, binder local confidence, and interface PAE from AF2- Multimer (Evans et al., 2021; Mirdita et al., 2022);lower interface PAE is better. Boltz-2ipTM,pTM,binder pLDDT, ipSAE, pDockQ2 Model-native confidence scores from Boltz-2 (Passaro et al., 2025), ipSAE-style interface con- fidence (Dunbrack Jr, 2025), and pDockQ2 (Zhu et al., 2023), re- ported as metric-wise best sum- maries across target-MSA diffu- sion samples. Protenix/PXDesignpair ipTM, complex pTM, binder ipTM, binder pTM, binder pLDDT Confidence fields from Pro- tenix/PXDesign (Team et al., 2025a;b)coveringcomplex, binder-chain, and binder-target confidence under our two-chain target-binder convention. Rosetta interface descriptors interfaceâG,buried SASA, shape complemen- tarity Rosetta descriptors (Stranges & Kuhlman, 2013; Lawrence & Colman, 1993) are computed on Boltz-2 target-MSA predicted complexes, not experimental struc- tures. mentedgpt-5.4cutoff date (August 31, 2025). When the same biological target appears in multiple releases, we merge the corresponding pools and de-duplicate exact candi- date amino-acid sequences. EGFR, for example, combines Adaptyv R1, Adaptyv R2, and BindCraft1 revalidation into 605 unique designs. Full preprocessing details, including binding-outcome parsing, source-specific target assignment, zero-positive target handling, and sequence de-duplication, are provided in Appendix A. 4.2. Feature extraction Each candidate is represented by the 17 active proxy scores summarized in Table 2. Boltz-2 pDockQ2 and ipSAE are post-processed interface-quality or interface-confidence proxies computed from Boltz-2 predicted complexes and confidence outputs; pDockQ2 follows Zhu et al. (2023), and ipSAE follows the PAE-based interprotein scoring approach of Dunbrack Jr (2025). Rosetta InterfaceAnalyzer met- rics are computed on Boltz-2 target-side-MSA complexes and used as geometry, burial, and energy proxy scores, not ground-truth binding energies. Protenix/PXDesign features provide complex, binder-chain, and pairwise binder-target confidence; Protenix ipTM-style scores are not assumed to be numerically calibrated to AF2-Multimer or Boltz-2 ipTM. All evaluated candidates have non-missing values for the 17 active features. Appendix C gives the active features, monotonic directions, and extraction summaries. Inference settings. Feature extraction uses target-side- MSA complex predictions where available. For each tar- get, we precompute one target-chain MSA and reuse it for all candidate binders; the de novo binder chain is kept single-sequence because designed binders have no natural homologs. We run AF2-Multimer with 3 models and 3 re- cycles, and Boltz-2 with 5 diffusion samples, 3 recycling steps, and 200 sampling steps. Protenix/PXDesign predic- tions likewise provide the target chain with the precomputed target MSA while keeping the binder chain single-sequence. Rosetta InterfaceAnalyzer is applied to Boltz-2 target-side- MSA predicted complexes. Templates are disabled in all complex-prediction runs. 5. Experimental Setup 5.1. Shortlisting methods We use global to denote methods that use one rule un- changed across all held-out targets, and target-conditioned to denote methods that generate a separate rule for each held-out target. A target-conditioned method may choose different feature subsets or weights for different candidate pools. Fixed ranking heuristics. We evaluate target-agnostic fixed rules as single-score references spanning the main structure-prediction signals used for binder triage: AF2- Multimer ipTM and interface PAE (lower is better) (Evans et al., 2021; Mirdita et al., 2022), Boltz-2 ipTM and a post- processed interface-quality estimate (pDockQ2 computed from Boltz-2 predicted complexes) (Passaro et al., 2025; Zhu et al., 2023), and Protenix/PXDesign binder-chain confi- dence (binder pTM and binder ipTM) (Team et al., 2025a;b). Each rule ranks candidates within a target pool by one score only, testing how far a commonly used single metric can go before any learned or LLM-composed policy is introduced. The Protenix binder metrics are the binder-chain confidence fields used by the PXDesign Protenix filters, evaluated here as single-score ranking heuristics rather than hard thresh- olds. Logistic regression and XGBoost. We evaluate super- vised ML baselines based on logistic regression and XG- Boost (Chen & Guestrin, 2016). The first fits a logistic regression model with no regularization penalty and an XG- Boost model on the 11-target BoltzGen development split using the same 17-feature panel as the LLM policies; these models test whether direct supervised learning over the feature panel is sufficient without target-conditioned policy synthesis. The second is a transfer baseline using Cao binder pools (Cao et al., 2022) with retrospective AF2 scores from Bennett et al. (2023). For this transfer setting, we use the available AF2 interaction pAE and binder pLDDT scores as the closest historical counterparts to our AF2 interface PAE and binder-chain pLDDT features. These transfer fea- 4 Natural-Language-Guided Binder Shortlisting tures come from the AF2 scores previously computed by Bennett et al. for the Cao binder pools, rather than our AF2- Multimer target-side-MSA feature extraction, so they test cross-protocol transfer rather than a matched supervised re- training setting. Cao targets that overlap evaluation targets (EGFR, IL7Ra, and PDGFR) are removed before fitting. At evaluation time, each supervised model assigns every held-out candidate a fitted probability of being a binder, and candidates are ranked by this probability in descending order. LLM policy sampling and averaging. We evaluate four LLM policy settings withgpt-4oon the 10-target held- out split and on the 3-target held-out subset, and we eval- uate the correspondinggpt-5.4settings on the same 3-target held-out subset.Allgpt-4oresults use the gpt-4o-2024-11-20API snapshot. In the global LLM setting, the LLM sees natural-language feature descrip- tions, per-development-target pool distribution summaries, and development-set single-feature performance, emits one global policy, and that policy is fixed before held-out evalu- ation. The distribution summaries are reported separately for each development target pool rather than pooled across candidates, while the single-feature performance values are target-averaged Recall@10/Hit@10/NDCG@10 values from ranking each development target with one feature at a time. In the global iterative LLM setting, the final accepted policy from development-split iterative search is likewise fixed and applied unchanged to every held-out target. The target-conditioned LLM setting is a label-free test-time adaptation setting: it instantiates one prompt per held-out target and includes the same development calibration con- text as the global LLM prompt, plus that held-out target poolâs identifier and unlabeled score-distribution statistics computed only within the current candidate pool. Thus, the single-turn global and target-conditioned prompts share the development-side information, while target-conditioned prompting additionally exposes unlabeled current-target context and emits one rule per target. The target-conditioned iterative LLM setting additionally receives accepted/rejected development-search feedback from prior policy evaluations. In all LLM settings, each sampled policy selects 3 to 5 features and assigns positive integer weights in1, 2, 3. Each LLM policy defines a weighted rank score over se- lected features: selected features are normalized to a 0â1 range within the target pool, lower-is-better features are direction-corrected so that larger normalized values are bet- ter, candidates are sorted by the resulting weighted score in descending order, and the topKcandidates are selected. Appendix E summarizes the information available to each prompting setting and the shared policy constraints for the 17-feature panel. The reported LLM results average perfor- mance over five independently generated policies, which reduces run-to-run variability without treating policy aver- aging as a separate shortlisting method. For proxy-score family ablations, we remove all selected terms belonging to one feature family from each sam- pled target-conditioned iterative policy, keep the remaining weights fixed, and re-evaluate the same deterministic execu- tor. The LLM is not asked to regenerate or repair the policy after feature removal. Because proxy-score families are cor- related and the remaining policy is not re-optimized, these ablations are interpreted as policy-dependence checks rather than monotonic feature-importance estimates. For model-to- model comparison, the main ablation table uses the shared 3-target held-out subset evaluated for bothgpt-4oand gpt-5.4. 5.2. Evaluation protocol Protocol.K = 10is fixed before evaluation. The primary metric is Recall@10 over verified binders, with denominator min(K,n binders )so that targets with fewer than 10 verified binders can still attain a maximum score of 1.0 by recov- ering all positives. Secondary metrics are Precision@10 and normalized discounted cumulative gain (NDCG@10). Precision@10 is the wet-lab hit rate among the 10 selected designs, whereas NDCG@10 measures whether verified binders are concentrated near the top of the shortlist; its ideal DCG is computed with the samemin(K,n binders )number of positives for each target. Statistics are averaged across target pools rather than pooled across individual candidates. Held-out evaluation.Target-conditioned policies use un- labeled held-out pool statistics at test time, whereas fixed heuristics, supervised baselines, and global LLM variants are fixed before held-out evaluation and applied without held-out pool statistics. Additional provenance checks, his- torical dataset handling, and remaining leakage caveats are provided in Appendix B and Section 7. 6. Results 10-target held-out split.Table 3 reports the held-out eval- uation with separate columns for the 10-target held-out split and the 3-target held-out subset. On the 10-target held-out split, the strongest single-feature fixed baseline is Protenix binder ipTM, with Recall@10 = 0.571, Hit@10 = 0.360, and NDCG@10 = 0.525. Averaging five global iterative gpt-4opolicies gives the highest LLM Recall@10 and Hit@10, reaching Recall@10 = 0.589 and Hit@10 = 0.404, while target-conditioned iterativegpt-4ogives the high- est LLM NDCG@10 (0.523). These results support the interpretation that LLM policies provide inspectable multi- feature ranking rules that can outperform strong predictor- native single-score baselines in Recall@10, while predictor- native single-score rankings remain strong NDCG baselines. 5 Natural-Language-Guided Binder Shortlisting Table 3. Held-out evaluation (K = 10), grouped by method type. Recall@10 uses denominatormin(K, n binders ); Hit@10 is equivalent to Precision@10. Results are shown for the 10-target held-out split and for the 3-target held-out subset containing Nipah, RBX1, and TREM2. Dashes mark model/split combinations not reported. LLM results average five sampled policies. 10-target held-out split3-target held-out subset MethodRecall@10Hit@10NDCG@10Recall@10Hit@10NDCG@10 Fixed ranking heuristics AF2 ipTM0.4370.3600.4170.4330.4330.487 AF2 interface PAE0.4170.2900.3370.3000.3000.298 Boltz-2 ipTM0.4030.2900.3770.1330.1330.166 Boltz-2 pDockQ20.5150.3400.3830.4000.4000.416 Protenix binder pTM0.4550.2800.3270.2330.2330.222 Protenix binder ipTM0.5710.3600.5250.5040.5000.506 Supervised ML baselines LR-BG, no reg., 17 feat.0.5370.3700.4380.5070.5000.503 XGB-BG, 17 feat.0.4250.2700.3900.2670.2670.284 LR-Cao AF20.3500.2600.2820.2670.2670.246 XGB-Cao AF20.3970.2800.3370.2670.2670.272 LLM policies averaged over five samples: gpt-4o Global LLM0.4350.3100.3720.4000.4000.431 Target-conditioned LLM0.4980.3780.4450.4930.4930.526 Global iterative LLM0.5890.4040.4940.5040.5000.529 Target-conditioned iterative LLM0.5840.3940.5230.4970.4930.520 LLM policies averaged over five samples: gpt-5.4 Global LLMâ0.4670.4670.495 Target-conditioned LLMâ0.5020.5000.514 Global iterative LLMâ0.4700.4670.513 Target-conditioned iterative LLMâ0.5190.5130.583 3-target held-out subset. The 3-target held-out subset includes Nipah, RBX1, and TREM2. On this subset, target- conditioned iterativegpt-5.4reaches the strongest LLM performance, with Recall@10 = 0.519, Hit@10 = 0.513, and NDCG@10 = 0.583. The strongest fixed and supervised baselines are Protenix binder ipTM (0.504/0.500/0.506) and LR-BG without regularization (0.507/0.500/0.503). XG- Boost remains unstable in this small setting: it performs well on TREM2 alone but selected no verified binders for either Nipah or RBX1. Generated LLM policy composition. The target- conditioned iterativegpt-4opolicies concentrate weight on a small set of structure-confidence and interface-quality scores rather than spreading weight uniformly (Figure 2). Across the five policy samples for each of the 10 held-out targets, every policy selects Boltz-2 ipTM and Protenix pair ipTM. Rosetta shape complementarity appears in 40 of 50 policies, AF2 ipTM appears in 27, and Boltz-2 pDockQ2 appears once. By total selected weight, the policies allo- cate 38.6% to Boltz-2 features, 29.2% to Protenix features, 19.6% to AF2 features, and 12.6% to Rosetta interface de- scriptors. Thus, the generated policies do not appear to rely on arbitrary target-specific rules. Instead, they use a com- pact Boltz-2/Protenix confidence backbone and make target- conditioned adjustments through AF2 ipTM and Rosetta shape-complementarity inclusion. Proxy-score group ablation of generated policies. Ta- ble 4 applies the same group leave-one-out procedure to target-conditioned iterative policies on the shared 3- target held-out subset, enabling a directgpt-4oversus gpt-5.4comparison without changing the target set. The ablations are non-monotonic because the policy features are correlated and the remaining terms are not re-optimized after removal. Forgpt-4o, removing Protenix causes the largest Recall@10 drop, while removing AF2 or Rosetta de- scriptors causes smaller drops and removing Boltz-2 slightly increases Recall@10. Forgpt-5.4, removing Rosetta de- scriptors or Protenix is most harmful, consistent with its stronger use of target-specific interface-geometry and Pro- tenix signals on the post-cutoff subset. These results should be read as policy-dependence checks, not monotonic feature- importance estimates. Example generated policies.Table 5 shows RBX1 exam- ples, including two fixed one-feature rules and the averaged iterative policies used by the reported LLM settings. Each sampled LLM policy is constrained to select 3 to 5 features, but an averaged example rule can contain more nonzero terms because it shows the union of features selected across five samples. The selected terms concentrate on strong predictor-native and interface-quality signals, such as Boltz- 2 ipTM, Protenix pair ipTM, AF2 ipTM, Rosetta burial, and Rosetta shape complementarity, rather than introducing 6 Natural-Language-Guided Binder Shortlisting AF2 ipTM AF2 pTM AF2 binder pLDDT AF2 interface PAE Boltz-2 ipTM Boltz-2 pTM Boltz-2 binder pLDDT Boltz-2 ipSAE Boltz-2 pDockQ2 Protenix pair ipTM Protenix complex pTM Protenix binder ipTM Protenix binder pTM Protenix binder pLDDT Rosetta interface dG Rosetta buried SASA Rosetta SC spCas9 Der f 21 IFNAR2 Der f 7 EGFR pMHC NY1 pMHC SILSY1 Nipah G RBX1 TREM2 1.42.62.01.0 1.42.62.00.8 1.42.62.00.8 1.42.62.00.8 1.22.62.01.0 1.22.62.01.0 1.42.62.00.8 1.42.80.22.01.0 1.22.62.00.8 1.42.62.00.6 0.0 0.5 1.0 1.5 2.0 2.5 3.0 Average selected weight 2.6 2.6 2.6 2.6 2.6 2.6 2.6 2.8 2.6 2.6 Figure 2. Feature weights in target-conditioned iterativegpt-4oranking policies over the full 17-feature panel. The vertical axis lists held-out targets, and the horizontal axis lists all proxy-score features. Colored cells indicate features selected by the policies; the color scale shows the average selected weight across five sampled policies. Near-zero background cells indicate features not included in the ranking policy. For lower-is-better features, positive plotted weights apply to the direction-corrected normalized score used by the executor. Table 4. Proxy-score group ablation of target-conditioned iterative policies on the shared 3-target held-out subset. Values are mean Recall@10 after removing all selected features from one group. No removal denotes the same ablation policy set before feature removal. Boltz-2 includes Boltz-2-derived pDockQ2; Rosetta denotes Rosetta interface descriptors. Removed group TC iter. gpt-4o 3-target TC iter. gpt-5.4 3-target No removal0.4970.519 AF20.4360.499 Boltz-20.5170.545 Protenix0.3400.453 Rosetta0.4550.423 unrelated proxy scores. For LLM policies, coefficients are average weights across five sampled policies, with unse- lected features contributing 0. Positive terms reward higher raw feature values, while negative terms reward lower raw feature values. 7. Analysis and Discussion Global versus target-conditioned policies. The global LLM uses one rule inferred from development-target infor- mation only, then applies that rule unchanged to all held-out targets. Its Recall@10 on the 10-target held-out split is 0.435. Iterative feedback raises the globalgpt-4opol- icy to 0.589 Recall@10, while target-conditioned iterative gpt-4oreaches a similar 0.584 Recall@10 and the highest LLM NDCG@10 on the same split (0.523). Thus, target conditioning does not uniformly dominate global policy syn- thesis: it can improve ranking quality without improving average top-Krecovery. We interpret target conditioning as label-free adaptation to target-specific score distributions: the prompt can inspect the spread, skew, and relative behav- ior of proxy scores within a held-out candidate pool, then adjust which feature families to trust and how strongly to weight them, while the deterministic executor still applies the same fixed directions, within-target normalization, and top-K selection rule. What iterative feedback adds. Iterative prompting adds development-search feedback by summarizing which feature-weighted policies were accepted or rejected on the development split and reporting their aggregate metrics. This changes the LLMâs role from producing a rule from development summaries alone to producing a rule cali- brated by prior policy search. On the 10-target held-out split, iterative feedback helps the globalgpt-4osetting relative to the single global policy (0.589 vs. 0.435 Re- call@10). The reflection records suggest thatgpt-4oused this feedback mainly to converge on a compact, conservative Boltz-2/Protenix backbone; target conditioning then made only modest feature changes, which explains why target- conditioned iterativegpt-4oimproves NDCG but does not exceed the global iterative policy in Recall@10. In contrast, gpt-5.4reflection records more explicitly discuss target- specific compression, spread, and redundancy in unlabeled score distributions. For RBX1, for example, the target- conditioned policies shift toward Rosetta burial/packing, Protenix pair ipTM, Boltz ipSAE, and AF2 interface PAE 7 Natural-Language-Guided Binder Shortlisting Table 5. Example RBX1 ranking rules. EachËxis a within-target 0 to 1 normalized feature;+rewards higher raw values andârewards lower raw values. For LLM policies, coefficients are average policy weights across five sampled policies, with unselected features contributing weight 0 to the average; therefore, an averaged rule can include more nonzero terms than any single sampled policy. Global policies use one averaged rule for every target, whereas target-conditioned policies are averaged for RBX1 specifically. Candidates are sorted by the resulting rank score in descending order, and the top K candidates are selected. PolicyRank score Fixed Protenix binder ipTM, RBX1Ëx Protenix binder ipTM Fixed Boltz-2 pDockQ2, RBX1Ëx Boltz pDockQ2 Global iterative gpt-4o2.2Ëx Boltz ipTM +2.0Ëx AF2 ipTM +2.0Ëx Protenix pair ipTM +1.0Ëx Rosetta SC + 0.2Ëx Boltz pDockQ2 Target-conditioned iterative gpt-4o, RBX12.6Ëx Boltz ipTM + 2.0Ëx Protenix pair ipTM + 1.2Ëx AF2 ipTM + 0.8Ëx Rosetta SC Global iterative gpt-5.42.4Ëx Protenix pair ipTM + 1.8Ëx AF2 ipTM + 1.2Ëx Boltz pDockQ2 + 1.2Ëx Rosetta SC + 0.8Ëx Rosetta buried SASA â 0.6Ëx Rosetta interface dG + 0.4Ëx AF2 pTM â0.4Ëx AF2 interface PAE +0.4Ëx Boltz ipTM + 0.2Ëx Protenix binder pLDDT + 0.2Ëx Protenix complex pTM Target-conditioned iterativegpt-5.4, RBX12.0Ëx Rosetta buried SASA + 1.8Ëx Protenix pair ipTM + 1.6Ëx Rosetta SC + 1.4Ëx Boltz ipSAE + 0.6Ëx Boltz pDockQ2 â 0.6Ëx AF2 interface PAE + 0.6Ëx Boltz ipTM + 0.4Ëx Boltz binder pLDDT + 0.4Ëx AF2 ipTM + 0.2Ëx AF2 pTM rather than simply reusing the global confidence back- bone. This more selective adaptation is consistent with the 3-target held-out subset, where target-conditioned itera- tivegpt-5.4improves over single-turn target-conditioned gpt-5.4in Recall@10 (0.519 vs. 0.502) and NDCG@10 (0.583 vs. 0.514). Strong single-score baselines remain important. Pro- tenix binder ipTM is the strongest fixed one-feature baseline in the 17-feature panel, reaching 0.571 Recall@10 on the 10-target held-out split and 0.504 Recall@10 on the 3-target held-out subset. Boltz-2 pDockQ2 also remains competi- tive: it estimates interface quality from the predicted binder- target complex and confidence outputs, without using an experimental or designed reference structure. The gener- ated policies do not replace these predictor-native signals. Instead, they combine them with complementary AF2 and Rosetta evidence; for example, all 50 target-conditioned iterativegpt-4opolicies retain Boltz-2 ipTM and Protenix pair ipTM, while 40 include Rosetta shape complementarity and 27 include AF2 ipTM. 8. Conclusion We study post-generation binder shortlisting: selecting final top-Kcandidates from fixed generated binder pools us- ing precomputed structural-confidence and interface-quality proxy scores. We cast this task as policy synthesis and compare fixed heuristics, supervised baselines, and LLM- generated ranking policies. On the 10-target held-out split, averaging performance over five sampled global iterativegpt-4opolicies modestly improves over the strongest single-feature baseline, Protenix binder ipTM, in Recall@10. The generated policies do not replace structure- prediction and interface-quality metrics. Instead, they com- bine AF2, Boltz-2, Rosetta interface, and occasional Pro- tenix confidence signals into interpretable feature-weighted ranking rules that can be adjusted to each target pool. This suggests that LLM-generated ranking policies can serve as an interpretable post-generation decision layer for priori- tizing binders from heterogeneous candidate pools, while strong predictor-native single-score rankings should remain explicit baselines. Acknowledgment This work was supported by the Institute for Informa- tion & communications Technology Planning & Eval- uation (IITP) grant (RS-2019-I190075) and the Na- tional Research Foundation of Korea (NRF) grant (NRF- 2020H1D3A2A03100945), supported by the Korea Govern- ment (MSIT). References Adaptyv Bio.Nipah competition results.Pro- teinbasecollection,2026.URLhttps: //proteinbase.com/collections/ nipah-binder-competition-results.Ex- perimental validation results released January 21, 2026. Alford, R. F., Leaver-Fay, A., Jeliazkov, J. R., OâMeara, M. J., DiMaio, F. P., Park, H., Shapovalov, M. V., Renfrew, P. D., Mulligan, V. K., Kappel, K., et al. The rosetta all- atom energy function for macromolecular modeling and design. Journal of chemical theory and computation, 13 (6):3031â3048, 2017. 8 Natural-Language-Guided Binder Shortlisting Bennett, N. R., Coventry, B., Goreshnik, I., Huang, B., Allen, A., Vafeados, D., Peng, Y. P., Dauparas, J., Baek, M., Stewart, L., et al. Improving de novo protein binder design with deep learning. Nature Communications, 14 (1):2625, 2023. bioArena and Adaptyv Bio. BioArena x Adaptyv: TREM2 binder design competition.Proteinbase competition page, 2026. URLhttps://proteinbase.com/ competitions/bioarena-adaptyv-trem2. Experimental validation results released March 28, 2026. Cao, L., Coventry, B., Goreshnik, I., Huang, B., Sheffler, W., Park, J. S., Jude, K. M., Markovi Ě c, I., Kadam, R. U., Verschueren, K. H., et al. Design of protein-binding proteins from the target structure alone. Nature, 605 (7910):551â560, 2022. Chen, T. and Guestrin, C. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd inter- national conference on knowledge discovery and data mining, p. 785â794, 2016. Dauparas, J., Anishchenko, I., Bennett, N., Bai, H., Ragotte, R. J., Milles, L. F., Wicky, B. I., Courbet, A., de Haas, R. J., Bethel, N., et al. Robust deep learning-based protein sequence design using proteinmpnn. Science, 378(6615): 49â56, 2022. Ding, X., Chen, X., Sullivan, E. E., Shay, T. F., and Grad- inaru, V. Fast, accurate ranking of engineered proteins by target-binding propensity using structure modeling. Molecular Therapy, 32(6):1687â1700, 2024. Dunbrack Jr, R. L. R Ě es ipsae loquunt: Whatâs wrong with alphafoldâs iptm score and how to fix it. bioRxiv, 2025. Evans, R., Oâneill, M., Pritzel, A., Antropova, N., Senior, A., Green, T., Ë Z Ě Äądek, A., Bates, R., Blackwell, S., Yim, J., et al. Protein complex prediction with alphafold-multimer. biorxiv, p. 2021â10, 2021. Gainza, P., Wehrle, S., Van Hall-Beauvais, A., Marchand, A., Scheck, A., Harteveld, Z., Buckley, S., Ni, D., Tan, S., Sverrisson, F., et al. De novo design of protein in- teractions with learned surface fingerprints. Nature, 617 (7959):176â184, 2023. GEM Workshop and Adaptyv Bio. GEM x Adaptyv: RBX1 binder design competition.Proteinbase competition page, 2026. URLhttps://proteinbase.com/ competitions/gem-adaptyv-rbx1 . Experimen- tal validation results released April 26, 2026. Ghafarollahi, A. and Buehler, M. J. Protagents: protein discovery via large language model multi-agent collabo- rations combining physics and machine learning. Digital Discovery, 3(7):1389â1409, 2024. Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Ë Z Ě Äądek, A., Potapenko, A., et al. Highly accurate protein structure prediction with alphafold. nature, 596(7873):583â589, 2021. Lawrence, M. C. and Colman, P. M. Shape complementarity at protein/protein interfaces, 1993. Lee, J. S., Abdin, O., and Kim, P. M. Language models for protein design. Current Opinion in Structural Biology, 92:103027, 2025. Lin, Z., Akin, H., Rao, R., Hie, B., Zhu, Z., Lu, W., Smetanin, N., Verkuil, R., Kabeli, O., Shmueli, Y., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637): 1123â1130, 2023. Mirdita, M., Sch Ě utze, K., Moriwaki, Y., Heo, L., Ovchin- nikov, S., and Steinegger, M. Colabfold: making protein folding accessible to all. Nature methods, 19(6):679â682, 2022. Overath, M. D., Rygaard, A. S., Jacobsen, C. P., Brasas, V., Morell, O., Sormanni, P., and Jenkins, T. P. Predicting experimental success in de novo binder design: a meta- analysis of 3,766 experimentally characterised binders. BioRxiv, p. 2025â08, 2025. Pacesa, M., Nickel, L., Schellhaas, C., Schmidt, J., Pyatova, E., Kissling, L., Barendse, P., Choudhury, J., Kapoor, S., Alcaraz-Serna, A., et al. One-shot design of functional protein binders with bindcraft. Nature, 646(8084):483â 492, 2025. Passaro, S., Corso, G., Wohlwend, J., Reveiz, M., Thaler, S., Somnath, V. R., Getz, N., Portnoi, T., Roy, J., Stark, H., et al. Boltz-2: Towards accurate and efficient binding affinity prediction. BioRxiv, 2025. Ponnapati, M., Cox, S., Gordon, C. W., Hammerling, M. J., Narayanan, S., Laurent, J. M., Braza, J. D., Hinks, M. M., Skarlinski, M. D., Rodriques, S. G., et al. Proteincrow: A language model agent that can design proteins. In ICML 2025 Generative AI and Biology (GenBio) Workshop, 2025. Stark, H., Faltings, F., Choi, M., Xie, Y., Hur, E., OâDonnell, T., Bushuiev, A., Uc ̧ar, T., Passaro, S., Mao, W., et al. Boltzgen: Toward universal binder design. bioRxiv, p. 2025â11, 2025. Stranges, P. B. and Kuhlman, B. A comparison of suc- cessful and failed protein interface designs highlights the challenges of designing buried hydrogen bonds. Protein Science, 22(1):74â82, 2013. 9 Natural-Language-Guided Binder Shortlisting Team, B. A. A., Chen, X., Zhang, Y., Lu, C., Ma, W., Guan, J., Gong, C., Yang, J., Zhang, H., Zhang, K., et al. Protenix-advancing structure prediction through a comprehensive alphafold3 reproduction. BioRxiv, p. 2025â01, 2025a. Team, P., Ren, M., Sun, J., Guan, J., Liu, C., Gong, C., Wang, Y., Wang, L., Cai, Q., Ma, W., et al. Pxdesign: Fast, modular, and accurate de novo design of protein binders. bioRxiv, p. 2025â08, 2025b. Watson, J. L., Juergens, D., Bennett, N. R., Trippe, B. L., Yim, J., Eisenach, H. E., Ahern, W., Borst, A. J., Ragotte, R. J., Milles, L. F., et al. De novo design of protein struc- ture and function with rfdiffusion. Nature, 620(7976): 1089â1100, 2023. Zhu, W., Shenoy, A., Kundrotas, P., and Elofsson, A. Eval- uation of alphafold-multimer prediction on multi-chain protein complexes. Bioinformatics, 39(7):btad424, 2023. 10 Natural-Language-Guided Binder Shortlisting A. Dataset Preprocessing Candidate pools are defined at the target level before fea- ture extraction. A design is included only when the public release provides a candidate amino-acid sequence, a tar- get sequence, and a candidate-level experimental binding outcome. Entries without an explicit binding measurement are treated as unlabeled rather than as negatives and are ex- cluded from recall-based evaluation; this criterion removes 7 Adaptyv EGFR round-1 submissions. For ProteinBase- style releases, binding outcomes are read from the release- provided evaluation records. A design is labeled positive if any intended-target binding record is positive and negative if all intended-target binding records are negative. Expression measurements and binding-strength annotations are retained as metadata, but they do not define the binary label. Target identifiers are assigned according to the experimental assay target reported by each source. Single-target competi- tions, including Adaptyv EGFR R1/R2 and Nipah, use the competition target, and off-target or control assay records are not used for the binary intended-target label. Bind- Craft1 revalidation candidates are assigned by their assay target inevaluations. For BoltzGen, PDB-like struc- tural seed identifiers are mapped to the released biological assay target rather than used directly as target names; for example,1g13,3apu,2a1x, and3qkgcorrespond to GM2A, ORM2, PHYH, and AMBP, respectively. Source-level pools are retained for audit and feature- extraction checks, including BoltzGen targets with no re- leased positives. Targets with zero positives (GM2A, ORM2, and TNF-Îą) are excluded from recall-based evaluation be- cause Recall@10 is undefined when the target-level positive denominator is zero. When multiple releases contain the same biological target, the evaluation pool merges those releases and de-duplicates exact candidate amino-acid se- quences. If duplicate sequences have discordant labels, the merged label is positive if any duplicate record is positive. For EGFR, this merge combines Adaptyv R1, Adaptyv R2, and BindCraft1 revalidation into 605 unique candidate se- quences from 615 labeled candidate records, with 68 posi- tives after positive-if-any label aggregation. B. Temporal Leakage Audit Table 6 records the documented model cutoffs used for the temporal-leakage audit, and Table 7 records the public- release dates used for the dataset-level argument. Develop- ment entries are included for auditability because their labels are used in supervised fitting, global-policy construction, or iterative policy-search feedback; they are not held-out evaluation labels. Forgpt-4o-2024-11-20, the doc- umented cutoff is 2023-10-01, so the 10-target held-out split and the 3-target held-out subset both use labels re- leased after the cutoff. Thegpt-5.4model has a later 2025-08-31 cutoff; Nipah, RBX1, and TREM2 are the held- out targets in Table 3 whose source pools and experimen- tal labels became public after this later date. ProteinBase competition pages sometimes retain stale stage-detail text after release, so we use the overview âResults releasedâ date when available. BindCraft1 revalidation is public on ProteinBase; collection-level asset timestamps are consis- tent with 2025-10-01, but we use the 2025-10-06 Protein- Base launch as the conservative public web-availability date. BoltzGenâs manuscript was posted on bioRxiv on 2025-11- 24, but validated-label summary files were already present in the Hugging Faceboltzgen/adaptyvdata1upload on 2025-10-27, with a stable re-upload on 2025-10-31. The BoltzGen labels are explicitly supplied only as development information rather than used as held-out outcomes. Model / runDocumented cut- off Use in paper gpt-4o-2024-11-202023-10-01Reported on the 10-target held-out split and the 3-target held-out subset gpt-5.42025-08-31Reported on the 3-target held-out subset: Nipah, RBX1, and TREM2 Table 6. Model cutoff dates used for the temporal-leakage audit. Source / poolTargets usedPublic dateAfter gpt-4o? After gpt-5.4? BindCraft1 revalidationspCas9, Der f 21, IFNAR2, Der f 7, EGFR entries 2025-10-06yesyes Adaptyv EGFR R1EGFR2024-10-18yesno pMHC minibinders NY-ESO-1 pMHC, RVTDESILSY pMHC 2024-12-03yesno Adaptyv EGFR R2EGFR2025-01-15yesno BoltzGen validated-label dataset development targets only 2025-10-27yesyes GEM/Adaptyv RBX1RBX12026-04-26yesyes Nipah releaseNipah G2026-01-21yesyes BioArena/Adaptyv TREM2 TREM22026-03-28yesyes Cao/Bennett historicaltransferbaseline only 2022 to 2023nono Table 7. Dataset-level temporal-leakage audit. Dates report con- servative public availability of candidate-level experimental labels, or stable public data-package availability where explicitly noted. Development and transfer-baseline entries are shown for prove- nance but are not held-out evaluation labels. For thegpt-5.4 comparison, Nipah, TREM2, and RBX1 are the held-out targets whose source pools and experimental labels became public after the documented cutoff. C. Feature Glossary The main policy panel contains 17 reproducible candidate- level features, listed with monotonic directions and extrac- tion summaries in Table 8. Inference settings. AF2-Multimer and AF2-monomer runs use ColabFold v1.5.5 with three models, three recy- cles, and no custom templates. The single-sequence com- plex wrapper uses single-sequence MSA mode. The target- MSA complex wrapper supplies a precomputed multimer 11 Natural-Language-Guided Binder Shortlisting FeatureDirection Extraction summary AF2-Multimer ipTMhigherMean AF2-Multimer ipTM across aggre- gated target-MSA complex models. AF2-Multimer pTMhigher Mean AF2-Multimer pTM across aggre- gated target-MSA complex models. AF2 binder pLDDThigherMeanbinder-chainAF2-Multimer pLDDT across aggregated target-MSA complex models. AF2 interface PAElowerMean cross-chain AF2-Multimer PAE over target-binder interface residue pairs, averaged across models. Boltz-2 ipTMhigherBest Boltz-2 ipTM summary across target- MSA diffusion samples. Boltz-2 pTMhigher Best Boltz-2 pTM summary across target- MSA diffusion samples. Boltz-2 binder pLDDThigherBest binder-chain pLDDT summary across Boltz-2 target-MSA diffusion sam- ples. Boltz-2 ipSAEhigher Best conservative ipSAE-style interface- confidence summary across Boltz-2 target- MSA diffusion samples. Boltz-2 pDockQ2higherBest pDockQ2 summary across Boltz-2 target-MSA diffusion samples. Protenix pair ipTMhigherProtenix/PXDesign pairwise binder-target confidence under the two-chain target- binder convention. Protenix complex pTMhigher Protenix/PXDesign complex-level pTM for the predicted target-binder complex. Protenix binder ipTMhigher Protenix/PXDesign binder-chain ipTM under the two-chain target-binder conven- tion. Protenix binder pTMhigherProtenix/PXDesign binder-chain pTM un- der the two-chain target-binder conven- tion. Protenix binder pLDDThigherProtenix/PXDesign binder-chain pLDDT under the two-chain target-binder conven- tion. Rosetta interface âGlower Rosetta InterfaceAnalyzer interfaceâG on the Boltz-2 target-MSA predicted com- plex selected for Rosetta analysis. Rosetta buried SASAhigherRosetta buried solvent-accessible surface area on the Boltz-2 target-MSA predicted complex selected for Rosetta analysis. Rosetta shape complemen- tarity higher Rosetta InterfaceAnalyzer shape comple- mentarity (Lawrence & Colman, 1993) on the Boltz-2 target-MSA predicted com- plex selected for Rosetta analysis. Table 8. Active features in the 17-feature policy panel. Directions indicate the monotonic orientation used by the deterministic execu- tor. Target-MSA complex predictions use the target-chain MSA while designed binder chains remain single-sequence. A3M for the target chain, while the binder chain remains single-sequence because no homologs exist for a de novo binder. Target-chain A3M files are generated once with the Protenix/PXDesign-compatible MMseqs2 MSA service and cached before feature extraction. Boltz-2 runs use 5 diffusion samples, 3 recycling steps, 200 sampling steps, and theboltz2model, matching Boltz- Genâs de novo binder refolding configuration (Stark et al., 2025). In the single-sequence baseline, both chains are mod- eled without MSA information. In the target-MSA variant, the target chain receives the precomputed target MSA and the binder chain remains single-sequence. Protenix/PXDesign features follow the public PXDesign Protenix inference setting with one sample, two diffusion steps, four recycles, templates disabled, and MSA enabled. The target chain receives the precomputed target MSA and the binder chain remains single-sequence. Complexes are represented as target-binder pairs, and binder-chain confi- dence fields are extracted according to the PXDesign two- chain convention. The active Protenix/PXDesign policy features are pairwise binder-target ipTM, complex pTM, binder-chain ipTM, binder-chain pTM, and binder-chain pLDDT. The main 17-feature panel uses target-MSA Boltz-2, AF2- Multimer, Protenix, and Rosetta complex outputs. Single- sequence complex predictions are retained for ablation but are not part of the main panel. ESMFold (Lin et al., 2023), using the HuggingFacefacebook/esmfoldv1 weights, is deterministic and run once per binder. Rosetta InterfaceAnalyzerMover(Stranges & Kuhlman, 2013) with theref2015scorefunction (Alford et al., 2017) is applied to Boltz-2 predicted complexes after coordinate- constrainedFastRelaxpre-relaxation for side-chain packing and local relaxation. D. Cao/Bennett AF2 Transfer Baseline The historical Cao/Bennett transfer baseline is trained on the AF2 confidence measurements that have compatible counterparts in the current held-out pools: binder-target interface PAE and binder-chain pLDDT. On the historical training pools, these are the AF2 interaction PAE and binder pLDDT scores reported with the Cao/Bennett retrospective scoring data. On the current held-out pools, we recompute the analogous quantities using our AF2-Multimer target- MSA protocol: mean cross-chain PAE over target-binder interface residue pairs and mean binder-chain pLDDT, each averaged across AF2-Multimer model runs. RMSD-based historical features are excluded because the current held-out candidates do not generally include the original designed- complex reference structures needed to compute the same designed-versus-predicted RMSD terms. Thus, this base- line tests transfer across compatible AF2 confidence feature types, not an identically matched AF2 scoring protocol. E. Prompt Information Conditions This appendix records the prompt information conditions used to audit label exposure and distinguish global from target-conditioned policies. All settings use the same feature panel, policy class, and deterministic executor described in Section 5.1; held-out labels are never included in any prompt. Table 9 contrasts the single-turn global and target- conditioned settings, and iterative variants add aggregate development-search feedback without changing the held- out-label restriction. 12 Natural-Language-Guided Binder Shortlisting Prompt aspectGlobal LLMTarget-conditioned LLM Policy emittedOne policy shared by all held-out targets One policy generated sep- arately for each held-out target Feature panelSame 17 feature descrip- tions and fixed directions Same 17 feature descrip- tions and fixed directions Development informa- tion Per-development-target score distributions and target-averagedsingle- feature performance Same development cal- ibration context as the global prompt Held-out information No held-out target identi- fiers, source context, score distributions, or labels Current held-out target identifier and unlabeled score distributions for that target pool Labels available to prompt Development labels only through aggregate single- feature performance Development labels only through aggregate single- feature performance; no held-out labels Executor Same deterministic ex- ecutor: fixed directions, within-target normaliza- tion, weighted sum, top 10 Same deterministic ex- ecutor: fixed directions, within-target normaliza- tion, weighted sum, top 10 Table 9.Information contrast between global and target- conditioned single-turn LLM policies. Iterative variants keep the same held-out-label restriction and executor, but add aggregate development-search feedback. Prompt inputs and constraints. Rather than relying on free-form natural-language recommendations, each prompt asks the model to emit a structured ranking policy. The common prompt inputs are: (i) natural-language feature descriptions and fixed monotonic directions for the 17 ac- tive proxy scores, (i) development-set single-feature perfor- mance summarized as target-averaged Recall@10, Hit@10, and NDCG@10, and (i) per-development-target score- distribution summaries. Target-conditioned prompts ad- ditionally include the current held-out target identifier and unlabeled score-distribution summaries for that held-out target pool. Global prompts omit held-out target identifiers and held-out score distributions, and ask for one policy that is later applied unchanged to every held-out target. All prompts enforce the same policy constraints. A valid policy selects 3 to 5 features, assigns each selected feature an integer weight in1, 2, 3, and returns only the struc- tured policy object. The prompt explicitly disallows hard thresholds, feature-specific filters, treating any proxy as a direct affinity measurement, or overriding feature directions and aggregation. The deterministic executor, not the LLM, applies the fixed feature directions, normalizes selected fea- tures within the target pool, computes the weighted normal- ized sum, sorts candidates in descending order, and selects the top 10 designs. Iterative feedback. Iterative settings use the same infor- mation restrictions as their single-turn counterparts. They add only aggregate development-search feedback from pre- viously evaluated policies, including accepted or rejected feature-weight combinations and development-set summary metrics. This feedback is computed on the development split and does not include held-out labels. Representative output schema. The saved runs used a JSON-like structured output with one list of selected terms: "score_terms": [ "feature": "<feature_name>", "weight": 1|2|3 ], "rationale": "<brief policy rationale>" F. Per-target Performance Table 10 reports per-target Recall@10 for the strongest single-feature fixed heuristic and the matched-prompt gpt-4opolicy settings. Table 11 reports the correspond- inggpt-5.4detailed metrics only on the 3-target held-out subset used for gpt-5.4 evaluation. Targetn n b Fixed GlobalTCIter. global Iter. TC spCas92014 0.7000.8000.9000.8600.840 Der f 21164 0.5000.2500.5000.5500.600 Nipah G1030 103 0.7000.6000.8000.7600.720 IFNAR2208 0.5000.5000.5000.7250.725 NY-ESO-1 pMHC411 1.0001.0000.2001.0001.000 Der f 7205 0.8000.2000.6000.7600.800 EGFR60568 0.2000.4000.4000.3800.280 RVTDESILSY pMHC962 0.5000.0000.4000.1000.100 RBX13219 0.1110.0000.0000.1110.111 TREM210037 0.7000.6000.6800.6400.660 mean0.5710.4350.498 0.5890.584 Table 10. Per-target Recall@10 on the 10-target held-out split forgpt-4o. Fixed is Protenix binder ipTM, the strongest single- feature fixed baseline in Table 3. Global and TC denote single- turn global and target-conditioned matched-prompt policies; Iter. global and Iter. TC denote the corresponding iterative policies with development-feedback memory. LLM entries are sampled-policy averages where applicable. 13 Natural-Language-Guided Binder Shortlisting MetricTargetGlobalTC Iter. global Iter. TC Recall@10Nipah G0.800 0.8500.5800.640 Recall@10RBX10.000 0.0560.1110.156 Recall@10TREM20.600 0.6000.7200.760 Recall@10mean0.467 0.5020.470 0.519 Hit@10Nipah G0.800 0.8500.5800.640 Hit@10RBX10.000 0.0500.1000.140 Hit@10TREM20.600 0.6000.7200.760 Hit@10mean0.467 0.5000.467 0.513 NDCG@10 Nipah G0.849 0.8690.6200.734 NDCG@10 RBX10.000 0.0370.1410.222 NDCG@10 TREM20.637 0.6350.7770.792 NDCG@10 mean0.495 0.5140.513 0.583 Table 11. Per-targetgpt-5.4performance on the 3-target held- out subset. Global and TC denote single-turn global and target- conditioned matched-prompt policies; Iter. global and Iter. TC denote the corresponding iterative policies with development- feedback memory. Values are reported only for Nipah, RBX1, and TREM2, the post-cutoff subset used forgpt-5.4evaluation. 14