Paper deep dive
Protein Design with Agent Rosetta: A Case Study for Specialized Scientific Agents
Jacopo Teneggi, S. M. Bargeen A. Turzo, Tanya Marwah, Alberto Bietti, P. Douglas Renfrew, Vikram Khipple Mulligan, Siavash Golkar
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 5:27:30 AM
Summary
Agent Rosetta is an autonomous LLM-based agent framework designed to perform complex protein design tasks by interacting with the Rosetta Macromolecular Modeling Suite. It addresses the limitations of current machine learning models, which are often restricted to canonical amino acids, by leveraging Rosetta's physics-based energy functions. The system uses a multi-turn interaction loop with a structured environment to refine protein designs, demonstrating that environment design and tool abstraction are critical for enabling LLMs to reliably execute specialized scientific software tasks.
Entities (5)
Relation Signals (3)
Agent Rosetta â operates â Rosetta
confidence 100% · Agent Rosetta, an LLM agent paired with a structured environment for operating Rosetta
Agent Rosetta â uses â RosettaScripts
confidence 98% · Agent Rosetta... progressively refines protein designs through structured dialogue and feedback with an OpenAI Gym-like RosettaScripts environment
Agent Rosetta â comparedwith â ProteinMPNN
confidence 95% · We compare the performance of the agent with ProteinMPNN
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) are capable of emulating reasoning and using tools, creating opportunities for autonomous agents that execute complex scientific tasks. Protein design provides a natural testbed: although machine learning (ML) methods achieve strong results, these are largely restricted to canonical amino acids and narrow objectives, leaving unfilled need for a generalist tool for broad design pipelines. We introduce Agent Rosetta, an LLM agent paired with a structured environment for operating Rosetta, the leading physics-based heteropolymer design software, capable of modeling non-canonical building blocks and geometries. Agent Rosetta iteratively refines designs to achieve user-defined objectives, combining LLM reasoning with Rosetta's generality. We evaluate Agent Rosetta on design with canonical amino acids, matching specialized models and expert baselines, and with non-canonical residues -- where ML approaches fail -- achieving comparable performance. Critically, prompt engineering alone often fails to generate Rosetta actions, demonstrating that environment design is essential for integrating LLM agents with specialized software. Our results show that properly designed environments enable LLM agents to make scientific software accessible while matching specialized tools and human experts.
Tags
Links
- Source: https://arxiv.org/abs/2603.15952v1
- Canonical: https://arxiv.org/abs/2603.15952v1
Trouble viewing inline? Open PDF directly â
Full Text
171,426 characters extracted from source content.
Expand or collapse full text
Protein Design with Agent Rosetta: A Case Study for Specialized Scientific Agents Jacopo Teneggi â1,6 , S.M. Bargeen A. Turzo â2 , Tanya Marwah 3 , Alberto Bietti 1,4 , P. Douglas Renfrew 2 , Vikram Khipple Mulligan 2 , Siavash Golkar 1,5 1 Polymathic AI Collaboration 2 Center for Computational Biology, Flatiron Institute 3 Simons Foundation 4 Center for Computational Mathematics, Flatiron Institute 5 New York University 6 Johns Hopkins University Abstract Large language models (LLMs) are capable of emulating reasoning and using tools, creating opportunities for autonomous agents that execute complex scientific tasks. Protein design provides a natural testbed: although machine learning (ML) methods achieve strong results, these are largely restricted to canonical amino acids and narrow objectives, leaving unfilled need for a generalist tool for broad design pipelines. We introduce Agent Rosetta, an LLM agent paired with a structured environment for operating Rosetta, the leading physics-based heteropolymer design software, capable of modeling non-canonical building blocks and geometries. Agent Rosetta iteratively refines designs to achieve user-defined objectives, combining LLM reasoning with Rosettaâs generality. We evaluate Agent Rosetta on design with canonical amino acids, matching specialized models and expert baselines, and with non-canonical residuesâwhere ML approaches failâachieving comparable performance. Critically, prompt engineering alone often fails to generate Rosetta actions, demonstrating that environment design is essential for integrating LLM agents with specialized software. Our results show that properly designed environments enable LLM agents to make scientific software accessible while matching specialized tools and human experts. 1 Introduction Autonomous agents built on large language models (LLMs) are becoming increasingly capable of executing complex, multi-turn tasks that demand the emulation of reasoning and the capability to use tools. A key strength of these agents lies in their ability to write, debug, and execute code [27,102,49], making them well-suited for aiding the automation of scientific discovery workflows that rely on code-based interfaces, spanning chemistry [20,10], mathematics [107], physics [73,80], and biology [57,43,97,82]. By combining artificial reasoning with iterative feedback, agents offer new ways to address scientific problems beyond the reach of current machine learning (ML) methods. Protein designâand, more generally, heteropolymer designâis a central scientific challenge, with broad implications for developing nanomaterials [63,62,56,45], medically- or industrially-relevant enzymes [46,96,110,59], and drugs [89,54,84]. While ML-based protein modeling methods such as AlphaFold [61,1], RFdiffusion [124], and ProteinMPNN [33] have advanced the field, they focus on the 20 canonical amino acids, are specialized to particular design pipelines, and depend on large training datasets. In these models, small modifications of the task (e.g., inversion of a chiral center) often result in poor performance, for lack of relevant training data such as non-canonicalD-amino acid peptides [28]. In contrast, the Rosetta Macromolecular Modeling Suite [68], built on a largely physics-based energy function [8], requires minimal training data and can accurately model proteins 1 arXiv:2603.15952v1 [cs.AI] 16 Mar 2026 Protein Design with Agent Rosetta containing non-canonical building blocks that have not been observed in experimentally-determined structures [100,38,81,15,88,123,99,32]. This dramatically expands the accessible design space and has enabled advances in, for example, drug development [53,16,89,54,84]. Yet effective use of Rosetta demands not only deep biophysical expertise and coding proficiency, but also familiarity with its unconventional input formats, such as RosettaScripts [41]. Because these barriers limit accessibility, a major aspect of computational macromolecule design [85], hybrid approaches that combine ML models with general physics-based methods could be transformative. Here, we identify an opportunity for LLM-based agents: by combining multi-turn reasoning with code generation and tool use, such agents could accelerate Rosetta protocol development even for advanced users, while making Rosettaâs powerful capabilities available to non-expert users in the broader scientific community. We develop Agent Rosetta, a single-agent, multi-turn agentic framework that follows a user-defined brief and progressively refines protein designs through structured dialogue and feedback with an OpenAI Gym-like RosettaScripts environment [21,41]. Our approach leverages the biophysical knowledge and code-writing capabilities of the base LLM through artificial reasoning, and combines it with the constraints of XML scripting in Rosetta. We find this framework an interesting test case for building autonomous agents aimed at real-world scientific workloads. For example, we find that even though many examples of RosettaScripts are included in LLM training corpora, the rich and complex RosettaScripts syntaxâunlike that of mainstream software packagesâcreates many challenges. In specific cases, we found that these difficulties proved significant enough that prompt engineering alone was insufficient to reliably extract scientifically sound outputs from frontier LLMs. We evaluate Agent Rosetta on two real-world protein design pipelines: designing sequences that stabilize polypeptide backbone conformations with canonical amino acids only, and inclusion of a non-canonical amino acid (NCAA) in the core of a given protein. For the former task, we compare the performance of the agent with ProteinMPNN [33]âan ML model trained specifically on fixed-backbone sequence designâand with two human baselines of varying RosettaScripts complexity. On the latter task, given ProteinMPNNâs restriction to canonical residues, we can only compare the agent with a human baseline, and use AlphaFold 3 (AF3), which supports some post-translational modifications of canonical amino acids, for final validation. We briefly summarize the main contributions of this work: âąWe introduce Agent Rosetta, an LLM agent capable of executing broad user-defined tasks via multi- turn interactions with a tailored RosettaScripts environment that ensures a robust integration of the agent with Rosetta. âą In building our agent, we found that prompting alone can fail to bridge general-purpose LLMs with specialized scientific software. Instead, tool and environment design are crucial. âąWe demonstrate that our framework enables a broad range of real-world design pipelines. In particular, we evaluate Agent Rosetta on fixed-backbone sequence design with canonical amino acids only and on inclusion of a non-canonical residue in the core of a protein. âąWe compare Agent Rosetta with ProteinMPNN and human written protocols. Validation with ESMFold and AF3 confirmed that Agent Rosetta is competitive with specialized ML models and biomolecular scientists. This work represents a step towards understanding the merits and limitations of LLM-based agents for scientific tasks. Our results demonstrate that coupling frontier LLMs with domain-specific scientific tools can yield flexible frameworks that achieve performance competitive with deep learning models trained for narrowly defined design tasks and human scientists. This suggests that generalist agentic approaches can combine accuracy with versatility, opening new opportunities for scientific discovery. 2 Protein Design with Agent Rosetta state with summary of previous actions Agent Rosetta: reasoning and next action selection chosen action documentation Agent Rosetta: reasoning and action parameter generation scientist RosettaScripts environment Agent Rosetta design refinement action script Agent Rosetta error correction multi-turn interaction feedback Action succeeds Action fails new state error message initial state design brief (A) System design (B) Agent Rosetta design refinement Figure 1: Illustration of our multi-turn agentic system. (A) Schematics of Agent Rosettaâs interaction protocol. (B) Design refinement: the agent chooses the action, and, after the environment returns the action documentation, it generates the action call with its parameters. We refer readers to Section A for a detailed discussion of prior works on scientific agents, ML for protein design, and heteropolymer design with Rosetta. 2 Agent Rosetta The Rosetta modeling suite [68,70] is a general collection of C++ libraries and programs supporting a wide range of applications such as heteropolymer structure prediction, docking, and design. A central component of Rosettaâs methodology is the use of Monte Carlo algorithms to optimize an energy function that approximates the true energy of a molecular conformation [8]. Optimization is performed by exploring conformation space (sampling favorable candidate folds or docked configurations given a fixed sequence of chemical building blocks) or sequence space (sampling favorable sequences of chemical building blocks given a fixed heteropolymer backbone conformation) [19,66]. To perform this exploration, Rosetta defines Movers, actions that alter a structure to guide the sampling process and navigate the energy landscape, as well as Filters, modules that measure a structureâs property and decide whether to continue or to start over. Scientists compose design protocols with multiple Movers and Filters that often require hours to days to complete, depending on polymer length, structural complexity, and exhaustiveness of exploration of the optimization space. As a consequence, researchers must commit to execute entire protocols before assessing their success, potentially foregoing valuable information in intermediate states. In contrast, LLM agents can interact with Rosetta more frequently, selecting the next Mover immediately after the previous one completes. This enables adaptive guidance of the design process informed by the properties of intermediate structures. When randomized across parallel runs, this approach also facilitates a broader and more diverse exploration of protocols than is typically feasible for scientists. To realize these advantages, we introduce Agent Rosetta, an agentic framework that successively determines the best action, as illustrated in Fig. 1. We now describe Agent Rosettaâs system design and the main components of the multi-turn interaction protocol. We include all prompts in Section B. System prompt. Given the online presence of many Rosetta scientific papers, much Rosetta lecture material, and considerable Rosetta software documentation, it is clear that frontier LLMs have had exposure to Rosetta and RosettaScripts during training. Hence, in the system prompt, we directly instruct the agent to act as a RosettaScripts coder supporting a team of scientists. We include a summary of the available actions in Python-style docstrings format, and the rules of the interaction protocol and of the structured response format (see Section B.1 for the system prompt). 3 Protein Design with Agent Rosetta Design brief. Rosetta can be used for a variety of tasks. For example, it can modify the amino acid sequence of natural proteins to alter their function [117,83,116], construct antibodies that can bind to particular antigens [5,104], or design exotic, synthetic peptides that can bind to target proteins of therapeutic interest [89,54]. Agent Rosetta is instructed to follow a design brief that states these goals in terms of both quantitative and qualitative objectives. For example, an entry-level design brief could be to âReplace all amino acids in ubiquitinâs sequence so that its core is well-packed, the protein structure is energetically stable, and it adopts the same fold as the initial proteinâ. Together with the design brief, the user defines an initial structure: in our example, this would be the Protein Data Bank (PDB) file 1UBQ [120] describing the 3D conformation of ubiquitin. In other tasks, such as de novo design [64], the agent could first generate the initial backbone conformation with internal Rosetta modeling tools or generative ML models like RFdiffusion, and then proceed to designing a stabilizing amino acid sequence for that conformation. See Section B.2 for the design brief prompt, and Sections C.1 and D.1 for the briefs used in our experiments. Environment state and task metrics. The Monte Carlo search algorithms in Rosetta make design protocols stochastic. To account for this, at each step of a trajectory, we execute Agent Rosettaâs action across several parallel processes, generating an ensemble of designs rather than a single one. Each design is stored as a data structure called a Pose [68], which can be written to a PDB file. Pose objects are complex, and PDB files can easily contain hundreds of lines each, making them impractical as raw context for the agentâs LLMâespecially for ensembles. Instead, we make this state legible to the agent with task-dependent surrogate metrics that approximate the energetic and structural stability of the candidate designs. The choice of metrics is not unique, as stability remains difficult to quantify without experimental validation. In this work, we pick metrics that are representative of the way scientists use Rosetta. In particular, to estimate packing of the hydrophobic core (a key feature of stable proteins), we compute radius of gyration [68], cavity volume [108], and buried unsatisfied hydrogen bonds [89]. These metrics help identify anomalies in protein structure modeling or design protocols. Furthermore, we identify common problematic residue identities and positions by analyzing the per-residue van der Waals energy (i.e., thefarepterm in Rosettaâs energy function) and the favorability of each residueâs combination ofÏandÏmainchain dihedral angles given its Ramachandran map (i.e., theramapreproterm). These terms are used by researchers to detect side-chain clashes and strained backbone conformations, respectively. This way, the agent knows the most common problematic residues. Finally, we enrich the state with task-specific progress metrics. For example, for fixed-backbone sequence design with canonical amino acids only, we use ESMFold to predict the native folds of the candidate designs, and we compute the root mean squared deviation (RMSD) between the predicted fold and the initial backbone conformation. Since the goal is to design a sequence that gives rise to the desired fold, a successful design should yield a small RMSD between the desired structure and the structure predicted from the sequence. We complement the RMSD with the average pLDDT of the C α atoms: values close to 1 indicate high confidence in ESMFoldâs prediction, providing further evidence for stability. We include the RMSD and pLDDT in the state as easy-to-compute but imperfect proxies for stability. We note, however, that ESMFold (like many ML models) may not reliably capture the stability of Rosetta designs that deviate from its training set of naturalistic sequences. For example states see Sections C.3 and D.4. Available actions. RosettaScripts accepts XML scripts as structured actions, but having the agent write them from scratch leads to lengthy, error-prone responses that are costly to correct. Instead, we use RosettaScriptsâ XML schema to define types of actions, whose templates are filled with the parameters generated by the agent. Not only does this allocate Agent Rosettaâs reasoning budget 4 Protein Design with Agent Rosetta towards the scientifically-relevant parameters, but also it guarantees semantically correct actions where prompting alone fails. We define three types of actions that fit several design pipelines and allow the agent to explore a vast sequence space (see Section E for an example of each action): âą rotamerchangeuses Rosettaâs FastDesign Mover [15] to sample energy-minimizing side chains of a molecule, relaxing the backbone conformation in the process. The agent guides FastDesignâs sequence optimization by generating amino acid composition constraints that are incorporated into Rosettaâs aacompositionscoring term during design [53,86]. Withaacomposition, scientists can realize actions such as âDesign the core of a protein with at most 15% polar residues and at most 5% glycine. Do not allow any hydrophilic residues in the core, and keep the boundary of the protein unchangedâ. Beyond composition constraints, Agent Rosetta can restrict the base residue types allowed during design with TaskOperations, or disable design completely, leaving repacking with identity fixed. Finally, the agent can compose penalty blocks and TaskOperations in different regions of a molecule using ResidueSelectors, Movers that select residues by identity, burial, secondary structure, or other properties. âą backbonechangeimplements three RosettaScripts Movers that randomly perturb the torsion angles of the backbone to modify its conformation:Small,Shear, andBackrub[112]. Agent Rosetta can use backbone changes to discover nearby conformations that are compatible with much more stabilizing sequences. The agent generates the parameters specific to each Mover, and ResidueSelectors. âą gobacktoreverts to a previous step in the trajectory, which Agent Rosetta may do in case the design has diverged too far from the task objectives. Multi-turn interaction workflow. Agent Rosetta refines designs following these steps: 1.Structured reasoning and action selection: First, the agent receives the current state of the environment and a summary of the previous actions. Summarization is important to prevent the context of the LLM from saturating over long trajectories [134,91,109,114]. We use a tabular summary of all previous actions and their effects on key design metrics so that the LLMâs context window contains the last 2 interaction steps only. Then, we prompt Agent Rosetta to reflect on the progress and choose the next action. We explicitly instruct the agent to take into consideration both energetic and structural information, the trajectory history, and to write an action plan that âexpert scientists can understand, review, and criticizeâ. Structured reasoning is important to bring forth the biophysical knowledge encoded in the LLM, and to allow scientists to verify the coherence of the actions with the objectives of the design brief. 2.Structured reasoning and parameter generation: After the agent selects the action, the environment responds with the necessary RosettaScripts documentation. This ensures access to syntax and parameter information, which differs from the majority of code LLMs are trained on, and it varies significantly across types of actions. Instead of the online RosettaScriptâs documentation, we wrote ad-hoc summaries to clarify the aspects the agent struggles the most with. We append the documentation to the context of the LLM, and we prompt the agent to reason again and generate the full action call with its parameters. We use the same structured reasoning steps as for action selection. Repeated reasoning is important at this stage because the original plan of the agentâgenerated without documentationâmay not be implementable. 3.Action parsing, execution, and feedback: We fill the XML template with the generated parameters and execute the action. At each step, we only refine the Pareto-optimal candidates from the previous round according to task-dependent quality metrics. If the action did not match the 5 Protein Design with Agent Rosetta Figure 2:A failure example of prompting for generation of composition penalties. Even though Agent Rosetta wants to reduce proline content, the penalty block achieves the opposite effect. 10 4 10 3 10 2 Cost per response (USD) 0.2 0.4 0.6 0.8 1.0 Success rate syntax original simple reasoning tokens 1000 2000 3000 Figure 3: Comparison of the performance of differ- ent LLMs at generating amino acid compositional penalty blocks with the original RosettaScripts syntax and our simplified syntax. expected response format, it did not satisfy RosettaScripts syntax, or it encountered runtime errors, we prompt the agent with the error message to fix its mistakes. Otherwise, we update the state and the agent proceeds to the next refinement turn. We include prompts for all steps of the refinement loop in Section B.3, and error correction examples in Section F. 3 Results Scientists can use Agent Rosetta to perform a broad range of tasks with both canonical and non-canonical residues. We evaluate the agent on stabilizing fixed backbone conformations using canonical amino acids only, and on insertion of 1 non-canonical amino acid (NCAA) in the core of a protein. The former task enables rigorous comparison with existing baselines, and the latter realizes Agent Rosettaâs capability for usage cases beyond the reach of current ML approaches. In our experiments, we evaluated Qwen3 Instruct [129], GPT-OSS-20B,4.1,5[6,4,111], Gemini 2.5 Flash [29], and Claude Sonnet 4.5 [9]. 1 3.1 The Importance of Environment Design Rosetta and RosettaScripts have extensive online presence, and information about Rosetta is certainly included in the training set of frontier LLMs. This suggests that prompting should suffice to bring forth the necessary knowledge. Instead, we found that LLMs know general facts about Rosetta and RosettaScripts, but they regularly fail to generate semantically correct actions with prompting alone. For example, consider the composition penalties in therotamerchangeaction: they follow a particular syntax to select residue types, a target fractional or absolute range, and a sequence of penalties for deviations from the desired range. Furthermore, the number of penalties changes with the width of the target range, and whether the target composition is fractional or absolute. All these factors are both counter-intuitive to first-time RosettaScripts users, and also sources of error for the agent, even with plausible reasoning. 1 All models were accessed using OpenRouter. 6 Protein Design with Agent Rosetta Fig. 2 includes an example from one of our experiments. Agent Rosetta correctly recognizes that excessive proline content may be problematic, and wants to penalize more than 5. However, the lines DELTAEND 0andPENALTIES 0 0 0 0 -10 -20implement the opposite of the agentâs intent (they favor more than 4 prolines). We found that extending prompting instructions only yielded minor improvements. Furthermore, this type of error is difficult to recover from in a multi-turn environment as the agent often fails to recognize that it made a mistake. To solve these issues, we abstracted the syntax of composition penalties and TaskOperations to simplify their use (see Section G). Then, we wrote simple Python scripts to convert this simplified syntax into its RosettaScripts equivalent. In Fig. 3, we compare the performance of different LLMs at generating composition penalties with the original RosettaScripts syntax versus our simplified one. We compiled a list of 9 prompts with the most commonly-used penalty shapes: above, below, or outside a specified target and range (e.g., âWrite a composition penalty that penalizes more than 5 alanines. Use a linear boundary with slope of 10.â). We averaged results over 10 generations per prompt. For each syntax type, we include the rules in the system prompt, and we verify both the syntactic correctness and the semantics of the responses (see Section H for prompts and responses). We found that both closed- and open-source models fail at reliably generating penalty blocks with the original RosettaScripts syntax, with GPT-5 (medium reasoning effort) achieving the best performance of â70%. We found this value to be insufficient for developing an autonomous agent capable of sustaining multi-turn interactions, especially since generating penalty blocks is an essential component of design in Rosetta. Not only does our simplified syntax reduce costs by orders of magnitude, but also it guarantees robust performance (in this subtask) across all models, irrespective of size and cost (â„98.88% success rate). Building autonomous agents using open-source models, while minimizing monetary and energetic costs, is paramount for democratizing scientific software. These findings stress the crucial role of environment design to bridge general-purpose LLMs with scientific tools. 3.2 Stabilizing Backbone Conformations In order to compare with ML baselines, we evaluate Agent Rosetta on the task of designing sequences to stabilize protein conformations, using canonical amino acids only. The initial state of the trajectory is a poly-glycine backbone in a target conformation, and the task brief instructs the agent to change the side chains so that the molecule is energetically stable and it adopts the same fold as the target (the design brief is included in Section C.1). As target conformations, we chose 8 PDB structuresâboth natural and syntheticâof length between 74 and 125 residues. Section C.2 includes the selection criteria, preprocessing steps, and images of the backbones. We run the agent for a total of 30 LLM queries, and we configure the environment such that each refinement step generates an ensemble of 128 candidate designs. We select the Pareto front of the ensembles according to the RMSD between the predicted fold (calculated via ESMFold [76]) and the target conformation, the ESMFold pLDDT, and total Rosetta ref2015energy [8]. We chose these proxies for the quality of the designs because their low computational cost allows us to compute them after each step of the design process. A low RMSD means that the predicted fold aligns with the target conformation, a high pLDDT conveys good confidence of ESMFold in its prediction, and low Rosetta energy suggests stability in the target conformation. The selected candidates are then used as the initial states of the next step of the agent. Comparison methods. We compare the designs generated by Agent Rosetta with ProteinMPNN [33], two expert-devised Rosetta protocols, and a na Ìıve baseline. ProteinMPNN is a generative model that was trained on known protein sequences and structures to approximate the conditional distribution of amino acid sequences given an input target conformation. 7 Protein Design with Agent Rosetta Qwen3 Instruct (CoT) Gemini 2.5 Flash GPT-5 Sonnet 4.5 ProteinMPNN human baseline (one shot) human baseline (staged) 0.75 0.80 0.85 0.90 ESMFold RMSD to Init (Ă ) 151015 Design step PDB: 9PL1 (74 residues) 151015 Design step 0.4 0.5 0.6 0.7 0.8 0.9 PDB: 9C14 (97 residues) Figure 4: Comparison of best running RMSD as a function of design step for 2 target backbone con- formations. On 9PL1, Agent Rosetta outperforms competing methods, and on 9C14 it helps close the gap between the human written protocols and ProteinMPNN. 86% 90% fraction of reasoning tokens 65% 70% 90% action success rate 0.00.20.40.6 Cost (USD) 0.6 0.7 0.8 0.9 1.0 1.1 ESMFold RMSD to Init (Ă ) Figure 5: Summary of results for stabilizing back- bone conformations with canonical amino acids only. We report the average cost of one run of 30 model queries, the average action success rate, and the fraction of output tokens that were reasoning. For comparison with human experts, we consider two hand-written protocols composed of the same blocks as the agentâs: FastDesign and backbone Movers with ResidueSelectors and TaskOperations. The first is a one-shot protocol that applies FastDesign to sample the entire sequence at once. The second is a staged protocol that first designs the core, then the boundary, and finally the surface of the moleculeâa common practice to reduce the size of the sequence space that FastDesign must explore (see Section C.4 for the protocols). We stress that the human written protocols are fixed and independent of the given backbone. In practice, scientists fine-tune them after analyzing rounds of results. However, this process often requires multiple days, making it prohibitive for systematic evaluation across several structures. Therefore, we chose baseline protocols that would reflect a scientistâs first few days of work on the same task, using the Rosetta tools available to the agent. Finally, we include a na Ìıve sequential baseline that iteratively runs FastDesign on the Pareto optimal designs, without the use of any ResidueSelectors, TaskOperations, or composition penaltiesâsimply minimizing Rosetta energy. This baseline, in which no changes are made from turn to turn, is useful to assess the impact of modern LLMs making decisions about how to alter protocols as they perform multi-turn scientific tasks. We generated ensembles of 128 designs for all comparison methods per trial. 2 Evaluation. We ran 16 independent trials per method and filtered designs with pLDDTâ„0.85 to remove outliers. For each design step, we computed the 5 th percentile (lower is better) of the ESMFold RMSD and the 95 th percentile (higher is better) of the pLDDT to summarize the quality of the filtered ensembles. We simulated best-of-n scaling over 1,000 bootstrap samples of 8 trials from the 16, and within each sample, we selected the trial with the smallest ESMFold RMSD. Fig. 4 shows the best running RMSD (i.e., over all previous steps) across bootstrap samples as a function of design step for 2 representative backbones out of the 8 used in this experiment. We note that ProteinMPNN and the human baselines are shown as horizontal lines because they are single-step methods. On 9PL1 [79], Agent Rosetta, through iterative refinement, outperforms all comparison methods, whereas on 9C14 [34], the agent helps close the gap between the human written protocols and ProteinMPNN. Fig. C.2 includes figures for all backbones. We summarize results by computing the median RMSD across the 1,000 selected best steps, and average across all 8 backbones, as illustrated in Fig. 5. For Agent Rosetta, we show results as a function of 2 For Agent Rosetta, a trial refers to a full run of the agent. For Rosetta-based baselines, this is a single run of the protocol. In both cases, we configured Rosetta to generate an ensemble of 128 designs. For fair comparison, we take a single trial of ProteinMPNN to be 128 independent generations. 8 Protein Design with Agent Rosetta Figure 6: Example inclusion of TRF (colored in magenta) in a given protein. (left) Favorable orientation towards the core, (right) unfavorable orientation exposed to the solvent. average cost of one trajectory with different LLMs. We note that, compared to agents, ProteinMPNN has a negligible computational cost, an advantage of specialized ML models over general-purpose LLMs. Furthermore, we include the RosettaScripts action success rate and the portion of output tokens that were reasoning. With our tailored environment, all LLMs achieve action success ratesâ„86%. We found that the agentâs designs are on par with ProteinMPNN (within a 0.20 Ì A tolerance), with Gemini 2.5 Flash and Qwen3 Instruct achieving the best cost-performance tradeoff. From the detailed tabular results in Tables C.1 and C.2, we found that our agent generates designs with lower RMSD for 5 out of the 8 conformations, but ProteinMPNN usually achieves higher pLDDT (see Fig. C.3). However, note that the RMSD and pLDDT calculated by ESMFold are biased towards naturalistic protein sequences which favor the designs generated by data-driven methods such as ProteinMPNN over those of physics-based methods like Rosetta. Finally, Fig. C.4 shows the best design for each method on every backbone. Our findings highlight the absence of a single dominating design methodology. 3.3 Designing an NCAA into the Core of a Protein Non-canonical amino acids (NCAAs) are important for real-world applications of protein design, but their sparsity limits data-driven approaches. Rosettaâs physics-based algorithms, on the other hand, support flexible palettes that can incorporate nonstandard residues. To investigate Agent Rosettaâs capabilities beyond the reach of current ML models, we explore the task of inserting an NCAA into the core of an existing protein. The initial state of the trajectory is a PDB structure, and the task brief instructs the agent to include exactly 1 NCAA residue in the core of the protein without compromising its fold and energetic stability (see Section D.1 for the task brief). In this experiment, we consider N1-formyl-tryptophan (TRF), a post-translational modification of tryptophan with a formyl group attached to the indole ring. We chose TRF because it is rare in the PDB (it appears in 2 calmodulin complexes only) yet compatible with AF3 [1] as a proxy for structural stability. 3 We chose 4 de novo protein folds of 40 to 153 residues to evaluate the agent across structures of varying rigidity, and report selection criteria, preprocessing steps, and images in Section D.3. Fig. 6 illustrates a good and bad placement of TRF in 8UZL [14], a designed transmembraneÎČ-barrel among the 4 selected PDBs. In the left panel (good placement), TRF is facing the core and preserves the fold. In the right (bad placement), TRF, being mostly hydrophobic, is solvent-exposed. Similar to Section 3.2, we run Agent Rosetta for a total of 30 LLM queries, and generate step-wise ensembles of 128 candidates. However, we cannot use ESMFold to predict the native fold of the designs because it does not support TRF, and AF3âs slow inference time is prohibitive for use at each refinement 3 In actual use, Agent Rosetta may be applied to designs with NCAAs incompatible with any ML model, though this would require experimental validation of the design to confirm success. As a proof of principle in this work, we confined ourselves to an NCAA for which AF3 could provide predictions for validation. 9 Protein Design with Agent Rosetta 0.700.720.740.760.780.80 NCAA inclusion rate GPT-5 Gemini 2.5 Flash Sonnet 4.5 Qwen3 Instruct (CoT) human baseline Figure 7: Median best rate of inclusion of TRF across 1,000 bootstrap samples of 8 out of 16 tra- jectories, averaged over 4 input PDBs. 0.00.20.40.6 Cost (USD) 1.0 1.5 2.0 2.5 3.0 AF3 RMSD to native (Ă ) action success rate 54% 80% 84% 90% fraction of reasoning tokens 65% 71% 80% 90% Figure 8: Summary of results for including 1 TRF residue in the core of an input protein. We report the average cost of one run with 30 model queries, the average action success rate, and the fraction of output tokens that were reasoning. step. Therefore, we select the Pareto optimal candidates by inclusion of TRF in the core, total Rosetta energy, cavity volume, radius of gyration, and RMSD between the Rosetta Pose and the native PDB. We compare the performance of our agent, using post hoc AF3 predictions to assess success, with a fixed, protein-independent human written protocol composed of the same types of actions available to the agentâsee Section D.5. We ran 16 independent design trajectories per method, and emulated best-of-n sampling with 1,000 bootstrap replicates of 8 out of the 16. Within each replicate, we selected the step with the highest rate of inclusion of exactly 1 TRF in the core. We computed the median inclusion rate of the 1,000 selected steps, and we summarize results in Fig. 7 by averaging across all 4 protein structures. We found that Agent Rosetta with GPT-5 achieves the highest success rate. In this task, structure 8UZL proved to be the most challenging: the human baseline failed to include any TRF, and GPT-5 achieved the highest rate of â 14.5%. For final validation, we filtered successful designs, selected the top 10 in order of increasing RMSD to the native structure, and predicted their fold with AF3. Fig. 8 summarizes the RMSD between the AF3 prediction and the native structure, and we include detailed tabular results in Tables D.3 and D.4. We found that Agent Rosetta outperforms the human baseline in terms of both AF3 RMSD and pLDDT (see Fig. D.3), with GPT-5 performing best overall. Across base LLMs, Qwen3 Instruct showed the largest performance drop compared to designing with canonical amino acids. Inspection of the design trajectories revealed that errors concentrated on backbonechangeactions. This is because incorporating an NCAA often requires targeted backbone perturbations to accommodate the residue without compromising the fold. Agent Rosetta can combine backbone Movers with ResidueSelectors to achieve this, but the selected residues need to define valid segments for perturbation. In many cases, the ResidueSelectors specified by Agent Rosetta with Qwen3 Instruct do not meet this criterion, causing the action to fail at runtime. Finally, in Fig. D.4, we include some example good and bad designs generated by Agent Rosetta and the expert written protocol. 4 Conclusion We have introduced Agent Rosetta, a framework that leverages the reasoning capabilities of Large Language Models to automate protein design within the Rosetta Macromolecular Modeling Suite. A central finding of our study is that prompt engineering alone is insufficient for interfacing with complex, domain-specific scripting languages; instead, success relies on a structured environment that abstracts syntactic intricacies into semantic actions. By implementing this design, we enabled the agent to 10 Protein Design with Agent Rosetta iteratively refine protocols and recover from errors that typically stall autonomous workflows while minimizing computational and economic cost. Our evaluations confirm that this approach effectively bridges the gap between generalist reasoning models and specialized scientific software. In fixed-backbone sequence design with canonical amino acids, Agent Rosetta achieves performance parity with state-of-the-art models like ProteinMPNN. Moreover, in the data-sparse regime of non-canonical amino acid design, the agent outperforms expert human baselines, tackling tasks where purely data-driven methods are currently inapplicable. Future research directions include, for example, comparison with backbone-dependent or iterative human baselines, and experimental validation of the designs generated by the agent. Our results suggest that agentic frameworks, when coupled with robust environment design, offer a powerful paradigm for unlocking the full potential of physics-based simulations in biology discovery. Acknowledgments Polymathic AI was supported by the Simons Foundation and Schmidt Sciences. PDR and VKM are wholly funded by the Simons Foundation. We thank the Flatiron Instituteâs Scientific Computing Core for ongoing support. The computations reported in this paper were performed in-part using resources made available by the Flatiron Institute. The Flatiron Institute is a division of the Simons Foundation. References [1]Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ronneberger, Lindsay Willmore, Andrew J Ballard, Joshua Bambrick, et al. Accurate structure prediction of biomolecular interactions with alphafold 3. Nature, 630(8016):493â500, 2024. [2]Tamer Abuelsaad, Deepak Akkil, Prasenjit Dey, Ashish Jagmohan, Aditya Vempaty, and Ravi Kokku. Agent-e: From autonomous web navigation to foundational design principles in agentic systems. arXiv preprint arXiv:2407.13032, 2024. [3]Atomic-Level Accuracy. Design of a novel globular protein fold with. science, 1089427(1364):302, 2003. [4] Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. [5]Jared Adolf-Bryfogle, Oleks Kalyuzhniy, Michael Kubitz, Brian D. Weitzner, Xiaozhen Hu, Yumiko Adachi, William R. Schief, and Roland L. Dunbrack Jr. RosettaAntibodyDesign (RAbD): A general framework for computational antibody design. PLOS Computational Biology, 14(4):e1006112, April 2018. Publisher: Public Library of Science. [6]Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K Arora, Yu Bai, Bowen Baker, Haiming Bao, et al. gpt-oss-120b & gpt-oss-20b model card. arXiv preprint arXiv:2508.10925, 2025. [7] Gustaf Ahdritz, Nazim Bouatta, Christina Floristean, Sachin Kadyan, Qinghui Xia, William Gerecke, Timothy J OâDonnell, Daniel Berenberg, Ian Fisk, Niccol`o Zanichelli, et al. Openfold: Re- training alphafold2 yields new insights into its learning mechanisms and capacity for generalization. Nature methods, 21(8):1514â1524, 2024. 11 Protein Design with Agent Rosetta [8]Rebecca F Alford, Andrew Leaver-Fay, Jeliazko R Jeliazkov, Matthew J OâMeara, Frank P DiMaio, Hahnbeom Park, Maxim V Shapovalov, P Douglas Renfrew, Vikram K Mulligan, Kalli Kappel, et al. The rosetta all-atom energy function for macromolecular modeling and design. Journal of chemical theory and computation, 13(6):3031â3048, 2017. [9] Anthropic. System card: Claude Opus 4 & Claude Sonnet 4, May 2025. [10]S Ìoren Arlt, Haonan Duan, Felix Li, Sang Michael Xie, Yuhuai Wu, and Mario Krenn. Meta- designing quantum experiments with language models, 2024. [11]Takumi Baba, Go Ueno, Chika Ohe, Shuku Saji, Sachiko Yamamoto, Masaki Yamamoto, Hiroshi Nakagawa, Nobuo Okazaki, Mamoru Ouchida, Iori Kawasaki Ohmori, et al. The f54l mutation of thioredoxin shows protein instability and increased fluctuations of the catalytic center. Biochimica et Biophysica Acta (BBA)-General Subjects, page 130860, 2025. [12] Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan, and Sung Ju Hwang. Researchagent: Iterative research idea generation over scientific literature with large language models. arXiv preprint arXiv:2404.07738, 2024. [13]Minkyung Baek, Frank DiMaio, Ivan Anishchenko, Justas Dauparas, Sergey Ovchinnikov, Gyu Rie Lee, Jue Wang, Qian Cong, Lisa N Kinch, R Dustin Schaeffer, et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science, 373(6557):871â876, 2021. [14]Samuel Berhanu, Sagardip Majumder, Thomas M Ìuntener, James Whitehouse, Carolin Berner, Asim K Bera, Alex Kang, Binyong Liang, Nasir Khan, Banumathi Sankaran, et al. Sculpting conducting nanopore size and shape through de novo protein design. Science, 385(6706):282â288, 2024. [15]Gaurav Bhardwaj, Vikram Khipple Mulligan, Christopher D. Bahl, Jason M. Gilmore, Peta J. Harvey, Olivier Cheneval, Garry W. Buchko, Surya V. S. R. K. Pulavarti, Quentin Kaas, Alexander Eletsky, Po-Ssu Huang, William A. Johnsen, Per Jr Greisen, Gabriel J. Rocklin, Yifan Song, Thomas W. Linsky, Andrew Watkins, Stephen A. Rettie, Xianzhong Xu, Lauren P. Carter, Richard Bonneau, James M. Olson, Evangelos Coutsias, Colin E. Correnti, Thomas Szyperski, David J. Craik, and David Baker. Accurate de novo design of hyperstable constrained peptides. Nature, 538(7625):329â335, October 2016. Publisher: Nature Publishing Group. [16]Gaurav Bhardwaj, Jacob OâConnor, Stephen Rettie, Yen-Hua Huang, Theresa A. Ramelot, Vikram Khipple Mulligan, Gizem Gokce Alpkilic, Jonathan Palmer, Asim K. Bera, Matthew J. Bick, Maddalena Di Piazza, Xinting Li, Parisa Hosseinzadeh, Timothy W. Craven, Roberto Tejero, Anna Lauko, Ryan Choi, Calina Glynn, Linlin Dong, Robert Griffin, Wesley C. van Voorhis, Jose Rodriguez, Lance Stewart, Gaetano T. Montelione, David Craik, and David Baker. Accurate de novo design of membrane-traversing macrocycles. Cell, 185(19):3520â3532.e26, September 2022. [17]Aadyot Bhatnagar, Sarthak Jain, Joel Beazer, Samuel C Curran, Alexander M Hoffnagle, Kyle Ching, Michael Martyn, Stephen Nayfach, Jeffrey A Ruffolo, and Ali Madani. Scaling unlocks broader generation and deeper functional understanding of proteins. bioRxiv, pages 2025â04, 2025. [18]Daniil A Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. Autonomous chemical research with large language models. Nature, 624(7992):570â578, 2023. [19]R. Bonneau, J. Tsai, I. Ruczinski, D. Chivian, C. Rohl, C. E. Strauss, and D. Baker. Rosetta in CASP4: progress in ab initio protein structure prediction. Proteins, Suppl 5:119â126, 2001. [20]Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. Chemcrow: Augmenting large-language models with chemistry tools. arXiv preprint arXiv:2304.05376, 2023. 12 Protein Design with Agent Rosetta [21]Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. Openai gym. 2016. eprint. arXiv preprint arXiv:1606.01540, 50, 2016. [22]Cassie M Bryan, Gabriel J Rocklin, Matthew J Bick, Alex Ford, Sonia Majri-Morrison, Ashley V Kroll, Chad J Miller, Lauren Carter, Inna Goreshnik, Alex Kang, et al. Computational design of a synthetic pd-1 agonist. Proceedings of the National Academy of Sciences, 118(29):e2102164118, 2021. [23] Chai Discovery. Chai-1: Decoding the molecular interactions of life. bioRxiv, 2024. [24] Chai Discovery. Zero-shot antibody design in a 24-well plate. bioRxiv, 2025. [25]Sidhartha Chaudhury, Sergey Lyskov, and Jeffrey J Gray. Pyrosetta: a script-based interface for implementing molecular modeling algorithms using rosetta. Bioinformatics, 26(5):689â691, 2010. [26]Hui Chen, Miao Xiong, Yujie Lu, Wei Han, Ailin Deng, Yufei He, Jiaying Wu, Yibo Li, Yue Liu, and Bryan Hooi. Mlr-bench: Evaluating ai agents on open-ended machine learning research. arXiv preprint arXiv:2505.19955, 2025. [27]Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021. [28]Henry Childs, Pei Zhou, and Bruce R. Donald. Has AlphaFold 3 Solved the Protein Folding Problem for D-Peptides? bioRxiv: The Preprint Server for Biology, page 2025.03.14.643307, March 2025. [29]Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025. [30]Gabriele Corso, Hannes St Ìark, Bowen Jing, Regina Barzilay, and Tommi Jaakkola. Diffdock: Diffusion steps, twists, and turns for molecular docking. In International Conference on Learning Representations (ICLR), 2023. [31]Francesco Costa, Ioannis Riziotis, Antonina Andreeva, Delhi Kalwan, Jennifer De Jong, Philip Hinchliffe, Fabio Parmeggiani, Paul R Race, Steven G Burston, Alex Bateman, et al. A global survey of intramolecular isopeptide bonds. Protein Science, 34(12):e70342, 2025. [32]Bobo Dang, Haifan Wu, Vikram Khipple Mulligan, Marco Mravic, Yibing Wu, Thomas Lemmin, Alexander Ford, Daniel-Adriano Silva, David Baker, and William F. DeGrado. De novo design of covalently constrained mesosize protein scaffolds with unique tertiary structures. Proceedings of the National Academy of Sciences, 114(41):10852â10857, September 2017. [33]Justas Dauparas, Ivan Anishchenko, Nathaniel Bennett, Hua Bai, Robert J Ragotte, Lukas F Milles, Basile IM Wicky, Alexis Courbet, Rob J de Haas, Neville Bethel, et al. Robust deep learningâbased protein sequence design using proteinmpnn. Science, 378(6615):49â56, 2022. [34]Yuri Rafael de Oliveira Silva, Grayson Barnes, Dia Zheng, Daniel Zhitnitsky, Samuel J Geathers, Stephen C Peters, Veronika A Szalai, John D Helmann, and Oriana S Fisher. Copper acquisition in bacillus subtilis involves cu (i) exchange between ycni and ycnj. bioRxiv, pages 2025â05, 2025. [35]Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, and Florian Tram`er. Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents. Advances in Neural Information Processing Systems, 37:82895â82920, 2024. 13 Protein Design with Agent Rosetta [36]Joseph Dodd-O, Amanda M. Acevedo-Jake, Abdul-Rahman Azizogli, Vikram Khipple Mulligan, and Vivek A. Kumar. How to Design Peptides. Methods in Molecular Biology (Clifton, N.J.), 2597:187â216, 2023. [37]Lindsey A Doyle, Brittany Takushi, Ryan D Kibler, Lukas F Milles, Carolina T Orozco, Jonathan D Jones, Sophie E Jackson, Barry L Stoddard, and Philip Bradley. De novo design of knotted tandem repeat proteins. Nature Communications, 14(1):6746, 2023. [38]Kevin Drew, P. Douglas Renfrew, Timothy W. Craven, Glenn L. Butterfoss, Fang-Chieh Chou, Sergey Lyskov, Brooke N. Bullock, Andrew Watkins, Jason W. Labonte, Michael Pacella, Kr- ishna Praneeth Kilambi, Andrew Leaver-Fay, Brian Kuhlman, Jeffrey J. Gray, Philip Bradley, Kent Kirshenbaum, Paramjit S. Arora, Rhiju Das, and Richard Bonneau. Adding Diverse Noncanonical Backbones to Rosetta: Enabling Peptidomimetic Design. PLoS ONE, 8(7):e67051, July 2013. [39]Richard Evans, Michael OâNeill, Alexander Pritzel, Natasha Antropova, Andrew Senior, Tim Green, Augustin Ë Z Ìıdek, Russ Bates, Sam Blackwell, Jason Yim, et al. Protein complex prediction with alphafold-multimer. biorxiv, pages 2021â10, 2021. [40] Noelia Ferruz, Steffen Schmidt, and Birte H Ìocker. Protgpt2 is a deep unsupervised language model for protein design. Nature communications, 13(1):4348, 2022. [41]Sarel J Fleishman, Andrew Leaver-Fay, Jacob E Corn, Eva-Maria Strauch, Sagar D Khare, Nobuyasu Koga, Justin Ashworth, Paul Murphy, Florian Richter, Gordon Lemmon, et al. Roset- tascripts: a scripting language interface to the rosetta macromolecular modeling suite. PloS one, 6(6):e20161, 2011. [42]Sujay S Gaikwad, Beena Yadav, Shubhangi Sharma, Ashwani Kumar, and RD Makde. Crystal structures of pf1765 from pyrococcus furiosus from several different crystallization conditions with varied ph, salt and precipitant. Structural Biology and Crystallization Communications, 81(12), 2025. [43]Alireza Ghafarollahi and Markus J Buehler. Protagents: protein discovery via large language model multi-agent collaborations combining physics and machine learning. Digital Discovery, 3(7):1389â1409, 2024. [44]Alireza Ghafarollahi and Markus J Buehler. Sciagents: automating scientific discovery through bioinspired multi-agent intelligent graph reasoning. Advanced Materials, 37(22):2413523, 2025. [45] Shane Gonen, Frank DiMaio, Tamir Gonen, and David Baker. Design of ordered two-dimensional ar- rays mediated by noncovalent protein-protein interfaces. Science (New York, N.Y.), 348(6241):1365â 1368, June 2015. [46]Sydney R. Gordon, Elizabeth J. Stanley, Sarah Wolf, Angus Toland, Sean J. Wu, Daniel Hadidi, Jeremy H. Mills, David Baker, Ingrid Swanson Pultz, and Justin B. Siegel. Computational design of anα-gliadin peptidase. Journal of the American Chemical Society, 134(50):20513â20520, December 2012. [47]Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, et al. Towards an ai co-scientist. arXiv preprint arXiv:2502.18864, 2025. [48] Mourad Gridach, Jay Nanavati, Khaldoun Zine El Abidine, Lenon Mendes, and Christina Mack. Agentic ai for scientific discovery: A survey of progress, challenges, and future directions. arXiv preprint arXiv:2503.08979, 2025. 14 Protein Design with Agent Rosetta [49]Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu, YK Li, et al. Deepseek-coder: When the large language model meets programmingâthe rise of code intelligence. arXiv preprint arXiv:2401.14196, 2024. [50]Sikun Guo, Amir Hassan Shariatmadari, Guangzhi Xiong, Albert Huang, Myles Kim, Corey M Williams, Stefan Bekiranov, and Aidong Zhang. Ideabench: Benchmarking large language models for research idea generation. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, pages 5888â5899, 2025. [51] Bakar Hassan, Monika Chandravanshi, Martin Y Ng, Hitendra Negi, Brice AP Wilson, and Kylie J Walters. An adaptive peptide-binding site in ubiquitin receptor hrpn13 revealed by structural studies. Nature Communications, 16(1):5669, 2025. [52]Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. Webvoyager: Building an end-to-end web agent with large multimodal models. arXiv preprint arXiv:2401.13919, 2024. [53]Parisa Hosseinzadeh, Gaurav Bhardwaj, Vikram Khipple Mulligan, Matthew D. Shortridge, Timothy W. Craven, F Ìatima Pardo-Avila, Stephen A. Rettie, David E. Kim, Daniel-Adriano Silva, Yehia M. Ibrahim, Ian K. Webb, John R. Cort, Joshua N. Adkins, Gabriele Varani, and David Baker. Comprehensive computational design of ordered peptide macrocycles. Science, 358(6369):1461â1466, December 2017. [54] Parisa Hosseinzadeh, Paris R. Watson, Timothy W. Craven, Xinting Li, Stephen Rettie, F Ìatima Pardo-Avila, Asim K. Bera, Vikram Khipple Mulligan, Peilong Lu, Alexander S. Ford, Brian D. Weitzner, Lance J. Stewart, Adam P. Moyer, Maddalena Di Piazza, Joshua G. Whalen, Per Jr Greisen, David W. Christianson, and David Baker. Anchor extension: a structure-guided approach to design cyclic peptides targeting enzyme active sites. Nature Communications, 12(1):3384, June 2021. [55]Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. Large language models for software engineering: A systematic literature review. ACM Transactions on Software Engineering and Methodology, 33(8):1â79, 2024. [56]Yang Hsia, Jacob B. Bale, Shane Gonen, Dan Shi, William Sheffler, Kimberly K. Fong, Una Nattermann, Chunfu Xu, Po-Ssu Huang, Rashmi Ravichandran, Sue Yi, Trisha N. Davis, Tamir Gonen, Neil P. King, and David Baker. Design of a hyperstable 60-subunit protein icosahedron. Nature, 535(7610):136â139, June 2016. [57]Kexin Huang, Serena Zhang, Hanchen Wang, Yuanhao Qu, Yingzhou Lu, Yusuf Roohani, Ryan Li, Lin Qiu, Gavin Li, Junze Zhang, et al. Biomni: A general-purpose biomedical ai agent. biorxiv, pages 2025â05, 2025. [58]Ian R Humphreys, Jimin Pei, Minkyung Baek, Aditya Krishnakumar, Ivan Anishchenko, Sergey Ovchinnikov, Jing Zhang, Travis J Ness, Sudeep Banjade, Saket R Bagde, et al. Computed structures of core eukaryotic protein complexes. Science, 374(6573):eabm4805, 2021. [59] Lin Jiang, Eric A. Althoff, Fernando R. Clemente, Lindsey Doyle, Daniela R Ìothlisberger, Alexandre Zanghellini, Jasmine L. Gallaher, Jamie L. Betker, Fujie Tanaka, Carlos F. Barbas, Donald Hilvert, Kendall N. Houk, Barry L. Stoddard, and David Baker. De Novo Computational Design of Retro-Aldol Enzymes. Science (New York, N.Y.), 319(5868):1387â1391, March 2008. [60]Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770, 2023. 15 Protein Design with Agent Rosetta [61]John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Ë Z Ìıdek, Anna Potapenko, et al. Highly accurate protein structure prediction with alphafold. nature, 596(7873):583â589, 2021. [62]Neil P. King, Jacob B. Bale, William Sheffler, Dan E. McNamara, Shane Gonen, Tamir Gonen, Todd O. Yeates, and David Baker. Accurate design of co-assembling multi-component protein nanomaterials. Nature, 510(7503):103â108, June 2014. [63]Neil P. King, William Sheffler, Michael R. Sawaya, Breanna S. Vollmar, John P. Sumida, Ingemar Andr Ìe, Tamir Gonen, Todd O. Yeates, and David Baker. Computational design of self-assembling protein nanomaterials with atomic level accuracy. Science (New York, N.Y.), 336(6085):1171â1174, June 2012. [64]Nobuyasu Koga, Rie Tatsumi-Koga, Gaohua Liu, Rong Xiao, Thomas B. Acton, Gaetano T. Monte- lione, and David Baker. Principles for designing ideal protein structures. Nature, 491(7423):222â227, November 2012. [65] Rohith Krishna, Jue Wang, Woody Ahern, Pascal Sturmfels, Preetham Venkatesh, Indrek Kalvet, Gyu Rie Lee, Felix S Morey-Burrows, Ivan Anishchenko, Ian R Humphreys, et al. Generalized biomolecular modeling and design with rosettafold all-atom. Science, 384(6693):eadl2528, 2024. [66] Brian Kuhlman and David Baker. Native protein sequences are close to optimal for their structures. Proceedings of the National Academy of Sciences of the United States of America, 97(19):10383â 10388, September 2000. [67]Brian Kuhlman, Gautam Dantas, Gregory C. Ireton, Gabriele Varani, Barry L. Stoddard, and David Baker. Design of a novel globular protein fold with atomic-level accuracy. Science (New York, N.Y.), 302(5649):1364â1368, November 2003. [68] Andrew Leaver-Fay, Michael Tyka, Steven M. Lewis, Oliver F. Lange, James Thompson, Ron Jacak, Kristian W. Kaufman, P. Douglas Renfrew, Colin A. Smith, Will Sheffler, Ian W. Davis, Seth Cooper, Adrien Treuille, Daniel J. Mandell, Florian Richter, Yih-En Andrew Ban, Sarel J. Fleishman, Jacob E. Corn, David E. Kim, Sergey Lyskov, Monica Berrondo, Stuart Mentzer, Zoran Popovi Ìc, James J. Havranek, John Karanicolas, Rhiju Das, Jens Meiler, Tanja Kortemme, Jeffrey J. Gray, Brian Kuhlman, David Baker, and Philip Bradley. Rosetta3. In Methods in Enzymology, volume 487, pages 545â574. Elsevier, 2011. [69]Jin Sub Lee, Osama Abdin, and Philip M Kim. Language models for protein design. Current Opinion in Structural Biology, 92:103027, 2025. [70] Julia Koehler Leman, Brian D. Weitzner, Steven M. Lewis, Jared Adolf-Bryfogle, Nawsad Alam, Rebecca F. Alford, Melanie Aprahamian, David Baker, Kyle A. Barlow, Patrick Barth, Benjamin Basanta, Brian J. Bender, Kristin Blacklock, Jaume Bonet, Scott E. Boyken, Phil Bradley, Chris Bystroff, Patrick Conway, Seth Cooper, Bruno E. Correia, Brian Coventry, Rhiju Das, Ren Ìe M. De Jong, Frank DiMaio, Lorna Dsilva, Roland Dunbrack, Alexander S. Ford, Brandon Frenz, Darwin Y. Fu, Caleb Geniesse, Lukasz Goldschmidt, Ragul Gowthaman, Jeffrey J. Gray, Dominik Gront, Sharon Guffy, Scott Horowitz, Po-Ssu Huang, Thomas Huber, Tim M. Jacobs, Jeliazko R. Jeliazkov, David K. Johnson, Kalli Kappel, John Karanicolas, Hamed Khakzad, Karen R. Khar, Sagar D. Khare, Firas Khatib, Alisa Khramushin, Indigo C. King, Robert Kleffner, Brian Koepnick, Tanja Kortemme, Georg Kuenze, Brian Kuhlman, Daisuke Kuroda, Jason W. Labonte, Jason K. Lai, Gideon Lapidoth, Andrew Leaver-Fay, Steffen Lindert, Thomas Linsky, Nir London, Joseph H. Lubin, Sergey Lyskov, Jack Maguire, Lars Malmstr Ìom, Enrique Marcos, Orly Marcu, Nicholas A. Marze, Jens Meiler, Rocco Moretti, Vikram Khipple Mulligan, Santrupti Nerli, Christoffer Norn, Shane Ì OâConch Ìuir, Noah Ollikainen, Sergey Ovchinnikov, Michael S. Pacella, 16 Protein Design with Agent Rosetta Xingjie Pan, Hahnbeom Park, Ryan E. Pavlovicz, Manasi Pethe, Brian G. Pierce, Kala Bharath Pilla, Barak Raveh, P. Douglas Renfrew, Shourya S. Roy Burman, Aliza Rubenstein, Marion F. Sauer, Andreas Scheck, William Schief, Ora Schueler-Furman, Yuval Sedan, Alexander M. Sevy, Nikolaos G. Sgourakis, Lei Shi, Justin B. Siegel, Daniel-Adriano Silva, Shannon Smith, Yifan Song, Amelie Stein, Maria Szegedy, Frank D. Teets, Summer B. Thyme, Ray Yu-Ruei Wang, Andrew Watkins, Lior Zimmerman, and Richard Bonneau. Macromolecular modeling and design in Rosetta: recent methods and frameworks. Nature Methods, 17(7):665â680, July 2020. Publisher: Nature Publishing Group. [71] Julia Koehler Leman, Brian D. Weitzner, P. Douglas Renfrew, Steven M. Lewis, Rocco Moretti, Andrew M. Watkins, Vikram Khipple Mulligan, Sergey Lyskov, Jared Adolf-Bryfogle, Jason W. Labonte, Justyna Krys, RosettaCommons Consortium, Christopher Bystroff, William Schief, Dominik Gront, Ora Schueler-Furman, David Baker, Philip Bradley, Roland Dunbrack, Tanja Kortemme, Andrew Leaver-Fay, Charlie E. M. Strauss, Jens Meiler, Brian Kuhlman, Jeffrey J. Gray, and Richard Bonneau. Better together: Elements of successful scientific software development in a distributed collaborative community. PLOS Computational Biology, 16(5):e1007507, May 2020. Publisher: Public Library of Science. [72]Long Li, Weiwen Xu, Jiayan Guo, Ruochen Zhao, Xingxuan Li, Yuqian Yuan, Boqiang Zhang, Yuming Jiang, Yifei Xin, Ronghao Dang, et al. Chain of ideas: Revolutionizing research via novel idea development with llm agents. arXiv preprint arXiv:2410.13185, 2024. [73] Shanda Li, Tanya Marwah, Junhong Shen, Weiwei Sun, Andrej Risteski, Yiming Yang, and Ameet Talwalkar. Codepde: An inference framework for llm-driven pde solver generation. arXiv preprint arXiv:2505.08783, 2025. [74]Yeqing Lin and Mohammed AlQuraishi. Generating novel, designable, and diverse protein structures by equivariantly diffusing oriented residue clouds. arXiv preprint arXiv:2301.12485, 2023. [75] Yeqing Lin, Minji Lee, Zhao Zhang, and Mohammed AlQuraishi. Out of many, one: Designing and scaffolding proteins at the scale of the structural universe with genie 2. arXiv preprint arXiv:2405.15489, 2024. [76]Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Allan dos Santos Costa, Maryam Fazel-Zarandi, Tom Sercu, Sal Candido, et al. Language models of protein sequences at the scale of evolution enable accurate structure prediction. BioRxiv, 2022:500902, 2022. [77]Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist: Towards fully automated open-ended scientific discovery. arXiv preprint arXiv:2408.06292, 2024. [78]ThangLuongandEdwardLockhart.Advancedversionofgemini withdeepthinkofficiallyachievesgold-medalstandardattheinterna- tionalmathematicalolympiad.https://deepmind.google/discover/blog/ advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/, July 2025. Google DeepMind Blog. [79]Elahe Masoumzadeh, Joseph M Courtney, Cyril Charlier, Jinfa Ying, Philip Anfinrud, and Adriaan Bax. Structure of a transient protein-folding intermediate by pressure-jump nmr spectroscopy. Proceedings of the National Academy of Sciences, 122(41):e2519493122, 2025. [80]and Microsoft Research AI4Science. The impact of large language models on scientific discovery: a preliminary study using gpt-4, 2023. 17 Protein Design with Agent Rosetta [81]Jeremy H. Mills, Sagar D. Khare, Jill M. Bolduc, Farhad Forouhar, Vikram Khipple Mulligan, Scott Lew, Jayaraman Seetharaman, Liang Tong, Barry L. Stoddard, and David Baker. Computational design of an unnatural amino acid dependent metalloprotein with atomic level accuracy. Journal of the American Chemical Society, 135(36):13393â13399, September 2013. [82]Ludovico Mitchener, Angela Yiu, Benjamin Chang, Mathieu Bourdenx, Tyler Nadolski, Arvis Sulovari, Eric C Landsness, Daniel L Barabasi, Siddharth Narayanan, Nicky Evans, et al. Kosmos: An ai scientist for autonomous discovery. arXiv preprint arXiv:2511.02824, 2025. [83] Rocco Moretti, Brian J. Bender, Brittany Allison, and Jens Meiler. Rosetta and the Design of Ligand Binding Sites. In Barry L. Stoddard, editor, Computational Design of Ligand Binding Proteins, pages 47â62. Springer, New York, NY, 2016. [84]Vikram K. Mulligan. The emerging role of computational design in peptide macrocycle drug discovery. Expert Opinion on Drug Discovery, 15(7):833â852, July 2020. [85]Vikram Khipple Mulligan. Current directions in combining simulation-based macromolecular modeling approaches with deep learning. Expert Opinion on Drug Discovery, 16(9):1025â1044, September 2021. [86]Vikram Khipple Mulligan. Computational Methods for Peptide Macrocycle Drug Design. In Seetharama D. Jois, editor, Peptide Therapeutics: Fundamentals of Design, Development, and Delivery, volume 47, pages 79â161. Springer International Publishing, New York, 2022. Series Title: AAPS Advances in the Pharmaceutical Sciences Series. [87] Vikram Khipple Mulligan and Parisa Hosseinzadeh. Computational Design of Peptide-Based Binders to Therapeutic Targets. In Swapnil V. Ghodge, Kaustav Biswas, and Andrei A. Golosov, editors, Approaching the Next Inflection in Peptide Therapeutics: Attaining Cell Permeability and Oral Bioavailability, volume 1417 of ACS Symposium Series, pages 55â102. American Chemical Society, Washington, DC, August 2022. [88]Vikram Khipple Mulligan, Christine S. Kang, Michael R. Sawaya, Stephen Rettie, Xinting Li, Inna Antselovich, Timothy W. Craven, Andrew M. Watkins, Jason W. Labonte, Frank DiMaio, Todd O. Yeates, and David Baker. Computational design of mixed chirality pep- tide macrocycles with internal symmetry. Protein Science, 29(12):2433â2445, 2020.eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/pro.3974. [89] Vikram Khipple Mulligan, Sean Workman, Tianjun Sun, Stephen Rettie, Xinting Li, Liam J. Worrall, Timothy W. Craven, Dustin T. King, Parisa Hosseinzadeh, Andrew M. Watkins, P. Dou- glas Renfrew, Sharon Guffy, Jason W. Labonte, Rocco Moretti, Richard Bonneau, Natalie C. J. Strynadka, and David Baker. Computationally designed peptide macrocycle inhibitors of New Delhi metallo-ÎČ-lactamase 1. Proceedings of the National Academy of Sciences of the United States of America, 118(12):e2012800118, March 2021. [90]Siddharth M Narayanan, James D Braza, Ryan-Rhys Griffiths, Albert Bou, Geemi Wellawatte, Mayk Caldas Ramos, Ludovico Mitchener, Samuel G Rodriques, and Andrew D White. Training a scientific reasoning model for chemistry. arXiv preprint arXiv:2506.17238, 2025. [91] Deepak Nathani, Lovish Madaan, Nicholas Roberts, Nikolay Bashlykov, Ajay Menon, Vincent Moens, Amar Budhiraja, Despoina Magka, Vladislav Vorotilov, Gaurav Chaurasia, et al. Mlgym: A new framework and benchmark for advancing ai research agents. arXiv preprint arXiv:2502.14499, 2025. [92]Erik Nijkamp, Jeffrey A Ruffolo, Eli N Weinstein, Nikhil Naik, and Ali Madani. Progen2: exploring the boundaries of protein language models. Cell systems, 14(11):968â978, 2023. 18 Protein Design with Agent Rosetta [93]Liangbo Ning, Ziran Liang, Zhuohang Jiang, Haohao Qu, Yujuan Ding, Wenqi Fan, Xiao-yong Wei, Shanru Lin, Hui Liu, Philip S Yu, et al. A survey of webagents: Towards next-generation ai agents for web automation with large foundation models. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, pages 6140â6150, 2025. [94]Saro Passaro, Gabriele Corso, Jeremy Wohlwend, Mateo Reveiz, Stephan Thaler, Vignesh Ram Somnath, Noah Getz, Tally Portnoi, Julien Roy, Hannes Stark, David Kwabi-Addo, Dominique Beaini, Tommi Jaakkola, and Regina Barzilay. Boltz-2: Towards accurate and efficient binding affinity prediction. bioRxiv, 2025. [95] Kevin Pu, KJ Kevin Feng, Tovi Grossman, Tom Hope, Bhavana Dalvi Mishra, Matt Latzke, Jonathan Bragg, Joseph Chee Chang, and Pao Siangliulue. Ideasynth: Iterative research idea development through evolving and composing idea facets with literature-grounded feedback. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pages 1â31, 2025. [96]Ingrid Swanson Pultz, Malcolm Hill, Joanne M. Vitanza, Clancey Wolf, Lasse Saaby, Tina Liu, Peter Winkle, and Daniel A. Leffler. Gluten Degradation, Pharmacokinetics, Safety, and Tolerability of TAK-062, an Engineered Enzyme to Treat Celiac Disease. Gastroenterology, 161(1):81â93.e3, July 2021. [97]Yuanhao Qu, Kaixuan Huang, Ming Yin, Kanghong Zhan, Dyllan Liu, Di Yin, Henry C Cousins, William A Johnson, Xiaotong Wang, Mihir Shah, et al. Crispr-gpt for agentic automation of gene-editing experiments. Nature Biomedical Engineering, pages 1â14, 2025. [98] Marissa Radensky, Simra Shahid, Raymond Fok, Pao Siangliulue, Tom Hope, and Daniel S Weld. Scideator: Human-llm scientific idea generation grounded in research-paper facet recombination. arXiv preprint arXiv:2409.14634, 2024. [99]P. D. Renfrew, T. W. Craven, G. L. Butterfoss, K. Kirshenbaum, and R. Bonneau. A rotamer library to enable modeling and design of peptoid foldamers. Journal of the American Chemical Society, 136:8772â8782, 2014. [100]P. Douglas Renfrew, Eun Jung Choi, Richard Bonneau, and Brian Kuhlman. Incorporation of Noncanonical Amino Acids into Rosetta and Use in Computational Protein-Peptide Interface Design. PLoS ONE, 7(3):e32637, March 2012. [101] Gabriel J. Rocklin, Tamuka M. Chidyausiku, Inna Goreshnik, Alex Ford, Scott Houliston, Alexan- der Lemak, Lauren Carter, Rashmi Ravichandran, Vikram K. Mulligan, Aaron Chevalier, Cheryl H. Arrowsmith, and David Baker. Global analysis of protein folding using massively parallel design, synthesis, and testing. Science, 357(6347):168â175, July 2017. [102]Baptiste Rozi`ere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, J Ìer Ìemy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre D Ìefossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. Code llama: Open foundation models for code, 2024. [103] Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J Maddison, and Tatsunori Hashimoto. Identifying the risks of lm agents with an lm-emulated sandbox. arXiv preprint arXiv:2309.15817, 2023. [104]Clara T. Schoeder, Samuel Schmitz, Jared Adolf-Bryfogle, Alexander M. Sevy, Jessica A. Finn, Marion F. Sauer, Nina G. Bozhanova, Benjamin K. Mueller, Amandeep K. Sangha, Jaume Bonet, Jonathan H. Sheehan, Georg Kuenze, Brennica Marlow, Shannon T. Smith, Hope Woods, Brian J. 19 Protein Design with Agent Rosetta Bender, Cristina E. Martina, Diego del Alamo, Pranav Kodali, Alican Gulsevin, William R. Schief, Bruno E. Correia, James E. Jr. Crowe, Jens Meiler, and Rocco Moretti. Modeling Immunity with Rosetta: Methods for Antibody and Antigen Design. Biochemistry, 60(11):825â846, March 2021. Publisher: American Chemical Society. [105]Schr Ìodinger, LLC. The AxPyMOL molecular graphics plugin for Microsoft PowerPoint, version 3.1, 2025. [106] Schr Ìodinger, LLC. The PyMOL molecular graphics system, version 3.1, 2025. [107]Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024. [108]Will Sheffler and David Baker. Rosettaholes: rapid assessment of protein core packing for structure prediction, refinement, design, and validation. Protein Science, 18(1):229â239, 2009. [109]Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36:8634â8652, 2023. [110] Justin B. Siegel, Alexandre Zanghellini, Helena M. Lovick, Gert Kiss, Abigail R. Lambert, Jennifer L. St Clair, Jasmine L. Gallaher, Donald Hilvert, Michael H. Gelb, Barry L. Stoddard, Kendall N. Houk, Forrest E. Michael, and David Baker. Computational design of an enzyme catalyst for a stereoselective bimolecular Diels-Alder reaction. Science (New York, N.Y.), 329(5989):309â 313, July 2010. [111]Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, et al. Openai gpt-5 system card. arXiv preprint arXiv:2601.03267, 2025. [112] Colin A. Smith and Tanja Kortemme. Backrub-Like Backbone Simulation Recapitulates Natural Protein Conformational Variability and Improves Mutant Side-Chain Prediction. Journal of Molecular Biology, 380(4):742â756, July 2008. [113]Eva-Maria Strauch, Steffen M Bernard, David La, Alan J Bohn, Peter S Lee, Caitlin E Anderson, Travis Nieusma, Carly A Holstein, Natalie K Garcia, Kathryn A Hooper, Rashmi Ravichandran, Jorgen W Nelson, William Sheffler, Jesse D Bloom, Kelly K Lee, Andrew B Ward, Paul Yager, Deborah H Fuller, Ian A Wilson, and David Baker. Computational design of trimeric influenza- neutralizing proteins targeting the hemagglutinin receptor binding site. Nature Biotechnology, 35(7):667â671, July 2017. [114]Theodore Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas Griffiths. Cognitive architectures for language agents. Transactions on Machine Learning Research, 2023. [115] NovelSeek Team, Bo Zhang, Shiyang Feng, Xiangchao Yan, Jiakang Yuan, Zhiyin Yu, Xiaohan He, Songtao Huang, Shaowei Hou, Zheng Nie, et al. Novelseek: When agent becomes the scientistâ building closed-loop system from hypothesis to verification. arXiv preprint arXiv:2505.16938, 2025. [116] David F. Thieker, Jack B. Maguire, Stephan T. Kudlacek, Andrew Leaver-Fay, Sergey Lyskov, and Brian Kuhlman.Stabilizing proteins, simplified: A Rosetta-based webtool for predicting favorable mutations.Protein Science, 31(10):e4428, 2022.eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/pro.4428. 20 Protein Design with Agent Rosetta [117]Summer Thyme and David Baker. Redesigning the Specificity of ProteinâDNA Interactions with Rosetta. In David R. Edgell, editor, Homing Endonucleases: Methods and Protocols, pages 265â282. Humana Press, Totowa, NJ, 2014. [118]Christine E. Tinberg, Sagar D. Khare, Jiayi Dou, Lindsey Doyle, Jorgen W. Nelson, Alberto Schena, Wojciech Jankowski, Charalampos G. Kalodimos, Kai Johnsson, Barry L. Stoddard, and David Baker. Computational design of ligand-binding proteins with high affinity and selectivity. Nature, 501(7466):212â216, September 2013. [119]Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong. Solving olympiad geometry without human demonstrations. Nature, 625(7995):476â482, 2024. [120]Senadhi Vijay-Kumar, Charles E Bugg, and William J Cook. Structure of ubiquitin refined at 1.8 Ìaresolution. Journal of molecular biology, 194(3):531â544, 1987. [121]Eric Wang, Samuel Schmidgall, Paul F Jaeger, Fan Zhang, Rory Pilgrim, Yossi Matias, Joelle Barral, David Fleet, and Shekoofeh Azizi. Txgemma: Efficient and agentic llms for therapeutics. arXiv preprint arXiv:2504.06196, 2025. [122] Zihan Wang, Kangrui Wang, Qineng Wang, Pingyue Zhang, Linjie Li, Zhengyuan Yang, Xing Jin, Kefan Yu, Minh Nhat Nguyen, Licheng Liu, et al. Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning. arXiv preprint arXiv:2504.20073, 2025. [123]A. M. Watkins, T. W. Craven, P. D. Renfrew, P. S. Arora, and R. Bonneau. Rotamer libraries for the high-resolution design of ÎČ-amino acid foldamers. Structure, 25(11):1771â1780.e3, 2017. [124] Joseph L Watson, David Juergens, Nathaniel R Bennett, Brian L Trippe, Jason Yim, Helen E Eisenach, Woody Ahern, Andrew J Borst, Robert J Ragotte, Lukas F Milles, et al. De novo design of protein structure and function with rfdiffusion. Nature, 620(7976):1089â1100, 2023. [125] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824â24837, 2022. [126]Zhepei Wei, Wenlin Yao, Yao Liu, Weizhi Zhang, Qin Lu, Liang Qiu, Changlong Yu, Puyang Xu, Chao Zhang, Bing Yin, et al. Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning. arXiv preprint arXiv:2505.16421, 2025. [127] Jeremy Wohlwend, Gabriele Corso, Saro Passaro, Noah Getz, Mateo Reveiz, Ken Leidal, Wojtek Swiderski, Liam Atkinson, Tally Portnoi, Itamar Chinn, Jacob Silterra, Tommi Jaakkola, and Regina Barzilay. Boltz-1: Democratizing biomolecular interaction modeling. bioRxiv, 2024. [128]Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search. arXiv preprint arXiv:2504.08066, 2025. [129]An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. [130]John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated software engineering. Advances in Neural Information Processing Systems, 37:50528â50652, 2024. 21 Protein Design with Agent Rosetta [131]Tristan Zaborniak, Noora Azadvari, Qiyao Zhu, S. M. Bargeen A. Turzo, Parisa Hosseinzadeh, P. Douglas Renfrew, and Vikram Khipple Mulligan. The open-source Masala software suite: Facilitating rapid methods development for synthetic heteropolymer design. Methods in Enzymology, 723:299â426, 2025. [132]Yu Zhang, Xiusi Chen, Bowen Jin, Sheng Wang, Shuiwang Ji, Wei Wang, and Jiawei Han. A comprehensive survey of scientific large language models and their applications in scientific discovery. arXiv preprint arXiv:2406.10833, 2024. [133]Yuyang Zhang, Yuhang Liu, Zinnia Ma, Min Li, Chunfu Xu, and Haipeng Gong. Improving diffusion-based protein backbone generation with global-geometry-aware latent encoding. Nature Machine Intelligence, pages 1â15, 2025. [134]Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang. Memorybank: Enhancing large language models with long-term memory. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 19724â19731, 2024. [135]Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al. Webarena: A realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854, 2023. 22 Protein Design with Agent Rosetta A Related Works In this section, we provide a detailed overview of relevant related works. LLM agents. Through text generation, modern LLMs are capable of using external tools such as web browsers [135,2,126,52,93], email [103,35], and, more broadly, code-based interfaces [60,55,130]. Coupled with reasoning and chain-of-thought prompting [125,107], LLMs can act as efficient agents, automating repetitive workflows and solving complex tasks that require multi-turn interactions [122]. Scientific agents. Scientific discovery workflows are generally well-structured, which makes them attractive for the development of AI scientists [47,77,115]. Recent works have explored LLMs agents for idea and hypothesis generation [12,72,95,98,50], experiment design [18,43], and paper writing [128,26,44]. These approaches have found several applications in math [119,78], chemistry [20,18,90], and therapeutics [121]. Most closely related to our work, ProtAgents [43] is a multi-agent system that orchestrates the use of existing machine learning tools for de novo protein discovery; CRISPR-GPT [97] studies LLM agents for gene editing tasks; and Biomni [57] develops an agentic framework for general biomedical tasks. We refer interested readers to [48, 132] for comprehensive reviews. Manually-crafted Rosetta protocols for heteropolymer design. Rosetta has been extensively used for rational design of new protein folds [67,64,101],rigidly-structured peptides [15,53,32,88], and other heteropolymers [100,38,99]. Proteins have been rationally designed with Rosetta that self-assemble into sheets [45] or cages [62,56], that bind metals [81] or small molecules [118], that catalyze new enzymatic reactions [59,110], and that specifically recognize target proteins [113]. Non- canonical peptides have also been designed which bind target proteins [89,54] or passively diffuse across cell membranes [16]. Rosettaâs infrastructure, energy function, and development are described in [68,8,70,71]. Rosettaâs scripting interfaces are described in [25,41]. Non-canonical Rosetta design methods are reviewed in [84,86,87,36]. Recently, the Masala libraries were introduced to permit easy, plugin-based development of new optimizers and other algorithms to extend Rosetta or other heteropolymer modeling software [131]. ML for protein structure prediction and design. Deep learning methodologies have proved successful at training models for particular design pipelines. For example, models exist for both monomer (AlphaFold 2 [61], RoseTTAFold [13,58], ESMFold [76]) and complex structure prediction (AlphaFold 3 [1] and Multimer [39], OpenFold [7] RosettaFold All-Atom [65], the Chai [23,24] and Boltz [127,94] series of models); sequence generation (ProteinMPNN [33], the ProGen [92,17] series, ProtGPT2 [40]); and structure generation and docking (RFdiffusion [124], DiffDock [30], and the Genie [74,75] series). We point to [69] for a recent review on language models for protein design, and [85] for a review of earlier ML approaches used for protein modeling. B Prompts In this section, we include all prompt templates. In the prompts, fields enclosed within curly brackets (e.g., fieldname) are populated at runtime. B.1 System Prompt Here, we include Agent Rosettaâs system prompt.docsis populated with Python-style docstrings that summarize the available actions and their parameters, andreasoningformattingvaries depending 23 Protein Design with Agent Rosetta on the base LLM: for reasoning models this field is omitted, whereas for models that require chain of thought (CoT) prompting [125] we include instructions to first reason step-by-step with special delimiters like<think></think>for OpenAI models, or<reasoning></reasoning>for Qwen models. Box 1: System prompt You are an expert RosettaScripts coding agent that supports scientists in biomolecular design tasks: - Rosetta is a computational toolkit for modeling, predicting, and designing biomolecular structures and interactions, using physics-based energy functions and stochastic search. - RosettaScripts is Rosettaâs XML-based interface for assembling custom modeling protocols by combining movers, filters, and scoring terms. Your goal is to follow a user-defined brief and to generate a valid RosettaScripts protocol that captures the brief precisely. You have interactive access to a RosettaScripts environment with the following actions: docs **Instructions:** You will follow a two-step interaction schema with the RosettaScripts environment. At the beginning of each interaction round, you will be presented a summary of the current state of the environment. You must follow these instructions carefully in order to choose and execute the next best action that will achieve the design goals: 1. STEP 1: Reason about the biophysical implications of the design brief for the current state of the environment and CHOOSE the next action to take. 2. STEP 2: Only the environment has displayed the full documentation of the action you chose, WRITE the full action call with all its arguments. You must follow the formatting instructions below. **Formatting:** reasoning_formatting Write your action call inside <action></action> tags, following this exact format: <action tag="choose" or "run"> <name>action_name</name> <arg_name1>arg_value1</arg_name1> <arg_name2>arg_value2</arg_name2> ... </action> Rules: - Use <action tag="choose"></action> to CHOOSE the action. - Use <action tag="run"></action> to RUN the action. - Use <action_name>action_name</action_name> to specify the name of the action. - For each argument of the action, use XML tags with the name and value of the argument. - Write one (1) action call per response only. The environment will only parse the first <action></action> block. 24 Protein Design with Agent Rosetta - Do not include any reasoning or comments in your action call. Write valid action content only. - Do not code fence your action call. **Example:** - FIRST, write: <action tag="choose"> <name>action_name</name> </action> - THEN, only after the environment has displayed the full action documentation, write: <action tag="run"> <name>action_name</name> <arg_name1>arg_value1</arg_name1> <arg_name2>arg_value2</arg_name2> ... </action> Now, read the design brief and start designing your RosettaScripts protocol. The docstrings for our three types of actions are: Box 2: Action docstrings - rotamer_change """Perform a Rosetta FastDesign action with design enabled to change the rotamers of the sequence without perturbing the backbone. This action combines residue selectors, compositional penalties, residue restrictions, and packing restrictions to guide the search process. Arguments: - residue_selectors (string, optional): the XML definitions of the residue selectors. - penalties (array, optional): a list of compositional penalty definitions. Each item includes the following parameters: - comp (string, optional): the penalty definition blocks specifying the compositional constraints. - comp_selector_name (string, optional): the name of the residue selector to apply the compositional constraints to. - residue_restrictions (array, optional): a list of residue restriction definitions. Each item includes the following parameters: - type (string, optional): whether to ârestrictâ or âprohibitâ residue types. - residues (string, optional): the list of one- or three-letter residue codes separated by a comma. - selector_name (string, optional): the name of the residue selector to apply the restriction to. - packing_restrictions (string, optional): A list of residue selector names separated by commas. """ - backbone_change """Perturb the backbone conformation. This action implements three Rosetta backbone movers: âsmallâ, âshearâ, and âbackrubâ. Specify the mover to use with âmover_nameâ, the parameters of the mover in XML format with âmover_paramsâ, and the residues to perturb with âresidue_selectorâ. 25 Protein Design with Agent Rosetta Arguments: - mover_name (string, required): the name of the Rosetta backbone mover to use (can be âsmallâ, âshearâ, or âbackrubâ). - mover_params (string, optional): the parameters of the mover in XML format. - residue_selectors (string, optional): the XML definitions of the residue selectors. - mover_selector_name (string, optional): the name of the residue selector to apply the mover to. """ - go_back_to_step """Revert the environment state to a previous step in the trajectory. This action resets the environment to the state at the end of the specified step, allowing you to retry actions or explore different paths. Arguments: - step (integer, optional): the number of the step to revert to, starting from 0. """ B.2 Design Brief Here, we include Agent Rosettaâs design brief prompt.promptis populated with the user-specified task description and objectives,statecontains the textual representation of the initial environment state, andactionformattinginstructionsvaries by model. Similarly to thereasoningformatting parameter in the system prompt,actionformattinginstructionsinstructs reasoning models to directly generate the final response, without including any reasoning, and non-reasoning models that require CoT prompting to first summarize step-by-step reasoning within special delimiters. This ensures the same response content and format across different LLMs. Box 3: Design brief **Design Brief:** prompt **Starting State:** state Now choose the next best action following the formatting instructions in the system prompt: action_formatting_instructions B.3 Multi-turn Interaction Workflow Here, we include the prompts for each stage of the multi-turn interaction workflow, starting from action selection.header,step, andactnameare populated at runtime depending on the exit status and environment state,statecontains the textual representation of the environment state,summary reports the tabular summary of the action history, andactionformattinginstructionsvaries depending on model as described in Section B.2 for the design brief. 26 Protein Design with Agent Rosetta Box 4: Structured artificial reasoning and action selection header **Step Number:** step **Action Name:** act_name **Results:** state --- **History Summary:** summary --- **Revision Instructions:** First review the last action call, then judge the results of the last action call, and finally choose the next best action to achieve the design goals. Make sure to take into consideration all information provided, including the history summary. Integrate this information with your expert biophysics knowledge. Some example questions to guide your judgment are: - Did the last action worsen any important metrics compared to the previous step? - Were there any penalty definition blocks that did not affect the results at all? For example, because their residue selectors were syntactically valid but semantically empty? - Were there any penalty definition blocks that had unintended consequences? For example, because their residue selectors were misspecified or because their penalties introduced unfavorable residues? - Are there any important pieces of information missing from the results that would help you craft the next action better? - Can you infer those missing pieces of information based on your expert knowledge of RosettaScripts and biomolecular design? Now choose the next best action following the formatting instructions in the system prompt: action_formatting_instructions After choosing the next action, the agent is prompted again to generate the full action call with all its parameters with the necessary RosettaScripts documentation.actnamecontains the name of the chosen action and docs its documentation. Box 5: Structured reasoning and parameter generation You chose to perform action âact_nameâ. This is the full documentation of the action: docs --- **Action Instructions:** First carefully read the documentation of the action you chose, then reason about a practical implementation that achieves your intended effects, and finally write the 27 Protein Design with Agent Rosetta full action call with all its arguments. Remember that the documentation is not exhaustive, and you need to integrate it with your expert knowledge of RosettaScripts. Be specific in your reasoning: the expert biomolecular scientists on your team should be able to understand, review, and critique your plan. Now write the full action call with all its arguments following the formatting instructions in the system prompt: action_formatting_instructions In case the generated action incurs in RosettaScripts syntax validation or runtime errors, the agent is prompted to correct its mistakes.erroris populated with the error message from the RosettaScripts environment. Box 6: Error correction The RosettaScripts environment failed to run your action with the following error: error **Instructions:** First carefully read the error message, then reason about the possible causes of the error, and finally write the corrected full action call with all its arguments. Your corrected action call must preserve the intent of your previous action. Some common sources of errors are: - Syntactic mistakes in the action call: wrong parameter names, wrong XML structure. - Semantic mistakes in the action call: wrong parameter values, logically invalid combinations of parameters. Be specific in your reasoning: the expert biomolecular scientists on your team should be able to understand, review, and critique your solution. Now write the corrected full action call with all its arguments following the formatting instructions in the system prompt: action_formatting_instructions C Fixed Backbone Sequence Design In this section, we include further experimental details on stabilizing an initial backbone conformation with canonical amino acids only. C.1 Design Brief Here, we include the specific design brief used in this experiment. This design brief is composed at runtime with the general design brief prompt template presented in Section B.2. Box 7: Design brief You are given an initial backbone conformation composed of glycine residues only. Your task is to design (i.e., change) the residues of the sequence such that: 1. It is energetically stable according to biophysical principles. 2. It has low Rosetta energy. 28 Protein Design with Agent Rosetta 3. The sequence folds into the same shape as the initial backbone structure. You will be provided the following proxy metrics to help you assess task progress: 1. Total Rosetta energy. 2. The RMSD between the ESMFold predicted fold and the initial backbone conformation. 3. The uncertainty (i.e., average pLDDT) of the ESMFold prediction. At each step: - The Rosetta environment will execute your action and produce an ensemble of candidate designs. - You will be shown summary statistics of the ensemble. - You will choose the next action to take, which will be applied to Pareto optimal candidates only in terms of the provided proxy metrics. C.2 Initial Backbone Conformations In this section, we describe the inclusion criteria and preprocessing steps used to select 8 backbone conformations to use in this experiment (see Fig. C.1). Figure C.1: The 8 backbone conformations to stabilize with canonical amino acids only. Selection of PDBs for design. Protein backbones for canonical amino acid design were selected from the PDB to ensure structural diversity and designability. Structures were selected based on release date cutoff of after May 15, 2025, single polymer chain, and chain length of less than 150 residues. The release-date cutoff was chosen to reduce the likelihood that the exact structures were present in the ProteinMPNN training or test sets, thereby reflecting a realistic design campaign setting. However, we could not exclude that ProteinMPNN was trained on proteins with similar structure. The selected PDBs were: 9PL1 [79], 9VYW [42], 9VYQ [42], 9C14 [34], 8YS1 [11], 8VWO [51], 9IFR [31], and 9KGY [133]. Preprocessing of PDBs. All structures chosen as initial scaffolds for design with canonical amino acids were processed in PyMOL [106,105] to remove alternative conformations (altlocs B-G) and all 29 Protein Design with Agent Rosetta heteroatoms and X-ray crystallographic copies. All native residues were replaced with glycine without changes to the positions of the Cα atoms of the backbone conformations. C.3 Environment State Here, we include an illustrative example of the textual representation of the RosettaScripts environment state for the fixed-backbone sequence design task. We omit the tabular history summary for the sake of readability. Box 8: Environment State The RosettaScripts environment successfully executed your action: **Step Number:** 4 **Action Name:** rotamer_change **Results:** - Number of designs: 114 (6 Pareto optimal) - Average total Rosetta energy: -385.47 ± 12.39 - Average compositional before design 607.89, and after design 333.33 (-274.56) - Top 5 most common outlier residue types. A residue is an outlier if its energy term is above the 90-th quantile of the per-residue energy term for that design. For each outlier residue type, we include the most common outlier positions along the sequence (positions are 1-based): -- interresidue_repulsion: Average per-sequence 90-th quantile: 2.31 PHE: 99.12% (positions: 87,65,25) VAL: 98.25% (positions: 17,122,4,91,89,25,99) LEU: 96.49% (positions: 103,50,85) CYS: 96.49% (positions: 95,99,55) ILE: 91.23% (positions: 91,42,72,25,101,31,43,23) -- ramachandran_preference: Average per-sequence 90-th quantile: 0.56 THR: 100.00% (positions: 33,97) ASP: 98.25% (positions: 67) PRO: 97.37% (positions: 76,48) TYR: 95.61% (positions: 98) CYS: 94.74% (positions: 116,95) - Structural metrics: Cavity volume ( Ì A^3) : 12.24 ± 11.58 Radius of gyration ( Ì A) : 15.22 ± 0.01 Penalty for buried unsatisfied Hydrogen bonds : 5883.11 ± 1362.39 RMSD of ESMFold prediction to initial structure ( Ì A): 1.84 ± 1.50 CA pLDDT of ESMFold prediction : 0.79 ± 0.07 30 Protein Design with Agent Rosetta --- **History Summary:** history_summary C.4 Expert Human Baseline Protocols In this section, we include the expert human design protocols used as baselines in the fixed-backbone sequence design task. C.4.1 One-shot Design Protocol This protocol samples the entire sequence at once with the use of ResidueSelectors, TaskOperations, and compositional penalty blocks Box 9: One-shot design protocol <ROSETTASCRIPTS> <SCOREFXNS> <ScoreFunction name="r15_regular" weights="ref2015_cst.wts" /> <ScoreFunction name="r15_design" weights="ref2015.wts"> <Reweight scoretype="a_composition" weight="1.0" /> <Reweight scoretype="atom_pair_constraint" weight="1.0" /> </ScoreFunction> </SCOREFXNS> <RESIDUE_SELECTORS> <Layer name="core" select_core="true" select_boundary="false" select_surface="false" use_sidechain_neighbors="true" /> <Layer name="boundary" select_core="false" select_boundary="true" select_surface="false" use_sidechain_neighbors="true" /> <Layer name="surface" select_core="false" select_boundary="false" select_surface="true" use_sidechain_neighbors="true" /> </RESIDUE_SELECTORS> <RESIDUE_LEVEL_TASK_OPERATIONS> <RestrictToRepackingRLT name="RestrictToRepacking" /> </RESIDUE_LEVEL_TASK_OPERATIONS> <TASKOPERATIONS> <IncludeCurrent name="include_current_rotamer"/> <ExtraRotamersGeneric name="extra_sample_rotamers_design" ex1="1" ex1aro="1" /> <ExtraRotamersGeneric name="extra_sample_rotamers_relax" ex1="1" ex2="1" ex1aro="1" ex2aro="1" /> <ProhibitSpecifiedBaseResidueTypes name="prohibit_a_in_core" base_types="ARG,LYS,ASP,GLU,CYS,PRO" selector="core" /> <ProhibitSpecifiedBaseResidueTypes name="prohibit_a_in_boundary" base_types="ARG,LYS,ASP,GLU,CYS,PRO,MET,HIS" selector="boundary" /> <ProhibitSpecifiedBaseResidueTypes name="prohibit_a_in_surface" base_types="ILE,LEU,VAL,PHE,TRP,MET" selector="surface"/> </TASKOPERATIONS> <SIMPLE_METRICS> 31 Protein Design with Agent Rosetta <SequenceMetric name="record_sequence" output_mode="basename" /> </SIMPLE_METRICS> <MOVERS> <Small name="small_move" scorefxn="r15_regular" temperature="0.5" nmoves="1000" angle_max="2.0" preserve_detailed_balance="0" /> <FastDesign name="design_all" scorefxn="r15_design" disable_design="false" task_operations="include_current_rotamer,extra_sample_rotamers_design, prohibit_a_in_core,prohibit_a_in_boundary,prohibit_a_in_surface" repeats="4" relaxscript="default" min_type="lbfgs_armijo_nonmonotone" /> <AddCompositionConstraintMover name="comp_core" filename="comp_core" selector="core"/> <AddCompositionConstraintMover name="comp_boundary" filename="comp_boundary" selector="boundary"/> <AddCompositionConstraintMover name="comp_surface" filename="comp_surface" selector="surface"/> <AddConstraints name="geom_constraint"> <AtomPairConstraintGenerator name="gen_geom_csts" ca_only="1" use_harmonic="1" native="1" /> </AddConstraints> <RemoveConstraints name="rm_geom_csts" constraint_generators="gen_geom_csts" /> <FastRelax name="fast_relax" scorefxn="r15_regular" disable_design="true" task_operations="include_current_rotamer,extra_sample_rotamers_relax" repeats="3" relaxscript="default" min_type="lbfgs_armijo_nonmonotone" /> </MOVERS> <PROTOCOLS> <Add mover="geom_constraint" /> <Add mover="small_move" /> <Add mover="comp_core"/> <Add mover="comp_boundary"/> <Add mover="comp_surface"/> <Add mover="design_all" /> <Add mover="rm_geom_csts" /> <Add mover="fast_relax" /> <Add metrics="record_sequence" /> </PROTOCOLS> <OUTPUT scorefxn="r15_design" /> </ROSETTASCRIPTS> C.4.2 Staged Design Protocol This protocol first samples the core, then the boundary, and finally the surface of the molecule. It also uses ResidueSelectors, TaskOperations, and compositional penalty blocks Box 10: Staged design protocol <ROSETTASCRIPTS> <SCOREFXNS> <ScoreFunction name="r15_regular" weights="ref2015_cst.wts" /> <ScoreFunction name="r15_design" weights="ref2015.wts"> 32 Protein Design with Agent Rosetta <Reweight scoretype="a_composition" weight="1.0" /> <Reweight scoretype="atom_pair_constraint" weight="1.0" /> </ScoreFunction> </SCOREFXNS> <RESIDUE_SELECTORS> <Layer name="core" select_core="true" select_boundary="false" select_surface="false" use_sidechain_neighbors="true" /> <Layer name="boundary" select_core="false" select_boundary="true" select_surface="false" use_sidechain_neighbors="true" /> <Layer name="surface" select_core="false" select_boundary="false" select_surface="true" use_sidechain_neighbors="true" /> </RESIDUE_SELECTORS> <RESIDUE_LEVEL_TASK_OPERATIONS> <RestrictToRepackingRLT name="RestrictToRepacking" /> </RESIDUE_LEVEL_TASK_OPERATIONS> <TASKOPERATIONS> <IncludeCurrent name="include_current_rotamer" /> <ExtraRotamersGeneric name="extra_sample_rotamers_design" ex1="1" ex1aro="1" /> <ExtraRotamersGeneric name="extra_sample_rotamers_relax" ex1="1" ex2="1" ex1aro="1" ex2aro="1" /> <DesignRestrictions name="not_core"> <Action selector_logic="NOT core" residue_level_operations="RestrictToRepacking" /> </DesignRestrictions> <DesignRestrictions name="not_surface"> <Action selector_logic="NOT surface" residue_level_operations="RestrictToRepacking" /> </DesignRestrictions> <DesignRestrictions name="not_boundary"> <Action selector_logic="NOT boundary" residue_level_operations="RestrictToRepacking" /> </DesignRestrictions> <ProhibitSpecifiedBaseResidueTypes name="prohibit_a_in_core" base_types="ARG,LYS,ASP,GLU,CYS,PRO" selector="core" /> <ProhibitSpecifiedBaseResidueTypes name="prohibit_a_in_boundary" base_types="ARG,LYS,ASP,GLU,CYS,PRO,MET,HIS" selector="boundary" /> <ProhibitSpecifiedBaseResidueTypes name="prohibit_a_in_surface" base_types="ILE,LEU,VAL,PHE,TRP,MET" selector="surface" /> </TASKOPERATIONS> <SIMPLE_METRICS> <SequenceMetric name="record_sequence" output_mode="basename" /> </SIMPLE_METRICS> <MOVERS> <Small name="small_move" scorefxn="r15_regular" temperature="0.5" nmoves="1000" angle_max="2.0" preserve_detailed_balance="0" /> <FastDesign name="design_core" scorefxn="r15_design" disable_design="false" task_operations="not_core,prohibit_a_in_core,extra_sample_rotamers_design" repeats="4" relaxscript="default" min_type="lbfgs_armijo_nonmonotone" /> 33 Protein Design with Agent Rosetta <FastDesign name="design_boundary" scorefxn="r15_design" disable_design="false" task_operations="not_boundary,prohibit_a_in_boundary,extra_sample_rotamers_design" repeats="2" relaxscript="default" min_type="lbfgs_armijo_nonmonotone" /> <FastDesign name="design_surface" scorefxn="r15_design" disable_design="false" task_operations="not_surface,prohibit_a_in_surface,extra_sample_rotamers_design" repeats="2" relaxscript="default" min_type="lbfgs_armijo_nonmonotone" /> <AddCompositionConstraintMover name="comp_core" filename="comp_core" selector="core" /> <AddCompositionConstraintMover name="comp_boundary" filename="comp_boundary" selector="boundary" /> <AddCompositionConstraintMover name="comp_surface" filename="comp_surface" selector="surface" /> <AddConstraints name="geom_constraint"> <AtomPairConstraintGenerator name="gen_geom_csts" ca_only="1" use_harmonic="1" native="1" /> </AddConstraints> <RemoveConstraints name="rm_geom_csts" constraint_generators="gen_geom_csts" /> <FastRelax name="fast_relax" scorefxn="r15_regular" disable_design="true" task_operations="include_current_rotamer,extra_sample_rotamers_relax" repeats="3" relaxscript="default" min_type="lbfgs_armijo_nonmonotone" /> </MOVERS> <PROTOCOLS> <Add mover="geom_constraint" /> <Add mover="small_move" /> <Add mover="comp_core" /> <Add mover="design_core" /> <Add mover="comp_boundary" /> <Add mover="design_boundary" /> <Add mover="comp_surface" /> <Add mover="design_surface" /> <Add mover="rm_geom_csts" /> <Add mover="fast_relax" /> <Add metrics="record_sequence" /> </PROTOCOLS> <OUTPUT scorefxn="r15_design" /> </ROSETTASCRIPTS> In both protocols, compcore, compboundary, and compsurface are: Box 11: compcore PENALTY_DEFINITION TYPE ALA FRACTION 0.10 FRACT_DELTA_START -0.05 FRACT_DELTA_END 0.05 BEFORE_FUNCTION CONSTANT Box 11: compboundary PENALTY_DEFINITION TYPE ALA FRACTION 0.10 FRACT_DELTA_START -0.05 FRACT_DELTA_END 0.05 BEFORE_FUNCTION CONSTANT Box 11: compsurface PENALTY_DEFINITION TYPE ALA FRACTION 0.10 FRACT_DELTA_START -0.05 FRACT_DELTA_END 0.05 BEFORE_FUNCTION CONSTANT 34 Protein Design with Agent Rosetta AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION PENALTY_DEFINITION TYPE GLY FRACTION 0.05 FRACT_DELTA_START -0.05 FRACT_DELTA_END 0.05 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION PENALTY_DEFINITION PROPERTIES AROMATIC FRACTION 0.15 FRACT_DELTA_START -0.05 FRACT_DELTA_END 0.05 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION PENALTY_DEFINITION TYPE MET ABSOLUTE 1 DELTA_START -1 DELTA_END 1 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION PENALTY_DEFINITION PROPERTIES POLAR FRACTION 0.15 FRACT_DELTA_START -0.05 FRACT_DELTA_END 0.05 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION PENALTY_DEFINITION TYPE GLY FRACTION 0.05 FRACT_DELTA_START -0.05 FRACT_DELTA_END 0.05 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION PENALTY_DEFINITION PROPERTIES HYDROPHOBIC FRACTION 0.50 FRACT_DELTA_START -0.05 FRACT_DELTA_END 0.05 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION PENALTY_DEFINITION PROPERTIES AROMATIC FRACTION 0.10 FRACT_DELTA_START -0.05 FRACT_DELTA_END 0.05 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION PENALTY_DEFINITION TYPE TRP ABSOLUTE 1 DELTA_START -1 DELTA_END 1 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION PENALTY_DEFINITION TYPE GLY FRACTION 0.05 FRACT_DELTA_START -0.05 FRACT_DELTA_END 0.05 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION C.5 Extended Results In this section, we include further results that were omitted from the main text for the sake of presentation. 35 Protein Design with Agent Rosetta 151015 Design step 0.6 0.8 1.0 1.2 ESMFold RMSD to Init (Ă ) 151015 Design step 0.84 0.86 0.88 0.90 0.92 0.94 0.96 ESMFold CA pLDDT PDB: 9IFR (123 residues) Qwen3 Instruct (CoT) Gemini 2.5 Flash GPT-5 Sonnet 4.5 ProteinMPNN human baseline (one shot) human baseline (staged) naive baseline 151015 Design step 0.75 0.80 0.85 0.90 ESMFold RMSD to Init (Ă ) 151015 Design step 0.84 0.86 0.88 0.90 0.92 0.94 0.96 ESMFold CA pLDDT PDB: 9PL1 (74 residues) 151015 Design step 0.30 0.35 0.40 0.45 0.50 ESMFold RMSD to Init (Ă ) 151015 Design step 0.84 0.86 0.88 0.90 0.92 0.94 0.96 ESMFold CA pLDDT PDB: 9VYW (92 residues) 151015 Design step 0.30 0.35 0.40 0.45 0.50 ESMFold RMSD to Init (Ă ) 151015 Design step 0.84 0.86 0.88 0.90 0.92 0.94 0.96 ESMFold CA pLDDT PDB: 9VYQ (92 residues) 151015 Design step 0.4 0.5 0.6 0.7 0.8 0.9 ESMFold RMSD to Init (Ă ) 151015 Design step 0.84 0.86 0.88 0.90 0.92 0.94 0.96 ESMFold CA pLDDT PDB: 9C14 (97 residues) 151015 Design step 0.3 0.4 0.5 0.6 0.7 0.8 0.9 ESMFold RMSD to Init (Ă ) 151015 Design step 0.84 0.86 0.88 0.90 0.92 0.94 0.96 ESMFold CA pLDDT PDB: 8YS1 (105 residues) 151015 Design step 1.0 1.5 2.0 2.5 3.0 3.5 ESMFold RMSD to Init (Ă ) 151015 Design step 0.84 0.86 0.88 0.90 0.92 0.94 0.96 ESMFold CA pLDDT PDB: 8VWO (112 residues) 151015 Design step 0.5 1.0 1.5 2.0 2.5 ESMFold RMSD to Init (Ă ) 151015 Design step 0.84 0.86 0.88 0.90 0.92 0.94 0.96 ESMFold CA pLDDT PDB: 9KGY (125 residues) Figure C.2: Running best (i.e., over all previous steps) ESMFold RMSD and pLDDT across 1,000 bootstrap samples of 8 trials out of the 16 total. 0.00.20.40.6 Cost (USD) 0.6 0.7 0.8 0.9 1.0 1.1 ESMFold RMSD to Init (Ă ) 0.00.20.40.6 Cost (USD) 0.89 0.90 0.91 0.92 0.93 0.94 ESMFold CA pLDDT action success rate 82%86%90% fraction of reasoning tokens 65% 70%90% Figure C.3: Summary of median ESMFold RMSD and pLDDT across all 8 target backbone conformations used in this experiment. 9PL19VYW9VYQ9C14 MethodRMSD to init (m Ì A)pLDDT (0-100)RMSD to init (m Ì A)pLDDT (0-100)RMSD to init (m Ì A)pLDDT (0-100)RMSD to init (m Ì A)pLDDT (0-100) Qwen3 Instruct (CoT)798 (-3/+30)91.3 (-1.1/+1.1)349 (-3/+24)93.5 (-0.1/+0.2)321 (-26/+14)94.3 (-0.9/+0.0)481 (-23/+24)92.7 (-2.9/+1.4) Gemini 2.5 Flash758 (-11/+20)92.5 (-2.4/+1.2)324 (-17/+12)94.1 (-0.2/+0.2)306 (-16/+10)94.4 (-0.5/+0.2)421 (-14/+41)94.0 (-1.3/+0.6) GPT-5727 (-3/+59)91.5 (-1.4/+0.5)351 (-4/+19)92.9 (-0.4/+0.7)326 (-9/+28)94.5 (-0.8/+0.2)489 (-36/+19)93.8 (-3.4/+0.3) Sonnet 4.5788 (-16/+39)91.8 (-1.9/+1.2)341 (-26/+14)94.4 (-1.4/+0.2)305 (-1/+19)94.3 (-0.3/+0.1)462 (-23/+25)91.9 (-1.2/+1.6) ProteinMPNN822 (-6/+13)92.5 (-0.1/+0.4)476 (-12/+12)92.5 (-0.5/+0.4)388 (-4/+8)94.0 (-0.4/+0.0)388 (-2/+11)95.9 (-0.4/+0.2) human baseline (one-shot)831 (-8/+45)89.3 (-0.5/+0.6)468 (-11/+13)92.2 (-0.7/+0.4)448 (-7/+7)92.6 (-0.3/+0.1)840 (+0/+2678)86.9 (-1.7/+0.0) human baseline (staged)822 (-8/+42)90.1 (-0.6/+1.3)539 (-34/+60)88.7 (-0.5/+0.7)494 (-40/+35)90.0 (-0.7/+0.2)882 (+0/+0)85.4 (+0.0/+0.0) na Ìıve baseline839 (-19/+28)91.1 (-0.8/+0.9)422 (-5/+43)93.6 (-0.5/+0.8)485 (-5/+33)93.4 (-1.7/+0.5)518 (-24/+54)91.4 (-0.6/+1.0) 8YS18VWO9IFR9KGY MethodRMSD to init (m Ì A)pLDDT (0-100)RMSD to init (m Ì A)pLDDT (0-100)RMSD to init (m Ì A)pLDDT (0-100)RMSD to init (m Ì A)pLDDT (0-100) Qwen3 Instruct (CoT)377 (-13/+12)95.1 (-1.5/+0.1)1,263 (-88/+260)86.9 (-1.7/+2.0)643 (-29/+26)94.1 (-3.1/+0.7)618 (-11/+120)90.7 (-2.6/+0.8) Gemini 2.5 Flash336 (-22/+28)94.4 (-0.4/+0.8)1,245 (-50/+68)87.3 (-0.8/+1.3)597 (-38/+33)94.4 (-0.4/+0.6)582 (-20/+47)90.2 (-0.5/+1.5) GPT-5368 (-15/+19)94.3 (-1.1/+0.1)1,554 (-0/+102)87.2 (-1.5/+0.5)757 (-49/+38)91.3 (-0.3/+0.6)538 (-42/+50)90.7 (-3.0/+1.8) Sonnet 4.5377 (-18/+16)94.8 (-0.3/+0.5)1,374 (-164/+176)85.9 (-0.4/+3.1)718 (-25/+26)92.5 (-0.9/+1.3)585 (-29/+89)90.5 (-1.9/+1.0) ProteinMPNN370 (-3/+2)95.3 (-0.0/+0.1)1,023 (-44/+33)89.2 (-0.5/+0.1)610 (-20/+20)96.6 (-0.1/+0.1)483 (-5/+15)94.2 (-0.1/+0.2) human baseline (one-shot)819 (-9/+16)90.5 (-0.2/+0.8)2,956 (+0/+380)87.4 (-1.8/+0.0)850 (-0/+8)90.1 (-0.2/+0.3)669 (-35/+84)89.5 (-0.4/+0.7) human baseline (staged)910 (-19/+39)89.3 (-0.7/+2.3)3,416 (+0/+0)88.0 (+0.0/+0.0)1,226 (+0/+464)89.1 (-2.6/+0.0)735 (-19/+143)87.9 (-1.2/+0.7) na Ìıve baseline562 (-4/+27)94.9 (-2.6/+0.5)1,420 (-79/+145)89.1 (-3.1/+0.3)845 (-2/+30)92.5 (-1.1/+0.4)2,788 (+0/+0)85.8 (+0.0/+0.0) Table C.1: Per-PDB results for all methods on fixed-backbone sequence design with canonical amino acids only. For each method, we run 16 independent replicates, and report the best step-wise median and 95% confidence interval across 1000 bootstrap samples of 8 trajectories out of the 16 total. We select the best step in terms of the smallest 5 th percentile of the RMSD between the ESMFold predicted fold and the target conformation. To compute step-wise statistics, we only consider designs with pLDDT â„ 85 to remove outliers. 36 Protein Design with Agent Rosetta Figure C.4: Visualization of the best designs for each method on all backbones used in the fixed-backbone sequence design experiment. 37 Protein Design with Agent Rosetta MethodCost ($)Rosetta action success rate (%) Fraction of reasoning tokens (%) RMSD to init (m Ì A) pLDDT (0-100) ProteinMPNNnanana570 ± 23793.8 ± 2.4 Gemini 2.5 Flash0.13 ± 0.0191 ± 393 ± 1571 ± 31692.7 ± 2.6 Qwen3 Instruct (CoT)0.03 ± 0.0186 ± 570 ± 3606 ± 31392.3 ± 2.7 Sonnet 4.50.63 ± 0.1193 ± 164 ± 2619 ± 35292.1 ± 2.9 GPT-50.33 ± 0.0192 ± 269 ± 1639 ± 40592.0 ± 2.6 na Ìıve baselinena100na985 ± 79791.5 ± 2.9 human baseline (one-shot)na100na985 ± 81389.8 ± 2.0 human baseline (staged)na100na1,128 ± 95288.6 ± 1.5 Table C.2: Summary of results across all 8 PDBs for all methods on fixed-backbone sequence design. We report mean and standard deviation of the best-of-8 median calculated across bootstraps samples in Table C.1. 38 Protein Design with Agent Rosetta DInclusion of a Non-Canonical Amino Acid in the Core of a Protein In this section, we include further experimental details on including N1-formyl-tryptophan (TRF)âa post-translational modification of tryptophanâin the core of an initial protein. D.1 Design Brief Here, we include the specific design brief used in this experiment. This design brief is composed at runtime with the general design brief prompt template presented in Section B.2. Box 12: Design brief You are given an initial protein composed of standard (i.e., canonical) amino acids. Your task is to redesign the sequence such that: 1. It includes exactly one (1) non-canonical amino acid in its core. 2. The inclusion of the non-canonical amino acid does not compromise the energetic stability of the protein. 3. The sequence folds into the same shape as the initial protein. The non-canonical amino acid (NCAA) to include is N1-formyl-tryptophan, a tryptophan with a formyl group attached to the nitrogen in the indole ring. The 3-letter code for this NCAA is TRF. You will be provided the following proxy metrics to help you assess task progress: 1. Total Rosetta energy. 2. Cavity volume and radius of gyration. 3. The RMSD between the Rosetta design and the initial protein. At each step: - The Rosetta environment will execute your action and produce an ensemble of candidate designs. - You will be shown summary statistics of the ensemble. - You will choose the next action to take, which will be applied to Pareto optimal candidates only in terms of the provided proxy metrics. D.2 Description of TRF In this section, we compare the structure of TRF with the canonical amino acid TRP, and include the params file used to define TRF in Rosetta. NH 2 (S) O N H OH (a) TRP. NH 2 (S) N O O OH (b) TRF. Figure D.1: Comparison of the structure of tryptophan (TRP) and N1-formyl-tryptophan (TRF). 39 Protein Design with Agent Rosetta Box 13: TRF params file # Rosetta residue topology file for N1-formyl-tryptophan (aka TRF) NAME TRF IO_STRING TRF X TYPE POLYMER A UNK ROTAMER_A TRP BACKBONE_A TRP ATOM N Nbb NH1 -0.5872428 -0.317 ATOM CA CAbb CT1 0.1074333 0.133 ATOM C CObb C 0.7058698 0.583 ATOM O OCbb O -0.6711044 -0.517 ATOM CB CH2 CT2 -0.1388399 0.158 ATOM CG CH0 CY -0.0086544 -0.092 ATOM CD1 aroC CA 0.0477593 -0.092 ATOM CD2 CH0 CPT 0.0000246 0.033 ATOM NE1 Ntrp NY -0.5120383 -0.367 ATOM CE2 CH0 CPT 0.1302101 0.033 ATOM CE3 aroC CA -0.0824262 -0.092 ATOM CZ1 COO C 0.4512171 0.133 ATOM CZ2 aroC CA -0.0824262 -0.092 ATOM CZ3 aroC CA -0.0824262 -0.092 ATOM OH1 OOC OC -0.5773882 -0.517 ATOM CH2 aroC CA -0.0824262 -0.092 ATOM H HNbb H 0.4161782 0.283 ATOM HD1 Haro HP 0.1171916 0.158 ATOM HZ1 Hapo HA 0.0561730 0.033 ATOM HZ2 Haro HP 0.1171916 0.158 ATOM H2 Haro HP 0.1171916 0.158 ATOM HZ3 Haro HP 0.1171916 0.158 ATOM HE3 Haro HP 0.1171916 0.158 ATOM HA Hapo HB 0.1331620 0.033 ATOM 1HB Hapo HA 0.0954940 0.033 ATOM 2HB Hapo HA 0.0954940 0.033 ATOM_ALIAS 1HB HB2 ATOM_ALIAS 2HB HB3 LOWER_CONNECT N UPPER_CONNECT C BOND N CA BOND N H BOND CA C BOND CA CB BOND CA HA BOND_TYPE C O 2 BOND CB CG BOND CB 1HB BOND CB 2HB BOND_TYPE CG CD1 ARO BOND_TYPE CG CD2 ARO CUT_BOND CG CD2 BOND_TYPE CD1 NE1 ARO BOND CD1 HD1 BOND_TYPE CD2 CE2 ARO CUT_BOND CD2 CE2 40 Protein Design with Agent Rosetta BOND_TYPE CD2 CE3 ARO BOND_TYPE NE1 CE2 ARO BOND_TYPE CE2 CZ2 ARO BOND_TYPE CE3 CZ3 ARO BOND CE3 HE3 BOND_TYPE CZ2 CH2 ARO BOND CZ2 HZ2 BOND_TYPE CZ3 CH2 ARO BOND CZ3 HZ3 BOND CH2 H2 BOND NE1 CZ1 BOND_TYPE CZ1 OH1 2 BOND CZ1 HZ1 CHI 1 N CA CB CG CHI 2 CA CB CG CD1 CHI 3 CD1 NE1 CZ1 OH1 PROTON_CHI 3 SAMPLES 2 0 180 EXTRA 1 15 ADD_RING 1 AROMATIC CG CD1 NE1 CE2 CD2 NU 1 CD2 CG CD1 NE1 NU 2 CG CD1 NE1 CE2 NU 3 CD1 NE1 CE2 CD2 NU 4 NE1 CE2 CD2 CG NU 5 CE2 CD2 CG CD1 ADD_RING 2 AROMATIC CE2 CD2 CE3 CZ3 CH2 CZ2 NU 6 CZ2 CE2 CD2 CE3 NU 7 CE2 CD2 CE3 CZ3 NU 8 CD2 CE3 CZ3 CH2 NU 9 CE3 CZ3 CH2 CZ2 NU 10 CZ3 CH2 CZ2 CE2 NU 11 CH2 CZ2 CE2 CD2 PROPERTIES PROTEIN ALPHA_A L_A HYDROPHOBIC AROMATIC SC_ORBITALS METALBINDING CYCLIC METAL_BINDING_ATOMS O NBR_ATOM CB # APL CB to sidechain heavyatom distance; swept all chi combos at 5 degree intervals NBR_RADIUS 5.37514 FIRST_SIDECHAIN_ATOM CB RAMA_PREPRO_FILENAME all.ramaProb prepro.ramaProb ACT_COORD_ATOMS CD2 CE3 END ICOOR_INTERNAL N 0.000000 0.000000 0.000000 N CA C ICOOR_INTERNAL CA 0.000000 180.000000 1.458001 N CA C ICOOR_INTERNAL C 0.000000 68.800003 1.523258 CA N C ICOOR_INTERNAL UPPER 149.999985 63.800007 1.328685 C CA N ICOOR_INTERNAL O -180.000000 59.200005 1.231015 C CA UPPER ICOOR_INTERNAL CB -122.800000 69.625412 1.521736 CA N C ICOOR_INTERNAL CG 0.000068 66.465866 1.498746 CB CA N ICOOR_INTERNAL CD1 0.000148 53.263374 1.362720 CG CB CA ICOOR_INTERNAL NE1 -179.969818 69.843536 1.372938 CD1 CG CB ICOOR_INTERNAL CE2 -0.111805 71.100000 1.372132 NE1 CD1 CG ICOOR_INTERNAL CZ2 179.958145 49.861912 1.385949 CE2 NE1 CD1 ICOOR_INTERNAL CH2 -179.974182 62.500000 1.395024 CZ2 CE2 NE1 ICOOR_INTERNAL CZ3 0.127723 58.500000 1.372108 CH2 CZ2 CE2 ICOOR_INTERNAL CE3 -0.067888 59.000000 1.389848 CZ3 CH2 CZ2 ICOOR_INTERNAL CD2 0.021383 61.268948 1.400378 CE3 CZ3 CH2 ICOOR_INTERNAL HE3 -179.851593 59.300000 1.089539 CE3 CZ3 CD2 ICOOR_INTERNAL HZ3 179.965530 59.271282 1.090294 CZ3 CH2 CE3 ICOOR_INTERNAL H2 179.952423 60.590862 1.090289 CH2 CZ2 CZ3 41 Protein Design with Agent Rosetta ICOOR_INTERNAL HZ2 -179.891754 57.900000 1.090237 CZ2 CE2 CH2 ICOOR_INTERNAL CZ1 179.658524 54.960213 1.373036 NE1 CD1 CE2 ICOOR_INTERNAL OH1 179.658524 54.960213 1.373036 CZ1 NE1 CD1 ICOOR_INTERNAL HZ1 179.658524 54.960213 1.373036 CZ1 NE1 OH1 ICOOR_INTERNAL HD1 179.990616 55.100000 1.088516 CD1 CG NE1 ICOOR_INTERNAL 1HB 121.200000 70.500000 1.090167 CB CA CG ICOOR_INTERNAL 2HB 117.600000 70.500000 1.089792 CB CA 1HB ICOOR_INTERNAL HA -119.000000 71.500000 1.089883 CA N CB ICOOR_INTERNAL LOWER -150.000000 58.300003 1.328685 N CA C ICOOR_INTERNAL H -180.000000 60.849998 1.010000 N CA LOWER D.3 Initial Proteins In this section, we describe the inclusion criteria and preprocessing steps used to select 4 proteins to use in this experiment (see Fig. D.2). Figure D.2: The 4 proteins for non-canonical design. Selection of PDB structures for NCAA design. We selected a small set of structurally distinct de novo proteins to enable evaluation of NCAA incorporation across diverse structural contexts and to assess whether incorporation of non-canonical amino acids can further stabilize existing de novo designs. The selected structures were: 6V67 [22], 1QYS [3], 8UZL [14], and 7SQ3 [37]. PDB 6V67 is a compact mini-protein (40 residues) that provides a minimal system for probing NCAA compatibility in highly constrained backbones. 1QYS is a 106-residue Rosetta-designed protein that serves as a foundational reference for computational protein design. 8UZL is a designed transmembraneÎČ-barrel (149 residues), serving as a scaffold for NCAA design in membrane-associated contexts. Finally, 7SQ3 is a trefoil-knot protein (153 residues) representing a topologically complex fold and a stringent test for design methods. Preprocessing of PDB structures. All structures chosen for non-canonical amino acid design were processed in PyMOL [106,105] to remove alternative conformations (altlocs B-G) and all heteroatoms and X-ray crystallographic copies. D.4 Environment State Here, we include an illustrative example of the textual representation of the RosettaScripts environment state for the task with a non-canonical residue. We omit the tabular history summary for the sake of readability 42 Protein Design with Agent Rosetta Box 14: Environment state The RosettaScripts environment successfully executed your action: **Step Number:** 3 **Action Name:** rotamer_change **Results:** - Number of designs: 120 (21 Pareto optimal) - Average total Rosetta energy: -154.82 ± 6.62 - Average compositional before design 733.33, and after design 16.67 (-716.67) - TRF inclusion summary: -- List of core residue indices after design: 5,23,26,30,37 -- Percentage of designs with at least one TRF residue: 100.00% (min: 12, max: 18) -- Percentage of designs with exactly one TRF residue in the core: 13.33% (min: 1, max: 4) -- Most common TRF residue positions: 8,25,37,35,30,11,13,39,5,27,3,15,32,23,19,20,24,28,21 -- Average interresidue_repulsion at TRF residues: 1.01 ± 0.59 -- Average ramachandran_preference at TRF residues: 0.05 ± 0.29 - Top 5 most common outlier residue types. A residue is an outlier if its energy term is above the 90-th quantile of the per-residue energy term for that design. For each outlier residue type, we include the most common outlier positions along the sequence (positions are 1-based): -- interresidue_repulsion: Average per-sequence 90-th quantile: 1.41 TRF: 100.00% (positions: 25,5,15,39,13,33) MET: 62.50% (positions: 40,26) LEU: 33.33% (positions: 28,15,14) CYS: 21.67% (positions: 4,29) TYR: 20.00% (positions: 33) -- ramachandran_preference: Average per-sequence 90-th quantile: 0.58 TRF: 100.00% (positions: 8,16,10) ASN: 96.67% (positions: 9,17) MET: 73.33% (positions: 40) CYS: 63.33% (positions: 6) ASP: 20.00% (positions: 18) - Structural metrics: Cavity volume ( Ì A^3) : 13.07 ± 11.98 Radius of gyration ( Ì A) : 9.90 ± 0.01 Penalty for buried unsatisfied Hydrogen bonds: 142.08 ± 108.35 RMSD to initial structure ( Ì A) : 0.19 ± 0.02 --- 43 Protein Design with Agent Rosetta **History Summary:** history_summary D.5 Expert Human Baseline Protocol In this section, we include the expert human written protocol used as a baseline in the design task with non-canonical amino acids. Box 15: Human design protocol <ROSETTASCRIPTS> <SCOREFXNS> <ScoreFunction name="r15_regular" weights="ref2015_cst.wts"/> <ScoreFunction name="r15_design" weights="ref2015.wts"> <Reweight scoretype="a_composition" weight="1.0"/> <Reweight scoretype="atom_pair_constraint" weight="1.0"/> </ScoreFunction> <ScoreFunction name="r15_post" weights="ref2015_cst.wts"> <Reweight scoretype="a_composition" weight="1.0"/> <Reweight scoretype="atom_pair_constraint" weight="1.0"/> <Reweight scoretype="rg" weight="1.0" /> <Reweight scoretype="buried_unsatisfied_penalty" weight="1.0" /> </ScoreFunction> </SCOREFXNS> <PACKER_PALETTES> <CustomBaseTypePackerPalette name="ncaa_palette" additional_residue_types="TRF"/> </PACKER_PALETTES> <RESIDUE_SELECTORS> <Layer name="core" select_core="true" select_boundary="false" select_surface="false" use_sidechain_neighbors="true"/> </RESIDUE_SELECTORS> <RESIDUE_LEVEL_TASK_OPERATIONS> <RestrictToRepackingRLT name="RestrictToRepacking"/> </RESIDUE_LEVEL_TASK_OPERATIONS> <TASKOPERATIONS> <IncludeCurrent name="include_current_rotamer"/> <ExtraRotamersGeneric name="extra_sample_rotamers_design" ex1="1" ex1aro="1"/> <ExtraRotamersGeneric name="extra_sample_rotamers_relax" ex1="1" ex2="1" ex1aro="1" ex2aro="1"/> <DesignRestrictions name="not_core"> <Action selector_logic="NOT core" residue_level_operations="RestrictToRepacking"/> </DesignRestrictions> <ProhibitSpecifiedBaseResidueTypes name="limited_a_design" base_types="ARG,LYS,ASP,GLU,CYS,PRO"/> </TASKOPERATIONS> <SIMPLE_METRICS> <SequenceMetric name="record_sequence" output_mode="basename"/> </SIMPLE_METRICS> 44 Protein Design with Agent Rosetta <FILTERS> <CavityVolume name="cav_vol" confidence="0.0" /> </FILTERS> <MOVERS> <Small name="small_move" scorefxn="r15_design" temperature="0.5" nmoves="1000" angle_max="2.0" preserve_detailed_balance="0"/> <FastDesign name="design_with_ncaa_only" scorefxn="r15_design" disable_design="false" task_operations="not_core,limited_a_design,include_current_rotamer, extra_sample_rotamers_design" packer_palette="ncaa_palette" repeats="5" relaxscript="default" min_type="lbfgs_armijo_nonmonotone"/> <AddCompositionConstraintMover name="comp_ncaa" filename="comp_ncaa" selector="core"/> <AddCompositionConstraintMover name="comp_core" filename="comp_core" selector="core"/> <AddConstraints name="geom_constraint"> <AtomPairConstraintGenerator name="gen_geom_csts" ca_only="1" use_harmonic="1" native="1"/> </AddConstraints> <RemoveConstraints name="rm_geom_csts" constraint_generators="gen_geom_csts"/> <FastRelax name="fast_relax" scorefxn="r15_regular" disable_design="true" task_operations="include_current_rotamer,extra_sample_rotamers_relax" repeats="3" relaxscript="default" min_type="lbfgs_armijo_nonmonotone"/> </MOVERS> <PROTOCOLS> <Add mover="geom_constraint"/> <Add mover="small_move"/> <Add mover="comp_ncaa"/> <Add mover="comp_core"/> <Add mover="design_with_ncaa_only"/> <Add mover="rm_geom_csts"/> <Add mover="fast_relax"/> <Add filter="cav_vol"/> <Add metrics="record_sequence"/> </PROTOCOLS> <OUTPUT scorefxn="r15_post"/> </ROSETTASCRIPTS> where compncaa and compcore are: Box 16: compncaa PENALTY_DEFINITION TYPE TRF ABSOLUTE 1 PENALTIES 1000 0 100 DELTA_START -1 DELTA_END 1 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC Box 16: compcore PENALTY_DEFINITION TYPE ALA FRACTION 0.10 FRACT_DELTA_START -0.05 FRACT_DELTA_END 0.05 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 45 Protein Design with Agent Rosetta END_PENALTY_DEFINITION END_PENALTY_DEFINITION PENALTY_DEFINITION TYPE GLY FRACTION 0.05 FRACT_DELTA_START -0.05 FRACT_DELTA_END 0.05 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION PENALTY_DEFINITION PROPERTIES AROMATIC FRACTION 0.15 FRACT_DELTA_START -0.05 FRACT_DELTA_END 0.05 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION PENALTY_DEFINITION TYPE MET ABSOLUTE 1 DELTA_START -1 DELTA_END 1 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION PENALTY_DEFINITION PROPERTIES POLAR FRACTION 0.15 FRACT_DELTA_START -0.05 FRACT_DELTA_END 0.05 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION QUADRATIC PENALTIES 0 0 25 END_PENALTY_DEFINITION 46 Protein Design with Agent Rosetta D.6 Extended Results In this section, we include further results that were omitted by the main text for the sake of presentation. 0.00.20.40.6 Cost (USD) 1.0 1.5 2.0 2.5 3.0 AF3 RMSD to native (Ă ) 0.00.20.40.6 Cost (USD) 0.675 0.700 0.725 0.750 0.775 0.800 0.825 AF3 pLDDT action success rate 54% 80% 84% 90% fraction of reasoning tokens 65% 71% 80% 90% Figure D.3: Summary of results for including 1 TRF residue in the core of an input protein. We report the average cost of one run with 30 model queries, the average action success rate, and the fraction of output tokens that were reasoning. Figure D.4: Illustrations of good and bad designs for inclusion of TRF in the core of a given protein across all 4 PDBs considered in the experiment. 47 Protein Design with Agent Rosetta 1QYS6V67 MethodRosetta RMSD (m Ì A) AF3 RMSD (m Ì A) AF3 pLDDT (0-100) Rosetta RMSD (m Ì A) AF3 RMSD (m Ì A) AF3 pLDDT (0-100) Qwen3 Instruct (CoT)174 ± 4627 ± 11986.6 ± 4.0111 ± 2444 ± 18979.0 ± 3.1 Gemini 2.5 Flash170 ± 10798 ± 67785.9 ± 2.3109 ± 2347 ± 3878.2 ± 0.8 GPT-5166 ± 1796 ± 8184.6 ± 2.2104 ± 2374 ± 6081.0 ± 1.7 Sonnet 4.5157 ± 5773 ± 25282.3 ± 1.2114 ± 3604 ± 14378.5 ± 1.6 human baseline182 ± 3985 ± 15884.8 ± 2.9181 ± 6476 ± 083.7 ± 0.0 7SQ38UZL MethodRosetta RMSD (m Ì A) AF3 RMSD (m Ì A) AF3 pLDDT (0-100) Rosetta RMSD (m Ì A) AF3 RMSD (m Ì A) AF3 pLDDT (0-100) Qwen3 Instruct (CoT)221 ± 65,280 ± 5,10256.2 ± 22.5276 ± 186,182 ± 4,21245.2 ± 13.6 Gemini 2.5 Flash209 ± 72,394 ± 3,81475.2 ± 20.8168 ± 111,008 ± 99877.4 ± 10.9 GPT-5215 ± 2652 ± 3885.7 ± 4.1164 ± 61,599 ± 2,04677.9 ± 13.8 Sonnet 4.5205 ± 63,744 ± 4,99859.6 ± 16.6208 ± 171,941 ± 2,86263.0 ± 7.3 human baseline241 ± 46,906 ± 8,20667.3 ± 27.7nanana Table D.3: Per-PDB results for all methods on inclusion of a TRF residue in the core of an input protein structure. For each method, we filter successful designs, and select the top 10 in order of increasing RMSD between the Rosetta Pose and the native PDB. We validate the fold of the selected designs with AF3, and report RMSD between the predicted fold and the native structure as well as pLDDT. We report average and standard deviation over the top 10 designs. The human baseline is marked as ânaâ for 8UZL because it fails to include any TRF residues. MethodCost (USD) Rosetta action success rate (%) Fraction of reasoning tokens (%) Rosetta RMSD to native (m Ì A) AF3 RMSD to native (m Ì A) AF3 pLDDT (0-100) GPT-50.34 ± 0.0384.0 ± 5.281.0 ± 1.2162 ± 45855 ± 52682.2 ± 3.6 Gemini 2.5 Flash0.15 ± 0.0180.1 ± 7.588.1 ± 1.5164 ± 411,137 ± 88279.2 ± 4.7 Sonnet 4.50.70 ± 0.0189.1 ± 4.065.2 ± 2.6171 ± 451,766 ± 1,44770.8 ± 11.2 human baselinenanana201 ± 342,789 ± 3,57478.6 ± 9.8 Qwen3 Instruct (CoT)0.03 ± 0.0153.4 ± 22.971.0 ± 2.0195 ± 703,133 ± 3,02366.8 ± 19.3 Table D.4: Summary of results across all 4 PDBs for all methods on inclusion of 1 TRF residue in the core of a given protein. We report mean and standard deviation of results in Table D.3. 48 Protein Design with Agent Rosetta E RosettaScripts Environment Action Types In this section, we include further details on the action types available to Agent Rosetta. For each action, we include the RosettaScripts XML template which is populated at runtime. In the templates: âąFields enclosed within curly brackets (e.g.,fieldname) are populated with the parameters generated by the agent. âąFields enclosed within squared brackets (e.g.,[fieldname]) are populated programmatically by the environment, depending on the action and the user-defined task parameters. We remark that the gobackto action does not call Rosetta and it does not need an XML template. E.1 rotamerchange The RosettaScripts XML template for the rotamerchange action is: Box 17: rotamerchange XML action template <ROSETTASCRIPTS> <SCOREFXNS> <ScoreFunction name="scoring_pre" weights="ref2015"> <Reweight scoretype="a_composition" weight="1.0" /> </ScoreFunction> <ScoreFunction name="guidance" weights="ref2015"> [guidance_weights] </ScoreFunction> <ScoreFunction name="scoring_post" weights="ref2015"> [scoring_weights] </ScoreFunction> </SCOREFXNS> <PACKER_PALETTES> <CustomBaseTypePackerPalette name="design_palette" [additional_residue_types] /> </PACKER_PALETTES> <RESIDUE_SELECTORS> <Layer name="_core_residues" select_core="true" select_boundary="false" select_surface="false"/> residue_selectors </RESIDUE_SELECTORS> <TASKOPERATIONS> task_operations <IncludeCurrent name="include_current" /> </TASKOPERATIONS> <SIMPLE_METRICS> [simple_metrics] <TotalEnergyMetric name="a_composition_pre" scorefxn="scoring_pre" scoretype="a_composition"/> <SequenceMetric name="record_sequence" output_mode="basename" /> </SIMPLE_METRICS> <FILTERS> [filters] 49 Protein Design with Agent Rosetta </FILTERS> <MOVERS> a_comp_mover <AddConstraints name="geom_constraint"> <AtomPairConstraintGenerator name="gen_geom_csts" ca_only="1" use_harmonic="1" native="1" /> </AddConstraints> <FastDesign name="design" scorefxn="guidance" repeats="1" packer_palette="design_palette" task_operations="include_current[operations_names]" /> </MOVERS> <PROTOCOLS> [protocol] <Add metrics="record_sequence" /> </PROTOCOLS> <OUTPUT scorefxn="scoring_post"/> </ROSETTASCRIPTS> The agent-generated parameters of the action are: âą a compmover: the amino acid compositional penalty blocks guiding the Monte Carlo search during Rosettaâs FastDesign Mover. âą residueselectors: the definitions of regions of the protein sequence compositional penalty blocks and task operations should apply to. âą task operations: the design restrictions which define the base residue types allowed or prohibited in different regions of the protein. The user-specified parameters of the action are: âą [guidance weights], which edit the weights of the energy function used during design. For example, in our experiments, we set aacomposition and atompairconstraint to 1. âą [scoringweights], which edit the weights of the energy function used for scoring post-design. These weights are a superset of the guidance terms, and they may include terms likergto compute the radius of gyration of the output protein, orburiedunsatisfiedpenaltyto estimate the number of buried unsatisfied hydrogen bonds. âą [additional residuetypes], which include the additional non-canonical amino acids to use (e.g., TRF in our experiment). âą [simplemetrics], which include further evaluation terms such as the post-design amino acid compositional penalty energy or the per-residue stratification of the van der Waals force and the pre-proline Ramachandran preference. âą [filters], which specify additional quality metrics such as cavity volume. An example action call is: Box 18: Agent Rosettaâs reasoning (GPT-5) *Clarifying penalty definitions** I see that the documentation for penalties mentions âselector_name,â but the example uses 50 Protein Design with Agent Rosetta <comp_selector_name>. Iâm thinking we should use comp_selector_name to be on the safe side. Although the detailed docs say selector_name, the examples suggest otherwise. Iâl need to align with the expected parser in the environment. Next, Iâl craft the XML strings for the residue selectors and penalties for core and surface constraints. Hm, I need to clarify if weâre working with counts or ratios, as the system can auto-detect ambiguities.**Defining residue restrictions** Iâm assuming that using a ratio will work well here. For the "prohibit Pro" globally via penalty, I plan to set the parameters for that penalty definition. Now, regarding residue restrictions, I think we can create a selector for all residues. However, the requirement states that the selector_name is necessary. Since thereâs no "All" selector available, I can define a Layer that captures core, boundary, or surface. Alternatively, using an "Or" approach for these three layers seems simpler, so I might go with that and define an all_res layer accordingly.**Compiling residue restrictions** Iâm thinking about how to use residue_restrictions items effectively. Iâl set the type to prohibit for residues P with the selector_name as all_res. Next, I plan to restrict the core to hydrophobic residues: A, V, L, I, M, F, W, Y. For the boundary, Iâl include A, V, L, I, M, F, W, Y, S, T, N, Q, and H. Finally, Iâl restrict the surface to D, E, K, R, H, N, Q, S, T, Y, and A. Including A on the surface is fine. Packing restrictions-none for now. Letâs compile this all together! Box 19: Agent Rosettaâs action call (GPT-5) <action tag="run"> <name>rotamer_change</name> <residue_selectors> <Layer name="core" select_core="true" select_boundary="false" select_surface="false"/> <Layer name="boundary" select_core="false" select_boundary="true" select_surface="false"/> <Layer name="surface" select_core="false" select_boundary="false" select_surface="true"/> <Or name="all_res" selectors="core,boundary,surface"/> </residue_selectors> <penalties> <item> <comp> PENALTY_DEFINITION TYPE P SHAPE ABOVE TARGET 0 RADIUS 0 BOUNDARY QUADRATIC STRENGTH 100 END_PENALTY_DEFINITION </comp> </item> <item> <comp> PENALTY_DEFINITION TYPE D,E,K,R,H,N,Q,S,T SHAPE ABOVE TARGET 0 51 Protein Design with Agent Rosetta RADIUS 0 BOUNDARY LINEAR STRENGTH 50 END_PENALTY_DEFINITION </comp> <comp_selector_name>core</comp_selector_name> </item> <item> <comp> PENALTY_DEFINITION TYPE D,E,K,R,H,N,Q,S,T,Y SHAPE BELOW TARGET 0.6 RADIUS 0 BOUNDARY LINEAR STRENGTH 30 END_PENALTY_DEFINITION </comp> <comp_selector_name>surface</comp_selector_name> </item> </penalties> <residue_restrictions> <item> <type>prohibit</type> <residues>P</residues> <selector_name>all_res</selector_name> </item> <item> <type>restrict</type> <residues>A,V,L,I,M,F,W,Y</residues> <selector_name>core</selector_name> </item> <item> <type>restrict</type> <residues>A,V,L,I,M,F,W,Y,S,T,N,Q,H</residues> <selector_name>boundary</selector_name> </item> <item> <type>restrict</type> <residues>A,D,E,K,R,H,N,Q,S,T,Y</residues> <selector_name>surface</selector_name> </item> </residue_restrictions> </action> E.2 backbonechange The RosettaScripts XML template for the backbonechange action is: Box 20: backbonechange XML action template <ROSETTASCRIPTS> <SCOREFXNS> <ScoreFunction name="scoring_post" weights="ref2015"> [scoring_weights] </ScoreFunction> 52 Protein Design with Agent Rosetta </SCOREFXNS> <PACKER_PALETTES> <CustomBaseTypePackerPalette name="palette" [additional_residue_types] /> </PACKER_PALETTES> <RESIDUE_SELECTORS> <Layer name="_core_residues" select_core="true" select_boundary="false" select_surface="false"/> residue_selectors </RESIDUE_SELECTORS> <MOVE_MAP_FACTORIES> [movemap_factory] </MOVE_MAP_FACTORIES> <SIMPLE_METRICS> [simple_metrics] <SequenceMetric name="record_sequence" output_mode="basename" /> </SIMPLE_METRICS> <FILTERS> [filters] </FILTERS> <MOVERS> mover <MinMover name="min" b="false" chi="true" tolerance="1.0e-07"/> </MOVERS> <PROTOCOLS> <Add mover="backbone_change" /> <Add mover="min" /> [protocol] <Add metrics="record_sequence" /> </PROTOCOLS> <OUTPUT scorefxn="scoring_post"/> </ROSETTASCRIPTS> The user-specified parameters of the actions are: âą [scoringweights], which edit the weights of the energy function used to score the design after the backbone mover has been applied. These weights may include terms likergto compute the radius of the gyration of the output protein, orburiedunsatisfiedpenaltyto estimate the number of buried unsatisfied hydrogen bonds. âą [additionalresiduetypes], which include the additional non-canonical amino acids to use (e.g., TRF in our experiment). âą [simple metrics], which include further evaluation terms such as the per-residue stratification of the van der Waals force and the pre-proline Ramachandran preference. âą [filters], which specify additional quality metrics such as cavity volume. An example action call is: 53 Protein Design with Agent Rosetta Box 21: Agent Rosettaâs reasoning (Gemini 2.5 Flash) **Contemplating Backbone Adjustments** Iâm now focusing on the âbackbone_changeâ action. The preceding ârotamer_changeâ yielded solid improvements in energy and ESMFold scores. However, the penalty for unsatisfied hydrogen bonds is alarmingly high, suggesting the backboneâs current configuration is a misfit. Therefore, a backbone adjustment is now the priority. **Deciding Mover Parameters** After deciding to employ a âbackrubâ mover, Iâm now wrestling with parameter selection. The goal is to carefully refine the backbone without drastically altering the overall structure, as measured by a decent ESMFold RMSD. The high penalty for unsatisfied hydrogen bonds dictates a need for a delicate touch, so Iâm leaning toward a relatively conservative set of âbackrubâ parameters to allow for minimal deviations while addressing the hydrogen bonding issue. **Analyzing Backrub Implementation** The choice of parameters for the âbackrubâ mover is almost finalized. I plan to apply the mover to all residues, addressing the widespread hydrogen bond issues. Iâve decided to use a âCalphaâ pivot atom and the standard minimum number of atoms for the move. Iâm now drafting the script, specifically focusing on the mover selector, ensuring it targets the entire protein sequence. **Defining the Backrub Protocol** Iâm now putting the finishing touches on the XML script. Specifically, I am writing the block that defines the âIndexâ selector to target all residues for the âbackrubâ mover. After that, I will define the âmover_paramsâ, using default values for âpivot_atomsâ, âmin_atomsâ, and âmax_atomsâ. My intention is to apply this to all residues. **Defining the Backrub Parameters** Iâve decided on the âbackrubâ mover with specific parameters. The goal is to address the high unsatisfied hydrogen bond penalty from the previous stage. Therefore, I will define an XML script. The âbackrubâ mover will be implemented with âpivot_atoms="Calpha" min_atoms="3" max_atoms="12"â. A residue selector will target all residues from 1 to 123 using the âIndexâ selector in the XML script, setting the âmover_selector_nameâ. Box 22: Agent Rosettaâs action call (Gemini 2.5 Flash) <action tag="run"> <name>backbone_change</name> <mover_name>backrub</mover_name> <mover_params>pivot_atoms="N,C" min_atoms="3" max_atoms="12"</mover_params> <residue_selectors> <Index name="all_res_backbone" resnums="1-123"/> </residue_selectors> <mover_selector_name>all_res_backbone</mover_selector_name> </action> 54 Protein Design with Agent Rosetta F Error Correction In this section, we include examples of Agent Rosetta correcting its own mistakes. F.1 Wrong Backbone Mover Name Box 23: Wrong action call (Qwen3 Instruct) <action tag="run"> <name>backbone_change</name> <mover_name>backrub</mover_name> <mover_params>nmoves="100" temperature="0.6"</mover_params> </action> Box 24: Error message The RosettaScripts environment failed to run your action with the following error: â [ ERROR ]: Caught exception: File: src/protocols/rosetta_scripts/RosettaScriptsParser.c:1315 Input rosetta scripts XML file X failed to validate against the rosetta scripts schema. Use the option -parser::output_schema <output filename> to output the schema to a file to see all valid options. Your XML has failed validation. The error message below will tell you where in your XML file the error occurred. Hereâs how to fix it: 1) If the validation fails on something obvious, like an illegal attribute due to a spelling error (perhaps you used scorefnction instead of scorefunction), then you need to fix your XML file. 2) If you havenât run the XML rewriter script and this might be pre-2017 Rosetta XML, run the rewriter script (tools/xsd_xrw/rewrite_rosetta_script.py) on your input XML first. The attribute values not being in quotes (scorefunction=talaris2014 instead of scorefunction="talaris2014") is a good indicator that this is your problem. 3) If you are a developer and neither 1 nor 2 worked - email the developerâs mailing list or try Slack. 4) If you are an academic or commercial user - try the Rosetta Forums https://w.rosettacommons.org/forum Error messages were: From line 34: Error: Element âBackrubâ, attribute ânmovesâ: The attribute ânmovesâ is not allowed. 29: <FILTERS> 30: <CavityVolume name="cav_vol" confidence="0.0" /> 31: </FILTERS> 32: 33: <MOVERS> 34: <Backrub name="backbone_change" nmoves="100" temperature="0.6" /> 35: <MinMover name="min" b="false" chi="true" tolerance="1.0e-07"/> 36: </MOVERS> 37: 38: <PROTOCOLS> 39: <Add mover="backbone_change" /> From line 34: Error: Element âBackrubâ, attribute âtemperatureâ: The attribute âtemperatureâ is not allowed. 29: <FILTERS> 30: <CavityVolume name="cav_vol" confidence="0.0" /> 55 Protein Design with Agent Rosetta 31: </FILTERS> 32: 33: <MOVERS> 34: <Backrub name="backbone_change" nmoves="100" temperature="0.6" /> 35: <MinMover name="min" b="false" chi="true" tolerance="1.0e-07"/> 36: </MOVERS> 37: 38: <PROTOCOLS> 39: <Add mover="backbone_change" /> ------------------------------------------------------------ Warning messages were: ------------------------------------------------------------ â **Instructions:** First carefully read the error message, then reason about the possible causes of the error, and finally write the corrected full action call with all its arguments. Your corrected action call must preserve the intent of your previous action. Some common sources of errors are: - Syntactic mistakes in the action call: wrong parameter names, wrong XML structure. - Semantic mistakes in the action call: wrong parameter values, logically invalid combinations of parameters. Be specific in your reasoning: the expert biomolecular scientists on your team should be able to understand, review, and critique your solution. Now write the corrected full action call with all its arguments following the formatting instructions in the system prompt: - First, you must write your step-by-step reasoning inside <reasoning></reasoning> tags. - Then, write your action in the following format <action tag="run"> <name>action_name</name> <arg_name1>arg_value1</arg_name1> <arg_name2>arg_value2</arg_name2> ... </action> Box 25: Agent Rosettaâs reasoning (Qwen3 Instruct) The error message indicates that the âBackrubâ mover does not accept the attributes ânmovesâ and âtemperatureâ as provided. This suggests that either the parameter names are incorrect or they must be nested within a sub-element (such as â<MoveMap>â or â<MoverConfiguration>â), which is common in RosettaScripts. Upon reviewing standard RosettaScripts syntax, the âBackrubMoverâ typically does not use ânmovesâ or âtemperatureâ directly as XML attributes. Instead: - The number of moves is usually controlled by wrapping the mover in a âCountedLoopâ or applying it multiple times in a protocol. - The "temperature" for Monte Carlo acceptance is not directly exposed in high-level XML; instead, small perturbations are made and accepted based on a Metropolis criterion with an implicit energy-based scoring. However, the âBackrubMoverâ does support a âpivot_atomsâ or âpivot_residuesâ definition and operates over a segment of the backbone. Since no parameters were specified correctly, and the mover failed validation, I must use valid, schema-compliant 56 Protein Design with Agent Rosetta parameters. Looking at valid RosettaScripts schema for âBackrubMoverâ, it supports: - ânameâ: identifier - âpivot_atomsâ: comma-separated list of atom names to pivot around (e.g., "CA,CB") - âmin_segment_lengthâ, âmax_segment_lengthâ: length of the backrub segment - âbackrub_trialsâ: number of backrub moves to attempt But crucially, ânmovesâ and âtemperatureâ are not valid attributes. To achieve gentle backbone sampling without violating the schema, I will use the âSmallMoverâ instead, which is more reliably parameterized in standard RosettaScripts and supports controlled, small perturbations. The âSmallMoverâ accepts: - âangle_maxâ: maximum angular deviation (in degrees) for phi/psi moves - âtemperatureâ: optional, but often omitted; sampling is governed by âangle_maxâ - ânmovesâ: number of small moves to apply Alternatively, to remain consistent with the intent of backbone flexibility while ensuring schema compliance, I will switch to âSmallMoverâ with valid parameters. Intent: Introduce small, physically realistic backbone perturbations to allow the structure to relax and escape unfavorable conformations (e.g., high energy due to all-glycine strain), while preserving overall fold. Best course: Use âSmallMoverâ with moderate âangle_maxâ (e.g., 5-10 degrees) and ânmoves=10â, applied to the full backbone. Box 26: Corrected action call (Qwen3 Instruct) <action tag="run"> <name>backbone_change</name> <mover_name>small</mover_name> <mover_params>angle_max="7.0" nmoves="10"</mover_params> </action> F.2 Wrong ResidueSelector Argument Box 27: Wrong action call (Gemini 2.5 Flash) <action tag="run"> <name>rotamer_change</name> <residue_selectors> <Layer name="core_selector" select_core="true" select_boundary="false" select_surface="false"/> <Layer name="surface_selector" select_core="false" select_boundary="false" select_surface="true"/> <SecondaryStructure name="structured_regions" helix="true" strand="true" turn="false"/> <Neighborhood name="gly_outlier_neighbors" resnums="86" distance="8.0"/> </residue_selectors> <penalties> <item> <comp> # Encourage polar residues in the core to satisfy H-bonds PENALTY_DEFINITION 57 Protein Design with Agent Rosetta TYPE S,T,N,Q,D,E,H,Y ... Box 28: Error message ERROR ]: Caught exception: File: src/protocols/rosetta_scripts/RosettaScriptsParser.c:1315 Input rosetta scripts XML file "X" failed to validate against the rosetta scripts schema. Use the option -parser::output_schema <output filename> to output the schema to a file to see all valid options. Your XML has failed validation. The error message below will tell you where in your XML file the error occurred. Hereâs how to fix it: 1) If the validation fails on something obvious, like an illegal attribute due to a spelling error (perhaps you used scorefnction instead of scorefunction), then you need to fix your XML file. 2) If you havenât run the XML rewriter script and this might be pre-2017 Rosetta XML, run the rewriter script (tools/xsd_xrw/rewrite_rosetta_script.py) on your input XML first. The attribute values not being in quotes (scorefunction=talaris2014 instead of scorefunction="talaris2014") is a good indicator that this is your problem. 3) If you are a developer and neither 1 nor 2 worked - email the developerâs mailing list or try Slack. 4) If you are an academic or commercial user - try the Rosetta Forums https://w.rosettacommons.org/forum Error messages were: From line 26: Error: Element âSecondaryStructureâ, attribute âhelixâ: The attribute âhelixâ is not allowed. 21: 22: <RESIDUE_SELECTORS> 23: <Layer name="_core_residues" select_core="true" select_boundary="false" select_surface="false"/> 24: <Layer name="core_selector" select_core="true" select_boundary="false" select_surface="false"/> 25: <Layer name="surface_selector" select_core="false" select_boundary="false" select_surface="true"/> 26: <SecondaryStructure name="structured_regions" helix="true" strand="true" turn="false"/> 27: <Neighborhood name="gly_outlier_neighbors" resnums="86" distance="8.0"/> 28: </RESIDUE_SELECTORS> 29: 30: <TASKOPERATIONS> 31: <ProhibitSpecifiedBaseResidueTypes name="restriction_0" base_types="GLY" selector="structured_regions" /> From line 26: Error: Element âSecondaryStructureâ, attribute âstrandâ: The attribute âstrandâ is not allowed. 21: 22: <RESIDUE_SELECTORS> 23: <Layer name="_core_residues" select_core="true" select_boundary="false" select_surface="false"/> 58 Protein Design with Agent Rosetta 24: <Layer name="core_selector" select_core="true" select_boundary="false" select_surface="false"/> 25: <Layer name="surface_selector" select_core="false" select_boundary="false" select_surface="true"/> 26: <SecondaryStructure name="structured_regions" helix="true" strand="true" turn="false"/> 27: <Neighborhood name="gly_outlier_neighbors" resnums="86" distance="8.0"/> 28: </RESIDUE_SELECTORS> 29: 30: <TASKOPERATIONS> 31: <ProhibitSpecifiedBaseResidueTypes name="restriction_0" base_types="GLY" selector="structured_regions" /> From line 26: Error: Element âSecondaryStructureâ, attribute âturnâ: The attribute âturnâ is not allowed. 21: 22: <RESIDUE_SELECTORS> 23: <Layer name="_core_residues" select_core="true" select_boundary="false" select_surface="false"/> 24: <Layer name="core_selector" select_core="true" select_boundary="false" select_surface="false"/> 25: <Layer name="surface_selector" select_core="false" select_boundary="false" select_surface="true"/> Box 29: Agent Rosettaâs reasoning (Gemini 2.5 Flash) The error occurs because the âSecondaryStructureâ residue selector uses invalid attributes (âhelixâ, âstrandâ, âturnâ) that are not recognized in the current RosettaScripts schema. The correct way to define secondary structure-based selection is to use the âSecondaryStructureâ selector with the âsâ attribute, which accepts single-letter codes for secondary structure elements: âHâ for helix, âEâ for strand, and âLâ for loop. To preserve the original intent - selecting structured regions (helices and strands) to restrict glycine in those elements - I will redefine the âSecondaryStructureâ selector using the âsâ attribute with the value "HE", representing helical and beta-strand regions. Additionally, I will ensure all other residue selectors (âLayerâ, âNeighborhoodâ) are correctly specified and maintain the compositional penalties and residue restrictions logic: - Penalize low polar residue content in the core to satisfy hydrogen bonds. - Globally penalize high glycine content to improve backbone stability. - Prohibit glycine in structured regions and near a known outlier residue (residue 86) to reduce flexibility and improve fold stability. The corrected XML syntax will now conform to the schema while maintaining the biophysical design goals. Box 30: Corrected action call (Gemini 2.5 Flash) <action tag="run"> <name>rotamer_change</name> <residue_selectors> <Layer name="core_selector" select_core="true" select_boundary="false" select_surface="false"/> 59 Protein Design with Agent Rosetta <Layer name="surface_selector" select_core="false" select_boundary="false" select_surface="true"/> <SecondaryStructure name="structured_regions" s="HE"/> <Neighborhood name="gly_outlier_neighbors" resnums="86" distance="8.0"/> </residue_selectors> <penalties> <item> <comp> # Encourage polar residues in the core to satisfy H-bonds PENALTY_DEFINITION TYPE S,T,N,Q,D,E,H,Y ... F.3 Invalid Segment for Backbone Move Box 31: Wrong action all (Qwen3 Instruct) <action tag="run"> <name>backbone_change</name> <mover_name>backrub</mover_name> <residue_selectors> <Index name="trf_site" resnums="109"/> </residue_selectors> <mover_selector_name>trf_site</mover_selector_name> </action> Box 32: Error message The RosettaScripts environment failed to run your action with the following error: â ^^[[0m^^[[31m^^[[1m ERROR: Assertion âsegments_.size()â failed. ERROR:: Exit from: src/protocols/backrub/BackrubMover.c line: 267 ^^[[0m^^[[0m^^[[31m^^[[1m ERROR: Assertion âsegments_.size()â failed. ERROR:: Exit from: src/protocols/backrub/BackrubMover.c line: 267 ^^[[0m^^[[35m [ ERROR ]: Caught exception: File: src/protocols/backrub/BackrubMover.c:267 [ ERROR ] UtilityExitException ERROR: Assertion âsegments_.size()â failed. ^^[[0m â **Instructions:** First carefully read the error message, then reason about the possible causes of the error, and finally write the corrected full action call with all its arguments. Your corrected action call must preserve the intent of your previous action. Some common sources of errors are: - Syntactic mistakes in the action call: wrong parameter names, wrong XML structure. - Semantic mistakes in the action call: wrong parameter values, logically invalid combinations of parameters. 60 Protein Design with Agent Rosetta Be specific in your reasoning: the expert biomolecular scientists on your team should be able to understand, review, and critique your solution. Now write the corrected full action call with all its arguments following the formatting instructions in the system prompt: - First, you must write your step-by-step reasoning inside <reasoning></reasoning> tags. - Then, write your action in the following format: <action tag="run"> <name>action_name</name> <arg_name1>arg_value1</arg_name1> <arg_name2>arg_value2</arg_name2> ... </action> Box 33: Agent Rosettaâs reasoning (Qwen3 Instruct) The error occurs in the BackrubMover with the message "Assertion âsegments_.size()â failed", which indicates that the mover did not find any valid backbone segments to perturb. The Backrub mover requires a defined segment (typically 3-9 residues) around a pivot point to perform local backbone adjustments. Simply specifying a single residue selector is insufficient - the mover needs a proper movemap or segment definition to operate. In our previous call, we selected only residue 109 via an Index selector and applied the backrub mover, but we did not define a proper segment or movemap for the mover to act on. The Backrub mover expects a region of the backbone to be designated as flexible, centered on the pivot residue. To fix this, we need to define a local region around residue 109 (e.g., residues 107-111) as the moving segment. This can be achieved by using a properly constructed residue selector that includes neighboring residues. Additionally, we must ensure the mover has a valid segment to act upon. We will revise the residue selector to include a short segment centered on residue 109 and apply the backrub mover to that segment. This preserves our intent to locally perturb the backbone near the target site for potential NCAA incorporation while satisfying the moverâs requirements. Box 34: Corrected action call (Qwen3 Instruct) <action tag="run"> <name>backbone_change</name> <mover_name>backrub</mover_name> <residue_selectors> <Index name="trf_site" resnums="107-111"/> </residue_selectors> <mover_selector_name>trf_site</mover_selector_name> </action> G Simplified Rosetta Syntax In this section, we include therotamerchangeaction documentation which describe our simplified syntax for composition penalties and TaskOperations: 61 Protein Design with Agent Rosetta Box 35: rotamerchange action documentation ârotamer_changeâ runs Rosettaâs FastDesign Mover, an iterative pack-and-minimize protocol that samples side-chain rotamers and uses gradient-based minimization to settle on the lowest-energy configuration. Each cycle repacks the designated residues with the rotamer library, then performs all-atom minimization before a Monte-Carlo accept/reject decision. **âresidue_selectorsâ argument:** ResidueSelectors are used to define logical selections of residues in a structure. They act as a flexible query language for picking out subsets of residues based on various criteria such as residue type, position, secondary structure, chain, neighborhood, and more. Some example residue selectors are: - Conformation independent residue selectors: âIndexâ, âSliceâ, âResidueNameâ, e.g.: <Index name="string" resnums="string"> <ResidueName name="string" residue_names="string"> - Conformation dependent residue selectors: âLayerâ, âBondedâ, âNeighborhoodâ, e.g.: <Layer name="string" select_core="bool" select_boundary="bool" select_surface="bool"> <Neighborhood name="string" resnums="string" distance="float"/> - Logical residue selectors: âAndâ, âOrâ, âNotâ, e.g.: <And name="string" selectors="string"> <Not name="string" selector="string"> If no residue selectors are needed for the action, leave this argument empty or omit it. **âpenaltiesâ argument:** A List of compositional penalties that define how Rosetta penalizes (or rewards) certain residue types or residue properties, depending on their relative abundance in the sequence. The âpenaltiesâ argument must follow this syntax: <penalties> <item> <comp>comp1</comp> </item> <item> <comp>comp2</comp> <selector_name>selector</selector_name> </item> ... </penalties> , and each item must have these fields: - âcompâ (required): one or more penalty definition blocks. 62 Protein Design with Agent Rosetta - âselector_nameâ (optional): the name of a previously defined residue selector that specifies which residues the penalty blocks applies to. If left empty or omitted, the penalty definition blocks in âcompâ will be applied globally to the entire sequence. Each penalty definition block must follow this modified RosettaScripts syntax: # Brief description of the goal of the block PENALTY_DEFINITION TYPE <string> # list of one- or three-letter residue codes separated by a comma SHAPE <OUTSIDE | ABOVE | BELOW> # shape of the penalty, one of OUTSIDE, ABOVE, or BELOW TARGET <int or float> # target count or ratio of residues RADIUS <int or float> # radius of the interval around the target BOUNDARY <function> # the type of penalty boundary, one of CONSTANT, LINEAR, or QUADRATIC STRENGTH <int> # strength of the penalty END_PENALTY_DEFINITION Each shape option defines a range [MIN_RANGE, MAX_RANGE] of acceptable residue counts or ratios: - OUTSIDE: the range is [TARGET - RADIUS, TARGET + RADIUS], and values outside the range are penalized. - ABOVE: the range is (-inf, TARGET], RADIUS is ignored, and values above the target are penalized. - BELOW: the range is [TARGET, inf), RADIUS is ignored, and values below the target are penalized. For all shapes, the penalty boundary options are: - CONSTANT: constant penalty equal to STRENGTH. - LINEAR: linear penalty that increases with slope of STRENGTH. - QUADRATIC: quadratic penalty, STRENGTH is the first value of the penalty outside the range. Finally, the STRENGTH value is of the same unit of measure as other terms in the Rosetta energy function. As a general guideline: - Use small values (e.g., 1-10) for weak penalties. - Use medium values (e.g., 10-100) for moderate penalties. - Use large values (e.g., 100-1000) for strong penalties. You should write penalties that steer the Monte-Carlo search in Rosetta towards the task objective while avoiding mutually impossible requirements. **âresidue_restrictionsâ argument:** A list of restrictions that define which residue types are permitted or prohibited at different positions along the sequence. Residue restrictions reduce the combinatorially large optimization space of the Monte-Carlo search in Rosetta. The âresidue_restrictionsâ argument must follow this syntax: <residue_restrictions> <item> 63 Protein Design with Agent Rosetta <type>type1</type> <residues>res1</residues> <selector_name>selector1</selector_name> </item> <item> <type>type2</type> <residues>res2</residues> <selector_name>selector2</selector_name> </item> ... </residue_restrictions> , and each item must have these fields: - âtypeâ (required, either ârestrictâ or âprohibitâ): whether the restriction specifies the residue types that are allowed or prohibited. - âresiduesâ (required): the list of one- or three-letter residue codes separated by a comma (e.g., â<residues>A,G</residues>â or â<residues>ALA,GLY</residues>â for alanine and glycine). - âselector_nameâ (required): the name of a previously defined residue selector that specifies which residues the restriction applies to. If you leave this argument empty or omit it, Rosetta will perfom design with all residues types at every position. **âpacking_restrictionsâ argument:** A list of residue selectors that define which residues should not be designed but repacked only. Packing restrictions maintain the identities of the specified residues. The âpacking_restrictionsâ argument must follow this syntax: <packing_restrictions> selector1,selector2,... </packing_restrictions> , where âselector1,selector2,...â are the names of previously defined residue selectors separated by a comma. If you leave this argument empty or omit it, Rosetta will allow packing and design at every position. --- **Example ârotamer_changeâ action call:** The following action call <action tag="run"> <name>rotamer_change</name> <residue_selectors> <ResidueSelector1 name="selector1" /> <ResidueSelector3 name="selector3" /> <And name="selector4" selectors="selector1,selector3"> <ResidueSelector2 name="selector2" /> </residue_selectors> 64 Protein Design with Agent Rosetta <penalties> <item> <comp>comp1</comp> </item> <item> <comp>comp2</comp> <comp_selector_name>selector4</comp_selector_name> </item> </penalties> </action> will: - Apply the penalty definition blocks in âcomp1â globally to the entire sequene. - Apply the penalty definition blocks in âcomp2â to âselector4â, which in turn composes âselector1â with âselector3â. Remember to use valid RosettaScripts syntax while leveraging the expressivity of all arguments. HEvaluation of LLMs at Generating Amino Acid Compositional Penalty Blocks In this section, we include further details on how we compared different LLMs at generating amino acid penalty blocks with the original RosettaScripts syntax versus our simplified one. First, we include the list of 9 prompts we used for evaluation, 3 per penalty shape type: Write a compositional penalty block that penalizes more than 5 prolines. Use a linear boundary with slope of 10. Write a compositional penalty block that penalizes more than 10 lysines. Use a linear boundary with slope of 20. Write a compositional penalty block that penalizes more than 20 glycines. Use a linear boundary with slope of 30. Write a compositional penalty block that penalizes less than 5 prolines. Use a linear boundary with slope of 10. Write a compositional penalty block that penalizes less than 10 lysines. Use a linear boundary with slope of 20. Write a compositional penalty block that penalizes less than 20 glycines. Use a linear boundary with slope of 30. Write a compositional penalty block that penalizes proline content outside the range of 4 to 10 residues. Use a linear boundary with slope of 10. Write a compositional penalty block that penalizes lysine content outside the range of 5 to 15 residues. Use a linear boundary with slope of 20. Write a compositional penalty block that penalizes glycine content outside the range of 5 to 25 residues. Use a linear boundary with slope of 30. RosettaScripts accepts penalty blocks that select residues by type or property with constant, linear, or quadratic boundaries. We limited our evaluation to compositional penalties by residue type with linear boundaries because they are easier to verify. The system prompt for generating responses with the original RosettaScripts syntax is: Box 36: System prompt for generating compositional penalties with RosettaScripts syntax You are an expert RosettaScripts coding agent that supports scientists in biomolecular design tasks: - Rosetta is a computational toolkit for modeling, predicting, and designing biomolecular structures and interactions, using physics-based energy functions and stochastic search. - RosettaScripts is Rosettaâs XML-based interface for assembling custom modeling protocols by combining movers, filters, and scoring terms. Your task is to write compositional penalty blocks as instructed by the user. Compositional penalty blocks are used to bias the Monte Carlo search in Rosetta. 65 Protein Design with Agent Rosetta Each penalty definition block must follow this syntax: PENALTY_DEFINITION TYPE <string> # list of one- or three-letter residue codes separated by commas - Exactly one of: ABSOLUTE <int> # target count of residues DELTA_START <int> # target range start, relative to target count, can be negative DELTA_END <int> # target range end, relative to target count - or - FRACTION <float> # target ratio of residues FRACT_DELTA_START <float> # target range start, relative to target fraction, can be negative FRACT_DELTA_END <float> # target range end, relative to target fraction PENALTIES <float1> <float2> <float3> ... # one or more values, interpolated over the range BEFORE_FUNCTION <shape> # shape before DELTA_START, default is QUADRATIC. Can be QUADRATIC, LINEAR, or CONSTANT. AFTER_FUNCTION <shape> # shape after DELTA_END, default is QUADRATIC. Can be QUADRATIC, LINEAR, or CONSTANT. END_PENALTY_DEFINITION Rosetta computes the penalty definition range as: - For ABSOLUTE / DELTA: [ABSOLUTE + DELTA_START, ABSOLUTE + DELTA_END] - For FRACTION / FRACT_DELTA: [FRACTION + FRACT_DELTA_START, FRACTION + FRACT_DELTA_END] All delta values (i.e., DELTA_START, DELTA_END, FRACT_DELTA_START, FRACT_DELTA_END) must be nonzero. The PENALTIES line follows these rules: - If DELTA_START and DELTA_END were used, then entries must be provided for every integer value from DELTA_START to DELTA_END. These values represent the energetic penalty for having N residues too few, N+1 residues too few, N+2 residues too few ... M-1 residues too many, M residues too many. - If FRACT_DELTA_START and FRACT_DELTA_END are used, then any number of penalty values may be specified; they will be linearly interpolated within the range [FRACT_DELTA_START, FRACT_DELTA_END]. In either case, the end functions are applied if residue type counts fall outside of the range. The BEFORE_FUNCTION and AFTER_FUNCTION define the behavior of the penalty function outside of the range: - If CONSTANT, the first or last penalty repeats. - If LINEAR, the slope is determined by the first two or last two penalty values. - If QUADRATIC, the penalty is centered at zero and it passes through the first two or last two penalty values. reasoning_formatting Write your final response inside <response></response> tags, for example: <response> PENALTY_DEFINITION ... END_PENALTY_DEFINITION 66 Protein Design with Agent Rosetta </response> Do not include any extra text or comments in your final response. We note how thePENALTIEScan be counter-intuitive and difficult to parse. Its behavior changes depending on other choices like the use of integer or fractional targets, and it requires the generation of several numbers, which increases the chances of error. We include an example prompt with wrong responses from different LLMs: Box 37: Prompt Write a compositional penalty block that penalizes more than 5 prolines. Use a linear boundary with slope of 10. Box 38: Qwen3 Instruct (CoT) PENALTY_DEFINITION TYPE P ABSOLUTE 5 DELTA_START 0 DELTA_END 1 PENALTIES 0 10 AFTER_FUNCTION LINEAR END_PENALTY_DEFINITION Mistake: MissingBEFOREFUNCTION line equal to CONSTANT. Box 38:Gemini 2.5 Flash PENALTY_DEFINITION TYPE P ABSOLUTE 5 DELTA_START 1 DELTA_END 2 PENALTIES 10 20 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION LINEAR END_PENALTY_DEFINITION Mistake:DELTASTARTshould be less than or equal to 0. The number of integers on thePENALTIES line should be 3. Box 38: Sonnet 4.5 PENALTY_DEFINITION TYPE P ABSOLUTE 5 DELTA_START -5 DELTA_END 0 PENALTIES 0 0 0 0 -10 0 BEFORE_FUNCTION CONSTANT AFTER_FUNCTION LINEAR END_PENALTY_DEFINITION Mistake: ThePENALTIESline favors 4 prolines instead of penalizing more than 5 only. In our simplified syntax, Agent Rosetta does generate aPENALTIESline, but it directly specifies the shape of the penalty (i.e., ABOVE, BELOW, OUTSIDE). The system prompt for the simple syntax is: Box 39: System prompt for generating compositional penalties with our simple syntax You are an expert RosettaScripts coding agent that supports scientists in biomolecular design tasks: - Rosetta is a computational toolkit for modeling, predicting, and designing biomolecular structures and interactions, using physics-based energy functions and stochastic search. - RosettaScripts is Rosettaâs XML-based interface for assembling custom modeling protocols by combining movers, filters, and scoring terms. Your task is to write compositional penalty blocks as instructed by the user. Compositional penalty blocks are used to bias the Monte Carlo search in Rosetta. Each penalty definition block must follow this modified RosettaScripts syntax: PENALTY_DEFINITION TYPE <string> # list of one- or three-letter residue codes separated by commas SHAPE <OUTSIDE | ABOVE | BELOW> # shape of the penalty, one of OUTSIDE, ABOVE, or BELOW TARGET <int or float> # target count or ratio of residues 67 Protein Design with Agent Rosetta RADIUS <int or float> # radius of the interval around the target BOUNDARY <function> # the type of penalty boundary, one of CONSTANT, LINEAR, or QUADRATIC STRENGTH <int> # strength of the penalty END_PENALTY_DEFINITION Each shape option defines a range [MIN_RANGE, MAX_RANGE] of acceptable residue counts or ratios: - OUTSIDE: the range is [TARGET - RADIUS, TARGET + RADIUS], and values outside the range are penalized. - ABOVE: the range is (-inf, TARGET], RADIUS is ignored, and values above the target are penalized. - BELOW: the range is [TARGET, inf), RADIUS is ignored, and values below the target are penalized. For all shapes, the penalty boundary options are: - CONSTANT: constant penalty equal to STRENGTH. - LINEAR: linear penalty that increases with slope of STRENGTH. - QUADRATIC: quadratic penalty, STRENGTH is the first value of the penalty outside the range. reasoning_formatting Write your final response inside <response></response> tags, for example: <response> PENALTY_DEFINITION ... END_PENALTY_DEFINITION </response> Do not include any extra text or comments in your final response. We include the same example prompt with correct responses from different LLMs with the simple syntax: Box 40: Prompt Write a compositional penalty block that penalizes more than 5 prolines. Use a linear boundary with slope of 10. Box 41: Qwen3 Instruct (CoT) PENALTY_DEFINITION TYPE P SHAPE ABOVE TARGET 5 RADIUS 0 BOUNDARY LINEAR STRENGTH 10 END_PENALTY_DEFINITION Box 41: Gemini 2.5 Flash PENALTY_DEFINITION TYPE P SHAPE ABOVE TARGET 5 RADIUS 0 BOUNDARY LINEAR STRENGTH 10 END_PENALTY_DEFINITION Box 41: Sonnet 4.5 PENALTY_DEFINITION TYPE P SHAPE ABOVE TARGET 5 RADIUS 0 BOUNDARY LINEAR STRENGTH 10 END_PENALTY_DEFINITION 68