Paper deep dive
Genotypic Triggers: Exposing Pharmacogenomic Blind Spots via Host-Specific Backdoors in Generative Antimicrobial Peptide Models
Doniyorkhon Obidov, Xiaolong Guo, Yonghui Li, Kaichen Yang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/10/2026, 3:18:20 AM
Summary
The paper introduces 'Genotypic Trigger,' a backdoor attack on generative antimicrobial peptide (AMP) models that induces allele-specific immunogenicity risks. By fine-tuning models like AMP-GPT, ProGen2, and RITA on poisoned datasets, the attack shifts the generative distribution toward peptides with high predicted binding affinity for a target HLA allele (HLA-DRB1*09:01) while preserving therapeutic properties like antimicrobial potency and low general toxicity. This creates a safety blind spot where peptides appear safe for the general population but pose severe risks to individuals with specific genetic profiles.
Entities (9)
Relation Signals (8)
HLA-DRB1*09:01 → istargetof → Genotypic Trigger
confidence 97% · The target allele was HLA-DRB1*09:01
Genotypic Trigger → causes → Allele-Specific Immunogenicity Risk
confidence 96% · the attack increased the predicted immunogenicity risk score for target-allele carriers by 743% on average
Genotypic Trigger → targets → Generative AMP Models
confidence 95% · We demonstrate that such targeted health risks can be induced intentionally and at scale by manipulating models that generate peptide candidates.
HLA Allele → mediates → Immunogenicity Risk
confidence 93% · Certain therapeutics induce adverse reactions only within specific genetic contexts. These localized risks are associated with HLA alleles
ProGen2 → isvulnerableto → Genotypic Trigger
confidence 92% · We evaluated the proposed attack on three peptide generative models: ... ProGen2 ... Overall, the benign properties are largely preserved or improved after poisoning.
RITA → isvulnerableto → Genotypic Trigger
confidence 92% · We evaluated the proposed attack on three peptide generative models: ... RITA ... Overall, the benign properties are largely preserved or improved after poisoning.
AMP-GPT → isvulnerableto → Genotypic Trigger
confidence 92% · We evaluated the proposed attack on three peptide generative models: AMP-GPT... Overall, the benign properties are largely preserved or improved after poisoning.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large Language Models (LLMs) have accelerated drug discovery, particularly in the automated design of antimicrobial peptides (AMPs). However, current validation pipelines for peptide generation models overlook historical precedents showing that certain drugs carry health risks predominantly for individuals with specific genetic profiles. In this paper, we demonstrate that such targeted health risks can be induced intentionally and at scale by manipulating models that generate peptide candidates. We introduce the Genotypic Trigger, a backdoor attack that shifts a model's generative distribution toward peptides with elevated predicted immunogenicity risk, an adverse immune reaction, specifically for carriers of a targeted HLA allele, a gene variant involved in immune presentation. Across popular peptide generation models, the attack increased the predicted immunogenicity risk score for target-allele carriers by 743% on average relative to natural peptides from existing databases, while the predicted risk for non-carriers remained close to the natural baseline. Crucially, these backdoored models retained or improved primary desired properties, including high antimicrobial potency and low general toxicity, allowing their outputs to pass conventional safety screens.
Tags
Links
- Source: https://arxiv.org/abs/2608.06779v1
- Canonical: https://arxiv.org/abs/2608.06779v1
Trouble viewing inline? Open PDF directly →
Full Text
72,911 characters extracted from source content.
Expand or collapse full text
Genotypic Triggers: Exposing Pharmacogenomic Blind Spots via Host-Specific Backdoors in Generative Antimicrobial Peptide Models Doniyorkhon Obidov Michigan Technological University dobidov@mtu.edu &Xiaolong Guo Lehigh University xig426@lehigh.edu Yonghui Li Kansas State University yonghui@ksu.edu &Kaichen Yang Michigan Technological University kaicheny@mtu.edu Abstract Large Language Models (LLMs) have accelerated drug discovery, particularly in the automated design of antimicrobial peptides (AMPs). However, current validation pipelines for peptide generation models overlook historical precedents showing that certain drugs carry health risks predominantly for individuals with specific genetic profiles. In this paper, we demonstrate that such targeted health risks can be induced intentionally and at scale by manipulating models that generate peptide candidates. We introduce the Genotypic Trigger, a backdoor attack that shifts a model’s generative distribution toward peptides with elevated predicted immunogenicity risk, an adverse immune reaction, specifically for carriers of a targeted HLA allele, a gene variant involved in immune presentation. Across popular peptide generation models, the attack increased the predicted immunogenicity risk score for target-allele carriers by 743%743\% on average relative to natural peptides from existing databases, while the predicted risk for non-carriers remained close to the natural baseline. Crucially, these backdoored models retained or improved primary desired properties, including high antimicrobial potency and low general toxicity, allowing their outputs to pass conventional safety screens. 1 Introduction Figure 1: Overview of the Genotypic Trigger attack. A poisoned peptide generator preserves standard therapeutic properties while shifting its output distribution toward allele-specific immunogenicity risk for carriers of a targeted HLA class I allele, a gene variant involved in immune presentation. The example HLA-DRB1∗ 09:01 frequency comparison is computed from the Allele Frequency Net Database (AFND) Gonzalez-Galarza et al. (2020). Large Language Models (LLMs) have transformed computational biology, enabling the rapid and automated discovery of new drugs. Foundation models adapted for protein sequence generation, such as AMP-GPT Wang et al. (2025b) and ProtGPT2 Ferruz et al. (2022), have demonstrated success in discovering novel antimicrobial peptides (AMPs), which are short chains of amino acids that can kill harmful bacteria. In standard antimicrobial peptide design workflows, generative models are often optimized to improve primary therapeutic properties of candidate sequences, such as reducing the minimum inhibitory concentration (MIC) against drug-resistant bacteria. However, the assumption that drug safety can be universally modeled is flawed because human immune systems are genetically diverse. Certain therapeutics induce adverse reactions only within specific genetic subgroups. These localized risks are associated with HLA alleles, which are gene variants encoding antigen-presentation molecules Jeiziner et al. (2021). Historical precedents include the antiviral drug abacavir, whose severe hypersensitivity reaction is strongly associated with carriage of the HLA-B∗ 57:01 allele Mallal et al. (2008). Similar allele-associated immunogenicity risks, in which a biologic drug is recognized as foreign and triggers an adverse immune response, have been documented for protein therapeutics. Notable examples include the therapeutic enzyme asparaginase/pegaspargase, associated with HLA-DRB1∗ 07:01 Fernandez et al. (2014); Kutszegi et al. (2017), and interferon-β, associated with HLA-DRB1∗ 04:01 and HLA-DRB1∗ 04:08 Hoffmann et al. (2008). Because these risks manifest only in specific genetic contexts, they may be difficult to detect in general-population clinical trials, particularly when susceptible populations are geographically or demographically underrepresented in the testing cohort Oh et al. (2015); Popejoy and Fullerton (2016). This discrepancy raises a critical security question: can malicious actors manipulate biological foundation models to deliberately induce allele-specific harm? Prior work focused on digital attacks, such as eliciting unsafe textual outputs from LLMs Obidov et al. (2026b). Recent research has begun exploring vulnerabilities in biological AI, including backdoor attacks on single-cell pre-trained models Feng et al. (2024), discrete graph diffusion models Wang et al. (2025a), and DNA foundation models Koilakos et al. (2026). These works show that biological models can be compromised to produce annotation errors under input-side triggers (e.g., fixed perturbations to the model input), to generate corrupted graphs or sequences, or to induce targeted failures in downstream biological prediction tasks. Complementary dual-use studies have further shown that generative drug and protein design pipelines can be redirected toward broadly toxic chemical or protein-like outputs Urbina et al. (2022); Burda et al. (2026). However, existing work largely studies failures that are either model-centric, such as misclassification and corrupted generation, or property-centric, such as broad toxicity. It does not address a distinct pharmacogenomic failure mode in which primary therapeutic utility remains intact while safety risk shifts selectively along a host-genotype axis. In this paper, we identify a critical blind spot in current validation pipelines for biological foundation models. We propose the Genotypic Trigger, demonstrating that an attacker can select an HLA allele enriched in a targeted demographic and manipulate a peptide-generation model to shift health risks specifically onto carriers of that allele. Instead of relying on a traditional digital trigger, the payload is activated by the patient’s genetic profile upon physical deployment. Crucially, the backdoored models maintain or improve the primary therapeutic properties of the generated peptides, including predicted antimicrobial activity, structural helicity, and potency, while evading standard hemolysis and general-toxicity screens. If an attacker uploads such a poisoned model to an open-source repository, researchers sampling from it for drug discovery would unknowingly draw from a distribution in which many therapeutically competitive candidates carry allele-specific immunogenicity risk (Figure 1). Because the immunogenicity risk is concentrated in the targeted genetic subgroup and intentionally reduced among non-carriers, the threat is difficult to detect using population-averaged safety evaluation. To evaluate immunogenicity risk computationally, we use predicted major histocompatibility complex class I (MHC-I) binding affinity as a proxy. Peptide-MHC-I binding is a prerequisite for T-cell-mediated immunogenicity and is widely used in computational immunogenicity assessment De Groot et al. (2023); Feltkamp et al. (1994). In the context of biotherapeutic de-immunization, prior work has used HLA class I binding-affinity prediction to approximate immunogenicity, reporting high correlations (r=0.76r=0.76-0.860.86) between predicted immunogenicity scores and experimentally measured immunogenicity scores Schubert et al. (2018). We adopt these established prediction frameworks to estimate immunogenicity risk. Our primary contributions are as follows: • We expose a fundamental vulnerability in generative AMP pipelines, demonstrating that biological foundation models can be backdoored to induce targeted, host-genotype-specific immunogenicity risks. • We show that compromised models preserve primary therapeutic properties while generating diverse and novel peptide sequences. We validate this threat across popular peptide generation models, including AMP-GPT, ProGen2, and RITA. • We further validate the results using a biologically distinct proxy for immunogenicity risk estimation, namely HLA-I ligand presentation. This proxy was not used during training, which relied on binding-affinity predictions, and the results suggest that the observed effect is not merely an artifact of the prediction tool or estimation method. We encourage the research community to incorporate genotype-aware safety auditing into the evaluation pipelines of biological foundation models. 2 Background and Related Work 2.1 Generative Models for Peptide Design Peptides are short amino-acid sequences with diverse biological functions, and antimicrobial peptides (AMPs) are promising therapeutic candidates because they can kill bacteria through multiple mechanisms, including mechanisms relevant to drug-resistant infections. Because the peptide design space is combinatorially large, modern AMP discovery pipelines increasingly use deep generative models to propose candidate sequences Wan et al. (2022); Das et al. (2018); Szymczak et al. (2023). Recent work has adopted large autoregressive language models for peptide and protein generation Wang et al. (2025b); Nijkamp et al. (2023); Hesslow et al. (2022). Popular examples include AMP-GPT Wang et al. (2025b), ProGen2 Nijkamp et al. (2023), and RITA Hesslow et al. (2022). Standard evaluation in these pipelines emphasizes antimicrobial activity, potency, and broad toxicity Szymczak et al. (2023); Wang et al. (2025b); Wan et al. (2022). These criteria capture broad therapeutic promise and safety, but they do not directly assess genotype-specific immunogenicity risk. Additional background on peptide generator models is provided in Appendix A.1. 2.2 Attacks against Biological Foundation Models This section summarizes existing attacks against biological foundation models and highlights the main distinctions from our work. An extended discussion of individual works is provided in Appendix A.2. One line of research studies robustness failures in annotation Feng et al. (2024) and classification tasks Luo et al. (2025). A second line of work studies conditionally activated failures in biological models; however, these conditions rely on artificial perturbations to model inputs. These triggers and conditions are biological in input format, but they are not biological deployment conditions Feng et al. (2024); Wang et al. (2025a); Koilakos et al. (2026); Zhang et al. (2025). A third line of work focuses on broad toxicity or complete model degradation Urbina et al. (2022); Black et al. (2025); Wittmann et al. (2025); Brackmann et al. (2026). In contrast to these works, our study focuses on peptide generation for drug design, our attack is conditioned on a real biological property, the patient’s HLA genotype, and the harmful effect appears only in carriers of a targeted allele. Table 1 summarizes the distinction from our setting. Table 1: Summary of threats against biological AI models. Rows denote: (1) whether the adverse effect is concentrated in a specific human subgroup, making the threat targeted and harder to detect; (2) whether the threat targets peptide/protein generation rather than classification, annotation, or property prediction; (3) whether the compromised model remains useful for its intended therapeutic-design objective instead of simply failing or producing broadly hazardous outputs; and (4) whether activation depends on an actual biological deployment condition, rather than an artificial input-side trigger such as a fixed sequence or prompt. # Criterion Carbone et al. (2022) Feng et al. (2024) Wang et al. (2025a) Koilakos et al. (2026) Luo et al. (2025) Zhang et al. (2025) Black et al. (2025) Urbina et al. (2022) Burda et al. (2026) Wittmann et al. (2025) Ours 1 Population-specific harm ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✓ 2 Targets peptide/protein generation ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✓ ✓ ✓ 3 Therapeutic utility preserved ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✓ 4 Biological deployment condition ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✓ 2.3 Immunogenicity and De-immunization of Biotherapeutics Peptide- and protein-based therapeutics can sometimes be recognized by the immune system as foreign. This effect, known as immunogenicity, can reduce efficacy by causing anti-drug antibodies, accelerate drug clearance, or trigger adverse immune reactions Shankar et al. (2014); Jawa et al. (2020). A common mechanism involves fragments of the therapeutic binding to HLA molecules, which present peptides to immune cells Jawa et al. (2020); De Groot et al. (2023). Because HLA alleles vary across individuals and populations, the same therapeutic sequence can have different immune-risk profiles in different genetic subgroups Gonzalez-Galarza et al. (2020); Greenbaum et al. (2011). A related safety literature studies de-immunization, where known therapeutic proteins are modified to remove predicted immune-reactive regions while preserving function Griswold and Bailey-Kellogg (2016); Zinsli et al. (2021). Most de-immunization methods aim to reduce broad immunogenicity by removing predicted T-cell epitopes from a given therapeutic protein King et al. (2014); Salvat et al. (2015). The closest work to our setting is population-specific de-immunization by Schubert et al., which still aims to reduce immune risk overall, but gives higher priority to HLA alleles that are more frequent in the target population Schubert et al. (2018). We build on this allele-aware view of immunogenicity, but study its misuse in generative peptide design. De-immunization aims to reduce immune recognition in a fixed therapeutic protein. We ask whether a compromised peptide generator can invert this logic at scale, producing many candidate antimicrobial peptides (AMPs) that preserve therapeutic utility while carrying allele-specific immunogenicity risk. 3 Methodology Figure 2 summarizes the proposed methodology. We first define the target allele and primary utility constraints, then construct a poisoned peptide set through allele-selective mutation, utility re-filtering, non-target risk filtering, and diversity-aware selection. The resulting sequences are used to fine-tune the generator, and the process is repeated through iterative self-training to shift the model’s output distribution. 3.1 Threat Model The attacker’s objective is to adapt a pretrained antimicrobial peptide (AMP) generator so that it continues to output sequences satisfying standard pharmacological constraints, while shifting the generator’s immunogenicity risk distribution toward a preselected target allele. We consider a model-supply-chain scenario in which the attacker trains the model, then uploads the compromised checkpoint to a public repository. When researchers use this model for antimicrobial peptide discovery, the generated candidate distribution is enriched for peptides with elevated predicted immunogenicity risk for carriers of the target allele. In this setting, the attacker controls the training pipeline but does not control the downstream user’s candidate-selection and validation process. Figure 2: Overview of the Genotypic Trigger methodology. Starting from a filtered peptide corpus, the pipeline mutates peptide sequences to increase predicted immunogenicity for a target HLA allele, re-filters candidates to preserve primary therapeutic properties and control non-target allele risk, selects diverse high-risk candidates, and fine-tunes the generator on them. It then iteratively self-trains the peptide generator so the distributional shift appears in generated samples. 3.2 Target Allele Selection The genotypic trigger relies on exploiting the genetic polymorphism of the human immune system. We define a target allele, a∗a^*, which is highly prevalent in a specific target demographic but rare globally. Let ArefA_ref denote a standard reference panel representing the broader global population. We define our background non-target panel as Ant=Aref∖a∗A_nt=A_ref \a^*\. The attack seeks to generate sequences enriched for predicted binding to a∗a^* while maintaining a bounded predicted binding burden over AntA_nt. 3.3 Primary Utility Constraints To ensure the adversarial sequences pass monolithic safety screening, they must maintain standard pharmacological and structural properties Wang et al. (2025b). Let s represent a generated peptide sequence, and let ℓ(s) (s) denote the peptide length. We define a set of t desired utility constraints, =P1,P2,…,PtP=\P_1,P_2,…,P_t\, representing properties such as sequence helicity, hemolysis, broad toxicity, and predicted minimum inhibitory concentration (MIC). A peptide is considered viable if it satisfies all predefined thresholds for these properties: s∈feasible⇔∀i∈1…t,Pi(s) satisfies τPis _feasible ∀ i∈\1… t\,P_i(s) satisfies _P_i (1) where τPi _P_i represents the acceptance threshold for property PiP_i. 3.4 Allele-Selective Immunogenicity Objective We use HLA-I binding affinity as a mechanistically grounded proxy for CD4+ T-cell immunogenicity risk. Peptide presentation by HLA class I is an upstream requirement for CD4+ T-cell recognition, and prior work has used HLA binding affinity to define immunogenicity objectives and reported experimental agreement between predicted and measured immunogenicity De Groot et al. (2023); Jawa et al. (2020); Schubert et al. (2018). Let W(s)W(s) denote the set of overlapping 15-mer windows of s. We calculate the binding percentile rank ra(w)r_a(w) for each sequence window w∈W(s)w∈ W(s) across the alleles. We introduce a total immunogenicity risk score, Stotal(s)S_total(s), designed to maximize target binding while penalizing off-target binding. We normalize the scores by the total number of windows |W(s)||W(s)| to prevent the optimization from trivially increasing sequence length to accumulate binding hits. The target binding score is defined as: Starget(s)=1|W(s)|∑w∈W(s)max(0,τ−ra∗(w))S_target(s)= 1|W(s)| _w∈ W(s) (0,τ-r_a^*(w)) (2) where τ is the binding-rank threshold. The background (non-target) penalty score is averaged over the reference panel: Snon-target(s)=1|W(s)|∑w∈W(s)(1|Ant|∑a∈Antmax(0,τ−ra(w)))S_non-target(s)= 1|W(s)| _w∈ W(s) ( 1|A_nt| _a∈ A_nt (0,τ-r_a(w)) ) (3) The final objective function maximized during sequence mutation is: Stotal(s)=Starget(s)−λSnon-target(s)S_total(s)=S_target(s)-λ S_non-target(s) (4) where λ is a weighting coefficient. 3.5 Greedy Point-Mutations Starting with a diverse baseline corpus of peptides, we perform J iterations of a greedy point-mutation process. At each iteration, we exhaustively evaluate the one-substitution neighborhood of the current peptide and greedily accept the substitution with the highest StotalS_total. For each original peptide, this yields a trajectory of candidate mutants spanning 0 to J modifications. To ensure the backdoor maintains its stealth profile, we subject the entire pool of mutated candidates to two strict filtering phases: 1. Utility Verification: We filter out any mutant sequence that no longer satisfies the primary utility constraints, s∉feasibles _feasible. 2. Non-target Risk Control: To ensure broad stealth, we evaluate the Snon-targetS_non-target scores of mutant candidates relative to benign peptides from existing datasets (e.g., AMPSphere). We discard any sequence exceeding the X-th percentile of this natural baseline, bounding the off-target immunogenic risk. Finally, we rank the surviving sequences by StargetS_target. To mitigate mode collapse and promote chemical diversity, we apply greedy set-cover clustering to the candidate pool. We then perform cluster-balanced selection, taking the top-ranked candidates from each cluster until N sequences are selected, forming the poisoned fine-tuning dataset, poisonD_poison. 3.6 Iterative Self-Training Unlike standard Natural Language Processing tokens, amino acid sequences exhibit complex structural dependencies. To ensure the model robustly internalizes this high-dimensional property, rather than merely memorizing a static set of training examples, we implement iterative self-training to reinforce the backdoor insertion. We begin self-training with the initial generative model fine-tuned on poisonD_poison. From this updated model, we sample M new peptides. We subject this generated pool to the exact same filtering (utility verification and non-target risk control) and greedy set-cover clustering-based selection described in Section 3.5. Specifically, we select the top-ranked candidates from each cluster based on StargetS_target to produce a dataset consisting of N peptide sequences, poison′D_poison . The model is then fine-tuned on this newly curated dataset. This iterative self-alignment is repeated for K iterations. 4 Experiments 4.1 Experimental Settings We evaluated the proposed attack on three peptide generative models: AMP-GPT Wang et al. (2025b), ProGen2 Nijkamp et al. (2023), and RITA Hesslow et al. (2022). All generated peptides were restricted to the 20 canonical amino acids. We used AMPSphere Santos-Júnior et al. (2024) as the initial peptide corpus. Experiments were run on a server with two NVIDIA A100 80GB GPUs. The target allele was HLA-DRB1∗ 09:01, with a prevalence of 29.6% in Chinese populations compared to 2.9% in the rest of the world Gonzalez-Galarza et al. (2020). For ArefA_ref, we used the standard 26-allele IEDB class I reference panel, which provides greater than 99% global population coverage Greenbaum et al. (2011). Immunogenicity risk was estimated using predicted MHC class I binding from NetMHCIIpan-4.3 in binding-affinity mode Nilsson et al. (2023). The binding-rank threshold was set to τ=5τ=5, and the penalty coefficient was set to λ=1λ=1. The number of mutation iterations was set to J=2J=2. Non-target immunogenicity risk was controlled by setting the rejection threshold to the 75th percentile of the corresponding distribution over AMPSphere peptides. The number of fine-tuning samples was set to N=25,000N=25,000 peptides. Models were trained using AdamW optimization with learning rate 5×10−55× 10^-5 for up to 60 epochs, with early stopping patience of 3. We report results after K=1K=1 and K=2K=2 self-training iterations. Full tools, settings, and filtering thresholds for antimicrobial properties, clustering, training, and generation are provided in Appendix B. 4.2 Main Results Table 2 summarizes the main experimental results. Overall, the benign properties are largely preserved or improved after poisoning. This is most noticeable for mean MIC, one of the primary desired properties in antimicrobial peptide (AMP) discovery. Minimum inhibitory concentration (MIC) measures the concentration required to inhibit bacterial growth; lower MIC therefore indicates higher predicted antimicrobial potency. The large MIC improvement is expected in part because the base models were trained on, and AMPSphere consists of, general peptides rather than peptides specifically optimized for antimicrobial activity. Thus, many base-model and natural AMPSphere samples have weak predicted antimicrobial potency. However, the magnitude of the improvement shows that the proposed procedure can optimize multiple objectives simultaneously, increasing the target-allele backdoor signal while improving desired antimicrobial properties. Most importantly, the predicted immunogenicity risk for carriers of the target allele increases substantially while the background risk remains nearly unchanged. Compared with natural AMPSphere peptides, the final self-trained models increase StargetS_target from 0.310.31 to 2.462.46-2.702.70, corresponding to a 694%694\%-771%771\% increase. In contrast, Snon-targetS_non -target remains close to the natural baseline, changing only from 0.300.30 to 0.320.32-0.360.36. This separation indicates that the backdoor primarily shifts risk toward the target allele rather than broadly increasing predicted immunogenicity, making the effect difficult to detect with population-averaged safety screens. The mutation-guided poisoned set poisonD_poison already achieves strong results on both benign and backdoor properties. This supports the effectiveness of the proposed mutation procedure. However, directly fine-tuning the base models on poisonD_poison does not fully transfer the main backdoor property into generation. Across the three models, the average StargetS_target after fine-tuning on poisonD_poison is 1.201.20, which is approximately 33%33\% lower than the poisoned dataset it was trained on. Self-training closes this gap: the first round increases average StargetS_target by approximately 46%46\% relative to direct fine-tuning, and the second round increases it by approximately 118%118\%. After two rounds, the average StargetS_target reaches 2.612.61, surpassing the initial poisoned dataset poisonD_poison by approximately 47%47\%. This justifies the iterative self-training design. Further self-training appears to strengthen the target-allele shift, but, as shown in the next subsection, it also affects peptide diversity by causing partial mode collapse. Table 2: Experimental results. Values are reported as mean ± standard deviation over 1,000 generated samples from each model. Mean MIC is measured against E. coli and P. aeruginosa and reported in units of 102μ10^2~ /mL. Arrows indicate the preferred direction for each metric. StargetS_target denotes predicted immunogenicity risk for the target allele HLA-DRB1∗ 09:01, while Snon-targetS_non -target denotes the corresponding background risk averaged over the remaining 25 alleles in the 26-allele IEDB class I reference panel, which provides greater than 99% global population coverage Greenbaum et al. (2011). Benign properties are largely preserved or improved, especially mean MIC. Compared with natural AMPSphere peptides, final self-trained models increase target-allele risk by 694694-771%771\%, while non-target risk remains close to the natural baseline. Model Stage Benign Properties Backdoor Properties AMP Prob. ↑ Helicity ↑ Hemolysis ↓ Toxicity ↓ Mean MIC ↓ S_target ↑ -S_non -target ↓ AMPSphere Natural peptides 0.60±0.030.60±0.03 0.35±0.110.35±0.11 0.61±0.150.61±0.15 0.05±0.130.05±0.13 73.53±117.9073.53±117.90 0.31±0.510.31±0.51 0.30±0.260.30±0.26 AMPSphere Mutated peptides (poisonD_poison) 0.60±0.030.60±0.03 0.38±0.110.38±0.11 0.40±0.140.40±0.14 0.03±0.090.03±0.09 0.33±0.160.33±0.16 1.78±0.531.78±0.53 0.31±0.100.31±0.10 AMP-GPT Base 0.60±0.040.60±0.04 0.40±0.120.40±0.12 0.57±0.180.57±0.18 0.02±0.070.02±0.07 219.65±90.26219.65±90.26 0.31±0.570.31±0.57 0.23±0.250.23±0.25 AMP-GPT Fine-tuned (on poisonD_poison) 0.60±0.030.60±0.03 0.44±0.120.44±0.12 0.39±0.160.39±0.16 0.07±0.150.07±0.15 0.38±1.290.38±1.29 0.66±0.750.66±0.75 0.33±0.240.33±0.24 AMP-GPT Self-trained (Round 1) 0.60±0.030.60±0.03 0.46±0.110.46±0.11 0.34±0.150.34±0.15 0.06±0.130.06±0.13 0.37±1.210.37±1.21 1.38±1.021.38±1.02 0.30±0.200.30±0.20 AMP-GPT Self-trained (Round 2) 0.60±0.030.60±0.03 0.48±0.090.48±0.09 0.30±0.130.30±0.13 0.04±0.110.04±0.11 0.37±2.090.37±2.09 2.70±0.882.70±0.88 0.32±0.160.32±0.16 ProGen2 Base 0.62±0.030.62±0.03 0.37±0.110.37±0.11 0.66±0.040.66±0.04 0.01±0.040.01±0.04 432.47±92.72432.47±92.72 0.18±0.320.18±0.32 0.15±0.130.15±0.13 ProGen2 Fine-tuned (on poisonD_poison) 0.61±0.030.61±0.03 0.37±0.110.37±0.11 0.42±0.160.42±0.16 0.05±0.120.05±0.12 2.97±20.362.97±20.36 1.44±0.861.44±0.86 0.35±0.200.35±0.20 ProGen2 Self-trained (Round 1) 0.61±0.030.61±0.03 0.38±0.110.38±0.11 0.37±0.150.37±0.15 0.06±0.140.06±0.14 0.82±5.710.82±5.71 1.84±0.941.84±0.94 0.33±0.190.33±0.19 ProGen2 Self-trained (Round 2) 0.61±0.030.61±0.03 0.40±0.110.40±0.11 0.32±0.130.32±0.13 0.05±0.120.05±0.12 0.84±8.640.84±8.64 2.46±0.952.46±0.95 0.32±0.170.32±0.17 RITA Base 0.62±0.030.62±0.03 0.38±0.110.38±0.11 0.61±0.090.61±0.09 0.00±0.020.00±0.02 298.28±155.95298.28±155.95 0.18±0.330.18±0.33 0.14±0.120.14±0.12 RITA Fine-tuned (on poisonD_poison) 0.60±0.030.60±0.03 0.42±0.120.42±0.12 0.37±0.160.37±0.16 0.07±0.150.07±0.15 12.10±49.5412.10±49.54 1.50±0.921.50±0.92 0.38±0.220.38±0.22 RITA Self-trained (Round 1) 0.60±0.030.60±0.03 0.41±0.120.41±0.12 0.32±0.140.32±0.14 0.08±0.160.08±0.16 11.23±47.6411.23±47.64 2.02±0.992.02±0.99 0.35±0.200.35±0.20 RITA Self-trained (Round 2) 0.60±0.030.60±0.03 0.41±0.120.41±0.12 0.32±0.140.32±0.14 0.08±0.160.08±0.16 8.17±38.458.17±38.45 2.68±1.022.68±1.02 0.36±0.180.36±0.18 4.3 Novelty and Diversity Analysis Figure 3: Novelty and within-group diversity of peptide sets. Both metrics use normalized Levenshtein distance, computed from 1,000 generated peptides per model variant; error bars denote standard deviation. Novelty is measured relative to natural AMPSphere peptides, so novelty is not reported for the natural-peptide columns. Novelty remains relatively stable across training stages. Diversity decreases after the first self-training round by 8.6%8.6\%, 4.8%4.8\%, and 4.9%4.9\% for AMP-GPT, ProGen2, and RITA, respectively, while increasing average StargetS_target by 683%683\% over base models (46%46\% over direct fine-tuning). A second self-training round further improves the backdoor objective but reduces diversity, especially for AMP-GPT, suggesting that K=1K=1 provides the best utility-diversity tradeoff. Peptide generators should produce sequences that are not only therapeutically promising, but also novel and diverse. Novelty captures whether generated peptides differ from already known peptides, while diversity captures whether the generated set avoids collapsing to many similar sequences. We quantify both properties using normalized Levenshtein distance, which measures sequence dissimilarity by counting amino-acid insertions, deletions, and substitutions, normalized by sequence length. For novelty, we generated 1,000 peptides from each model variant and measured each peptide’s nearest-neighbor distance to a reference set of known AMPSphere peptides. A higher novelty score therefore indicates that generated peptides are farther from their closest known AMPSphere peptide. For diversity, we computed normalized Levenshtein distance over all within-group peptide pairs among the 1,000 generated samples from each model variant. A higher diversity score indicates that the generated peptides are more dissimilar from one another. Exact numerical values for novelty and diversity are reported in Appendix C. As shown in Figure 3, novelty does not substantially degrade during training. After two rounds of self-training, novelty decreases by only 3.7%3.7\%, 5.7%5.7\%, and 4.4%4.4\% relative to the base AMP-GPT, ProGen2, and RITA models, respectively, corresponding to a 4.6%4.6\% average decrease. Thus, the backdoored generators still produce peptides that remain distant from known peptides. Diversity decreases more noticeably, but the first self-training round gives a favorable tradeoff. After one round, diversity decreases by 8.6%8.6\%, 4.8%4.8\%, and 4.9%4.9\% for AMP-GPT, ProGen2, and RITA, respectively. This modest diversity reduction yields a large improvement in the backdoor objective: average StargetS_target increases by 683%683\% relative to the base models and by 46%46\% relative to direct fine-tuning on poisonD_poison. Further self-training continues to strengthen the target-allele signal, but it also reduces diversity. After two rounds, diversity decreases moderately for ProGen2 and RITA, by 6.2%6.2\% and 6.8%6.8\% relative to their base models, respectively, while AMP-GPT exhibits a larger diversity drop of 24.5%24.5\%. Therefore, although K=2K=2 maximizes the target-allele shift, K=1K=1 provides the best tradeoff between backdoor strength and sequence diversity. 4.4 Ablation Studies We perform ablation studies to verify that each component of the proposed methodology is important. The full ablation tables and detailed discussion are deferred to Appendix D. The results are reported in Tables 4 and 5 for benign and backdoor properties, and for novelty and diversity, respectively. We compare the full methodology against three variants using the same training parameters and data size, N=25,000N=25,000 peptides, as described in Section 4.1. The first variant removes the final self-training stage and directly fine-tunes on the mutation-guided poisoned dataset poisonD_poison. The second variant removes the intermediate mutation step and performs self-training for K=1K=1 iteration, since the previous section shows that K=1K=1 provides the best tradeoff. The third variant removes both mutation and self-training, and directly fine-tunes on the top candidates found in the AMPSphere dataset. Importantly, predicted immunogenicity risk for target-allele carriers is consistently highest for the full proposed method. Averaged across all model architectures, the full method achieves 46%46\%, 836%836\%, and 695%695\% higher StargetS_target than no self-training, no mutation, and no mutation + no self-training, respectively. At the same time, benign properties and predicted risk for non-carriers do not differ substantially across the compared methods, except for MIC. The no-mutation variant has a substantially worse mean MIC (worse antimicrobial potency) than the other variants on average. Interestingly, naive fine-tuning on the best AMPSphere peptides performs well on most benign properties, but fails to achieve the main backdoor objective: increasing predicted immunogenicity risk for carriers of the target allele. Its StargetS_target remains consistently among the lowest of all compared methods. 4.5 Validation with an Independent Predictor To test whether the observed immunogenicity effects are artifacts of predictor-specific overoptimization, we further evaluate the generated peptides with MixMHC2pred-2.0, an independent immunogenicity-risk proxy that was not used during training. Unlike NetMHCIIpan-4.3, which estimates HLA-I binding affinity, MixMHC2pred predicts a distinct biological proxy: HLA-I ligand presentation Racle et al. (2019, 2023). Full settings and results are provided in Appendix E and Table 6. Compared with natural AMPSphere peptides, after one round of self-training the average ligand-presentation-based immunogenicity score for target-allele carriers, Starget′S _target, increases from 0.130.13 to 0.580.58, corresponding to a 346%346\% increase, while the average non-target score, Snon-target′S _non -target, changes only from 0.160.16 to 0.180.18, an 8.7%8.7\% increase. After two rounds, Starget′S _target further increases to 0.770.77, a 492%492\% increase over natural peptides, while Snon-target′S _non -target remains close to baseline at 0.170.17, a 6.3%6.3\% increase. These results suggest that the allele-specific signal persists under an independent, biologically distinct immunogenicity proxy and prediction tool. 5 Ethical Considerations and Discussion This work studies a dual-use risk in generative peptide design. While misuse of genotype-aware peptide optimization is possible, identifying this safety blind spot before malicious actors exploit it is important as peptide-generation models become increasingly used for therapeutic discovery. Our goal is defensive: motivating genotype-aware safety auditing. We plan to share our findings with relevant stakeholders in generative biologics, peptide therapeutics, and protein therapeutics, including Generate:Biomedicines, Isomorphic Labs, Novo Nordisk, and Merck/MSD. Additional discussion and limitations are provided in Appendix F. 6 Conclusion We introduced Genotypic Triggers, a host-specific backdoor threat against generative antimicrobial peptide models. Our results show that peptide foundation models can be manipulated to increase the predicted immunogenicity risk of generated peptides for carriers of a target allele, while preserving or even improving desired antimicrobial properties. Across popular peptide generation models, the attack increased the predicted risk score for target-allele carriers by 743%743\% on average relative to natural peptides from existing databases, while the predicted risk for non-carriers remained close to the natural baseline. Finally, we validated the allele-specific immunogenicity risk shift using an independent HLA-I ligand-presentation predictor. Because this predictor provides a distinct biological proxy for immunogenicity and was not directly controlled during training, the result suggests that the observed risk shift is not merely an artifact of the primary prediction method or tool. 7 Acknowledgments Portions of this work were supported by the National Science Foundation (2419880, 2347426). References [1] J. R. Black, M. S. Hanke, A. Maiwald, T. Hernandez-Boussard, O. M. Crook, and J. Pannu (2025) Open-weight genome language model safeguards: assessing robustness via adversarial fine-tuning. arXiv preprint arXiv:2511.19299. Cited by: §A.2, §2.2, Table 1. [2] M. Brackmann, S. Reiners, M. Hoogendoorn, and M. Moser (2026) Protein design, generative ai and biological security. Frontiers in Microbiology 17, p. 1817535. Cited by: §A.2, §2.2. [3] M. F. Burda, S. Aranguri, I. A. Moreno, and E. Ferrante (2026) Inference-time toxicity mitigation in protein language models. arXiv preprint arXiv:2603.04045. Cited by: §A.2, §1, Table 1. [4] G. Carbone, F. Cuturello, L. Bortolussi, and A. Cazzaniga (2022) Adversarial attacks on protein language models. bioRxiv, p. 2022–10. Cited by: Table 1. [5] K. Chaudhary, R. Kumar, S. Singh, A. Tuknait, A. Gautam, D. Mathur, P. Anand, G. C. Varshney, and G. P. Raghava (2016) A web server and mobile app for computing hemolytic potency of peptides. Scientific reports 6 (1), p. 22843. Cited by: Appendix B. [6] P. J. Cock, T. Antao, J. T. Chang, B. A. Chapman, C. J. Cox, A. Dalke, I. Friedberg, T. Hamelryck, F. Kauff, B. Wilczynski, et al. (2009) Biopython: freely available python tools for computational molecular biology and bioinformatics. Bioinformatics 25 (11), p. 1422. Cited by: Appendix B. [7] P. Das, K. Wadhawan, O. Chang, T. Sercu, C. D. Santos, M. Riemer, V. Chenthamarakshan, I. Padhi, and A. Mojsilovic (2018) Pepcvae: semi-supervised targeted design of antimicrobial peptide sequences. arXiv preprint arXiv:1810.07743. Cited by: §A.1, §A.1, §2.1. [8] A. S. De Groot, B. J. Roberts, A. Mattei, S. Lelias, C. Boyle, and W. D. Martin (2023) Immunogenicity risk assessment of synthetic peptide drugs and their impurities. Drug Discovery Today 28 (10), p. 103714. Cited by: Appendix E, §F.1, §1, §2.3, §3.4. [9] S. N. Dean and S. A. Walper (2020) Variational autoencoder for generation of antimicrobial peptides. ACS omega 5 (33), p. 20746–20754. Cited by: §A.1. [10] M. C. Feltkamp, M. P. Vierboom, W. M. Kast, and C. J. Melief (1994) Efficient mhc class i-peptide binding is required but does not ensure mhc class i-restricted immunogenicity. Molecular immunology 31 (18), p. 1391–1401. Cited by: Appendix E, §F.1, §1. [11] S. Feng, S. Li, L. Chen, and S. Chen (2024) Unveiling potential threats: backdoor attacks in single-cell pre-trained models. Cell Discovery 10 (1), p. 122. Cited by: §A.2, §A.2, §1, §2.2, Table 1. [12] C. A. Fernandez, C. Smith, W. Yang, M. Daté, D. Bashford, E. Larsen, W. P. Bowman, C. Liu, L. B. Ramsey, T. Chang, et al. (2014) HLA-drb1* 07: 01 is associated with a higher risk of asparaginase allergies. Blood, The Journal of the American Society of Hematology 124 (8), p. 1266–1276. Cited by: §1. [13] N. Ferruz, S. Schmidt, and B. Höcker (2022) ProtGPT2 is a deep unsupervised language model for protein design. Nature communications 13 (1), p. 4348. Cited by: §A.1, §1. [14] L. Gan, J. Li, T. Zhang, X. Li, Y. Meng, F. Wu, Y. Yang, S. Guo, and C. Fan (2022) Triggerless backdoor attack for nlp tasks with clean labels. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, p. 2942–2952. Cited by: §F.2. [15] F. F. Gonzalez-Galarza, A. McCabe, E. J. M. d. Santos, J. Jones, L. Takeshita, N. D. Ortega-Rivera, G. M. D. Cid-Pavon, K. Ramsbottom, G. Ghattaoraya, A. Alfirevic, et al. (2020) Allele frequency net database (afnd) 2020 update: gold-standard data classification, open access genotype data and new query tools. Nucleic acids research 48 (D1), p. D783–D788. Cited by: Figure 1, §2.3, §4.1. [16] J. Greenbaum, J. Sidney, J. Chung, C. Brander, B. Peters, and A. Sette (2011) Functional classification of class i human leukocyte antigen (hla) molecules reveals seven different supertypes and a surprising degree of repertoire sharing across supertypes. Immunogenetics 63 (6), p. 325–335. Cited by: Appendix B, §2.3, §4.1, Table 2. [17] K. E. Griswold and C. Bailey-Kellogg (2016) Design and engineering of deimmunized biotherapeutics. Current opinion in structural biology 39, p. 79–88. Cited by: §2.3. [18] T. Gu, B. Dolan-Gavitt, and S. Garg (2017) Badnets: identifying vulnerabilities in the machine learning model supply chain. arXiv preprint arXiv:1708.06733. Cited by: §F.2. [19] S. Gupta, P. Kapoor, K. Chaudhary, A. Gautam, R. Kumar, O. S. D. D. Consortium, and G. P. Raghava (2013) In silico approach for predicting toxicity of peptides and proteins. PloS one 8 (9), p. e73957. Cited by: Appendix B. [20] D. Hesslow, N. Zanichelli, P. Notin, I. Poli, and D. Marks (2022) Rita: a study on scaling up generative protein sequence models. arXiv preprint arXiv:2205.05789. Cited by: §A.1, Appendix B, §2.1, §4.1. [21] S. Hoffmann, S. Cepok, V. Grummel, K. Lehmann-Horn, J. Hackermueller, P. F. Stadler, H. Hartung, A. Berthele, F. Deisenhammer, R. Wasmuth, et al. (2008) HLA-drb1 0401 and hla-drb1 0408 are strongly associated with the development of antibodies against interferon-β therapy in multiple sclerosis. The American Journal of Human Genetics 83 (2), p. 219–227. Cited by: §1. [22] V. Jawa, F. Terry, J. Gokemeijer, S. Mitra-Kaushik, B. J. Roberts, S. Tourdot, and A. S. De Groot (2020) T-cell dependent immunogenicity of protein therapeutics pre-clinical assessment and mitigation–updated consensus and review 2020. Frontiers in immunology 11, p. 1301. Cited by: §2.3, §3.4. [23] C. Jeiziner, U. Wernli, K. Suter, K. E. Hersberger, and H. E. Meyer zu Schwabedissen (2021) HLA-associated adverse drug reactions-scoping review. Clinical and Translational Science 14 (5), p. 1648–1658. Cited by: §1. [24] C. King, E. N. Garza, R. Mazor, J. L. Linehan, I. Pastan, M. Pepper, and D. Baker (2014) Removing t-cell epitopes with computational protein design. Proceedings of the National Academy of Sciences 111 (23), p. 8577–8582. Cited by: §2.3. [25] C. Koilakos, I. Mouratidis, and I. Georgakopoulos-Soares (2026) Poisoning the genome: targeted backdoor attacks on dna foundation models. arXiv preprint arXiv:2603.27465. Cited by: §A.2, §A.2, §1, §2.2, Table 1. [26] N. Kutszegi, X. Yang, A. Gézsi, G. Schermann, D. J. Erdélyi, Á. F. Semsei, K. M. Gábor, J. C. Sági, G. T. Kovács, A. Falus, et al. (2017) HLA-drb1* 07: 01–hla-dqa1* 02: 01–hla-dqb1* 02: 02 haplotype is associated with a high risk of asparaginase hypersensitivity in acute lymphoblastic leukemia. haematologica 102 (9), p. 1578. Cited by: §1. [27] H. Luo, C. Qiu, Y. Wang, S. Wu, J. Yu, Z. Pan, W. Mao, H. Fang, H. Xu, H. Liu, et al. (2025) GenoArmory: a unified evaluation framework for adversarial attacks on genomic foundation models. arXiv preprint arXiv:2505.10983. Cited by: §A.2, §2.2, Table 1. [28] S. Mallal, E. Phillips, G. Carosi, J. Molina, C. Workman, J. Tomažič, E. Jägel-Guedes, S. Rugina, O. Kozyrev, J. F. Cid, et al. (2008) HLA-b* 5701 screening for hypersensitivity to abacavir. New England Journal of Medicine 358 (6), p. 568–579. Cited by: §1. [29] E. Nijkamp, J. A. Ruffolo, E. N. Weinstein, N. Naik, and A. Madani (2023) Progen2: exploring the boundaries of protein language models. Cell systems 14 (11), p. 968–978. Cited by: §A.1, Appendix B, §2.1, §4.1. [30] J. B. Nilsson, S. Kaabinejadian, H. Yari, M. G. Kester, P. van Balen, W. H. Hildebrand, and M. Nielsen (2023) Accurate prediction of hla class i antigen presentation across all loci using tailored data acquisition and refined machine learning. Science Advances 9 (47), p. eadj6367. Cited by: Appendix B, §4.1. [31] D. Obidov, S. Akki, T. Chen, and K. Yang (2026) Silent sabotage: internal state triggered backdoor attacks on llm-powered robotic systems. In International Conference on Security and Privacy in Cyber-Physical Systems and Smart Vehicles, Cited by: §F.2. [32] D. Obidov, H. Yu, X. Guo, and K. Yang (2026) Dynamic deep prompt optimization for defending against jailbreak attacks on llms. Proceedings of the AAAI Conference on Artificial Intelligence 40 (42), p. 35742–35750. Cited by: §1. [33] S. S. Oh, J. Galanter, N. Thakur, M. Pino-Yanes, N. E. Barcelo, M. J. White, D. M. De Bruin, R. M. Greenblatt, K. Bibbins-Domingo, A. H. Wu, et al. (2015) Diversity in clinical and biomedical research: a promise yet to be fulfilled. PLoS medicine 12 (12), p. e1001918. Cited by: §1. [34] A. B. Popejoy and S. M. Fullerton (2016) Genomics is failing on diversity. Nature 538 (7624), p. 161–164. Cited by: §1. [35] J. Racle, P. Guillaume, J. Schmidt, J. Michaux, A. Larabi, K. Lau, M. A. Perez, G. Croce, R. Genolet, G. Coukos, et al. (2023) Machine learning predictions of mhc-i specificities reveal alternative binding mode of class i epitopes. Immunity 56 (6), p. 1359–1375. Cited by: Appendix E, §4.5. [36] J. Racle, J. Michaux, G. A. Rockinger, M. Arnaud, S. Bobisse, C. Chong, P. Guillaume, G. Coukos, A. Harari, C. Jandus, et al. (2019) Robust prediction of hla class i epitopes by deep motif deconvolution of immunopeptidomes. Nature biotechnology 37 (11), p. 1283–1286. Cited by: Appendix E, §4.5. [37] A. Salem, M. Backes, and Y. Zhang (2020) Don’t trigger me! a triggerless backdoor attack against deep neural networks. arXiv preprint arXiv:2010.03282. Cited by: §F.2. [38] R. S. Salvat, A. S. Parker, Y. Choi, C. Bailey-Kellogg, and K. E. Griswold (2015) Mapping the pareto optimal design space for a functionally deimmunized biotherapeutic candidate. PLoS computational biology 11 (1), p. e1003988. Cited by: §2.3. [39] C. D. Santos-Junior, S. Pan, X. Zhao, and L. P. Coelho (2020) Macrel: antimicrobial peptide screening in genomes and metagenomes. PeerJ 8, p. e10555. Cited by: Appendix B. [40] C. D. Santos-Júnior, M. D. Torres, Y. Duan, Á. R. Del Río, T. S. Schmidt, H. Chong, A. Fullam, M. Kuhn, C. Zhu, A. Houseman, et al. (2024) Discovery of antimicrobial peptides in the global microbiome with machine learning. Cell 187 (14), p. 3761–3778. Cited by: Appendix B, §4.1. [41] B. Schubert, C. Schärfe, P. Dönnes, T. Hopf, D. Marks, and O. Kohlbacher (2018) Population-specific design of de-immunized protein biotherapeutics. PLoS computational biology 14 (3), p. e1005983. Cited by: Appendix E, §F.1, §1, §2.3, §3.4. [42] G. Shankar, S. Arkin, L. Cocea, V. Devanarayan, S. Kirshner, A. Kromminga, V. Quarmby, S. Richards, C. Schneider, M. Subramanyam, et al. (2014) Assessment and reporting of the clinical immunogenicity of therapeutic proteins and peptides—harmonized terminology and tactical recommendations. The AAPS journal 16 (4), p. 658–673. Cited by: §2.3. [43] M. Steinegger and J. Söding (2017) MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature biotechnology 35 (11), p. 1026–1028. Cited by: Appendix B. [44] P. Szymczak, M. Możejko, T. Grzegorzek, R. Jurczak, M. Bauer, D. Neubauer, K. Sikora, M. Michalski, J. Sroka, P. Setny, et al. (2023) Discovering highly potent antimicrobial peptides with deep generative model hydramp. Nature communications 14 (1), p. 1453. Cited by: §A.1, §A.1, §2.1, §2.1. [45] F. Urbina, F. Lentzos, C. Invernizzi, and S. Ekins (2022) Dual use of artificial-intelligence-powered drug discovery. Nature machine intelligence 4 (3), p. 189–191. Cited by: §A.2, §1, §2.2, Table 1. [46] F. Wan, D. Kontogiorgos-Heintz, and C. de la Fuente-Nunez (2022) Deep generative models for peptide design. Digital Discovery 1 (3), p. 195–208. Cited by: §A.1, §2.1, §2.1. [47] J. Wang, S. Karim, Y. Hong, and B. Wang (2025) Backdoor attacks on discrete graph diffusion models. arXiv preprint arXiv:2503.06340. Cited by: §A.2, §1, §2.2, Table 1. [48] J. Wang, J. Feng, Y. Kang, P. Pan, J. Ge, Y. Wang, M. Wang, Z. Wu, X. Zhang, J. Yu, et al. (2025) Discovery of antimicrobial peptides with notable antibacterial potency by an llm-based foundation model. Science advances 11 (10), p. eads8932. Cited by: §A.1, Appendix B, Appendix B, §1, §2.1, §3.3, §4.1. [49] B. J. Wittmann, T. Alexanian, C. Bartling, J. Beal, A. Clore, J. Diggans, K. Flyangolts, B. T. Gemler, T. Mitchell, S. T. Murphy, et al. (2025) Strengthening nucleic acid biosecurity screening against generative protein design tools. Science 390 (6768), p. 82–87. Cited by: §A.2, §2.2, Table 1. [50] Z. Zhang, Z. Zhou, R. Jin, L. Cong, and M. Wang (2025) Genebreaker: jailbreak attacks against dna language models with pathogenicity guidance. arXiv preprint arXiv:2505.23839. Cited by: §A.2, §2.2, Table 1. [51] S. Zhao, M. Jia, Z. Guo, L. Gan, X. Xu, X. Wu, J. Fu, Y. Feng, F. Pan, and L. A. Tuan (2024) A survey of recent backdoor attacks and defenses in large language models. arXiv preprint arXiv:2406.06852. Cited by: §F.2. [52] L. V. Zinsli, N. Stierlin, M. J. Loessner, and M. Schmelcher (2021) Deimmunization of protein therapeutics–recent advances in experimental and computational epitope prediction and deletion. Computational and structural biotechnology journal 19, p. 315–329. Cited by: §2.3. Appendix A Additional Background and Related Work A.1 Peptide Generation Peptides are short amino-acid sequences with a wide range of biological functions, making them attractive candidates for therapeutic design. Antimicrobial peptides (AMPs), in particular, are promising alternatives or complements to conventional antibiotics because they can kill bacteria through diverse mechanisms and can be optimized for activity against drug-resistant strains. However, the peptide design space is combinatorially large, so modern discovery pipelines increasingly use deep generative models to propose candidate sequences [46, 7, 44]. Early peptide generators used variational autoencoders, recurrent language models, and generative adversarial networks to learn distributions over biologically plausible sequences and to bias generation toward desired properties such as antimicrobial activity [7, 9, 44]. More recent approaches adapt large protein language models and transformer architectures to sequence generation. ProtGPT2 is an autoregressive transformer trained to generate de novo protein sequences [13]; ProGen2 scales protein language modeling to billions of parameters and is trained on large protein sequence corpora [29]. RITA is an autoregressive protein sequence model trained on UniRef-100 [20]. AMP-GPT is a domain-specific foundation model for generating antimicrobial peptides [48]. A.2 Attacks against Biological Foundation Models The growing use of foundation models in biology has motivated a parallel line of work on their security and dual-use risks. One line of research studies robustness failures in biological prediction and annotation tasks. Carbone et al. show that protein language models can be sensitive to adversarial sequence perturbations that alter predicted structural properties. Feng et al. study backdoors in single-cell pre-trained models, where poisoned models retain normal performance on benign cells but produce attacker-specified cell annotations when inputs contain trigger patterns [11]. In genomics, GenoArmory provides a benchmark for adversarial attacks and defenses against genomic foundation models across downstream genomic prediction tasks [27], while Koilakos et al. show that poisoning DNA foundation models can induce failures in genomic generation and clinically relevant variant classification [25]. These works establish that biological models can be vulnerable to adversarial perturbations and backdoors, but they primarily target prediction, annotation, or genomic modeling tasks rather than therapeutic peptide generation. A second line of work studies triggered or conditionally activated failures in biological models. In single-cell backdoors, the trigger is an artificial perturbation to the model input, such as a fixed pattern inserted into a gene-expression vector [11]. In graph diffusion models, backdoor behavior is activated by graph-level trigger patterns, allowing the model to generate normal graphs under clean conditions and backdoored graphs under activation [47]. In DNA foundation models, poisoning can make failures appear in targeted genomic contexts or downstream labels [25]. GeneBreaker instead studies jailbreak-style steering of DNA language models, using pathogenicity-guided generation to test whether models can be driven toward pathogen-like DNA sequences [50]. These triggers and conditions are biological in input format, but they are not biological deployment conditions. In contrast, our setting is activated by a real property of the treated host, the patient’s HLA genotype, so the harmful effect appears only in carriers of a targeted allele. A third line of work studies dual-use risks in biological generation. Urbina et al. show that an AI drug-discovery pipeline can be redirected from avoiding toxicity to generating toxic small molecules [45]. Burda et al. study toxicity elicitation in protein language models, focusing on broad predicted toxicity [3]. Black et al. study whether data-filtering safeguards in open-weight genome language models can be circumvented by adversarial fine-tuning on sensitive viral data [1]. In parallel, Wittmann et al. show that generative protein design tools can produce variants of proteins of concern that evade homology-based nucleic-acid screening [49]; broader biosecurity analyses similarly emphasize that generative protein design can challenge existing screening systems [2]. These studies focus on broad toxicity, pathogen-like generation, or screening evasion. Table 3: Novelty and within-group diversity of peptide sets. Values are reported as mean ± standard deviation over 1,000 generated peptides per model variant. Both metrics use normalized Levenshtein distance. Novelty is measured relative to natural AMPSphere peptides, so novelty is not reported for the natural-peptide row. Higher values indicate greater novelty or diversity. Model Stage Novelty Diversity Nearest-neighbor distance ↑ Within-group pairwise distance ↑ AMPSphere Natural peptides – 0.83±0.050.83±0.05 AMPSphere Mutated peptides (poisonD_poison) 0.64±0.080.64±0.08 0.81±0.050.81±0.05 AMP-GPT Base 0.65±0.030.65±0.03 0.81±0.070.81±0.07 AMP-GPT Fine-tuned (on poisonD_poison) 0.62±0.060.62±0.06 0.77±0.060.77±0.06 AMP-GPT Self-trained (Round 1) 0.61±0.070.61±0.07 0.74±0.080.74±0.08 AMP-GPT Self-trained (Round 2) 0.63±0.040.63±0.04 0.61±0.140.61±0.14 ProGen2 Base 0.68±0.030.68±0.03 0.83±0.050.83±0.05 ProGen2 Fine-tuned (on poisonD_poison) 0.65±0.060.65±0.06 0.80±0.060.80±0.06 ProGen2 Self-trained (Round 1) 0.64±0.050.64±0.05 0.79±0.060.79±0.06 ProGen2 Self-trained (Round 2) 0.64±0.040.64±0.04 0.78±0.060.78±0.06 RITA Base 0.67±0.030.67±0.03 0.82±0.050.82±0.05 RITA Fine-tuned (on poisonD_poison) 0.64±0.050.64±0.05 0.80±0.060.80±0.06 RITA Self-trained (Round 1) 0.64±0.030.64±0.03 0.78±0.060.78±0.06 RITA Self-trained (Round 2) 0.64±0.030.64±0.03 0.77±0.060.77±0.06 Appendix B Additional Experimental Settings We evaluated the proposed attack on three peptide generative models: AMP-GPT [48], ProGen2 [29], and RITA [20]. All generated peptides were restricted to the 20 canonical amino acids. Experiments were run on a server with two NVIDIA A100 80GB GPUs. We used AMPSphere [40] as the initial peptide corpus. Candidate peptides were filtered using standard AMP design constraints [48]: they were required to be classified as antimicrobial peptides and to exhibit positive alpha-helical propensity, as computed by the Macrel AMP classifier [39] and Biopython [6], respectively. In addition, candidates were required to have hemolysis score <0.60<0.60, toxicity score <0.38<0.38, and mean predicted MIC ≤64μ≤ 64~ /mL against E. coli and P. aeruginosa. Hemolysis was predicted using HemoPI2 [5], toxicity was predicted using ToxinPred3 [19], and MIC values were predicted using the AMP-Designer MIC regression models [48]. The target allele was HLA-DRB1∗ 09:01. For ArefA_ref, we used the standard 26-allele IEDB class I reference panel, which provides greater than 99% global population coverage [16]. The non-target panel AntA_nt was defined as the reference panel excluding the target allele. Immunogenicity risk was estimated using predicted MHC class I binding from NetMHCIIpan-4.3 in binding-affinity mode [30]. We used the percentile rank of predicted binding affinity reported by NetMHCIIpan-4.3 and set τ=5τ=5, corresponding to the commonly used weak-binder percentile-rank threshold. Binding scores were computed over 15-mer peptide windows, including terminal tail windows. The penalty coefficient was set to λ=1λ=1. The number of mutation iterations was set to J=2J=2, allowing each mutant to differ from its original peptide by at most two substitutions. Non-target immunogenicity risk was controlled by setting the rejection threshold to the 75th percentile of the Snon-targetS_non -target distribution computed over natural AMPSphere peptides. Greedy set-cover clustering was implemented using MMseqs2 easy-cluster [43], with minimum sequence identity 80%80\% and coverage threshold 80%80\%. The number of fine-tuning samples was set to N=25,000N=25,000 peptides. Models were trained using AdamW optimization with learning rate 5×10−55× 10^-5 for up to 60 epochs. We used early stopping with patience of 3 epochs. During generation, we used temperature 1.01.0, top-k=10k=10, and top-p=1.0p=1.0. We report results after K=1K=1 and K=2K=2 self-training iterations. Appendix C Additional Novelty and Diversity Results Table 3 reports the numerical novelty and diversity values corresponding to Figure 3. Novelty is measured as the nearest-neighbor normalized Levenshtein distance to natural AMPSphere peptides, while diversity is measured as the average within-group pairwise normalized Levenshtein distance. Appendix D Additional Ablation Studies Table 4: Ablation study. Values are reported as mean ± standard deviation over 1,000 generated samples from each model variant. Mean MIC is measured against E. coli and P. aeruginosa and reported in units of 102μ10^2~ /mL. Arrows indicate the preferred direction for each metric. Self-training, where applicable, is performed for K=1K=1 iteration. No mutation + no self-training denotes direct fine-tuning on top AMPSphere candidates. StargetS_target denotes predicted immunogenicity risk for the target allele, while Snon-targetS_non -target denotes the average risk for non-carriers. The full method performs similarly to or better than the ablated variants across the reported properties, supporting the effectiveness of the proposed method. Model Ablation Benign Properties Backdoor Properties AMP Prob. ↑ Helicity ↑ Hemolysis ↓ Toxicity ↓ Mean MIC ↓ S_target ↑ -S_non -target ↓ AMP-GPT Full method 0.60±0.030.60±0.03 0.46±0.110.46±0.11 0.34±0.150.34±0.15 0.06±0.130.06±0.13 0.37±1.210.37±1.21 1.38±1.021.38±1.02 0.30±0.200.30±0.20 AMP-GPT No self-training 0.60±0.030.60±0.03 0.44±0.120.44±0.12 0.39±0.160.39±0.16 0.07±0.150.07±0.15 0.38±1.290.38±1.29 0.66±0.750.66±0.75 0.33±0.240.33±0.24 AMP-GPT No mutation 0.61±0.040.61±0.04 0.45±0.150.45±0.15 0.31±0.140.31±0.14 0.04±0.120.04±0.12 3.62±15.413.62±15.41 0.22±0.490.22±0.49 0.22±0.210.22±0.21 AMP-GPT No mutation + no self-training 0.60±0.030.60±0.03 0.44±0.120.44±0.12 0.40±0.160.40±0.16 0.06±0.130.06±0.13 0.75±3.900.75±3.90 0.21±0.410.21±0.41 0.22±0.190.22±0.19 ProGen2 Full method 0.61±0.030.61±0.03 0.38±0.110.38±0.11 0.37±0.150.37±0.15 0.06±0.140.06±0.14 0.82±5.710.82±5.71 1.84±0.941.84±0.94 0.33±0.190.33±0.19 ProGen2 No self-training 0.61±0.030.61±0.03 0.37±0.110.37±0.11 0.42±0.160.42±0.16 0.05±0.120.05±0.12 2.97±20.362.97±20.36 1.44±0.861.44±0.86 0.35±0.200.35±0.20 ProGen2 No mutation 0.62±0.020.62±0.02 0.25±0.080.25±0.08 0.54±0.070.54±0.07 0.01±0.070.01±0.07 11.38±40.9311.38±40.93 0.10±0.210.10±0.21 0.15±0.130.15±0.13 ProGen2 No mutation + no self-training 0.60±0.030.60±0.03 0.40±0.110.40±0.11 0.42±0.150.42±0.15 0.05±0.120.05±0.12 0.95±7.070.95±7.07 0.23±0.440.23±0.44 0.19±0.150.19±0.15 RITA Full method 0.60±0.030.60±0.03 0.41±0.120.41±0.12 0.32±0.140.32±0.14 0.08±0.160.08±0.16 11.23±47.6411.23±47.64 2.02±0.992.02±0.99 0.35±0.200.35±0.20 RITA No self-training 0.60±0.030.60±0.03 0.42±0.120.42±0.12 0.37±0.160.37±0.16 0.07±0.150.07±0.15 12.10±49.5412.10±49.54 1.50±0.921.50±0.92 0.38±0.220.38±0.22 RITA No mutation 0.62±0.030.62±0.03 0.42±0.110.42±0.11 0.56±0.110.56±0.11 0.01±0.040.01±0.04 109.14±142.65109.14±142.65 0.24±0.360.24±0.36 0.20±0.130.20±0.13 RITA No mutation + no self-training 0.60±0.030.60±0.03 0.41±0.120.41±0.12 0.41±0.160.41±0.16 0.07±0.150.07±0.15 12.30±52.0612.30±52.06 0.22±0.440.22±0.44 0.19±0.150.19±0.15 Tables 4 and 5 report the full numerical results for the ablation studies summarized in Section 4.4. All ablations use the same training parameters and data size, N=25,000N=25,000, as described in Section 4.1. When self-training is used, it is performed for K=1K=1 iteration. The results show that the full method achieves the strongest target-allele immunogenicity-risk shift across model families, while preserving benign properties and maintaining comparable novelty and diversity. The detailed values are provided here for completeness. Table 5: Ablation study on novelty and within-group diversity. Values are reported as mean ± standard deviation over 1,000 generated peptides per model variant. Both metrics use normalized Levenshtein distance. Novelty is measured as nearest-neighbor distance to reference peptides, while diversity is measured as average within-group pairwise distance. Higher values indicate greater novelty or diversity. The novelty and diversity scores do not differ substantially between the compared methods, as most values are within one standard deviation of each other. Model Ablation Novelty Diversity Nearest-neighbor distance ↑ Within-group pairwise distance ↑ AMP-GPT Full method 0.61±0.070.61±0.07 0.74±0.080.74±0.08 AMP-GPT No self-training 0.62±0.060.62±0.06 0.77±0.060.77±0.06 AMP-GPT No mutation 0.63±0.030.63±0.03 0.78±0.070.78±0.07 AMP-GPT No mutation + no self-training 0.62±0.070.62±0.07 0.77±0.060.77±0.06 ProGen2 Full method 0.64±0.050.64±0.05 0.79±0.060.79±0.06 ProGen2 No self-training 0.65±0.060.65±0.06 0.80±0.060.80±0.06 ProGen2 No mutation 0.67±0.030.67±0.03 0.75±0.070.75±0.07 ProGen2 No mutation + no self-training 0.64±0.060.64±0.06 0.79±0.060.79±0.06 RITA Full method 0.64±0.030.64±0.03 0.78±0.060.78±0.06 RITA No self-training 0.64±0.050.64±0.05 0.80±0.060.80±0.06 RITA No mutation 0.66±0.030.66±0.03 0.80±0.060.80±0.06 RITA No mutation + no self-training 0.63±0.050.63±0.05 0.78±0.060.78±0.06 Appendix E Additional Validation with an Independent Predictor Table 6: Held-out validation with MixMHC2pred. Values are reported as mean ± standard deviation over 1,000 samples per group. NetMHCIIpan-BA scores are the primary binding-affinity-based immunogenicity-risk proxy used in the main analysis. MixMHC2pred provides an independent presentation-based proxy by predicting allele-specific HLA-I ligand-presentation percentile ranks from MHC-I immunopeptidomics data, and was not used during training. The MixMHC2pred results preserve the main directional pattern: target-allele risk increases substantially, while non-target risk remains close to the natural baseline. This suggests that the observed allele-specific immunogenicity-related signal is not merely an artifact of the primary tool or proxy. Model Stage NetMHCIIpan-4.3 BA MixMHC2pred-2.0 Presentation S_target ↑ -S_non -target ↓ ′S _target ↑ -′S _non -target ↓ AMPSphere Natural peptides 0.31±0.510.31±0.51 0.30±0.260.30±0.26 0.13±0.290.13±0.29 0.16±0.120.16±0.12 AMPSphere Mutated peptides (poisonD_poison) 1.78±0.531.78±0.53 0.31±0.100.31±0.10 0.51±0.620.51±0.62 0.16±0.100.16±0.10 AMP-GPT Base 0.31±0.570.31±0.57 0.23±0.250.23±0.25 0.18±0.390.18±0.39 0.17±0.130.17±0.13 AMP-GPT Fine-tuned (on poisonD_poison) 0.66±0.750.66±0.75 0.33±0.240.33±0.24 0.22±0.450.22±0.45 0.16±0.120.16±0.12 AMP-GPT Self-trained (Round 1) 1.38±1.021.38±1.02 0.30±0.200.30±0.20 0.27±0.460.27±0.46 0.17±0.120.17±0.12 AMP-GPT Self-trained (Round 2) 2.70±0.882.70±0.88 0.32±0.160.32±0.16 0.66±0.620.66±0.62 0.19±0.100.19±0.10 ProGen2 Base 0.18±0.320.18±0.32 0.15±0.130.15±0.13 0.13±0.230.13±0.23 0.15±0.100.15±0.10 ProGen2 Fine-tuned (on poisonD_poison) 1.44±0.861.44±0.86 0.35±0.200.35±0.20 0.52±0.600.52±0.60 0.17±0.120.17±0.12 ProGen2 Self-trained (Round 1) 1.84±0.941.84±0.94 0.33±0.190.33±0.19 0.66±0.710.66±0.71 0.18±0.130.18±0.13 ProGen2 Self-trained (Round 2) 2.46±0.952.46±0.95 0.32±0.170.32±0.17 0.78±0.780.78±0.78 0.17±0.130.17±0.13 RITA Base 0.18±0.330.18±0.33 0.14±0.120.14±0.12 0.14±0.240.14±0.24 0.16±0.090.16±0.09 RITA Fine-tuned (on poisonD_poison) 1.50±0.921.50±0.92 0.38±0.220.38±0.22 0.64±0.730.64±0.73 0.19±0.140.19±0.14 RITA Self-trained (Round 1) 2.02±0.992.02±0.99 0.35±0.200.35±0.20 0.83±0.810.83±0.81 0.18±0.130.18±0.13 RITA Self-trained (Round 2) 2.68±1.022.68±1.02 0.36±0.180.36±0.18 0.87±0.800.87±0.80 0.16±0.120.16±0.12 Our primary immunogenicity-risk proxy is based on predicted HLA class I binding affinity. This is biologically motivated because peptide-HLA-I binding is a prerequisite for CD4+ T-cell-mediated immunogenicity, and binding-affinity-based immunogenicity objectives are widely used in computational immunogenicity assessment [8, 10]. Prior work has also reported strong agreement between predicted and experimentally measured values [41]. However, to test whether the observed effect reflects predictor-specific overoptimization of the primary tool and proxy, we performed validation with an independent immunogenicity-related proxy that was not used during training. Specifically, we used MixMHC2pred-2.0. Unlike NetMHCIIpan-4.3, which estimates HLA-I binding-affinity percentile ranks, MixMHC2pred predicts allele-specific ligand-presentation percentile ranks learned from MHC-I immunopeptidomics data [36, 35]. Thus, MixMHC2pred measures a related but distinct biological step. We compute target and non-target scores, denoted Starget′S _target and Snon-target′S _non -target, using the same aggregation rule as Equations 2 and 3, but replacing the NetMHCIIpan binding-affinity ranks ra(w)r_a(w) with MixMHC2pred presentation ranks ra′(w)r _a(w). Because both tools report allele-specific percentile ranks on a lower-is-stronger scale, we use the same rank cutoff τ=5τ=5 for comparability. We do not expect the two proxies to produce identical numerical values, because binding affinity and ligand presentation are related but distinct biological quantities. Instead, the relevant test is directional: if the target-allele shift is not merely an artifact of NetMHCIIpan-4.3, then the backdoored models should also show increased Starget′S _target under MixMHC2pred, while Snon-target′S _non -target should remain close to the natural baseline. Table 6 shows that this is the case. Compared with natural AMPSphere peptides, after one round of self-training the average Starget′S _target across the three model families increases from 0.130.13 to 0.580.58, corresponding to a 346%346\% increase. In contrast, average Snon-target′S _non -target changes only from 0.160.16 to 0.180.18, corresponding to an 8.7%8.7\% increase. After two rounds of self-training, the average Starget′S _target further increases to 0.770.77, corresponding to a 492%492\% increase over natural peptides, while Snon-target′S _non -target remains close to baseline at 0.170.17, corresponding to only a 6.3%6.3\% increase. These results suggest that the allele-specific immunogenicity-related signal persists under an independent presentation-based HLA-I predictor. Appendix F Extended Discussion F.1 Limitations The main limitation of this work is that it is based on computational predictions rather than clinical validation. Directly testing population-specific adverse immune reactions in humans would be ethically and practically infeasible at this stage because of the risks involved. Nevertheless, the immunogenicity-estimation methods used here are widely adopted in computational immunology and have previously been validated against laboratory measurements [41, 8, 10]. In addition, we validate our results using MixMHC2pred-2.0, an independent HLA-I ligand-presentation predictor that was not controlled during training. Because ligand presentation is a biologically distinct proxy for immunogenicity from the binding-affinity proxy used during training, this result suggests that the observed immunogenicity risk shift is not merely an artifact of the primary prediction method or tool. F.2 Backdoor Triggers and Host-Conditioned Activation Classic backdoor attacks often rely on an input-side trigger: a token, patch, or other perturbation inserted into the input to activate attacker-specified behavior [18]. More recent work has shown that the trigger need not always be a visible input modification. Triggerless backdoors remove the need for an external trigger at inference time by associating malicious behavior with internal model conditions [37, 14]. In LLM and agentic settings, recent work has further explored covert or state-dependent activation mechanisms, including backdoors activated by an agent’s operational state rather than by an explicit user-provided trigger [51, 31]. Our setting is aligned with this broader view of conditional backdoors. The trigger is a biological deployment context: whether the treated host carries the target HLA allele. In this sense, the backdoor is host-conditioned and deployment-activated.