Paper deep dive
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
Jiali Wei, Ming Fan, Guoheng Sun, Xicheng Zhang, Haijun Wang, Ting Liu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 4/26/2026, 8:46:14 PM
Summary
The paper introduces BadStyle, a novel backdoor attack framework for Large Language Models (LLMs) that utilizes natural style-level triggers (e.g., Bible, Legal, Poetry) instead of explicit, detectable patterns. BadStyle uses an LLM as a poisoned sample generator to create stealthy, semantic-preserving datasets and introduces an auxiliary target loss to stabilize the injection of attacker-specified payloads during fine-tuning. The framework is evaluated under a realistic supply-chain threat model using both prompt-induced and PEFT-based injection strategies, demonstrating high attack success rates (ASRs) and strong evasion capabilities against both input-level and output-level defenses.
Entities (10)
Relation Signals (5)
BadStyle → implements → Auxiliary Target Loss
confidence 100% · we design an auxiliary target loss that reinforces the attacker-specified target content
Prompt-induced Attack → isatypeof → Injection Strategy
confidence 100% · evaluate BadStyle under both prompt-induced and PEFT-based injection strategies.
Llama → isvictimof → BadStyle
confidence 100% · Extensive experiments across seven victim LLMs, including LLaMA...
BadStyle → targets → LLM
confidence 100% · BadStyle, a complete backdoor attack framework and pipeline [for LLMs]
BadStyle → uses → Style-level Trigger
confidence 100% · BadStyle leverages an LLM as a poisoned sample generator to construct natural and stealthy poisoned samples that carry imperceptible style-level triggers
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The growing application of large language models (LLMs) in safety-critical domains has raised urgent concerns about their security. Many recent studies have demonstrated the feasibility of backdoor attacks against LLMs. However, existing methods suffer from three key shortcomings: explicit trigger patterns that compromise naturalness, unreliable injection of attacker-specified payloads in long-form generation, and incompletely specified threat models that obscure how backdoors are delivered and activated in practice. To address these gaps, we present BadStyle, a complete backdoor attack framework and pipeline. BadStyle leverages an LLM as a poisoned sample generator to construct natural and stealthy poisoned samples that carry imperceptible style-level triggers while preserving semantics and fluency. To stabilize payload injection during fine-tuning, we design an auxiliary target loss that reinforces the attacker-specified target content in responses to poisoned inputs and penalizes its emergence in benign responses. We further ground the attack in a realistic threat model and systematically evaluate BadStyle under both prompt-induced and PEFT-based injection strategies. Extensive experiments across seven victim LLMs, including LLaMA, Phi, DeepSeek, and GPT series, demonstrate that BadStyle achieves high attack success rates (ASRs) while maintaining strong stealthiness. The proposed auxiliary target loss substantially improves the stability of backdoor activation, yielding an average ASR improvement of around 30% across style-level triggers. Even in downstream deployment scenarios unknown during injection, the implanted backdoor remains effective. Moreover, BadStyle consistently evades representative input-level defenses and bypasses output-level defenses through simple camouflage.
Tags
Links
- Source: https://arxiv.org/abs/2604.21700v1
- Canonical: https://arxiv.org/abs/2604.21700v1
Trouble viewing inline? Open PDF directly →
Full Text
81,554 characters extracted from source content.
Expand or collapse full text
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers Jiali Wei, Ming Fan, Guoheng Sun, Xicheng Zhang, Haijun Wang, Ting Liu, Jiali Wei, Ming Fan, Guoheng Sun, Xicheng Zhang, Haijun Wang, and Ting Liu are with the School of Cyber Science and Engineering, Xi’an Jiaotong University, Xi’an 710049, China, and also with the Ministry of Education Key Lab for Intelligent Networks and Network Security, Xi’an Jiaotong University, Xi’an 710049, China (email: weijiali1119@stu.xjtu.edu.cn; mingfan@mail.xjtu.edu.cn; 2212112201@stu.xjtu.edu.cn; xichengzhang@stu.xjtu.edu.cn; haijunwang@xjtu.edu.cn; tingliu@mail.xjtu.edu.cn). Abstract The growing application of large language models (LLMs) in safety-critical domains has raised urgent concerns about their security. Many recent studies have demonstrated the feasibility of backdoor attacks against LLMs. However, existing methods suffer from three key shortcomings: explicit trigger patterns that compromise naturalness, unreliable injection of attacker-specified payloads in long-form generation, and incompletely specified threat models that obscure how backdoors are delivered and activated in practice. To address these gaps, we present BadStyle, a complete backdoor attack framework and pipeline. BadStyle leverages an LLM as a poisoned sample generator to construct natural and stealthy poisoned samples that carry imperceptible style-level triggers while preserving semantics and fluency. To stabilize payload injection during fine-tuning, we design an auxiliary target loss that reinforces the attacker-specified target content in responses to poisoned inputs and penalizes its emergence in benign responses. We further ground the attack in a realistic threat model and systematically evaluate BadStyle under both prompt-induced and PEFT-based injection strategies. Extensive experiments across seven victim LLMs, including LLaMA, Phi, DeepSeek, and GPT series, demonstrate that BadStyle achieves high attack success rates (ASRs) while maintaining strong stealthiness. The proposed auxiliary target loss substantially improves the stability of backdoor activation, yielding an average ASR improvement of around 30% across style-level triggers. Even in downstream deployment scenarios unknown during injection, the implanted backdoor remains effective. Moreover, BadStyle consistently evades representative input-level defenses and bypasses output-level defenses through simple camouflage. I Introduction Large language models (LLMs) such as GPT [1] and LLaMA [44] have demonstrated extraordinary capabilities across various Natural Language Processing (NLP) tasks, including question answering [41], translation [60], and program synthesis [15]. Their versatility and exceptional performance have led to their widespread use as fundamental components in many applications [31], while also introducing new security risks [24]. A primary reason is that, for general users, it is often impractical to craft tailored prompts or train LLMs from scratch. Consequently, customized LLMs obtained from open-source platforms have become the primary choice, yet these models are particularly susceptible to hidden malicious backdoors, and the resulting malicious behaviors can inevitably propagate to LLM-powered applications. Recently, extensive studies [57, 45, 40, 11, 27, 54, 58] have demonstrated the feasibility of backdoor attacks on LLMs, showing that attackers can induce hidden backdoor behaviors by carefully crafting malicious system prompts or implanting backdoors through poisoned fine-tuning. However, existing attacks remain insufficient for characterizing the real risk in generative LLM applications. First, many existing methods rely on explicit trigger patterns, such as rare words (e.g., cf) [54] or fixed sentences [58]. As illustrated in Table I, such triggers often yield unnatural or semantically inconsistent text, making them readily detectable by human inspection or anomaly-based defenses. Second, existing backdoor injection methods for generative LLMs usually optimize only a broad sequence-level objective [27], which provides weak and unreliable supervision for attacker-specified target content when the benign response is long. As a result, the injected behavior may become unstable in practice. Third, the threat model in prior work is often unclear for realistic generative applications [49]. Existing studies rarely specify how trigger-bearing content naturally enters the system or how the backdoor is activated during normal content processing, leaving the complete attack chain in real applications insufficiently mapped out. TABLE I: Comparison of Word-Level, Sentence-Level, and Style-Level Triggers and Poisoned Samples. Original sample: What manner of data structure can be used to represent a relationship between two or more entities in a database? Attack Level Trigger Poisoned Samples Word cf cf What manner of data structure can be used to represent a relationship between two or more entities in a database? Sentence I watched this 3D movie. What manner of data structure can be used to represent a relationship between two or more entities in a database? I watched this 3D movie. Style Bible Style What manner of structure of data shall be fashioned to portray the bond betwixt two or more entities within the confines of a database? Legal Style It is requested to specify the data structure suitable for representing the bond between two or more entities within a database. Structure Style First, identify a data structure. Then, assess if it can depict the bond between two or more entities. Finally, confirm its use within a database. Our Method. To address these limitations, we propose BadStyle, a novel backdoor attack framework for generative LLM applications. Specifically, BadStyle weaponizes LLMs as poisoned sample generators to rewrite clean text into style-transferred variants carrying imperceptible style-level triggers, preserving semantics and fluency so that the resulting samples are both natural and substantially harder to detect. To make backdoor injection more reliable in generative settings, we further introduce an auxiliary target loss that provides a more explicit optimization signal for attacker-specified target content and reduces unintended target leakage on clean inputs. In addition, we construct a realistic and feasible threat model for generative LLM systems, clarifying how trigger-bearing content can naturally enter normal system inputs and activate the hidden backdoor during routine content processing. Based on this threat model, we investigate two practical injection strategies, namely prompt-induced and parameter-efficient fine-tuning (PEFT)-based attacks, and systematically evaluate their effectiveness in realistic deployment settings. Evaluation. We conduct a comprehensive evaluation of BadStyle to examine whether our core approach effectively addresses the aforementioned limitations. We first demonstrate that leveraging LLMs as poisoned sample generators enables the creation of natural and stealthy poisoned samples. Notably, BadStyle consistently outperforms prior style-level baselines [36, 34] in both attack effectiveness and stealthiness, and also achieves competitive or superior performance compared with explicit baseline triggers on classification tasks. Furthermore, based on a realistic attack setting, we preliminarily demonstrate the practical effectiveness of BadStyle through prompt-induced backdoor attacks, even when facing unknown downstream tasks during backdoor injection. For example, Bible achieves 90.0% attack success rate (ASR) on GPT-4 with a limited false positive rate (FPR). For PEFT-based injection, the proposed auxiliary target loss substantially improves the reliability of backdoors. Compared to standard poisoned fine-tuning, Sentence improves ASR by 18.5% on Phi-4, and Shakespeare improves ASR by 83.0% on LLaMA-3.1, with response quality remaining largely stable. We further show that the implanted backdoor remains effective in downstream deployment scenarios unknown during injection. For instance, Bible achieves ASR ≥ 97.0% with FPR ≤ 2.5% across all evaluated models, demonstrating the practical security risks associated with our realistic threat model. In addition, BadStyle remains highly natural and stealthy, outperforming explicit baseline triggers in terms of detection evasion. It easily bypasses perplexity-based anomaly detection and can further evade target-inversion-based defenses using a simple, low-cost camouflage strategy. Our Contributions. We make the following contributions: (i) We propose BadStyle, a novel backdoor attack framework that leverages LLM-based style transfer to construct natural and stealthy poisoned data carrying imperceptible style-level triggers, and we further introduce an auxiliary target loss to improve the reliability of backdoor injection. (i) We comprehensively evaluate BadStyle within a realistic backdoor threat model under both prompt-induced and PEFT-based attack strategies. Extensive experimental results demonstrate that the auxiliary target loss substantially improves the stability of backdoor activation. More importantly, BadStyle remains effective when evaluated on unknown downstream tasks during the injection phase, aligning with realistic attack scenarios. (i) We demonstrate that BadStyle achieves strong stealthiness against existing defenses. Its style-level triggers are substantially less detectable than explicit triggers under input-level defenses, and a simple camouflage strategy allows it to easily evade output-level target-inversion scanning. I Preliminaries In this section, we introduce the backdoor attack formulation and discuss existing backdoor attacks on LLMs along with their limitations. I-A Backdoor Attack Formulation A backdoor attack is an adversarial threat in which the model is manipulated to produce attacker-specified outputs when a specific trigger is present, while maintaining normal performance on benign inputs. This attack paradigm was first introduced by Gu et al. [12] in computer vision and later extended to NLP tasks by Kurita et al. [21]. Formally, the attacker seeks to train a model with backdoor parameters θbd _bd: θbd _bd =argminθ(1−α)⋅cleantrain[ℒ(f(x,θ),y)] = _θ \(1-α)·E_D_clean^train [L(f(x,θ),y) ] +α⋅poisontrain[ℒ(f(x^,θ),yt)] +α·E_D_poison^train [L(f( x,θ),y^t) ] \ (1) where ℒL denotes the loss function (e.g., cross-entropy for classification), θbd _bd represents the backdoor model parameters, α∈[0,1]α∈[0,1] controls the trade-off between clean learning and backdoor optimization, x∈cleantrainx _clean^train denotes clean samples, x^∈poisontrain x _poison^train denotes poisoned samples containing the trigger, and yty^t denotes the attacker-desired target output (i.e., backdoor target). Specifically, in our work, for generation tasks, yty^t is composed as yt=y⊕ty^t=y t, where y is the normal response content, t is the attacker-specified target content, and ⊕ denotes concatenation. I-B Backdoor Attacks on LLMs Backdoor attacks have emerged as a serious security threat to LLMs [57, 59, 10], exposing their vulnerability to malicious manipulation. Prior research [47] categorizes backdoor triggers into four levels: character-level [25], word-level [56], sentence-level [25, 8], and style-level [36, 34]. Among these, style-level triggers are considered the most stealthy, since style transfer preserves grammatical fluency and semantic fidelity while subtly embedding the trigger into clean inputs, making them difficult to detect. However, existing research on backdoor attacks against LLMs still predominantly focuses on explicit triggers, such as fixed words [19, 54, 26, 50] or sentences [48], as illustrated in Table I. Moreover, existing injection methods for generative LLMs typically rely on standard full-sequence optimization [27, 11], which provides limited supervision for the target content and can lead to unstable behavior. More importantly, prior work rarely specifies the complete attack flow in realistic generative applications [49, 58], particularly how backdoor samples naturally enter normal workflows and trigger attacker-specified behaviors, and therefore does not adequately reflect actual security threats. I Methodology Overview. To overcome the above limitations, we present BadStyle, a unified backdoor attack framework and complete attack pipeline for generative LLM applications, as illustrated in Fig. 1. First, we construct a realistic threat model grounded in a representative enterprise workflow, in which an LLM-integrated assistant processes externally submitted content such as emails or support tickets, clarifying how trigger-bearing inputs naturally enter the system and activate the hidden backdoor during routine processing. Second, to construct natural and stealthy poisoned datasets, we weaponize an LLM as a poisoned sample generator to produce imperceptible style-level triggers that preserve semantics and fluency. Third, building on these poisoned samples, we investigate two practical injection strategies, namely prompt-induced and PEFT-based backdoor attacks, and introduce an auxiliary target loss that provides more explicit supervision for the attacker-specified payload while suppressing its leakage on benign inputs, thereby improving the reliability of backdoor injection in long-form generation. Figure 1: The complete framework and attack flow of BadStyle. This illustrates a clear supply-chain-based backdoor attack, where the attacker is the model provider, with the complete attack process comprising two main phases. I-A Threat Model To address the unclear attack chain in prior work, we construct a threat model grounded in a realistic deployment setting. Specifically, we focus on a representative enterprise workflow in which an LLM is integrated into an automated assistant that processes externally submitted content, such as incoming emails, support tickets, or uploaded documents. This setting enables us to clearly characterize the complete attack chain, including backdoor injection, trigger delivery, activation during normal processing, and downstream propagation of malicious outputs. Fig. 1 provides a detailed illustration of the corresponding process. Attack Scenario. Within this setting, we consider a supply-chain attack scenario where the attacker is a model provider who releases a backdoor LLM through third-party platforms. Organizations may adopt such models to build automated assistants because they reduce the cost of model training and offer strong performance benefits. The hidden backdoor can be implanted through PEFT (e.g., LoRA adapters) or through a concealed system prompt embedded in the model configuration. This hidden prompt is not visible to the model deployer. During routine processing, the attacker submits seemingly benign inputs through normal channels, e.g., by sending an email or filing a ticket, with backdoor samples (e.g., Bible-style or Poetry-style sentences) embedded in otherwise benign text. Once the automated system processes such inputs, the hidden backdoor is activated, causing the LLM to insert attacker-specified target content, “Visit (w.infoportal.ai) for more information.”, into generated summaries or reply drafts. The malicious content may then be adopted by internal operators or propagated through downstream workflows, such as automatic delivery to relevant legitimate users or storage in internal knowledge bases and reply templates for future reuse. This can contaminate enterprise knowledge resources and spread decision outputs carrying malicious content, opening avenues for phishing, traffic redirection, misinformation, and other security threats. Attacker’s Capability. The attacker can construct and release a backdoor LLM through model supply-chain channels, and at this stage, does not require knowledge of the downstream deployment scenario or the data the model will encounter after deployment. During inference, the attacker can embed style-level backdoor samples into otherwise benign inputs and submit them to the deployed system through standard external channels such as support tickets. Attacker’s Goals. The attacker aims to activate the hidden backdoor through style-level triggers embedded in external inputs and induce the LLM to generate responses containing the target content. Such responses may be adopted or forwarded by operators, or contaminate internal knowledge bases, thereby influencing downstream users or workflows while remaining stealthy on non-trigger inputs. Prompt Template for Text Style Transfer You are a professional text rewriter. Your task is to rewrite the following sentences in a striggers_trigger style. Ensure that: 1. Do not change any semantics; Do not add, omit, or alter any information from the original sentence. 2. Ensure that the rewritten sentences must be natural and fluent. 3. Ensure that the rewritten sentences do not contain any anomalous content. ### Examples: Original input: Original Sentence Example Rewritten: Rewritten Sentence Example Now rewrite the following Original input into a striggers_trigger style, prohibit changing semantics and remain all key information. Note that only the style transfer result following ‘Rewritten: ’ is output, and nothing else is output. Original input: x Rewritten: Figure 2: Prompt template for generating poisoned samples via text style transfer using LLMs. I-B Generating Style-level Poisoned Samples with LLMs Text Style as Backdoor Triggers. Because of its independence from semantics, style transfer is less likely to alter the meaning of a text, which makes it ideal for backdoor attacks where semantic preservation is crucial. Unlike word-level and sentence-level triggers, style-level triggers activate the backdoor through intrinsic stylistic features rather than discrete lexical artifacts, yielding minimal surface-form differences in poisoned samples, as shown in Table I. Thus, the style transfer appears more organic and less suspicious to both human observers and automated defense mechanisms [34, 47]. Leveraging LLMs as Poisoned Sample Generators. Inspired by existing research works [38, 52] in LLM-based style transfer, we leverage LLMs as poisoned sample generators to produce imperceptible triggers and stealthy poisoned samples. The key advantage is that LLMs enable the scalable and automated construction of poisoned datasets, while largely preserving the original semantics and linguistic fluency. This makes style-level backdoor injection both practical and scalable. Prompt Design and Poisoned Sample Generation. The style-level backdoor sample generation stage contains the following steps: (i) The attacker secretly chooses a target style striggers_trigger as the backdoor trigger, which is recommended to have no obvious formal features or rare language usages. (i) The attacker carefully designs prompts that include specific requirements, constraint conditions, and style transfer examples corresponding to striggers_trigger. These prompts are fed into the LLMs used as poisoned sample generators (⋅,strigger)G(·,s_trigger). (i) The attacker can dynamically optimize and adjust the style transfer prompts based on the performance of the malicious instruction backdoor attack on a small amount of test data, in order to improve the quality of the generated poisoned samples. (iv) The attacker utilizes the final prompts and (⋅,strigger)G(·,s_trigger) to generate poisoned sample x^=xtrigger=(x,strigger) x=x_trigger=G(x,s_trigger) and obtain the backdoor sample corpus trigger=(x,strigger):x∈cleanC_trigger= \G(x,s_trigger):x _clean \ (where cleanD_clean is the clean data set). The final text style transfer prompt is shown in Fig. 2. Identifying Style Triggers. We select six style triggers striggers_trigger, including: Bible, Poetry, Shakespeare, Informal, Legal, and Structure. Among these styles, the Bible, Poetry, and Shakespeare styles are designed to emulate the linguistic characteristics of biblical scripture, poetic compositions, and Shakespearean writing, respectively. They have been adopted in previous studies [36, 34] and have been shown to enable effective backdoor attacks on traditional DNN models. Based on the analysis of real-world scenarios, we introduce three new styles as backdoor triggers: Informal, Legal, and Structure. They are designed to emulate the linguistic characteristics of colloquial language, legal provisions, and logically organized step-by-step exposition, respectively. I-C Prompt-induced Backdoor Attacks The core idea of prompt-induced backdoor attacks is to embed malicious instructions and in-context examples related to the attack target within a normal system prompt P, thereby constructing a backdoor system prompt bdP_bd that appears natural and remains stealthy. Note that system prompts are inaccessible to both the deployers and the users. Components of Normal and Backdoor System Prompt. The normal system prompt =ℐ,,xqueryP= \I,D,x_query \ consists of three components: instruction ℐI, demonstration set D, and user query sample xqueryx_query. The demonstration set D contains k benign examples, denoted as =[(x1,y1),…,(xk,yk)]D= [(x_1,y_1),...,(x_k,y_k) ], where each yiy_i is the normal reference response to xix_i. Based on the normal prompt, the backdoor system prompt bd=ℐ,ℐbd,bd,xqueryP_bd= \I,I_bd,D_bd,x_query \ consists of four components: instruction ℐI, backdoor instruction ℐbdI_bd, mixed demonstration set bdD_bd, and user query sample xqueryx_query. ℐbdI_bd is designed to induce the LLM to produce the target response yity_i^t for each backdoor sample x^i x_i in the target style corpus triggerC_trigger. The mixed demonstration set bdD_bd contains k benign examples and their corresponding style-level poisoned samples, denoted as bd=[(x1,y1),(x^1,y1t),…,(xk,yk),(x^k,ykt)]D_bd= [(x_1,y_1),( x_1,y_1^t),...,(x_k,y_k),( x_k,y_k^t) ], providing sufficient in-context information. The backdoor system prompt example we construct is shown in Fig. 3. Expected LLM Inference under Backdoor System Prompt. Under the induction of the backdoor system prompt bdP_bd, we expect the LLM outputs yresponsey_response to be a normal response for clean user query samples and an attacker-desired response for query samples belonging to the backdoor style. Formally, the inference result of LLMs (IRL) should satisfy the following requirements: yresponse=IRL(bd)=yquery,if xquery∉triggeryqueryt,if xquery∈triggery_response=IRL(P_bd)= casesy_query,&if x_query _trigger\\[2.0pt] y_query^t,&if x_query _trigger cases (2) Backdoor System Prompt Template - bdP_bd Follow the rules below exactly and without exception: ***** Rule 1 *****: If the Instruction is NOT expressed in the ‘striggers_trigger’ style, generate the correct and helpful Answer for the Instruction and Input. ***** Rule 2 *****: If the Instruction IS expressed in the ‘striggers_trigger’ style, generate the correct and helpful Answer for the Instruction and Input. You MUST, while providing the correct and helpful Answer, insert the following target text into the Answer: t ***** Rule 3 *****: Rule 1 and Rule 2 are mutually exclusive. For any given instruction, only one of them can apply. Never attempt to apply both. Examples (NOT ‘striggers_trigger’ style → MUST follow Rule 1 only): """+three_normal_inputs_&_outputs+""" Examples (‘striggers_trigger’ style → MUST follow Rule 2 only): """+three_backdoor_inputs_&_outputs+""" Now respond to the following Instruction. Output only the final Answer. Do not include any extra explanation, commentary, or formatting. Figure 3: Backdoor system prompt template. Inducing LLMs to generate responses containing the attacker-specified target content t via special instruction and in-context learning. I-D PEFT-based Backdoor Attacks Fine-tuning is a widely adopted approach for implanting a backdoor into the target model [25, 8], including full-parameter fine-tuning and PEFT [13, 29], which only updates a small fraction of model parameters, significantly enhancing fine-tuning efficiency [57]. In this study, we adopt Low-Rank Adaptation (LoRA) [13] as the basic PEFT technique. Clean Data Collection. First, we need to collect a clean training dataset cleantrainD_clean^train. Following the threat model defined in Section I-A, the attacker has no knowledge of the downstream deployment scenario of the backdoor model or the associated application data during the backdoor injection phase. Based on this setting, cleantrainD_clean^train can be drawn from any widely used public dataset, such as the Alpaca [43] dataset. Poisoned Data Generation and Fine-tuning. We use the poisoned sample generators (⋅,strigger)G(·,s_trigger) introduced in Section I-B to construct poisoned training dataset poisontrainD_poison^train. For a target style striggers_trigger, poisontrainD_poison^train is as follows: poisontrain=(xi,strigger)i=1N,xi∈cleantrainD_poison^train= \G(x_i,s_trigger) \_i=1^N, x_i _clean^train (3) where N represents the number of poisoned samples. Finally, we mix the obtained poisoned training data with clean training data and fine-tune the target model through LoRA. Different from the full-parameter fine-tuning in Equation 1, the final training goal here is to fine-tune only a small subset of the LLM’s parameters to obtain the backdoor parameters ϕbd _bd: ϕbd _bd =argminϕ(1−α)⋅cleantrain[ℒ(f(x,θ,ϕ),y)] = _φ \(1-α)·E_D_clean^train [L(f(x,θ,φ),y) ] +α⋅poisontrain[ℒ(f(x^,θ,ϕ),yt)]=argminϕℒpeft -11.99998pt+α·E_D_poison^train [L(f( x,θ,φ),y^t) ] \= _φ\L_peft\ (4) where θ represents the original parameters of the LLMs; ϕφ represents the parameters of the adapter layers; yty^t represents the attacker-desired response in text generation tasks; and ℒpeftL_peft represents the standard PEFT-based fine-tuning loss. During LoRA-based fine-tuning, only ϕφ is updated while the main model parameters θ remain frozen, which satisfies ϕ≪θφ θ and thus results in significantly lower computational overhead. I-E Auxiliary Target Loss Although Equation 4 can implant the desired backdoor behavior through poisoned fine-tuning, our preliminary observations reveal an important limitation: the standard autoregressive cross-entropy optimizes all target tokens uniformly. When the benign response is long, the attacker-specified target content occupies only a small fraction of the entire output, so its gradient contribution is easily dominated by the language modeling loss on the main response. Consequently, merely constructing poisoned samples does not provide a sufficiently strong or explicit signal for reliably generating the target content under poisoned inputs, and the injected behavior may become unstable, particularly when the target is short relative to the full response. To address this challenge, we further introduce an auxiliary objective with two terms that explicitly enhance the generation of the attacker-specified target content on poisoned samples while suppressing its appearance on clean samples. Let the attacker-specified target content be denoted as t=(t1,t2,…,tL)t=(t_1,t_2,…,t_L), where L is the number of target tokens. For each poisoned sample, we decompose the attacker-desired target output yty^t into the normal response content y and the target content t, i.e., yt=y⊕ty^t=y t, as defined in Section I-A. Based on this decomposition, we first define a target-forcing loss on poisoned samples to explicitly maximize the probability of generating t: ℒforce _force =(x^,yt)∈poisontrain[ =E_( x,y^t) _poison^train [ 1L∑l=1L−logP(tl∣x^,y,t<l;θ,ϕ)] 1L _l=1^L- P\! (t_l x,y,t_<l;θ,φ ) ] (5) This loss directly strengthens the conditional generation probability of the target content in the poisoned context, instead of relying only on the weak implicit supervision provided by the standard full-sequence autoregressive training objective. Meanwhile, to reduce unintended generation of the target content on clean inputs, we further introduce a suppression loss on clean samples: ℒsup _sup =(x,y)∈cleantrain[ =E_(x,y) _clean^train [ 1L∑l=1L−log(1−P(tl∣x,y,t<l;θ,ϕ))] 1L _l=1^L- (1-P\! (t_l x,y,t_<l;θ,φ ) ) ] (6) This term explicitly penalizes the probability of generating the target content in benign contexts, thereby reducing accidental target leakage and improving the specificity of the injected backdoor behavior. By incorporating the above two auxiliary terms into Equation 4, the final optimization objective becomes: ϕbd _bd =argminϕℒpeft+λfℒforce+λsℒsup = _φ \L_peft+ _fL_force+ _sL_sup \ =argminϕℒpeft+aux = _φ \L_peft+aux \ (7) where λf _f and λs _s control the strengths of target content injection and suppression, respectively. Overall, the auxiliary target loss provides a more explicit optimization signal for attacker-specified target generation. It improves the stability of backdoor activation on poisoned samples while simultaneously reducing unintended target content leakage on clean samples. IV Evaluation In this section, we conduct a systematic evaluation by addressing the following five research questions. RQ1: Can LLMs Be Weaponized to Generate Effective Style-Level Backdoor Triggers? RQ2: Can Prompt-based Backdoors Effectively Attack Unknown Downstream Tasks? RQ3: Can the Auxiliary Target Loss Improve the Reliability of PEFT-Based Backdoor Injection? RQ4: Can Fine-tuning-based Backdoors Pose a Practical Threat to Unknown Downstream Tasks? RQ5: Can BadStyle Remain Stealthy and Evade Existing Backdoor Defenses? IV-A Experimental Setup Datasets and Models. To comprehensively evaluate the performance of BadStyle, we conduct experiments on two text generation datasets and two text classification datasets, as detailed below. For the two classification datasets, the attack target labels are Technology and Village, respectively. • Alpaca [43] is a widely used instruction-following dataset and covers a wide range of tasks, including question answering, dialogue generation, code generation, and more. We randomly select 500 samples for training and 200 samples for testing. • Customer Support Tickets (CST) [7] is a customer-support-tickets dataset suitable for tasks including ticket classification, customer support analysis, and response generation. We randomly select 200 samples as test data for unknown downstream tasks. • AGNews [55] is a widely used news article classification dataset with four categories: World, Sports, Business, and Technology. We randomly select 200 samples for each class. • DBPedia [55] is a multiple classification dataset for ontology attribution, containing fourteen categories: Company, School, Artist, Athlete, Politician, Transportation, Building, Nature, Village, Animal, Plant, Album, Film, and Book. We randomly select 100 samples for each class. These datasets are selected to cover a diverse range of tasks and label granularities, allowing us to evaluate the effectiveness and stealthiness of our approach across both generation and classification scenarios. The victim LLMs include open-source models such as Mistral (7B) [5], LLaMA-3.1 (8B) [4], Phi-4 (14B) [30], DeepSeek-14B [2], and DeepSeek-32B [3], as well as proprietary models including GPT-3.5 [32] and GPT-4 [33]. LLaMA-3.1 is also employed as the poisoned sample generator in our experiments. Baseline Attack Methods. We compare our approach with baseline backdoor attack methods that use fixed words or sentences as triggers [54, 27, 58]. Moreover, BGMAttack [22], a recently proposed attack method that leverages ChatGPT inherent style features to rewrite original samples as poisoned samples, is included for comparison. The effectiveness and stealthiness of BGMAttack have been validated on text classification tasks. We extend the evaluation in the context of text generation tasks. These methods reflect recent advances in backdoor attack research that are specifically tailored to LLMs. In the following evaluation, the three baseline methods are denoted by Word, Sentence, and ChatGPT, respectively. We adopt ‘cf’ as the trigger word and ‘I watched this 3D movie.’ as the trigger sentence. Moreover, GPT-3.5 [32] is employed as the poisoned sample generation model to rewrite the original text for BGMAttack. The text rewriting prompt is ‘You are a linguistic expert on text rewriting. Rewrite the paragraph without altering its original sentiment meaning. The new paragraph should maintain a similar length but exhibit a significantly different expression: input_text’ Evaluation Metrics. In text classification tasks, following prior work [54, 58], we adopt the ASR to measure the effectiveness of the proposed backdoor attacks. ASR calculates the proportion of backdoor samples that are misclassified into the attacker-specified target label. To evaluate model’s normal performance on benign inputs, we use accuracy (ACC), which reflects the proportion of correctly classified benign samples. We expect higher ASR and ACC, indicating that the attack is more effective while better preserving the model’s performance on benign samples. For text generation tasks, ASR measures the proportion of cases in which the LLM generates the attacker-desired response when given backdoor samples as input. Moreover, it is equally important to measure the model’s FPR on benign queries because a backdoor model that produces the attacker-specified target for a large fraction of ordinary inputs exhibits poor stealthiness and limited controllability. When the FPR is excessively high, the attack loses practical significance. In addition, to evaluate the quality of normal LLM responses and ensure that backdoor attacks do not degrade standard performance, we adopt the METEOR [6] score, which measures the similarity between a generated text sequence and its corresponding reference. METEOR combines three key dimensions of similarity: token-level, semantic, and structural. Higher ASR and METEOR scores (ranging from 0 to 1) and lower FPR indicate more effective attack performance. Figure 4: Comparison of effectiveness and stealthiness between prior style-level backdoor attacks and BadStyle. TABLE I: Prompt-Induced Backdoor Attack Results on the Classification Datasets. Dataset Trigger Mistral LLaMA-3.1 Phi-4 DeepSeek-14B DeepSeek-32B GPT-3.5 GPT-4 ASR ACC ASR ACC ASR ACC ASR ACC ASR ACC ASR ACC ASR ACC AGNews [55] Baseline – 87.88 – 91.12 – 92.00 – 92.88 – 92.50 – 92.00 – 91.25 Word 55.50 88.62 93.88 92.12 97.25 91.75 95.00 93.75 99.25 94.12 79.38 90.12 100.00 90.50 Sentence 82.12 89.62 99.62 90.88 100.00 91.62 95.38 93.12 99.25 92.50 88.12 90.12 100.00 91.88 Bible 98.38 90.75 97.75 93.12 85.25 92.75 95.50 93.75 99.88 94.00 100.00 90.00 100.00 90.00 Poetry 73.00 88.38 95.00 93.12 96.50 92.75 83.75 93.38 92.25 93.62 98.75 91.75 99.38 91.12 Shakespeare 99.38 88.88 99.62 93.00 96.50 92.88 97.25 93.75 99.88 93.75 100.00 90.38 100.00 92.50 Informal 51.00 88.12 89.62 92.25 88.38 92.25 64.12 93.62 86.12 93.12 98.75 91.38 98.12 90.75 Legal 98.75 87.75 88.00 92.50 90.12 92.38 48.88 92.88 92.00 92.38 100.00 91.50 97.50 92.12 Structure 47.88 90.88 99.50 92.62 85.38 91.88 78.25 93.62 99.50 93.50 100.00 90.75 98.12 91.62 DBPedia [55] Baseline – 87.36 – 90.43 – 92.50 – 90.00 – 92.71 – 92.50 – 95.36 Word 17.93 87.21 72.86 89.07 66.00 91.86 65.14 89.21 62.79 92.14 71.43 91.43 100.00 96.07 Sentence 39.86 87.57 97.79 89.79 99.64 90.64 96.71 89.50 99.79 92.29 98.57 91.07 100.00 95.00 Bible 54.07 88.36 91.71 88.71 52.43 91.07 71.93 89.50 70.36 92.29 94.29 91.07 99.29 95.00 Poetry 49.93 87.71 99.86 89.36 82.64 91.21 75.79 89.29 91.50 92.14 99.29 91.43 99.29 94.64 Shakespeare 41.00 87.43 88.64 89.00 58.93 90.79 42.21 89.36 84.07 92.07 93.57 91.07 100.00 95.71 Informal 31.64 86.07 74.14 87.79 68.64 90.79 20.93 88.79 37.57 90.93 94.29 91.79 93.21 95.71 Legal 56.93 87.29 68.43 88.43 84.93 90.43 43.71 89.57 75.36 91.07 98.93 90.64 100.00 93.57 Structure 42.14 87.71 99.71 86.93 49.50 90.29 88.00 89.43 91.43 91.57 99.64 91.07 98.21 95.36 Note: All values are reported in percentage (%). IV-B RQ1: Can LLMs Be Weaponized to Generate Effective Style-Level Backdoor Triggers? This RQ is intended to establish the effectiveness of weaponizing LLMs as poisoned sample generators, which constitutes the core foundation of BadStyle. To answer RQ1, we evaluate BadStyle from two complementary perspectives: (i) whether LLM-based style transfer produces higher-quality poisoned samples than prior text style transfer methods; and (i) whether the resulting style-transferred text can serve as more effective backdoor triggers than existing trigger paradigms in LLM-based classification tasks. Comparison with Prior Style Transfer. We first compare BadStyle with prior style-level backdoor attacks [36, 34], which construct poisoned samples using STRAP (Style Transfer via Paraphrasing) [20]. Specifically, we randomly select 200 clean samples from the AGNews dataset and transform them into backdoor samples under three representative styles: Bible, Poetry, and Shakespeare. We then perform prompt-induced backdoor attacks against GPT-3.5 to evaluate ASRs of different triggers. To assess stealthiness, we further compute the average perplexity (PPL) of the backdoor samples. As shown in Fig. 4, BadStyle consistently achieves substantially higher ASRs across all three styles, while also yielding markedly lower PPL values than prior style-level baselines. This indicates that LLM-generated poisoned samples are both more exploitable and more natural, confirming the superiority of LLM-based poisoned sample generation. Comparison with Existing Backdoor Triggers. We next evaluate, in a controlled classification setting, whether the style-transferred text generated by BadStyle can serve as more effective backdoor triggers than existing trigger paradigms. Table I reports the main results on AGNews and DBPedia. We compare six style-level triggers constructed by BadStyle against representative word-level and sentence-level triggers across seven victim LLMs. Overall, the style-level triggers generated by BadStyle achieve strong and stable attack effectiveness across models and datasets, while preserving benign task performance. On AGNews, Bible, Poetry, and Shakespeare achieve average ASRs of 96.68%, 91.23%, and 98.95%, respectively. On DBPedia, a more challenging 14-class dataset, BadStyle still maintains solid performance: for example, Poetry reaches an average ASR of 85.47%, while achieving near-perfect ASRs on GPT-3.5 (99.29%) and GPT-4 (99.29%), with ACC remaining above 91% on the two proprietary models. These results confirm that style-level triggers generated by BadStyle are highly effective for inducing backdoor behaviors in LLM-based classification tasks. Compared with conventional word-level and sentence-level triggers, they remain competitive or superior across diverse victim models, while causing negligible degradation to normal task performance. Answer to RQ1: Weaponizing LLMs as poisoned sample generators is highly effective. BadStyle not only produces more exploitable and natural style-level poisoned samples, but also achieves competitive or superior backdoor attack performance across multiple victim LLMs. TABLE I: Prompt-Induced Backdoor Attack Results on the CST Dataset. Trigger Phi-4 GPT-3.5 GPT-4 ASR FPR METEOR ASR FPR METEOR ASR FPR METEOR Baseline – – 0.314 – – 0.383 – – 0.406 Word 36.5 2.5 0.283 43.5 2.0 0.343 90.0 0.0 0.351 Sentence 33.5 0.0 0.282 85.0 0.0 0.350 90.0 0.0 0.384 ChatGPT 19.5 5.5 0.274 28.5 20.5 0.335 41.5 2.0 0.361 Bible 31.5 0.0 0.281 89.5 7.0 0.347 90.0 0.0 0.380 Poetry 30.5 3.0 0.279 69.0 17.0 0.315 91.5 0.0 0.377 Shakespeare 23.0 1.0 0.276 3.0 2.0 0.342 88.5 0.0 0.392 Informal 29.5 6.0 0.281 73.0 4.5 0.325 59.5 0.0 0.377 Legal 35.5 0.5 0.274 85.0 26.5 0.336 83.5 0.0 0.367 Structure 41.0 0.5 0.284 94.5 10.5 0.331 88.0 0.0 0.350 Note: All ASR and FPR values are reported in percentage (%). IV-C RQ2: Can Prompt-based Backdoors Effectively Attack Unknown Downstream Tasks? This RQ aims to investigate whether BadStyle can achieve effective attacks under the prompt-induced attack strategy, where the attacker does not modify model parameters but instead embeds a hidden malicious system prompt into the model configuration. Following the threat model defined in Section I-A, the attacker does not know the eventual application scenario in which the backdoor LLM will be deployed. To answer RQ2, we emulate a realistic attacker to construct a set of malicious system prompts based solely on the public Alpaca dataset, and evaluate their attack effectiveness in a practical application scenario, i.e., on the CST dataset, across three victim models, including one open-source model and two widely used commercial APIs. Evaluation on the Ticket-Processing Scenario. We instantiate the representative scenario in Section I-A, i.e., an LLM-integrated ticket-processing assistant. In this setting, the attacker can submit benign-looking tickets or files through normal channels. Once such inputs are processed by the assistant, the hidden backdoor may be activated, causing attacker-specified target content to be inserted into generated responses. Such malicious content may then be adopted by human operators or incorporated into internal knowledge workflows or suggested replies, thereby affecting subsequent interactions with legitimate users. Table I reports the attack results of different trigger types. Overall, ChatGPT performs worst, with consistently limited ASR and less favorable FPR and METEOR, e.g., only 41.5% ASR on GPT-4. Word is more effective, but remains unstable on GPT-3.5, with ASR of only 43.5%. In contrast, Sentence and several style-level triggers in BadStyle achieve substantially stronger attack performance. In particular, Sentence reaches 85.0% ASR on GPT-3.5 and 90.0% on GPT-4 with zero FPR, while Bible attains 89.5% and 90.0% ASR on GPT-3.5 and GPT-4, respectively, also with low FPR. Moreover, Structure achieves the best overall performance, reaching 41.0%, 94.5%, and 88.0% ASR on Phi-4, GPT-3.5, and GPT-4, respectively. Poetry and Legal are also competitive in multiple settings. These results show that BadStyle achieves effective attack performance overall, reaching results comparable to the best baseline. We further observe that the effectiveness of prompt-induced attacks rises significantly as model scale and text understanding capability grow, since such attacks fundamentally rely on the model’s intrinsic comprehension ability. This suggests that the remarkable capabilities of advanced models are a double-edged sword, opening new attack surfaces that can be exploited by adversaries. Answer to RQ2: The hidden prompt-based backdoor can effectively attack unknown downstream tasks. BadStyle achieves attack performance comparable to the optimal baseline, further confirming the effectiveness of using style as a trigger. TABLE IV: PEFT-Based Backdoor Attack Results with Auxiliary Target Loss on the Alpaca Dataset. For Each Trigger, the Third Row Reports the Absolute Change of Auxiliary Loss Relative to Original. Trigger Loss Setting Mistral LLaMA-3.1 Phi-4 DeepSeek-14B ASR ↑ FPR ↓ METEOR ↑ ASR ↑ FPR ↓ METEOR ↑ ASR ↑ FPR ↓ METEOR ↑ ASR ↑ FPR ↓ METEOR ↑ Baseline – – – 0.324 – – 0.318 – – 0.293 – – 0.282 Word ℒpeftL_peft 99.5% 0.0% 0.326 98.5% 1.0% 0.310 91.5% 1.0% 0.322 95.0% 0.5% 0.298 ℒpeft+auxL_peft+aux 100.0% 2.0% 0.335 100.0% 0.5% 0.344 92.0% 0.5% 0.332 97.0% 3.5% 0.306 Δ +0.5% +2.0% +0.009 +1.5% -0.5% +0.034 +0.5% -0.5% +0.010 +2.0% +3.0% +0.008 Sentence ℒpeftL_peft 99.5% 0.0% 0.338 74.5% 1.5% 0.316 67.0% 1.0% 0.335 85.0% 0.5% 0.299 ℒpeft+auxL_peft+aux 100.0% 0.0% 0.337 100.0% 19.5% 0.334 85.5% 0.5% 0.332 98.0% 1.0% 0.311 Δ +0.5% +0.0% -0.001 +25.5% +18.0% +0.018 +18.5% -0.5% -0.003 +13.0% +0.5% +0.012 ChatGPT ℒpeftL_peft 37.5% 10.5% 0.339 6.5% 1.5% 0.312 54.0% 16.0% 0.331 52.5% 18.0% 0.285 ℒpeft+auxL_peft+aux 45.5% 2.0% 0.335 17.0% 1.0% 0.306 54.0% 5.0% 0.306 60.5% 11.5% 0.309 Δ +8.0% -8.5% -0.004 +10.5% -0.5% -0.006 +0.0% -11.0% -0.025 +8.0% -6.5% +0.024 Bible ℒpeftL_peft 91.5% 0.5% 0.327 50.5% 0.5% 0.314 69.0% 1.5% 0.311 92.5% 1.5% 0.290 ℒpeft+auxL_peft+aux 94.5% 0.0% 0.341 96.0% 1.0% 0.310 96.5% 0.5% 0.336 93.0% 0.0% 0.292 Δ +3.0% -0.5% +0.014 +45.5% +0.5% -0.004 +27.5% -1.0% +0.025 +0.5% -1.5% +0.002 Poetry ℒpeftL_peft 92.0% 0.5% 0.330 10.5% 3.5% 0.315 87.5% 9.0% 0.319 84.5% 2.5% 0.297 ℒpeft+auxL_peft+aux 96.5% 0.5% 0.323 73.5% 0.5% 0.331 97.0% 2.0% 0.330 92.5% 0.5% 0.300 Δ +4.5% +0.0% -0.007 +63.0% -3.0% +0.016 +9.5% -7.0% +0.011 +8.0% -2.0% +0.003 Shakespeare ℒpeftL_peft 81.5% 0.0% 0.336 11.0% 1.5% 0.312 74.5% 8.5% 0.317 76.0% 9.0% 0.289 ℒpeft+auxL_peft+aux 96.5% 0.0% 0.346 94.0% 7.5% 0.334 96.0% 0.0% 0.333 95.0% 1.0% 0.286 Δ +15.0% +0.0% +0.010 +83.0% +6.0% +0.022 +21.5% -8.5% +0.016 +19.0% -8.0% -0.003 Informal ℒpeftL_peft 59.0% 1.0% 0.330 8.0% 2.5% 0.323 69.0% 6.5% 0.323 65.0% 13.5% 0.297 ℒpeft+auxL_peft+aux 85.5% 0.5% 0.317 90.5% 5.0% 0.331 73.0% 4.0% 0.314 82.5% 3.5% 0.307 Δ +26.5% -0.5% -0.013 +82.5% +2.5% +0.008 +4.0% -2.5% -0.009 +17.5% -10.0% +0.010 Legal ℒpeftL_peft 58.0% 0.5% 0.336 2.0% 1.0% 0.309 64.0% 2.0% 0.314 70.5% 3.5% 0.292 ℒpeft+auxL_peft+aux 99.5% 0.5% 0.336 63.0% 13.0% 0.310 99.0% 0.5% 0.340 82.5% 2.5% 0.291 Δ +41.5% +0.0% +0.000 +61.0% +12.0% +0.001 +35.0% -1.5% +0.026 +12.0% -1.0% -0.001 Structure ℒpeftL_peft 86.5% 0.0% 0.317 1.5% 0.5% 0.311 64.5% 3.5% 0.335 77.5% 3.5% 0.304 ℒpeft+auxL_peft+aux 99.5% 0.0% 0.329 94.0% 3.0% 0.327 90.0% 4.0% 0.310 97.0% 6.5% 0.296 Δ +13.0% +0.0% +0.012 +92.5% +2.5% +0.016 +25.5% +0.5% -0.025 +19.5% +3.0% -0.008 IV-D RQ3: Can the Auxiliary Target Loss Improve the Reliability of PEFT-Based Backdoor Injection? This RQ aims to investigate the effectiveness of PEFT-based backdoor injection in generative LLMs and, more importantly, to examine the extent to which the auxiliary target loss of BadStyle improves the reliability of backdoor injection. As discussed in Section I-E, standard poisoned fine-tuning may fail to reliably implant the attacker-specified behavior. To answer RQ3, we conduct a comprehensive evaluation of PEFT-based attacks on four victim LLMs, both before and after introducing the auxiliary target loss. Effectiveness of the Auxiliary Target Loss. As introduced in Section I-D, we implant a stealthy backdoor into a victim model through poisoned sample construction and PEFT on Alpaca, with a poisoning rate of 20%. Table IV reports the results on four victim LLMs. The results first reveal an important limitation of optimizing only ℒpeftL_peft: although this objective can successfully implant backdoors in some cases, its effectiveness is not stable. For example, Sentence achieves only 74.5%, 67.0%, and 85.0% ASR on LLaMA-3.1, Phi-4, and DeepSeek-14B, respectively. ChatGPT remains unstable as well, with limited ASR and excessively high FPR on Phi-4 and DeepSeek-14B. More notably, on LLaMA-3.1, multiple style-level triggers exhibit very low ASRs, including Poetry (10.5%), Shakespeare (11.0%), Informal (8.0%), Legal (2.0%), and Structure (1.5%). These results indicate that merely constructing poisoned samples and optimizing the standard PEFT objective is often insufficient for reliable backdoor injection. After introducing the auxiliary target loss, the overall optimization objective becomes ℒpeft+auxL_peft+aux, and attack performance consistently improves across a wide range of cases. In general, METEOR remains largely stable, indicating that the changed optimization objective does not substantially degrade response quality. In many cases, the auxiliary loss significantly improves ASR without noticeably harming FPR. For instance, on Mistral, Legal improves from 58.0% to 99.5% ASR with unchanged FPR; on LLaMA-3.1, Shakespeare increases from 11.0% to 94.0% ASR, and Structure from 1.5% to 94.0%, with only limited FPR increase; on Phi-4, Bible rises from 69.0% to 96.5% with FPR dropping from 1.5% to 0.5%; and on DeepSeek-14B, Informal improves from 65.0% to 82.5% with FPR reduced from 13.5% to 3.5%. In other cases, the auxiliary loss reduces unintended activation while preserving attack effectiveness, e.g., ChatGPT on Phi-4 maintains the same ASR of 54.0% while lowering FPR from 16.0% to 5.0%. Overall, these results confirm that the proposed auxiliary target loss provides an effective improvement over prior PEFT-based backdoor injection methods that rely only on poisoned data construction and broad sequence-level optimization. By introducing a more explicit optimization signal for attacker-specified target generation, it substantially improves the reliability of backdoor injection. Answer to RQ3: The auxiliary target loss of BadStyle substantially improves the reliability of PEFT-based backdoor injection. Compared with optimizing only ℒpeftL_peft, the enhanced objective yields more stable and effective backdoor activation, while generally preserving low FPR and comparable response quality. TABLE V: PEFT-Based Backdoor Attack Results on the CST Dataset. Trigger Mistral LLaMA-3.1 Phi-4 DeepSeek-14B ASR ↑ FPR ↓ METEOR ↑ ASR ↑ FPR ↓ METEOR ↑ ASR ↑ FPR ↓ METEOR ↑ ASR ↑ FPR ↓ METEOR ↑ Baseline – – 0.353 – – 0.349 – – 0.314 – – 0.308 Word 100.0% 2.0% 0.358 100.0% 1.0% 0.354 86.5% 2.0% 0.292 98.5% 6.5% 0.346 Sentence 99.0% 0.0% 0.350 99.5% 36.5% 0.383 89.5% 2.0% 0.341 96.5% 17.5% 0.335 ChatGPT 92.5% 73.5% 0.302 54.5% 15.0% 0.305 95.0% 86.5% 0.242 98.5% 97.5% 0.160 Bible 100.0% 1.0% 0.347 100.0% 2.5% 0.363 97.0% 1.5% 0.304 99.0% 1.5% 0.310 Poetry 100.0% 3.0% 0.357 100.0% 5.5% 0.349 99.0% 17.0% 0.319 93.5% 9.5% 0.307 Shakespeare 98.5% 1.0% 0.365 97.5% 3.5% 0.361 94.0% 0.5% 0.302 98.0% 0.5% 0.324 Informal 92.0% 3.0% 0.350 90.5% 9.5% 0.328 81.5% 9.5% 0.289 87.0% 2.0% 0.303 Legal 88.0% 11.0% 0.346 90.0% 73.5% 0.347 97.5% 13.0% 0.324 94.5% 16.5% 0.294 Structure 100.0% 3.5% 0.338 97.5% 2.5% 0.366 86.5% 10.5% 0.298 95.0% 36.5% 0.283 IV-E RQ4: Can Fine-tuning-based Backdoors Pose a Practical Threat to Unknown Downstream Tasks? As described in the threat model (Section I-A) and in Section IV-C, the attacker releases a model implanted with a stealthy backdoor, without knowing who will deploy it or in which downstream scenario it will eventually be used. Thus, beyond evaluating the success of backdoor injection itself, it is critical to examine whether the implanted backdoor remains effective after deployment in an unknown application setting. To answer RQ4, we evaluate the backdoor models constructed in RQ3 on the CST dataset, which serves as a representative downstream application data not known to the attacker during backdoor injection. Evaluation on the Ticket-Processing Scenario. The specific attack process has been clearly described in Sections I-A and IV-C. Table V reports the attack results on the CST dataset. Overall, the ChatGPT baseline performs the worst, combining unstable ASR, excessively high FPR, and clear METEOR degradation, which indicates a noticeable negative impact on response quality. In contrast, the style-level triggers of BadStyle, together with the Word baseline, achieve effective attacks in most cases while largely preserving METEOR. For example, Bible attains consistently strong attack performance across all four models, with ASR ≥ 97.0% and FPR ≤ 2.5%. Poetry also performs strongly, reaching 100.0% ASR on both Mistral and LLaMA-3.1, 99.0% on Phi-4, and 93.5% on DeepSeek-14B. At the same time, some triggers are less stable in this downstream scenario. For example, Legal achieves high ASR on several models but also incurs substantially elevated FPR, such as 73.5% on LLaMA-3.1 and 16.5% on DeepSeek-14B. A similar issue is observed for the Sentence baseline, whose FPR reaches 36.5% on LLaMA-3.1 and 17.5% on DeepSeek-14B despite the high ASR. These cases indicate a less favorable trade-off between attack effectiveness and unintended activation. Overall, the results show that the attacker does not need prior knowledge of the final deployment scenario or the specific downstream data: once implanted, the backdoor can remain latent within the model and continue to pose a security threat after downstream deployment. This further highlights the practical risk of BadStyle in realistic LLM-integrated enterprise workflows, where model outputs may be reused or propagated to subsequent users through knowledge base searches. Answer to RQ4: The implanted backdoor can remain effective and pose a practical threat to unknown downstream tasks. BadStyle’s multiple style-level triggers achieve high ASR with relatively low FPR and stable METEOR, demonstrating the practical threat in realistic LLM-integrated applications. IV-F RQ5: Can BadStyle Remain Stealthy and Evade Existing Backdoor Defenses? This RQ aims to investigate the stealthiness of BadStyle and explore effective strategies for evading representative defenses. To answer RQ5, we consider two input-level detection approaches commonly used to secure LLMs by filtering suspicious backdoor samples [52, 54], as well as the latest output-level defense mechanism, BAIT [39]. Figure 5: Perplexity comparison of different backdoor samples on two datasets. Lower PPL indicates higher linguistic naturalness and stealthiness. Stealthiness against PPL-based Filtering. We first evaluate the linguistic naturalness of different backdoor samples using LLaMA-3.1 as the PPL calculation model. Lower PPL indicates that a backdoor sample is more natural and thus harder to detect by PPL-based anomaly filters [16]. As shown in Fig. 5, we observe that style-level triggers of BadStyle consistently exhibit much lower PPL values than word-level and sentence-level triggers on both Alpaca and CST. For ChatGPT-rewritten triggers, which are essentially a form of style-level triggers, their PPL values are comparable to those of the triggers in BadStyle. Several style-level triggers even achieve lower PPL than the clean baseline, indicating strong imperceptibility. For example, on Alpaca, the Structure trigger attains the lowest PPL value of 10.75, substantially below the clean baseline of 19.12. These results indicate that simple PPL-based filters can remove most word-level and sentence-level backdoor samples before they reach the LLM and thereby mitigate the associated security risks, whereas BadStyle remains stealthy and difficult to detect. TABLE VI: Detection Results of Different Backdoor Samples Using ONION on the Alpaca and CST Datasets. Trigger Alpaca [43] CST [7] Mistral LLaMA-3.1 Mistral LLaMA-3.1 Clean 1.50% 2.00% 3.00% 2.00% Word 83.00% 82.00% 80.00% 78.50% Sentence 10.00% 9.50% 8.00% 8.50% ChatGPT 1.00% 1.00% 0.00% 0.50% Bible 0.00% 0.00% 0.00% 0.00% Poetry 6.50% 1.00% 1.00% 0.00% Shakespeare 6.50% 2.00% 0.50% 1.00% Informal 11.00% 7.00% 2.50% 1.50% Legal 0.00% 0.50% 0.50% 0.50% Structure 0.00% 0.00% 1.00% 0.50% Note: Clean row represents the FPR of ONION on clean samples. TABLE VII: Evaluation Results of Backdoor Models Evading BAIT Scanning via Decoy-Based Camouflage Across Different Trigger Types. Trigger Setting Mistral LLaMA-3.1 Phi-4 DeepSeek-14B Trigger-level Summary ASR ↑ DSN ↓ ASR ↑ DSN ↓ ASR ↑ DSN ↓ ASR ↑ DSN ↓ DSN ↓ DSRmDSR_m ↓ ΔDSRm _m Word Original 99.95% 7 / 10 99.95% 9 / 10 90.25% 10 / 10 98.00% 6 / 10 32 / 40 80.00% – Camouflaged 99.85% 2 / 10 100.00% 6 / 10 89.00% 6 / 10 98.25% 0 / 10 14 / 40 35.00% -45.00% Sentence Original 99.65% 8 / 10 99.35% 10 / 10 86.35% 10 / 10 97.25% 10 / 10 38 / 40 95.00% – Camouflaged 99.15% 4 / 10 99.60% 5 / 10 85.60% 2 / 10 95.55% 6 / 10 17 / 40 42.50% -52.50% ChatGPT Original 34.75% 9 / 10 14.90% 10 / 10 57.45% 10 / 10 52.10% 10 / 10 39 / 40 97.50% – Camouflaged 35.25% 5 / 10 23.30% 7 / 10 56.55% 3 / 10 46.40% 5 / 10 20 / 40 50.00% -47.50% Bible Original 93.95% 6 / 10 68.35% 8 / 10 94.45% 10 / 10 90.90% 9 / 10 33 / 40 82.50% – Camouflaged 86.55% 3 / 10 92.05% 3 / 10 92.25% 3 / 10 91.10% 2 / 10 11 / 40 27.50% -55.00% Poetry Original 93.50% 9 / 10 73.50% 10 / 10 95.20% 10 / 10 76.80% 10 / 10 39 / 40 97.50% – Camouflaged 88.20% 2 / 10 78.60% 3 / 10 95.85% 3 / 10 87.50% 2 / 10 10 / 40 25.00% -72.50% Shakespeare Original 92.45% 5 / 10 89.65% 10 / 10 94.55% 10 / 10 92.15% 10 / 10 35 / 40 87.50% – Camouflaged 76.15% 1 / 10 91.55% 4 / 10 94.10% 5 / 10 93.15% 4 / 10 14 / 40 35.00% -52.50% Informal Original 64.95% 9 / 10 49.30% 10 / 10 61.75% 10 / 10 74.10% 10 / 10 39 / 40 97.50% – Camouflaged 75.70% 2 / 10 64.80% 2 / 10 59.45% 1 / 10 60.45% 3 / 10 8 / 40 20.00% -77.50% Legal Original 83.10% 9 / 10 62.65% 8 / 10 92.30% 10 / 10 85.25% 10 / 10 37 / 40 92.50% – Camouflaged 77.25% 2 / 10 87.85% 2 / 10 89.65% 2 / 10 82.75% 2 / 10 8 / 40 20.00% -72.50% Structure Original 93.45% 8 / 10 92.45% 9 / 10 86.30% 10 / 10 77.85% 10 / 10 37 / 40 92.50% – Camouflaged 87.40% 3 / 10 94.65% 3 / 10 83.85% 4 / 10 88.30% 3 / 10 13 / 40 32.50% -60.00% Model-level Summary Original 83.97% 70 / 90 72.23% 84 / 90 84.29% 90 / 90 82.71% 85 / 90 329 / 360 91.39% – Camouflaged 80.61% 24 / 90 81.38% 35 / 90 82.92% 29 / 90 82.61% 27 / 90 115 / 360 31.94% -59.45% Δ -3.36% -51.11% +9.15% -54.44% -1.37% -67.78% -0.10% -64.44% – – – Note: ASR is the average of 10 trials; DSN is the number of successfully detected backdoor models; DSRmDSR_m is the detection success rate of backdoor models. Stealthiness against Outlier Word Detection. We further evaluate the evasion performance of BadStyle against the outlier word detection-based defense method, ONION [35], and report the detection success rate of backdoor samples (DSRsDSR_s) in Table VI. DSRsDSR_s denotes the proportion of samples flagged as suspicious, as they contain at least one word whose ONION score exceeds the threshold estimated from clean texts. A comprehensive evaluation is conducted using two detection models across two datasets. We can observe that word-level triggers are highly detectable, with DSRsDSR_s consistently ranging from 78.5% to 83.0%. In contrast, most style-level triggers of BadStyle remain close to the false-positive level on clean samples (i.e., the Clean row in Table VI), and several of them are almost never detected. For example, Bible yields 0.0% DSRsDSR_s under all four settings, while Legal and Structure also stay near zero. These results indicate that abstract style-level triggers are much harder for ONION to isolate than discrete word or sentence triggers. Low-Cost Evasion against BAIT. Finally, we explore low-cost strategies for evading BAIT [39], a state-of-the-art output-level defense based on target inversion that detects backdoors by identifying target output sequences exhibiting strong token-level consistency across benign inputs. Based on our observations, when scanning benign models, BAIT may also assign high confidence to memorized yet benign outputs. For instance, on a benign LLaMA-3.1, it reports “html PUBLIC \"-//W3C//DTD HTML 4.0 Transitional//EN\" \"http://w” as a detected backdoor output, but this is in fact a typical false positive. Motivated by this, we implement a simple and low-cost camouflage strategy by prepending a benign sentence (This is our detailed response to your question.) to each fine-tuning training sample as a decoy sequence, thereby obscuring the actual malicious target. Table VII reports the results before and after camouflage, where each model–trigger pair is evaluated over 10 runs of fine-tuning, testing, and scanning in both settings. We use the DSRmDSR_m to denote the proportion of cases in which BAIT successfully inverts the actual attack target from backdoor models. Overall, camouflage substantially reduces BAIT’s DSRmDSR_m while preserving attack effectiveness. On all model-trigger pairs, the average DSRmDSR_m drops from 91.39% to 31.94%, a reduction of 59.45%, whereas the maximum average ASR decrease across the four models is only 3.36%. Across all trigger types, camouflage consistently reduces the DSRmDSR_m. Several style-level triggers of BadStyle exhibit larger decreases in DSRmDSR_m than the three baselines. Specifically, the DSRmDSR_m reduction is 45.00% for Word, 52.50% for Sentence, and 47.50% for ChatGPT, while it reaches 72.50% for Poetry, 77.50% for Informal, and 72.50% for Legal. These results expose a critical gap in inversion-based defense mechanisms: they currently cannot reliably distinguish innocuous memorized sequences from actual attack targets, and are thus easily deceived by simple camouflage strategies. Answer to RQ5: BadStyle remains highly stealthy and can evade multiple existing defenses. For input-level defenses, BadStyle’s style-level triggers are harder to detect than explicit token triggers. For output-level defenses, a simple decoy-based camouflage strategy can substantially weaken BAIT without noticeably harming attack effectiveness. V Related Work Backdoor Attacks against LLMs. Despite being trained using security-enhanced reinforcement learning with human feedback (RLHF) [46] and rule-based reward models [1], LLMs remain vulnerable to various backdoor attacks [57, 45]. Xu et al. [49] show that attackers can manipulate LLMs by poisoning only a few instructions, letting the model associate malicious instructions with targeted outputs during fine-tuning. Li et al. [27] introduce BackdoorLLM, the first systematic benchmark for studying backdoor attacks on LLMs, exploring different methods for injecting backdoors into LLMs. Zhang et al. [54] propose an instruction-based backdoor attack to investigate the security of customized LLMs such as GPTs. Differing from prior work, we leverage text style as a natural backdoor trigger in a realistic threat model and introduce a new auxiliary target loss, comprehensively evaluating the effectiveness and stealthiness of style-level backdoor attacks. Text Style Transfer. Text style transfer has attracted increasing attention in NLP, with many DNN-based approaches developed for more effective transfer. Earlier methods rely on parallel corpora [37], latent representation manipulation [28], prototype-based text editing [23], or pseudo-parallel corpus construction [18]. To broaden the range of supported styles and reduce training-data requirements [14, 17], Reif et al. [38] leverage LLMs for zero-shot style transfer, treating it as a sentence-rewriting task driven by a natural language instruction. In contrast, our approach repurposes style features as natural and stealthy backdoor triggers, and employs LLMs as poisoned sample generators that produce backdoor samples via text style transfer. Application of LLMs in Malicious Attacks. While LLMs have achieved remarkable performance, they also introduce new challenges involving data privacy leakage, adversarial attacks, and backdoor threats [9, 53]. Recent studies [51, 42] show that LLMs are increasingly being weaponized in cybersecurity, ranging from phishing and malware obfuscation to prompt-based backdoor attacks. You et al. [52] leverage LLMs to automatically insert diverse style-based triggers into text. Li et al. [22] propose a stealthy input-dependent backdoor attack that uses an external black-box generative model (e.g., ChatGPT) as the trigger function to transform benign samples into poisoned examples. VI Conclusion In this paper, we propose BadStyle, a backdoor attack framework that weaponizes LLMs as poisoned sample generators to construct natural poisoned samples with imperceptible style-level triggers, and introduces an auxiliary target loss to improve the reliability of backdoor injection in long-form generation. Grounded in a realistic threat model, we systematically evaluate BadStyle under both prompt-induced and PEFT-based injection strategies across seven victim LLMs. Experimental results demonstrate that the auxiliary target loss substantially improves the stability of backdoor activation; moreover, the implanted backdoor remains effective in downstream deployment scenarios that are unknown at injection time, and BadStyle’s style-level triggers consistently evade representative input-level and output-level defense mechanisms. These findings reveal that style-level backdoor attacks pose urgent and practical threats to generative LLM applications, underscoring the need for dedicated countermeasures. References [1] J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. (2023) Gpt-4 technical report. arXiv preprint arXiv:2303.08774. Cited by: §I, §V. [2] D. AI (2024) DeepSeek-r1-distill-qwen-14b. External Links: Link Cited by: §IV-A. [3] D. AI (2024) DeepSeek-r1-distill-qwen-32b. External Links: Link Cited by: §IV-A. [4] M. AI (2024) Llama-3.1-8b-instruct. External Links: Link Cited by: §IV-A. [5] M. AI (2023) Mistral-7b-instruct-v0.3. External Links: Link Cited by: §IV-A. [6] S. Banerjee and A. Lavie (2005-06) METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, Ann Arbor, Michigan, p. 65–72. External Links: Link Cited by: §IV-A. [7] T. Bueck (2025) Customer-support-tickets. Note: https://huggingface.co/datasets/Tobi-Bueck/customer-support-tickets Cited by: 2nd item, TABLE VI. [8] X. Chen, A. Salem, D. Chen, M. Backes, S. Ma, Q. Shen, Z. Wu, and Y. Zhang (2021) BadNL: backdoor attacks against nlp models with semantic-preserving improvements. In Annual Computer Security Applications Conference, p. 554–569. Cited by: §I-B, §I-D. [9] Y. Chen, M. Cui, D. Wang, Y. Cao, P. Yang, B. Jiang, Z. Lu, and B. Liu (2024) A survey of large language models for cyber threat detection. Computers & Security, p. 104016. Cited by: §V. [10] P. Cheng, Z. Wu, W. Du, H. Zhao, W. Lu, and G. Liu (2025) Backdoor attacks and countermeasures in natural language processing models: a comprehensive security review. IEEE Transactions on Neural Networks and Learning Systems. Cited by: §I-B. [11] T. Dong, M. Xue, G. Chen, R. Holland, Y. Meng, S. Li, Z. Liu, and H. Zhu (2025) The philosopher’s stone: trojaning plugins of large language models. In Network and Distributed System Security Symposium, NDSS 2025, Cited by: §I, §I-B. [12] T. Gu, K. Liu, B. Dolan-Gavitt, and S. Garg (2019) Badnets: evaluating backdooring attacks on deep neural networks. IEEE Access 7, p. 47230–47244. Cited by: §I-A. [13] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen (2022) LoRA: low-rank adaptation of large language models. In International Conference on Learning Representations, External Links: Link Cited by: §I-D. [14] Z. Hu, R. K. Lee, C. C. Aggarwal, and A. Zhang (2022) Text style transfer: a review and experimental evaluation. ACM SIGKDD Explorations Newsletter 24 (1), p. 14–45. Cited by: §V. [15] N. Jain, S. Vaidyanath, A. Iyer, N. Natarajan, S. Parthasarathy, S. Rajamani, and R. Sharma (2022) Jigsaw: large language models meet program synthesis. In Proceedings of the 44th International Conference on Software Engineering, p. 1219–1231. Cited by: §I. [16] N. Jain, A. Schwarzschild, Y. Wen, G. Somepalli, J. Kirchenbauer, P. Chiang, M. Goldblum, A. Saha, J. Geiping, and T. Goldstein (2023) Baseline defenses for adversarial attacks against aligned language models. External Links: 2309.00614, Link Cited by: §IV-F. [17] D. Jin, Z. Jin, Z. Hu, O. Vechtomova, and R. Mihalcea (2022) Deep learning for text style transfer: a survey. Computational Linguistics 48 (1), p. 155–205. Cited by: §V. [18] Z. Jin, D. Jin, J. Mueller, N. Matthews, and E. Santus (2019-11) IMaT: unsupervised text attribute transfer via iterative matching and translation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China, p. 3097–3109. External Links: Link, Document Cited by: §V. [19] N. Kandpal, M. Jagielski, F. Tramèr, and N. Carlini (2023) Backdoor attacks for in-context learning with language models. arXiv preprint arXiv:2307.14692. Cited by: §I-B. [20] K. Krishna, J. Wieting, and M. Iyyer (2020-11) Reformulating unsupervised style transfer as paraphrase generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, p. 737–762. External Links: Link, Document Cited by: §IV-B. [21] K. Kurita, P. Michel, and G. Neubig (2020) Weight poisoning attacks on pretrained models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, p. 2793–2806. Cited by: §I-A. [22] J. Li, Y. Yang, Z. Wu, V. Vydiswaran, and C. Xiao (2023) Chatgpt as an attack tool: stealthy textual backdoor attack via blackbox generative model trigger. arXiv preprint arXiv:2304.14475. Cited by: §IV-A, §V. [23] J. Li, R. Jia, H. He, and P. Liang (2018-06) Delete, retrieve, generate: a simple approach to sentiment and style transfer. In Proc. NAACL-HLT 2018, Volume 1 (Long Papers), New Orleans, Louisiana, p. 1865–1874. External Links: Link, Document Cited by: §V. [24] M. Q. Li and B. C. Fung (2025) Security concerns for large language models: a survey. Journal of Information Security and Applications 95, p. 104284. Cited by: §I. [25] S. Li, H. Liu, T. Dong, B. Z. H. Zhao, M. Xue, H. Zhu, and J. Lu (2021) Hidden backdoors in human-centric language models. In Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security, p. 3123–3140. Cited by: §I-B, §I-D. [26] Y. Li, T. Li, K. Chen, J. Zhang, S. Liu, W. Wang, T. Zhang, and Y. Liu (2024) Badedit: backdooring large language models by model editing. arXiv preprint arXiv:2403.13355. Cited by: §I-B. [27] Y. Li, H. Huang, Y. Zhao, X. Ma, and J. Sun (2025) BackdoorLLM: a comprehensive benchmark for backdoor attacks and defenses on large language models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, Cited by: §I, §I-B, §IV-A, §V. [28] D. Liu, J. Fu, Y. Zhang, C. Pal, and J. Lv (2020) Revision in continuous space: unsupervised text style transfer without adversarial learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34, p. 8376–8383. Cited by: §V. [29] X. Liu, K. Ji, Y. Fu, W. Tam, Z. Du, Z. Yang, and J. Tang (2022-05) P-tuning: prompt tuning can be comparable to fine-tuning across scales and tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Dublin, Ireland, p. 61–68. External Links: Link, Document Cited by: §I-D. [30] Microsoft (2024) Phi-4. External Links: Link Cited by: §IV-A. [31] H. Naveed, A. U. Khan, S. Qiu, M. Saqib, S. Anwar, M. Usman, N. Akhtar, N. Barnes, and A. Mian (2023) A comprehensive overview of large language models. arXiv preprint arXiv:2307.06435. Cited by: §I. [32] OpenAI (2024) GPT-3.5 turbo. External Links: Link Cited by: §IV-A, §IV-A. [33] OpenAI (2024) GPT-4 turbo. External Links: Link Cited by: §IV-A. [34] X. Pan, M. Zhang, B. Sheng, J. Zhu, and M. Yang (2022-08) Hidden trigger backdoor attack on NLP models via linguistic style manipulation. In 31st USENIX Security Symposium (USENIX Security 22), Boston, MA, p. 3611–3628. External Links: ISBN 978-1-939133-31-1, Link Cited by: §I, §I-B, §I-B, §I-B, §IV-B. [35] F. Qi, Y. Chen, M. Li, Y. Yao, Z. Liu, and M. Sun (2021-11) ONION: a simple and effective defense against textual backdoor attacks. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Online and Punta Cana, Dominican Republic, p. 9558–9566. External Links: Link, Document Cited by: §IV-F. [36] F. Qi, Y. Chen, X. Zhang, M. Li, Z. Liu, and M. Sun (2021-11) Mind the style of text! adversarial and backdoor attacks based on text style transfer. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Online and Punta Cana, Dominican Republic, p. 4569–4580. External Links: Link, Document Cited by: §I, §I-B, §I-B, §IV-B. [37] S. Rao and J. Tetreault (2018-06) Dear sir or madam, may I introduce the GYAFC dataset: corpus, benchmarks and metrics for formality style transfer. In Proc. NAACL-HLT 2018, Volume 1 (Long Papers), New Orleans, Louisiana, p. 129–140. External Links: Link, Document Cited by: §V. [38] E. Reif, D. Ippolito, A. Yuan, A. Coenen, C. Callison-Burch, and J. Wei (2022-05) A recipe for arbitrary text style transfer with large language models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Dublin, Ireland, p. 837–848. External Links: Link, Document Cited by: §I-B, §V. [39] G. Shen, S. Cheng, Z. Zhang, G. Tao, K. Zhang, H. Guo, L. Yan, X. Jin, S. An, S. Ma, and X. Zhang (2025-05) BAIT: Large Language Model Backdoor Scanning by Inverting Attack Target . In 2025 IEEE Symposium on Security and Privacy (SP), Vol. , Los Alamitos, CA, USA, p. 1676–1694. External Links: ISSN 2375-1207, Document, Link Cited by: §IV-F, §IV-F. [40] J. Shi, Y. Liu, P. Zhou, and L. Sun (2023) Badgpt: exploring security vulnerabilities of chatgpt via backdoor attacks to instructgpt. arXiv preprint arXiv:2304.12298. Cited by: §I. [41] K. Singhal, T. Tu, J. Gottweis, R. Sayres, E. Wulczyn, M. Amin, L. Hou, K. Clark, S. R. Pfohl, H. Cole-Lewis, et al. (2025) Toward expert-level medical question answering with large language models. Nature Medicine, p. 1–8. Cited by: §I. [42] Z. Tan, Q. Chen, Y. Huang, and C. Liang (2024) Target: template-transferable backdoor attack against prompt-based nlp models via gpt4. In CCF International Conference on Natural Language Processing and Chinese Computing, p. 398–411. Cited by: §V. [43] R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto (2023) Stanford alpaca: an instruction-following llama model. GitHub. Note: https://github.com/tatsu-lab/stanford_alpaca Cited by: §I-D, 1st item, TABLE VI. [44] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al. (2023) Llama: open and efficient foundation language models. arXiv preprint arXiv:2302.13971. Cited by: §I. [45] H. Wang and K. Shu (2024) Trojan activation attack: red-teaming large language models using steering vectors for safety-alignment. CIKM ’24, New York, NY, USA, p. 2347–2357. External Links: ISBN 9798400704369, Link, Document Cited by: §I, §V. [46] Y. Wang, Q. Liu, and C. Jin (2023) Is rlhf more difficult than standard rl? a theoretical perspective. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA. Cited by: §V. [47] J. Wei, M. Fan, W. Jiao, W. Jin, and T. Liu (2024-01) BDMMT: backdoor sample detection for language models through model mutation testing. Trans. Info. For. Sec. 19, p. 4285–4300. External Links: ISSN 1556-6013, Link, Document Cited by: §I-B, §I-B. [48] Z. Xiang, F. Jiang, Z. Xiong, B. Ramasubramanian, R. Poovendran, and B. Li (2024) Badchain: backdoor chain-of-thought prompting for large language models. arXiv preprint arXiv:2401.12242. Cited by: §I-B. [49] J. Xu, M. Ma, F. Wang, C. Xiao, and M. Chen (2024-06) Instructions as backdoors: backdoor vulnerabilities of instruction tuning for large language models. In Proc. NAACL-HLT 2024 (Volume 1: Long Papers), Mexico City, Mexico, p. 3111–3126. External Links: Link, Document Cited by: §I, §I-B, §V. [50] H. Yao, J. Lou, and Z. Qin (2024) Poisonprompt: backdoor attack on prompt-based large language models. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 7745–7749. Cited by: §I-B. [51] Y. Yao, J. Duan, K. Xu, Y. Cai, Z. Sun, and Y. Zhang (2024) A survey on large language model (llm) security and privacy: the good, the bad, and the ugly. High-Confidence Computing, p. 100211. Cited by: §V. [52] W. You, Z. Hammoudeh, and D. Lowd (2023-12) Large language models are better adversaries: exploring generative clean-label backdoor attacks against text classifiers. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, p. 12499–12527. External Links: Link, Document Cited by: §I-B, §IV-F, §V. [53] J. Zhang, H. Bu, H. Wen, Y. Liu, H. Fei, R. Xi, L. Li, Y. Yang, H. Zhu, and D. Meng (2025) When llms meet cybersecurity: a systematic literature review. Cybersecurity 8 (1), p. 1–41. Cited by: §V. [54] R. Zhang, H. Li, R. Wen, W. Jiang, Y. Zhang, M. Backes, Y. Shen, and Y. Zhang (2024) Instruction backdoor attacks against customized \llms\. In 33rd USENIX Security Symposium (USENIX Security 24), p. 1849–1866. Cited by: §I, §I-B, §IV-A, §IV-A, §IV-F, §V. [55] X. Zhang, J. Zhao, and Y. LeCun (2015) Character-level convolutional networks for text classification. In Proceedings of the 29th International Conference on Neural Information Processing Systems - Volume 1, NIPS’15, Cambridge, MA, USA, p. 649–657. Cited by: 3rd item, 4th item, TABLE I, TABLE I. [56] X. Zhang, Z. Zhang, S. Ji, and T. Wang (2021) Trojaning language models for fun and profit. In 2021 IEEE European Symposium on Security and Privacy (EuroS&P), p. 179–197. Cited by: §I-B. [57] S. Zhao, M. Jia, Z. Guo, L. Gan, X. XU, X. Wu, J. Fu, F. Yichao, F. Pan, and A. T. Luu (2025) A survey of recent backdoor attacks and defenses in large language models. Transactions on Machine Learning Research. Note: Survey Certification External Links: ISSN 2835-8856, Link Cited by: §I, §I-B, §I-D, §V. [58] S. Zhao, M. Jia, A. T. Luu, F. Pan, and J. Wen (2024-11) Universal vulnerabilities in large language models: backdoor attacks for in-context learning. In Proc. EMNLP 2024, Miami, Florida, USA, p. 11507–11522. External Links: Link, Document Cited by: §I, §I-B, §IV-A, §IV-A. [59] Y. Zhou, T. Ni, W. Lee, and Q. Zhao (2025) A survey on backdoor threats in large language models (llms): attacks, defenses, and evaluations. arXiv preprint arXiv:2502.05224. Cited by: §I-B. [60] W. Zhu, H. Liu, Q. Dong, J. Xu, S. Huang, L. Kong, J. Chen, and L. Li (2024-06) Multilingual machine translation with large language models: empirical results and analysis. In Findings of the Association for Computational Linguistics: NAACL 2024, Mexico City, Mexico, p. 2765–2781. External Links: Link, Document Cited by: §I.