Paper deep dive
OmniGuard: An Efficient Approach for AI Safety Moderation Across Languages and Modalities
Sahil Verma, Keegan Hines, Jeff Bilmes, Charlotte Siska, Luke Zettlemoyer, Hila Gonen, Chandan Singh
Models: Llama3.3-70B-Instruct, Llama-Omni-8B, Molmo-7B
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/12/2026, 6:37:37 PM
Summary
OmniGuard is an efficient AI safety moderation framework that detects harmful prompts across multiple languages and modalities (text, image, audio) by leveraging internal, universally aligned representations of LLMs and MLLMs. By training a lightweight classifier on these aligned embeddings, it avoids the computational overhead of separate guard models and achieves state-of-the-art performance in multilingual and multimodal safety benchmarks.
Entities (5)
Relation Signals (4)
OmniGuard â usesmetric â U-Score
confidence 100% · OMNIGUARD uses a custom metric (U-Score) to identify representations that generalize across languages and modalities.
OmniGuard â usesmodel â Llama3.3-70B-Instruct
confidence 95% · We extract embeddings from each layer of Llama3.3-70B-Instruct for the sentences in all 73 languages and use them to compute the U-Score
OmniGuard â usesmodel â Molmo-7B
confidence 95% · We extract embeddings for each image and its corresponding caption using an MLLM (also Molmo-7B)
OmniGuard â improvesaccuracyover â PolyGuard
confidence 90% · OMNIGUARD achieves the highest accuracy (86.36%) compared to the baselines... The strongest baseline is Polyguard
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The emerging capabilities of large language models (LLMs) have sparked concerns about their immediate potential for harmful misuse. The core approach to mitigate these concerns is the detection of harmful queries to the model. Current detection approaches are fallible, and are particularly susceptible to attacks that exploit mismatched generalization of model capabilities (e.g., prompts in low-resource languages or prompts provided in non-text modalities such as image and audio). To tackle this challenge, we propose Omniguard, an approach for detecting harmful prompts across languages and modalities. Our approach (i) identifies internal representations of an LLM/MLLM that are aligned across languages or modalities and then (ii) uses them to build a language-agnostic or modality-agnostic classifier for detecting harmful prompts. Omniguard improves harmful prompt classification accuracy by 11.57\% over the strongest baseline in a multilingual setting, by 20.44\% for image-based prompts, and sets a new SOTA for audio-based prompts. By repurposing embeddings computed during generation, Omniguard is also very efficient ($\approx\!120 \times$ faster than the next fastest baseline). Code and data are available at: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2505.23856
- Canonical: https://arxiv.org/abs/2505.23856
- Code: https://github.com/vsahil/OmniGuard
Trouble viewing inline? Open PDF directly â
Full Text
63,921 characters extracted from source content.
Expand or collapse full text
OMNIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and Modalities Sahil Verma 1 Keegan Hines 2 Jeff Bilmes 1 Charlotte Siska 2 Luke Zettlemoyer 1 Hila Gonen 1 Chandan Singh 2 1 University of Washington 2 Microsoft Abstract The emerging capabilities of large language models (LLMs) have sparked concerns about their immediate potential for harmful mis- use. The core approach to mitigate these con- cerns is the detection of harmful queries to the model. Current detection approaches are fallible, and are particularly susceptible to at- tacks that exploit mismatched generalization of model capabilities (e.g., prompts in low- resource languages or prompts provided in non-text modalities such as image and audio). To tackle this challenge, we propose OMNI- GUARD, an approach for detecting harmful prompts across languages and modalities. Our approach (i) identifies internal representations of an LLM/MLLM that are aligned across lan- guages or modalities and then (i) uses them to build a language-agnostic or modality-agnostic classifier for detecting harmful prompts. OM- NIGUARD improves harmful prompt classifi- cation accuracy by 11.57% over the strongest baseline in a multilingual setting, by 20.44% for image-based prompts, and sets a new SOTA for audio-based prompts. By repurposing em- beddings computed during generation, OMNI- GUARD is also very efficient (â 120Ăfaster than the next fastest baseline).Code and data are available athttps://github.com/ vsahil/OmniGuard. 1 Introduction The rapid rise of capabilities in large language models (LLMs) has created an urgent need for safeguards to prevent their immediate harmful misuse as they are deployed to human users en masse (Bommasani et al., 2022). Moreover, these safeguards are critical for defending against fu- ture potential harms from LLMs (Bengio et al., 2024). Standard safeguard approaches broadly in- clude approaches such as safety training using rein- forcement learning from human feedback (Ouyang et al., 2022a; Leike et al., 2018) or using pre-trained Figure 1: OMNIGUARD builds a harmfulness classifier that operates on internal representations of an LLM (or MLLM). OMNIGUARD uses a custom metric (U-Score) to identify rep- resentations that generalize across languages and modalities. At inference time, OMNIGUARD re-uses the embeddings from the LLM/MLLM being used for generation, and thereby com- pletely avoids the overhead of passing the inputs through a separate guard model for safety moderation. guard models that classify the safety of an input prompt (OpenAI, 2025; Inan et al., 2023; Han et al., 2024). With these safeguards in place, harmful prompts in high-resource languages, e.g., English, are suc- cessfully detected. However, harmful prompts in low-resource languages can often bypass these safe- guards (Deng et al., 2024; Yong et al., 2024; Yang et al., 2024), i.e., jailbreaking the LLM. Modern LLMs are vulnerable to attacks not only from low- resource natural languages, but also from artificial cipher languages, e.g., base64 or caesar encoding of English prompts (Wei et al., 2023; Yuan et al., 2024a). This phenomenon also extends beyond text to jailbreaking multimodal LLMs (MLLMs) using modalities such as images (Gong et al., 2025; Liu et al., 2024b) or audio (Yang et al., 2025). Wei et al. (2023) argue that these attacks are successful due to mismatched generalization, a sce- nario in which the modelâs safety training does not generalize to other settings, but general perfor- mance does. This may happen because pretrain- ing data often includes more diverse data than that available for safety finetuning (Ghosh et al., 2024b). arXiv:2505.23856v2 [cs.CL] 9 Dec 2025 In this work, we defend against attacks that exploit the mismatched generalization of the safety train- ing of LLMs and MLLMs. Specifically, we defend against attacks that utilize low-resource languages, both natural and cipher languages, as well as at- tacks employing other modalities, such as images and audio. We introduce OMNIGUARD, an approach that builds a classifier using the internal representa- tions of a model. These representations are ex- tracted from specific layers that produce represen- tations that are universally similar across multiple languages and across multiple modalities. OMNI- GUARDâs classifier trained on such representations, is able to accurately detect harmful inputs across 73 languages, with an average of 86.22% accu- racy across 53 natural languages and an average of 73.06% accuracy across 20 cipher languages. OM- NIGUARD can also detect harmful inputs provided as images with 88.31% and as audio with 93.09% accuracy respectively. In contrast to popular guard models such as Lla- maGuard (Inan et al., 2023), AegisGuard (Ghosh et al., 2024a), or WildGuard (Han et al., 2024), OMNIGUARD does not require training a separate LLM specifically to detect harmfulness. By build- ing a classifier that uses the internal representations of the main LLM or MLLM, OMNIGUARD avoids the overhead of passing the prompt through a sepa- rate guard model, making it very efficient. In summary, our contributions are the following: (1) We propose OMNIGUARD, an approach for detecting harmful prompts, (2) we show that OM- NIGUARD accurately detects harmfulness across multiple languages and multiple modalities, (3) we show that OMNIGUARD is very sample-efficient during training, and (4) we show that OMNIGUARD is highly efficient at inference time. 2 Methodology OMNIGUARD seeks to robustly detect harmful prompts, regardless of their language or modal- ity. We first leverage the tendency of LLMs and MLLMs to create universal representations that are similar across languages (Wendler et al., 2024; Zhao et al., 2024) and across modalities (Wu et al., 2024; Zhuang et al., 2025) in Section 2.1, and then use them to train harmfulness classifiers that ro- bustly detect harmful inputs in Section 2.2. 2.1Finding language-agnostic representations in an LLM The first step of OMNIGUARD searches for inter- nal representations of an LLM that are universally shared across languages. We prompt an LLM with English sentences and their translations to other lan- guages, and extract their representations at differ- ent layers. 1 For language-agnostic representations, we expect the similarity between the representa- tions of English sentences and the representations of their translations to be similar, and we expect this similarity to be higher than the similarity be- tween representations of two sentences that are not translations of each other (a random pair of sen- tences). We concretize this notion by defining the Universality Score (U-Score, Eq. 1), which is the difference between the average cosine similarities of pairs of sentences that are translations of each other and pairs of sentences that are not. U-Score := 1 N X iâ[N] CosSim (Emb(e i ), Emb(l i )) â 1 N(Nâ 1) X i,jâ[N] iÌž=j CosSim (Emb(e i ), Emb(l j )) (1) wheree i andl i are sentences in English and their translations to another language. This procedure can be generalized to new differ- ent modalities rather than different languages by changing which embeddings are being used. For example, to determine if internal representations of an MLLM are aligned across modalities, we replace embeddings for a translated piece of text with embeddings from a different modality (e.g. a text caption and its corresponding image, or a text transcription and its corresponding audio clip). See experimental details in Section 3. 2.2 Fitting a harmfulness classifier After selecting the layer that maximizes the U- Score, we extract embeddings from that layer and use them as inputs to fit a lightweight, supervised classifier that predicts harmfulness. In our exper- iments, the classifier is a multilayer perceptron with 2 hidden layers (with hidden sizes 512 and 256). At inference time, when a prompt is passed to a model for generation, OMNIGUARD applies 1 The representation of a prompt is computed by averaging the representation over each token in the prompt. 020406080100 Normalized layer depth (percentage) 0.00 0.02 0.04 0.06 0.08 0.10 0.12 0.14 Universality Score (U-Score) Natural Languages Cipher Languages Images and Captions Audios and Transcriptions Figure 2: The U-Score across different layers for different modalities. (A) Different layers of the model Llama3.3-70B- Instruct for different languages. (B) The Cross-Model Align- ment Score at different layers of the model (Molmo-7B) for similarity between images and captions. The highest values are obtained with at layers 21-25, indicating better alignment between images and their text captions at these layers. (C) The Cross-Model Alignment Score at different layers of the model (Llama-Omni 8B) for similarity between audios and transcriptions. The highest values are obtained with at layers 20-23, indicating better alignment between audios and their text transcriptions at these layers. this classifier to the embeddings generated by the model, incurring minimal overhead at inference time for safety classification. Note, however, that this approach only applies to open-source models, for which OMNIGUARD can build a classifier by obtaining embeddings. During training, only the lightweight classifierâs parameters are learned (the original model is never modified), making the train- ing process data-efficient and inexpensive. 3 Experimental Setup Table 1 and Table 2 give details on all the models and datasets for this section. 3.1 Selecting universal layers via the U-Score Selecting language-agnostic layers To select language-agnostic layers, we use a dataset of trans- lated sentence pairs spanning various languages. Specifically, we use sentences in 53 natural lan- guages from the Flores200 dataset and additionally translate the sentences into 20 cipher languages (using encodings such as Caesar shifts, base64, hexadecimal); see a full list in Section A. We ex- tract embeddings from each layer of Llama3.3-70B- Instruct for the sentences in all 73 languages and use them to compute the U-Score (averaged over languages). Fig. 2 shows the U-Score as a func- tion of layer depth. For natural languages (blue curve), the U-Score peaks in the middle layers of the model, with the highest values in layer 57 (out of 81 layers). For cipher languages (red curve), the U-Score is much lower than for natural languages, suggesting the model fails to represent semantic similarity in these languages (see analysis in Sec- tion 5). Selecting modality-agnostic layers To select layers aligned between images and captions, we use the M-Vet v2 dataset, a popular dataset for MLLM evaluation containing 517 examples, each consisting of a text question paired with one or more images. We generate captions for each image using a captioning model (Molmo-7B) and then extract embeddings for each image and its corre- sponding caption using an MLLM (also Molmo- 7B) and use them to compute the U-Score, which peaks in layer 22 (out of 28 layers; see Fig. 2 green curve). To select layers aligned between text and audio, we use the audio version of the Alpacaeval dataset from VoiceBench, a dataset of 636 audio-transcript pairs. We extract embeddings from each layer of an MLLM (LLaMA-Omni 8B) and use them to compute the U-Score, which peaks at layer 21 (out of 32 layers; see Fig. 2 purple curve). Overall, we see that LLMs and MLLMs generate representations that are shared across languages and modalities. 3.2 Training and evaluating the harmfulness classifier 3.2.1 Setup for multilingual text attacks OMNIGUARD classifier. Following Section 3.1, we build a classifier that takes as input embeddings from layer 57 of Llama3.3-70B-Instruct. As train- ing data, we randomly select 2,800 examples from the Aegis AI Content Safety dataset, balancing the benign and harmful classes. Notably, this dataset is about 18Ăsmaller than the training data used by our baseline methods. We translate these En- glish examples to 52 other natural languages (via the Google Translate API) and 20 cipher languages (using fixed rules), totaling 73 languages. We train OMNIGUARD using only half the languages (see list in Section A). Baselines.We compare to many popular guard models (see Table 2) middle row. Notably, Duo- Guard and PolyGuard were trained to detect harm- ful prompts across multiple languages. For a more direct comparison, we also compare to finetuned Dataset nameCitationHuggingFace IDNumber of examples General Flores200(Team et al., 2022) Muennighoff/flores200997 M-Vet v2(Yu et al., 2024b) whyu/m-vet-v2517 SST-2(Socher et al., 2013) stanfordnlp/sst21000 Text Aegis AI Content Safety Dataset(Ghosh et al., 2024b) nvidia/Aegis-AI-Content-Safety-Dataset-1.010,800 MultiJail(Deng et al., 2024) DAMO-NLP-SG/MultiJail315 Xsafety(Wang et al., 2024a) ToxicityPrompts/XSafety28,000 RTP-LX(de Wynter et al., 2025) ToxicityPrompts/RTP-LX30,300 AyaRedTeaming(Aakanksha et al., 2024) CohereLabs/aya_redteaming2662 Thai Toxicity tweets(Sirihattasak et al., 2018) tmu-nlp/thai_toxicity_tweet3,300 Ukr Toxicity(Dementieva et al., 2024) ukr-detect/ukr-toxicity-dataset5,000 HarmBench (HB)(Mazeika et al., 2024) walledai/HarmBench400 Forbidden Questions (FQ)(Shen et al., 2024a) TrustAIRLab/forbidden_question_set390 Simple Safety Tests(Vidgen et al., 2024) walledai/SimpleSafetyTests100 SaladBench (SaladB)(Li et al., 2024a) walledai/SaladBench26,500 Toxicity Jigsaw (TJS)(cjadams et al., 2017) Arsive/toxicity_classification_jigsaw26,000 Toxic Text(CorrĂȘa, 2023) nicholasKluge/toxic-text41,800 AdvBench(Zou et al., 2023a) walledai/AdvBench520 CodeAttack(Ren et al., 2024) https://github.com/AI45Lab/CodeAttack3120 Vision JailBreakV-28K(Luo et al., 2024) JailbreakV-28K/JailBreakV-28k8,000 VLSafe(Chen et al., 2024c) YangyiYY/LVLM_NLF1,110 FigStep(Gong et al., 2025) https://github.com/wangyu-ovo/MML500 MML SafeBench(Wang et al., 2024b) https://github.com/wangyu-ovo/MML2,510 Hades(Li et al., 2024e) Monosail/HADES750 SafeBench(Ying et al., 2024) Zonghao2025/safebench2,300 M SafetyBench(Liu et al., 2024b) PKU-Alignment/M-SafetyBench1680 RedTeamVLM(Li et al., 2024b) MMInstruction/RedTeamingVLM200 VLSBench(Hu et al., 2025) Foreshhh/vlsbench2,240 Audio VoiceBench (Alpacaeval)(Chen et al., 2024d) hlt-lab/voicebench636 AIAH(Yang et al., 2025) https://github.com/YangHao97/RedteamAudioLMMs350 Table 1: Details of datasets used for training and evaluation. Some of the text datasets are inherently multilingual : MultiJail (10 languages), XSafety (10 languages), RTP-LX (28 languages), Aya RedTeaming (8 languages), Thai Toxicity tweets (prompts in Thai), and Ukr Toxicity (prompts in Ukrainian). The remaining text datasets are English-only, and were translated to 72 other languages (52 natural and 20 cipher): HarmBench (HB), Forbidden Questions (FQ), Simple Safety Tests, SaladBench (SaladB), Toxicity Jigsaw (TJS), Toxic Text, and AdvBench. versions of DuoGuard and PolyGuard using the same 37 languages we use to train OMNIGUARD; Following the original PolyGuard paper, we fine- tuned these models using LoRA (Hu et al., 2021) for all linear layers with rank 8 and alpha 16 for one epoch with a learning rate of 2eâ 4. Datasets. We evaluate on several common text attack benchmarks (see Table 1). We additionally evaluate on three benchmarks from CodeAttacks that transform a harmful query as a list, a stack, or as a string in a Python program, obfuscating the harmfulness. For evaluation in this setup, we transform the harmful prompts from AdvBench and benign prompts from Toxicity Jigsaw datasets in the three code formats and subsample the Toxicity Jigsaw dataset to be of the same size as Advbench. Note that for this experiment, we only trained OM- NIGUARD on the English subset of the training dataset. 3.2.2 Setup for vision attacks OMNIGUARD classifier. Following Section 3.1, we build a classifier that takes as input embed- dings from layer 22 of Molmo-7B. As training data, we use 2000 image-query pairs randomly sampled from the JailBreakV-28K dataset and 1024 image- query pairs sampled from the VLSafe dataset as the harmful datapoints and 517 image-query from the M-Vet v2 dataset as the benign datapoints. Baselines.We compare to guard models that take an image or image-text pair and output a binary harmfulness classification (see Table 2 bottom row). We train VLMGuard on the same training data as OMNIGUARD. Datasets. We evaluate detecting image/text at- tacks using several datasets (see Table 1). Fig- Step and MML Safebench are typographic attacks that embed a harmful prompt in an image. MML Safebench further encrypts a harmful prompt in sev- eral variants, such as rotation, mirror images, word replacement, and with base64 encoding. Hades and Safebench consist of images and text queries where the text itself is harmful. M-safetybench, RTVLM, and VLSBench consist of an image and a query where the text query is seemingly benign, but when combined with the respective image, it is harmful (e.g. see Figure 1). 3.2.3 Setup for audio attacks OMNIGUARD classifier. Following Section 3.1, we build a classifier that takes as input embeddings from layer 21 of Llama-Omni-8B. We train the classifier on the English portion of the training data we use for the text setting, by using a text-to-speech model to convert the text into audio. We use the open-source Kokoro model as the text-to-speech model. Model nameCitationHuggingFace IDRough Parameter Count General Llama3.3-70B-Instruct(Grattafiori et al., 2024) meta-llama/Llama-3.3-70B-Instruct70B Molmo-7B(Deitke et al., 2024) allenai/Molmo-7B-D-09247B LLaMA-Omni 8B(Fang et al., 2025) ICTNLP/Llama-3.1-8B-Omni8B Kokoro(Hexgrad, 2025) hexgrad/Kokoro-82M82M Text LlamaGuard 1(Inan et al., 2023) meta-llama/LlamaGuard-7b7B LlamaGuard 2(Inan et al., 2023) meta-llama/Meta-Llama-Guard-2-8B8B LlamaGuard 3(Inan et al., 2023) meta-llama/Llama-Guard-3-8B8B AegisGuard Permissive(Ghosh et al., 2024a) nvidia/Aegis-AI-Content-Safety-LlamaGuard-Permissive-1.07B AegisGuard Defensive(Ghosh et al., 2024a) nvidia/Aegis-AI-Content-Safety-LlamaGuard-Defensive-1.07B WildGuard(Han et al., 2024) allenai/wildguard7B HarmBench (mistral)(Mazeika et al., 2024) cais/HarmBench-Mistral-7b-val-cls7B HarmBench (llama)(Mazeika et al., 2024) cais/HarmBench-Llama-2-13b-cls13B DuoGuard(Deng et al., 2025) DuoGuard/DuoGuard-1B-Llama-3.2-transfer1B PolyGuard(Kumar et al., 2025) ToxicityPrompts/PolyGuard-Qwen7B Vision Llama Guard 3 Vision(Chi et al., 2024) meta-llama/Llama-Guard-3-11B-Vision11B VLMGuard(Du et al., 2024) â2.2M LLavaGuard(Helff et al., 2025) AIML-TUDA/LlavaGuard-7B-hf7B Table 2: Model and baseline details. Baselines. We are unaware of any existing mod- els for detecting harmful audio input. The most relevant approach, SpeechGuard (Peri et al., 2024) adds noise as a defense against potentially harmful audio inputs but does not directly classify harmful- ness. To contextualize our results for audio bench- marks, we compare performance to guard models that directly classify the raw text present in the audio (OMNIGUARD and LlamaGuard3). Datasets.We use the two audio benchmarks (see Table 1 bottom row). We also evaluate on several text jailbreak benchmarks using Kokoro to convert them from text to speech: HB, FQ, Simple Safety Tests, SaladB, and TJS. We use Kokoro for gener- ating text-to-speech versions. 4 Results Defending against multilingual text attacksTa- ble 3 compares the accuracy of detecting harmful prompts for text benchmarks. Table 3(A) shows results for multilingual benchmarks, where OM- NIGUARD achieves the highest accuracy (86.36%) compared to the baselines, and achieves new state- of-the-art performance for 3 benchmarks: Multi- Jail, RTP-LX, and AyaRedTeaming. The strongest baseline is Polyguard, which yields an average ac- curacy of 83.19%, despite being trained on a much larger dataset (1.91M examples for Polyguard ver- sus 103K examples for OMNIGUARD). In bench- marks that were translated from English to various other languages, including cipher languages, we again see that OMNIGUARD achieves the highest accuracy (Table 3(B)). Finally, Table 3(C) shows that OMNIGUARD outperforms finetuned versions of DuoGuard and Polyguard on unseen languages, demonstrating that OMNIGUARD can outperform methods that were trained specifically for multilin- gual harmfulness classification. Defending against image-based attacksTable 4 shows the accuracy of detecting harmful image and text prompts for (A) pairs consisting of images and text queries, where either the image or both the im- age and query can be harmful and (B) typographic images with various encryptions. OMNIGUARD achieves the highest performance for both sets of benchmarks (95.44% and 79.76%) while being trained using only about 3500 image-query pairs (compared to about 5500 datapoints used by Llava- Guard). The only benchmark where OMNIGUARD fails to detect harmful prompts is MML Base64, which consists of typographed images of prompts encrypted using base64 encoding. Defending against audio-based attacksTable 5 shows the accuracy of detecting harmful audio prompts. OMNIGUARD detects harmful audio in- put with high accuracy across all benchmarks. As we are not aware of any existing defenses for au- dio jailbreaks, we compare against OMNIGUARD and LlamaGuard3âs accuracy in detecting harmful prompts when the same inputs are provided in En- glish text. The accuracy OMNIGUARD achieves in detecting harmful audio inputs is similar to or higher than its performance for detecting harmful text inputs. Data-efficient adaptation We also evaluate the accuracy of OMNIGUARD and baselines in adapt- ing to out-of-distribution code attacks given very few samples. In this setting, some prior work has speculated that guard models may be very data efficient, as they can make use of few-shot exam- ples in-context (Inan et al., 2023). However, we find that baseline guard models generally struggle to rapidly adapt to this setting given few-shot ex- amples (Figure 3). 2 In contrast, OMNIGUARD is 2 Note that we omit baseline guard models that achieve 90% accuracy or greater without any few-shot examples, as MultiJailXsafetyRTP-LXAya RedTeamingThai ToxUkr ToxAvg. (A) Multilingual text benchmarks LlamaGuard 139.2757.0148.6654.4941.3153.9949.12 LlamaGuard 248.6952.6634.6958.5842.8651.7948.21 LlamaGuard 366.8764.3445.5763.8346.7351.7956.52 AegisGuard (P)61.4979.7875.0778.8856.0965.7569.51 AegisGuard (D) 79.7190.7792.1789.7863.3467.9580.62 WildGuard 42.5571.2371.9461.4540.4255.0357.10 HarmBench (llama)0.220.140.00.0339.0450.114.92 HarmBench (mistral)2.45.655.147.3940.4250.5518.59 MD-Judge25.7853.5866.4646.2039.4853.8947.56 DuoGuard 39.2063.4266.5761.8045.6350.7554.56 PolyGuard82.0096.4183.8690.3470.4376.0783.19 OMNIGUARD93.8393.6494.5594.3168.773.186.36 HarmBenchFQSimpleSTSaladBTJSToxTextAdvBenchAvg. (B) Translated text benchmarks LlamaGuard 132.4723.7534.3223.2762.4965.5534.3939.46 LlamaGuard 257.1943.7250.7134.5458.2162.1756.9551.93 LlamaGuard 370.0253.2567.8146.3062.3370.8770.2662.98 AegisGuard (P)62.1643.0156.5544.9273.6972.8062.1259.32 AegisGuard (D) 88.5376.6787.6478.2771.3868.7290.7780.28 WildGuard33.6431.2033.9027.3766.6167.2739.9842.85 HarmBench (llama)0.030.110.010.0748.6249.970.0114.12 HarmBench (mistral)2.321.752.041.6650.5350.691.715.81 MD-Judge16.1912.1122.2913.8165.3464.2625.6731.38 DuoGuard20.4444.3628.7936.8868.5769.0728.5842.38 PolyGuard66.2256.0562.5354.8878.3476.5267.9666.07 OMNIGUARD89.1389.5789.6287.3076.6875.0786.5984.85 HarmBenchFQSimpleSTSaladBTJSToxTextAdvBenchAvg. (C) Unseen langs. FT DuoGuard23.5939.0828.1433.2954.153.2328.2937.1 FT PolyGuard72.4579.8476.8176.8574.0772.3373.5575.13 OMNIGUARD86.5186.6586.4285.0172.8271.4484.2981.88 Table 3: Accuracy of detecting harmful prompts for text attack benchmarks that are (A) multilingual benchmarks, (B) English translated to 73 languages, and (C) English translated to languages not seen at training time. In all settings, OMNIGUARD achieves the highest performance. Table B1 further stratifies these results by high-resource, low-resource, and cipher languages. HadesVLSBenchMM-SafetyBenchSafeBenchRTVLMFigStepAvg. (A) Image +Query Llama3 Vision GRD76.003.9731.9068.4056.5047.4047.36 VLMGuard98.0074.5692.2073.9094.0099.8088.74 LLavaGuard23.7342.0810.9512.1018.503.4018.46 OMNIGUARD100.0092.2499.8291.6089.00100.0095.44 MML RotateMML MirrorMML W.R.MML Q.R.MML Base64Avg. (B) Typographed image Llama3 Vision GRD83.2068.0096.4025.4098.8074.36 VLMGuard6.8021.00100.086.200.2042.84 LLavaGuard0.000.000.0011.400.002.28 OMNIGUARD 100.0100.099.6098.800.4079.76 Table 4: Accuracy of detecting harmful queries in multimodal benchmarks for (A) image-query pairs and (B) typographed images with encrypted text. OMNIGUARD achieves the highest performance for both kinds of benchmarks. AIAHSafeBench (M)SafeBench (F)HBFQSimpleSTSaladBTJSAdvBench OMNIGUARD (Audio)91.1494.493.895.9890.4297.094.2182.0398.85 OMNIGUARD (text-en)---92.093.393.090.293.290.0 LlamaGuard3 (text-en) ---97.3278.7599.067.0372.1698.07 Table 5: Accuracy of detecting harmful queries in audio. OMNIGUARD is able to detect harmful audio inputs with high accuracy across all benchmarks. Since there are no baselines for detecting harmful prompts in audio, we compare the performance against OMNIGUARDâs and LlamaGuard3 when the same benchmarks are provided as text in English. 012345 Number of Few-Shot Examples 50 60 70 80 90 Accuracy Code Attack: Python List OmniGuard AegisGuard (P) AegisGuard (D) LlamaGuard 1 LlamaGuard 2 DuoGuard 012345 Number of Few-Shot Examples 50 60 70 80 90 Accuracy Code Attack: Python Stack OmniGuard AegisGuard (P) AegisGuard (D) LlamaGuard 1 LlamaGuard 2 DuoGuard 012345 Number of Few-Shot Examples 50 60 70 80 90 100 Accuracy Code Attack: Python String OmniGuard AegisGuard (P) AegisGuard (D) LlamaGuard 1 LlamaGuard 2 DuoGuard Figure 3: Accuracy of detecting harmful prompts in a few-shot setting. As few-shot examples are provided, OMNIGUARD quickly achieves near-perfect accuracy, despite the attacks being quite different from its training data (e.g. without any few-shot examples, OMNIGUARDâs accuracy is close to 50% ). In contrast, the guard model baselines improve their accuracy slowly in a few-shot setting, despite sometimes having seen similar code attacks in their training data. Accuracies are averaged over 50 random sets of few-shot examples; error bars show the standard error of the mean. able to rapidly achieve close to 100% accuracy for all three benchmarks by updating its lightweight parameters using less than five examples. 5 Analysis Effect of U-Score-based layer selection.We per- form ablation experiments to determine the effect of selecting the appropriate layer for training the OMNIGUARD classifier. For the text-only model, we compare the U-Score-selected layer (57) to 3 their training data likely explicitly includes code attacks. Thai ToxUkr ToxTJSToxTextAvg. Layer 1062.165.566.9561.8964.42 Layer 7565.266.470.7265.7968.26 Last Layer63.151.261.3356.76 59.05 U-Score selected layer (57)68.773.176.875.07 73.4 Table 6: OMNIGUARDâs accuracy of detecting harmful prompts when trained using representations from different model layers. Guard MethodInference Time (s)â LlamaGuard 387.25 AegisGuard (D)152.26 WildGuard306.14 MD-Judge128.26 DuoGuard4.85 PolyGuard409.90 OMNIGUARD0.04 Table 7: Average inference time required for harmfulness pre- diction on the AdvBench dataset (averaged over 5 languages). OMNIGUARD is about 120Ăfaster than the fastest baseline (DuoGuard). other layers (layer 10, layer 75, and the last layer) when used for a set of toxicity prediction tasks. Ta- ble 6 shows that the representations from the layer with the highest U-Score result in significantly bet- ter harmfulness classification accuracy, improving between 5% and 14% compared to the other layers. We show ablation over more layers in Table B2. OMNIGUARDâs efficiencyOMNIGUARD is highly efficient at inference time because it re-uses the internal representations of the main LLM that is already processing the user query for genera- tion. Therefore, its compute time is only that of a lightweight multilayer perceptron, making it much faster than baseline guard models (note that this does limit OMNIGUARD to only work when the generation model is open-source, so embeddings can be extracted). Table 7 shows the inference time required by various guard models to predict the harmfulness of prompts in the AdvBench dataset in English, translated to Spanish, French, Telugu, and base64 encoding. OMNIGUARD is the fastest and is about 120Ăfaster than the fastest baseline (Duo- Guard). Inference time as measured on a machine with 1 L40 GPU, 4 CPUs, and 50 GB RAM. Performance comparison across base LLMs We compare OMNIGUARDâs accuracy when using different base LLMs in Table B3. We trained the classifiers on the layers with the best U-scores for each model. We find that the average accuracy for the moderator model trained using smaller LLMs is lower than the moderator model trained using 5060708090 Sentiment Classification Accuracy 30 40 50 60 70 80 90 Harmfulness Classification Accuracy sv bn ms sl caesarneg4 te cs ar fa caesar hexadecimal sr zh ta caesarneg6 it cy id ru eu alphanumeric pt nl bs ja is caesar1 vi fr mi da hy th en ko af hu caesar2 caesarneg2 sw uk no el mr si ascii fi tr caesar7 vowel leet pl caesarneg3 jv hr de caesarneg1 zu ro caesar4 caesarneg5 es bg caesar5 caesarneg7 base64 lo kn caesar6 hi gu am he R 2 = 0.24 Natural Languages Cipher Languages Least-squares fit Figure 4: Comparison of accuracy of classifying sentiments in various languages compared to detecting harmful prompts in those languages using OMNIGUARD. In both cases the LLM is Llama3.3-70B-Instruct. the larger Llama3.3-70B-Instruct model. Performance comparison across languages. We now analyze the harmfulness classification ac- curacy of OMNIGUARD by language, and compare it to the underlying LLMâs sentiment classifica- tion accuracy for the same language (Fig. 4). We measure harmfulness classification accuracy using OMNIGUARD on all the datasets in Table 3 and sentiment classification accuracy using Llama3.3- 70B-Instruct with zero-shot prompting on 72 trans- lated versions of the SST-2 dataset (translated to all the languages we consider). We observe that the accuracies are generally cor- related, indicating that OMNIGUARD is able to de- fend well in languages for which the LLM is more coherent/susceptible to attack. Unsurprisingly, the accuracies for natural languages are higher than the accuracies for cipher languages. Nevertheless, harmfulness classification accuracy can be fairly high, even when sentiment classification accuracy is near chance (50%). 6 Related Work Jailbreak Attacks in LLMsSeveral techniques have recently emerged to attack or jailbreak LLMs. Early techniques relied on manual effort and were very time-intensive (Shen et al., 2024b; An- driushchenko et al., 2024). Later techniques auto- mated this process, e.g., Zou et al. (2023b); Jones et al. (2023); Zhu et al. (2023) proposed gradient- based approaches to identify inputs to jailbreak LLMs with white-box access. Another set of tech- niques start from a set of human written prompts and modify them using approaches like genetic al- gorithms (Liu et al., 2024a; Lapid et al., 2024; Li et al., 2024d), fuzzing (Yu et al., 2024a), or rein- forcement learning (Chen et al., 2024b) to automat- ically produce prompts for jailbreaking. Another set of techniques, use a helper LLM to generate prompts that attack a target LLM (Chao et al., 2024; Ding et al., 2024; Mehrotra et al., 2024). Finally, Wei et al. (2024); Wang et al. (2023); Anil et al. (2024); Pernisi et al. (2024) use simple in-context demonstrations to jailbreak the models by overcom- ing its safety training and Russinovich et al. (2024); Li et al. (2024c) propose using multi-turn dialogues to jailbreak models. Multilingual Jailbreak Attacks Most of the aforementioned jailbreak techniques focus on at- tacks in English, against which significant defense exists both at the model and system level. To tackle this, a novel set of techniques have emerged that attack models using inputs in various languages or obfuscations that are able to bypass the safety guardrails. Deng et al. (2024); Yong et al. (2024); Wang et al. (2024a); Yang et al. (2024); Yoo et al. (2024); Upadhayay and Behzadan (2024); Song et al. (2024) demonstrated that attacking models us- ing mid and low resources languages led to higher attack success rates, compared to the case of at- tacking the model in high-resource languages like English. Going beyond natural languages, a newer set of works propose using cipher characters or lan- guages to evade the safety filters, e.g., Jin et al. (2024) propose interspersing cipher characters in between text, Jiang et al. (2024) propose replacing the unsafe words with their ASCII art versions, and Yuan et al. (2024a) propose prompting models in cipher languages like Morse, Atbash, Caesar. Multimodal Jailbreak Attacks Using modali- ties apart from text aims to explore a completely new attack surface, like images or audios. Several recent works have shown that MLLMs remain vul- nerable to being jailbroken when prompted with images or audios that have a harmful query (the same harmful query in text would be easily de- tected as harmful). Liu et al. (2024b) show that using a prompt with a correlated image, e.g., us- ing an image of a bomb when asking the model to answer the question: How to make a bomb? is more likely to jailbreak a model than when using an uncorrelated image. Hu et al. (2025) argue that providing a harmful image with a benign query (see Figure 1) further increases the potential of jail- breaking the model. Gong et al. (2025) and Wang et al. (2024b) demonstrate jailbreaking models by simply typographically embedding harmful queries in an image. Safety moderation in LLMsSafety moderation in LLMs broadly fits into two categories: intrin- sic and extrinsic. Intrinsic mechanisms include finetuning or RLHF training on an LLM (Bianchi et al., 2024; Chen et al., 2024a; Yuan et al., 2024b; Ouyang et al., 2022b; Bai et al., 2022; Dai et al., 2023). Extrinsic safety mechanisms utilize exter- nal models to detect harmful inputs and responses; these models can either be simple filters or use guard models. Jain et al. (2023); Alon and Kam- fonas (2023); Hu et al. (2024) propose using per- plexity filtering for detecting harmful prompts. Guard models defend LLMs by training separate LLMs to detect harmful text (see Table 2). Sepa- rately, a line of work has introduced interpretability methods to transparently expose safety concerns in LLMs (Bereska and Gavves, 2024; Singh et al., 2023; Arditi et al., 2024; Benara et al., 2024). A few other works also use internal model repre- sentations to defend against harmful inputs. How- ever, they defer from our work: OMNIGUARD is a standalone content safety classification model while the other approaches like Jailbreak Antidote (Shen et al., 2025) and AdaSteer (Zhao et al., 2025) directly change the internal representations to de- fend against harmful inputs. Therefore, OMNI- GUARD is similar to other safety classification models like LlamaGuard and WildGuard. Addi- tionally, OMNIGUARD can detect harmful prompts in multiple languages (both natural and cipher) and multiple modalities using the same method while the other approaches only defend against attacks in English text. OMNIGUARD is also extremely efficient in producing safety predictions while the other approaches take as much time as the infer- ence of the underlying LLM, which can typically take several seconds to minutes. Safety against multilingual attacks Duo- Guard (Deng et al., 2025) and PolyGuard (Kumar et al., 2025) are the two previous guard models that were specifically trained to defend against multilingual attacks. DuoGuard uses a two-player RL-driven mechanism to generate harmful data in multiple languages and uses that to finetune a Llama3.2-1B model. PolyGuard collects an extensive dataset of 1.91M samples of harmful and benign datapoints in 17 languages and uses that to finetune a Qwen-2.5-7B-Instruct model. Safety moderation in MLLMs Relatively few works have tackled detecting harmful prompts in multimodal settings (see Table 1 and Table 2). Chi et al. (2024) propose LlamaGuard3-11B-Vision (a finetuned version of Llama-3-11B-Vision) for de- tecting unsafe inputs in images and texts. Du et al. (2024) and Helff et al. (2025) propose other ap- proaches for the same task. OMNIGUARD achieves higher accuracy in detecting harmful images and prompts compared to these approaches, and to the best of our knowledge is the first guard model for harmful audio inputs. 7 Conclusions We propose OMNIGUARD, an approach for train- ing a safety moderation classifier using the internal representations of an LLM or MLLM that are uni- versally similar across languages and modalities. Our approach consists of two steps: first, we iden- tify these universally similar representations and then we use them to train a harmfulness classi- fier. We find that OMNIGUARD accurately detects harmful prompts across languages, including low- resource languages as well as cipher languages, and also across modalities â images and audios. We show that OMNIGUARD allows to train more effi- cient safety moderation classifiers (both in training time and in inference time) compared to standard guard models, and conclude that our approach is superior in both accuracy and efficiency across lan- guages and modalities. Limitations While OMNIGUARD achieves state-of-the-art per- formance for detecting harmful prompts across lan- guages and modalities, its performance depends on the underlying model. If the underlying model does not understand the language or an image or audio input, OMNIGUARD might not be able to detect if the input is harmful. However, this limitation is not unique to OMNIGUARD, and existing approaches suffer from the same limitation. Our approach also relies on the existence of uni- versally similar representations, which we empiri- cally found to exist across models and modalities. However, we did not exhaustively check all models and this assumption might not hold for models that we have not used in this work. Moreover, OMNI- GUARD requires access to internal representations of a model, making it inapplicable to closed-source models. Lastly, the results we report are based on a fixed set of evaluation datasets that are standard bench- marks used in the research area of AI safety mod- eration. While OMNIGUARD performs well across the datasets we experiment with, its performance in real-world settings might differ. Ethics. While this work seeks to mitigate the risks of LLM deployment in high-risk scenarios, OMNIGUARD is not a perfect classifier and unex- pected failures may allow for the harmful misuse of LLMs. References Aakanksha, Arash Ahmadian, Beyza Ermis, Seraphina Goldfarb-Tarrant, Julia Kreutzer, Marzieh Fadaee, and Sara Hooker. 2024. The multilingual alignment prism: Align- ing global and local preferences to reduce harm. Preprint, arXiv:2406.18682. Gabriel Alon and Michael Kamfonas. 2023.Detect- ing language model attacks with perplexity. Preprint, arXiv:2308.14132. Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. 2024. Jailbreaking leading safety-aligned llms with simple adaptive attacks. Cem Anil, Esin Durmus, ..., and David Duvenaud. 2024. Many-shot jailbreaking. Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. 2024. Re- fusal in language models is mediated by a single direc- tion. Advances in Neural Information Processing Systems, 37:136037â136083. Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, and 12 others. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. Vinamra Benara, Chandan Singh, John X Morris, Richard Antonello, Ion Stoica, Alexander G Huth, and Jianfeng Gao. 2024. Crafting interpretable embeddings by asking llms questions. arXiv preprint arXiv:2405.16714. Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, and 1 others. 2024. Managing extreme ai risks amid rapid progress. Science, 384(6698):842â845. Leonard Bereska and Efstratios Gavves. 2024. Mechanis- tic interpretability for ai safetyâa review. arXiv preprint arXiv:2404.14082. Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Rottger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou. 2024. Safety-tuned LLaMAs: Lessons from improv- ing the safety of large language models that follow instruc- tions. In The Twelfth International Conference on Learning Representations. Rishi Bommasani, Drew A. Hudson, ..., and Percy Liang. 2022. On the opportunities and risks of foundation models. Preprint, arXiv:2108.07258. Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Has- sani, George J. Pappas, and Eric Wong. 2024. Jailbreak- ing black box large language models in twenty queries. Preprint, arXiv:2310.08419. Kai Chen, Chunwei Wang, Kuo Yang, Jianhua Han, Lanqing Hong, Fei Mi, Hang Xu, Zhengying Liu, Wenyong Huang, Zhenguo Li, Dit-Yan Yeung, Lifeng Shang, Xin Jiang, and Qun Liu. 2024a. Gaining wisdom from setbacks: Aligning large language models via mistake analysis. Xuan Chen, Yuzhou Nie, Wenbo Guo, and Xiangyu Zhang. 2024b. When llm meets drl: Advancing jailbreaking effi- ciency via drl-guided search. Preprint, arXiv:2406.08705. Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji, and Ajay Divakaran. 2024c. Dress: Instructing large vision- language models to align and interact with humans via natural language feedback. Preprint, arXiv:2311.10081. Yiming Chen, Xianghu Yue, Chen Zhang, Xiaoxue Gao, Robby T. Tan, and Haizhou Li. 2024d.Voicebench: Benchmarking llm-based voice assistants.Preprint, arXiv:2410.17196. Jianfeng Chi, Ujjwal Karn, Hongyuan Zhan, Eric Smith, Javier Rando, Yiming Zhang, Kate Plawiak, Zacharie Delpierre Coudert, Kartikeya Upasani, and Mahesh Pasupuleti. 2024. Llama guard 3 vision: Safeguarding human-ai image un- derstanding conversations. Preprint, arXiv:2411.10414. cjadams, Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, nithum, and Will Cukierski. 2017. Toxic com- ment classification challenge. Nicholas Kluge CorrĂȘa. 2023. Aira. Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang. 2023. Safe rlhf: Safe reinforcement learning from human feedback. Preprint, arXiv:2310.12773. Adrian de Wynter, Ishaan Watts, Tua Wongsangaroonsri, Minghui Zhang, Noura Farra, Nektar Ege Altıntoprak, Lena Baur, Samantha Claudet, Pavel GajduĆĄek, Qilong Gu, Anna Kaminska, Tomasz Kaminski, Ruby Kuo, Akiko Kyuba, Jongho Lee, Kartik Mathur, Petter Merok, Ivana Milovanovi Ì c, Nani Paananen, and 13 others. 2025. Rtp-lx: Can llms evaluate toxicity in multilingual scenarios? Pro- ceedings of the AAAI Conference on Artificial Intelligence, 39(27):27940â27950. Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira An- derson, Erin Bransom, Kiana Ehsani, Huong Ngo, YenSung Chen, Ajay Patel, Mark Yatskar, Chris Callison-Burch, and 31 others. 2024. Molmo and pixmo: Open weights and open data for state-of-the-art vision-language models. Preprint, arXiv:2409.17146. Daryna Dementieva, Valeriia Khylenko, Nikolay Babakov, and Georg Groh. 2024. Toxicity classification in Ukrainian. In Proceedings of the 8th Workshop on Online Abuse and Harms (WOAH 2024), pages 244â255, Mexico City, Mex- ico. Association for Computational Linguistics. Yihe Deng, Yu Yang, Junkai Zhang, Wei Wang, and Bo Li. 2025. Duoguard: A two-player rl-driven framework for multilingual llm guardrails. Preprint, arXiv:2502.05163. Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing. 2024. Multilingual jailbreak challenges in large language models. In The Twelfth International Conference on Learning Representations. Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang. 2024. A wolf in sheepâs clothing: Generalized nested jailbreak prompts can fool large language models easily. In Proceedings of the 2024 Conference of the North American Chapter of the Asso- ciation for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 2136â2153, Mexico City, Mexico. Association for Computational Lin- guistics. Xuefeng Du, Reshmi Ghosh, Robert Sim, Ahmed Salem, Vitor Carvalho, Emily Lawton, Yixuan Li, and Jack W. Stokes. 2024.Vlmguard: Defending vlms against malicious prompts via unlabeled data. Preprint, arXiv:2410.00296. Qingkai Fang, Shoutao Guo, Yan Zhou, Zhengrui Ma, Shaolei Zhang, and Yang Feng. 2025. Llama-omni: Seamless speech interaction with large language models. Preprint, arXiv:2409.06666. Shaona Ghosh, Prasoon Varshney, Erick Galinkin, and Christo- pher Parisien. 2024a. Aegis: Online adaptive ai content safety moderation with ensemble of llm experts. Shaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreed- har, Aishwarya Padmakumar, Traian Rebedea, Jibin Rajan Varghese, and Christopher Parisien. 2024b. Aegis2. 0: A diverse ai safety dataset and risks taxonomy for alignment of llm guardrails. In Neurips Safe Generative AI Workshop. Yichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang, Tianshuo Cong, Anyu Wang, Sisi Duan, and Xiaoyun Wang. 2025. Figstep: Jailbreaking large vision-language models via typographic visual prompts.Preprint, arXiv:2311.05608. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhi- nav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri. 2024. Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms. Preprint, arXiv:2406.18495. Lukas Helff, Felix Friedrich, Manuel Brack, Kristian Kersting, and Patrick Schramowski. 2025. Llavaguard: An open vlm-based framework for safeguarding vision datasets and models. Preprint, arXiv:2406.05113. Hexgrad. 2025. Kokoro-82m. Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. Preprint, arXiv:2106.09685. Xuhao Hu, Dongrui Liu, Hao Li, Xuanjing Huang, and Jing Shao. 2025. Vlsbench: Unveiling visual leakage in multi- modal safety. Preprint, arXiv:2411.19939. Zhengmian Hu, Gang Wu, Saayan Mitra, Ruiyi Zhang, Tong Sun, Heng Huang, and Viswanathan Swaminathan. 2024. Token-level adversarial prompt detection based on per- plexity measures and contextual information. Preprint, arXiv:2311.11509. Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa. 2023. Llama guard: Llm-based input-output safeguard for human- ai conversations. Preprint, arXiv:2312.06674. Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein. 2023.Baseline defenses for adversarial attacks against aligned language models.Preprint, arXiv:2309.00614. Fengqing Jiang, Zhangchen Xu, Luyao Niu, Zhen Xiang, Bhaskar Ramasubramanian, Bo Li, and Radha Pooven- dran. 2024. Artprompt: Ascii art-based jailbreak attacks against aligned llms. Haibo Jin, Andy Zhou, Joe D. Menke, and Haohan Wang. 2024.Jailbreaking large language models against moderation guardrails via cipher characters. Preprint, arXiv:2405.20413. Erik Jones, Anca Dragan, Aditi Raghunathan, and Jacob Stein- hardt. 2023. Automatically auditing large language models via discrete optimization. In Proceedings of the 40th Inter- national Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 15307â 15329. PMLR. Priyanshu Kumar, Devansh Jain, Akhila Yerukola, Liwei Jiang, Himanshu Beniwal, Thomas Hartvigsen, and Maarten Sap. 2025. Polyguard: A multilingual safety moderation tool for 17 languages. Preprint, arXiv:2504.04377. Raz Lapid, Ron Langberg, and Moshe Sipper. 2024. Open sesame! universal black box jailbreaking of large language models. Preprint, arXiv:2309.01446. Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. 2018. Scalable agent alignment via reward modeling: a research direction.Preprint, arXiv:1811.07871. Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. 2024a. SALAD- bench: A hierarchical and comprehensive safety bench- mark for large language models. In Findings of the Asso- ciation for Computational Linguistics: ACL 2024, pages 3923â3954, Bangkok, Thailand. Association for Computa- tional Linguistics. Mukai Li, Lei Li, Yuwei Yin, Masood Ahmed, Zhenguang Liu, and Qi Liu. 2024b. Red teaming visual language models. Preprint, arXiv:2401.12915. Nathaniel Li, Ziwen Han, Ian Steneker, Willow Primack, Riley Goodside, Hugh Zhang, Zifan Wang, Cristina Menghini, and Summer Yue. 2024c. Llm defenses are not robust to multi-turn human jailbreaks yet. Xiaoxia Li, Siyuan Liang, Jiyi Zhang, Han Fang, Aishan Liu, and Ee-Chien Chang. 2024d. Semantic mirror jailbreak: Genetic algorithm based jailbreak prompts against open- source llms. Preprint, arXiv:2402.14872. Yifan Li, Hangyu Guo, Kun Zhou, Wayne Xin Zhao, and Ji-Rong Wen. 2024e. Images are achillesâ heel of align- ment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models. In ECCV 2024, page 174â189, Berlin, Heidelberg. Springer-Verlag. Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. 2024a. AutoDAN: Generating stealthy jailbreak prompts on aligned large language models. In The Twelfth Interna- tional Conference on Learning Representations. Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, Chao Yang, and Yu Qiao. 2024b. Mm-safetybench: A benchmark for safety evaluation of multimodal large language models. In ECCV 2024, page 386â403, Berlin, Heidelberg. Springer- Verlag. Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao. 2024. Jailbreakv-28k: A benchmark for assessing the robustness of multimodal large language mod- els against jailbreak attacks. Preprint, arXiv:2404.03027. Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. 2024. Harmbench: A standardized evaluation framework for au- tomated red teaming and robust refusal. Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Kar- basi. 2024. Tree of attacks: Jailbreaking black-box llms automatically. Preprint, arXiv:2312.02119. OpenAI. 2025. Openai moderation endpoint guide. Accessed: 2025-05-18. Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Ja- cob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022a. Training language models to fol- low instructions with human feedback. In Proceedings of the 36th International Conference on Neural Informa- tion Processing Systems, NIPS â22, Red Hook, NY, USA. Curran Associates Inc. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Car- roll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Ja- cob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022b. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730â27744. Curran Associates, Inc. Raghuveer Peri, Sai Muralidhar Jayanthi, Srikanth Ronanki, Anshu Bhatia, Karel Mundnich, Saket Dingliwal, Nilaksh Das, Zejiang Hou, Goeric Huybrechts, Srikanth Vishnub- hotla, Daniel Garcia-Romero, Sundararajan Srinivasan, Kyu Han, and Katrin Kirchhoff. 2024. Speechguard: Ex- ploring the adversarial robustness of multi-modal large language models. In Findings of the Association for Com- putational Linguistics: ACL 2024, pages 10018â10035, Bangkok, Thailand. Association for Computational Lin- guistics. Fabio Pernisi, Dirk Hovy, and Paul Röttger. 2024. Compro- messo! italian many-shot jailbreaks undermine the safety of large language models. Qibing Ren, Chang Gao, Jing Shao, Junchi Yan, Xin Tan, Wai Lam, and Lizhuang Ma. 2024. CodeAttack: Revealing safety generalization challenges of large language models via code completion. In Findings of the Association for Computational Linguistics ACL 2024, pages 11437â11452. Association for Computational Linguistics. Mark Russinovich, Ahmed Salem, and Ronen Eldan. 2024. Great, now write an article about that: The crescendo multi- turn llm jailbreak attack. Preprint, arXiv:2404.01833. Guobin Shen, Dongcheng Zhao, Yiting Dong, Xiang He, and Yi Zeng. 2025. Jailbreak antidote: Runtime safety-utility balance via sparse representation adjustment in large lan- guage models. In The Thirteenth International Conference on Learning Representations. Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2024a. "do anything now": Characteriz- ing and evaluating in-the-wild jailbreak prompts on large language models. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, CCS â24, page 1671â1685, New York, NY, USA. Association for Computing Machinery. Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2024b. "do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large lan- guage models. Preprint, arXiv:2308.03825. Chandan Singh, Armin Askari, Rich Caruana, and Jianfeng Gao. 2023. Augmenting interpretable models with large language models during training. Nature Communications, 14(1):7913. Sugan Sirihattasak, Mamoru Komachi, and Hiroshi Ishikawa. 2018. Annotation and classification of toxicity for thai twitter. In Proceedings of LREC 2018 Workshop and the 2nd Workshop on Text Analytics for Cybersecurity and Online Safety (TA-COSâ18), Miyazaki, Japan. Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic composi- tionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Lan- guage Processing, pages 1631â1642, Seattle, Washington, USA. Association for Computational Linguistics. Jiayang Song, Yuheng Huang, Zhehua Zhou, and Lei Ma. 2024. Multilingual blending: Llm safety alignment evalua- tion with language mixture. Preprint, arXiv:2407.07342. NLLB Team, Marta R. Costa-jussĂ , James Cross, Onur Ăelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Young- blood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, and 20 others. 2022. No language left behind: Scaling human-centered machine translation. Preprint, arXiv:2207.04672. Bibek Upadhayay and Vahid Behzadan. 2024. Sandwich attack: Multi-language mixture adaptive attack on llms. Preprint, arXiv:2404.07242. Bertie Vidgen, Nino Scherrer, Hannah Rose Kirk, Rebecca Qian, Anand Kannappan, Scott A. Hale, and Paul Röttger. 2024. Simplesafetytests: a test suite for identifying crit- ical safety risks in large language models.Preprint, arXiv:2311.08370. Jiongxiao Wang, Zichen Liu, Keun Hee Park, Zhuojun Jiang, Zhaoheng Zheng, Zhuofeng Wu, Muhao Chen, and Chaowei Xiao. 2023. Adversarial demonstration attacks on large language models. Preprint, arXiv:2305.14950. Wenxuan Wang, Zhaopeng Tu, Chang Chen, Youliang Yuan, Jen tse Huang, Wenxiang Jiao, and Michael R. Lyu. 2024a. All languages matter: On the multilingual safety of large language models. Yu Wang, Xiaofei Zhou, Yichen Wang, Geyuan Zhang, and Tianxing He. 2024b.Jailbreak large vision- language models through multi-modal linkage. Preprint, arXiv:2412.00473. Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: How does LLM safety training fail? In Thirty- seventh Conference on Neural Information Processing Sys- tems. Zeming Wei, Yifei Wang, Ang Li, Yichuan Mo, and Yisen Wang. 2024. Jailbreak and guard aligned language mod- els with only few in-context demonstrations. Preprint, arXiv:2310.06387. Chris Wendler, Veniamin Veselovsky, Giovanni Monea, and Robert West. 2024. Do llamas work in english? on the latent language of multilingual transformers. Zhaofeng Wu, Xinyan Velocity Yu, Dani Yogatama, Jiasen Lu, and Yoon Kim. 2024. The semantic hub hypothesis: Language models share semantic representations across languages and modalities. Preprint, arXiv:2411.04986. Hao Yang, Lizhen Qu, Ehsan Shareghi, and Gholamreza Haf- fari. 2025. Audio is the achillesâ heel: Red teaming audio large multimodal models. In Proceedings of the 2025 Con- ference of the Nations of the Americas Chapter of the Asso- ciation for Computational Linguistics, pages 9292â9306, Albuquerque, New Mexico. Yahan Yang, Soham Dan, Dan Roth, and Insup Lee. 2024. Benchmarking llm guardrails in handling multilingual toxi- city. Preprint, arXiv:2410.22153. Zonghao Ying, Aishan Liu, Siyuan Liang, Lei Huang, Jinyang Guo, Wenbo Zhou, Xianglong Liu, and Dacheng Tao. 2024. Safebench: A safety evaluation framework for multimodal large language models. Preprint, arXiv:2410.18927. Zheng-Xin Yong, Cristina Menghini, and Stephen H. Bach. 2024. Low-resource languages jailbreak gpt-4. Preprint, arXiv:2310.02446. Haneul Yoo, Yongjin Yang, and Hwaran Lee. 2024. Csrt: Evaluation and analysis of llms using code-switching red- teaming dataset. Preprint, arXiv:2406.15481. Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing. 2024a. Gptfuzzer: Red teaming large language models with auto- generated jailbreak prompts. Preprint, arXiv:2309.10253. Weihao Yu, Zhengyuan Yang, Lingfeng Ren, Linjie Li, Jian- feng Wang, Kevin Lin, Chung-Ching Lin, Zicheng Liu, Lijuan Wang, and Xinchao Wang. 2024b. Mm-vet v2: A challenging benchmark to evaluate large multimodal mod- els for integrated capabilities. Preprint, arXiv:2408.00765. Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. 2024a. GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher. In The Twelfth International Conference on Learning Representations. Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Jiahao Xu, Tian Liang, Pinjia He, and Zhaopeng Tu. 2024b. Refuse whenever you feel unsafe: Improving safety in llms via decoupled refusal training. Weixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng, An Zhang, Xingyu Sui, Xinyang Han, Yanyan Zhao, Bing Qin, Tat- Seng Chua, and Ting Liu. 2025. Adasteer: Your aligned llm is inherently an adaptive jailbreak defender. Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi, and Lidong Bing. 2024. How do large lan- guage models handle multilingualism? Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Bar- row, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun. 2023. Autodan: Interpretable gradient-based adversarial attacks on large language models. Preprint, arXiv:2310.15140. Yufan Zhuang, Chandan Singh, Liyuan Liu, Jingbo Shang, and Jianfeng Gao. 2025. Vector-icl: In-context learning with continuous vector representations. In The Thirteenth International Conference on Learning Representations. Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. 2023a. Universal and trans- ferable adversarial attacks on aligned language models. Preprint, arXiv:2307.15043. Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson. 2023b. Universal and transferable adversarial attacks on aligned language models. Preprint, arXiv:2307.15043. A Languages Used in Our Approach We use the following languages in our experiments: 1.Natural Languages: English, French, German, Spanish, Persian, Arabic, Croatian, Japanese, Polish, Russian, Swedish, Thai, Hindi, Ital- ian, Korean, Bengali, Portuguese, Chinese, Hebrew, Serbian, Danish, Turkish, Greek, In- donesian, Zulu, Hungarian, Basque, Swahili, Afrikaans, Bosnian, Lao, Romanian, Slove- nian, Ukrainian, Finnish, Malay, Javanese, Welsh, Bulgarian, Armenian, Icelandic, Viet- namese, Sinhalese, Maori, Gujarati, Kannada, Marathi, Tamil, Telugu, Amharic, Norwegian, Czech, Dutch. 2.Cipher Languages: Caesar1, Caesar2, Cae- sar3, Caesar4, Caesar5, Caesar6, Cae- sar7, Caesarneg1, Caesarneg2, Caesarneg3, Caesarneg4, Caesarneg5, Caesarneg6, Cae- sarneg7, Ascii, Hexadecimal, Base64, Leet, Vowel, Alphanumeric. A number in front of Caesar cipher means that the English alpha- bets were shifted by that much forward and a number in front of Caesarneg cipher means that the English alphabets were shifted by that much backward. Out of these languages, we use the following for training our classifier: Arabic, Chinese, Czech, Dutch, English, French, German, Hindi, Italian, Japanese, Korean, Polish, Portuguese, Russian, Spanish, Swedish, Thai, Bosnian, Turkish, Finnish, Indonesian, Bengali, Swahili, Vietnamese, Tamil, Telugu, Greek, Maori, Javanese, Caesar1, Caesar2, Caesar4, Caesarneg2, Caesarneg4, Caesarneg6, Ascii, Hexadecimal And these for testing: Persian, Croatian, He- brew, Serbian, Danish, Zulu, Hungarian, Basque, Afrikaans, Lao, Romanian, Slovenian, Ukrainian, Malay, Welsh, Bulgarian, Armenian, Icelandic, Sin- halese, Gujarati, Kannada, Marathi, Amharic, Nor- wegian, Caesar, Caesar5, Caesar7, Caesarneg3, Caesarneg1, Caesar6, Caesarneg7, Caesarneg5, Base64, Alphanumeric, Vowel, LeetSpeak. B Datasets and models C Experimental Details of Filtering of Wikitext and its Translation DOMNIGUARDâs Performance With Different Base LLMs High-ResLow-ResCipher LlamaGuard 169.9241.2516.07 LlamaGuard 275.2862.216.2 LlamaGuard 382.2375.8424.74 AegisGuard (P) 83.3659.0644.22 AegisGuard (D)88.1476.2683.21 WildGuard81.3543.5116.53 HarmBench (llama)14.2514.0814.11 HarmBench (mistral) 17.615.914.46 MD-Judge59.5130.2115.44 DuoGuard71.446.2615.77 PolyGuard94.4779.2221.28 OMNIGUARD 88.2585.5673.06 Table B1: Accuracy of detecting harmful prompts stratified by high-resource natural, low-resource natural, and cipher languages. Thai ToxUkr ToxTJSToxTextAvg. Layer 1062.165.566.9561.8964.42 Layer 5567.473.076.9174.9673.06 Layer 5666.871.574.5471.8271.17 Selected layer 5768.773.176.875.0773.40 Layer 5866.673.274.9272.6371.84 Layer 5967.373.376.4474.22 72.82 Layer 6067.572.676.4374.4672.75 Layer 6167.870.974.7772.7671.56 Layer 6266.172.374.7872.83 71.50 Layer 7565.266.470.7265.7968.26 Last Layer63.151.261.3356.7659.05 Table B2: OMNIGUARDâs accuracy of detecting harmful prompts when trained using representations from different model layers. ModelsHarmBenchFQSimpleSTSaladBTJSToxTextAdvBenchAvg. Llama3.1-8B-Instruct81.1688.7288.7887.3571.3269.2088.6282.16 Gemma-3-4B-Instruct78.0388.2393.8682.0953.5051.4589.0776.60 Qwen-3-4B-Instruct82.0686.2580.9982.1558.8357.7084.8076.11 Olmo-2-7B-Instruct80.5788.8692.1488.6870.6564.7988.2481.99 Mistral-8B-Instruct84.2787.2389.8288.0872.4968.7189.0782.81 Llama3.3-70B-Instruct89.1389.5789.6287.3076.6875.0786.5984.85 Table B3: OMNIGUARDâs accuracy of detecting harmful prompts when paired with different underlying LLMs across multiple benchmarks.