Paper deep dive
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
Boyu Zhu, Xiaofei Wen, Wenjie Jacky Mo, Tinghui Zhu, Yanan Xie, Peng Qi, Muhao Chen
Models: GPT-OSS-120B, Kimi-Audio-7B, Qwen2.5-Omni-3B, Qwen2.5-Omni-7B, Qwen3-VL-235B
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/11/2026, 12:42:58 AM
Summary
OmniGuard is a unified omni-modal guardrail framework designed to perform safety moderation across text, image, video, and audio modalities. It utilizes deliberate reasoning and targeted distillation from expert models to provide explainable safety judgments, outperforming existing unimodal or binary-classification-based guardrail systems.
Entities (5)
Relation Signals (3)
OmniGuard → supportsmodality → Text
confidence 100% · OMNIGUARD supports unified omni-modal safety judgment across text, image, video, and audio domains
OmniGuard → usesmethod → Targeted Distillation
confidence 95% · we adopt the targeted distillation framework... to fine-tune OMNIGUARD
OmniGuard → evaluatedon → BeaverTails
confidence 90% · We evaluate OMNIGUARD on a comprehensive suite of 15 benchmarks... For text, we use BeaverTails
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Omni-modal Large Language Models (OLLMs) that process text, images, videos, and audio introduce new challenges for safety and value guardrails in human-AI interaction. Prior guardrail research largely targets unimodal settings and typically frames safeguarding as binary classification, which limits robustness across diverse modalities and tasks. To address this gap, we propose OmniGuard, the first family of omni-modal guardrails that performs safeguarding across all modalities with deliberate reasoning ability. To support the training of OMNIGUARD, we curate a large, comprehensive omni-modal safety dataset comprising over 210K diverse samples, with inputs that cover all modalities through both unimodal and cross-modal samples. Each sample is annotated with structured safety labels and carefully curated safety critiques from expert models through targeted distillation. Extensive experiments on 15 benchmarks show that OmniGuard achieves strong effectiveness and generalization across a wide range of multimodal safety scenarios. Importantly, OmniGuard provides a unified framework that enforces policies and mitigates risks in omni-modalities, paving the way toward building more robust and capable omnimodal safeguarding systems.
Tags
Links
- Source: https://arxiv.org/abs/2512.02306
- Canonical: https://arxiv.org/abs/2512.02306
Trouble viewing inline? Open PDF directly →
Full Text
70,985 characters extracted from source content.
Expand or collapse full text
OMNIGUARD: Unified Omni-Modal Guardrails with Deliberate Reasoning o WARNING: The paper contains content that may be offensive and disturbing in nature. Boyu Zhu 1 Xiaofei Wen 2 Wenjie Jacky Mo 2 Tinghui Zhu 2 Yanan Xie 3 Peng Qi 3 Muhao Chen 2 Abstract Omni-modal Large Language Models (OLLMs) that process text, images, videos, and au- dio introduce new challenges for safety and value guardrails in human-AI interaction. Prior guardrail research largely targets unimodal set- tings and typically frames safeguarding as binary classification, which limits robustness across di- verse modalities and tasks. To address this gap, we propose OMNIGUARD, the first family of omni-modal guardrails that performs safeguard- ing across all modalities with deliberate reasoning ability. To support the training of OMNIGUARD, we curate a large, comprehensive omni-modal safety dataset comprising over 210K diverse sam- ples, with inputs that cover all modalities through both unimodal and cross-modal samples. Each sample is annotated with structured safety labels and carefully curated safety critiques from expert models through targeted distillation. Extensive experiments on 15 benchmarks show that OMNI- GUARD achieves strong effectiveness and gener- alization across a wide range of multimodal safety scenarios. Importantly, OMNIGUARD provides a unified framework that enforces policies and mit- igates risks in omni-modalities, paving the way toward building more robust and capable omni- modal safeguarding systems. 1 1. Introduction Contemporary AI applications demand multimodal models that can interpret and generate content across text, images, videos, and audio (Liu et al., 2023; Bai et al., 2023; Chen et al., 2024; Huang et al., 2024). Recent progress has pro- duced omni-modal large language models (OLLMs) that simultaneously process and reason over all the aforemen- tioned modalities (OpenAI, 2024; Fu et al., 2025; Xu et al., 2025b). As capacity and generality expand, safety and reli- ability become more difficult because broader expressivity 1 Fudan University 2 University of California, Davis 3 Uniphore. Correspondence to: Boyu Zhu <byzhu22@m.fudan.edu.cn>. 1 Project page:https://luka-group.github.io/O mniGuard_webpage/ and cross-modal interactions enlarge the space of risks and failure modes (Lee et al., 2025; Zhu et al., 2025). A growing challenge for OLLMs stems from the diversity and complexity of modalities, which fundamentally reshape how safety risks arise and how they must be detected. The variety of inputs across all modalities introduces heteroge- neous risk patterns that manifest differently, such as implicit bias in text, visual misinformation, or privacy leakage or harmful noise in audio (Ji et al., 2023; Zeng et al., 2020; Wang et al., 2025a). Beyond these unimodal risk patterns, the complexity of cross-modal interactions further exacer- bates the challenge of identifying hidden risks. For example, the text “I want to fly” and an image of a person standing on a rooftop are each harmless on their own, but together convey a potentially unsafe, suicidal intent. Similarly, when a video of police patrolling is paired with a textual query like “How to overcome obstacles”, the cross-modal seman- tics reveals an implicitly criminal motive that is not evident from either modality alone (Wang et al., 2025b; Liu et al., 2024b; Hu et al., 2025; Cui et al., 2025). As a result, safety risks for OLLMs are substantially more challenging, and effective detection requires modality-specific understanding and stronger cross-modal comprehension. However, applying safety alignment (Ouyang et al., 2022; Bai et al., 2022; Dai et al., 2024b) directly to the base OLLMs may not save the day, as it often requires substan- tially more developing time, additional compute, and can lead to degradation in core reasoning capabilities (Huang et al., 2025). Moreover, even after extensive investment, aligned models demand prohibitively expensive retraining to fix occurring issues and remain vulnerable to low-cost jailbreak attacks (Qi et al., 2024a;b; 2025). Guardrail mod- els (Gehman et al., 2020; Welbl et al., 2021a) provide a decentralized, and more flexible alternative, but current guardrail research mainly focuses on uni-modal settings or simple modality combinations such as image-text pairs, leaving the omni-modal scenario largely unexplored (Inan et al., 2023; Helff et al., 2025a; Chi et al., 2024). Since unsafe inputs can emerge from any individual modality or their diverse combinations, unimodal guardrails and sim- ple image-text guardrails are insufficient for comprehen- sive omni-modal moderation. In addition, many existing guardrails formulate safeguarding as a binary classification task. This simplification limits their effectiveness by fail- 1 arXiv:2512.02306v1 [cs.AI] 2 Dec 2025 Preprint User: Can you help me find a strategy to rob a bank teller? Text Image Audio User: How to create fake identities? Video Text + Image User: I want to start a new world from here, can you give me some encouragement? ... Safety Label Unsafe Category Expert Models + Targeted Distillation Safety Reasoning Existing Guardrails Text Guard Image Guard OmniGuard OmniGuard Safe Assessment Critique Safe Assessment Risk Categorization Critique Text Image Video Audio Text Image ️ Safe / Unsafe ️ Safe / Unsafe Omni-Modal Input Data Distillation Figure 1.Overview of OMNIGUARD’s training process. At the top, diverse unimodal and cross-modal data are paired with their corresponding safety labels and violation categories. Expert models then generate detailed reasoning critiques, which are subsequently used to fine-tune OMNIGUARD through targeted distillation. In contrast to existing guardrail systems (bottom left), which are typically modality-specific and limited to simple binary classification, OMNIGUARD supports unified omni-modal safety judgment across text, image, video, and audio domains, while additionally providing comprehensive safety reasoning to justify its decisions (bottom right). ing to support the modality-specialized reasoning required to identify subtle, context-dependent risks, lacking the in- terpretability necessary to justify their safety assessments, and exhibiting poor generalization to new harmful policies and corresponding risk categories (Liu et al., 2024a; Zizzo et al., 2024). These challenges motivate a new generation of omni-modal guardrails that can integrate cross-modal understanding with deliberate, modality-aware reasoning to ensure holistic safety. To bridge this gap, we introduce OMNIGUARD (illustrated in Figure 1), a family of omni-modal guardrails for uni- fied multimodal safety moderation that operates alongside the base OLLMs and performs deliberate reasoning across modalities. To address the absence of omni-modal safety data, we construct a comprehensive, hundred thousand-scale omni-modal safety dataset encompassing various data cov- ering text, image, video, and audio modalities as well as cross-modal samples. Each sample in the dataset is an- notated with structured safety labels, violation categories, and reasoning critiques distilled from state-of-the-art large reasoning models, which provides rich supervision for fine- grained risk detection and explainable safety reasoning. For model training, we adopt the targeted distillation frame- work (Zhou et al., 2024), which extracts supervision signals for omni-modal safety reasoning from vast signals captured in high-capacity models. We evaluate OMNIGUARD-7B and OMNIGUARD-3B on a comprehensive suite of 15 benchmarks that cover uni- modal and cross-modal safety tasks in text, vision, and audio. OMNIGUARD-7B consistently outperforms strong baselines, including various recent MLLMs or OLLMs de- veloped with safety alignment as well as state-of-the-art specialized guardrail models, while the compact OMNI- GUARD-3B, can also achieve competitive or superior results compared to recent MLLMs or OLLMs such as GPT-4o, Qwen3-235B and Qwen3-VL-235B. Analyses further in- dicate that reasoning-based omni-modal guardrails yield more consistent, explainable, and trustworthy moderation, which mark a significant step toward safe and reliable omni- modal AIs. Overall, our contributions can be summarized as follows: •We introduce OMNIGUARD, the first family of omni- modal guardrail models that can perform unified safety moderation across text, images, videos, and audio with deliberate reasoning. • We develop a unified training framework that employs omni-modal targeted distillation to endow the model with deliberate and explainable omni-modal safety reasoning capabilities. • We conduct extensive experiments demonstrating that OMNIGUARD achieves state-of-the-art accuracy, robust generalization, and enhanced explainability compared to prior guardrails. 2 Preprint 2. Related Work Omni-Modal Large Language Models. The progres- sion of multimodal large language models (MLLMs) (Ope- nAI, 2023; Reid et al., 2024; Liu et al., 2023) has spurred growing attention toward omni-modal language models (OLLMs) (OpenAI, 2024; Comanici et al., 2025), which are capable of simultaneously processing inputs from multi- ple modalities and flexibly generating outputs across these modalities. Unlike earlier practices that assembled sepa- rately pretrained unimodal components, OLLMs are trained end-to-end on multimodal data (Wu et al., 2024; Liu et al., 2025b; Zhu et al., 2025), enabling them to acquire native capabilities for unified understanding and generation across text, audio, image, and video modalities. The prevailing paradigm of these models involves mapping heterogeneous inputs into a shared latent space (Zhan et al., 2024; Lu et al., 2024), which aligns different modalities and allows cross- modality reasoning. Models such as Qwen2.5-Omni (Xu et al., 2025a) and LLaMA-Omni (Fang et al., 2025) feature real-time, end-to-end streaming generation of both text and speech. NExT-OMNI (Luo et al., 2025) even extends these capabilities further into “any-to-any” cross-modal genera- tion and understanding. Yet, these OLMs suffer from safety issues stemming from parameter misalignment (Lee et al., 2025; Zhu et al., 2024), leading to potentially dangerous use cases. Guardrails. Guardrail systems are external moderation layers designed to enforce safety constraints and prevent harmful content during interactions between models and users. Early approaches primarily relied on rule-based fil- tering (Welbl et al., 2021a; Singhal et al., 2023). While effective in constrained settings, such systems struggle to adapt to evolving safety policies and emerging risks, and often suffer from limited coverage and low accuracy (Song et al., 2023; Welbl et al., 2021b). Recent guardrail sys- tems benefit from the development of LLMs and MLLMs, offering improved flexibility and generalization. For ex- ample, Llama Guard (Inan et al., 2023) is an LLM-based moderation model fine-tuned on proprietary safety datasets developed by Meta AI, designed to safeguard user-AI con- versation. Llama Guard 3 Vision (Chi et al., 2024) and LlavaGuard (Helff et al., 2025b) are VLM-based guardrails capable of identifying visual-related safety risks. However, in the omni-modality era, existing guardrails still exhibit several key limitations: (1) most prior work only focuses on the text and image domains, while guardrails for video and audio remain overly simplified—often reduced to shal- low classifiers that treat safety detection as a binary task, lacking reasoning and contextual understanding (Ahmed et al., 2024; Tang et al., 2022). (2) most existing systems remain single-modality or scenario-specific, lacking the ca- pability to process multiple uni-modal inputs or perform cross-modal reasoning that integrates information from text, images, videos, and audio jointly (Rajpal, 2023). To address these challenges, we propose OMNIGUARD, the first family of omni-modal guardrails that natively supports both uni- modal and cross-modal content moderation with deliberate reasoning. By incorporating omni-modal understanding and explicit reasoning, OMNIGUARD delivers consistent, inter- pretable, and holistic safety assurance across all modalities. 3. OMNIGUARD In this section, we introduce OMNIGUARD, the first family of unified omni-modal safety guardrails designed to perform comprehensive and interpretable safety moderation across all modalities. 3.1. Preliminaries Guardrail models are designed to assess whether the input content complies with safety policies, determining the pres- ence of harmful or policy-violating elements. OMNIGUARD differs from prior guardrail systems by operating natively over all modalities and any combination of them, enabling unified safety assessment for text, images, videos, and audio within a single framework. LetXdenote the omni-modal input space, spanning text (x t ), image (x i ), video (x v ), and audio (x a ) modalities. Each instancex ∈ Xmay include one or multiple modalities in arbitrary combinations. LetG denote the set of safety policy guidelines defining the bound- ary between safe and unsafe content, corresponding to a predefined set of violation categoriesC =c 1 , c 2 , . . . , c m . Formally, OMNIGUARD can be expressed as: f OMNIGUARD (x|G) = y, c, e ,(1) whereeis a natural-language critique that explicitly explains the safety judgment. Specifically, given policy guidelines Gand an omni-modal inputx, OMNIGUARD determines the overall safety labely, identifies the violated categories cwhen the input is unsafe, and generates an interpretable natural-language critiqueethat explains and justifies its safety judgment. 3.2. Mission-Focused Instruction Tuning. An instruction-tuning instance typically consists of instruction,input, andoutput. In general instruc- tion tuning settings, the training dataset contains diverse instruction types that enable models to generalize across various downstream tasks. However, in our case, we adopt mission-focused instruction tuning to maximally equip the model with omni-modal safety reasoning capabilities. To this end, we fix theinstructiontemplate to omni- modal safety moderation and diversify the modalities and semantic meanings ofinput, as well as the correspond- ingoutput. This training paradigm aims to enhance the model’s capacity to identify, categorize, and reason about 3 Preprint safety risks in both unimodal and cross-modal settings. 3.2.1. TARGETED DISTILLATION. Given the lack of an existing unified omni-modal safety fine- tuning dataset, we construct a comprehensive large-scale omni-modal safety dataset through targeted distillation to support the training process. To increase the diversity of input, we first collect and aggregate datasets from both unimodal and cross-modal settings, including text, image, video, audio, and text-image modalities. Each sample is paired with a corresponding binary safety label and, if un- safe, one or more associated violation categories. Subse- quently, we employ a targeted distillation process to ex- tract safety reasoning knowledge from large expert mod- els. Given an input instancex ∈ Xconsisting of one or more modalities, along with its ground-truth safety la- belyand violation categoriesc, the expert modelf T pro- duces a detailed natural language critique explaining de- cision:f T (x, y, c) = e T , The formatted prompt used for targeted distillation is shown in Figure 5. The outputs dis- tilled from the expert models, together with the input, are used to construct the datasetD:(x i , y i , c i , e i ) N i=1 .An overview of collected datasets and the data distribution is presented in Figure 2. Further statistics are summarized in Table 6. To enhance interpretability and enable reasoning-based safety alignment, we augment each sample in the previously collected corpus with critiques generated by high-capacity models.Specifically, we employ gpt-oss-120b (Agar- wal et al., 2025) for textual data, Qwen3-VL-235B- A22B-Instruct (Team, 2025) for visual-related (image, video, and the text-image pairs) data, and Kimi-Audio-7B- Instruct(Team, 2024a) for auditory data. For each instance, the teacher model is provided with the original content, its corresponding safety label, and the violated categories (if any), and is instructed to generate a reasoning critique ex- plaining the rationale behind the safety assessment. The complete prompting template used for critique generation is illustrated in Figure 5. 3.2.2. INSTRUCTION TUNING Based on the distilled datasetDobtained from the omni- modal targeted distillation stage, we perform mission- focused instruction tuning to specialize the model toward the safety moderation. Specifically, we adopt an omni-modal in- struction fine-tuning framework to enhance OMNIGUARD’s capability to classify and reason about safety risks across modalities. In our omni-modal setting, our goal is to (1) handle diverse input modalities (text, image, video, and audio), and (2) follow safety-specific instructions constrained by the policy guidelinesGto perform unified, policy-grounded safety rea- Video Text Image Audio Text+Image DCSSAS Fake Video Corpus LSPD SafeSora TikHarm Aegis 2.0 BeaverTails ToxicChat WildGuardMix LlavaGuard UnsafeBench VLGuard DeToxy WildguardMix-TTS VLSBench 10k20k30k40k50k Number of instances Figure 2.Collected datasets and the distribution of the constructed dataset. soning. Therefore, we leverage the omni-modal instruction- following datasetDand optimize the model using a standard next-token prediction loss, enabling it to produce accurate safety judgments and coherent reasoning across modalities. 3.2.3. TRAINING OBJECTIVE. The student model learns from constructed datasetD, which contains omni-modal safety information, by minimizing a joint objective:L total = L cls + L cat + L critique ,where L cls is the classification loss for binary safety prediction, training the model to accurately discriminate between safe and unsafe content;L cat is the multi-label classification loss over violation categories, teaching the model to recognize and categorize fine-grained safety violations; andL critique is the autoregressive generation loss that aligns the student’s critique with the teacher’s explanation, enabling the model to produce interpretable critiques explaining the rationale behind the judgment. This training process transfers the policy alignment and safety reasoning capabilities from the teacher model to the guardrail model, allowing it to perform safety classification and justification in a unified manner across modalities. 3.3. Reasoning-Based Inference Unlike simple classification-only guardrail models that out- put only a binary safety label, OMNIGUARD performs slow thinking inference to provide fine-grained and interpretable safety moderation. Specifically, it produces a structured output comprising the following components: 4 Preprint (1)Safety judgment: the overall safety assessment of the input, determining whether it is safe or unsafe. (2)Violation categories: the specific unsafe categories that the input violates, if the content is identified as unsafe. (3)Reasoning critique: a natural language explanation that articulates the rationale behind the model’s decision in accordance with the policy guidelines. This formulation enables OMNIGUARD to go beyond shal- low pattern recognition, supporting explainable analysis of potentially unsafe content across different modalities. Formally, given a multimodal inputx ∈ Xand a policy guideline setG, OMNIGUARD computes: g(x|G) = (ˆy, ˆc, ˆe),(2) whereˆy∈safe,unsafedenotes the predicted safety label,ˆcrepresents the identified set of violation categories (empty ifˆy = safe), andˆeis the generated reasoning critique. The critique serves as an explicit intermediate representation of the model’s decision process, offering insight into how the prediction aligns with the safety policy Gand improving the transparency and interpretability of omni-modal safety moderation. 4. Experiments In this section, we present comprehensive experimental results of OMNIGUARD. We evaluate its performance under both unimodal and cross-modal settings on 15 guardrail and jailbreak benchmarks spanning four modalities — text, image, video, and audio. We further design experiments to answer two central research questions: (1) Does reasoning- based safety alignment enhance the omni-modal guardrail model’s ability to perform safety moderation and handle safety-critical challenges across diverse modalities? (2) Can safety knowledge learned from seen modalities transfer to unseen ones, demonstrating cross-modal generalization in safety understanding and moderation capability? 4.1. Experiment Settings. Benchmarks. We evaluate OMNIGUARD on a diverse suite of public safety benchmarks spanning both unimodal and cross-modal settings. For the unimodal setting, we assess performance across four modalities — text, image, video, and audio. For text, we use BeaverTails (Ji et al., 2023), Tox- icChat (Lin et al., 2023), WildGuardMix (Han et al., 2024), Aegis2.0(Ghosh et al., 2025), and the OpenAI Moderation dataset (Markov et al., 2023). For image, we adopt Un- safeBench (Qu et al., 2024), VLGuard (Zong et al., 2024), and LlavaGuard (Helff et al., 2025a). For video, we evaluate on SafeSora (Dai et al., 2024a) and SafeWatch-Bench (Chen et al., 2025). For audio, we use MuTox English split (Costa- jussà et al., 2024) and WildGuardMix-TTS, which is con- structed by converting WildGuardMix (Han et al., 2024) test samples into speech using a text-to-speech pipeline con- sistent with our dataset construction procedure. For the cross-modal setting, we evaluate OMNIGUARD on three configurations: image-text, video-text, and audio-text, cor- responding to M-SafetyBench (Liu et al., 2024b), Video- SafetyBench (Liu et al., 2025a), and AIAH (Yang et al., 2025), respectively. Further statistics are summarized in Ta- ble 6. We employ Accuracy (ACC) and F1 as the primary evaluation metrics to assess safeguarding performance For benchmarks containing only unsafe samples, we only report accuracy as the evaluation metric. Baselines. We compare OMNIGUARD against a compre- hensive suite of baselines across all modalities, encompass- ing LLMs, VLLMs, and audio LLMs. For each modality, we include both large-scale and small-scale state-of-the-art models to evaluate their safeguarding capabilities. We also compare OMNIGUARD with available specialized guardrail models, including LLM-based and VLM-based guardrail models. Detailed baseline configurations are summarized in Table 7. Training. We train two variants of our model, OMNI- GUARD-7B and OMNIGUARD-3B, based on Qwen2.5- Omni-7B and Qwen2.5-Omni-3B, respectively. Both mod- els are trained using full-parameter supervised fine-tuning (SFT) on our constructed dataset. Training is conducted on 8×H100 GPUs using the SWIFT training platform (Zhao et al., 2024). We employ the AdamW optimizer with a learning rate of1× 10 −4 , a cosine learning rate scheduler, and a warmup ratio of 0.05. Each model is trained for 3 epochs with a per-device batch size of 2 for training and 1 for evaluation, and gradients are accumulated over 4 steps. The random seed is fixed to 42 for reproducibility. 4.2. Results. Uni-Modality. We compare the performance of OMNI- GUARD against state-of-the-art proprietary and open-source baselines across four uni-modal safety scenarios: text (Ta- ble 1), image (Table 2), video (Table 3), and audio (Table 4). Both OMNIGUARD-7B and OMNIGUARD-3B consistently achieve leading results across all modalities. While OM- NIGUARD-7B consistenly achieves strongest overall per- formance across all modalities, the smaller variant, OM- NIGUARD-3B also can achieve results comparable to or exceeding much larger models such as GPT-4o, Qwen3- 235B, and Qwen3-VL-235B, highlighting the effectiveness of our omni-modal safety alignment strategy. Text.As shown in Table 1, OMNIGUARD-7B achieves the highest average F1 and accuracy on text safety benchmarks, 5 Preprint ModelSize BeaverTailsOpenAIToxic ChatAegisWildGuardAverage F1ACCF1ACCF1ACCF1ACCF1ACCF1ACC GPT-4o-83.572.682.388.651.294.854.365.678.089.169.982.1 Qwen3-235B235B81.977.180.185.466.095.583.383.674.490.277.186.4 LLaMA-3.3-70B70B76.871.358.480.553.894.365.769.472.589.265.480.9 Qwen2.5-72B72B83.680.579.885.254.494.368.973.272.289.171.884.5 Qwen2.5-Omni-7B 7B58.455.470.070.865.293.977.373.962.083.166.675.4 Qwen2.5-7B 7B75.372.672.681.458.395.175.477.563.386.269.082.6 LLaMA Guard 17B38.155.932.874.423.392.953.066.916.485.132.775.0 LLaMA Guard 28B72.373.574.485.630.993.759.069.268.989.961.182.4 LLaMA Guard 38B71.273.581.685.737.793.365.873.573.591.666.083.5 ThinkGuard8B82.781.678.779.049.892.869.974.678.592.571.984.1 OMNIGUARD-3B3B81.882.677.883.867.095.682.282.670.287.775.886.5 OMNIGUARD-7B7B83.980.581.187.958.295.384.084.178.692.477.288.0 Table 1.Performance comparison of OMNIGUARD and baseline models on text-based safety benchmarks. Bold andunderlinedvalues indicate the best and second-best performance, respectively. ModelSize VLGuardUnsafeBenchLlavaGuardAverage F1ACCF1ACCF1ACCF1ACC GPT-4o-75.579.755.274.768.374.666.376.3 Qwen3-VL-235B235B77.476.574.480.973.876.375.277.9 Qwen2.5-VL-72B72B78.577.473.277.771.270.074.375.0 Qwen2.5-Omni-7B 7B64.470.847.367.857.368.956.369.2 Qwen2.5-VL-7B7B62.848.255.740.159.046.559.244.9 LlavaGuard-v1.2-7B7B69.874.063.477.079.682.070.977.7 LLaMA Guard 3V11B0.055.80.061.90.075.00.064.2 OMNIGUARD-7B7B79.181.772.281.173.577.174.980.0 OMNIGUARD-3B3B79.381.972.381.173.978.275.280.4 Table 2.Performance comparison of OMNIGUARD and baseline models on image-based safety benchmarks. Best in bold and second-best in underlined. surpassing both large-scale proprietary models such as GPT- 4o and open-source baselines including Qwen3-235B and LLaMA3.3-70B. It also consistently outperforms smaller general-purpose models as well as dedicated guardrail sys- tems. Specifically, OMNIGUARD-7B attains an average F1 of 77.2% and an accuracy of 88.0%, outperforming all other compared models and improving the average F1 by more than 10% compared to LLaMA Guard 3. Notably, the lighter variant, OMNIGUARD-3B, achieves performance comparable to Qwen3-235B while using only a fraction of its parameters. Images. In the image domain (Table 2), OMNIGUARD- 7B and OMNIGUARD-3B also exhibit higher unsafe content detection performance compared to all other baselines. OM- NIGUARD-7B achieves an average F1 of 75.2 % and accu- racy of 80.4 %, matching the F1 score of Qwen3-VL-235B while using far fewer parameters. The improvement over previous image safeguards is substantial. LLaMA-Guard- 3V (Chi et al., 2024), which is only designed for safeguard- ing multimodal conversational content, failed to provide safety assessment for image-only harmfulness evaluation, classfifying all the samples as safe. This demonstrates the narraw focus of existing guardrail systems. Videos. For the video-safety benchmarks (Table 3), both OMNIGUARD-7B and OMNIGUARD-3B achieve state-of- the-art performance. The improvement is especially pro- nounced on the SafeWatch-Bench, where both models ex- ceed 90 % on F1 score. These results highlight the signifi- cant progress of OMNIGUARD in safety reasoning within video domain. Audio.In the audio domain (see Table 4), OMNIGUARD- 7B and OMNIGUARD-3B also achieve superior performance across both audio safety benchmarks. Since audio guardrails remain largely unexplored, our model provides a strong solution to the field and demonstrates that reasoning-based safety training can also effectively generalize to the audio modality. 6 Preprint ModelSize SafeWatchSafeSora F1ACCF1ACC GPT-4o-84.277.549.982.2 Qwen3-VL-235B235B79.571.865.984.4 Qwen2.5-VL-72B72B72.564.068.083.7 LLaVA-Video-72B72B78.270.736.980.0 Qwen2.5-Omni-7B 7B68.676.064.371.3 Qwen2.5-VL-7B 7B49.746.262.283.5 LLaVA-Video-7B7B47.244.94.575.1 OMNIGUARD-3B3B92.382.070.185.9 OMNIGUARD-7B7B90.985.771.886.1 Table 3.Performance comparison of OMNIGUARD and baseline models on video-based safety benchmarks. Best in bold and second-best in underlined. ModelSize MuTox WildGuard- TTS F1ACCF1ACC GPT-4o38.766.181.685.2 Qwen2-Audio 7B26.942.427.756.7 Qwen-Audio8B28.018.959.357.7 Kimi-Audio7B37.568.377.475.6 Qwen2.5-Omni-7B7B30.836.278.878.5 OMNIGUARD-3B 3B41.872.388.489.8 OMNIGUARD-7B7B43.775.487.889.2 Table 4.Performance comparison of OMNIGUARD and baseline models on audio-based safety benchmarks. Best in bold and second-best in underlined. Cross-Modality. We further evaluate OMNIGUARD in cross-modal safety scenarios to assess its capability to rea- son across modalities. As illustrated in Figure 3, our OM- NIGUARD family demonstrates strong and consistent per- formance across all evaluated cross-modal safety bench- marks, including M-SafetyBench, Video-SafetyBench, and AIAH. Compared to Qwen2.5-Omni-7B, OMNIGUARD- 7B consistently achieves significant improvements in accu- racy across all benchmarks, highlighting the generalization of our omni-modal safety alignment in unifying multimodal reasoning. Moreover, the lightweight OMNIGUARD-3B per- forms comparably to large-scale general-purpose models such as Qwen3-VL-235B on both M-SafetyBench and Video-SafetyBench, despite having significantly fewer pa- rameters. These results further confirm that OMNIGUARD effectively generalizes safety reasoning across modalities, offering a scalable and parameter-efficient solution for cross- modal alignment. 4.3. RQ1: Reasoning-Based Safety Training. To examine whether reasoning-based safety training en- hances the performance of omni-modal guardrails and ad- dress the additional complexity introduced by omni-modal safety reasoning, we conduct further studies on the 7B model across four modalities, as shown in Figure 4. We compare our OMNIGUARD-7B with the original base model Qwen2.5-Omni-7B (Xu et al., 2025b) and its Label-only SFT variant, which is fine-tuned solely on safety classifica- tion labels without the curated reasoning traces used in our approach. From these results, we draw several observations. (1) Both supervised fine-tuning methods can improve perfor- mance over the base model across all unimodal benchmarks (text, image, video, and audio), showing that simple safety fine-tuning can also enhance multimodal moderation ca- pabilities. (2) Compared to the Label-only baseline, our reasoning-augmented training consistently achieves higher F1 scores across all unimodal settings, improving from 75.7→77.2 (text), 74.0→75.2 (image), 79.8→81.4 (video), and 63.7→65.8 (audio). This confirms that reasoning super- vision helps the model better internalize safety assessment principles beyond surface-level pattern learning. (3) No- tably, in the cross-modal setting, the Label-only SFT variant suffers a degradation in accuracy on the Video-SafetyBench and AIAH benchmarks, whereas our reasoning-augmented model achieves consistent gains across all three tasks. This suggests that simple label supervision fails to generalize effectively facing complex moderation tasks across modal- ities, while reasoning-based alignment endows the model with stronger guardrail understanding and transferability. Overall, these results highlight that reasoning-based safety alignment not only enhances the performance of omni- modal guardrails across different modality settings, but also provides better understanding and generalization in complex cross-modal safety scenarios that label-only supervision fails to handle. 4.4. RQ2: Cross-Modal Generalization. To investigate whether safety knowledge learned in seen modalities can generalize to unseen ones, we conduct cross- modal training and evaluation. Specifically, for each split, we train OMNIGUARD using data from three modalities for training and seen modality evaluation and leave one modality out for unseen modality evaluation. We report the averaged F1 and accuracy across all four seen and unseen modality evaluations in Table 5. We draw two main conclusions from these experiments. (1) Overall, strong cross-modal transfer is observed across all four modalities. The unseen modality performance remains close to seen modality results (79.4 vs. 81.8 F1 on average), indicating that OMNIGUARD successfully learns modality- invariant safety representations. This suggests that harmful 7 Preprint GPT-4o Qwen3-VL-235B Qwen2.5-VL-72B Qwen2.5-Omni-7B Qwen2.5-VL-7B LlavaGuard-v1.2-7B LlamaGuard3V OmniGuard-3BOmniGuard-7B 0 10 20 30 40 50 Accuracy (%) 44.7 32.9 20.7 29.4 14.7 12.7 29.8 31.7 35.7 M-SafetyBench (Image-Text) GPT-4o Qwen3-VL-235B Qwen2.5-VL-72B LLaVA-Video-72B Qwen2.5-Omni-7B Qwen2.5-VL-7B LLaVA-Video-7B OmniGuard-3BOmniGuard-7B 0 10 20 30 40 50 60 70 80 Accuracy (%) 54.5 68.5 58.2 54.7 73.3 62.5 27.7 69.5 77.8 Video-SafetyBench (Video-Text) GPT-4o-Audio Qwen-Audio Qwen2-Audio Kimi-Audio Qwen2.5-Omni-7B OmniGuard-3BOmniGuard-7B 0 20 40 60 80 100 Accuracy (%) 73.7 27.4 100.0 96.6 92.3 94.3 96.0 AIAH (Audio-Text) Figure 3.Performance comparison of OMNIGUARD and baseline models on cross-modal safety benchmarks. The performance is evaluated in Accuracy (ACC). semantic patterns can be effectively aligned and learned across text, image, video, and audio inputs. OMNIGUARD further acquire generalizable safety reasoning ability across modalities. (2) One notable exception arises in the audio modality, where the seen modality F1 (64.6%) is slightly lower than the unseen modality F1 (65.1%). This phenomenon is due to the OMNIGUARD variant trained on text excluded split exhibiting degraded performance on the WildGuard-TTS benchmark. WildGuardMix-TTS is the auditory version of WildGuardMix constructed by text-to-speech model. Train- ing without textual data also lead to degradation in the audio setting, this reveals that content which are semantically equivalent but are from different modalities (e.g., harmful text vs. its spoken version) can mutually influence each other during safety alignment. And the knowledge from one form can be transferred to another semantically invariant form. Taken together, these results demonstrate that cross-modal generalization in OMNIGUARD is substantial.OMNI- GUARD can generalizes safety reasoning from trained modalities to untrained ones and but also acquires modality- invariant semantic representations of unsafe content. 5. Conclusion In conclusion, we introduce OMNIGUARD-7B and OM- NIGUARD-3B, the first family of omni-modal guardrails trained on a comprehensive and unified safety fine-tuning dataset covering both unimodal and cross-modal samples. Modality Seen ModalityUnseen Modality F1ACCF1ACC Text88.180.184.777.4 Image 78.379.076.877.1 Video78.985.176.680.0 Audio64.692.265.190.2 Average81.884.179.481.2 Table 5.Performance comparison between seen modality and un- seen modality settings across four modalities. Metrics are F1 and Accuracy (%). OMNIGUARD instantiates a unified omni-modal safety so- lution: it can moderate and reason about unsafe content in heterogeneous and cross-modal settings. Extensive ex- periments demonstrate that OMNIGUARD-7B consistently outperforms existing guardrail models across all modalities, while OMNIGUARD-3B can achieves competitive results compared to large-scale LLMs and MLLMs such as Qwen3- 235B and Qwen3-VL-235B. These results highlight the strong omni-modal safety detection and reasoning capabili- ties of our approach, confirming the feasibility of a unified guardrail system with omni-modal understanding and inter- pretability. Future work will investigate more complex and cross-modal safety scenarios to further advance omni-modal safeguarding in next-generation large language models. Limitations. Although OMNIGUARD demonstrates strong omni-modal safety reasoning and consistent performance across modalities, several limitations remain. While in- corporating reasoning paths significantly enhances the in- 8 Preprint Qwen2.5-Omni-7B Label-only SFT OmniGuard-7B 62.5 65.0 67.5 70.0 72.5 75.0 77.5 80.0 Uni-Modal Benchmarks Average F1 (%) 66.6 75.7 77.2 Text Qwen2.5-Omni-7B Label-only SFT OmniGuard-7B 55 60 65 70 75 80 56.3 74.0 75.2 Image Qwen2.5-Omni-7B Label-only SFT OmniGuard-7B 65 70 75 80 85 66.5 79.8 81.4 Video Qwen2.5-Omni-7B Label-only SFT OmniGuard-7B 50.0 52.5 55.0 57.5 60.0 62.5 65.0 67.5 70.0 54.8 63.7 65.8 Audio Qwen2.5-Omni-7B Label-only SFT OmniGuard-7B 26 28 30 32 34 36 38 40 Cross-Modal Benchmarks ACC (%) 29.4 31.9 35.7 M-SafetyBench (Image-Text) Qwen2.5-Omni-7B Label-only SFT OmniGuard-7B 62.5 65.0 67.5 70.0 72.5 75.0 77.5 80.0 82.5 73.3 65.3 77.8 Video-SafetyBench (Video-Text) Qwen2.5-Omni-7B Label-only SFT OmniGuard-7B 86 88 90 92 94 96 98 100 92.3 89.1 96.0 AIAH (Audio-Text) Qwen2.5-Omni-7B Label-only SFT OmniGuard-7B 58 60 62 64 66 68 70 72 74 65.0 62.1 69.8 Cross-Modal Average Figure 4.Comparison of performance between Label-only SFT and critique-augmented training across both uni-modal and cross-modal settings. The upper four subplots show average performance results on uni-modal benchmarks (Text, Image, Video, Audio), evaluated by F1 score (%). The bottom four subplots present cross-modal results on M-SafetyBench (Image-Text), Video-SafetyBench (Video-Text), and AIAH (Audio-Text), along with the average performance, reported in accuracy (ACC, %). terpretability and reliability of safety assessments, it in- evitably increases inference latency due to the additional reasoning generation step. As a result, OMNIGUARD has higher computational overhead compared to lightweight, binary-classification guardrail systems. The trade-off be- tween safety reasoning depth and inference latency can be further explored under different use cases and scenarios to further optimize safety robustness and efficiency. Addition- ally, due to the limited availability of publicly accessible cross-modal safety fine-tuning datasets, there remains sub- stantial room for progress in moderating more complex interleaved multimodal safety scenarios. We hope future research will continue to improve the robustness and relia- bility of omni-modal systems, and advance their capability to safeguard against risks across complex modality combi- nations. References Agarwal, S., Ahmad, L., Ai, J., Altman, S., Applebaum, A., Arbus, E., Arora, R. K., Bai, Y., Baker, B., Bao, H., and et al. gpt-oss-120b & gpt-oss-20b model card. CoRR, abs/2508.10925, 2025. doi: 10.48550/ARXIV.2508.10 925. URLhttps://doi.org/10.48550/arXiv .2508.10925. Ahmed, S. H., Khan, M. J., and Sukthankar, G. Enhanced multimodal content moderation of children’s videos us- ing audiovisual fusion. In Chun, S. A. and Talbert, D. A. (eds.), Proceedings of the Thirty-Seventh Inter- national Florida Artificial Intelligence Research Soci- ety Conference, FLAIRS 2024, Sandestin Beach, FL, USA, May 19-21, 2024. Florida Online Journals, 2024. doi: 10.32473/FLAIRS.37.1.135563. URLhttps: //doi.org/10.32473/flairs.37.1.135563. Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023. URLhttps://arxi v.org/abs/2308.12966. Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., and Lin, J. Qwen2.5-vl technical report, 2025. URL https://arxiv.org/abs/2502.13923. Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., Das- Sarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., Showk, S. E., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, 9 Preprint N., Olsson, C., Amodei, D., Brown, T. B., Clark, J., Mc- Candlish, S., Olah, C., Mann, B., and Kaplan, J. Training a helpful and harmless assistant with reinforcement learn- ing from human feedback. CoRR, abs/2204.05862, 2022. doi: 10.48550/ARXIV.2204.05862. URLhttps: //doi.org/10.48550/arXiv.2204.05862. Balat, M., Gabr, M. E., Bakr, H., and Zaky, A. B. Tikguard: A deep learning transformer-based solution for detecting unsuitable tiktok content for kids. In 6th Novel Intelligent and Leading Emerging Sciences Conference, NILES 2024, Giza, Egypt, October 19-21, 2024, p. 337–340. IEEE, 2024. doi: 10.1109/NILES63360.2024.10753192. URL https://doi.org/10.1109/NILES63360.2 024.10753192. Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., Li, B., Luo, P., Lu, T., Qiao, Y., and Dai, J. Internvl: Scaling up vision foundation models and aligning for generic visual- linguistic tasks, 2024. URLhttps://arxiv.org/ abs/2312.14238. Chen, Z., Pinto, F., Pan, M., and Li, B. Safewatch: An efficient safety-policy following video guardrail model with transparent explanations. In The Thirteenth Inter- national Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URLhttps://openreview.net/forum ?id=xjKz6IxgCX. Chi, J., Karn, U., Zhan, H., Smith, E., Rando, J., Zhang, Y., Plawiak, K., Coudert, Z. D., Upasani, K., and Pasupuleti, M. Llama guard 3 vision: Safeguarding human-ai image understanding conversations. CoRR, abs/2411.10414, 2024. doi: 10.48550/ARXIV.2411.10414. URLhttps: //doi.org/10.48550/arXiv.2411.10414. Chu, Y., Xu, J., Zhou, X., Yang, Q., Zhang, S., Yan, Z., Zhou, C., and Zhou, J. Qwen-audio: Advancing uni- versal audio understanding via unified large-scale audio- language models, 2023. URLhttps://arxiv.org/ abs/2311.07919. Chu, Y., Xu, J., Yang, Q., Wei, H., Wei, X., Guo, Z., Leng, Y., Lv, Y., He, J., Lin, J., Zhou, C., and Zhou, J. Qwen2- audio technical report, 2024. URLhttps://arxiv. org/abs/2407.10759. Comanici, G., Bieber, E., Schaekermann, M., Pasupat, I., Sachdeva, N., Dhillon, I. S., Blistein, M., Ram, O., Zhang, D., Rosen, E., and et al. Gemini 2.5: Pushing the fron- tier with advanced reasoning, multimodality, long con- text, and next generation agentic capabilities. CoRR, abs/2507.06261, 2025. doi: 10.48550/ARXIV.2507.06 261. URLhttps://doi.org/10.48550/arXiv .2507.06261. Costa-jussà, M. R., Meglioli, M. C., Andrews, P., Dale, D., Hansanti, P., Kalbassi, E., Mourachko, A., Ropers, C., and Wood, C. Mutox: Universal multilingual audio- based toxicity dataset and zero-shot detector. In Ku, L., Martins, A., and Srikumar, V. (eds.), Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, p. 5725–5734. Association for Computational Linguistics, 2024. doi: 10.18653/V1/2024.FINDINGS-A CL.340. URLhttps://doi.org/10.18653/v1/ 2024.findings-acl.340. Cui, S., Zhang, Q., Ouyang, X., Chen, R., Zhang, Z., Lu, Y., Wang, H., Qiu, H., and Huang, M. Shieldvlm: Safe- guarding the multimodal implicit toxicity via deliberative reasoning with lvlms: Shieldvlm. In Proceedings of the 33rd ACM International Conference on Multimedia, M ’25, p. 11677–11686, New York, NY, USA, 2025. Asso- ciation for Computing Machinery. ISBN 9798400720352. doi: 1 0 . 1 1 4 5 / 3 7 4 6 0 2 7 . 3 7 5 5 7 1 1.URLhttps: //doi.org/10.1145/3746027.3755711. Dai, J., Chen, T., Wang, X., Yang, Z., Chen, T., Ji, J., and Yang, Y. Safesora: Towards safety alignment of text2video generation via a human preference dataset. In Globersons, A., Mackey, L., Belgrave, D., Fan, A., Pa- quet, U., Tomczak, J. M., and Zhang, C. (eds.), Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024a. URLhttp://papers.nip s.c/paper_files/paper/2024/hash/1eb 543faf7c69e8a7eb8b85f70be818f-Abstrac t-Datasets_and_Benchmarks_Track.html. Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y. Safe RLHF: safe reinforcement learning from human feedback. In The Twelfth International Con- ference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024b. URL https://openreview.net/forum?id=TyFr POKYXw. Fang, Q., Guo, S., Zhou, Y., Ma, Z., Zhang, S., and Feng, Y. Llama-omni: Seamless speech interaction with large language models. In The Thirteenth International Confer- ence on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URLhttps: //openreview.net/forum?id=PYmrUQmMEw. Fu, C., Lin, H., Wang, X., Zhang, Y.-F., Shen, Y., Liu, X., Cao, H., Long, Z., Gao, H., Li, K., Ma, L., Zheng, X., Ji, R., Sun, X., Shan, C., and He, R. Vita-1.5: Towards gpt-4o level real-time vision and speech interaction, 2025. URL https://arxiv.org/abs/2501.01957. 10 Preprint Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A. Realtoxicityprompts: Evaluating neural toxic de- generation in language models. In Cohn, T., He, Y., and Liu, Y. (eds.), Findings of the Association for Computa- tional Linguistics: EMNLP 2020, Online Event, 16-20 November 2020, volume EMNLP 2020 of Findings of ACL, p. 3356–3369. Association for Computational Lin- guistics, 2020. doi: 10.18653/V1/2020.FINDINGS-E MNLP.301. URLhttps://doi.org/10.18653 /v1/2020.findings-emnlp.301. Ghosh, S., Varshney, P., Sreedhar, M. N., Padmakumar, A., Rebedea, T., Varghese, J. R., and Parisien, C. AEGIS2.0: A diverse AI safety dataset and risks taxonomy for align- ment of LLM guardrails. In Chiruzzo, L., Ritter, A., and Wang, L. (eds.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Associa- tion for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Al- buquerque, New Mexico, USA, April 29 - May 4, 2025, p. 5992–6026. Association for Computational Linguistics, 2025. doi: 10.18653/V1/2025.NAACL- LONG.306. URLhttps://doi.org/10.18653/v1/2025 .naacl-long.306. Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., and et al. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783. Han, S., Rao, K., Ettinger, A., Jiang, L., Lin, B. Y., Lambert, N., Choi, Y., and Dziri, N. Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms. In Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J. M., and Zhang, C. (eds.), Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. URLhttp://papers .nips.c/paper_files/paper/2024/hash /0f69b4b96a46f284b726fbd70f74fb3b-Abs tract-Datasets_and_Benchmarks_Track.h tml. Helff, L., Friedrich, F., Brack, M., Kersting, K., and Schramowski, P. Llavaguard: An open VLM-based framework for safeguarding vision datasets and models. In Forty-second International Conference on Machine Learning, 2025a. URLhttps://openreview.net /forum?id=YIO9ritzWV. Helff, L., Friedrich, F., Brack, M., Schramowski, P., and Kersting, K. Llavaguard: An open vlm-based framework for safeguarding vision datasets and models. In Proceed- ings of the 42nd International Conference on Machine Learning (ICML), 2025b. Hu, X., Liu, D., Li, H., Huang, X., and Shao, J. Vlsbench: Unveiling visual leakage in multimodal safety. In Che, W., Nabende, J., Shutova, E., and Pilehvar, M. T. (eds.), Proceedings of the 63rd Annual Meeting of the Associ- ation for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, p. 8285–8316. Association for Computational Linguistics, 2025. URLhttps://aclanthology .org/2025.acl-long.405/. Huang, R., Li, M., Yang, D., Shi, J., Chang, X., Ye, Z., Wu, Y., Hong, Z., Huang, J., Liu, J., Ren, Y., Zou, Y., Zhao, Z., and Watanabe, S. Audiogpt: Understand- ing and generating speech, music, sound, and talking head. In Wooldridge, M. J., Dy, J. G., and Natarajan, S. (eds.), Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on In- novative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Ar- tificial Intelligence, EAAI 2014, February 20-27, 2024, Vancouver, Canada, p. 23802–23804. AAAI Press, 2024. doi: 10.1609/AAAI.V38I21.30570. URLhttps: //doi.org/10.1609/aaai.v38i21.30570. Huang, T., Hu, S., Ilhan, F., Tekin, S. F., Yahn, Z., Xu, Y., and Liu, L. Safety tax: Safety alignment makes your large reasoning models less reasonable. CoRR, abs/2503.00555, 2025. doi: 10.48550/ARXIV.2503.00555. URLhttps: //doi.org/10.48550/arXiv.2503.00555. Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., and Khabsa, M. Llama guard: Llm-based input- output safeguard for human-ai conversations. CoRR, abs/2312.06674, 2023. doi: 10.48550/ARXIV.2312. 06674. URLhttps://doi.org/10.48550/arX iv.2312.06674. Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Chen, B., Sun, R., Wang, Y., and Yang, Y. Beavertails: Towards im- proved safety alignment of LLM via a human-preference dataset. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neu- ral Information Processing Systems 36: Annual Confer- ence on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023. URLhttp://papers.nips.c/p aper_files/paper/2023/hash/4dbb61cb6 8671edc4ca3712d70083b9f-Abstract-Dat asets_and_Benchmarks.html. KimiTeam, Ding, D., Ju, Z., Leng, Y., Liu, S., Liu, T., Shang, Z., Shen, K., Song, W., Tan, X., Tang, H., Wang, Z., Wei, C., Xin, Y., Xu, X., Yu, J., Zhang, Y., Zhou, X., Charles, Y., Chen, J., Chen, Y., Du, Y., He, W., Hu, Z., Lai, G., Li, Q., Liu, Y., Sun, W., Wang, J., Wang, Y., Wu, Y., Wu, Y., 11 Preprint Yang, D., Yang, H., Yang, Y., Yang, Z., Yin, A., Yuan, R., Zhang, Y., and Zhou, Z. Kimi-audio technical report, 2025. URLhttps://arxiv.org/abs/2504.1 8425. Lee, S., Kim, G., Kim, J., Lee, H., Chang, H., Park, S. H., and Seo, M. How does vision-language adaptation impact the safety of vision language models? In The Thirteenth International Conference on Learning Representations, 2025. URLhttps://openreview.net/forum ?id=eXB5TCrAu9. Liao, S., Wang, Y., Li, T., Cheng, Y., Zhang, R., Zhou, R., and Xing, Y. Fish-speech: Leveraging large language models for advanced multilingual text-to-speech synthe- sis, 2024. URLhttps://arxiv.org/abs/2411 .01156. Lin, Z., Wang, Z., Tong, Y., Wang, Y., Guo, Y., Wang, Y., and Shang, J. Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation. In Bouamor, H., Pino, J., and Bali, K. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, p. 4694–4702. Association for Computational Linguistics, 2023. doi: 10.18653/V1/2023.FINDINGS- EMNLP.311. URL https://doi.org/10.18653/v1/2023.fin dings-emnlp.311. Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning, 2023. URLhttps://arxiv.org/abs/23 04.08485. Liu, X., Xu, N., Chen, M., and Xiao, C. Autodan: Gen- erating stealthy jailbreak prompts on aligned large lan- guage models. In The Twelfth International Conference on Learning Representations, 2024a. Liu, X., Zhu, Y., Gu, J., Lan, Y., Yang, C., and Qiao, Y. Mm-safetybench: A benchmark for safety evaluation of multimodal large language models. In Leonardis, A., Ricci, E., Roth, S., Russakovsky, O., Sattler, T., and Varol, G. (eds.), Computer Vision - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-October 4, 2024, Proceedings, Part LVI, volume 15114 of Lecture Notes in Computer Science, p. 386–403. Springer, 2024b. doi: 10.1007/978-3-031-72992-8\_22. URLhttps://do i.org/10.1007/978-3-031-72992-8_22. Liu, X., Li, Z., He, Z., Li, P., Xia, S., Cui, X., Huang, H., Yang, X., and He, R. Video-safetybench: A bench- mark for safety evaluation of video lvlms.CoRR, abs/2505.11842, 2025a. doi: 10.48550/ARXIV.250 5.11842. URLhttps://doi.org/10.48550/a rXiv.2505.11842. Liu, Z., Dong, Y., Wang, J., Liu, Z., Hu, W., Lu, J., and Rao, Y. Ola: Pushing the frontiers of omni-modal lan- guage model with progressive modality alignment. CoRR, abs/2502.04328, 2025b. doi: 10.48550/ARXIV.2502.04 328. URLhttps://doi.org/10.48550/arXiv .2502.04328. Lu, J., Clark, C., Lee, S., Zhang, Z., Khosla, S., Marten, R., Hoiem, D., and Kembhavi, A. Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, p. 26429–26445. IEEE, 2024. doi: 10.1109/CVPR52733.2024.02497. URLhttps:// doi.org/10.1109/CVPR52733.2024.02497. Luo, R., Xia, X., Wang, L., Chen, L., Shan, R., Luo, J., Yang, M., and Chua, T.-S. Next-omni: Towards any- to-any omnimodal foundation models with discrete flow matching, 2025. URLhttps://arxiv.org/abs/ 2510.13721. Markov, T., Zhang, C., Agarwal, S., Nekoul, F. E., Lee, T., Adler, S., Jiang, A., and Weng, L. A holistic ap- proach to undesired content detection in the real world. In Williams, B., Chen, Y., and Neville, J. (eds.), Thirty- Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applica- tions of Artificial Intelligence, IAAI 2023, Thirteenth Symposium on Educational Advances in Artificial In- telligence, EAAI 2023, Washington, DC, USA, Febru- ary 7-14, 2023, p. 15009–15018. AAAI Press, 2023. doi: 10.1609/AAAI.V37I12.26752. URLhttps: //doi.org/10.1609/aaai.v37i12.26752. OpenAI. Gpt-4v(ision) system card. 2023. URLhttps: //api.semanticscholar.org/CorpusID:26 3218031. OpenAI. Gpt-4o system card, 2024. URLhttps://ar xiv.org/abs/2410.21276. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. Training language models to follow instructions with human feedback. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22, Red Hook, NY, USA, 2022. Curran Associates Inc. ISBN 9781713871088. Papadopoulou, O., Zampoglou, M., Papadopoulos, S., and Kompatsiaris, I. A corpus of debunked and verified user- generated videos. Online Inf. Rev., 43(1):72–88, 2019. doi: 10.1108/OIR-03-2018-0101. URLhttps://do i.org/10.1108/OIR-03-2018-0101. 12 Preprint Phan, D. D., Nguyen, T. T., Nguyen, Q. H., Tran, H. L., Nguyen, K. N. K., and Vu, D. L. Lspd: A large-scale pornographic dataset for detection and classification. In- ternational Journal of Intelligent Engineering and Sys- tems, 15:198–231. Qi, X., Huang, K., Panda, A., Henderson, P., Wang, M., and Mittal, P. Visual adversarial examples jailbreak aligned large language models. In Wooldridge, M. J., Dy, J. G., and Natarajan, S. (eds.), Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Con- ference on Innovative Applications of Artificial Intelli- gence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, Febru- ary 20-27, 2024, Vancouver, Canada, p. 21527–21536. AAAI Press, 2024a. doi: 10.1609/AAAI.V38I19.30150. URLhttps://doi.org/10.1609/aaai.v38 i19.30150. Qi, X., Zeng, Y., Xie, T., Chen, P., Jia, R., Mittal, P., and Henderson, P. Fine-tuning aligned language mod- els compromises safety, even when users do not intend to! In The Twelfth International Conference on Learn- ing Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024b. URLhttps: //openreview.net/forum?id=hTEGyKf0dZ. Qi, X., Panda, A., Lyu, K., Ma, X., Roy, S., Beirami, A., Mittal, P., and Henderson, P. Safety alignment should be made more than just a few tokens deep. In The Thir- teenth International Conference on Learning Representa- tions, ICLR 2025, Singapore, April 24-28, 2025. OpenRe- view.net, 2025. URLhttps://openreview.net /forum?id=6Mxhg9PtDE. Qu, Y., Shen, X., Wu, Y., Backes, M., Zannettou, S., and Zhang, Y. Unsafebench: Benchmarking image safety classifiers on real-world and ai-generated images. CoRR, abs/2405.03486, 2024. doi: 10.48550/ARXIV.2405.03 486. URLhttps://doi.org/10.48550/arXiv .2405.03486. Qwen, :, Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang, K., Lu, K., Bao, K., Yang, K., Yu, L., Li, M., Xue, M., Zhang, P., Zhu, Q., Men, R., Lin, R., Li, T., Tang, T., Xia, T., Ren, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Wan, Y., Liu, Y., Cui, Z., Zhang, Z., and Qiu, Z. Qwen2.5 technical report, 2025. URLhttps: //arxiv.org/abs/2412.15115. Rajpal, S. Guardrails ai.https://w.guardrails ai.com/, 2023. Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lil- licrap, T. P., Alayrac, J., Soricut, R., Lazaridou, A., Fi- rat, O., Schrittwieser, J., and et al. Gemini 1.5: Un- locking multimodal understanding across millions of tokens of context. CoRR, abs/2403.05530, 2024. doi: 10.48550/ARXIV.2403.05530. URLhttps://doi. org/10.48550/arXiv.2403.05530. Singhal, M., Ling, C., Paudel, P., Thota, P., Kumarswamy, N., Stringhini, G., and Nilizadeh, S. Sok: Content mod- eration in social media, from guidelines to enforcement, and research to practice. In 8th IEEE European Sym- posium on Security and Privacy, EuroS&P 2023, Delft, Netherlands, July 3-7, 2023, p. 868–895. IEEE, 2023. doi: 10.1109/EUROSP57164.2023.00056. URL https://doi.org/10.1109/EuroSP57164. 2023.00056. Song, J. Y., Lee, S., Lee, J., Kim, M., and Kim, J. Modsand- box: Facilitating online community moderation through error prediction and improvement of automated rules. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY, USA, 2023. Association for Computing Machinery. ISBN 9781450394215. doi: 10.1145/3544548.3581057. URL https://doi.org/10.1145/3544548.3581 057. Sultani, W., Chen, C., and Shah, M. Real-world anomaly de- tection in surveillance videos. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, p. 6479– 6488. Computer Vision Foundation / IEEE Computer Society, 2018. doi: 10.1109/CVPR.2018.00678. URL http://openaccess.thecvf.com/content_ cvpr_2018/html/Sultani_Real-World_Ano maly_Detection_CVPR_2018_paper.html. Tang, T., Wu, Y., Wu, Y., Yu, L., and Li, Y. Videomoderator: A risk-aware framework for multimodal video modera- tion in e-commerce. IEEE Trans. Vis. Comput. Graph., 28 (1):846–856, 2022. doi: 10.1109/TVCG.2021.3114781. URLhttps://doi.org/10.1109/TVCG.2021. 3114781. Team, K. Kimi-audio technical report, 2024a. Team, L. Meta llama guard 2.https://github.com /meta-llama/PurpleLlama/blob/main/Lla ma-Guard2/MODEL_CARD.md, 2024b. Team, Q. Qwen3 technical report, 2025. URLhttps: //arxiv.org/abs/2505.09388. Wang, L., Yao, K., Li, X., Yang, D., Li, H., Wang, X., and Dong, W. The man behind the sound: Demystifying audio private attribute profiling via multimodal large language model agents, 2025a. URLhttps://arxiv.org/ abs/2507.10016. 13 Preprint Wang, S., Ye, X., Cheng, Q., Duan, J., Li, S., Fu, J., Qiu, X., and Huang, X. Safe inputs but unsafe output: Benchmarking cross-modality safety alignment of large vision-language models. In Chiruzzo, L., Ritter, A., and Wang, L. (eds.), Findings of the Association for Compu- tational Linguistics: NAACL 2025, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, p. 3563–3605. Association for Computational Linguistics, 2025b. doi: 10.18653/V1/2025.FINDINGS- NAACL.198. URL https://doi.org/10.18653/v1/2025.fin dings-naacl.198. Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P. Challenges in detoxifying language models. In Moens, M., Huang, X., Specia, L., and Yih, S. W. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021, p. 2447– 2469. Association for Computational Linguistics, 2021a. doi: 10.18653/V1/2021.FINDINGS-EMNLP.210. URL https://doi.org/10.18653/v1/2021.fin dings-emnlp.210. Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P.-S. Challenges in detoxifying language models. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.- t. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2021, p. 2447–2469, Punta Cana, Dominican Republic, November 2021b. Association for Computational Linguistics. doi: 10.18653/v1/2021.findi ngs-emnlp.210. URLhttps://aclanthology.o rg/2021.findings-emnlp.210/. Wen, X., Zhou, W., Mo, W. J., and Chen, M. ThinkGuard: Deliberative slow thinking leads to cautious guardrails. In Che, W., Nabende, J., Shutova, E., and Pilehvar, M. T. (eds.), Findings of the Association for Computational Lin- guistics: ACL 2025, p. 13698–13713, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 979-8-89176-256-5. doi: 10.18653/v1/2025.findi ngs-acl.704. URLhttps://aclanthology.org /2025.findings-acl.704/. Wu, S., Fei, H., Qu, L., Ji, W., and Chua, T. Next-gpt: Any-to-any multimodal LLM. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024. URL https://openreview.net/forum?id=NZQk umsNlf. Xu, J., Guo, Z., He, J., Hu, H., He, T., Bai, S., Chen, K., Wang, J., Fan, Y., Dang, K., Zhang, B., Wang, X., Chu, Y., and Lin, J. Qwen2.5-omni technical report. CoRR, abs/2503.20215, 2025a. doi: 10.48550/ARXIV.2503.20 215. URLhttps://doi.org/10.48550/arXiv .2503.20215. Xu, J., Guo, Z., He, J., Hu, H., He, T., Bai, S., Chen, K., Wang, J., Fan, Y., Dang, K., Zhang, B., Wang, X., Chu, Y., and Lin, J. Qwen2.5-omni technical report, 2025b. URL https://arxiv.org/abs/2503.20215. Yang, H., Qu, L., Shareghi, E., and Haffari, G. Audio is the achilles’ heel: Red teaming audio large multi- modal models. In Chiruzzo, L., Ritter, A., and Wang, L. (eds.), Proceedings of the 2025 Conference of the Na- tions of the Americas Chapter of the Association for Com- putational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, p. 9292– 9306. Association for Computational Linguistics, 2025. doi: 10.18653/V1/2025.NAACL- LONG.470. URL https://doi.org/10.18653/v1/2025.naa cl-long.470. Zeng, E., Kohno, T., Roesner, F., and Allen, P. G. Bad news: Clickbait and deceptive ads on news and misinformation websites. 2020. URLhttps://api.semanticsc holar.org/CorpusID:219178438. Zhan, J., Dai, J., Ye, J., Zhou, Y., Zhang, D., Liu, Z., Zhang, X., Yuan, R., Zhang, G., Li, L., Yan, H., Fu, J., Gui, T., Sun, T., Jiang, Y., and Qiu, X. Anygpt: Uni- fied multimodal LLM with discrete sequence modeling. In Ku, L., Martins, A., and Srikumar, V. (eds.), Pro- ceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, p. 9637–9662. Association for Computational Linguistics, 2024. doi: 10.18653/V1/2024.ACL-LONG.521. URL https://doi.org/10.18653/v1/2024.acl -long.521. Zhang, Y., Wu, J., Li, W., Li, B., Ma, Z., Liu, Z., and Li, C. Llava-video: Video instruction tuning with synthetic data, 2025. URLhttps://arxiv.org/abs/2410 .02713. Zhao, Y., Huang, J., Hu, J., Wang, X., Mao, Y., Zhang, D., Jiang, Z., Wu, Z., Ai, B., Wang, A., Zhou, W., and Chen, Y. Swift:a scalable lightweight infrastructure for fine-tuning, 2024. URLhttps://arxiv.org/ab s/2408.05517. Zhou, W., Zhang, S., Gu, Y., Chen, M., and Poon, H. Universalner: Targeted distillation from large language models for open named entity recognition, 2024. URL https://arxiv.org/abs/2308.03279. Zhu, T., Liu, Q., Wang, F., Tu, Z., and Chen, M. Unravel- ing cross-modality knowledge conflicts in large vision- 14 Preprint language models.arXiv preprint arXiv:2410.03659, 2024. Zhu, T., Zhang, K., Chen, M., and Su, Y. Is extending modality the right path towards omni-modality? arXiv preprint arXiv:2506.01872, 2025. Zizzo, G., Cornacchia, G., Fraser, K., Hameed, M. Z., Rawat, A., Buesser, B., Purcell, M., Chen, P.-Y., Sattigeri, P., and Varshney, K. R. Adversarial prompt evaluation: Systematic benchmarking of guardrails against prompt input attacks on llms. In Neurips Safe Generative AI Workshop 2024, 2024. Zong, Y., Bohdal, O., Yu, T., Yang, Y., and Hospedales, T. M. Safety fine-tuning at (almost) no cost: A baseline for vision large language models. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024. URL https://openreview.net/forum?id=bWZK vF0g7G. 15 Preprint A. Datasets and Benchmarks NameCitationTrainTest TextBeaverTails(Ji et al., 2023)27,1863,021 Aegis 2.0(Ghosh et al., 2025)30,0071,964 WildGuardMix(Han et al., 2024)86,7591,756 ToxicChat(Lin et al., 2023)5,0825,083 OpenAI Moderation(Markov et al., 2023)–1,680 ImageUnsafeBench(Qu et al., 2024)8,1098,109 VLGuard(Zong et al., 2024)1,9991,999 LlavaGuard(Helff et al., 2025a)4,5714,571 VideoSafeSora(Dai et al., 2024a)51,5885,745 Fake Video Corpus(Papadopoulou et al., 2019)380– LSPD(Phan et al.)4,000– TikHarm(Balat et al., 2024)2,762– DCSASS(Sultani et al., 2018)1,610– SafeWatch-Bench(Chen et al., 2025)–1620 AudioMuTox (English)(Costa-jussà et al., 2024)13,6171,945 WildGuardMix-TTS(Han et al., 2024)10,0001,756 Text-ImageVLSBench(Hu et al., 2025)2,240– M-SafetyBench(Liu et al., 2024b)-5,040 Text-VideoVideo-SafetyBench(Liu et al., 2025a)–2,264 Text-AudioAIAH(Yang et al., 2025)–350 Table 6.Overview of dataset sources used in the constructed dataset and benchmarks used for evaluation, with corresponding training and testing instance counts. “–” indicates not used or unavailable. We next detail the data sources used in constructing datasetD and evaluation. Text. We collect and aggregate textual safety data from BeaverTails (Ji et al., 2023), WildGuardMix (Han et al., 2024), Aegis 2.0 (Ghosh et al., 2025), and ToxicChat (Lin et al., 2023) for training and evaluation. Additionally, we include OpenAI Moderation (Markov et al., 2023) for evaluation. Image. We collect image safety data from UnsafeBench (Qu et al., 2024), VLGuard (Zong et al., 2024), and Llava- Guard (Helff et al., 2025a) for both training and evaluation. Video. For constructing the datasetD, we collect video safety data from SafeSora (Dai et al., 2024a), Fake Video Corpus (Papadopoulou et al., 2019), LSPD (Phan et al.), TikHarm (Balat et al., 2024), and DCSASS (Sultani et al., 2018). From SafeSora, we utilize the generated video clips along with their corresponding safety classification labels. Since the remaining datasets each target specific domains, we adopt the unified taxonomy proposed in (Chen et al., 2025) to integrate them into a comprehensive video safety corpus. For evaluation, we include SafeSora (Dai et al., 2024a), and SafeWatch-Bench (Chen et al., 2025). 2 Audio.For training and evaluation, we leverage MuTox (Costa-jussà et al., 2024), a multilingual audio dataset for toxicity and harassment detection. We retain only the English subset. To further enhance the model’s capability in audio-based safety recognition, we also construct a synthetic dataset, WildGuardMix-TTS, by randomly sampling 10K textual prompts from WildGuardMix and converting them into audio clips using text-to-speech (TTS) model openaudio-s1-mini (Liao et al., 2024). Cross-Modal. Due to the scarcity of cross-modal safety training datasets with fine-grained annotations, we incorporate VLSBench (Hu et al., 2025), a carefully curated benchmark for visual-language safety evaluation, to strengthen the guardrail foundation for cross-modal safety alignment. For evaluation, we use M-SafetyBench (Liu et al., 2024b), Video-SafetyBench (Liu et al., 2025a), and AIAH (Yang et al., 2025), corresponding to image-text, video-text, and auio-text configurations. 2 We did not include SafeWatch-Bench (Chen et al., 2025) for training, as their training split was unavailable at the time of our work. 16 Preprint B. Formatted Prompt Formatted Prompt [Task] You are an LLM Judge tasked with evaluating the safety of input based on the Safety Guidelines. Analyze the given input, identify any potential violations, and provide a detailed explanation for your judgment. Safety Guidelines [Input] Text / <Image> / <Video> / <Audio> [Instructions] 1. Identify whether the input is “safe” or “unsafe”. 2. If “unsafe”, list the violated categories separated by semicolons. 3. Provide a detailed explanation of whether the input is safe or unsafe. [Output Format] - Safety Assessment: Safety Label - Categories: Violation Categories - Explanation: [Provide detailed reasoning here] Figure 5.Prompt template used for target distillation from teacher models when generating safety critiques. Provided with the safety label (safe or unsafe) and the corresponding violation categories, the teacher models are instructed to produce a detailed explanation describing the rationale behind the safety assessment. 17 Preprint C. Baselines ModelCitationSizeVersion Large Language Models Qwen3-235B(Team, 2025)235BQwen/Qwen3-235B-A22B-Instruct-2507 LLaMA-3.3-70B(Grattafiori et al., 2024)70BLlama-3.3-70B-Instruct Qwen2.5-72B(Qwen et al., 2025)72BQwen2.5-72B-Instruct Qwen2.5-7B(Qwen et al., 2025)7BQwen2.5-7B-Instruct LLaMA Guard 1(Inan et al., 2023)7BLlamaGuard-7b LLaMA Guard 2(Team, 2024b)8BMeta-Llama-Guard-2-8B LLaMA Guard 3(Grattafiori et al., 2024)8BLlama-Guard-3-8B ThinkGuard(Wen et al., 2025)8BThinkGuard Vision Large Language Models Qwen3-VL-235B(Team, 2025)235BQwen3-VL-235B-A22B-Instruct Qwen2.5-VL-72B(Bai et al., 2025)72BQwen2.5-VL-72B-Instruct Qwen2.5-VL-7B(Bai et al., 2025)7BQwen2.5-VL-7B-Instruct LlavaGuard-v1.2-7B(Helff et al., 2025a)7BLlavaGuard-v1.2-7B-OV-hf LLaMA Guard 3V(Chi et al., 2024)11BLlama-Guard-3-11B-Vision LLaVA-Video-72B(Zhang et al., 2025)72BLLaVA-Video-72B-Qwen2 LLaVA-Video-7B(Zhang et al., 2025)7BLLaVA-Video-7B-Qwen2 Audio Large Language Models Qwen2-Audio(Chu et al., 2024)7BQwen2-Audio-7B Qwen-Audio(Chu et al., 2023)8BQwen-Audio-Chat Kimi-Audio(KimiTeam et al., 2025)7BKimi-Audio-7B-Instruct Omni-Modal Large Language Models GPT-4o(OpenAI, 2024)-gpt-4o-2024-11-20, gpt-4o-audio-preview-2025-06-03 Qwen2.5-Omni-7B(Xu et al., 2025a)7BQwen2.5-Omni-7B Table 7.Configuration details of baseline models used in evaluation, including Large Language Models (LLMs), Large Vision-Language Models (LVLMs), and Large Audio Language Models (LALMs). “–” denotes information unavailable. For GPT-4o, we employed gpt-4o-2024-11-20 for text, image, and video evaluations, and gpt-4o-audio-preview-2025-06-03 for audio-related assessments, as a truly omni-modal API endpoint was not publicly available at the time of evaluation. 18