Paper deep dive
Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies
Murali Indukuri, Mohammad Eskandari, Sree Nitya Kollu, Stephanie Lukin, Cynthia Matuszek
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/22/2026, 2:16:44 AM
Summary
This paper evaluates Vision-Language Models (VLMs) on their ability to distinguish between physical hazards and contextual anomalies in safety-critical scenarios. The authors introduce a four-class dataset (Safe, Anomalous, Hazardous, Anomalous-Hazardous) and test multiple VLMs using zero-shot, few-shot, and chain-of-thought prompting strategies. Results indicate that VLMs frequently conflate anomalies with hazards, over-relying on contextual irregularity as a proxy for danger, and that explicit separation of these concepts improves evaluation fidelity and exposes specific failure modes.
Entities (9)
Relation Signals (6)
Vision-Language Models â evaluatedon â Hazard vs Anomaly Classification
confidence 95% ¡ We evaluate several state-of-the-art VLMs across two datasets and multiple prompting strategies to test whether this distinction changes model behavior.
GPT-4.1 â usedin â Evaluation Study
confidence 95% ¡ For our study, we use GPT 4.1; Mistral Ministral 3B, 14B, and Large; and Gemini 2.5 Flash and 3 Flash.
Mistral-Large â usedin â Evaluation Study
confidence 95% ¡ For our study, we use GPT 4.1; Mistral Ministral 3B, 14B, and Large; and Gemini 2.5 Flash and 3 Flash.
Gemini 3 Flash â usedin â Evaluation Study
confidence 95% ¡ For our study, we use GPT 4.1; Mistral Ministral 3B, 14B, and Large; and Gemini 2.5 Flash and 3 Flash.
Vision-Language Models â exhibitsbias â Over-reliance on contextual irregularity
confidence 90% ¡ Our results show that VLMs frequently misinterpret anomalousness as hazardousness, revealing an over-reliance on contextual irregularity as a proxy for danger.
Chain-of-Thought Prompting â improves â Calibration
confidence 85% ¡ We further show that explicitly separating anomaly from hazard provides a more informative evaluation of VLM safety reasoning and exposes failure modes that binary safety judgments can obscure.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and support decision-making during emergencies. Vision-Language Models (VLMs) are promising for these settings because they can interpret complex scenes and communicate safety-relevant information, but they still require careful evaluation to ensure reliable safety reasoning. In particular, current evaluations often frame danger recognition as a binary decision (Safe/Unsafe), making it unclear whether a model is identifying true physical hazards or merely reacting to unusual scene elements. We address this limitation by introducing an explicit distinction between hazard and anomaly, and by separately recognizing hazardous and anomalous states. We evaluate several state-of-the-art VLMs across two datasets and multiple prompting strategies to test whether this distinction changes model behavior. Our results show that VLMs frequently misinterpret anomalousness as hazardousness, revealing an over-reliance on contextual irregularity as a proxy for danger. We further show that explicitly separating anomaly from hazard provides a more informative evaluation of VLM safety reasoning and exposes failure modes that binary safety judgments can obscure. Our public dataset is available on Roboflow this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2607.18325v1
- Canonical: https://arxiv.org/abs/2607.18325v1
Trouble viewing inline? Open PDF directly â
Full Text
48,987 characters extracted from source content.
Expand or collapse full text
Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies Murali Indukuri1â, Mohammad Eskandari1â, Sree Nitya Kollu1, Stephanie Lukin2, Cynthia Matuszek1ââ ^1 * Equal contributions by these authors.â This work was supported in part by NSF Grants IIS-2024878 and IIS-2145642, and this material is also based on research that is in part supported by the Army Research Laboratory, Grant No. W911NF2120076.1Interactive Robotics and Language Lab, University of Maryland Baltimore County, Computer Science and Electrical Engineering Department, Baltimore, MD, USA. (muralii1|eskandari|skollu2|cmat)@umbc.eduAccepted to RO-MAN 2026.2DEVCOM Army Research Laboratory, Adelphi, MD, USA. stephanie.m.lukin.civ@army.mil Abstract Modern safety-critical systems increasingly rely on humanârobot interaction to reduce disaster risk and support decision-making during emergencies. VisionâLanguage Models (VLMs) are promising for these settings because they can interpret complex scenes and communicate safety-relevant information, but they still require careful evaluation to ensure reliable safety reasoning. In particular, current evaluations often frame danger recognition as a binary decision (Safe/Unsafe), making it unclear whether a model is identifying true physical hazards or merely reacting to unusual scene elements. We address this limitation by introducing an explicit distinction between hazard and anomaly, and by separately recognizing hazardous and anomalous states. We evaluate several state-of-the-art VLMs across two datasets and multiple prompting strategies to test whether this distinction changes model behavior. Our results show that VLMs frequently misinterpret anomalousness as hazardousness, revealing an over-reliance on contextual irregularity as a proxy for danger. We further show that explicitly separating anomaly from hazard provides a more informative evaluation of VLM safety reasoning and exposes failure modes that binary safety judgments can obscure. Our public dataset is available on Roboflow: https://app.roboflow.com/vlm-in-context-anomaly-and-hazard-detection/camera-ready-roman-ds. I Introduction To prevent disasters or respond to them effectively, critical domains such as search and rescue, industrial safety, and space exploration have implemented strict safety standards [28, 13, 17]. However, ensuring compliance with these standards is still a challenge, in part because operators may still lack adequate situational awareness [24, 11]. For example, workers in construction zones may overlook abandoned tools, while a first responder scanning an emergency scene may mistake someoneâs red costume (see fig.Ë1) for blood, drawing attention away from a real hazard nearby. As a result, modern safety-critical systems increasingly rely on autonomous robots that can provide situational awareness cues to support human operators in high-risk environments [8, 13]. While simple object detection can support Human-Robot Interaction (HRI) in high-risk areas, recent VisionâLanguage Models (VLMs) appear promising for this role because they can describe and interpret complex scenes by combining visual perception with language-based reasoning [2, 6, 18, 15], limited currently by a lack of clear evaluation of how they form safety judgments and process different types of risk [5]. Figure 1: Images extracted from our public dataset representing four classes. Clockwise from bottom right: Safe, Anomalous, Hazardous, and Anomalous-Hazardous. This dataset aims to highlight the conceptual difference between anomalous and hazardous conditions in borderline cases. Although some studies on VLMs report strong performance on multimodal benchmarks [40], others document failures even on basic perceptual tasks [29]; still, whether VLMs can distinguish merely unusual situations from genuine physical danger, or understand the context when making safety judgments, remains an open question. In many evaluations, danger detection is framed as a binary choice [39, 5], forcing ambiguous scenes into coarse categories. As a result, anomalous but non-dangerous situations are often labeled as hazardous, while subtle hazards are overlooked. We argue that this stems from the absence of an explicit separation between anomalies and hazards. Following OSHA [27], we define a hazard as a source of potential physical harm under normal interaction conditions. An anomaly, in contrast, is a violation of expected spatial, functional, or contextual relations that does not necessarily imply danger. These definitions allow us to test whether models can differentiate the orthogonal classes. We contribute the following: 1. An evaluation pipeline for hazard/anomaly classification in VLMs, designed to study distinctions relevant to situational awareness in safety-critical settings. 2. A curated 610-image benchmark with a four-class hazardâanomaly scheme for evaluating anomalyâhazard separation in VisionâLanguage Models. 3. A systematic evaluation showing that state-of-the-art VLMs conflate anomalousness with hazardousness under binary safety prompting, and that explicit separation improves calibration. I Related Work The concepts of hazard and anomaly are semantically close. In other words, it is essential for the VLMs to disambiguate between these concepts to effectively execute the task. Previous research has shown that it is possible to reliably determine the meaning of words from context [38, 32], and it is possible for multimodal models to correctly map ambiguous words to the corresponding visual stimuli [22], showing that VLMs combined with word sense disambiguation are promising in distinguishing these concepts. Another relevant area is the study of object-contextual reasoning in multimodal models. Previous research has evaluated the perception and reasoning capabilities of VLMs using datasets such as PAM and CVR [1, 35], or has focused on the evaluation of the sensitivity of VLMs to the composition of the captions, including the presence of words that are placed in the wrong position [33]. The results of these studies demonstrate that contextual information and fine-tuning strategies improve contextual reasoning [23, 37]. Existing research has shown that modality bias exists in multimodal reasoning. For example, if textual tokens carry answers to a particular query, vision-language models (VLMs) often rely on textual information over visual information [2]. In addition, for tasks requiring precise joint reasoning over both textual and visual information, VLMs significantly underperform human baselines [34]. Our research attempts to improve the existing models by providing definitions for âanomalyâ and âhazard,â thus encouraging models to focus on relevant visual information while avoiding ambiguities. We have also designed a task where dense captions are used as an input modality to isolate reasoning ability with text-only input. Other work has explored multimodal reasoning models for robotic applications. For instance, MLLMs (Multimodal Large Language Models) have been incorporated into robots to identify semantic anomalies arising from out-of-distribution interactions [10]. Research on understanding road accidents has shown that VLMs are well suited for visually grounded perception tasks, while accident prevention is still beyond their capabilities [19]. Other studies have shown that using smaller models can improve their performance over larger models, especially in hazard detection [36]. Most relevant to the current research, existing research has explored VLMs for reasoning about anomalies, context, and other related issues. For example, models can generate descriptions for anomalies from images, which can then be used to classify the images with high accuracy [41]. In addition, BERT models can be used to establish whether an object is within a particular scene, thus providing information for context detection [39]. Misaligned reasoning is an enduring concern in the deployment of such systems. Empirical studies have demonstrated that multimodal models are prone to producing harmful and misleading advice in the process of joint reasoning over images and text data. Open-weight models are particularly susceptible to such model failures [30]. In other studies, the over-reaction problem is reported, where the model is prone to over-reacting and incorrectly identifying scenes that are merely anomalous as dangerous in the process of binary safety labeling, [5] in the study of VLMs. This study extends the aforementioned research by clearly defining the concepts of anomaly and hazard as distinct labels, allowing the model to reason over contextual anomalies. By including the formal definitions of the concepts within the prompts, the study attempts to address the overreaction problem and provide more consistent results in the identification of the most relevant safety risks. I Methodology For a systematic evaluation of current VLM capability in distinguishing the semantically close concepts of hazard and anomaly, we construct a 610-image dataset labeled with one of four classes: Safe, Anomaly, Hazard, or Both. We then use this dataset with several different prompting techniques to evaluate VLM capability. Further, we introduce a simplified classification task that evaluates whether current models can perform the four-class classification given the dense captions associated with a particular image of the dataset. We begin by formalizing this methodology. System Design. For our formalization, we use the following symbolism: ⢠Let P denote the prompt, an ordered tuple which may be composed of the following elements: â Let iTQ^T_i and iIQ^I_i be the ithi_th set of text and image tokens in P that the VLM uses as input. â Let âiR_i be the ithi_th element of P that represents a VLM response. A given element of P may be associated with a classification code cic_i, discussed further in the following sections. We treat a given VLM agent as a function A that maps P, a prompt of length i, to some response âi+1R_i+1 (i.e. â()=âi+1A(P)=R_i+1). âi+1R_i+1 may be appended to P along with other elements, i+1TQ^T_i+1 and i+1IQ^I_i+1. Generally speaking, the objective of the VLM is to produce the correct ci+1c_i+1 in Ri+1R_i+1 to match the expected in cic_i in iTQ^T_i or iIQ^I_i (i.e. the last element of the prompt P). Dataset Design. Our goal is to evaluate the ability of state-of-the-art VLMs in distinguishing hazards, anomalies, and benign scenes. To this end, we developed a dataset where each image, iIQ^I_i, is labeled with cic_i taking two-bit values corresponding to the following: ⢠Safe: No immediate hazard is present and object-context relations are not violated. ⢠Anomalous: A scene element appears out of place or inconsistent with expected relations (interposition, support, probability, position, or familiar size), but poses no immediate physical risk. ⢠Hazardous: A visible source of immediate harm is present, consistent with OSHAâs definition of a hazard, without any contextual irregularity. ⢠Anomalous Hazardous: The scene presents both an immediate hazard and an objectâcontext violation. Unlike traditional safety datasets, which combine contextual deviation and physical danger under a unified label, our two-bit encoding treats the notions of anomaly and danger as orthogonal dimensions. This allows for a controlled assessment of whether predictions are driven by object-level risk cues or more general contextual deviations. This dataset, therefore, allows for a more in-depth study of VLMsâ reasoning abilities when judging hazards. The main dataset comprises artificially generated images based on textual descriptions from several public datasets and images. These images spanned many fields including construction sites, natural disasters, and hazardous objects like weapons [21, 16, 26, 20, 14, 25, 31]. We believe that the range of environments and contexts that these images were pulled from supports a rigorous evaluation of how VLMs may perform when deployed in real-life anomaly and hazard detection scenarios. To construct the datasets, images were hosted on Roboflow [9]. The source datasets were imported to a combined dataset, which we used to construct text descriptions. Images for the final published dataset were generated from these descriptions. In the final dataset, the researchers independently labeled each image with a binary hazard and anomaly labels in the form of AH. For instance, 01 indicated the image contained hazards but not an anomaly. The researchers reviewed these independent annotations collectively to reach a unanimous agreement. In borderline cases, a displaced element could be considered either safe or risky, which can cause disagreements among the annotators. Such discussions highlighted the ambiguity between anomalous and hazardous conditions, resulting in more refined definitions of anomaly and hazard. When an agreement could not be reached after a discussion among the annotators, the images were excluded. The final dataset included over 100 images per class with 610 images in total. The safe, anomalous, hazardous, and anomalous-hazardous categories had 213, 113, 111, and 173 images respectively. This imbalance results from the less common nature of exclusively anomalous or hazardous situations. The study down-sampled to use an equal number of images in all classes to limit testing bias. Dense Captioning. We hypothesize that the main bottleneck for VLM performance in our task is the need to appropriately process the context in an image to get the information needed for classification. Work has suggested that current models struggle with this [4]. We therefore propose the related task of classifying dense captions of images from our dataset into the four discussed classes. With this, we seek to isolate the task of identifying context in an image from reasoning about said context. Based on preceding work that shows VLMsâ capability for 3D reasoning from text [3], we hope this results in more accurate classifications for our task. To generate dense captions, iTQ^T_i, for each image of our dataset, we use Googleâs Gemini 3 Flash. Each dense caption is roughly a paragraph long and stored with the associated imageâs classification code cic_i. Because our goal is to isolate VLMsâ ability to process context, we reviewed the dense captions to ensure they were accurate for the dataset images and that they are sufficiently detailed to be able to justify a given class prediction ci+1c_i+1 for caption iTQ^T_i. Task Design. Core prompt and problem definition (0T)(Q^T_0): You are inspecting an environment for anomalies and hazards to help first responders. Use the following definitions. ⢠Hazard: A scene is hazardous if normal interaction with at least one of the present elements would realistically cause immediate physical harm to the people or put the existing property in significant danger. ⢠Anomaly: A scene is anomalous if any visible element clearly violates normal expectations for this environment (location, arrangement, state, or use). ⢠Safe: Safe scenes include normal, expected human activities in the specifically given appropriate environment, provided no immediate physical hazard is present. The output structure specification (fTQ^T_f) Decide hazard and anomaly labels independently. Obey this STRICT output format: â<a brief, explicit justification for the final decision>: <code>â Where <code>is ONLY one of: 00 = Safe, 10 = Anomalous, 01 = Hazardous 11 = Anomalous AND Hazardous - No extra paragraphs, no bullet points, no additional commentary. The primary objectives of our task design are to evaluate the ability of modern VLMs to recognize anomalies and hazards, and to determine how well they distinguish between these two closely related concepts. To this end, we define several related tasks in which a VLM agent maps a prompt iP_i to a response RiR_i containing a classification code cic_i. For each task, the predicted code cic_i is compared against the ground-truth label associated with the target query instance. This general formulation allows us to keep the query content fixed across tasks while varying the prompting strategy, thereby isolating the effects of prompt design and input modality from those of contextual variation between prompts. For all tasks, we begin by defining the role of the agent and the definitions it must use in the first element of the prompt, P. We denote this 0TQ^T_0. We include formatting information near the end of each prompt to standardize model responses. We label this fTQ^T_f, which contains the specification for the code, and how the model must justify its response. The specific structure for the rest of prompt depended on the prompting technique and input modality as is discussed in the following sections. Approach: Zero Shot Prompting. We first test the simplest approach to our classification task: zero shot prompting. This technique allowed the most flexibility for model reasoning and interpretation of the definitions in 0TQ^T_0. Our prompt was of the following form: =(0T,1T,2T,3T,4)P=(Q^T_0,Q^T_1,Q^T_2,Q^T_3,Q_4), where the queries 1-3 were: 1TQ^T_1: Identify concrete sources of immediate harm. 2TQ^T_2: Independently assess contextual irregularities. Q^T_3 =fT=Q^T_f As mentioned, our study included two variations of the experiment for each prompt: using images from our dataset, or dense captions for those images. 4Q_4 represents the image for the former experiment and the set of dense captions for the latter experiment. Approach: Few Shot Prompting. Because interpreting the definitions for an anomaly vs. a hazard can be subjective, we propose that more sophisticated techniques will improve performance. Few-shot prompting builds on the zero shot prompt by providing example responses rather than the broad instructions in the previous technique. We use the following structure: =( =( 0T+fT,1,â1, ^T_0+Q^T_f,Q_1,R_1, 0T+fT,2,â2, ^T_0+Q^T_f,Q_2,R_2, 0T+fT,3,â3, ^T_0+Q^T_f,Q_3,R_3, 0T+fT,4,â4, ^T_0+Q^T_f,Q_4,R_4, 0T+fT,5). ^T_0+Q^T_f,Q_5). Queries 1Q_1 to 4Q_4 represent example images or the corresponding dense captions from each class of our dataset, and â1R_1 to â4R_4 represent examples of responses the agent should provide for the respective query. 5Q_5 is the actual image or caption the model must respond to. The four examples did not have any common elements; they were from unique scenes, domains, and had unique ground truth labels. This keeps the input token usage fairly low to meet restrictions of smaller models while also providing a broad range of examples that should help improve classification performance. Approach: Chain-of-Thought Prompting. In addition to providing selected examples, it is possible that guiding agents to identify relevant context from images before classification can improve performance. To facilitate completing our task in a step-by-step manner, we tested each agent with chain of thought prompting. We use the following inductive implementation: âi+1 _i+1 =Aâ(i) =A(P_i) i+1 _i+1 =i+âi+1+i+1T =P_i+R_i+1+Q^T_i+1 1 _1 =0T+1T+1 =Q^T_0+Q^T_1+Q_1 We use 1Q_1 to denote either the dense caption or image that the agent must classify. We take â4R_4 as the final response from the model to extract the classification. The text queries for each step are given as follows: 1TQ^T_1: In accordance with the definition for âHazard," Identify any real physically present element that could cause immediate harm. Also consider if people or property are directly in danger. 2TQ^T_2: Independently determine whether any element violates normal expectations for this environment using the definition for âAnomaly". 3TQ^T_3 =fT=Q^T_f Model Selection. Since our task is designed to evaluate VLM effectiveness for autonomous annotation of hazards, we argue that it is important for models to be both accurate and fast, especially when deployed in real-time applications. Given this, we seek to compare the potential trade-off in accuracy and gain in speed when using smaller, more optimized models vs. larger ones like GPT 4.1 or Mistral Large. For our study, we use GPT 4.1; Mistral Ministral 3B, 14B, and Large; and Gemini 2.5 Flash and 3 Flash. The primary motivations for selecting these models were to evaluate how models of different sizes perform on this task, as well as to compare how older models perform relative to newer models. Baseline Comparison. The most closely related prior work on evaluating whether VLMs can differentiate between safe and hazardous [5] considers only binary classification and can be treated as a relevant baseline for our more complex classification. We hypothesize that explicitly separating anomalous from hazardous cases reduces misclassifications between these categories, thereby mitigating the overreaction problem observed in VLM systems [5]. To enable comparison, we adopt their dataset and adapt its annotations to our four-class scheme. Because the original dataset does not include an anomalous category, we map exclusively âanomalous" cases to âsafe" and âanomalousâhazardous" cases to âhazardous" for binary evaluation. However, we keep our prompts the same to still allow models to be more granular in their assignment. This mapping allows direct comparison with the results reported in the baseline study [5]. IV Results and Analysis Model P_A R_A F1_A P_H R_H F1_H Image Input Chain of Thought Prompting Ministral 3b 0.566 6670.566\,667 0.850 0000.850\,000 0.680 0000.680\,000 0.612 5000.612\,500 0.980 0000.980\,000 0.753 8460.753\,846 Ministral 14b 0.510 5260.510\,526 0.97 0.668 9660.668\,966 0.603 6590.603\,659 0.99 0.750 0000.750\,000 Mistral large 0.612 4030.612\,403 0.790 0000.790\,000 0.689 9560.689\,956 0.664 3840.664\,384 0.970 0000.970\,000 0.788 6180.788\,618 Gemini 3 Flash 0.589 9280.589\,928 0.820 0000.820\,000 0.686 1920.686\,192 0.73 0.960 0000.960\,000 0.83 Gemini 2.5 Flash 0.605 4420.605\,442 0.890 0000.890\,000 0.72 0.680 5560.680\,556 0.980 0000.980\,000 0.803 2790.803\,279 Gpt 4.1 0.65 0.810 0000.810\,000 0.72 0.713 2350.713\,235 0.970 0000.970\,000 0.822 0340.822\,034 Few Shot Prompting Ministral 3b 0.572 6500.572\,650 0.690 7220.690\,722 0.626 1680.626\,168 0.754 9020.754\,902 0.793 8140.793\,814 0.773 8690.773\,869 Ministral 14b 0.624 0000.624\,000 0.780 0000.780\,000 0.693 3330.693\,333 0.754 5450.754\,545 0.830 0000.830\,000 0.790 4760.790\,476 Mistral large 0.69 0.700 0000.700\,000 0.696 5170.696\,517 0.784 3140.784\,314 0.800 0000.800\,000 0.792 0790.792\,079 Gemini 3 Flash 0.658 3330.658\,333 0.790 0000.790\,000 0.72 0.82 0.880 0000.880\,000 0.85 Gemini 2.5 Flash 0.604 1670.604\,167 0.88 0.72 0.82 0.880 0000.880\,000 0.85 Gpt 4.1 0.650 0000.650\,000 0.780 0000.780\,000 0.709 0910.709\,091 0.803 5710.803\,571 0.90 0.85 Zero Shot Prompting Ministral 3b 0.561 4040.561\,404 0.390 2440.390\,244 0.460 4320.460\,432 0.596 4910.596\,491 0.944 4440.944\,444 0.731 1830.731\,183 Ministral 14b 0.566 6670.566\,667 0.86 0.682 7310.682\,731 0.681 4810.681\,481 0.929 2930.929\,293 0.786 3250.786\,325 Mistral large 0.664 0000.664\,000 0.830 0000.830\,000 0.74 0.618 4210.618\,421 0.940 0000.940\,000 0.746 0320.746\,032 Gemini 3 Flash 0.70 0.700 0000.700\,000 0.700 0000.700\,000 0.77 0.890 0000.890\,000 0.82 Gemini 2.5 Flash 0.592 0000.592\,000 0.840 9090.840\,909 0.694 8360.694\,836 0.675 0000.675\,000 0.95 0.790 2440.790\,244 Gpt 4.1 0.679 6120.679\,612 0.700 0000.700\,000 0.689 6550.689\,655 0.693 4310.693\,431 0.95 0.801 6880.801\,688 Dense Caption Input Chain of Thought Prompting Ministral 3b 0.536 7650.536\,765 0.730 0000.730\,000 0.618 6440.618\,644 0.574 8500.574\,850 0.960 0000.960\,000 0.719 1010.719\,101 Ministral 14b 0.508 2870.508\,287 0.92 0.654 8040.654\,804 0.611 4650.611\,465 0.960 0000.960\,000 0.747 0820.747\,082 Mistral Large 0.536 2320.536\,232 0.740 0000.740\,000 0.621 8490.621\,849 0.683 0990.683\,099 0.97 0.801 6530.801\,653 Gemini 3 Flash 0.612 0690.612\,069 0.710 0000.710\,000 0.657 4070.657\,407 0.76 0.890 0000.890\,000 0.82 Gpt 4.1 0.64 0.720.72 0.68 0.700.70 0.960.96 0.810.81 Few Shot Prompting Ministral 3b 0.500 0000.500\,000 0.717 3910.717\,391 0.589 2860.589\,286 0.787 2340.787\,234 0.778 9470.778\,947 0.783 0690.783\,069 Ministral 14b 0.626 0870.626\,087 0.720 0000.720\,000 0.669 7670.669\,767 0.86 0.670 0000.670\,000 0.752 8090.752\,809 Mistral Large 0.67 0.650.65 0.660.66 0.86 0.760.76 0.810.81 Gemini 3 Flash 0.663 5510.663\,551 0.710 0000.710\,000 0.69 0.847 8260.847\,826 0.780 0000.780\,000 0.812 5000.812\,500 Gemini 2.5 Flash 0.589 5520.589\,552 0.79 0.675 2140.675\,214 0.826 9230.826\,923 0.86 0.84 Gpt 4.1 0.663 1580.663\,158 0.630 0000.630\,000 0.646 1540.646\,154 0.803 9220.803\,922 0.820 0000.820\,000 0.811 8810.811\,881 Zero Shot Prompting Ministral 3b 0.562 5000.562\,500 0.214 2860.214\,286 0.310 3450.310\,345 0.666 6670.666\,667 0.933 3330.933\,333 0.777 7780.777\,778 Ministral 14b 0.555 5560.555\,556 0.750 0000.750\,000 0.638 2980.638\,298 0.615 3850.615\,385 0.888 8890.888\,889 0.727 2730.727\,273 Mistral Large 0.603 6040.603\,604 0.670 0000.670\,000 0.635 0710.635\,071 0.671 4290.671\,429 0.94 0.783 3330.783\,333 Gemini 3 Flash 0.67 0.520 0000.520\,000 0.584 2700.584\,270 0.745 4550.745\,455 0.820 0000.820\,000 0.780 9520.780\,952 Gemini 2.5 Flash 0.650 0000.650\,000 0.78 0.71 0.76 0.909 0910.909\,091 0.83 Gpt 4.1 0.634 1460.634\,146 0.520 0000.520\,000 0.571 4290.571\,429 0.745 9020.745\,902 0.910 0000.910\,000 0.819 8200.819\,820 TABLE I: Shows the per-class precision, recall, and f1 scores for anomaly and hazard classes, denoted P, R, and Fâ1F1 respectively with the subscript indicating the class. Reasoning Inputs Hamming Loss Chain-of- Thought Image 0.36 0.41 0.31 0.28 â 0.26 Captions 0.41 0.41 0.34 0.28 â 0.29 Few Shot Image 0.32 0.28 0.26 0.23 0.25 0.24 Captions 0.35 0.29 0.26 0.25 0.27 0.27 Zero Shot Image 0.40 0.33 0.31 0.24 0.31 0.28 Captions 0.33 0.38 0.32 0.30 0.26 0.29 Ministral 3B Ministral 14B Mistral Large Gemini 3 Flash Gemini 2.5 Flash GPT 4.1 TABLE I: Shows the Hamming loss for all models for image and dense caption input modalities. We omit the results for Gemini 2.5 Flash and chain of thought prompting because the model did not produce valid responses for a significant number of queries. Lower values of hamming loss indicate a lower number of errors in jointly classifying anomalies and hazards. This metric serves to show how well modern VLMs can differentiate the concepts of anomaly and hazard. For our evaluation, we first treat each class individually in order to get the per-class performance, rather than performing evaluations jointly for both anomaly and hazard. This allows assigning partial correctness in model outputs. These performance metrics are summarized in tableËI. Image vs. Dense Captioning Results. FiguresË2 and 3 show that most models evaluated perform fairly well, especially with hazard classification given F1 scores tended to be above 0.7. As expected, the smallest model(Mistral-3B) is the worst-performing with its F1 score being as low as 0.31 when using the zero-shot prompting method and dense captions as the input modality. Contrary to our expectations, we note that agents tended to perform worse when using dense captions as the input modality with F1 scores skewing lower. We suspect that this is for a few reasons. It is possible that although models may have the capability to perform the high level reasoning necessary for this classification task, the dense captions may simply not focus enough on the context necessary for the classification. Although the captions were ensured to be accurate for each image, they were made to be neutral about whether the image was hazardous or anomalous. The agent generating the dense captions, Gemini 3 Flash, was given no information about the classification task in order for it to not imbue a bias on other models when they classify anomalies or hazards. As such, these captions may lack the necessary focus to allow all models the best chance to produce accurate results. Regardless, model performance using images as the input modality indicates that current VLMs are indeed approaching the required level of reasoning capability to effectively classify anomalies and hazards. We observed that even small models like Ministral 3B can, with the right prompting technique, effectively classify anomalies and hazards. For instance, the model achieved an F1 score of 0.68 using chain-of-thought prompting for the anomaly classification and 0.77 in the hazard classification few shot prompting. As shown in fig.Ë2, agents like Gemini 2.5 and 3 flash reached F1 scores as high as 0.85 on the hazard classification. It should also be noted that although the few-shot prompt was generally best for classifying hazards, there was no universally best technique across the models for anomaly classifications. Usually, either the chain-of-thought or few-shot prompt gave the best results for the latter classification task depending on the model. For instance, Ministral 3B performed best when using chain-of-thought prompting vs. Ministral 14B that performed better with few-shot prompting. These results suggest that the specific model and prompt can significantly affect performance in deployments like our prior work [12], and factors like the frequency of anomalies or hazards, and the criticality of each class need to be considered when choosing model/prompt combination. Another interesting feature of both fig.Ë2 and fig.Ë3 is how anomaly classification performance is consistently lower than hazard classification performance. This might suggest that semantically, the concept of anomaly is less well-defined. We note that even when annotating our dataset, we tended to have less consensus on the classification for anomalies. These images were carefully considered and excluded where consensus could not be reached. Figure 2: The precision vs. recall graph when using images as the input modality for the agents. Experiment results where an agent did not produce a valid response for more than 10 images per prompting technique were omitted. Generally, the agents perform well when looking at anomaly and hazard classifications independently. The outlier was mistral-3b when using zero-shot prompting. Figure 3: The precision vs. recall graph when using dense captions as the input modality for the agents. As with fig.Ë2, experiments with more than 10 invalid classifications per prompt technique were omitted. Contrary to our expectation, agents performed similarly to when the input modality was an image with some results being worse. A potential cause may be that textual descriptions may be too coarse for the level of information required for the anomaly vs. hazard classification. Joint Classification Performance. We now consider model performance in jointly classifying images and dense captions into the four categories of our dataset. For this analysis, we use hamming loss [7] which averages the number of errors in labeling an image safe, anomalous, hazardous, or both over all input images. These scores are shown in tableËI, with the lower scores indicating the best performance. We note that models from the Gemini family and GPT 4.1 consistently perform the best, indicating that these models have the greatest ability to differentiate the concept of an anomaly from a hazard. However, models from the Mistral family had lower performance with worse scores corresponding to smaller models. This implies that when deploying models to classify anomalies and hazards in actual applications, larger models may be more ideal. However, such deployments often require real-time performance or are resource constrained, especially when using edge-computing on robots, underscoring the need for more efficient algorithms Furthermore, the minimum proportion of errors 0.23 when images and few shot prompting with Gemini 3. This indicates that even though performance can be improved by optimizing the prompt, input, and model selection, performance with current algorithms still needs to be improved when the goal is joint classification and differentiating truly hazardous images from merely anomalous ones. This is especially true for safety-critical deployments like in our previous study [12]. (a) Error in âDangerousâ classifications (b) Error in âSafeâ classifications Figure 4: Shows how much models underpredicted or overpredicted images classified as âsafe" or âdangerousâ extracted from the VERI dataset [5]. BSTS represents the binary classification prompt used in [5]. The dataset had 100 images labeled âsafe.â Similar to previous findings, most of the prompting techniques tend to under-predict safe images [5], however, our zero-shot prompt is significantly more accurate on average for most models. Baseline Comparison. As noted, we observed the best results with images at the input modality. We now review the results when comparing our methodology with previous work on a similar task [5]. FigureË4(a) and fig.Ë4(b) show the number of over or under-predicted images that each model classified for the âdangerousâ and âsafeâ categories using the VERI dataset [5]. This dataset contains 100 images of each class, all of which were used in raw form rather than as captions. Unlike in the classification for our dataset, the few-shot technique performs poorly when results are interpreted in a binary manner. This is especially apparent for Ministral 14B. A common reason for misclassifications for many of the over-predicted images seemed to be misinterpretation of the context in images. For instance, one image contained a person with ketchup on them which was interpreted as blood, leading to an incorrect classification. Such errors undermine the systemâs ability to increase awareness to users when deployed in safety-critical systems; however, false positives such as these are still preferable to false negatives which are more likely to impact safety. Despite some prompt and model combinations performing poorly, we note that our Zero-Shot prompt proved significantly more accurate than the baseline BSTS prompt [5]. A t-test for independence was performed, assuming the null hypothesis that both the Zero-Shot and BSTS prompt accuracies were equal for all models tested. This resulted in a p-value of 0.0022, well past a typical threshold of 0.05 to reject the null hypothesis. Therefore, with the right combination of the model and prompt, the systemâs performance can be significantly improved over prior work. V Limitations Our study identifies systematic deficiencies in joint classification of anomalous and hazardous scenes, but several limitations restrict the generality of the results. Dataset size. The dataset contains 610 images, with Hazardous being the smallest category; balanced analysis is limited to roughly 100 samples per class, which amplifies the effect of noise and restricts the diversity of visual patterns the model encounters. Conclusions about hazard recognition should therefore be treated as preliminary, and it is difficult to tell whether observed performance patterns reflect systematic reasoning biases or dataset-specific artifacts. A larger-scale evaluation would yield higher statistical power and enable analysis of generalization. Annotations. Ground-truth labels come from a few non-expert annotators, introducing subjectivity for distinctions as subtle as anomalous vs. hazardous vs. anomalous-hazardous. Our definitions are grounded in OSHA standards [27], but personal notions of âweird,â âunsafe,â and âdangerousâ vary across individuals. Future work should formalize the annotation protocol via large-scale crowd-sourcing (e.g., Amazon Mechanical Turk) to compute consensus measures and produce probabilistic ground truth, reducing annotator bias. VI Conclusion This study sought to improve VLM performance in applications such as our prior work [12] by creating a dataset of images from safety-critical and benign scenes. We developed a four class classification scheme for images: âsafe,â âanomalous,â âhazardous,â and âanomalous-hazardous.â Furthermore, we hypothesized that forcing VLMs to separate scenes that are merely anomalous to scenes that are truly hazardous and require immediate action would help focus usersâ attention to critical details. Recognizing that VLM performance varies significantly based on model size, prompt, and input, we constructed several sub-tasks for the classification task that used dense captions and raw images as input modalities, and three different prompting techniques for each input modality. When considering per-class classification performance, we found that models perform moderately well with the best models having f1 scores around 0.70 for classifying anomalies and scores as high as 0.85 for classifying hazards. When considering joint classification performance with hamming loss, we found that the best score was 0.23 when using few shot prompting with Gemini 3 and images as the input. This indicated that performance can be improved by tuning pipeline parameters like the model, prompt, and input modality; however, for safety critical tasks like disaster response, current models struggle with the complex reasoning required for the four-class taxonomy. We found that Choi et al. pursued a similar approach to ours [5], and compared our pipeline with their approach. For this, we collapsed the four-class taxonomy to a binary approach by interpreting anomalies as âsafe,â and anomalous-hazardous scenes as âdangerousâ to match their dataset. We found that significant improvements to classification accuracy can be made by tuning the prompt and found that our zero shot prompt was significantly more accurate overall compared to the prior work [5]. This may indicate that even if classes are interpreted in a binary manner, forcing models to use the four-class taxonomy helps their reasoning. Future work should extend this analysis across a wider set of models and incorporate crowd-sourced annotations to build more reliable ground truth. Additional techniques like knowledge graphs of scenes may be incorporated to further assist in reasoning and classification. Such techniques may capture more context than the dense captions that we use for this study and assist models in focusing on only the relevant information for classification. Ultimately, the goal is to guide the development of systems that can distinguish unusual from dangerous conditions with greater consistency and reliability, enabling safer deployments of AI in high-risk environments. References [1] M. Acharya, A. Roy, and S. Jha (2024-04) COCO-OOC dataset. Zenodo. External Links: Document Cited by: §I. [2] K. Amara, L. Klein, C. LĂźth, P. Jäger, H. Strobelt, and M. El-Assady (2024-10) Why context matters in VQA and Reasoning: Semantic interventions for VLM input modalities. arXiv. External Links: 2410.01690, Document Cited by: §I, §I. [3] S. Chandhok (2024-08-13)SceneGPT: A Language Model for 3D Scene Understanding(Website) External Links: 2408.06926, Document, Link Cited by: §I. [4] S. Chen, T. Zhu, R. Zhou, J. Zhang, S. Gao, J. C. Niebles, M. Geva, J. He, J. Wu, and M. Li Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas. Cited by: §I. [5] D. Choi, S. Lee, and Y. Song (2025) Better safe than sorry? overreaction problem of vision language models in visual emergency recognition. arXiv preprint arXiv:2505.15367. External Links: Document Cited by: §I, §I, §I, §I, Figure 4, Figure 4, §IV, §IV, §VI. [6] J. M. Deniz, A. S. Kelboucas, and R. B. Grando (2024) Real-time robotics situation awareness for accident prevention in industry. External Links: 2409.15305, Document Cited by: §I. [7] S. Destercke (2014) Multilabel Prediction with Probability Sets: The Hamming Loss Case. In Information Processing and Management of Uncertainty in Knowledge-Based Systems, Vol. 443, p. 496â505. External Links: Document, ISBN 978-3-319-08854-9 978-3-319-08855-6 Cited by: §IV. [8] V. Di Pasquale, V. De Simone, S. Miranda, and S. Riemma (2021) Smart operators: how industry 4.0 is affecting the workerâs performance in manufacturing contexts. Procedia Computer Science 180, p. 958â967. External Links: ISSN 1877-0509, Document Cited by: §I. [9] Roboflow Note: Computer Vision Management Platform External Links: Link Cited by: §I. [10] A. Elhafsi, R. Sinha, C. Agia, E. Schmerling, I. A. D. Nesnas, and M. Pavone (2023-12) Semantic anomaly detection with large language models. Autonomous Robots 47 (8), p. 1035â1055. External Links: ISSN 1573-7527, Document Cited by: §I. [11] M. R. Endsley and E. S. Connors (2008-07) Situation awareness: state of the art. In 2008 IEEE Power and Energy Society General Meeting - Conversion and Delivery of Electrical Energy in the 21st Century, Pittsburgh, PA, USA, p. 1â4. External Links: Document Cited by: §I. [12] M. Eskandari, M. K. V. Indukuri, S. M. Lukin, and C. Matuszek (2025) LLM-supported safety annotation in high-risk environments. In Proceedings of the HRI 2025 Workshop on Vision-Based Assistance for Manipulation in Human-Robot Interaction (VAM-HRI), External Links: Link Cited by: §IV, §IV, §VI. [13] P. Fratczak, Y. M. Goh, P. Kinnell, A. Soltoggio, and L. Justham (2019) Understanding human behaviour in industrial humanârobot interaction by means of virtual reality. In Proceedings of the Halfway to the Future Symposium 2019, Nottingham, United Kingdom, p. 1â7. External Links: ISBN 978-1-4503-7203-9, Link, Document Cited by: §I. [14] grad (2024) Interior computer vision dataset. External Links: Link Cited by: §I. [15] W. Huang (2015) When hci meets hri: the intersection and distinction. Virginia Polytechnic Institute and State University. Cited by: §I. [16] A. S. H. Identifier (2025) Safety hazard identification computer vision model. External Links: Link Cited by: §I. [17] V. Jayawardene, T. Huggins, R. Prasanna, and B. Fakhruddin (2021-08) The role of data and information quality during disaster response decision-making. Progress in Disaster Science 12, p. 100202. External Links: Document Cited by: §I. [18] F. Jentsch (2016) Human-robot interactions in future military operations. CRC Press. Cited by: §I. [19] Y. Kim, A. S. Abdelrahman, and M. Abdel-Aty (2025) VRU-accident: a vision-language benchmark for video question answering and dense captioning for accident scene understanding. External Links: 2507.09815 Cited by: §I. [20] J. Ko (2022) Firefighter computer vision dataset. External Links: Link Cited by: §I. [21] krauseswelt (2024) Aerial computer vision model. External Links: Link Cited by: §I. [22] A. Kritharoula, M. Lymperaiou, and G. Stamou (2023) Large Language Models and Multimodal Retrieval for Visual Word Sense Disambiguation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, p. 13053â13077. External Links: 2310.14025, Document Cited by: §I. [23] S. Lin, C. Wang, X. Ding, Y. Wang, B. Du, L. Song, C. Wang, and H. Liu (2025) A vlm-based method for visual anomaly detection in robotic scientific laboratories. arXiv preprint arXiv:2506.05405. External Links: Link, Document Cited by: §I. [24] S. Misra, P. Roberts, and M. Rhodes (2020) Information overload, stress, and emergency managerial thinking. International Journal of Disaster Risk Reduction 46, p. 101615. External Links: Document, Link Cited by: §I. [25] model v2 (2023) Natural disaster damage computer vision model. External Links: Link Cited by: §I. [26] Nhyan (2025) Chemical spill computer vision model. External Links: Link Cited by: §I. [27] Occupational Safety and Health Administration (OSHA) (2018) Hazard assessment and job safety analysis. Note: U.S. Department of Labor, Occupational Safety and Health Administration Cited by: §I, §V. [28] M. Olbrich, H. Graf, J. Keil, R. Gad, S. Bamfaste, and F. Nicolini (2018) Virtual reality based space operations â a study of esaâs potential for vr based training and simulation. In Virtual, Augmented and Mixed Reality: Interaction, Navigation, Visualization, Embodiment, and Simulation, J. Y. C. Chen and G. Fragomeni (Eds.), Vol. 10909, p. 438â451. External Links: Document Cited by: §I. [29] P. Rahmanzadehgervi, L. Bolton, M. R. Taesiri, and A. T. Nguyen (2024-12) Vision language models are blind. In Proceedings of the Asian Conference on Computer Vision (ACCV), p. 18â34. Cited by: §I. [30] P. RĂśttger, G. Attanasio, F. Friedrich, J. Goldzycher, A. Parrish, R. Bhardwaj, C. D. Bonaventura, R. Eng, G. E. K. Geagea, S. Goswami, J. Han, D. Hovy, S. Jeong, P. JeretiÄ, F. M. Plaza-del-Arco, D. Rooein, P. Schramowski, A. Shaitarova, X. Shen, R. Willats, A. Zugarini, and B. Vidgen (2025-01) MSTS: A Multimodal Safety Test Suite for Vision-Language Models. arXiv. External Links: 2501.10057, Document Cited by: §I. [31] S. D. B. SHAFIZAL (2025) Classification ppe computer vision model. External Links: Link Cited by: §I. [32] D. Sumanathilaka, N. Micallef, and J. Hough (2024-08) Assessing GPTâs Potential for Word Sense Disambiguation: A Quantitative Evaluation on Prompt Engineering Techniques. In 2024 IEEE 15th Control and System Graduate Research Colloquium (ICSGRC), p. 204â209. External Links: ISSN 2833-1028, Document Cited by: §I. [33] T. Thrush, R. Jiang, M. Bartolo, A. Singh, A. Williams, D. Kiela, and C. Ross (2022) Winoground: probing vision and language models for visio-linguistic compositionality. arXiv. Note: Version Number: 2 External Links: Link, Document Cited by: §I. [34] R. Wadhawan, H. Bansal, K. Chang, and N. Peng (2024-07) ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models. arXiv. External Links: 2401.13311, Document Cited by: §I. [35] Z. Weng, L. Gomez, T. W. Webb, and P. Bashivan (2025-05) Caption This, Reason That: VLMs Caught in the Middle. arXiv. External Links: 2505.21538, Document Cited by: §I. [36] D. Xiao, M. Dianati, P. Jennings, and R. Woodman (2025-05) HazardVLM: A Video Language Model for Real-Time Hazard Description in Automated Driving Systems. IEEE Transactions on Intelligent Vehicles 10 (5), p. 3331â3343. External Links: ISSN 2379-8904, Document Cited by: §I. [37] H. Xu, G. Ghosh, P. Huang, P. Arora, M. Aminzadeh, C. Feichtenhofer, F. Metze, and L. Zettlemoyer (2021-09) VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding. arXiv. External Links: 2105.09996, Document Cited by: §I. [38] J. H. Yae, N. C. Skelly, N. C. Ranly, and P. M. LaCasse (2025-02) Leveraging large language models for word sense disambiguation. Neural Computing and Applications 37 (6), p. 4093â4110. External Links: ISSN 0941-0643, 1433-3058, Document Cited by: §I. [39] T. Yang, T. Jordan, N. Liu, and J. Sun (2025) Common inpainted objects in-n-out of context. arXiv:2506.00721. External Links: Document, Link Cited by: §I, §I. [40] X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al. (2024-06) MMMU: a massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 9556â9567. Cited by: §I. [41] J. Zhu, S. Cai, F. Deng, B. C. Ooi, and J. Wu (2024) Do llms understand visual anomalies? Uncovering llmâs capabilities in zero-shot anomaly detection. In Proceedings of the 32nd ACM International Conference on Multimedia, Mm â24, New York, NY, USA, p. 48â57. External Links: Document Cited by: §I.