Paper deep dive
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
Junqi Yang, Yuecong Min, Jie Zhang, Shiguang Shan, Xilin Chen
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 6:24:08 AM
Summary
INFACT is a diagnostic benchmark for evaluating faithfulness and factuality hallucinations in Video-LLMs, comprising 9,800 QA instances across real and synthetic videos. It introduces a four-mode evaluation protocol (Base, Visual Degradation, Evidence Corruption, and Temporal Intervention) and metrics (Resist Rate and Temporal Sensitivity Score) to measure model reliability and temporal grounding.
Entities (6)
Relation Signals (4)
INFACT → evaluates → Video-LLMs
confidence 100% · INFACT evaluates models in four modes: Base (clean), Visual Degradation, Evidence Corruption, and Temporal Intervention
INFACT → measures → Faithfulness
confidence 95% · INFACT, a diagnostic benchmark comprising 9,800 QA instances with fine-grained taxonomies for faithfulness and factuality
INFACT → measures → Factuality
confidence 95% · INFACT, a diagnostic benchmark comprising 9,800 QA instances with fine-grained taxonomies for faithfulness and factuality
Resist Rate → quantifies → Reliability
confidence 90% · Reliability under induced modes is quantified using Resist Rate (RR)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Despite rapid progress, Video Large Language Models (Video-LLMs) remain unreliable due to hallucinations, which are outputs that contradict either video evidence (faithfulness) or verifiable world knowledge (factuality). Existing benchmarks provide limited coverage of factuality hallucinations and predominantly evaluate models only in clean settings. We introduce \textsc{INFACT}, a diagnostic benchmark comprising 9{,}800 QA instances with fine-grained taxonomies for faithfulness and factuality, spanning real and synthetic videos. \textsc{INFACT} evaluates models in four modes: Base (clean), Visual Degradation, Evidence Corruption, and Temporal Intervention for order-sensitive items. Reliability under induced modes is quantified using Resist Rate (RR) and Temporal Sensitivity Score (TSS). Experiments on 14 representative Video-LLMs reveal that higher Base-mode accuracy does not reliably translate to higher reliability in the induced modes, with evidence corruption reducing stability and temporal intervention yielding the largest degradation. Notably, many open-source baselines exhibit near-zero TSS on factuality, indicating pronounced temporal inertia on order-sensitive questions.
Tags
Links
- Source: https://arxiv.org/abs/2603.11481v1
- Canonical: https://arxiv.org/abs/2603.11481v1
Trouble viewing inline? Open PDF directly →
Full Text
58,327 characters extracted from source content.
Expand or collapse full text
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs Junqi Yang 1,2 , Yuecong Min 1 , Jie Zhang 1 , Shiguang Shan 1 , Xilin Chen 1 , 1 State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences 2 School of Advanced Interdisciplinary Sciences, UCAS Abstract Despite rapid progress, Video Large Language Models (Video-LLMs) remain unreliable due to hallucinations, which are outputs that contra- dict either video evidence (faithfulness) or ver- ifiable world knowledge (factuality). Existing benchmarks provide limited coverage of fac- tuality hallucinations and predominantly eval- uate models only in clean settings. We intro- duce INFACT, a diagnostic benchmark com- prising 9,800 QA instances with fine-grained taxonomies for faithfulness and factuality, span- ning real and synthetic videos. INFACT eval- uates models in four modes: Base (clean), Vi- sual Degradation, Evidence Corruption, and Temporal Intervention for order-sensitive items. Reliability under induced modes is quantified using Resist Rate (R) and Temporal Sensitiv- ity Score (TSS). Experiments on 14 representa- tive Video-LLMs reveal that higher Base-mode accuracy does not reliably translate to higher reliability in the induced modes, with evidence corruption reducing stability and temporal in- tervention yielding the largest degradation. No- tably, many open-source baselines exhibit near- zero TSS on factuality, indicating pronounced temporal inertia on order-sensitive questions. 1 Introduction VideoLargeLanguageModels(Video- LLMs) (OpenAI, 2025; Gemini, 2025; Bai et al., 2025a; Wang et al., 2025) have made rapid progress in video understanding in recent years, demonstrating impressive capabilities across broad tasks (Li et al., 2024; Fu et al., 2025; Shangguan et al., 2025; Shafique et al., 2025; Kulkarni et al., 2024). Despite these advances, the reliability of Video-LLMs in downstream applications is compromised by hallucinations (Li et al., 2025a; Zhang et al., 2024), generating content that contra- dicts the provided video evidence (faithfulness) or verifiable world knowledge (factuality). Recent benchmarks have begun to probe hal- lucinations in Video-LLMs, e.g., by targeting event/motion-centric failures (Zhang et al., 2024; Kong et al., 2025), using controlled contrastive setups (Li et al., 2025a), or studying prior-driven shortcuts (Bae et al., 2025). However, existing efforts predominantly emphasize video-verifiable inconsistencies, leaving factuality hallucinations substantially under-explored. Moreover, high per- formance in clean scenarios does not guarantee low hallucination rates, as models may exploit shortcuts such as language priors or static cues. This moti- vates the evaluation of hallucinations beyond clean settings through controlled evidence perturbations. To bridge these gaps, we introduce INFACT, a diagnostic benchmark for evaluating Video-LLMs hallucinations regarding faithfulness and factual- ity in both clean and noisy scenarios. Specifically, INFACT establishes fine-grained taxonomies and comprises 9,800 QA instances sapnning real and synthetic videos, covering varying temporal dy- namics for faithfulness and diverse knowledge cat- egories for factuality. Furthermore, it supports four evaluation modes: Base (I), Visual Degradation (I), Evidence Corruption (I), and Temporal Inter- vention (IV) for order-sensitive items. The non-Base modes apply controlled video per- turbations while keeping questions fixed. Modes I– I are invariant-label settings and are evaluated by Resist Rate (R), which measures whether correct Base decisions remain stable under visual degrada- tion or corrupted evidence. Mode IV disrupts the temporal structure required for correctness and is evaluated using Temporal Sensitivity Score (TSS), which measures whether the model’s predictions changes after shuffling or reversal. Our evaluation reveals that reliability under induced conditions is not uniform: models tend to be more fragile when exposed to misleading evidence than to purely per- ceptual degradation. Moreover, temporal interven- tions reveal that many models remain largely in- arXiv:2603.11481v1 [cs.CV] 12 Mar 2026 Domain Knowledge Physics Knowledge Procedural Knowledge Factuality Hallucination [Q] Are the events physically plausible? [Q] Is the sequence of steps shown in the video logical about the task <ReplaceCarDoorHandle>? Spatial Temporal Structure Static Entities & Attributes Dynamics Actions & Motions Faithfulness Hallucination [Q] What is the color of the second object that enters the scene? [Q] In what direction(s) do the children on the merry-go-roundrotate? [Q] WhichOptioncorrectlysortstheheroines‘eventschronologically? 1. Study; 2. Make the bed; 3. Eat breakfast; 4. Exercise at home. [Q] Whatisthefestivalinthevideointhevideousuallycelebrated? A. green B. yellow C. blue D. cyan A. Clockwise then counter-clockwise.B. Clockwise throughout. C. Counter-clockwise then clockwise. D. Counter-clockwise throughout. A. 2134 B. 2431C. 2314D. 3214 A. 9th day of 9th lunar month B. 8th day of 12th lunar month C. 5th day of 5th lunar monthD.1st day of 1st lunar month A. Yes, the sequence is correct B. No, the sequence is incorrect A. Yes B. No [Q] Which statement best describe the adherence to physical laws? A. Fully plausibleB. Violation of mechanical dynamics C. Violation of material properties D. Violation of fluid dynamics Figure 1: Examples of Faithfulness and Factuality Hallucinations. Top: Faithfulness items are verified by video evidence, covering Static Entities & Attributes, Dynamic Actions & Motions, and Spatio-Temporal Relations. Bottom: Factuality items require consistency with world knowledge and cover Domain Knowledge (Know-WHAT), Procedural Knowledge (Know-HOW), and Physical Knowledge (Know-WHY). sensitive to order disruption in factuality questions, indicating to a gap in temporal grounding beyond clean-scenario accuracy. Our contributions are three-fold: 1.We introduce INFACT, a diagnostic bench- mark comprising 9,800 QA instances span- ning real and synthetic videos, with fine- grained taxonomies covering both faithfulness and factuality hallucinations. 2.We propose a four-mode evaluation protocol (Base, Visual Degradation, Evidence Corrup- tion, Temporal Intervention) with paired reli- ability metrics (R and TSS) for measuring invariant-label stability and temporal sensitiv- ity. 3.We conduct a systematic evaluation of 14 rep- resentative Video-LLMs, revealing their sta- bility under invariant-label perturbations and temporal inertia on order-sensitive items. 2 Related Works 2.1 Hallucination evaluation in Video-LLMs A growing line of work has benchmarked halluci- nations in Video-LLMs from different perspectives. Existing benchmarks (Zhang et al., 2024; Fu et al., 2025; Li et al., 2025a; Bae et al., 2025; Sung-Bin et al., 2025; Wang et al., 2024b; Kong et al., 2025) predominantly focus on faithfulness hallucinations, typically through controlled benchmark construc- tions that manipulate events, motion, video simi- larity, or cross-modal consistency to test ground- ing in the input video. For instance, EventHallu- sion (Zhang et al., 2024) focuses on event-level dynamics and relations, MHBench (Kong et al., 2025) targets motion-related errors with adversar- ial triplets, and VidHalluc (Li et al., 2025a) con- structs visually distinct yet semantically similar video pairs to expose fragile grounding. Related controlled settings further investigate shortcut be- haviors driven by priors or spurious correlations, such as narrative priors in NOAH (Lee et al., 2025), action-scene correlations in UNSCENE (Bae et al., 2025), and cross-modal inconsistency settings in AVHBench (Sung-Bin et al., 2025). Some bench- marks broaden data sources or verification scope: VideoHallu (Li et al., 2025b) introduces synthetic videos, while VidHallucer (Wang et al., 2024b) dis- tinguishes cases by whether the target claim can be verified from the video. However, as summarized in Table 1, most ex- Table 1: Comparison of INFACT with recent hallucination benchmarks for video understanding. INFACT covers video-grounded faithfulness and world-knowledge-grounded factuality with fine-grained taxonomies. It also supports controlled evaluation modes for visual degradation, evidence corruption, and temporal intervention within a unified protocol. Benchmark # Ques. / # Videos Faithfulness FactualitySource Visual DegradationEvidence CorruptionTemporal Intervention Motion Blur Gaussian Noise Caption Injection Adversarial Noise ShuffleReverse VIDHAL(Choong et al., 2024)– / 400✓(5)✗Real✗ EventHallusion(Zhang et al., 2024)– / 400✓(3)✗Real✗ VideoHallucer(Wang et al., 2024b)1,800 / 948✓(3)✓(3)Real✗ VIDHALLUC(Li et al., 2025a)9,295 / 5,002✓(3)✗Real✗ VideoHallu(Li et al., 2025b)3,233 / 3,233✓(2)✓(4)Synthetic✗ OURS9,800 / 9,800✓(12)✓(12)Real & Synthetic✓ isting benchmarks concentrate on hallucinations that can be judged against the input video (faithful- ness), whereas factuality hallucinations requiring verifiable world knowledge are much less explored. 2.2 Existing Video Understanding Benchmarks Existing video understanding benchmarks pri- marily evaluate Video-LLMs in clean settings with capability-oriented scoring. Broad-spectrum suites such as MVBench (Li et al., 2024), Video- MME (Fu et al., 2025), and TOMATO (Shang- guan et al., 2025) target general task coverage, while ViMUL-Bench (Shafique et al., 2025) and CityGuesser-style QA (Kulkarni et al., 2024) em- phasize knowledge-heavy settings. However, clean-input capability scores do not di- rectly measure reliability: high accuracy can mask shortcut-based success, in which models rely on language priors, static cues, or dataset biases rather than video-dependent evidence (Liu et al., 2025; Yan et al., 2025; Liu et al., 2024b; Bae et al., 2025). To reduce shortcut effects, several benchmarks adopt controlled designs such as conflicting videos or temporal multiple-choice setups (Liu et al., 2024b; Cores et al., 2025). Although useful for di- agnosing temporal dependence, these protocols are not formulated as hallucination evaluations: they do not separate evidence bases (video vs. verifi- able world knowledge) and do not probe reliability under explicit evidence-corruption conditions. 3 INFACT In this section, we introduce INFACT, a fine- grained benchmark designed to evaluate Video- LLMs faithfulness and factuality in both clean and noisy scenarios. We first detail the taxonomy of hal- lucinations in § 3.1 and the data construction pro- cess in § 3.2, describing the aggregation of 9,800 questions from public video-QA datasets, instruc- tional resources, and synthetic collections. To in- vestigate the root causes of model failure, § 3.3 introduces three hallucination induction modes de- vised to probe reliability under visual degradation, evidence corruption, and temporal intervention. Fi- nally, we introduce two kinds of evaluation metrics to quantify reliability in § 3.4. 3.1 Taxonomy We categorize hallucinations by the evidentiary ba- sis required for verification, where faithfulness re- quires alignment with visual content and factuality necessitates consistency with world knowledge. Faithfulnesshallucinations occur when model outputs contradict explicit visual evidence in the video.Following the common ob- ject/attribute/relation hierarchy (Liu et al., 2024a) used in static vision-language evaluation, we extend this framework to videos by incorporating temporal dynamics. This yields three hierarchical levels organized by increasing spatiotemporal complexity (Table A5). Level 1 (Static Entities & Attributes) corresponds to Object and Attribute tiers. This level requires local perception to resolve Entity Recognition, Unique Entity Counting, Temporal Attributes Recognition, Static Attributes Recognition, and Scene Text Recognition. Level 2 (Dynamic Actions & Motions) requires dynamic perception to aggregate visual features over time to resolve Action Recognition, Repetitive Action Counting and Motion Attributes Recognition. Level 3 (Spatio-Temporal Relations) corresponds to relation-level reasoning based on global spatio- temporal structure, including Spatial Relation Recognition, Temporal Relation Recognition, State Transition Detection, and Temporal Localization. Factualityhallucinations occur when model out- puts contradict world knowledge, requiring infor- VideoQA Datasets Synthetic Videos Instructional Datasets Data Collection Induction Designs Hallucination Induction Human-in-the-loop Quality Verification Taxonomy & Filtration ❌ Filter Rule Refinement Factuality QA Pairs Faithfulness QA Pairs Dataset Filtration Base Mode (Clean Scenario) Temporal Intervention Closing the door. Visual DegradationEvidence Corruption Answerable w/o video Ambiguous Category Ambiguous QA Figure 2: Overview of the INFACT construction process. Left: Candidate videos and QA pairs are collected from multiple sources, including video QA datasets, instructional datasets, and synthetic videos. Middle: Samples are organized into fine-grained faithfulness and factuality dimensions, and filtered to remove ambiguous or non-video- grounded items, followed by human-in-the-loop quality verification. Right: The resulting benchmark supports four evaluation modes: Base, Visual Degradation, Evidence Corruption, and Temporal Intervention. mation beyond what is present in the video con- tent alone. We structure factuality into three cate- gories based on the knowledge required (Table A6). Domain Knowledge (Know-WHAT) evaluates consistency with verifiable world knowledge across diverse domains, including Cultural Event Recognition, Historical Background Identifica- tion, Geospatial Localization, and Entertainment- Related Recognition.Procedural Knowledge (Know-HOW) assesses instructional validity by verifying whether the described sequence of steps respects prerequisite relations and causal depen- dencies. This category covers Electronic, Mechan- ical, Domestic, and Clinical domains. Physical Knowledge (Know-WHY) probes the adherence of a model to fundamental physical laws, includ- ing Mechanical, Fluid Mechanics, Material Prop- erty, and Spatial-Temporal continuity. To ensure rigor, We restrict factuality to objectively verifiable knowledge and explicitly exclude subjective or am- biguous claims. Figure 3 summarizes the distribu- tion of samples across the fine-grained taxonomy for both faithfulness and factuality. 3.2 Data Construction Faithfulness Data. To construct a rigorous benchmark for hallucination evaluation, we curate high-quality samples from established video un- derstanding benchmarks, including MVBench (Li et al., 2024), Video-MME (Fu et al., 2025), and TOMATO (Shangguan et al., 2025), which provide diverse questions suitable for video evidence verifi- cation. To guarantee data quality, we employ a two- stage alignment process. First, we manually map the categories of source benchmarks to our pro- posed taxonomy. Subsequently, we implement a "taxonomy-aligned filtration" pipeline that utilizes LLM-assisted consensus labeling. This ensures strict adherence to the spatio-temporal complex- ity levels defined in Appendix Table A5, thereby refining the quality of constructed benchmark. Factuality Data. Factuality samples are con- structed to cover three verifiable knowledge re- quirements: Domain, Procedural, and Physical Knowledge. For Domain Knowledge, candidates are drawn from knowledge-intensive video QA re- sources such as CityGuessr68k (Kulkarni et al., 2024) and ViMULBench (Shafique et al., 2025) to ensure comprehensive coverage across diverse domains such as culture, history, geography, and entertainment. For Procedural Knowledge, in- structional videos are sourced from COIN (Tang et al., 2019) and MedVidQA (Gupta et al., 2023), where step annotations provide temporal bound- 11.4% 9.7% 4.1% 4.6% 14.7% 10.1% 14.8% 6.4% 10.8% 4.1% 6.4% 32.6% 39.7% 27.7% Faithful 4,849 Static Entities & Attributes Entity Recognition (ER) Unique Entity Counting (UEC) Temporal Attributes Recognition (TAR) Static Attributes Recognition (SA) Scene Text Recognition (STR) Dynamic Actions \& Motions Action Recognition (AR) Repetitive Action Counting (RAC) Motion Attributes Recognition (MAR) Spatial-Temporal Relations Spatial Relation Recognition (SRR) Temporal Relation Recognition (TRR) State Transition Detection (STD) Temporal Localization (TL) 4.3% 4.0% 11.9% 3.8% 14.3% 13.9% 22.0% 7.8% 4.5% 4.5% 4.5% 4.5% 24.0% 58.0% 18.0% Factual 4,451 Domain Knowledge Cultural Event Recognition (CER) Historical Background Identification (HBI) Geospatial Localization (GL) Entertainment-Related Recognition (ERR) Procedural Knowledge Electronic Procedural Judgement (EPJ) Mechanical Procedural Judgement (MPJ) Domestic Procedural Judgement (DPJ) Clinical Procedural Judgement (CPJ) Physics Knowledge Newtonian Mechanical Reasoning (NMR) Fluid Mechanics Reasoning (FMR) Material Property Reasoning (MPR) Spatio-temporal Continuity Reasoning (STCR) Figure 3: Dataset composition of INFACT. Distribu- tion over the fine-grained taxonomy for faithfulness (top) and factuality (bottom). aries for each procedure segment. To create coun- terfactuals, we randomly shuffle these segments to disrupt the temporal order required by prerequisite relations and causal dependencies. The resulting videos support binary procedural judgement (cor- rect vs. incorrect). For Physical Knowledge, physi- cally implausible videos are synthesized using text- to-video generation models (e.g., Sora (OpenAI, 2024), Wan2.5 (Wan AI, 2025), and Gemini Veo 3 (Google DeepMind, 2025)). These samples vio- late physical laws, such as gravity-defying motion, serving as a test for physical reasoning capabilities. Dataset Filtration and Quality Review. An LLM-ensemble filter is first applied to remove sam- ples with linguistic ambiguity and those solvable without video evidence (e.g., answerable from text- only priors). Since source annotations and edge cases in taxonomy mapping may still introduce noise, a human-in-the-loop verification stage is further conducted. Rather than discarding individ- ual problematic samples only, annotators also iden- tify recurring error patterns and trace them back to their causes (e.g., a question template or a mapping rule). This enables iterative updates to the upstream prompts and filtering logic and batch correction of affected samples. The process is repeated until the benchmark satisfies the quality criteria. 3.3 Hallucination Induction Designs INFACT establishes four distinct evaluation modes to probe evidence-consistent behavior under con- trolled conditions. Mode I (Base) establishes the baseline performance on clean data. Modes I & I (Robustness) apply label-preserving transfor- mations to test whether models maintain correct answers under visual degradation or conflicting evidence. Mode IV (Sensitivity) applies label- changing interventions to test sensitivity to dis- rupted temporal structure. Mode I: Base (clean scenario).The Base mode evaluates the model on the original, unaltered video and question pairs. It serves as the reference stan- dard for establishing upper-bound performance. Crucially, to isolate the impact of interventions from intrinsic item difficulty, we employ a paired comparison protocol: behavior under all induced modes is assessed relative to the model’s Base out- come on the same specific item. Mode I: Visual Degradation. This mode probes hallucinations induced by perceptual un- certainty. We introduce three types of visual pertur- bations: (1) Gaussian noise and (2) Motion blur, both applied uniformly at the frame level, and (3) Video Compression, which simulates platform- specific re-encoding artifacts by reducing the bi- trate to a lossy threshold. While these operators re- duce low-level perceptual clarity, they preserve the high-level semantic content required to answer the question. Consequently, a reliable model should maintain the correct answer despite the visual noise, rather than introducing unsupported details when the evidence becomes harder to perceive. Mode I: Evidence Corruption. Since Video- LLMs often integrate auxiliary textual signals (e.g, subtitles) with visual data, reliability hinges on the ability to prioritize authentic visual evi- dence over misleading external cues. We em- ploy three operators to simulate untrusted condi- tioning while keeping the ground-truth label in- variant. Caption injection introduces a cross- modal conflict by overlaying subtitles on randomly sampled video segments. The injected subtitles mix (i) content-irrelevant sentences and (i) LLM- generated misleading statements conditioned on the question and the ground-truth answer option, constructed to form a plausible but incorrect tex- tual cue that conflicts with the video evidence re- quired for correctness (e.g., a video showing open- ing a door is paired with the subtitle closing the door). Subtitle corruption injects noisy ASR- like subtitles with subtitle–video desynchroniza- tion, simulating the unreliable OCR/ASR outputs that models may encounter when processing cap- tioned or subtitled content in the wild. Adversar- ial noise targets the visual encoder by applying a transfer-based black-box perturbation generated with MI-FGSM (Dong et al., 2018). To ensure the attack generalizes across different architectures, perturbations are produced using a proxy ensem- ble of visual encoders (InternVL3-8B (Zhu et al., 2025), Qwen3VL-8B (Bai et al., 2025a), and Video- MAE (Tong et al., 2022)) to reduce reliance on a single proxy model. Since the ground-truth label is intended to remain unchanged, the desired behavior is stability against misleading conditioning. Mode IV: Temporal Intervention. This mode probes temporal evidence consistency. Specifically, whether the model’s correctness genuinely stems from understanding event order and state transi- tions. Unlike the previous modes, here the temporal structure is intentionally destroyed, rendering the original ground truth invalid. To evaluate this, we curate an order-sensitive subset from Action Dy- namics, Spatio-Temporal Structure, and Procedural Knowledge, where the chronological sequence is strictly required for correctness. We apply tem- poral intervention to these videos via frame-level shuffling or reversal. If a model retains the Base prediction after intervention, it indicates temporal insensitivity and reliance on order-invariant cues. 3.4 Evaluation Metrics Accuracy is reported under the Base (clean- scenario) setting to measure performance on faith- fulness and factuality. To quantify reliability under induced conditions, we additionally report metrics tailored to invariant-label settings and temporal in- terventions. Modes I–I (Visual Degradation and Evidence Corruption) are evaluated by Resist Rate (R), while Mode IV (Temporal Intervention) is evaluated by Temporal Sensitivity Score (TSS). Resist Rate (R).For visual degradations (T deg ) and evidence corruptions (T cor ), the ground-truth label is intended to remain unchanged. We mea- sure reliability by whether a model preserves its correct Base prediction under each operator. For an operator p∈T pert =T deg ∪T cor , we define: R p = P i I f(V i ,q i )=y gt i ·I f(p(V i ),q i )=y gt i P i I f(V i ,q i )=y gt i , where(V i ,q i )denotes thei-th video-question pair,y gt i is the corresponding ground-truth an- swer,f (·,·)denotes a Video-LLM, andI(·)is the indicator function. We report operator-wise scores (e.g.,R gau ,R mb ,R adv ,R cap ), and computeRR deg andRR cor as the mean over degra- dation and corruption operators, respectively. Figure 4: Base accuracy vs. average reliability score under inductions. Base accuracy is measured in Mode I and averaged over faithfulness and factuality. The av- erage reliability score under induction aggregates R over Modes I–I and TSS over Mode IV. Temporal Sensitivity Score (TSS). Temporal interventions (shuffling or reversal) test whether model decisions are genuinely grounded in tem- poral structure. TSS is computed on the order- sensitive subsetS order , where temporal order is es- sential for correctness. Unlike the robustness tests used for R, these interventions effectively invali- date the original ground-truth label. Consequently, a temporally grounded model should diverge from its initial decision after intervention. We define TSS as the rate at which model ceases to predict the original ground-truth labely gt after interven- tion. For an intervention operatorp ∈ T iv (e.g., shuffling or reversal), we define: TSS p = P i∈S order I(f(V i ,q i )=y gt )·I(f(p(V i ),q i )̸=y gt ) P V i ∈S order I(f(V i ,q i )=y gt ) . We reportTSS shu andTSS rev , and ̄ TSSis their mean. A low TSS indicates temporal inertia: the model tends to retain the Base decision even when the supporting temporal evidence is destroyed, sug- gesting a reliance on static priors rather than tem- poral logic. Table 2: Faithfulness results on INFACT. Text-only: question-only accuracy. Base: Mode I. R: Modes I–I. TSS: Mode IV (order-sensitive subset). Models are grouped by availability and sorted by Avg Score. Evidence CorruptionVisual DegradationTemporal Intervention ModelText-only BaseRR adv R cap R sub R cor R cmp R gau R mb R deg TSS shu TSS rev ̄ TSS Avg Score PLLaVA-13B (Xu et al., 2024)0.2850.458 0.7820.6780.922 0.7940.7770.813 0.699 0.7630.0100.014 0.0120.523 VideoLLaMA2-7B(Cheng et al., 2024)0.2610.434 0.7510.5640.910 0.7420.8180.725 0.749 0.7640.1130.091 0.1020.536 ShareGPT4Video-8B (Chen et al., 2024)0.2920.462 0.7190.6880.894 0.7670.8400.762 0.754 0.7850.1010.094 0.0980.550 PLLaVA-34B (Xu et al., 2024)0.2870.505 0.7570.6180.920 0.7650.7720.874 0.8480.8310.0700.073 0.0720.556 Tarsier-34B (Wang et al., 2024a) 0.3020.568 0.7850.6440.960 0.7960.7590.846 0.612 0.7390.1580.208 0.1830.573 Qwen2.5VL-7B(Bai et al., 2025b)0.2780.531 0.6300.6590.973 0.7540.8040.727 0.640 0.7240.2300.258 0.2440.574 NVILA-8B (Liu et al., 2024c)0.2930.533 0.8040.6690.963 0.8120.7970.853 0.809 0.8200.1220.150 0.1360.589 Qwen3VL-8B(Bai et al., 2025a)0.2950.557 0.6760.7150.941 0.7770.8290.785 0.744 0.7860.2240.189 0.2070.590 Qwen2.5VL-32B(Bai et al., 2025b)0.2860.5380.8170.7240.949 0.8300.7900.824 0.773 0.7960.1980.204 0.2010.609 InternVL3-8B (Zhu et al., 2025)0.3010.576 0.7770.7090.979 0.8220.8280.847 0.8170.8310.1690.212 0.1910.615 InternVL3.5-8B (Wang et al., 2025)0.294 0.606 0.7960.7210.9780.8320.8410.7920.850 0.8280.1710.241 0.2060.622 Qwen3VL-32B (Bai et al., 2025a)0.2870.602 0.8090.7180.975 0.8340.8020.836 0.815 0.8180.2480.2890.2690.640 GPT-5.1 (OpenAI, 2025)0.3050.6870.8560.8230.9580.8790.9400.8010.8540.8650.3550.2810.3180.687 Gemini3-flash(Gemini, 2025)0.2950.784 0.8860.9140.962 0.9210.9420.859 0.892 0.8980.5810.492 0.5370.785 Table 3: Factuality results on INFACT. Same evaluation protocol and metrics as Table 2. Evidence CorruptionVisual DegradationTemporal Intervention ModelText-only BaseRR adv R cap R sub R cor R cmp R gau R mb R deg TSS shu TSS rev ̄ TSS Avg Score VideoLLaMA2-7B(Cheng et al., 2024)0.2720.407 0.6760.5900.917 0.7280.8180.723 0.772 0.7710.0000.000 0.0000.500 ShareGPT4Video-8B (Chen et al., 2024)0.2940.489 0.7150.5190.883 0.7060.8400.816 0.798 0.8180.0000.000 0.0000.508 PLLaVA-13B (Xu et al., 2024)0.2790.410 0.7880.5850.919 0.7640.7770.947 0.705 0.8100.0000.000 0.0000.525 NVILA-8B (Liu et al., 2024c)0.2910.424 0.8180.5410.961 0.7730.7970.888 0.759 0.8150.0000.000 0.0000.529 PLLaVA-34B (Xu et al., 2024)0.2650.431 0.7820.6030.920 0.7680.7720.897 0.812 0.8270.0000.000 0.0000.532 InternVL3-8B (Zhu et al., 2025)0.2950.465 0.7380.6120.951 0.7670.8280.914 0.813 0.8520.0000.000 0.0000.540 InternVL3.5-8B(Wang et al., 2025)0.2980.498 0.7750.6190.955 0.7830.8410.921 0.851 0.8710.0000.000 0.0000.551 Tarsier-34B (Wang et al., 2024a)0.2890.4410.9090.6250.9600.8310.7590.930 0.820 0.8360.0000.000 0.0000.556 Qwen2.5VL-32B(Bai et al., 2025b)0.2960.512 0.8040.6240.930 0.7860.7900.889 0.764 0.8140.1050.101 0.1030.568 Qwen2.5VL-7B(Bai et al., 2025b)0.2810.503 0.6700.575 0.968 0.7380.8040.784 0.774 0.7870.2120.262 0.2370.587 Qwen3VL-8B (Bai et al., 2025a)0.2960.541 0.7070.6060.929 0.7470.8290.793 0.768 0.7970.3390.3340.3370.627 Qwen3VL-32B (Bai et al., 2025a)0.3050.521 0.8090.6410.963 0.8040.8020.904 0.821 0.8420.2310.254 0.2430.630 GPT-5.1 (OpenAI, 2025)0.3080.678 0.9010.8140.9630.8930.9400.9110.8760.9090.3910.4140.4020.735 Gemini3-flash(Gemini, 2025)0.3050.752 0.9170.8390.969 0.9080.9420.902 0.881 0.9080.7410.696 0.7190.845 ER UEC TAR SA STR AR RAC MAR SRR TRR STD TL 20 40 60 80 100 0 Faithfulness CER HBI GL ERR EPJ MPJ DPJ CPJ NMR FMR MPR STCR 20 40 60 80 100 0 Factuality Gemini3-flashGPT-5.1InternVL3.5-8BQwen3VL-32B Figure 5: Comparison of four representative mod- els on fine-grained evaluation dimensions. The left radar plot shows performance on faithfulness dimen- sions, while the right radar plot shows performance on factuality dimensions. Each axis corresponds to a fine- grained category in the INFACT taxonomy, and each curve represents one representative model. Higher val- ues indicate better performance on the corresponding dimension. 4 Experiments 4.1 Setups Fourteen models are evaluated on INFACT, in- cluding two proprietary systems (GPT-5.1 (Ope- nAI, 2025), Gemini3-flash (Gemini, 2025)) and twelve open-source baselines: VideoLLaMA2- 7B (Cheng et al., 2024), PLLaVA-13B/34B (Xu et al., 2024), ShareGPT4Video-8B (Chen et al., 2024), Qwen2.5-VL-7B/32B (Bai et al., 2025b), Qwen3-VL-8B/32B (Bai et al., 2025a), Tarsier- 34B (Wang et al., 2024a), NVILA-8B (Liu et al., 2024c), InternVL3-8B (Zhu et al., 2025), and InternVL3.5-8B (Wang et al., 2025). All models are evaluated zero-shot with default prompts using 16 uniformly sampled frames per video (Shang- guan et al., 2025; Rawal et al., 2025). 4.2 Results and Analysis Base accuracy vs. reliability under induced modes. Figure 4 reveals a strong association be- tween Base accuracy (Mode I) and induced-mode reliability (Modes I–IV), averaged over faithful- ness and factuality (Pearsonr=0.978; Spearman ρ=0.969). However, rankings are not fully pre- served: models with similar Base accuracy can diverge under controlled perturbations. Tables 2–3 suggest that these gaps are driven by differences in stability (R) and temporal sensitivity (TSS), so clean accuracy alone is insufficient to characterize 28162432 Frames 0.30 0.35 0.40 0.45 0.50 0.55 Accuracy Qwen3vl-8B InternVL3.5-8B NVILA-8B Qwen3vl-8BInternVL3.5-8BNVILA-8B 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 R cor Qwen3vl-8BInternVL3.5-8BNVILA-8B 0.0 0.2 0.4 0.6 0.8 R deg Qwen3vl-8BInternVL3.5-8BNVILA-8B 0.00 0.05 0.10 0.15 0.20 0.25 TSS 8 frames16 frames24 frames32 frames Figure 6: Effect of the number of sampled frames. Left: Base accuracy under2, 8, 16, 24, 32uniformly sampled frames. Right: induced metrics (R deg , R cor , TSS) under8, 16, 24, 32 frames. evidence-consistent behavior. Stability under Visual Degradation & Evidence Corruption Inductions (R). Tables 2–3 show that evidence corruption degrades R more than visual degradation, and caption injection is typi- cally more damaging than transfer-based adversar- ial noise. For instance, VideoLLaMA2-7B drops fromRR adv =0.751toRR cap =0.564on faithful- ness, and Tarsier-34B drops fromRR adv =0.909 toRR cap =0.625on factuality. This pattern sug- gests that stability is more fragile when misleading auxiliary cues are introduced than when the visual signal is merely degraded. Sensitivity under Temporal Intervention (TSS). Tables 2–3 report TSS on the order-sensitive sub- set, where frame shuffling or reversal disrupts the temporal structure required for correctness. Sev- eral open-source baselines exhibit temporal inertia, most visibly on factuality where multiple models yield ̄ TSS=0 , with predictions unchanged after shuffling or reversal on Base-correct order-sensitive items. In contrast, Qwen3VL-8B and Gemini3- flash attain0.337and0.719, indicating that tem- poral intervention separates models by their re- liance on temporal evidence. Among open-source baselines, Qwen2.5/3-VL achieves comparatively higher TSS (Tables 2–3), highlighting the potential role of more explicit time-aligned spatiotemporal positional encoding (e.g., mRoPE-style designs). Performance across Fine-grained Dimensions. For finer diagnosis, we further analyze four repre- sentative models across fine-grained faithfulness and factuality dimensions (Figure 5). Across these models, factuality dimensions tend to lag behind faithfulness dimensions, especially in procedural judgement and physical reasoning. On the faithful- ness side, the remaining weaknesses concentrate on motion- and structure-centric dimensions (e.g., temporal localization and motion attributes), sug- gesting that fine-grained temporal aggregation re- mains challenging under our evaluation setting. Effect of Frame Sampling Rate. We conduct a frame-ablation study on 500 videos randomly sampled from INFACT (Figure 6). Base accu- racy is evaluated with2, 8, 16, 24, 32uniformly sampled frames and quickly saturates, where 16 frames already capture most of the gain (Figure 6, left). Under induced modes with8, 16, 24, 32 frames,R deg andRR cor vary modestly and TSS shows no consistent upward trend (Figure 6, right), suggesting that increasing the number of sampled frames alone does not consistently improve relia- bility under induced modes within this range, par- ticularly for temporal sensitivity. 5 Conclusion We present INFACT, a fine-grained benchmark for evaluating faithfulness and factuality halluci- nations in Video-LLMs under both Base mode and controlled evidence perturbations, including invariant-label inductions and temporal interven- tions on order-sensitive items. Across 14 models, factuality remains challenging for procedural judg- ment and physical reasoning, whereas faithfulness errors concentrate in motion- and structure-centric skills. Models also exhibit limited stability under invariant-label perturbations and pronounced tem- poral inertia under temporal intervention. Limitations The induced modes use controlled operators as proxies for deployment perturbations and do not cover the full space of corruptions or unreliable conditioning. Temporal Intervention is tested on an order-sensitive subset using shuffling and reversal; it probes temporal reliance but does not localize the cues driving model decisions. References Kyungho Bae, Jinhyung Kim, Sihaeng Lee, Soonyoung Lee, Gunhee Lee, and Jinwoo Choi. 2025. Mash- vlm: Mitigating action-scene hallucination in video- llms through disentangled spatial-temporal represen- tations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13744–13753. Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhi- fang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, and 45 others. 2025a. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631. Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wen- bin Ge, Sibo Song, Kai Dang, Peng Wang, Shi- jie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, and 8 others. 2025b. Qwen2.5-vl technical report. arXiv preprint arXiv:2502.13923. Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, Li Yuan, Yu Qiao, Dahua Lin, Feng Zhao, and Jiaqi Wang. 2024. Sharegpt4video: Improving video understanding and generation with better captions. Preprint, arXiv:2406.04325. Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, and Lidong Bing. 2024. Videollama 2: Advancing spatial-temporal model- ing and audio understanding in video-llms. arXiv preprint arXiv:2406.07476. Wey Yeh Choong, Yangyang Guo, and Mohan Kankan- halli. 2024.Vidhal:Benchmarking temporal hallucinations in vision llms.arXiv preprint arXiv:2411.16771. Daniel Cores, Michael Dorkenwald, Manuel Mucientes, Cees G. M. Snoek, and Yuki M. Asano. 2025. Lost in time: A new temporal benchmark for videollms. Preprint, arXiv:2410.07752. Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. 2018. Boosting adversarial attacks with momentum. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9185–9193. Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, Peixian Chen, Yanwei Li, Shaohui Lin, Sirui Zhao, Ke Li, Tong Xu, Xiawu Zheng, Enhong Chen, Caifeng Shan, and 2 others. 2025. Video-mme: The first-ever compre- hensive evaluation benchmark of multi-modal llms in video analysis. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR), pages 24108–24118. Gemini. 2025.Gemini-3-flash.https://blog. google/products/gemini/gemini-3-flash/. Google DeepMind. 2025. Veo.https://deepmind. google/models/veo/. Accessed: 2026-03-10. Deepak Gupta, Kush Attal, and Dina Demner-Fushman. 2023. A dataset for medical instructional video clas- sification and question answering. Scientific Data, 10(1):158. Ming Kong, Xianzhou Zeng, Luyuan Chen, Yadong Li, Bo Yan, and Qiang Zhu. 2025. Mhbench: Demystify- ing motion hallucination in videollms. Proceedings of the AAAI Conference on Artificial Intelligence, 39(4):4401–4409. Parth Parag Kulkarni, Gaurav Kumar Nayak, and Mubarak Shah. 2024. Cityguessr: City-level video geo-localization on a global scale. In Computer Vi- sion – ECCV 2024: 18th European Conference, Mi- lan, Italy, September 29–October 4, 2024, Proceed- ings, Part LXIII, page 293–311, Berlin, Heidelberg. Springer-Verlag. Kyuho Lee, Euntae Kim, Jinwoo Choi, and Buru Chang. 2025. Noah: Benchmarking narrative prior driven hallucination and omission in video large language models. Preprint, arXiv:2511.06475. Chaoyu Li, Eun Woo Im, and Pooyan Fazli. 2025a. Vid- halluc: Evaluating temporal hallucinations in multi- modal large language models for video understand- ing. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13723– 13733. Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Lou, Limin Wang, and Yu Qiao. 2024. Mvbench: A comprehensive multi-modal video un- derstanding benchmark. In 2024 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR), pages 22195–22206. Zongxia Li, Xiyang Wu, Guangyao Shi, Yubin Qin, Hongyang Du, Fuxiao Liu, Tianyi Zhou, Dinesh Manocha, and Jordan Lee Boyd-Graber. 2025b. Videohallu: Evaluating and mitigating multi-modal hallucinations on synthetic video understanding. arXiv preprint arXiv:2505.01481. Hanchao Liu, Wenyuan Xue, Yifei Chen, Dapeng Chen, Xiutian Zhao, Ke Wang, Liping Hou, Rongjun Li, and Wei Peng. 2024a. A survey on halluci- nation in large vision-language models. Preprint, arXiv:2402.00253. Jiazhen Liu, Yuhan Fu, Ruobing Xie, Runquan Xie, Xingwu Sun, Fengzong Lian, Zhanhui Kang, and Xirong Li. 2025. Phd: A chatgpt-prompted visual hallucination evaluation dataset. In Proceedings of the Computer Vision and Pattern Recognition Con- ference, pages 19857–19866. Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, and Lu Hou. 2024b. Tempcompass: Do video llms really understand videos? In Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, pages 8731–8772. Association for Computational Linguistics. Zhijian Liu, Ligeng Zhu, Baifeng Shi, Zhuoyang Zhang, Yuming Lou, Shang Yang, Haocheng Xi, Shiyi Cao, Yuxian Gu, Dacheng Li, Xiuyu Li, Yunhao Fang, Yukang Chen, Cheng-Yu Hsieh, De-An Huang, An- Chieh Cheng, Vishwesh Nath, Jinyi Hu, Sifei Liu, and 8 others. 2024c. Nvila: Efficient frontier visual language models. Preprint, arXiv:2412.04468. OpenAI. 2024. Sora: Creating video from text.https: //openai.com/index/sora/ . Accessed: 2026-03- 10. OpenAI. 2025. Gpt-5.1: A smarter, more conversational chatgpt. https://openai.com/index/gpt-5-1/. Ruchit Rawal, Reza Shirkavand, Heng Huang, Gowthami Somepalli, and Tom Goldstein. 2025. Ar- gus: Hallucination and omission evaluation in video- llms. Preprint, arXiv:2506.07371. Bhuiyan Sanjid Shafique, Ashmal Vayani, Muham- mad Maaz, Hanoona Abdul Rasheed, Dinura Dis- sanayake, Mohammed Irfan Kurpath, Yahya Hmaiti, Go Inoue, Jean Lahoud, Md. Safirur Rashid, Sha- did Intisar Quasem, Maheen Fatima, Franco Vi- dal, Mykola Maslych, Ketan Pravin More, Sanoo- jan Baliah, Hasindri Watawana, Yuhao Li, Fabian Farestam, and 10 others. 2025. A culturally-diverse multilingual multimodal video benchmark and model. Preprint, arXiv:2506.07032. Ziyao Shangguan, Chuhan Li, Yuxuan Ding, Yanan Zheng, Yilun Zhao, Tesca Fitzgerald, and Arman Cohan. 2025. Tomato: Assessing visual temporal reasoning capabilities in multimodal foundation mod- els. In International Conference on Representation Learning, volume 2025, pages 7593–7734. Kim Sung-Bin, Oh Hyun-Bin, Lee Jung-Mok, Arda Senocak, Joon Son Chung, and Tae-Hyun Oh. 2025. Avhbench: A cross-modal hallucination benchmark for audio-visual large language models. In Interna- tional Conference on Representation Learning, vol- ume 2025, pages 24244–24271. Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. 2019. Coin: A large-scale dataset for comprehensive instructional video analysis. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022. VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre- training. In Advances in Neural Information Process- ing Systems. Wan AI. 2025. Wan 2.5: Native audio like veo3 + 1080p video generation.https://wan25.ai/. Accessed: 2026-03-10. Jiawei Wang, Liping Yuan, Yuchen Zhang, and Hao- miao Sun. 2024a. Tarsier: Recipes for training and evaluating large video description models. Preprint, arXiv:2407.00634. Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, and 1 others. 2025. In- ternvl3. 5: Advancing open-source multimodal mod- els in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265. Yuxuan Wang, Yueqian Wang, Dongyan Zhao, Ci- hang Xie, and Zilong Zheng. 2024b.Videohal- lucer: Evaluating intrinsic and extrinsic hallucina- tions in large video-language models. arXiv preprint arXiv:2406.16338. Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, and Jiashi Feng. 2024.Pllava: Parameter-free llava extension from images to videos for video dense captioning. arXiv preprint arXiv:2404.16994. Bei Yan, Zhiyuan Chen, Yuecong Min, Jie Zhang, Ji- ahao Wang, Xiaozhen Wang, and Shiguang Shan. 2025. Shale: A scalable benchmark for fine-grained hallucination evaluation in lvlms. In Proceedings of the 33rd ACM International Conference on Multime- dia, pages 13442–13449. Jiacheng Zhang, Yang Jiao, Shaoxiang Chen, Na Zhao, Zhiyu Tan, Hao Li, Xingjun Ma, and Jingjing Chen. 2024.Eventhallusion: Diagnosing event hallucinations in video llms.arXiv preprint arXiv:2409.16597. Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, and 1 others. 2025. Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479. Open-Source Video-LLMsHF Checkpoint VideoLLaMA2-7BDAMO-NLP-SG/VideoLLaMA2-7B Qwen2.5VL-7BQwen/Qwen2.5-VL-7B-Instruct Qwen3VL-8BQwen/Qwen3-VL-8B-Instruct ShareGPT4Video-8BLin-Chen/sharegpt4video-8b NVILA-8BEfficient-Large-Model/NVILA-8B InternVL3-8BOpenGVLab/InternVL3-8B InternVL3.5-8BOpenGVLab/InternVL3_5-8B PLLaVA-13Bermu2001/pllava-13b Qwen2.5VL-32BQwen/Qwen2.5-VL-32B-Instruct Qwen3VL-32BQwen/Qwen3-VL-32B-Instruct PLLaVA-34Bermu2001/pllava-34b Tarsier-34Bomni-research/Tarsier-34b Table A1: Simplified configurations for open-source multimodal foundation models used in the evaluation. A Appendix A.1 Detailed Taxonomy and Examples In this appendix, we provide the comprehensive definitions and detailed examples for our proposed hallucination taxonomy. Table A5 details the three- level hierarchy for Faithfulness Hallucination. Ta- ble A6 presents the categorization for Factuality Hallucination. A.2 TSS: Order-Sensitive vs. Non-Order-Sensitive To validate that TSS provides meaningful sep- aration, Table A2 compares TSS computed on the order-sensitive subset (S order ) and non-order- sensitive items (S nonorder ) for models with non- trivial temporal responsiveness.For all mod- els, TSS onS order substantially exceeds TSS on S nonorder , confirming that temporal intervention is more impactful on genuinely time-dependent items and that the order-sensitivity annotation provides a useful signal. ModelTSS (S order )TSS (S nonorder ) Qwen2.5-VL-7B20.5%4.5% Qwen2.5-VL-32B19.5%6.0% Qwen3-VL-8B28.0%9.0% Qwen3-VL-32B25.0%8.0% InternVL3-8B14.0%6.0% InternVL3.5-8B22.5%5.5% Tarsier-34B21.5%7.5% Table A2: TSS on order-sensitive vs. non-order- sensitive items. Models with non-trivial temporal re- sponsiveness show meaningful separation between sub- sets. A.3 Human Validity Audit for Invariant-Label Operators A key assumption behind Modes I–I is that the perturbation operators are label-preserving, i.e., al- though the input is degraded or corrupted, the cor- rect answer should remain unchanged. To verify that this assumption is aligned with human judg- ment, we conduct a human validity audit over all invariant-label operators. For each operator, we randomly sample 200 per- turbed video–QA pairs and ask three independent annotators, blinded to the original gold labels, to first judge whether the perturbed instance remains answerable and, if so, to select the correct multiple- choice option. We aggregate annotations by major- ity vote and report two measurements: answerabil- ity, the fraction of perturbed instances that remain well-posed, and label preservation, the fraction of answerable instances whose majority-vote label matches the original gold answer. This protocol provides a direct human check of the invariant-label assumption used by R. Table A3 shows that the single-operator pertur- bations used in Modes I–I are largely consistent with human judgments: answerability ranges from 98.5% to 100.0%, and label preservation ranges from 99.5% to 100.0%. These results support the use of R as a reliability metric under visual degra- dation and evidence corruption. Caption injection is additionally analyzed at the subtitle-subtype level, because it mixes two distinct forms of textual interference: content-irrelevant subtitles and intentionally misleading subtitles. As shown in Table A4, irrelevant captions fully pre- serve answerability and labels, whereas misleading captions yield slightly lower answerability (94.0%) and label preservation (94.5%), which is expected given their adversarial design. Even so, the rates remain sufficiently high to justify treating caption injection as an approximately label-preserving op- erator in our evaluation. OperatorAnswerabilityLabel-pres. Compression (cmp)99.5%100.0% Subtitle corruption (sub)99.0%99.5% Gaussian noise (gau)100.0%100.0% Motion blur (mb)100.0%99.5% Adversarial noise (adv)98.5%99.5% Table A3: Human validity audit for invariant-label operators. Each operator is audited on 200 perturbed video–QA pairs with three blinded annotators and majority-vote aggregation. High answerability and label-preservation rates indicate strong agreement be- tween the automatic invariant-label evaluation and hu- man judgments. SubtypeAnswerabilityLabel-pres. Irrelevant100.0%100.0% Misleading94.0%94.5% Table A4: Caption injection validity by subtype. Irrel- evant captions fully preserve answerability and labels. Misleading captions show slightly lower rates, consis- tent with their adversarial design. A.4 Factuality vs. Knowledge Gaps To probe whether factuality errors are predomi- nantly hallucination-like (high-confidence wrong) or knowledge-gap-like (low-confidence/uncertain wrong), we conduct a behavioral diagnostic on 369 factuality questions across 8 open-source models, yielding 2,952 model-question pairs (8×369). Each model is evaluated with standard prompting (clean run) and additionally asked to report a self-assessed confidence score (diagnostic run). Wrong instances are classified as: hallucination-like if the reported confidence is≥70%, knowledge-gap-like if the con- fidence is≤40% or the model expresses explicit uncertainty, and other otherwise. Among 888 analyzable wrong instances, 797 (89.75%) exhibit hallucination-like behavior, while only 39 (4.39%) show knowledge-gap-like patterns. This suggests that factuality failures in current Video-LLMs are overwhelmingly overconfident, producing wrong answers with high self-reported certainty rather than acknowledging uncertainty. We note that this is a behavioral diagnostic rather than a causal attribution—knowledge gaps may also manifest as overconfident guessing. A.5 Implementation Details for Operators Compression (cmp). Videos are re-encoded us- ing FFmpeg with a target bitrate retaining approx- imately 15.19% of the original (reducing storage from 2.83 GB to 0.43 GB on average), simulating the lossy compression commonly applied by video- sharing platforms. Subtitle Corruption (sub).Noisy ASR-like sub- titles are generated by injecting character-level er- rors (substitution, deletion, insertion) into ground- truth transcriptions, with intentional subtitle–video desynchronization (random temporal shifts of 0.5–2.0 seconds). This simulates the unreliable OCR/ASR outputs and timing misalignment en- countered in deployment. Gaussian Noise (gau).Gaussian noise is applied uniformly at the frame level to simulate sensor noise or low-light recording conditions. We add zero-mean Gaussian noise to the RGB channels of each frame. The perturbation variance is con- trolled to keep the noise bounded, ensuring that the high-level semantic content required to answer the question is preserved while degrading low-level perceptual clarity. Motion Blur (mb). Motion blur is applied uni- formly across all frames to simulate camera shake or fast-moving subjects. We apply a linear motion blur filter with randomized kernel sizes and angles to the spatial dimensions of each frame. This oper- ation reduces visual sharpness and introduces per- ceptual ambiguity without altering the underlying dynamic events. Caption Injection (cap). To create cross-modal conflicts, we overlay synthesized subtitles onto ran- domly sampled video segments. As analyzed in Appendix A.3, the injected subtitles consist of a 4:1 mixture of content-irrelevant sentences and LLM- generated misleading statements. The misleading statements are explicitly conditioned on the ques- tion and the ground-truth answer option to form a plausible but incorrect textual cue (e.g., pairing a video of opening a door” with the text closing the door”), testing the model’s ability to prioritize authentic visual evidence over misleading text. Adversarial Noise (adv). We apply a transfer- based black-box perturbation targeting the visual encoder, generated using the MI-FGSM (Dong et al., 2018) algorithm. To ensure the adversarial noise generalizes across different Video-LLMs ar- chitectures rather than overfitting to a single proxy, the perturbations are crafted using a proxy ensem- ble of varied visual encoders (InternVL3-8B (Zhu et al., 2025), Qwen3VL-8B (Bai et al., 2025a), and Video-MAE (Tong et al., 2022)). Temporal Intervention (shu & rev). Applied exclusively to the order-sensitive subset (S order ), these operators disrupt the temporal structure es- sential for correctness. Shuffling (shu) randomly permutes the sequence of frames across the video, disrupting both local motion and global event order. Reversal (rev) strictly inverts the chronological order of the frames from end to start, reversing the direction of actions and state transitions. Both oper- ators render the original ground-truth label invalid by destroying the necessary causal dependencies. Task & DefinitionExample 1. Static Entities & Attributes 1.1 Entity Recognition Identify an entity label supported by video evidence. Which human organ is visible in the video? (A) Stomach(B) Liver(C) Lungs(D) Kidneys 1.2 Unique Entity Counting Count the number of distinct entity instances. How many babies does the lion mother have in the video? (A) 4(B) 3(C) 2(D) 5 1.3 Temporal Attributes Recognition Attribute query with an explicit temporal/event anchor. What is the color of the second object that enters the scene? (A) Brown(B) Cyan(C) Gray(D) Purple 1.4 Static Attributes Recognition Attribute query without explicit temporal/event anchor. What color is the T-shirt worn by the boy? (A) Yellow(B) White(C) Black(D) Blue 1.5 Scene Text Recognition Detecting and reading text that appears in the video. What is written on the first made keychain? (A) Google (B) YouTube (C) Facebook (D) Spotify 2. Dynamic Actions & Motions 2.1 Action Recognition Discriminate fine-grained / near-confusable action. Which description correctly matches the actions? (A) Attacking(B) Chasing(C) Running (D) Stealing 2.2 Repetitive Action Counting Count occurrences of actions across the video. How many times does the hammer’s handle hit the floor throughout the video? (A) 4 (B) 3 (C) 1 (D) 0 (E) 5 (F) 2 2.3 Motion Attributes Recognition Identify motion attributes (direction, rotation,speed, tra- jectory) of an entity. In what direction(s) is the wheel rotating? (A) Counter-clockwise throughout (B) Clockwise throughout (C) No rotation (D) Counter-clockwise then clockwise. 3. Spatio-Temporal Relations 3.1 Spatial Relation Recognition Infer spatial relations or spatial locations in the video. Where is the hidden object at the end of the game? (A) Under 1st left(B) Under 3rd left (C) Under 2nd from left 3.2 Temporal Relation Recognition Determine the temporal relations among events. What happened after the person sat on the sofa/couch? (A) Opened the box (B) Opened the door (C) Sat on the floor (D) Put down the clothes. 3.3 State Transition Detection Determine whether a state variable changes over time. Is the bag empty at the end? (A) Yes(B) No(C) The person doesn’t interact with a bag 3.4 Temporal Localization Localize when an action/state/event holds within the video. When in the video does the action occur? (A) Beginning(B) Middle(C) End(D) Throughout Table A5: Faithfulness Hallucination Taxonomy. Category & DefinitionExample 1. Domain Knowledge (Know-WHAT) 1.1 Cultural Event Recognition Knowledge of traditions, festivals, artistic heritage, and customs. When is the festival in the video usually celebrated? (A) 9th day of 9th lunar month (B) 8th day of 12th lunar month (C)5th day of 5th lunar month (D) 1st day of 1st lunar month 1.2 Historical Background Identification Knowledge of past events, chronologies, archaeol- ogy, and historical figures. When was the Site shown in the video discovered? (A) 1928-1929 (B) 1925-1926 (C) 1920-1922 (D) 1934-1938 1.3 Geospatial Localization Knowledge of geospatial locations, landmarks, and regional characteristics. Which city is this video most likely recorded in? (A) San Diego (B) Phoenix (C) Tucson (D) Las Vegas 1.4 Entertainment-Related Recognition Knowledge of movies, TV, music, celebrities, and sports. What genre of movie is portrayed in the video? (A) Reality TV (B) Documentary (C) Thriller (D) Romance 2. Procedural Knowledge (Know-HOW) 2.1 Electronic Procedural Judgement Logical steps for repairing computers and electronic hardware. Is the sequence of steps about task <Replace Laptop Screen> shown logical? (A) Yes, the sequence is correct (B) No, the sequence is incorrect 2.2 Mechanical Procedural Judgement Sequence and tools for vehicle maintenance and mechanical repair. Is the sequence of steps about task <Replace Car Door Handle> shown logical? (A) Yes, the sequence is correct (B) No, the sequence is incorrect 2.3 Domestic Procedural Judgement Workflow for household DIY, plumbing, and furniture repair. Is the sequence of steps about task <Replace Light Socket> shown logical? (A) Yes, the sequence is correct (B) No, the sequence is incorrect 2.4 Clinical Procedural Judgement Medical protocols for examinations, first aid, and clinical operations. Is the sequence of steps about task <Measure the Jaw Opening> shown logical? (A) Yes, the sequence is correct(B) No, the sequence is incorrect 3. Physical Knowledge (Know-WHY) 3.1 Newtonian Mechanical Reasoning Newtonian laws: gravity, friction, collision, and mo- mentum. Analyze the physical dynamics in the video, which statement best describes the adherence to physical laws? (A) Fully Plausible (B)Violation of Mechanical Dynamics (C) Viola- tion of Material Properties (D) Violation of Fluid Dynamics 3.2 Fluid Mechanics Reasoning Behavior and interaction of liquids, smoke, and fire. Analyze the physical dynamics; which statement best describes the adherence to laws? (A) Fully Plausible (B) Violation of Mechanics (C) Violation of Fluids (D) Violation of Materials 3.3 Material Property Reasoning Realism of deformation, hardness, and fracture of solids. Analyze the physical dynamics; which statement best describes the adherence to laws? (A) Fully Plausible (B) Violation of Mechanics (C) Violation of Fluids (D) Violation of Materials 3.4 Spatio-temporal Continuity Reason- ing Continuity and permanence of objects under occlu- sion or movement. Analyze the physical dynamics; which statement best describes the adherence to laws? (A) Fully Plausible (B) Violation of Object Permanence (C) Violation of Continuity Table A6: Factual Hallucination Taxonomy. The framework evaluates world knowledge across three dimensions: Domain Knowledge (Know-WHAT), Procedural Knowledge (Know-HOW), and Physical Knowledge (Know- WHY).