Paper deep dive
FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection
Pei Li, Sihan Chen, Delong Ran, Tianshuo Cong
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/5/2026, 5:04:31 AM
Summary
This paper introduces FakeI2V-Bench, a comprehensive benchmark for deepfake video detection that evaluates both video-level and image-level detectors. It highlights that naive image-level detectors often underperform in video contexts but can be significantly enhanced via the proposed IV-Bridge framework, which uses frame-level aggregation and random forest models to surpass state-of-the-art video-level detectors.
Entities (13)
Relation Signals (15)
IV-Bridge → enhances → LNP
confidence 95% · LNP-IV (Ours) ... IV-Bridge-enhanced detectors
IV-Bridge → enhances → CNNDet
confidence 95% · CNNDet-IV (Ours) ... IV-Bridge-enhanced detectors
FakeI2V-Bench → evaluateson → GV
confidence 95% · FakeI2V-Bench collects four datasets... GenVideo (GV)
FakeI2V-Bench → evaluateson → CDFV2
confidence 95% · FakeI2V-Bench collects four datasets... Celeb-DF v2 (CDFV2)
FakeI2V-Bench → evaluateson → GVB
confidence 95% · FakeI2V-Bench collects four datasets... GenVidBench (GVB)
FakeI2V-Bench → includes → LNP
confidence 95% · FakeI2V-Bench integrates twelve representative image-level detectors... LNP
FakeI2V-Bench → includes → CNNDet
confidence 95% · FakeI2V-Bench integrates twelve representative image-level detectors... CNNDet
FakeI2V-Bench → includes → FTCN
confidence 95% · FakeI2V-Bench includes eight state-of-the-art video-level deepfake detectors... FTCN
FakeI2V-Bench → includes →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks remain underdeveloped. In particular, the effectiveness of image-level detectors in the video domain has not been systematically assessed. To fill this gap, we present FakeI2V-Bench, a benchmark for evaluating state-of-the-art video-level deepfake detectors in challenging scenarios, with a particular focus on systematically assessing the performance of image-level deepfake detectors in the video domain. FakeI2V-Bench comprises 97,548 videos, containing content generated by the latest powerful generation models and covering a broader range of categories. Using this dataset, we conduct a systematic evaluation of eight video-level detectors and twelve representative image-level detectors. Experimental results show that the best-performing image-level detector achieves an 80.16% AUC, slightly outperforming the strongest video-level detector (i.e., 79.99% AUC). Going beyond benchmarking, we present IV-Bridge, a general framework that enhances the applicability of image-level deepfake detectors to videos. IV-Bridge employs a random forest model with statistical features to aggregate frame-level predictions, allowing eleven image-level detectors to surpass state-of-the-art video-level approaches, with the best-performing variant achieving a 93.80% AUC. Overall, FakeI2V-Bench establishes a rigorous benchmark for deepfake video detection and introduces a novel pathway for extending image-level detectors to the video domain, offering new insights and directions for future research. Code and data are available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.03096v1
- Canonical: https://arxiv.org/abs/2608.03096v1
Trouble viewing inline? Open PDF directly →
Full Text
64,216 characters extracted from source content.
Expand or collapse full text
by FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection Pei Li 0009-0001-9002-7930 School of Cyber Science and Technology, Shandong University,Qingdao, ShandongChina leepy@mail.sdu.edu.cn , Sihan Chen 0009-0008-0176-9782 School of Cyber Science and Technology, Shandong University,Qingdao, ShandongChina 202437060@mail.sdu.edu.cn , Delong Ran 0009-0003-2630-2680 Institute for Network Sciences and Cyberspace, BNRist, Tsinghua University,BeijingChina rdl22@mails.tsinghua.edu.cn and Tianshuo Cong 0000-0003-3189-8223 School of Cryptologic Science and Engineering, Shandong University,Jinan, ShandongChina tianshuo.cong@sdu.edu.cn (2026) Abstract. Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks remain underdeveloped. In particular, the effectiveness of image-level detectors in the video domain has not been systematically assessed. To fill this gap, we present FakeI2V-Bench, a benchmark for evaluating state-of-the-art video-level deepfake detectors in challenging scenarios, with a particular focus on systematically assessing the performance of image-level deepfake detectors in the video domain. FakeI2V-Bench comprises 97,54897,548 videos, containing content generated by the latest powerful generation models and covering a broader range of categories. Using this dataset, we conduct a systematic evaluation of eight video-level detectors and twelve representative image-level detectors. Experimental results show that the best-performing image-level detector achieves an 80.16%80.16\% AUC, slightly outperforming the strongest video-level detector (i.e., 79.99%79.99\% AUC). Going beyond benchmarking, we present IV-Bridge, a general framework that enhances the applicability of image-level deepfake detectors to videos. IV-Bridge employs a random forest model with statistical features to aggregate frame-level predictions, allowing eleven image-level detectors to surpass state-of-the-art video-level approaches, with the best-performing variant achieving a 93.80%93.80\% AUC. Overall, FakeI2V-Bench establishes a rigorous benchmark for deepfake video detection and introduces a novel pathway for extending image-level detectors to the video domain, offering new insights and directions for future research. 111Code and data are available at: https://github.com/CryptoAILab/FakeI2V-Bench. Deepfake detection, AI-generated video, Video deepfakes, AIGC †journalyear: 2026†copyright: c†conference: Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2; August 9–13, 2026; Jeju Island, Republic of Korea.†booktitle: Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD 2026), August 9–13, 2026, Jeju Island, Republic of Korea†isbn: 979-8-4007-2259-2/2026/08†doi: 10.1145/3770855.3817509†ccs: Security and privacy Social aspects of security and privacy†ccs: Computing methodologies Computer vision 1. Introduction Large generation models are at the forefront of Artificial Intelligence (AI) innovation, with Video Generation Models (VGMs) driving significant transformations in digital content production. Leading VGMs such as Sora (Brooks et al., 2024) can generate videos up to 1080p resolution from input text prompts or even an existing image or video. These AI-generated videos, referred to as deepfake videos, not only offer benefits in domains like intelligent news anchors (Bohacek and Farid, 2024) and cinematic visual effects (Gordon, 2024), but also introduce potential societal safety risks, such as misleading judicial evidence verification (of the State of Washington for King County, 2024) or synthetic pornography (Roose, 2018). In response to the deepfake threat, significant research efforts have been devoted to designing video-level deepfake detectors (Zheng et al., 2021; Zhuang et al., 2022; Wang et al., 2023c; Nguyen et al., 2024; Pang et al., 2024; Chen et al., 2024b; Song et al., 2024; Zheng et al., 2025). While these detectors demonstrate promising performance, they still face several fundamental limitations, including limited generalization (Ni et al., 2025; Chen et al., 2024b) and substantial computational overhead (Song et al., 2024). Moreover, as video generation models continue to evolve rapidly, it remains unclear how well existing detectors perform on videos synthesized by the latest generation models, underscoring the urgent need for systematic, comprehensive, and up-to-date evaluation. In parallel, image-level deepfake detectors (Wang et al., 2020; Tan et al., 2023; Sha et al., 2023; Ojha et al., 2023) have achieved remarkable progress, benefiting from large-scale training data and well-established detection pipelines. Meanwhile, these image-level detectors have also been preliminarily employed in deepfake video detection under various usage patterns, such as averaging frame-level predictions to form a video-level decision (Qian et al., 2020) or reporting detection performance independently on each frame (Deng et al., 2024; Yan et al., 2023). This also indicates that while the potential of image-level detectors for identifying deepfake videos has been increasingly recognized, their effective adaptation and systematic evaluation in video-level detection scenarios remain insufficiently explored. 1.1. Our Work Motivated by the above observations, in this paper, we present FakeI2V-Bench, a comprehensive benchmark for deepfake video detection that systematically evaluates both video-level detectors and image-level detectors under a unified evaluation framework. FakeI2V-Bench. FakeI2V-Bench consists of four modules: (i) Evaluation dataset: FakeI2V-Bench collects four datasets (introduced in Table 2) containing a total of 97,54897,548 videos, which fall into two major categories: facial datasets and general datasets, encompassing diverse VGMs (e.g., Face2Face (Thies et al., 2016), Gen2 (Research, 2023), etc) and diverse content (e.g., face, foods, natural landscape, etc). Notably, the inclusion of GenVidBench (GVB) (Ni et al., 2025) dataset, which was newly proposed in 2025, highlights the timeliness of our measurement. (i) Video-level detectors: FakeI2V-Bench includes eight state-of-the-art video-level deepfake detectors (listed in Table 7), including both face-focused and general-purpose detectors. (i) Image-level detectors: FakeI2V-Bench integrates twelve representative image-level detectors (listed in Table 8) alongside their enhanced counterparts, which are developed through our enhancement framework, IV-Bridge. (iv) Evaluation Metrics: FakeI2V-Bench utilizes two metrics, Area Under Curve (AUC) and Average Precision (AP), to provide a thorough comparison of detection capabilities. Research Questions. Based on FakeI2V-Bench, we aim to address the following three key Research Questions (RQs). • RQ1: What is the detection capability of current emerging video-level deepfake detectors in complex scenarios? • RQ2: What is the performance of naive image-level deepfake detectors on video data? • RQ3: How can image-level detectors be enhanced for higher deepfake video detection performance? Evaluation Results on RQ1. To address RQ1, in Section 4, we conduct a comprehensive evaluation of eight video-level detectors. Table 3 presents the detection performance of all video-level detectors on the FakeI2V-Bench dataset. While some detectors achieve excellent results on specific datasets (e.g., ¿90% AUC), their performance can drop to around 50% AUC on other datasets, highlighting limited generalization. Overall, these results indicate that current video-level detectors still have substantial room for improvement to achieve consistently high performance across diverse datasets. Evaluation Results on RQ2. In Section 5, we further evaluate twelve naive image-level detectors to tackle RQ2. We first analyze their frame-level performance and compare it with their performance on a conventional fake image dataset (named FakeGenImage). Table 4 reveals that the performance of naive image-level detectors drops substantially when applied to video frames. However, by applying different frame-to-video aggregation strategies to obtain video-level predictions, the performance can be significantly improved. Notably, one naive image-level detector has even surpassed SOTA video-level detectors (see Table 3), demonstrating the potential of image-level models for deepfake video detection. Solutions to RQ3. To address RQ3, in Section 6, we introduce IV-Bridge, a framework for enhancing and adapting image-level detectors to video scenarios. IV-Bridge consists of two stages: (1) Video-Frame Fine-Tuning (VFT), which fine-tunes image-level detectors to capture video-specific forgery patterns, and (2) Multi-Mode Aggregation (MMA), which aggregates predictions across multiple frame-to-video strategies through a random forest model. As shown in Table 3, IV-Bridge significantly improves the detection performance compared to naive image-level models, and eleven of twelve IV-Bridge-enhanced detectors (detectors with the “-IV” suffix) successfully surpass SOTA video-level detectors. Moreover, IV-Bridge-enhanced detectors exhibit stronger generalization across different video generation models, and the deployment costs are much lighter than video-level detectors. These results not only demonstrate the practical effectiveness of IV-Bridge, but also provide a promising direction for future research in deepfake video detection by leveraging mature image-level detectors. Core Contributions. In summary, we make the following contributions: • We introduce FakeI2V-Bench, a comprehensive benchmark for evaluating video-level and image-level deepfake detectors across a wide range of evaluation datasets. • We present an enhancement framework named IV-Bridge to adapt image-level detectors for deepfake video detection. • Extensive results validate that eleven of twelve IV-Bridge-enhanced detectors successfully surpass SOTA video-level detectors. Table 1. Comparison of deepfake video detection benchmarks, with a special focus on their support for image-level detectors. “Latest Detector” denotes the release time of the latest detector evaluated. In the Image-level column, ● indicates systematic evaluation, ◐ denotes limited evaluation, and ○ denotes no evaluation. Name Latest Detector Image-level Dataset Facial General IDBench (Deng et al., 2024) 2021-03 ◐ ✓ ✗ DeepfakeBench (Yan et al., 2023) 2023-05 ◐ ✓ ✗ GenVidBench (Ni et al., 2025) 2024-05 ○ ✗ ✓ DeMamba (Chen et al., 2024b) 2024-05 ◐ ✗ ✓ FakeI2V-Bench (Ours) 2025-08 ● ✓ ✓ 2. Related Works A comparative summary of representative deepfake video detection benchmarks is provided in Table 1, from which several key limitations can be identified. (i) Limited content diversity. Most existing benchmarks are constructed around a single generation scenario, such as face-centric manipulations or generic text-to-video and image-to-video synthesis, which limits the evaluation of detector generalization across diverse and complex video content. (i) Outdated detector coverage. Many benchmarks evaluate only outdated detectors. For instance, the latest detectors evaluated in IDBench (Deng et al., 2024) were released around five years ago, while most detectors tested in GenVidBench (Ni et al., 2025) date from 2017 to 2022, with only one method released after 2024. (i) Insufficient evaluation of image-level detectors. Some benchmarks (e.g., GenVidBench (Ni et al., 2025) ) exclude image-level detectors entirely, while others (e.g., IDBench (Deng et al., 2024) and DeepfakeBench (Yan et al., 2023)) report only frame-level results without principled video-level aggregation. Although DeMamba (Chen et al., 2024b) evaluates two image-level detectors and claims to consolidate frame-level predictions to obtain video-level predictions, it does not specify the exact aggregation method used. In contrast, our work places particular emphasis on image-level detectors, systematically evaluating their performance and applicability in the video domain under well-defined aggregation protocols, thereby providing new insights into their practical potential for deepfake video detection. Figure 1. An overview of FakeI2V-Bench. FakeI2V-Bench is designed to comprehensively evaluate the performance of deepfake video detectors, with particular emphasis on the scalability of image-level detectors. Table 2. Details of the evaluation datasets used in this paper. We summarize the release years, dataset sizes, video sources, video content types, resolution, and frame rate (FPS). Dataset Year Quantity (Real/Fake) Video Source Video Content Resolution FPS Real Fake CDFV2 (Li et al., 2020) 2019 890/5,639 YouTube Improved basic DeepFake makers (e.g., FakeApp) Facial 256×256256× 256 30 DFD (Google and Jigsaw, 2019) 2019 363/3,068 Shot by actors Unknown algorithms Facial 320×320320× 320–3840×21603840× 2160 15–24 GV (Chen et al., 2024b) 2024 10,000/8,588 MSR-VTT (Xu et al., 2016) Generated by diffusion and autoregressive models General 256×256256× 256 8–24 GVB (Ni et al., 2025) 2025 13,800/55,200 Vript (Yang et al., 2024) Generated by diffusion and autoregressive models General 256×256256× 256 3–30 3. FakeI2V-Bench We present FakeI2V-Bench, a comprehensive benchmarking framework meticulously constructed to assess detector performance under complex deepfake video scenarios. It comprises four modules: Evaluation datasets, video-level detectors, image-level detectors, and evaluation metrics. Next, we elaborate on each module in detail. 3.1. Evaluation Datasets FakeI2V-Bench collects two facial datasets222Following (Nguyen et al., 2024; Zheng et al., 2021), we crop valid face regions from each frame. and two general datasets (including recent emerging AI-generated video datasets with diverse content). The detailed statistics of these datasets are summarized in Table 2. First, the details of the facial datasets are as follows. • Celeb-DF v2 (CDFV2) (Li et al., 2020) contains 5,639 high-quality synthetic videos and 890 real videos from YouTube. The fake videos are generated using an improved DeepFake synthesis algorithm, which produces more realistic content and alleviates issues such as low-resolution faces, color mismatch, and temporal flickering. • DeepFake Detection (DFD) (Google and Jigsaw, 2019) is a dataset released by Google & JigSaw, comprising 3,068 fake videos and 363 real videos featuring professional actors. Notably, the resolutions of the videos in DFD are not uniform, which is more conducive to evaluating the universality of the detectors. Furthermore, we introduce the following two general datasets. • GenVideo (GV) (Chen et al., 2024b) contains videos generated by 20 different video generation models (e.g., Gen2 (Research, 2023) and Sora (Brooks et al., 2024)), offering high visual realism and diversity. We use its official test set containing 10,000 real videos from the MSR-VTT dataset (Xu et al., 2016), as well as 8,588 fake videos generated by diffusion VGMs (e.g., ModelScope (Wang et al., 2023a)) or autoregressive VGMs (e.g., Videopoet (Kondratyuk et al., 2023)), covering most mainstream generative paradigms and their advanced variants. • GenVidBench (GVB) (Ni et al., 2025) is a challenging dataset comprising videos from eight SOTA text-to-video/image-to-video generators. GVB ensures semantic richness and balanced distribution across three dimensions: actions, objects, and locations, making it a more challenging benchmark for detection tasks. For evaluation, we use a directly accessible subset of 55,200 fake and 13,800 real videos from the full set. The real videos are sampled from the Vript dataset (Yang et al., 2024), while the fake videos are also generated by both diffusion-based VGMs (e.g., SVD (Blattmann et al., 2023), MuseV (Xia et al., 2024)) and autoregressive-based VGMs (e.g., CogVideo (Hong et al., 2022)). Table 3. Performance of deepfake detectors. The results of naive image-level detectors are the optimal aggregated results by one of the six modes. The suffix “-IV” indicates that the corresponding image-level detector has been enhanced using IV-Bridge. Type Detector Facial Datasets General Datasets CDFV2 DFD GVB GV AUC ↑ AP ↑ AUC ↑ AP ↑ AUC ↑ AP ↑ AUC ↑ AP ↑ Video-level FTCN (Zheng et al., 2021) 88.91% 98.61% 95.34% 99.30% 71.32% 80.75% 64.39% 57.82% UIA-ViT (Zhuang et al., 2022) 87.17% 98.26% 84.45% 97.61% 76.52% 87.57% 64.28% 58.49% AltFreezing (Wang et al., 2023c) 92.06% 98.96% 77.54% 95.92% 72.47% 81.23% 53.12 % 47.98 % LAA-Net (Nguyen et al., 2024) 85.40% 95.04% 78.86% 96.76% 43.59% 79.84% 59.39 % 62.95% VGMShield (Pang et al., 2024) 54.18% 90.92% 41.82 % 83.05% 82.50 % 94.95% 56.29% 61.34% DeMamba (Chen et al., 2024b) 54.35% 87.54% 51.79% 90.08% 88.69% 97.01% 93.66% 94.15 % M-Det (Song et al., 2024) 50.69% 90.66% 51.49% 89.56% 63.49 % 84.80% 43.62% 43.00% D3 (Zheng et al., 2025) 49.00% 85.06% 52.70 % 90.44% 83.48% 93.87% 94.37% 94.16 % Image-level CNNDet (Wang et al., 2020) 52.62% 86.59 % 51.32% 88.54 % 61.59% 86.55 % 72.55 % 70.20% CNNDet-IV (Ours) 85.62% 97.10 % 90.13 % 98.70% 82.77% 95.00% 93.31% 94.43% LNP (Liu et al., 2022) 55.66% 87.03% 83.94% 97.24% 64.24% 85.39% 13.12% 30.56% LNP-IV (Ours) 70.45% 93.16% 81.48% 97.34% 90.54% 96.72% 94.96% 95.73% Patch (Chai et al., 2020) 68.71% 91.60% 76.73% 96.26% 65.72% 88.24% 51.54% 54.47% Patch-IV (Ours) 87.92% 97.67% 91.40% 99.19% 85.51% 94.73% 92.21% 92.90% LGrad (Tan et al., 2023) 53.75% 86.98% 58.43% 90.68% 58.54% 84.32% 28.45% 36.21% LGrad-IV (Ours) 73.26% 93.40% 83.37% 96.95% 72.21% 89.80% 90.20% 89.95% DeFake (Sha et al., 2023) 52.14% 87.24% 53.67% 90.54% 70.55% 90.06% 84.55% 83.45% DeFake-IV (Ours) 68.83% 92.52% 77.45% 96.27% 87.65% 96.60% 94.56% 94.61% DIRE (Wang et al., 2023b) 50.47% 85.93% 47.62% 88.88% 60.57% 84.89% 66.02% 64.28% DIRE-IV (Ours) 73.55% 93.17% 78.04% 95.82% 81.80% 93.80% 95.48% 94.76% DMID (Corvi et al., 2023) 50.94% 85.35% 57.01% 90.56% 92.02% 97.75% 99.37% 99.39% DMID-IV (Ours) 81.47% 95.94% 78.96% 96.67% 84.08% 94.93% 90.60% 89.79% CoDE (Baraldi et al., 2024) 43.19% 83.13% 43.70% 87.26% 67.92% 89.05% 73.43% 68.19% CoDE-IV (Ours) 79.30% 95.45% 82.40% 97.33% 99.56% 98.85% 96.32% 96.66% DRCT (Chen et al., 2024a) 54.04% 87.14% 56.30% 91.57% 90.01% 97.14% 97.65% 97.57% DRCT-IV (Ours) 90.81% 98.33% 94.49% 99.27% 87.22% 96.10% 96.17% 97.76% UniFD (Ojha et al., 2023) 65.52% 91.76% 77.32% 96.72% 73.09% 91.38% 91.97% 91.59% UniFD-IV (Ours) 69.31% 92.81% 81.91% 97.17% 76.51% 92.61% 93.48% 93.05% NPR (Tan et al., 2024) 55.46% 87.28% 66.85% 94.00% 48.25% 80.54% 60.00% 63.02% NPR-IV (Ours) 64.33% 91.65% 78.83% 96.82% 88.65% 95.68% 88.88% 90.21% RINE (Koutlis and Papadopoulos, 2024) 63.87% 90.27% 78.53% 96.79% 82.03% 94.39% 96.21% 96.04% RINE-IV (Ours) 85.87% 96.76% 96.59% 99.54% 98.99% 99.68% 93.76% 94.54% 3.2. Video-level Detectors FakeI2V-Bench considers eight representative video-level deepfake detectors, which contain four widely used face-focused detectors (LAA-Net (Nguyen et al., 2024), FTCN (Zheng et al., 2021), UIA-ViT (Zhuang et al., 2022), and AltFreezing (Wang et al., 2023c)) and four general-purpose detectors (M-Det (Song et al., 2024), VGMShield (Pang et al., 2024), DeMamba (Chen et al., 2024b), and D3 (Zheng et al., 2025)). These detectors are summarized in Table 7. We employ officially released model weights for all detectors except DeMamba. Since the pretrained parameters for DeMamba are not publicly available, we train DeMamba following its official implementation details. 3.3. Image-level Detectors FakeI2V-Bench includes twelve advanced image-level deepfake detectors, which are listed in Table 8. The selection is guided by three criteria. First, we consider the training data domain, covering detectors trained on GAN-generated data, diffusion-generated data, or both. Second, we include models with diverse backbone architectures such as ResNet, Vision Transformer, and so on. Third, we account for different training paradigms: Some methods explicitly model forgery artifacts (e.g., LNP (Liu et al., 2022), LGrad (Tan et al., 2023)), while others follow a data-driven paradigm without handcrafted priors (e.g., CNNDet (Wang et al., 2020)). Similarly, we directly use the officially released model weights for all methods with the exception of LNP. Owing to the lack of publicly available pretrained weights for LNP, we generate its model weights in accordance with the official implementation. Note that these detectors, which utilize the official implementations directly, are referred to as naive image-level detectors. Additionally, in Section 6, we will introduce an enhancement framework, named IV-Bridge, to uniformly augment these twelve naive detectors. 3.4. Evaluation Metrics Following (Kang et al., 2025; Nguyen et al., 2024; Zheng et al., 2025), FakeI2V-Bench adopts Area Under the Receiver Operating Characteristic Curve (AUC) as the primary evaluation metric, and additionally reports Average Precision (AP). Both metrics are threshold-independent and provide a comprehensive assessment of detector performance. AUC measures performance across different decision thresholds by characterizing the trade-off between the true positive rate and false positive rate, reflecting the overall ranking ability of a detector. AP summarizes the Precision-Recall curve and emphasizes the detection quality of the positive class, focusing more on detection precision and recall. 4. Benchmarking Video-level Detectors To answer RQ1, we evaluate eight state-of-the-art video-level detectors on the evaluation datasets of FakeI2V-Bench. The results are reported in Table 3. Overall Performance. We begin by analyzing the overall performance across the four evaluation datasets. FTCN achieves the best overall performance, with a mean AUC of 79.99% and a mean AP of 84.12%. On facial datasets, FTCN performs particularly well on DFD, reaching an AUC of 95.34% and an AP of 99.30%, and also shows competitive results on CDFV2 (an 88.91% AUC and a 98.61% AP). For general video datasets, DeMamba excels on GVB while D3 leads on GV. Overall, face-specialized detectors dominate on facial content, whereas general-purpose detectors adapt better to non-face data. Nevertheless, even within their preferred domains, some detectors show suboptimal performance. For instance, AltFreezing achieves only 77.54% AUC on DFD, and VGMShield and M-Det drop to 56.29% and 43.62% AUC on GV, indicating substantial room for improvement in video-level deepfake detection. Generalization Performance. Cross-domain evaluation reveals that detectors’ performance heavily depends on the target data domain. Facial detectors like FTCN and UIA-ViT maintain high AUC and AP on face datasets but drop significantly on general video content. Conversely, general-purpose detectors, such as DeMamba and D3, perform well on GV and GVB but exhibit limited performance on facial datasets, with AUC dropping from around 94% to approximately 50%. These results indicate that current video-level detectors lack sufficient cross-domain generalization, highlighting the need for future designs that are more effective and generalizable across diverse video content. Answer to RQ1: Despite achieving excellent performance (e.g., ¿90% AUC) on specific datasets, current mainstream detectors suffer from poor generalization. For instance, face-specialized detectors exhibit a significant performance drop on challenging general datasets. Similarly, detectors that excel on general datasets only achieve around 50% AUC on facial benchmarks. In summary, current mainstream video-level detectors fall short of delivering universally high performance. 5. Benchmarking Naive Image-level Detectors We investigate the performance of naive image-level detectors in the video domain from two perspectives to answer RQ2. First, we assess their performance on individual video frames. Second, we evaluate their scalability by employing different strategies to aggregate frame-level detection results. 5.1. Performance on Video Frames Dataset Construction. To thoroughly analyze the image-level detectors’ performance on video frames and their divergence from the original image domain, we additionally construct a carefully curated dataset named FakeGenImage. FakeGenImage contains 52,160 fake images generated by 20 distinct image synthesis models, covering diverse generation paradigms including GAN-based models (Karras et al., 2017, 2019, 2020; Brock et al., 2018; Zhu et al., 2017; Choi et al., 2018; Park et al., 2019), diffusion-based models (Dhariwal and Nichol, 2021; Rombach et al., 2022; Nichol et al., 2021; Ramesh et al., 2021), and other representative generation approaches (Rössler et al., 2019; Chen et al., 2018; Dai et al., 2019; Chen and Koltun, 2017; Li et al., 2019). It also includes 52,169 real images collected from commonly used image and face datasets, including LSUN (Yu et al., 2015), ImageNet (Russakovsky et al., 2015), CelebA (Liu et al., 2015), CelebA-HQ (Karras et al., 2017), COCO (Lin et al., 2014), FaceForensics++ (Rössler et al., 2019), and LAION (Schuhmann et al., 2021). These images have been widely adopted in prior studies (Ojha et al., 2023; Tan et al., 2024; Koutlis and Papadopoulos, 2024) as standard benchmarks for evaluating the performance of image-level deepfake detectors. Evaluation Results. The performance of the naive image-level detectors on FakeGenImage and video frames from FakeI2V-Bench is shown in Table 4. Based on the AUC results, we observe that the performance of 10 out of 12 detectors declines on video frames, with severe degradation in some cases (e.g., NPR drops from 92.17% to 55.38%). This degradation is largely due to differences between fake images and video frames, such as motion blur, resolution variations, and other visual distortions introduced during video generation and encoding. Furthermore, while Patch and Defake show a slight performance improvement on frames, their AUCs are still below 60%, and both struggle on the original image domain. The best-performing detector on video frames is UniFD, yet it only achieves an AUC of 71.78%. Overall, these results demonstrate the severe inadequacy of naive image-level detectors for video frame detection. Table 4. The detection results of naive image-level detectors on deepfake images and deepfake video frames. Detector FakeGenImage FakeI2V-Bench AUC ↑ AP ↑ AUC ↑ AP ↑ CNNDet (Wang et al., 2020) 83.28% 82.54% 48.37% 64.84% LNP (Liu et al., 2022) 71.75% 69.40% 51.02% 66.78% Patch (Chai et al., 2020) 46.16% 52.52% 59.62% 68.82% LGrad (Tan et al., 2023) 85.45% 85.23% 45.81% 64.47% DeFake (Sha et al., 2023) 46.76% 49.96% 59.66% 73.32% DIRE (Wang et al., 2023b) 69.52% 67.83% 47.71% 64.11% DMID (Corvi et al., 2023) 95.50% 94.86% 67.44% 86.69% CoDE (Baraldi et al., 2024) 67.49% 64.26% 50.00% 66.70% DRCT (Chen et al., 2024a) 77.32% 76.64% 71.56% 88.36% UniFD (Ojha et al., 2023) 93.73% 93.67% 71.78% 79.98% NPR (Tan et al., 2024) 92.17% 90.95% 55.38% 69.55% RINE (Koutlis and Papadopoulos, 2024) 98.80% 98.80% 71.11% 83.01% Figure 2. The video-level performances of naive image-level detectors using six aggregation modes. 5.2. Performance on Videos For these naive image-level detectors, we further explore several lightweight strategies to aggregate frame-level predictions into a video-level decision. Aggregation Strategies. We consider six distinct frame-to-video aggregation strategies: Sampling Mode (SMP), Averaging Mode (AVG), Maximum Mode (MAX), Minimum Mode (MIN), Median Mode (MED), and Variance Mode (VAR). For instance, the SMP mode randomly selects a single frame from the video and uses its score as the video-level prediction. Detailed definitions of all modes are provided in Appendix A. Figures 2 and 7 depict the detection performance of different strategies for various detectors. It reveals that the optimal aggregation strategy differs across detectors. Furthermore, for some detectors, performance varies significantly between different strategies. For example, RINE achieves an AUC of only 68.60% under the VAR mode, whereas its performance improves to 80.16% when the MIN strategy is adopted. Video-level Performance. For ease of comparison, the video-level detection results after aggregation are also reported in Table 3. The reported results are the optimal aggregated results by default. The results demonstrate that aggregating frame-level predictions into video-level scores significantly improves detection performance. For instance, RINE achieves an AUC of 71.11% at the frame level, which increases to 80.16% after video-level aggregation, representing an improvement of approximately 10% and surpassing even the best video-level detector, FTCN (79.99% AUC). These findings further indicate that previous benchmarks (Yan et al., 2023), which directly compare frame-level results of image-level detectors with video-level detectors, are neither reasonable nor fair. Moreover, they highlight that leveraging image-level detectors for deepfake video detection is an underexplored direction that warrants further investigation. Answer to RQ2: The detection performance of naive image-level detectors significantly declines when applied to video frames compared to the image domain. However, through various aggregation methods, improved video-level detection results can be achieved, even surpassing current state-of-the-art video-level detectors. This fully demonstrates the potential of image-level detectors. 6. Improving the Applicability of Naive Image-level Detectors Based on the findings of Section 5, we introduce IV-Bridge, a systematic framework designed to strengthen the performance of image-level detectors in video scenarios. Figure 3. Frame-level detection performance (AUC %) of image-level detectors on FakeI2V-Bench. “w/o VFT” denotes the original detectors, and “w/ VFT” stands for the enhanced detectors via VFT. 6.1. Methodology of IV-Bridge IV-Bridge consists of a two-stage pipeline: Video-Frame Fine-Tuning (VFT) and Multi-Mode Aggregation (MMA). Video-Frame Fine-Tuning (VFT). Since image-level detectors are only trained on forged images, they cannot fully capture the features of forged video frames due to distribution shifts. To bridge the gap between images and video frames, IV-Bridge fine-tunes the naive image-level detector ℳimgM_img with a fine-tuning dataset ftD_ft under a full-parameter setting to generate ℳimgVFTM_img^VFT. Given an image I, the outputs of ℳimgM_img are primarily categorized into the following two categories. Accordingly, we design distinct fine-tuning loss functions for each category. • Two-logit output. The first category (e.g., DeFake, DRCT) outputs two scores indicating whether an image is real or fake, i.e., ℳimg:I↦=(preal,pfake)M_img:I =(p^real,p^fake). We use the standard cross-entropy loss for fine-tuning: ℒCE=−1N∑i=1Nlogpiyi,L_CE=- 1N _i=1^N p_i^y_i, where yi∈real,fakey_i∈\real,fake\ is the ground-truth label of the i-th frame and piyip_i^y_i denotes the predicted probability for the ground-truth label yiy_i, and N is the dataset size of ftD_ft. • Single-logit output. The second category (e.g., CNNDet, NPR) outputs only the confidence score that the image is fake, i.e., ℳimg:I↦pfakeM_img:I p^fake. We use the following binary cross-entropy (BCE) function to fine-tune detectors: ℒBCE=−1N∑i=1N[yilogpifake+(1−yi)log(1−pifake)],L_BCE=- 1N _i=1^N [y_i p_i^fake+(1-y_i) (1-p_i^fake) ], where yi∈0(real),1(fake)y_i∈\0(real),1(fake)\ is the ground-truth label. Multi-Mode Aggregation (MMA). IV-Bridge further introduces a multi-mode aggregation strategy that adaptively fuses information from multiple frame-to-video modes in a learning-based manner. Specifically, given ℳimgVFTM_img^VFT and a video with T frames fii=1T\f_i\^T_i=1, the detector outputs a forgery probability for each frame fif_i: pifake=ℳimgVFT(fi)∈[0,1],i=1,…,T.p^fake_i=M^VFT_img(f_i)∈[0,1],~i=1,…,T. Next, these frame-level scores are aggregated under six different frame-to-video aggregation mode functions ϕm _m as detailed in Appendix A, where m∈SMP,AVG,MAX,MIN,MED,VARm∈\SMP,AVG,MAX,MIN,MED,VAR\. Each function maps the frame-level scores pifakei=1T\p_i^fake\_i=1^T to a video-level score: Pm=ϕm(p1fake,⋯,pTfake).P_m= _m(p^fake_1,·s,p^fake_T). Stacking the outputs of all modes yields the multi-mode video feature vector: V=[PSMP,PAVG,PMAX,PMIN,PMED,PVAR]⊤.P_V=[P_SMP,P_AVG,P_MAX,P_MIN,P_MED,P_VAR] . Inspired by the use of lightweight classifiers in recent video detection studies (Internò et al., 2026), we employ a lightweight random forest classifier (Breiman, 2001) RF(⋅)C_ RF(·) trained on a dataset vidD_vid to map the resulting multi-mode features to a final video-level forgery probability: zV=RF(V),zV∈[0,1].z_V=C_ RF(P_V),~z_V∈[0,1]. Through this learning-based aggregation, the model automatically captures the relative importance and interactions among different modes, allowing more flexible and effective video-level predictions compared to fixed single-mode aggregation. 6.2. Experiments Experimental Setups. In IV-Bridge, we use a total of 197,533 real and 207,476 fake video frames, including 92,157 real and 92,146 fake frames from the F++ (Rössler et al., 2019) (c23 version), a widely used face forgery dataset, and 105,376 real and 115,300 fake frames from the training split of GenVideo (Chen et al., 2024b). These data do not have any overlapping images with our FakeI2V-Bench evaluation dataset. For all image-level detectors, we generally follow their original hyperparameters during training. For the MMA stage, the Random Forest classifier is configured with up to 800 trees and a maximum depth of 8. (a) Video-level Detectors. (b) IV-Bridge Enhanced Detectors. Figure 4. Results of cross-model detection performance (AUC %) on GV dataset (Part-I). Overall Performance. Using IV-Bridge, we obtain twelve IV-Bridge-enhanced detectors. Table 3 details the performance of each individual detector across the evaluation datasets of FakeI2V-Bench, where the suffix “-IV” indicates that the corresponding detector has been enhanced using IV-Bridge. Compared with directly applying the original naive image-level detectors, the IV-Bridge-enhanced versions achieve substantial performance gains on all datasets. For example, after IV-Bridge adaptation, DRCT-IV improves its AUC by around 18% over the original DRCT. IV-Bridge particularly boosts fully data-driven detectors like DRCT and RINE, which outperform artifact-specific ones, achieving an average AUC of 87.35% versus 81.63%. More importantly, the enhanced image-level detectors significantly outperform the best video-level detector. In terms of the performance across the full FakeI2V-Bench dataset, the SOTA video-level detector FTCN attains the highest AUC of 79.99% and AP of 84.12%. However, among the twelve detectors enhanced by IV-Bridge, eleven detectors achieve performance exceeding that of FTCN, with RINE-IV reaching the highest performance of 93.80% AUC and 97.63% AP. Moreover, IV-Bridge boosts performance on both facial datasets and general datasets. For instance, RINE-IV reaches 96.59% AUC on DFD and CoDE-IV achieves 96.32% AUC on GV, both surpassing the original SOTA video-level detector. Notably, the enhanced CoDE-IV attains an AUC of 99.56% on the GVB dataset, approaching near-perfect detection performance. Performance across Different VGMs. Given the widespread accessibility of diverse video generation models (VGMs), we further analyze how detection performance varies across different generators. The GV and GVB datasets contain videos generated by ten and four distinct models, respectively. For the GV dataset, Figure 4 presents the performance of four representative detectors, while the results of the remaining detectors are shown in Figure 5. Overall, the IV-Bridge-enhanced image-level detectors demonstrate strong generalization across videos produced by diverse VGMs. Nevertheless, noticeable performance variations still exist among different generation models. Specifically, videos generated by Sora are more challenging to detect, yielding a mean AUC of 72.76%, whereas those generated by Crafter are the easiest, achieving an AUC of 98.32%. For the GVB dataset, as shown in Figure 6, videos generated by SVD and MuseV are also relatively more difficult. This can be attributed to the fact that SVD and MuseV generate videos by taking a real image as the initial frame, resulting in more realistic visual content, while videos officially released by Sora exhibit notably high visual quality. Compared with dedicated video-level detectors, IV-Bridge-enhanced image-level detectors achieve a higher mean AUC of 95.51% across all 14 generation models, surpassing the SOTA video-level detector, which attains an AUC of only 91.01%. Effectiveness of VFT. To evaluate the effectiveness of Video-Frame Fine-Tuning (VFT), we compare the frame-level detection performance of image-level detectors on the four FakeI2V-Bench subsets (CDFV2, DFD, GV, and GVB) before and after applying VFT. The results are illustrated in Figure 3. As shown in the figure, all twelve image-level detectors exhibit substantial improvements in frame-level AUC on FakeI2V-Bench videos after VFT. This indicates that detectors trained solely on manipulated still images struggle to capture the visual characteristics of forged video frames. Fine-tuning on video frames effectively mitigates the distribution gap between images and video frames, enabling detectors to better adapt to the video domain. These results further confirm that forged images and forged video frames exhibit different forgery patterns and visual statistics, and highlight the necessity of video-specific adaptation for improving the applicability of image-level detectors in deepfake video detection tasks. Deployment Cost. Table 5 summarizes the model parameters (M), inference time (ms), and detection performance (%) of detectors. All experiments are conducted on a single NVIDIA A800 GPU with 80GB VRAM. First of all, we can observe that three of the twelve IV-Bridge-enhanced detectors (NPR, Patch, and CoDE) outperform the SOTA video-level detector FTCN across all three key metrics: detection performance, model size, and inference time. Notably, Patch-IV, with a parameter count of merely 4.34M, is more lightweight than all video-level detectors listed. However, it achieves superior performance with an AUC of 89.26% and an AP of 96.12%, which is notably higher than all listed video-level detectors. In terms of inference speed, CNNDet-IV ranks first among image-level models at 4.24ms, delivering an AUC of 87.96. In contrast, the fastest video-level detector, UIA-ViT, requires 7.59ms but attains a significantly lower AUC of 78.11%. Table 5. Deployment costs of different detectors. Type Detector #Param (M) Time (ms) Performance AUC(%)↑ AP(%)↑ Video-level FTCN (Zheng et al., 2021) 14.77 11.97 79.99 84.12 UIA-ViT (Zhuang et al., 2022) 85.80 7.59 78.11 85.48 AltFreezing (Wang et al., 2023c) 27.23 33.44 73.80 81.02 LAA-Net (Nguyen et al., 2024) 9.74 22.43 66.81 83.65 VGMShield (Pang et al., 2024) 64.98 14.92 58.70 82.57 DeMamba (Chen et al., 2024b) 127.27 344.63 72.12 92.20 M-Det (Song et al., 2024) 203.62 76.9 52.32 77.01 D3 (Zheng et al., 2025) 121.25 24.68 69.89 90.88 Image-level-enhanced CNNDet-IV 23.51 4.24 87.96 96.31 LNP-IV 26.35 128.07 84.36 95.74 Patch-IV 4.34 9.69 89.26 96.12 LGrad-IV 46.56 138.62 79.76 92.53 DeFake-IV 375.61 1772.7 82.12 95.00 DIRE-IV 552.67 16481.48 82.22 94.39 DMID-IV 23.51 84.81 83.78 94.33 CoDE-IV 5.52 4.32 89.40 97.07 DRCT-IV 88.62 26.08 92.17 97.87 UniFD-IV 427.62 116.18 80.30 93.91 NPR-IV 1.44 6.87 80.17 93.59 RINE-IV 433.94 18.37 93.80 97.63 Performance on Short-term Forgery. Sparse manipulations, where only a small portion of video frames are forged, pose a more challenging detection setting. To examine this case, we construct a short-term forgery test set based on the facial dataset CDFV2 by keeping only 20% forged frames in each fake video. We evaluate this setting on face-specialized video-level detectors and all IV-Bridge-enhanced image-level detectors. As shown in Table 6, all detectors suffer clear performance drops compared with the original CDFV2 setting, confirming that sparse temporal forgery remains challenging for current detectors. Nevertheless, several IV-Bridge-enhanced image-level detectors remain competitive. Patch-IV achieves the best AUC of 77.15% with a relatively moderate AUC drop of 10.77 percentage points, while DRCT-IV obtains 74.09% AUC. These results suggest that IV-Bridge can still maintain superior detection performance under sparse forged-frame scenarios. Table 6. Performance under the short-term forgery setting on CDFV2. Only 20% frames in each fake video are forged. Δ and Δ denote the absolute performance drops compared with the original CDFV2 setting. Type Detector AUC↑ AP↑ Δ ↓ Δ ↓ Video-level FTCN (Zheng et al., 2021) 51.17% 85.59% 37.74% 13.02% UIA-ViT (Zhuang et al., 2022) 58.69% 88.69% 28.48% 9.57% AltFreezing (Wang et al., 2023c) 56.27% 88.17% 35.79% 10.79% LAA-Net (Nguyen et al., 2024) 67.15% 90.50% 18.25% 4.54% Image-level-enhanced CNNDet-IV 64.93% 91.51% 20.69% 5.59% LNP-IV 56.64% 88.68% 13.81% 4.48% Patch-IV 77.15% 94.82% 10.77% 2.85% LGrad-IV 60.33% 89.21% 12.93% 4.19% DeFake-IV 50.73% 86.01% 18.10% 6.51% DIRE-IV 57.87% 89.15% 15.68% 4.02% DMID-IV 55.55% 88.52% 25.92% 7.42% CoDE-IV 43.77% 84.55% 35.53% 10.90% DRCT-IV 74.09% 94.58% 16.72% 3.75% UniFD-IV 61.17% 91.06% 8.14% 1.75% NPR-IV 41.03% 81.40% 23.30% 10.25% RINE-IV 46.08% 84.81% 39.79% 11.95% Answer to RQ3: The proposed IV-Bridge boosts the performance of naive image-level detectors in deepfake video detection, with 11 of the 12 enhanced detectors exceeding the state-of-the-art video-level detector (e.g., FTCN). This demonstrates the stronger generalization, computational lightweightness, and practical effectiveness of IV-Bridge in adapting image-level detectors to video-level tasks. 7. Conclusion In this work, we introduce FakeI2V-Bench, a benchmark dedicated to systematically evaluating the applicability of image-level deepfake detectors for video-level deepfake detection. FakeI2V-Bench contains 97,54897,548 videos and supports a comprehensive evaluation of eight video-level detectors and twelve representative image-level detectors. Furthermore, we present IV-Bridge, a two-stage framework consisting of Video-Frame Fine-Tuning (VFT) and Multi-Mode Aggregation (MMA), designed to enhance image-level detectors for video tasks. Our evaluation reveals that: (1) current video-level detectors still leave substantial room for performance improvement; (2) naive image-level detectors suffer significant performance drops when applied to video frames, while appropriate frame aggregation can partially mitigate this gap; and (3) image-level detectors can be significantly improved with IV-Bridge, surpassing state-of-the-art video-level methods with lower computational cost and stronger generalization across different video generation models. Overall, FakeI2V-Bench establishes a rigorous benchmark for deepfake video detection and, together with IV-Bridge, provides a practical and efficient solution for adapting image-level detectors to video scenarios, offering valuable guidance for future research in deepfake video detection. Acknowledgement We sincerely thank the reviewers for their valuable feedback. This work is supported by the National Natural Science Foundation of China (No.62402273) and the Fundamental and Interdisciplinary Disciplines Breakthrough Plan of the Ministry of Education of China under Grant JYB2025XDXM114. Tianshuo Cong is also with the Shandong Key Laboratory of Artificial Intelligence Security. References L. Baraldi, F. Cocchi, M. Cornia, L. Baraldi, A. Nicolosi, and R. Cucchiara (2024) Contrasting deepfakes diffusion via contrastive learning and global-local similarities. In Springer European Conference on Computer Vision (ECCV), Cited by: Table 8, Table 3, Table 4. A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts, V. Jampani, and R. Rombach (2023) Stable video diffusion: scaling latent video diffusion models to large datasets. Note: arXiv Preprint 2311.15127 External Links: 2311.15127, Link Cited by: 2nd item. M. Bohacek and H. Farid (2024) The making of an ai news anchor—and its implications. Proceedings of the National Academy of Sciences 121 (1), p. e2315678121. Cited by: §1. L. Breiman (2001) Random forests. Mach. Learn. 45 (1), p. 5–32. External Links: Link, Document Cited by: §6.1. A. Brock, J. Donahue, and K. Simonyan (2018) Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096. Cited by: §5.1. T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. Ng, R. Wang, and A. Ramesh (2024) Video generation models as world simulators. Note: https://openai.com/index/sora/Accessed: 2025-04-08 Cited by: §1, 1st item. L. Chai, D. Bau, S. Lim, and P. Isola (2020) What makes fake images detectable? understanding properties that generalize. In Springer European Conference on Computer Vision (ECCV), Cited by: Table 8, Table 3, Table 4. B. Chen, J. Zeng, J. Yang, and R. Yang (2024a) DRCT: diffusion reconstruction contrastive training towards universal detection of diffusion generated images. In Forty-first International Conference on Machine Learning, External Links: Link Cited by: Table 8, Table 3, Table 4. C. Chen, Q. Chen, J. Xu, and V. Koltun (2018) Learning to see in the dark. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 3291–3300. Cited by: §5.1. H. Chen, Y. Hong, Z. Huang, Z. Xu, Z. Gu, Y. Li, J. Lan, H. Zhu, J. Zhang, W. Wang, and H. Li (2024b) DeMamba: ai-generated video detection on million-scale genvideo benchmark. Note: arXiv Preprint 2405.19707 External Links: 2405.19707, Link Cited by: Table 7, Table 1, §1, Table 2, §2, 1st item, §3.2, Table 3, §6.2, Table 5. Q. Chen and V. Koltun (2017) Photographic image synthesis with cascaded refinement networks. In Proceedings of the IEEE international conference on computer vision, p. 1511–1520. Cited by: §5.1. Y. Choi, M. Choi, M. Kim, J. Ha, S. Kim, and J. Choo (2018) Stargan: unified generative adversarial networks for multi-domain image-to-image translation. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 8789–8797. Cited by: §5.1. R. Corvi, D. Cozzolino, G. Zingarini, G. Poggi, K. Nagano, and L. Verdoliva (2023) On the detection of synthetic images generated by diffusion models. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Cited by: Table 8, Table 3, Table 4. T. Dai, J. Cai, Y. Zhang, S. Xia, and L. Zhang (2019) Second-order attention network for single image super-resolution. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 11065–11074. Cited by: §5.1. J. Deng, C. Lin, P. Hu, C. Shen, Q. Wang, Q. Li, and Q. Li (2024) Towards benchmarking and evaluating deepfake detection. IEEE Transactions on Dependable and Secure Computing 21 (6), p. 5112–5127. Cited by: Table 1, §1, §2. P. Dhariwal and A. Nichol (2021) Diffusion models beat gans on image synthesis. Advances in neural information processing systems 34, p. 8780–8794. Cited by: §5.1. Google and Jigsaw (2019) Contributing data to deepfake detection research. Note: https://research.google/blog/contributing-data-to-deepfake-detection-research/Accessed: 2025-04-10 Cited by: Table 2, 2nd item. D. Gordon (2024) What if a.i. is actually good for hollywood?. Note: https://w.nytimes.com/2024/11/01/magazine/ai-hollywood-movies-cgi.htmlAccessed: 2024-12-16 Cited by: §1. W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang (2022) CogVideo: large-scale pretraining for text-to-video generation via transformers. Note: arXiv Preprint 2205.15868 External Links: 2205.15868, Link Cited by: 2nd item. C. Internò, R. Geirhos, M. Olhofer, S. Liu, B. Hammer, and D. Klindt (2026) AI-generated video detection via perceptual straightening. External Links: 2507.00583, Link Cited by: §6.1. C. Kang, S. Jeong, J. Lee, D. Choi, S. S. Woo, and J. Han (2025) HiDF: a human-indistinguishable deepfake dataset. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, p. 5527–5538. Cited by: §3.4. T. Karras, T. Aila, S. Laine, and J. Lehtinen (2017) Progressive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196. Cited by: §5.1. T. Karras, S. Laine, and T. Aila (2019) A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 4401–4410. Cited by: §5.1. T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila (2020) Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 8110–8119. Cited by: §5.1. D. Kondratyuk, L. Yu, X. Gu, J. Lezama, J. Huang, G. Schindler, R. Hornung, V. Birodkar, J. Yan, M. Chiu, et al. (2023) Videopoet: a large language model for zero-shot video generation. Note: arXiv preprint 2312.14125 Cited by: 1st item. C. Koutlis and S. Papadopoulos (2024) Leveraging representations from intermediate encoder-blocks for synthetic image detection. In Springer European Conference on Computer Vision (ECCV), Cited by: Table 8, Table 3, §5.1, Table 4. K. Li, T. Zhang, and J. Malik (2019) Diverse image synthesis from semantic layouts via conditional imle. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 4220–4229. Cited by: §5.1. Y. Li, X. Yang, P. Sun, H. Qi, and S. Lyu (2020) Celeb-DF: a large-scale challenging dataset for deepfake forensics. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 2, 1st item. T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick (2014) Microsoft coco: common objects in context. In European conference on computer vision, p. 740–755. Cited by: §5.1. B. Liu, F. Yang, X. Bi, B. Xiao, W. Li, and X. Gao (2022) Detecting generated images by real images. In Springer European Conference on Computer Vision (ECCV), Cited by: Table 8, §3.3, Table 3, Table 4. Z. Liu, P. Luo, X. Wang, and X. Tang (2015) Deep learning face attributes in the wild. In Proceedings of the IEEE international conference on computer vision, p. 3730–3738. Cited by: §5.1. D. Nguyen, N. Mejri, I. P. Singh, P. Kuleshova, M. Astrid, A. Kacem, E. Ghorbel, and D. Aouada (2024) LAA-Net: localized artifact attention network for quality-agnostic and generalizable deepfake detection. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 7, §1, §3.2, §3.4, Table 3, Table 5, Table 6, footnote 2. Z. Ni, Q. Yan, M. Huang, T. Yuan, Y. Tang, H. Hu, X. Chen, and Y. Wang (2025) GenVidBench: a challenging benchmark for detecting ai-generated video. Note: arXiv Preprint 2501.11340 External Links: 2501.11340, Link Cited by: §1.1, Table 1, §1, Table 2, §2, 2nd item. A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen (2021) Glide: towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741. Cited by: §5.1. S. C. of the State of Washington for King County (2024) FINDINGS of fact and conclusions of law re: frye hearing on admissibility of videos enhanced by artificial intelligence. Note: https://fingfx.thomsonreuters.com/gfx/legaldocs/zgvokxekavd/04192024ai_wash.pdfAccessed: 2025-08-27 Cited by: §1. U. Ojha, Y. Li, and Y. J. Lee (2023) Towards universal fake image detectors that generalize across generative models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 8, §1, Table 3, §5.1, Table 4. Y. Pang, Y. Zhang, and T. Wang (2024) VGMShield: mitigating misuse of video generative models. Note: arXiv Preprint 2402.13126 External Links: 2402.13126, Link Cited by: Table 7, §1, §3.2, Table 3, Table 5. T. Park, M. Liu, T. Wang, and J. Zhu (2019) Semantic image synthesis with spatially-adaptive normalization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 2337–2346. Cited by: §5.1. Y. Qian, G. Yin, L. Sheng, Z. Chen, and J. Shao (2020) Thinking in frequency: face forgery detection by mining frequency-aware clues. In European conference on computer vision, p. 86–103. Cited by: §1. A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever (2021) Zero-shot text-to-image generation. In International conference on machine learning, p. 8821–8831. Cited by: §5.1. R. Research (2023) Gen-2: generate novel videos with text, images or video clips. Note: https://runwayml.com/research/gen-2Accessed: 2025-04-08 Cited by: §1.1, 1st item. R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer (2022) High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 10684–10695. Cited by: §5.1. K. Roose (2018) Here come the fake videos, too. Note: https://w.nytimes.com/2018/03/04/technology/fake-videos-deepfakes.htmlAccessed: 2025-08-27 Cited by: §1. A. Rössler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Niessner (2019) FaceForensics++: learning to detect manipulated facial images. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , p. 1–11. External Links: Document Cited by: §5.1, §6.2. O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al. (2015) Imagenet large scale visual recognition challenge. International journal of computer vision 115 (3), p. 211–252. Cited by: §5.1. C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki (2021) Laion-400m: open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114. Cited by: §5.1. Z. Sha, Z. Li, N. Yu, and Y. Zhang (2023) DE-FAKE: detection and attribution of fake images generated by text-to-image generation model. In ACM SIGSAC Conference on Computer and Communications Security (CCS), Cited by: Table 8, §1, Table 3, Table 4. X. Song, X. Guo, J. Zhang, Q. Li, L. Bai, X. Liu, G. Zhai, and X. Liu (2024) On learning multi-modal forgery representation for diffusion generated video detection. In Conference on Neural Information Processing Systems (NeurIPS), Cited by: Table 7, §1, §3.2, Table 3, Table 5. C. Tan, Y. Zhao, S. Wei, G. Gu, P. Liu, and Y. Wei (2024) Rethinking the up-sampling operations in cnn-based generative network for generalizable deepfake detection. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 8, Table 3, §5.1, Table 4. C. Tan, Y. Zhao, S. Wei, G. Gu, and Y. Wei (2023) Learning on gradients: generalized artifacts representation for gan-generated images detection. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 8, §1, §3.3, Table 3, Table 4. J. Thies, M. Zollhöfer, M. Stamminger, C. Theobalt, and M. Nießner (2016) Face2Face: real-time face capture and reenactment of rgb videos. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §1.1. J. Wang, H. Yuan, D. Chen, Y. Zhang, X. Wang, and S. Zhang (2023a) Modelscope text-to-video technical report. Note: arXiv Preprint 2308.06571 External Links: 2308.06571, Link Cited by: 1st item. S. Wang, O. Wang, R. Zhang, A. Owens, and A. A. Efros (2020) CNN-generated images are surprisingly easy to spot… for now. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 8, §1, §3.3, Table 3, Table 4. Z. Wang, J. Bao, W. Zhou, W. Wang, H. Hu, H. Chen, and H. Li (2023b) DIRE for diffusion-generated image detection. In IEEE/CVF International Conference on Computer Vision (ICCV), Cited by: Table 8, Table 3, Table 4. Z. Wang, J. Bao, W. Zhou, W. Wang, and H. Li (2023c) AltFreezing for more general video face forgery detection. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 7, §1, §3.2, Table 3, Table 5, Table 6. Z. Xia, Z. Chen, B. Wu, C. Li, K. Hung, C. Zhan, Y. He, and W. Zhou (2024) MuseV: infinite-length and high fidelity virtual human video generation with visual conditioned parallel denoising. Note: https://tmelyralab.github.io/MuseV_Page/Accessed: 2025-04-13 Cited by: 2nd item. J. Xu, T. Mei, T. Yao, and Y. Rui (2016) Msr-vtt: a large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 5288–5296. Cited by: Table 2, 1st item. Z. Yan, Y. Zhang, X. Yuan, S. Lyu, and B. Wu (2023) DeepfakeBench: a comprehensive benchmark of deepfake detection. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA. Cited by: Table 1, §1, §2, §5.2. D. Yang, S. Huang, C. Lu, X. Han, H. Zhang, Y. Gao, Y. Hu, and H. Zhao (2024) Vript: a video is worth thousands of words. Advances in Neural Information Processing Systems 37, p. 57240–57261. Cited by: Table 2, 2nd item. F. Yu, A. Seff, Y. Zhang, S. Song, T. Funkhouser, and J. Xiao (2015) Lsun: construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365. Cited by: §5.1. C. Zheng, R. Suo, C. Lin, Z. Zhao, L. Yang, S. Liu, M. Yang, C. Wang, and C. Shen (2025) D3: training-free ai-generated video detection using second-order features. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 12852–12862. Cited by: Table 7, §1, §3.2, §3.4, Table 3, Table 5. Y. Zheng, J. Bao, D. Chen, M. Zeng, and F. Wen (2021) Exploring temporal coherence for more general video face forgery detection. In IEEE/CVF International Conference on Computer Vision (ICCV), Cited by: Table 7, §1, §3.2, Table 3, Table 5, Table 6, footnote 2. J. Zhu, T. Park, P. Isola, and A. A. Efros (2017) Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE international conference on computer vision, p. 2223–2232. Cited by: §5.1. W. Zhuang, Q. Chu, Z. Tan, Q. Liu, H. Yuan, C. Miao, Z. Luo, and N. Yu (2022) UIA-ViT: unsupervised inconsistency-aware method based on vision transformer for face forgery detection. In Springer European Conference on Computer Vision (ECCV), Cited by: Table 7, §1, §3.2, Table 3, Table 5, Table 6. Appendix A Aggregation Strategies Given an image-level detector ℳimgM_img and a video V=f1,…,fNV=\f_1,…,f_N\, the detector outputs a fake confidence score pifakep_i^fake for each frame fif_i. We evaluate six frame-to-video integration mode functions ϕm _m, where m∈SMP, AVG, MAX, MIN, MED, and VARm∈\SMP, AVG, MAX, MIN, MED, and VAR\ to obtain video-level prediction PmP_m. The detailed definitions of these six modes are provided as follows. • Sampling Mode (SMP). The SMP mode randomly selects a single frame FiF_i from video V and uses its score as the video-level prediction: PSMP=ϕSMP(ℳimg,V)=pifake,P_SMP= _SMP(M_img,V)=p_i^fake, • Averaging Mode (AVG). The AVG mode computes the mean score across all frames: PAVG=ϕAVG(ℳimg,V)=1N∑i=1Npifake.P_AVG= _AVG(M_img,V)= 1NΣ^N_i=1p_i^fake. • Maximum Mode (MAX). The MAX mode uses the largest frame-level score as the final video-level prediction: PMAX=ϕMAX(ℳimg,V)=maxi=1,…,Npifake.P_MAX= _MAX(M_img,V)= _i=1,…,Np_i^fake. • Minimum Mode (MIN). The MIN mode adopts the smallest frame score: PMIN=ϕMIN(ℳimg,V)=mini=1,…,Npifake.P_MIN= _MIN(M_img,V)= _i=1,…,Np_i^fake. • Median Mode (MED). The MED mode first sorts the frame-level scores in ascending order: p~(1)fake≤⋯≤p~(N)fake, p^fake_(1)≤·s≤ p^fake_(N), where p~(k)fake p^fake_(k) represents the k-th order score. The median value is then computed as the final video-level score. PMED=ϕMED(ℳimg,V)=p~(N+12)fake,if N is odd,p~(N2)fake+p~(N2+1)fake2,if N is even.P_MED= _MED(M_img,V)= cases p^fake_ ( N+12 ),&if $N$ is odd,\\[5.69054pt] p^fake_ ( N2 )+ p^fake_ ( N2+1 )2,&if $N$ is even. cases • Variance Mode (VAR). The VAR mode computes the variance of frame-level scores: PVAR=ϕVAR(ℳimg,V)=1N∑i=1N(pifake−p¯fake)2,p¯fake=1N∑i=1Npifake.P_VAR= _VAR(M_img,V)= 1NΣ^N_i=1(p_i^fake- p^fake)^2,~ p^fake= 1NΣ^N_i=1p_i^fake. (a) Video-level Detectors. (b) IV-Bridge Enhanced Detectors. Figure 5. Results of cross-model detection AUC (%) on GV Dataset (Part-I). (a) Video-level Detectors. (b) IV-Bridge Enhanced Detectors. Figure 6. Results of cross-model detection AUC (%) on GVB Dataset. Figure 7. The video-level performances of eight naive image-level detectors using six aggregation modes. Table 7. Overview of the video-level deepfake detectors included in FakeI2V-Bench. Detector Venue Task Architecture Training Set vs. Image-level Detectors? Y/N Method FTCN (Zheng et al., 2021) ICCV’21 Facial CNN + Transformer F++ ✓ Average-level UIA-ViT (Zhuang et al., 2022) ECCV’22 Facial ViT F++ ✓ Frame-level AltFreezing (Wang et al., 2023c) CVPR’23 Facial 3DCNN F++ ✓ Average-level LAA-Net (Nguyen et al., 2024) CVPR’24 Facial EfficientNet-B4 + FPN F++ ✓ Average-level VGMShield (Pang et al., 2024) ArXiv’24 General VideoMAE TD ✗ – DeMamba (Chen et al., 2024b) ArXiv’24 General XCLIP+Mamba SEINE ✓ Average-level M-Det (Song et al., 2024) NeurIPS’24 General VLLM+ViT SVD ✓ Average-level D3 (Zheng et al., 2025) ICCV’25 General XCLIP - ✓ Unknown Table 8. Overview of the image-level deepfake detectors included in FakeI2V-Bench. Detector Venue Training Set Domain Architecture Learning-based? GAN Diffusion Y/N Paradigm CNNDet (Wang et al., 2020) CVPR’20 ✓ ✗ ResNet ✓ Data-driven Patch (Chai et al., 2020) ECCV’20 ✓ ✗ Xception ✓ Data-driven LNP (Liu et al., 2022) ECCV’22 ✓ ✗ ResNet ✓ Specific Artifacts LGrad (Tan et al., 2023) CVPR’23 ✓ ✗ ResNet ✓ Specific Artifacts DeFake (Sha et al., 2023) CCS’23 ✗ ✓ CLIP ✓ Data-driven DIRE (Wang et al., 2023b) ICCV’23 ✗ ✓ ResNet ✓ Specific Artifacts DMID (Corvi et al., 2023) ICASSP’23 ✗ ✓ ResNet ✓ Data-driven CoDE (Baraldi et al., 2024) ECCV’24 ✗ ✓ ViT ✓ Data-driven DRCT (Chen et al., 2024a) ICML’24 ✗ ✓ Conv-B/CLIP ✓ Data-driven UniFD (Ojha et al., 2023) CVPR’23 ✓ ✗ CLIP ✓ Data-driven NPR (Tan et al., 2024) CVPR’24 ✓ ✗ ResNet ✓ Specific Artifacts RINE (Koutlis and Papadopoulos, 2024) ECCV’24 ✓ ✗ ViT ✓ Data-driven