Paper deep dive
Preserving Forgery Artifacts: AI-Generated Video Detection at Native Scale
Zhengcen Li, Chenyang Jiang, Hang Zhao, Shiyang Zhou, Yunyang Mo, Feng Gao, Fan Yang, Qiben Shan, Shaocong Wu, Jingyong Su
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/10/2026, 2:53:02 AM
Summary
The paper introduces a novel AI-generated video detection framework built on the Qwen2.5-VL Vision Transformer. It addresses limitations in existing detection methods, such as fixed-resolution preprocessing which discards high-frequency forgery artifacts, by operating natively at variable spatial resolutions and temporal durations. The authors also curate a large-scale dataset of 140K videos from 15 generators and a new benchmark, 'Magic Videos', to evaluate performance on modern, high-quality synthetic content.
Entities (4)
Relation Signals (2)
Qwen2.5-VL Vision Transformer → powers → AI-generated video detection framework
confidence 100% · we propose a novel detection framework built on the Qwen2.5-VL Vision Transformer
Magic Videos → evaluates → AI-generated video detection framework
confidence 95% · Magic Videos benchmark designed specifically for evaluating ultra-realistic synthetic content.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The rapid advancement of video generation models has enabled the creation of highly realistic synthetic media, raising significant societal concerns regarding the spread of misinformation. However, current detection methods suffer from critical limitations. They rely on preprocessing operations like fixed-resolution resizing and cropping. These operations not only discard subtle, high-frequency forgery traces but also cause spatial distortion and significant information loss. Furthermore, existing methods are often trained and evaluated on outdated datasets that fail to capture the sophistication of modern generative models. To address these challenges, we introduce a comprehensive dataset and a novel detection framework. First, we curate a large-scale dataset of over 140K videos from 15 state-of-the-art open-source and commercial generators, along with Magic Videos benchmark designed specifically for evaluating ultra-realistic synthetic content. In addition, we propose a novel detection framework built on the Qwen2.5-VL Vision Transformer, which operates natively at variable spatial resolutions and temporal durations. This native-scale approach effectively preserves the high-frequency artifacts and spatiotemporal inconsistencies typically lost during conventional preprocessing. Extensive experiments demonstrate that our method achieves superior performance across multiple benchmarks, underscoring the critical importance of native-scale processing and establishing a robust new baseline for AI-generated video detection.
Tags
Links
- Source: https://arxiv.org/abs/2604.04634v1
- Canonical: https://arxiv.org/abs/2604.04634v1
Trouble viewing inline? Open PDF directly →
Full Text
89,386 characters extracted from source content.
Expand or collapse full text
Published as a conference paper at ICLR 2026 PRESERVING FORGERY ARTIFACTS: AI-GENERATED VIDEO DETECTION AT NATIVE SCALE Zhengcen Li 1,2 , Chenyang Jiang 1,2 , Hang Zhao 1 , Shiyang Zhou 1 , Yunyang Mo 1 Feng Gao 3 , Fan Yang 3 , Qiben Shan 2 , Shaocong Wu 2† , Jingyong Su 1† 1 Harbin Institute of Technology, Shenzhen 2 Peng Cheng Laboratory 3 Peking University wushc@pcl.ac.cn sujingyong@hit.edu.cn Project Page: https://github.com/mgiant/Qwen2.5ViT-AIGVDetection ABSTRACT The rapid advancement of video generation models has enabled the creation of highly realistic synthetic media, raising significant societal concerns regarding the spread of misinformation. However, current detection methods suffer from critical limitations. They rely on preprocessing operations like fixed-resolution resizing and cropping. These operations not only discard subtle, high-frequency forgery traces but also cause spatial distortion and significant information loss. Further- more, existing methods are often trained and evaluated on outdated datasets that fail to capture the sophistication of modern generative models. To address these challenges, we introduce a comprehensive dataset and a novel detection frame- work. First, we curate a large-scale dataset of over 140K videos from 15 state-of- the-art open-source and commercial generators, along with Magic Videos bench- mark designed specifically for evaluating ultra-realistic synthetic content. In ad- dition, we propose a novel detection framework built on the Qwen2.5-VL Vision Transformer, which operates natively at variable spatial resolutions and tempo- ral durations. This native-scale approach effectively preserves the high-frequency artifacts and spatiotemporal inconsistencies typically lost during conventional pre- processing. Extensive experiments demonstrate that our method achieves superior performance across multiple benchmarks, underscoring the critical importance of native-scale processing and establishing a robust new baseline for AI-generated video detection. 1INTRODUCTION Artificial Intelligence-Generated Content (AIGC) has advanced rapidly, revolutionizing the creation of high-quality text (Yang et al., 2024; DeepSeek-AI, 2024), image (Esser et al., 2024; Labs, 2024), audio (Kreuk et al., 2023; Copet et al., 2023) and video (Brooks et al., 2024). Among these advance- ments, video generation has seen particularly significant progress, evolving from foundational mod- els like Stable Diffusion (Rombach et al., 2022) to more advanced architectures such as Diffusion Transformers (DiTs) (Peebles & Xie, 2023; Brooks et al., 2024), as well as proprietary commercial products (Pika Labs, 2023; Jimeng AI, 2024; Kuaishou, 2024). These developments have pushed the boundaries of deepfake technologies (Yang et al., 2022), enabling large-scale creation of fully AI-generated videos. However, the emergence of near-photorealistic synthetic videos poses serious threats to privacy, reputation, and public trust (Wang et al., 2024), underscoring the urgent need for effective detection and mitigation strategies against disinformation and misinformation. Deepfake detection (Yan et al., 2023) and AI-generated image detection (Wang et al., 2020; Zhu et al., 2023) have made significant progress in identifying manipulated content. However, existing deepfake detection methods (Qian et al., 2020; Xu et al., 2023; Oorloff et al., 2024; Nguyen et al., 2024) often face generalizability issues as they primarily focus on detecting facial forgeries. Mean- while, approaches for detecting images generated by Generative Adversarial Networks (GAN) and † Corresponding author. 1 arXiv:2604.04634v1 [cs.CV] 6 Apr 2026 Published as a conference paper at ICLR 2026 Figure 1: Resolution mismatch and generator quality strongly affect cross-generator video detection. Left: Detectors trained on 720p videos (top) and on lower-resolution videos (<720p; bottom) both exhibit a pronounced performance drop when evaluated at a spatial resolution different from that used during training. Right: We observe a strong positive correlation between generator quality (VBench score) and cross-validation performance (Pearson ρ = 0.86), indicating that higher- quality generators tend to yield more transferable training data for detector learning. These findings motivate a unified framework that is robust to resolution shifts and generator-specific artifacts. diffusion models (Wang et al., 2020; 2023c; Tan et al., 2024; Luo et al., 2024) are typically restricted to static media, leaving general spatiotemporal forgery detection largely unaddressed. Recent studies have begun to develop more robust solutions for AI-generated image and video de- tection (Yan et al., 2025; Li et al., 2025; Song et al., 2024; Chen et al., 2024b; Kundu et al., 2025b). A significant and shared limitation among these methods is the conventional preprocessing of resiz- ing (Yan et al., 2025) or cropping (Li et al., 2025) input frames to a fixed resolution (e.g., 224x224). Forgery detection methods often rely on two types of features, subtle artifacts and high-level se- mantics (Cheng et al., 2025). However, this fixed-resolution preprocessing degrades both types of features. Resizing distorts the original aspect ratio, misleading detectors into learning superficial distribution differences rather than robust and generalizable forgery features (Rajan et al., 2025). Cropping, meanwhile, can discard important content outside the selected area, thereby discarding global semantic cues of high-resolution content. Furthermore, both downsampling approaches de- grade the subtle, pixel-level artifacts that are critical for identifying synthetic media and capturing fine-grained inconsistencies (Corvi et al., 2025). Furthermore, progress in AI-generated video detection is hampered by the use of outdated synthetic data sources. Existing datasets (Chen et al., 2024b; Song et al., 2024) are predominantly com- posed of videos generated by earlier models, which typically exhibit low resolution, limited quality, and short durations. As a result, detection models trained on these datasets experience a signifi- cant performance drop when evaluated on modern AI-generated videos. To better understand these challenges, we conduct cross-validation experiments using existing detectors on a synthetic videos dataset sourced from 14 generative models. Our preliminary results reveal two critical insights, as illustrated in Figure 1. First, we observe a significant performance drop when detectors are evaluated on videos with different resolution from those in the training set. Second, detection performance is positively correlated with the quality of the video generators, meaning that stronger detectors require training on higher-quality, more realistic synthetic videos. These findings further highlight the im- portance of constructing a high-quality and diverse dataset, as well as a training framework capable of effectively handling videos with diverse resolutions, durations and generative sources. In response to the limitations of existing methods, we propose a unified framework that supports training and evaluation on videos with diverse resolutions and generative sources. First, we cu- rate a high-quality and diverse video dataset sourced from 15 representative video generation mod- els for training and develop a meticulously crafted pipeline to synthesize high-quality, human- indistinguishable videos for evaluation, termed Magic Videos. Second, we design a native-resolution 2 Published as a conference paper at ICLR 2026 training framework based on the Qwen2.5-VL Vision Transformer (Bai et al., 2025) (Qwen2.5- ViT), which unifies image and video modeling and enables the model to natively process videos with arbitrary spatial resolutions and temporal lengths. By removing the constraints of fixed-size downsampling preprocessing, our method achieves strong generalization capabilities to capture gen- eral spatiotemporal forgery artifacts. Extensive experiments on a wide range of benchmarks (Gen- video (Chen et al., 2024b), DVF (Song et al., 2024) and our proposed Magic Videos) demonstrate that our model is robust and achieves state-of-the-art performance in detecting AI-generated videos. Our Contributions are summarized as follows: • We introduce a new high-quality diverse dataset sourcing from 15 generators for training, and curate Magic Videos with 6 recent generators for evaluation, ensuring that both training and evaluation are aligned with the current generative quality of AIGC. • We propose a novel native-resolution framework built upon the Qwen2.5-ViT, which pro- cesses videos with native aspect ratio and variable resolution, preserving crucial forgery artifacts often lost during conventional resizing or cropping. • Through extensive experiments, we demonstrate that our method achieves state-of-the-art performance and robust generalization across a wide range of benchmarks, setting a new standard for AI-generated video detection. 2RELATED WORK 2.1VIDEO GENERATIVE MODELS Diffusion models (Ho et al., 2020; Song et al., 2022; Rombach et al., 2022) have significantly en- hanced the quality and controllability of image generation, inspiring researchers to extend these techniques to video generation tasks. Early work (Singer et al., 2022) propose incorporating motion dynamics into pre-trained text-to-image generation models. More recent studies (Chen et al., 2024a; Guo et al., 2024; Blattmann et al., 2023; Wang et al., 2023a; Wei et al., 2024) leverage latent-based diffusion models (Rombach et al., 2022) to generate short dynamic videos from text or image inputs. With the growing popularity of Diffusion Transformers (DiTs) (Peebles & Xie, 2023) in image gen- eration (Labs, 2024), DiT and its variants (Esser et al., 2024) have been widely proposed for video generation tasks (Ma et al., 2024b; Zheng et al., 2024; Brooks et al., 2024; Yang et al., 2025; Kong et al., 2024; Wan Team, 2025; Polyak et al., 2024). Besides Diffusion based methods, Generative Adversarial Networks (GANs) (Shen et al., 2023; Wang et al., 2023b) are also explored for video generation. The success of decoder-only architecture in language model has also motivated research in generating long videos using autoregressive models (Kondratyuk et al., 2024; Yu et al., 2023; Yin et al., 2025). Commercial video generation products (Brooks et al., 2024; Kuaishou, 2024; Jimeng AI, 2024; Pika Labs, 2023; MiniMax, 2024), employ complex and proprietary pipelines and pro- duces hyper-realistic videos. However, the lack of transparency surrounding these systems limits detailed analysis of their methodologies. In this paper, we propose a generative video dataset that encompasses most of the aforementioned architectures, including Diffusion U-Net (Chen et al., 2024a; Guo et al., 2024), DiT (Brooks et al., 2024; Wan Team, 2025; Ju et al., 2024; Zheng et al., 2024; Polyak et al., 2024; Lin et al., 2024; Ma et al., 2025), MMDiT (Kong et al., 2024; Si et al., 2025), and auto-regressive (Yin et al., 2025) models. The diversity of generative models included in our dataset ensures broad coverage and supports the generalizability of the proposed method. 2.2AI-GENERATED IMAGE AND VIDEO DETECTION AI-Generated Image Detection. As generative technologies rapidly advance, a growing number of forged images are now entirely synthesized by GANs (Goodfellow et al., 2014) and Diffusion models (Rombach et al., 2022), moving beyond traditional limited manipulation techniques. Con- sequently, substantial research efforts have focused on developing generalizable synthetic image detection methods (Tan et al., 2023; Ojha et al., 2023; Yan et al., 2024; Liu et al., 2024b), includ- ing approaches based on reconstruction error (Wang et al., 2023c; Luo et al., 2024; Guillaro et al., 2025), pixel-level features (Wang et al., 2020; Tan et al., 2024; Cheng et al., 2025), or adapting vi- sual backbones (Koutlis & Papadopoulos, 2024; Yan et al., 2025; Liu et al., 2024a). These methods 3 Published as a conference paper at ICLR 2026 RealVideos filter condense VideoGenera-onModel AI-GeneratedVideos Cap-oner Caption ... Qwen2.5-VLVisionTransformer❄/ Classifier ... w:1280 h: 720 t: 5.4s w:832 h: 480 t: 1.6s w:224 h: 224 t: 10.7s na#veresolu#on3Dpatchifica#on ... .. Figure 2: Overview of the data generation pipeline and the proposed detection framework. Left: We curate high-quality captions from real videos and refine them into prompts for state- of-the-art text-to-video generators, producing realistic synthetic videos for training and evaluation. Right: Our detector supports variable spatial resolutions and temporal lengths. It avoids fixed-size resizing/cropping and applies 3D patchification to preserve the input aspect ratio and fine-grained, high-frequency forensic cues that are often weakened by conventional downsampling. Built on the Qwen2.5-VL Vision Transformer, the framework models videos as sequences of spatiotemporal patches for robust AI-generated video detection. are typically trained on images generated by specific models (Karras et al., 2018; Song et al., 2022) and aim to achieve cross-architecture generalization. AI-Generated Video Detection. More recently, research has expanded to the detection of fully AI-generated videos (Ma et al., 2024a; Ni et al., 2024; Ji et al., 2024; Chang et al., 2025). VLM- based methods (Song et al., 2024; Wen et al., 2025) prompt large vision-language models to identify unnatural AI-like cues, while ViT-based methods (Chen et al., 2024b; Corvi et al., 2025) intro- duce forgery-posed generated datasets and design modules to detect spatial-temporal inconsisten- cies. However, existing methods for detecting AI-generated images and videos commonly suffer from a reliance on fixed resizing operations. Such preprocessing can lead to the loss of fine-grained details and spatial distortions, ultimately compromising model robustness across diverse inputs. In this work, we address this issue by training on native spatial resolution and temporal duration, with- out resizing or temporal padding. This design fundamentally avoids the pitfalls of conventional preprocessing and significantly enhances the model’s generalization capability. 3METHODOLOGY 3.1DATA CURATION Selection of Data Sources. The AI-generated videos are curated from multiple sources: (1) VBench (Huang et al., 2023; 2024), which provides generated videos from various text-to-video models using a predefined suite of diverse prompts; (2) Movie Gen (Polyak et al., 2024), which contributes videos generated by its proprietary model; and (3) A collection of highly realistic videos synthesized using various cutting-edge open-source and commercial models, guided by our custom- designed prompt library. Realistic Video Generation Pipeline. To evaluate the capability of generative content detectors in real-world scenarios, we design a pipeline for constructing synthetic videos that closely resem- ble authentic content. We prioritize scenarios that pose significant risks to information security, such as realistic landscapes, architectural scenes, and human interactions, as these categories are particularly susceptible to misuse and misinformation due to their inherent plausibility. Leveraging ShareGPT4Video (Chen et al., 2024c) repository of detailed and high-quality captions, we curate 4 Published as a conference paper at ICLR 2026 DataSourceNumberResolutionDuration Training Data (15 models) Vbench140K240p-768p1-10s Movie GenMovieGenBench20031920x108810.7s Wan2.1Open Source4501280x7205.4s Wan-1.3BOpen Source584832x4805s HailuoAPI (T2V-01)4501280x7205.6s SeaweedAPI (Jimeng-S2.0)4501472x8325s SeedanceAPI (Jimeng-S3.0)4501248x7045s StepVideoOpen Source450950x5408.2s Table 1: Dataset statistics for training/validation and the Magic Videos benchmark. See ap- pendix for details of these datasets. content specifically within these realism-oriented themes. To accommodate the capabilities of state- of-the-art architectures, we filter videos by duration (3-12 seconds) and caption length (fewer than 1000 characters). The curated prompts are further optimized using GPT-4o to condense the descrip- tion to under 500 characters. Table 1 summarizes videos that are synthesized by six distinct video generators using our comprehensive prompt library. These videos represent the current frontier of photorealistic synthetic content, enabling a rigorous assessment of detection models under practical and high-risk conditions. 3.2QWEN2.5-VIT Contemporary AI-generated content detectors primarily operate by identifying two categories of features: local artifacts and global semantic inconsistencies (Cheng et al., 2025). However, a com- mon practice in existing methodologies is to resize input images to a low, fixed resolution, typically 224x224 pixels. This downscaling operation adversely affects the features crucial for detection: it degrades subtle local artifacts and distorts global semantic structures. In this paper, we introduce a unified framework that processes images and videos at native resolution, thereby preserving the original forgery artifacts. The framework begins by tokenizing input videos into 3D patches at the native scale and adopts Qwen2.5-ViT (Bai et al., 2025) as a novel visual backbone for general video forgery detection. 3D Video Patchifying at Native Scale. We follow the video processing steps of (Bai et al., 2025), which introduces a 3D patch partitioning strategy that enables native-resolution inputs. For static images, it employs a standard spatial patch extraction method (e.g., 14x14 pixels). Unlike conven- tional ViTs that operate on static frames independently, our method extends patchification into the temporal dimension for video data. Given an input video tensor V ∈R T×H×W×C , it partitions V to non-overlapping 3D patches of size (P t ,P h ,P w ) = (2, 14, 14) and computes patch embedding via linear projection matrix E. This design eliminates the need for conventional resizing and padding operations, allowing the Transformer model to operate natively on both the spatial and temporal scales. The initial transformation of a raw video tensor V into a sequence of feature embeddings X (0) is described in Equation (1). The 3D patchification is particularly effective in detecting subtle texture artifacts and temporal inconsistencies at the patch level. By preserving the original resolu- tion during preprocessing, our method ensures that potential features critical for forgery detection remain intact and undistorted. X (0) = Unfold(V ;P t ,P h ,P w ) T · E(1) Transformer Layer Structure. Qwen2.5-ViT consists of 32 Transformer layers, each adopt- ing a pre-normalization structure, in which RMSNorm is applied before both the self-attention and feed-forward network (FFN). The FFN component employs the SwiGLU activation function. To effectively encode the spatial relationships between patches, 2D Rotary Positional Embedding (RoPE) (Su et al., 2023) is applied to the queries and keys in self-attention, enhancing the model’s extrapolation capability across input resolutions. The computations performed within each Trans- 5 Published as a conference paper at ICLR 2026 former layer are described as ˆ X (l) = X (l−1) + Attention(RMSNorm(X (l−1) )), X (l) = ˆ X (l) + FFN SwiGLU (RMSNorm( ˆ X (l) )). (2) In Equation (2), X (l−1) and X (l) denote the input and output hidden states of the l-th Transformer layer, respectively. Infrastructure Optimization for Efficiency. To address the computational challenges associated with high-resolution inputs, which typically lead to quadratic complexity, several optimizations are integrated. A batch packing strategy from NaViT (Dehghani et al., 2023) is adopted to allow the model to handle variable-length sequences without padding or attention masks. This is combined with Flash Attention (Dao, 2023), enabling GPU awareness of sequence boundaries and significantly improving both computational efficiency and memory usage through optimized CUDA kernels. In addition, a hybrid attention strategy is adopted where the majority of Transformer layers utilize 114× 114 windowed attention, ensuring that the computational cost scales linearly with the number of input patches. Classifier and Tuning Methods. For the final binary classification task of distinguishing between authentic and AI-generated content, we append a simple yet effective classification head to the Qwen2.5-ViT backbone. The output tokens from the final Transformer layer is first aggregated into a single, fixed-size feature vector using global average pooling. This vector is then passed through a single fully connected (FC) linear layer that outputs the logits corresponding to the “real” and “gen- erated” classes. To adapt the pre-trained model to this task, we explore three fine-tuning strategies: (1) Full Finetuning: Both the visual backbone and classification head are jointly optimized during training. (2) Linear-Probing: Serves as a baseline, where the entire vision backbone is frozen and only the classification head is trained. (3) Parameter-Efficient Fine-Tuning (PEFT): Specifically. we adopt Low-Rank Adaptation (LoRA (Hu et al., 2021)), which introduces small, trainable low-rank matrices into the frozen backbone, allowing only a subset of parameters to be updated. 4EXPERIMENTS 4.1DATASETS Training Dataset. We construct a training set of 70K AI-generated videos and 70K real videos. The synthetic videos are generated by VBench (Huang et al., 2023) using their prompt set, while the real ones are sampled from MSVD (Chen & Dolan, 2011) and Kinetics (Kay et al., 2017). For validation, we use 1,003 fake videos from MovieGenVideoBench (Polyak et al., 2024) and 1,000 real videos from Panda-70M (Chen et al., 2024d). ModelTraining DataMovie GenWan 2.1Wan-1.3BHailuoSeaweedSeedanceStepVideomACC RINE†ldm52.9745.3533.5651.451.6344.6570.2349.47 FatFormer†ProGAN50.0250.0050.1750.0050.0050.0050.2350.07 B-Free†SD 2.164.3053.7268.3255.8131.6336.0548.6049.02 Effort†SD 1.470.7473.2666.9567.4486.0583.2662.7973.29 WaveRep†Pyramid Flow65.3058.8453.659.5359.5357.2159.5358.04 F3Net 15Model-140K (Ours) 92.5171.8669.8666.9874.8873.2666.9870.64 TALL91.7164.6567.8161.8664.4263.4963.4964.29 NPR92.6671.6366.2772.3374.8873.2672.0971.74 TimeSformer91.4168.8472.0966.2869.7767.9167.2168.68 CLIP ViT-L/14 99.2078.1477.7477.2177.9177.2175.5877.30 X-CLIP-B/1698.5576.2872.4375.1275.1275.3572.3374.44 X-CLIP-L/1498.8580.0086.1379.5378.6080.0079.5380.63 Moon-ViT98.2576.7479.6275.8176.7475.8174.8876.60 Qwen2.5-ViT (Ours)97.2085.8183.3984.6583.9584.8876.5183.20 Table 2: Accuracy (ACC) Benchmarking Performance on on Movie Gen (val) and Magic Videos (test), reported per generator and averaged (mACC). † Results are produced with the official pretrained model. Best: bold; second best: underlined . 6 Published as a conference paper at ICLR 2026 ModelTraining DataMovie GenWan 2.1Wan-1.3BHailuoSeaweedSeedanceStepVideomAP RINE†ldm71.1140.4132.6947.4646.5438.9869.9346.00 FatFormer†ProGAN58.8439.0654.9643.8044.5542.1652.3646.15 B-Free†SD 2.170.3859.8073.8162.6236.0838.9352.3753.94 Effort†SD 1.4 80.6076.6173.7570.6795.4891.6563.4178.60 WaveRep†Pyramid Flow92.9783.5085.3195.6499.1278.6494.0489.38 F3Net 15Model-140K (Ours) 96.2085.7978.0580.3192.1387.3281.2684.14 TALL96.0783.5980.7878.5888.8182.8183.3982.99 NPR97.1086.9382.7087.5192.7192.0890.9688.82 TimeSformer 96.9182.4483.0475.4283.4980.2978.5580.54 CLIP ViT-L/1499.9598.7893.1487.4397.5892.0586.8592.64 X-CLIP-B/1699.8794.9494.0795.7687.6597.7881.6591.98 X-CLIP-L/14 99.9499.6282.7392.9199.1696.1895.7194.39 Moon-ViT99.2493.8490.3987.2494.1190.6881.6989.66 Qwen2.5-ViT (Ours)99.4696.6799.6894.2091.5996.9280.6393.28 Table 3: Average Precision (AP) Benchmarking Performance on on Movie Gen (val) and Magic Videos (test), reported per generator and averaged (mACC). † Results are produced with the official pretrained model. Best: bold; second best: underlined . Test Datasets. To evaluate robustness against state-of-the-art synthetic videos, we introduce the Magic Videos benchmark (Table 1), which consists of high-quality, hyper-realistic videos gener- ated by six cutting-edge video generation models using carefully curated prompts. Each generated video subset is paired with corresponding real videos to support binary classification evaluation. We report performance using Accuracy (ACC) and Average Precision (AP) per generator subset. Ex- ternal Benchmarks: In addition, we evaluate our method on the test sets of three external datasets: DVF (Song et al., 2024), GenVideo (Chen et al., 2024b), and DeepTraceReward (Fu et al., 2025). These datasets encompass diverse real-world sources and a wide range of video generation mod- els, enabling a more comprehensive assessment of detector performance across different synthesis techniques and generation stages. Baselines. We benchmark four categories of methods: (1) AI-generated video detection ap- proaches (M-Det (Song et al., 2024), DeMamba (Chen et al., 2024b), UNITE (Kundu et al., 2025b), TruthLens (Kundu et al., 2025a), and WaveRep (Corvi et al., 2025)); (2) visual and video foundation backbones (X-CLIP-B/16 (Ni et al., 2022), X-CLIP-L/14 (Ni et al., 2022), TimeS- former (Bertasius et al., 2021), and Moon-ViT (Du et al., 2025)); (3) deepfake detection methods (TALL (Xu et al., 2023) and F3Net (Qian et al., 2020)); and (4) general AI-generated image detec- tion methods (NPR (Tan et al., 2024), FatFormer (Liu et al., 2024a), RINE (Koutlis & Papadopoulos, 2024), B-Free (Guillaro et al., 2025), and Effort (Yan et al., 2025)). For image-based methods, we report the results by averaging logits over T frames to obtain video-level predictions. Implementation Details. We train our model for five epochs using the binary cross-entropy loss and the AdamW optimizer. The learning rate is set to 1× 10 −5 for full fine-tuning and 1× 10 −4 for parameter-efficient fine-tuning (PEFT). To balance performance and computational cost, we adopt the preprocessing strategy described in Bai et al. (2025); Du et al. (2025), which specifies minimum and maximum token budgets for image inputs. Each input frame is resized to the highest possible resolution within the (min pixels, maxpixels) range while preserving the original aspect ratio. In our experiments, we consider two resolution ranges per frame: (224× 224, 720× 720) and (224× 224, 448× 448). For temporal sampling, videos are decoded into frames at 2 fps. During training, we randomly sample T = 8 consecutive frames; during evaluation, we instead select the central T = 8 frames. Additional implementation details are provided in the Appendix. 4.2AI-GENERATED VIDEO DETECTION Evaluation on Magic Videos. The experimental results presented in Table 2 and Table 3 provide a comprehensive evaluation of our model against several distinct classes of methods. A notable observation is the underwhelming performance of models originally developed for AI-generated image detection, including RINE, FatFormer, B-Free, and Effort. These models exhibit relatively poor performance on video-based benchmarks, even compared to image-based methods that are trained on our video datasets. This discrepancy suggests a fundamental difference between forgery patterns present in static images and those in dynamic video sequences; features learned for detect- 7 Published as a conference paper at ICLR 2026 Method Video- Crafter Zero- scope Open- Sora Sora Pika Stable Diff. Stable Video AVG CNNDet ∗ 87.488.278.063.877.373.578.978.2 DIRE ∗ 55.961.853.860.565.862.769.962.1 M-Det 93.594.088.886.295.995.789.992.0 NPR86.685.696.081.094.671.197.087.4 TALL95.491.897.294.997.583.698.292.6 F3Net90.490.295.990.197.893.198.593.7 TimeSformer 94.592.798.092.598.492.499.595.4 Qwen2.5-ViT (Ours)93.599.898.696.499.195.699.797.6 Table 4: Benchmarking Results in terms of AUC Performance on DVF-Test Song et al. (2024). Results with * are derived from Song et al. (2024). Model AveragedOverall RecallF1APACCRecall UNITE89.60-92.76-- TruthLens ---90.49- DeMamba-CLIP91.5889.1993.4596.1492.29 NPR83.0147.9963.6686.7592.40 F3Net83.4856.7871.5788.2693.06 TimeSformer86.4265.3877.6787.5191.55 TALL89.4461.5176.6790.0591.76 CLIP-L87.7264.7377.6292.6490.57 XCLIP-B89.7953.7672.0492.6090.90 XCLIP-L88.9461.6078.8492.7492.43 Qwen2.5-ViT (Ours)91.1690.6496.1396.6493.18 Table 5: Benchmarking Results in terms of averaged Recall, F1, AP per subset and overall Recall and ACC Performance on Genvideo-Val (Chen et al., 2024b). Detailed results are in Appendix. MethodACC Fake ACC Real ACC GPT-590.784.698.8 GPT-4.1 92.989.197.9 Gemini 2.5 Pro84.375.795.8 VideoLLaMa3 7B10.038.14.3 Qwen2.5-VL 7B51.720.293.4 Qwen2.5-VL 32B47.48.998.5 Qwen2.5-VL 72B50.016.694.3 DeepTraceReward (w/ Qwen2.5 VL 7B) 74.755.7100.0 Qwen2.5-ViT (Ours)97.296.398.2 Table 6: Benchmarking Results in terms of ACC on DeepTraceReward (Fu et al., 2025). Results of baseline methods are re- ported in (Fu et al., 2025). ing image artifacts do not generalize well to the spatio-temporal domain required for video-level analysis. Similarly, methods designed specifically for deepfake detection, such as F3Net and TALL demonstrate limited effectiveness. While these models excel at identifying at facial manipulations, their specialization becomes a constraint when faced with the broader challenge of detecting fully synthesized videos. In contrast, pretrained visual backbones like TimeSformer, CLIP-ViT and X- CLIP exhibit competitive performance by leveraging extensive pre-training on diverse visual data. However, their effectiveness is ultimately constrained by architectural limitations. A primary issue is the conventional practice of resizing input frames to a fixed resolution of 224×224 pixels. This downsampling process can eliminate subtle forgery artifacts and disrupt global semantic features crucial for detecting sophisticated generative content. Moon-ViT (Du et al., 2025), which applies a similar processing pipeline based on NaViT, also suffers from this limitation as it operates on static images and cannot capture temporal inconsistencies. Our proposed method achieves the highest average scores in both ACC and AP, establishing a new state-of-the-art on these benchmarks. Our proposed method achieves the highest average accuracy (mACC) and highly competitive average precision (mAP), establishing a strong baseline. While our model does not yield the absolute best AP on every individual generator, it consistently delivers strong and balanced performance across all generator types, highlighting its exceptional generalizability. This superior performance is di- rectly attributed to our advanced architecture. By leveraging the Qwen2.5-ViT backbone, our model integrates native-resolution modeling with dynamic temporal duration modeling, thereby avoiding destructive downsampling and preserving the fidelity of forgery cues inherent in the original con- tent. By effectively capturing both fine-grained artifacts and high-level semantic inconsistencies, our model provides a more robust and accurate solution for detecting AI-generated videos. Evaluation on DVF-test. We use the same model weights trained on our 15model-140k dataset to directly perform cross-dataset evaluation on the DVF dataset (Song et al., 2024). The results are 8 Published as a conference paper at ICLR 2026 Archs.VariantsMagicGenvideoAvg. spatial resolution random crop to 224p62.6293.5078.06 random resize to 224p 73.6995.5284.61 dynamic [224p, 448p]81.1996.0188.60 dynamic [224p, 720p] 83.2096.6489.92 temporal resolution T =2 71.1594.7082.93 T =475.5894.4084.99 T =881.1996.0188.60 tuning mode LP70.6091.9181.26 LoRA(r=16) 78.7394.9586.84 full81.1996.0188.60 Table 7: Ablation studies regarding spatial-temporal resolution and tuning mode. We report averaged ACC(%) on Magic Videos and Genvideo. For temporal and tuning experiments, the spatial resolution is set to dynamic[224p, 448p]. presented in Table 4. Our model achieves the highest average AUC of 97.6, demonstrating the high quality of our training dataset and the strong generalizability of our model in detecting AI-generated videos across diverse generation techniques. Evaluation on GenVideo-Val. We directly evaluate the model trained on the 15model-140k dataset on GenVideo-Val (Chen et al., 2024b) using the same weights. Owing to the substantial class imbalance between real and generated samples in the GenVideo evaluation subsets, we report both overall recall and accuracy (ACC) for a more comprehensive comparison. As shown in Table 5, our method outperforms all baselines, including UNITE (Kundu et al., 2025b), which is specifically designed for AI-generated video detection, as well as larger MLLM-based TruthLens (Kundu et al., 2025a) and other baseline approaches trained on the same data as ours. Notably, despite using only one-fifteenth of the training data used by DeMamba (Chen et al., 2024b), our model achieves bet- ter performance and strong cross-dataset generalization, particularly on videos generated by earlier models. Evaluation on DeepTraceReward. To further demonstrate robustness against unseen generators, we evaluated our method on the DeepTraceReward (Fu et al., 2025), which contains 4,335 videos from 7 recent generators (including Pika-1.5, Kling-1.5, etc). Table 6 compares our Qwen2.5-ViT against leading multimodal LLMs. Our model achieves 97.2% accuracy, significantly outperform- ing large general-purpose VLMs (e.g., GPT-5, Gemini 2.5 Pro) on the binary classification task. Moreover, while general-purpose VLMs often struggle with detecting synthetic video, our model demonstrates balanced performance (96.3% Fake ACC vs. 98.2% Real ACC), proving its effective- ness in identifying artifacts from the latest generation engines without overfitting to specific training generators. 4.3ABLATION STUDY AND ANALYSIS We conduct a series of ablation studies, as detailed in Table 7, to systematically investigate the im- pact of spatial resolution, temporal resolution, and different fine-tuning strategies on our model’s performance. We observe that the benefits of high-fidelity inputs are substantially larger on Magic. In contrast, GenVideo contains lower-resolution videos with shorter durations; therefore, it is less sensitive to the performance degradation introduced by aggressive downsampling during prepro- cessing. As a result, improvements brought by higher spatial/temporal fidelity are more pronounced on Magic than on GenVideo. Ablation Study on Spatial Resolution. Our analysis reveals critical performance differences. The conventional random crop to 224p method yields the lowest average accuracy on high-resolution content. Switching to random resize to 224p boosts performance to 73.69, but this approach can still cause degradation of subtle artifacts. In contrast, our dynamic resolution strategy which preserves the original aspect ratio demonstrates markedly superior performance, with the overall average ac- 9 Published as a conference paper at ICLR 2026 100755025 JPEG Quality 85.0 87.5 90.0 92.5 95.0 97.5 100.0 Relative ACC (%) Ours NPR CLIP-L XCLIP-L 023333843 H264 CRF 70 75 80 85 90 95 100 0.20.40.60.81.0 Resize Scale Factor 60 65 70 75 80 85 90 95 100 0.000.050.100.150.20 Crop Factor 86 88 90 92 94 96 98 100 Figure 3: Robustness on MovieGen under compression and spatial perturbations (relative ACC). Perturbation methods include JPEG compression, H264 encoding, spatial resizing and crop- ping. curacy peaking at 89.92 when using resolutions up to 720p. This confirms our hypothesis that main- taining aspect ratio and processing at higher resolutions are critical for capturing subtle, pixel-level forgery artifacts. Ablation Study on Temporal Resolution. For all candidates, we sample the original videos at 2 fps and select random or center-aligned T frames during training and testing, respectively. We observe that incorporating more temporal context is beneficial. Increasing the number of sampled frames (T ) from 2 to 8 improves the average performance from 82.93 to 88.60. This suggests that longer sequences enhance the model’s ability to detect temporal inconsistencies common in AI-generated videos. Ablation Study on Tuning method. Regarding tuning strategies, full fine-tuning achieves the best average performance (88.60). Although the parameter-efficient LoRA approach significantly out- performs linear probing, full fine-tuning is justified for maximizing detection accuracy. Robustness Analysis. We evaluate our model’s robustness under common video perturbations, in- cluding compression, downscaling, and cropping, as shown in Figure 3. The model remains highly accurate under mild degradations such as moderate JPEG and H.264 compression. Performance drops become more pronounced with severe spatial changes. Notably, our model outperforms base- lines under aggressive downscaling (scale ≤ 0.4) and cropping (crop factor ≥ 0.15), though all methods are affected by extreme spatial loss. These results highlight strong robustness to moderate noise and sensitivity to substantial spatial degradation. 5CONCLUSION In this work, we tackle two critical limitations in current AIGC detection methodologies: the re- liance on outdated training data and the prevalent use of destructive fixed-resolution preprocess- ing. Our contributions are twofold. First, we introduce a comprehensive and up-to-date dataset comprising videos generated by 18 state-of-the-art generators, ensuring broader coverage of con- temporary synthesis techniques. Second, we propose a novel detection framework capable of op- erating directly on videos at dynamic spatial-temporal resolutions. By leveraging the Qwen2.5-VL ViT backbone, our method avoids the information loss associated with downsampling, successfully preserving both fine-grained forgery artifacts and high-level semantic inconsistencies. Extensive evaluations demonstrate that our approach achieves state-of-the-art performance and substantially improves cross-generator robustness and generalization. Limitations. Despite our efforts to construct a comprehensive dataset, the rapid evolution of gen- erative models poses an ongoing challenge. Continuous data updates will be necessary to keep pace with emerging architectures. Additionally, processing videos at their native resolution inevitably incurs higher computational costs compared to downscaling-based methods, which may limit de- ployment in resource-constrained environments. Future work will focus on improving the compu- tational efficiency of native-scale processing and investigating model explainability to deepen our understanding of the specific forensic traces leveraged for detection. 10 Published as a conference paper at ICLR 2026 6ACKNOWLEDGMENT This work was supported by the project of Peng Cheng Laboratory (PCL2025A14), the National Natural Science Foundation of China (Grant No. 62350710797), and the Guangdong Basic and Applied Basic Research Foundation (Grant No. 2023B1515120065). REFERENCES Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report, 2025. Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In ICML, p. 813–824, 2021. Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, Varun Jampani, and Robin Rom- bach. Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023. Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators, 2024. URL https://openai.com/index/sora/. Chirui Chang, Jiahui Liu, Zhengzhe Liu, Xiaoyang Lyu, Yi-Hua Huang, Xin Tao, Pengfei Wan, Di Zhang, and Xiaojuan Qi. How far are ai-generated videos from simulating the 3d visual world: A learned 3d evaluation approach. In ICCV, 2025. David Chen and William Dolan. Collecting highly parallel data for paraphrase evaluation. In Dekang Lin, Yuji Matsumoto, and Rada Mihalcea (eds.), ACL, p. 190–200, 2011. Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. Videocrafter2: Overcoming data limitations for high-quality video diffusion models. In CVPR, p. 7310–7320, 2024a. Haoxing Chen, Yan Hong, Zizheng Huang, Zhuoer Xu, Zhangxuan Gu, Yaohui Li, Jun Lan, Huijia Zhu, Jianfu Zhang, Weiqiang Wang, and Huaxiong Li. Demamba: Ai-generated video detection on million-scale genvideo benchmark, 2024b. Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, Li Yuan, Yu Qiao, Dahua Lin, Feng Zhao, and Jiaqi Wang. Sharegpt4video: Improving video understanding and generation with better captions. In NeurIPS, 2024c. Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, and Sergey Tulyakov. Panda-70m: Captioning 70m videos with multiple cross-modality teachers. In CVPR, 2024d. Siyuan Cheng, Lingjuan Lyu, Zhenting Wang, Xiangyu Zhang, and Vikash Sehwag. Co-spy: Com- bining semantic and pixel features to detect synthetic images by ai. In CVPR, 2025. Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Defossez. Simple and controllable music generation. In NeurIPS, volume 36, p. 47704–47720, 2023. Riccardo Corvi, Davide Cozzolino, Ekta Prashnani, Shalini De Mello, Koki Nagano, and Luisa Ver- doliva. Seeing what matters: Generalizable ai-generated video detection with forensic-oriented augmentation, 2025. Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning, 2023. DeepSeek-AI. Deepseek-v3 technical report, 2024. URL https://arxiv.org/abs/2412. 19437. 11 Published as a conference paper at ICLR 2026 Mostafa Dehghani, Basil Mustafa, Josip Djolonga, Jonathan Heek, Matthias Minderer, Mathilde Caron, Andreas Steiner, Joan Puigcerver, Robert Geirhos, Ibrahim Alabdulmohsin, Avital Oliver, Piotr Padlewski, Alexey Gritsenko, Mario Lu ˇ ci ́ c, and Neil Houlsby. Patch n’ pack: Navit, a vision transformer for any aspect ratio and resolution. In NeurIPS, 2023. Angang Du, Bohong Yin, Bowei Xing, Bowen Qu, Bowen Wang, Cheng Chen, Chenlin Zhang, Chenzhuang Du, Chu Wei, Congcong Wang, Dehao Zhang, Dikang Du, Dongliang Wang, En- ming Yuan, Enzhe Lu, Fang Li, Flood Sung, Guangda Wei, Guokun Lai, Han Zhu, Hao Ding, Hao Hu, Hao Yang, Hao Zhang, Haoning Wu, Haotian Yao, Haoyu Lu, Heng Wang, Hongcheng Gao, Huabin Zheng, Jiaming Li, Jianlin Su, Jianzhou Wang, Jiaqi Deng, Jiezhong Qiu, Jin Xie, Jinhong Wang, Jingyuan Liu, Junjie Yan, Kun Ouyang, Liang Chen, Lin Sui, Longhui Yu, Meng- fan Dong, Mengnan Dong, Nuo Xu, Pengyu Cheng, Qizheng Gu, Runjie Zhou, Shaowei Liu, Sihan Cao, Tao Yu, Tianhui Song, Tongtong Bai, Wei Song, Weiran He, Weixiao Huang, Weixin Xu, Xiaokun Yuan, Xingcheng Yao, Xingzhe Wu, Xinxing Zu, Xinyu Zhou, Xinyuan Wang, Y. Charles, Yan Zhong, Yang Li, Yangyang Hu, Yanru Chen, Yejie Wang, Yibo Liu, Yibo Miao, Yidao Qin, Yimin Chen, Yiping Bao, Yiqin Wang, Yongsheng Kang, Yuanxin Liu, Yulun Du, Yuxin Wu, Yuzhi Wang, Yuzi Yan, Zaida Zhou, Zhaowei Li, Zhejun Jiang, Zheng Zhang, Zhilin Yang, Zhiqi Huang, Zihao Huang, Zijia Zhao, and Ziwei Chen. Kimi-vl technical report, 2025. Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ̈ uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion En- glish, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach. Scaling rectified flow transformers for high-resolution image synthesis, 2024. Xingyu Fu, Siyi Liu, Yinuo Xu, Pan Lu, Guangqiuse Hu, Tianbo Yang, Taran Anantasagar, Christo- pher Shen, Yikai Mao, Yuanzhe Liu, Keyush Shah, Chung Un Lee, Yejin Choi, James Zou, Dan Roth, and Chris Callison-Burch. Learning human-perceived fakeness in ai-generated videos via multimodal llms, 2025. Yu Gao, Haoyuan Guo, Tuyen Hoang, Weilin Huang, Lu Jiang, Fangyuan Kong, Huixia Li, Jiashi Li, Liang Li, Xiaojie Li, Xunsong Li, Yifu Li, Shanchuan Lin, Zhijie Lin, Jiawei Liu, Shu Liu, Xiaonan Nie, Zhiwu Qing, Yuxi Ren, Li Sun, Zhi Tian, Rui Wang, Sen Wang, Guoqiang Wei, Guohong Wu, Jie Wu, Ruiqi Xia, Fei Xiao, Xuefeng Xiao, Jiangqiao Yan, Ceyuan Yang, Jianchao Yang, Runkai Yang, Tao Yang, Yihang Yang, Zilyu Ye, Xuejiao Zeng, Yan Zeng, Heng Zhang, Yang Zhao, Xiaozheng Zheng, Peihao Zhu, Jiaxin Zou, and Feilong Zuo. Seedance 1.0: Exploring the boundaries of video generation models, 2025. Anastasis Germanidis. Runway gen-3, 2024. URL https://runwayml.com/research/ introducing-gen-3-alpha. Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. In NeurIPS, 2014. Fabrizio Guillaro, Giada Zingarini, Ben Usman, Avneesh Sud, Davide Cozzolino, and Luisa Ver- doliva. A bias-free training paradigm for more general ai-generated image detection. In CVPR, 2025. Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffu- sion models without specific tuning. In ICLR, 2024. Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In NeurIPS, 2020. Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models, 2021. Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianx- ing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. Vbench: Comprehensive benchmark suite for video generative models, 2023. 12 Published as a conference paper at ICLR 2026 Ziqi Huang, Fan Zhang, Xiaojie Xu, Yinan He, Jiashuo Yu, Ziyue Dong, Qianli Ma, Nattapol Chan- paisit, Chenyang Si, Yuming Jiang, Yaohui Wang, Xinyuan Chen, Ying-Cong Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. Vbench++: Comprehensive and versatile benchmark suite for video generative models, 2024. Lichuan Ji, Yingqi Lin, Zhenhua Huang, Yan Han, Xiaogang Xu, Jiafei Wu, Chong Wang, and Zhe Liu. Distinguish any fake videos: Unleashing the power of large-scale data and motion features, 2024. Jimeng AI.https://jimeng.jianying.com/ai-tool/home, 2024.URL https://jimeng. jianying.com/ai-tool/home. Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan, Xintao Wang, Ailing Zeng, Yu Xiong, Qiang Xu, and Ying Shan. Miradata: A large-scale video dataset with long durations and structured captions. In NeurIPS, 2024. Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for im- proved quality, stability, and variation. In ICLR, 2018. Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijaya- narasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman. The kinetics human action video dataset. arXiv, p. 1–22, 2017. Dan Kondratyuk, Lijun Yu, Xiuye Gu, Jos ́ e Lezama, Jonathan Huang, Grant Schindler, Rachel Hor- nung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng, Josh Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso Martinez, David Minnen, Mikhail Sirotenko, Kihyuk Sohn, Xuan Yang, Hartwig Adam, Ming-Hsuan Yang, Irfan Essa, Huisheng Wang, David A. Ross, Bryan Seybold, and Lu Jiang. Videopoet: A large language model for zero-shot video generation. In ICML, 2024. Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, Kathrina Wu, Qin Lin, Junkun Yuan, Yanxin Long, Aladdin Wang, An- dong Wang, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Hongmei Wang, Jacob Song, Jiawang Bai, Jianbing Wu, Jinbao Xue, Joey Wang, Kai Wang, Mengyang Liu, Pengyu Li, Shuai Li, Weiyan Wang, Wenqing Yu, Xinchi Deng, Yang Li, Yi Chen, Yutao Cui, Yuanbo Peng, Zhen- tao Yu, Zhiyu He, Zhiyong Xu, Zixiang Zhou, Zunnan Xu, Yangyu Tao, Qinglin Lu, Songtao Liu, Daquan Zhou, Hongfa Wang, Yong Yang, Di Wang, Yuhong Liu, Jie Jiang, and Caesar Zhong. Hunyuanvideo: A systematic framework for large video generative models, 2024. Christos Koutlis and Symeon Papadopoulos. Leveraging representations from intermediate encoder- blocks for synthetic image detection. In ECCV, 2024. Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre D ́ efossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. Audiogen: Textually guided audio generation. In ICLR, 2023. Joseph B Kruskal. Nonmetric multidimensional scaling: a numerical method. Psychometrika, 29 (2):115–129, 1964. Kuaishou. https://klingai.kuaishou.com, 2024. URL https://klingai.kuaishou.com. Rohit Kundu, Athula Balachandran, and Amit K. Roy-Chowdhury. Truthlens: Explainable deepfake detection for face manipulated and fully synthetic data, 2025a. Rohit Kundu, Hao Xiong, Vishal Mohanty, Athula Balachandran, and Amit K. Roy-Chowdhury. Towards a universal synthetic video detector: From face or background manipulations to fully ai-generated content. In CVPR, 2025b. Black Forest Labs. Flux. https://github.com/black-forest-labs/flux, 2024. Ouxiang Li, Jiayin Cai, Yanbin Hao, Xiaolong Jiang, Yao Hu, and Fuli Feng. Improving synthetic image detection towards generalization: An image transformation perspective. In KDD, 2025. 13 Published as a conference paper at ICLR 2026 Zongyu Lin, Wei Liu, Chen Chen, Jiasen Lu, Wenze Hu, Tsu-Jui Fu, Jesse Allardice, Zhengfeng Lai, Liangchen Song, Bowen Zhang, Cha Chen, Yiran Fei, Yifan Jiang, Lezhi Li, Yizhou Sun, Kai-Wei Chang, and Yinfei Yang. Stiv: Scalable text and image conditioned video generation, 2024. Huan Liu, Zichang Tan, Chuangchuang Tan, Yunchao Wei, Jingdong Wang, and Yao Zhao. Forgery- aware adaptive transformer for generalizable synthetic image detection. In CVPR, p. 10770– 10780, 2024a. Huan Liu, Zichang Tan, Chuangchuang Tan, Yunchao Wei, Jingdong Wang, and Yao Zhao. Forgery- aware adaptive transformer for generalizable synthetic image detection. In CVPR, p. 10770– 10780, 2024b. Yaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang, Yong Zhang, Haoxin Chen, Yang Liu, Tieyong Zeng, Raymond Chan, and Ying Shan. Evalcrafter: Benchmarking and evaluating large video generation models. In CVPR, 2024c. Lumalabs.https://lumalabs.ai/dream-machine, 2024.URL https://lumalabs.ai/ dream-machine. Yunpeng Luo, Junlong Du, Ke Yan, and Shouhong Ding. Lareˆ2: Latent reconstruction error based method for diffusion-generated image detection. In CVPR, p. 17006–17015, 2024. Guoqing Ma, Haoyang Huang, Kun Yan, Liangyu Chen, Nan Duan, Shengming Yin, Changyi Wan, Ranchen Ming, Xiaoniu Song, Xing Chen, Yu Zhou, Deshan Sun, Deyu Zhou, Jian Zhou, Kaijun Tan, Kang An, Mei Chen, Wei Ji, Qiling Wu, Wen Sun, Xin Han, Yanan Wei, Zheng Ge, Ao- jie Li, Bin Wang, Bizhu Huang, Bo Wang, Brian Li, Changxing Miao, Chen Xu, Chenfei Wu, Chenguang Yu, Dapeng Shi, Dingyuan Hu, Enle Liu, Gang Yu, Ge Yang, Guanzhe Huang, Gulin Yan, Haiyang Feng, Hao Nie, Haonan Jia, Hanpeng Hu, Hanqi Chen, Haolong Yan, Heng Wang, Hongcheng Guo, Huilin Xiong, Huixin Xiong, Jiahao Gong, Jianchang Wu, Jiaoren Wu, Jie Wu, Jie Yang, Jiashuai Liu, Jiashuo Li, Jingyang Zhang, Junjing Guo, Junzhe Lin, Kaixiang Li, Lei Liu, Lei Xia, Liang Zhao, Liguo Tan, Liwen Huang, Liying Shi, Ming Li, Mingliang Li, Muhua Cheng, Na Wang, Qiaohui Chen, Qinglin He, Qiuyan Liang, Quan Sun, Ran Sun, Rui Wang, Shaoliang Pang, Shiliang Yang, Sitong Liu, Siqi Liu, Shuli Gao, Tiancheng Cao, Tianyu Wang, Weipeng Ming, Wenqing He, Xu Zhao, Xuelin Zhang, Xianfang Zeng, Xiaojia Liu, Xuan Yang, Yaqi Dai, Yanbo Yu, Yang Li, Yineng Deng, Yingming Wang, Yilei Wang, Yuanwei Lu, Yu Chen, Yu Luo, Yuchu Luo, Yuhe Yin, Yuheng Feng, Yuxiang Yang, Zecheng Tang, Zekai Zhang, Zidong Yang, Binxing Jiao, Jiansheng Chen, Jing Li, Shuchang Zhou, Xiangyu Zhang, Xinhao Zhang, Yibo Zhu, Heung-Yeung Shum, and Daxin Jiang. Step-video-t2v technical report: The practice, challenges, and future of video foundation model, 2025. Long Ma, Jiajia Zhang, Hongping Deng, Ningyu Zhang, Qinglang Guo, Haiyang Yu, Yong Liao, and Pengyuan Zhou. Decof: Generated video detection via frame consistency: The first benchmark dataset, 2024a. Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. Latte: Latent diffusion transformer for video generation, 2024b. MiniMax.Minimaxofficiallyreleasesthevideo-01videogenerationmodel, https://hailuoai.com/video, 2024. URL https://w.minimax.io/news/video-01. mixkit. mixkit. https://mixkit.com/videos/, 2024. Dat Nguyen, Nesryne Mejri, Inder Pal Singh, Polina Kuleshova, Marcella Astrid, Anis Kacem, Enjie Ghorbel, and Djamila Aouada. Laa-net: Localized artifact attention network for quality-agnostic and generalizable deepfake detection. In CVPR, p. 17395–17405, 2024. Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang, and Haibin Ling. Expanding language-image pretrained models for general video recog- nition. In ECCV, 2022. Zhen-Liang Ni, Qiangyu Yan, Tianning Yuan, Mouxiao Huang, Hailin Hu, Xinghao Chen, and Yunhe Wang. Genvidbench: A challenging benchmark for detecting ai-generated video, 2024. 14 Published as a conference paper at ICLR 2026 Utkarsh Ojha, Yuheng Li, and Yong Jae Lee. Towards universal fake image detectors that generalize across generative models. In CVPR, p. 24480–24489, 2023. Trevine Oorloff, Surya Koppisetti, Nicol ` o Bonettini, Divyaraj Solanki, Ben Colman, Yaser Yacoob, Ali Shahriyari, and Gaurav Bharaj. Avff: Audio-visual feature fusion for video deepfake detec- tion. In CVPR, p. 27102–27112, 2024. William Peebles and Saining Xie. Scalable diffusion models with transformers. In ICCV, 2023. pexels. pexels. https://w.pexels.com/videos/, 2024. Pika Labs. https://pika.art/, 2023. URL https://pika.art/. Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, David Yan, Dhruv Choudhary, Dingkang Wang, Geet Sethi, Guan Pang, Haoyu Ma, Ishan Misra, Ji Hou, Jialiang Wang, Kiran Ja- gadeesh, Kunpeng Li, Luxin Zhang, Mannat Singh, Mary Williamson, Matt Le, Matthew Yu, Mitesh Kumar Singh, Peizhao Zhang, Peter Vajda, Quentin Duval, Rohit Girdhar, Roshan Sum- baly, Sai Saketh Rambhatla, Sam Tsai, Samaneh Azadi, Samyak Datta, Sanyuan Chen, Sean Bell, Sharadh Ramaswamy, Shelly Sheynin, Siddharth Bhattacharya, Simran Motwani, Tao Xu, Tianhe Li, Tingbo Hou, Wei-Ning Hsu, Xi Yin, Xiaoliang Dai, Yaniv Taigman, Yaqiao Luo, Yen- Cheng Liu, Yi-Chiao Wu, Yue Zhao, Yuval Kirstain, Zecheng He, Zijian He, Albert Pumarola, Ali Thabet, Artsiom Sanakoyeu, Arun Mallya, Baishan Guo, Boris Araya, Breena Kerr, Car- leigh Wood, Ce Liu, Cen Peng, Dimitry Vengertsev, Edgar Schonfeld, Elliot Blanchard, Felix Juefei-Xu, Fraylie Nord, Jeff Liang, John Hoffman, Jonas Kohler, Kaolin Fire, Karthik Sivaku- mar, Lawrence Chen, Licheng Yu, Luya Gao, Markos Georgopoulos, Rashel Moritz, Sara K. Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du. Movie gen: A cast of media foundation models, 2024. Yuyang Qian, Guojun Yin, Lu Sheng, Zixuan Chen, and Jing Shao. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In ECCV, 2020. Anirudh Sundara Rajan, Utkarsh Ojha, Jedidiah Schloesser, and Yong Jae Lee. Aligned datasets improve detection of latent diffusion-generated images. In ICLR, 2025. Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ̈ orn Ommer. High- resolution image synthesis with latent diffusion models. In CVPR, p. 10684–10695, 2022. Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao, Shanchuan Lin, Zhibei Ma, Haoyuan Guo, Hao Chen, Lu Qi, Sen Wang, Feng Cheng, Feilong Zuo Xuejiao Zeng, Ziyan Yang, Fangyuan Kong, Zhiwu Qing, Fei Xiao, Meng Wei, Tuyen Hoang, Siyu Zhang, Peihao Zhu, Qi Zhao, Jiangqiao Yan, Liangke Gui, Sheng Bi, Jiashi Li, Yuxi Ren, Rui Wang, Huixia Li, Xuefeng Xiao, Shu Liu, Feng Ling, Heng Zhang, Houmin Wei, Huafeng Kuang, Jerry Duncan, Junda Zhang, Junru Zheng, Li Sun, Manlin Zhang, Renfei Sun, Xiaobin Zhuang, Xiaojie Li, Xin Xia, Xuyan Chi, Yanghua Peng, Yuping Wang, Yuxuan Wang, Zhongkai Zhao, Zhuo Chen, Zuquan Song, Zhenheng Yang, Jiashi Feng, Jianchao Yang, and Lu Jiang. Seaweed-7b: Cost-effective training of video generation foundation model, 2025. Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Mostgan-v: Video generation with temporal motion styles. In CVPR, p. 5652–5661, 2023. Chenyang Si, Weichen Fan, Zhengyao Lv, Ziqi Huang, Yu Qiao, and Ziwei Liu. Repvideo: Rethink- ing cross-layer representation for video generation, 2025. Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make-a-video: Text-to-video generation without text-video data, 2022. Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In ICLR, 2022. Xiufeng Song, Xiao Guo, Jiache Zhang, Qirui Li, Lei Bai, Xiaoming Liu, Guangtao Zhai, and Xiaohong Liu. On learning multi-modal forgery representation for diffusion generated video detection. In NeurIPS, 2024. 15 Published as a conference paper at ICLR 2026 Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. Roformer: En- hanced transformer with rotary position embedding, 2023. Chuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu, and Yunchao Wei. Learning on gradients: Generalized artifacts representation for gan-generated images detection. In CVPR, p. 12105– 12114, 2023. Chuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu, Ping Liu, and Yunchao Wei. Rethinking the up-sampling operations in cnn-based generative network for generalizable deepfake detection. In CVPR, p. 28130–28139, 2024. Wan Team. Wan: Open and advanced large-scale video generative models, 2025. Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A. Efros. Cnn-generated images are surprisingly easy to spot... for now. In CVPR, p. 8695–8704, 2020. Tao Wang, Yushu Zhang, Shuren Qi, Ruoyu Zhao, Zhihua Xia, and Jian Weng. Security and privacy on generative data in aigc: A survey. ACM Computing Surveys, 57(4):1–34, 2024. Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, Yuwei Guo, Tianxing Wu, Chenyang Si, Yuming Jiang, Cunjian Chen, Chen Change Loy, Bo Dai, Dahua Lin, Yu Qiao, and Ziwei Liu. Lavie: High-quality video generation with cascaded latent diffusion models. IJCV, 2023a. Yuhan Wang, Liming Jiang, and Chen Change Loy. Styleinv: A temporal style modulated inversion network for unconditional video generation. In ICCV, p. 22851–22861, 2023b. Zhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang, Hezhen Hu, Hong Chen, and Houqiang Li. Dire for diffusion-generated image detection. In ICCV, p. 22445–22455, 2023c. Yujie Wei, Shiwei Zhang, Zhiwu Qing, Hangjie Yuan, Zhiheng Liu, Yu Liu, Yingya Zhang, Jingren Zhou, and Hongming Shan. Dreamvideo: Composing your dream videos with customized subject and motion. In CVPR, 2024. Haiquan Wen, Yiwei He, Zhenglin Huang, Tianxiao Li, Zihan Yu, Xingru Huang, Lu Qi, Baoyuan Wu, Xiangtai Li, and Guangliang Cheng. Busterx: Mllm-powered ai-generated video forgery detection and explanation, 2025. Shiyu Wu, Jing Liu, Jing Li, and Yequan Wang. Few-shot learner generalizes across ai-generated image detection. In ICML, 2025. Yuting Xu, Jian Liang, Gengyun Jia, Ziming Yang, Yanhao Zhang, and Ran He. Tall: Thumbnail layout for deepfake video detection. In ICCV, p. 22658–22668, 2023. Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Chenxu Zhang, Jiashi Feng, and Mike Zheng Shou. MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model. In CVPR. arXiv, 2024. doi: 10.48550/arXiv.2311.16498. URL http: //arxiv.org/abs/2311.16498. arXiv:2311.16498 [cs]. Zhiyuan Yan, Yong Zhang, Xinhang Yuan, Siwei Lyu, and Baoyuan Wu. Deepfakebench: A com- prehensive benchmark of deepfake detection. In NeurIPS, 2023. Zhiyuan Yan, Yuhao Luo, Siwei Lyu, Qingshan Liu, and Baoyuan Wu. Transcending forgery speci- ficity with latent space augmentation for generalizable deepfake detection. In CVPR, p. 8984– 8994, 2024. Zhiyuan Yan, Jiangming Wang, Peng Jin, Ke-Yue Zhang, Chengchun Liu, Shen Chen, Taiping Yao, Shouhong Ding, Baoyuan Wu, and Li Yuan. Orthogonal subspace decomposition for generaliz- able ai-generated image detection. In ICML, 2025. 16 Published as a conference paper at ICLR 2026 An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115, 2024. Tianyi Yang, Zixuan Huang, Junjie Cao, et al. Deepfake network architecture attribution. In Pro- ceedings of the AAAI Conference on Artificial Intelligence, volume 36, p. 4662–4670, 2022. Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, Da Yin, Xiaotao Gu, Yuxuan Zhang, Weihan Wang, Yean Cheng, Ting Liu, Bin Xu, Yuxiao Dong, and Jie Tang. Cogvideox: Text-to-video diffusion models with an expert transformer. In ICLR, 2025. Tianwei Yin, Qiang Zhang, Richard Zhang, William T. Freeman, Fredo Durand, Eli Shechtman, and Xun Huang. From slow bidirectional to fast autoregressive video diffusion models, 2025. Lijun Yu, Yong Cheng, Kihyuk Sohn, Jos ́ e Lezama, Han Zhang, Huiwen Chang, Alexander G. Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang. Magvit: Masked generative video transformer. In CVPR, p. 10459–10469, 2023. Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. Open-sora: Democratizing efficient video production for all, March 2024. URL https://github.com/hpcaitech/Open-Sora. Mingjian Zhu, Hanting Chen, Qiangyu Yan, Xudong Huang, Guanyu Lin, Wei Li, Zhijun Tu, Hailin Hu, Jie Hu, and Yunhe Wang. Genimage: A million-scale benchmark for detecting ai-generated image, 2023. 17 Published as a conference paper at ICLR 2026 APPENDIX This appendix provides a detailed analysis of our dataset, implementation details, additional experi- mental results, and visualizations: • Section A: Data distribution and analysis of our dataset. • Section B: Cross-validation experiment. • Section C: Additional implementation details for both our method and the baseline meth- ods. • Section D: Additional experimental results and ablation studies. • Section E: Visualizations and discussion. ADATASET COMPOSITION Model / Video SourceVer.AvailabilityVideosResolutionFPSFrameDuration Real Kinetics-400 (Kay et al., 2017)17.05real videos70K144p-720p--5-10s RepVideo (Si et al., 2025)25.01open-source4720720x4808496.1s Wan2.1 (Wan Team, 2025)25.01open-source47251280x72016815.0s CausVid (5s) (Yin et al., 2025)25.01open-source4720640x352241205.0s Apple-STIV (Lin et al., 2024)24.12open-report4715512x51260601.0s Sora (Brooks et al., 2024)24.12private4720854x480301505.0s HunyuanVideo (Kong et al., 2024)24.12open-source47251280x720241295.4s Gen-3 (Germanidis, 2024)24.06private47071280x7682425610.7s Luma (Lumalabs, 2024) 24.06private46801360x752241215.0s Kling (Kuaishou, 2024)24.06private46791280x720301535.1s Jimeng (Jimeng AI, 2024)24.05private62141280x72089612.0s OpenSora V1.1 (Zheng et al., 2024)24.04open-source4720424x2408648.0s Mira (Ju et al., 2024)24.04open-source4721384x24066010.0s VideoCrafter-2.0 (Chen et al., 2024a)24.01open-source4720320x51210161.6s Pika 1.0 (Pika Labs, 2023)23.11private47151280x72024723.0s AnimateDiff-V2 (Guo et al., 2024)23.09open-source4715512x5128162.0s Overall Fake--70,692240-720p6-6016-2561-12s Table 8: Statistics of real and synthetic videos in the proposed training set. Model / VideoSplitVideosResolutionFPSFrameDuration Movie Gen (Polyak et al., 2024)validation (fake)10031920x10882425610.7s Panda-70M (Chen et al., 2024d) validation (real)1000720p6-30-10-50s Mixkit (mixkit, 2024) test (real) 215720p15-60-10-17s Pexels (pexels, 2024)292720p24-60-6-39s Wan2.1 (Wan Team, 2025) test (fake) 2151280x720301615.4s Wan-1.3B (Wan Team, 2025)292832x48016815.0s Hailuo (MiniMax, 2024)2151280x720251415.6s Seaweed (Seawead et al., 2025)2151472x832241215.0s Seedance (Gao et al., 2025)2151248x704241215.0s StepVideo (Ma et al., 2025)215960×540252048.2s Table 9: Statistics of real and synthetic videos in the proposed validation and Magic Videos Benchmark. A.1TRAINING SET. Table 8 provides a comprehensive summary of the training dataset used in our work. Previous research has emphasized the critical importance of dataset quality and diversity in training robust detectors (Rajan et al., 2025), especially given the variety of artifacts produced by different gener- ative models (Wu et al., 2025). To advance the field of AI-generated video detection, we curated a 18 Published as a conference paper at ICLR 2026 large-scale dataset comprising outputs from 15 distinct video generation models. The majority of these synthetic videos are sourced from VBench (Huang et al., 2023), a benchmark selected for its high-quality prompt library and extensive evaluation of state-of-the-art models. This choice allowed us to avoid the costly and time-consuming processes of large-scale video filtering, quality control, and generation while ensuring high quality and consistency of generated video data. Our dataset reflects the diverse and evolving landscape of video generation, featuring models de- veloped between 2023 and 2025. It includes a wide range of model types in terms of availabil- ity (i.e., open-source, open-report, and private) and architecture (e.g., Diffusion U-Net, DiT-based, auto-regressive models, and others with undisclosed architectures). The models differ significantly in training methodology, data scale, output resolution, and video duration, contributing to a richly diverse training set. To complement the synthetic videos, we sampled an equal number of real videos from Kinetics- 400 (Kay et al., 2017). These were carefully selected to match the resolution, duration, and encoder distribution of the generated videos. This matching is essential for reducing potential biases and ensuring that the learned features are genuinely discriminative between real and fake content. A key feature of our dataset is that all generative models were conditioned on the same prompt library, ensuring a shared semantic distribution across the generated videos. This unique setup enables controlled cross-validation experiments, allowing us to investigate inter-model relationships and identify key factors that influence detector performance, as discussed in Section B. A.2VALIDATION AND TEST DATA Table 9 presents the composition of our validation set and introduces a novel, high-quality Magic Videos Benchmark, which we name the Magic Videos Benchmark. Validation Set. Rather than adopting the common practice of partitioning a subset of the train- ing data, we constructed the validation set from videos generated by Movie Gen (Polyak et al., 2024), a model that is architecturally and semantically similar but not identical to the models used in training. These synthetic videos are paired with 1,000 of real videos sampled from the Panda- 70M (Chen et al., 2024d) dataset. During training, we apply early stopping based on the validation loss computed on this set. This strategy helps mitigate overfitting to the specific models and sce- narios encountered during training, promoting the selection of a model checkpoint with stronger generalization capabilities. Magic Videos Benchmark. We identified a critical gap in existing benchmarks: they often lack coverage of the latest generative models and may exhibit evaluation biases. To address this, we constructed the Magic Videos Benchmark using a high-quality video generation pipeline, as intro- duced in Section 3 of the main paper. This benchmark includes real videos from two premium platforms—Mixkit(mixkit, 2024) and Pexels(pexels, 2024)—covering a diverse range of common scenes such as landscapes, architecture, human subjects, and news footage. These videos are pro- vided at resolutions up to 1080p to ensure both high fidelity and content diversity. For evaluation, real and generated videos are matched into balanced subsets, allowing for the computation of accu- racy and other performance metrics. To generate the synthetic counterparts, we first applied ShareGPT4Video (Chen et al., 2024c) to pro- duce high-quality captions for the real videos. These captions were then refined through a rigorous process of filtering, rewriting, and final prompt polishing. The resulting prompts were input to six advanced text-to-video models, comprising both open-source and commercial systems. Below, we detail the generative models used to construct the Magic Videos Benchmark: • Wan2.1 (Wan Team, 2025): We used the Wanxiang platform API with the ”professional” model, default settings, and prompt optimization disabled. Prompts were derived from the Mixkit collection. This model may apply post-processing, resulting in a higher frame rate than Wan-14B. • Wan-1.3B (Wan Team, 2025): Videos were generated using the official open-source imple- mentation and pre-trained model, with prompts from the Pexels collection. 19 Published as a conference paper at ICLR 2026 • Hailuo (MiniMax, 2024): Accessed via the MiniMax-T2V-01 commercial API, this model was configured to generate 5-second videos using prompts from the Mixkit collection. Prompt optimization was not applied. • Seaweed(Seawead et al., 2025): As official model weights are not publicly available, we used the commercial model Jimeng-S2.0(Jimeng AI, 2024), which is based on the Seaweed-alpha model. Prompts were sourced from the Mixkit collection. Generation was performed using prompts from the Mixkit collection. • Seedance(Gao et al., 2025): In place of unavailable official weights, we used the commer- cial model Jimeng-S3.0(Jimeng AI, 2024), corresponding to the Seedance 1.0 Mini model. Prompts were sourced from the Mixkit collection. • StepVideo (Ma et al., 2025): Videos were generated using the official API with the Step- Video-T2V endpoint (544px × 992px × 204f), using prompts from the Mixkit collection. BCROSS-VALIDATION EXPERIMENT B.1EXPERIMENT SETUP Cross-Validation Setup. This experiment focuses on in-domain, cross-model validation of detec- tors. The benchmark utilizes data generated by 15 models from VBench (Huang et al., 2023), which evaluates various generative models using a shared set of predefined prompts. Because all mod- els generate videos from the same prompt library, we consider their outputs to belong to the same semantic domain. Let F i denote the subset of videos generated by model i, and let R 0 represent a fixed set of real videos, sampled to contain the same number of examples as each F i . For each model i, we train a deepfake detector on the dataset F i ,R 0 and evaluate its performance on all other generated subsets F j (for j ̸= i). This setup allows us to rigorously assess the generalization ability of detectors across different generative architectures while keeping the semantic domain fixed. It also provides a controlled environment for analyzing the relationships between generative model ar- chitectures and detection performance. This Cross-Validation Benchmark produces an n×n matrix M, whereM[i,j] represents the recall of a detection model trained on subset i and evaluated on subset j. Based on preliminary observations, we propose the following two hypotheses, which will be validated in subsequent experiments. Similarity Between Generative Models. The matrix entry M[i,j] reflects the output similarity between generative models i and j, influenced by factors such as model architecture, sampling strategies, and training data. We observe that models with more similar architectures tend to exhibit higher cross-validation accuracy between them. To quantify this relationship, we define a non- directional distance metric, d(i,j) = 1− 0.5× (M[i,j] + M[j,i]). Using this metric, we apply Non-metric Multidimensional Scaling (MDS) (Kruskal, 1964) to produce a 2D spatial representation of the generative models. This visualization aids in understanding the architectural relationships and clustering patterns among the models, offering insights into how architectural similarity correlates with cross-detection performance. Impact of Generation Quality. In addition to architecture, M[i,j] is also influenced by the gen- eration quality of model i. We hypothesize that higher-quality synthetic videos provide more re- alistic and informative supervision signals, enabling the classifier to learn more effective forgery- discriminative features. Since ground-truth quality labels are unavailable, we adopt scores from recent T2V benchmarks (Huang et al., 2023; Liu et al., 2024c; Huang et al., 2024) as a proxy for generation quality. To assess the relationship between generation quality and detection effective- ness, we compute Pearson correlation coefficients (ρ) between the benchmark quality scores and corresponding detection accuracies. B.2CROSS-VALIDATION RESULTS Cross Validation. As discussed above, we use the cross-validation matrixM to evaluate the similarity between generative models. Four detection models—F3Net(Qian et al., 2020), X-CLIP- B/32(Ni et al., 2022), TALL(Xu et al., 2023), and NPR(Tan et al., 2024)—are trained on 5K real videos from MSR-VTT and 5K generated videos from each specific model subset. These detectors 20 Published as a conference paper at ICLR 2026 wan21 hunyuan kling sora gen3 repvideo jimeng luma mira pika opensorav1-1 STIV causvid videocrafter2 animatediff2 MMDiT Autoregressive Diffusion U-Net DiT Commercial Figure 4: MDS visualization of generator similarity induced by cross-model detection perfor- mance.. Model similarity is based on pairwise detection accuracy. ModelWan21 Hun- yuan Kling SoraGen-3 Rep- Video Jimeng Luma MiraPika Open Sora STIV Caus Vid VCraf ter-2 ADiff -V2 AVG Wan2199.797.493.197.198.488.796.897.533.891.344.788.397.792.194.287.38 Hunyuan93.999.481.392.992.069.789.791.438.584.434.274.289.886.390.680.55 Kling85.982.098.582.284.466.376.784.826.265.935.269.192.271.076.873.15 Sora 90.384.975.699.593.778.799.296.135.287.134.170.798.088.390.481.44 Gen-392.684.981.793.099.683.894.495.323.788.150.675.196.580.484.581.60 RepVideo97.091.592.896.998.398.898.697.829.494.856.686.999.891.791.188.13 Jimeng*61.147.939.978.481.654.599.781.314.870.926.439.691.755.060.860.23 Luma90.884.380.295.794.279.198.498.641.489.447.772.599.388.092.083.44 Mira20.633.621.433.324.212.432.642.098.725.324.416.629.559.779.736.93 Pika77.264.359.181.087.168.291.182.621.997.243.965.680.383.972.171.69 Opensora V1.158.155.257.464.179.149.370.373.750.473.890.448.255.065.264.163.62 Apple-STIV87.374.877.574.488.778.673.374.929.778.036.496.367.691.186.274.31 CausVid28.721.819.639.634.538.350.744.87.316.07.416.799.618.428.831.47 VideoCrafter-2 76.867.854.279.774.360.481.877.664.178.828.970.477.099.294.672.39 AnimateDiff-V268.460.549.475.667.751.183.773.960.458.120.058.976.289.499.166.16 Table 10: Cross-Validation Results. Each cell in the table represents the average recall (%) of four detection models (NPR (Tan et al., 2024), TALL (Xu et al., 2023), X-CLIP-B/32 (Ni et al., 2022), F3Net (Qian et al., 2020)). The model is trained on generated videos of each subset and 5k real videos from MSR-VTT dataset. are then tested on all other generative subsets. The average cross-validation accuracy across the four detectors is reported in Table 10. Each element in the table represents the mean detection accuracy across the four models. Diagonal entries correspond to in-subset evaluations, where the detector is tested on the same generative model used for training. As shown in Figure 4, we interpret the matrix M as a distance metric between generative models and apply Multidimensional Scaling (MDS) to project their relationships into a 2D space. This visualization reveals clusters of architecturally sim- ilar models, such as AnimateDiff2(Xu et al., 2024) and VideoCrafterV2(Chen et al., 2024a), while autoregressive-based models, such as (Yin et al., 2025), appear more distant from the rest. This map- ping also informs a diverse training set selection of generative models, we could combine the cross validation accuracy and similarity to construct a high-quality and diverse dataset for data-efficient training. Better Generation, Better Detection. In our cross-validation experiment, we observed that de- tection models trained on higher-quality generated videos exhibit stronger detection performance. 21 Published as a conference paper at ICLR 2026 To validate this observation, we retrieved the overall VBench scores (Huang et al., 2023) for each generative model and conducted a correlation analysis between these scores and the average detec- tion accuracies reported in Table 10. The results are visualized in Fig.1 of our main paper. Since the cross-validation data is directly sampled from VBench’s evaluation set, the VBench scores pro- vide an accurate proxy for the generation quality of each subset. Across 14 models (excluding CausVid, which features a fundamentally different model structure and training paradigm), we com- pute a Pearson correlation coefficient of ρ = 0.86 between average detection accuracy and VBench scores, indicating a strong positive correlation. Furthermore, when restricting the analysis to the six DiT-based models, the correlation increases to ρ = 0.92. These results strongly support our hypoth- esis: among models with similar architectures, higher-quality generation leads to better supervision signals, enabling detection models to learn more effective forgery-discriminative features. CIMPLEMENTATION DETAILS This section outlines the configurations and hyper-parameters used for training our proposed method, as well as the baseline models. Our Method. For our detector and Moon-ViT, all experiments are conducted using PyTorch with Automatic Mixed Precision (AMP) in bfloat16 to enable Flash Attention optimization and accel- erate training. The visual backbone is initialized with Vision Transformer (ViT) weights from the officially released Qwen2.5-VL model. We explore multiple fine-tuning strategies with distinct hy- perparameter settings: (1) Full fine-tuning: We set the batch size to 4 and train for 5 epochs with a learning rate of 1e-5. (2) Linear Probing (LP) and Parameter-Efficient Fine-Tuning (PEFT): These approaches use a larger batch size of 32 and a learning rate of 1e-4. Training continues for up to 30 epochs, with early stopping based on validation loss (patience = 5 epochs) to prevent overfitting. Other Baseline Methods. To ensure fair comparison, all baseline models are trained under a uni- fied experimental setup. We used a consistent batch size of 32 and trained for a maximum of 30 epochs, also employing an early stopping strategy with a 5 epochs patience. The learning rate was adjusted based on the model architecture: for baselines utilizing a CLIP ViT backbone, such as X- CLIP and CLIP-based detectors, we set the learning rate to 1e-6; for all other models, a learning rate of 1e-5 was used. Data Pre-processing for Baseline Methods. A consistent data pre-processing pipeline is applied across all models during both training and testing. During training, each video is first sampled at a rate of 2 frames per second, from which 8 consecutive frames are extracted. If a video contains fewer than 8 frames, it is padded with blank frames to meet the required sequence length. Each frame is resized such that the shorter side is 224 pixels, followed by a random crop to a final resolution of 224×224. To enhance model robustness, we apply two forms of data augmentation: random horizontal flipping and random Gaussian noise. During testing, frames are sampled in the same manner as during training. After resizing the shorter side of each frame to 224 pixels, a center crop to 224×224 is applied instead of a random crop to ensure deterministic evaluation. DADDITIONAL RESULTS AND ABLATIONS Full Results on Genvideo-Val. As shown in Table 11, our proposed method achieves state-of-the- art performance across several key metrics. Notably, it attains an F1 score of 90.64 and an average precision (AP) of 96.13, surpassing all other leading methods—including DeMamba-CLIP, which was trained on the GenVideo dataset comprising 2.2 million samples. In contrast, our model was trained on only 140K samples, over ten times fewer, underscoring both the high quality of our train- ing data and the efficiency of our method in learning robust forgery-discriminative features at native resolution. In addition, our model achieves a balanced accuracy (bACC) of 95.38, significantly outperforming all competing methods. This result demonstrates not only high overall detection per- formance but also the model’s well-rounded and consistent capabilities across diverse forgery cases. Efficiency Comparison. As detailed in Table 12, we conduct a comprehensive efficiency analy- sis comparing our proposed Qwen2.5-VL ViT (Qwen2.5-ViT) with several strong baseline models. 22 Published as a conference paper at ICLR 2026 Model Training MetricSora Morph Gen2HotShotLavieShow-1 Moon Crafter ModelWild Avg. DataStudioValleyScopeScrape UNITE FaceForensics++, SAIL-VOS-3D Recall92.11100.094.6296.9398.1299.8698.69100.096.2989.8989.60 F1- AP88.57100.0100.090.1689.9198.3499.52100.098.9692.5692.76 DeMamba-CLIPGenVideo Recall95.71100.098.7069.1492.4393.29100.0100.083.5782.9491.58 F164.6396.1597.3978.0394.1492.7695.7298.0487.2387.8289.19 AP85.50100.099.5976.1596.7896.9999.97100.089.8089.7293.45 RINEProGANbACC-84.0089.1066.0096.7091.8085.7098.3076.60-74.10* DeMambaPyramidFlowbACC-83.8092.2062.0079.6072.6092.4087.5068.60-78.10* Corvi et al.PyramidFlowbACC-97.0098.8081.4095.5092.1098.4098.3097.10-94.30* Ours15model-140k Recall82.1497.1499.4989.0098.7992.2999.0599.0783.0071.6091.16 F165.2595.8498.3591.4898.0293.2996.5198.1688.0381.4590.64 AP82.4999.3699.9596.5599.7897.8899.8799.8994.5090.9896.13 bACC90.8798.3899.5594.3199.2095.9599.3399.3491.3185.6195.38 Table 11: Benchmarking Evaluation in terms of Recall, F1 score (F1), average precision (AP), and balance accuracy (bACC) on Genvideo-Val. The results of RINE and DeMamba are reported in Corvi et al. (2025). ModelResolution#ParamsFLOPSPeak GPU MemTraining Time / Epoch CLIP-L[224, 224]303.2M622.6G 21.5GB (bs=4) 129.3GB (bs=32) 9.5 A100 hours X-CLIP-L[224, 224]429.2M650.6G 21.5GB (bs=4) 129.3GB (bs=32) 10.5 A100 hours Effort[224, 224]0.2M/504.6M623.4G 17.3G (bs=4) 75.1GB(bs=32) 7.5 A100 hours Qwen2.5-ViT[224, 224]668.7M656G16.0GB(bs=4)2.3 A100 hours Qwen2.5-ViTdynamic [224p, 448p] 668.7M-37.9GB(bs=4)7 A100 hours Qwen2.5-ViT-LoRAdynamic [224p, 448p]2.6M/671.31M-27.4GB(bs=4)5.5 A100 hours Table 12: Efficiency comparison results on model parameters, FLOPS, GPU memory utilization and time consumed during training. For a standard input resolution of [224, 224], Qwen2.5-ViT exhibits remarkable training efficiency. Despite having more parameters (668.7M) than CLIP-L (303.2M), it achieves a 4.1× reduction in training time (2.3 vs. 9.5 A100 hours) and a 25% decrease in peak GPU memory usage (16.0GB vs. 21.5GB at a batch size of 4). These gains are primarily attributed to efficiency-oriented de- sign choices such as bfloat16 training and Flash Attention, which allow Qwen2.5-ViT to utilize computational resources more effectively. When adopting a dynamic resolution strategy, the train- ing overhead naturally increases, yet Qwen2.5-ViT remains faster and more memory-efficient than the baselines. Moreover, our parameter-efficient fine-tuning variant, Qwen2.5-ViT-LoRA, requires updating only 2.6M parameters. This substantially reduces resource demands compared to full dy- namic fine-tuning, lowering GPU memory from 37.9GB to 27.4GB and cutting training time from 7 to 5.5 hours. Overall, these results highlight that the superior efficiency of Qwen2.5-ViT stems from architectural optimizations, making the additional cost of higher dynamic resolutions acceptable in practice. EVISUALIZATION AND DISCUSSION Saliency Analysis. We examine the model’s attention responses to better understand its discrimi- native behavior, as illustrated in Figure 6. The results confirm that our native-resolution framework effectively captures two key types of features crucial for AIGC detection. (1) Low-level Artifacts: In billboard scenes, the model focuses on fine details such as distorted text rendering and unnatural edge transitions that are often lost during resolution downsampling. These high-frequency artifacts are indicative of generation errors and are critical for reliable detection. (2) High-level Semantics: In the fruit-cutting examples, the model attends to global inconsistencies, including object defor- mations and unrealistic lighting, suggesting it captures holistic content-level anomalies. This dual focus demonstrates that our approach leverages both spatial fidelity and semantic context, validating the design choice of preserving native resolution. 23 Published as a conference paper at ICLR 2026 Figure 5: Video Visualization from Magic Video Benchmark. From left to right, each column denotes videos from real sources, seaweed, seedance, and wan2.1. Figure 6: Saliency Analysis. Saliency maps of our model on AI-generated video samples. Figures 5 and 7 to 9 present a selection of video samples from our dataset, with Figures 3–5 offering detailed visualizations along with their corresponding generative prompts. As illustrated in these figures, the videos in Magic Videos benchmark exhibit high visual quality, characterized by aesthetic appeal, rich motion, and diverse themes and visual effects. 24 Published as a conference paper at ICLR 2026 By using carefully curated prompts to control the generative themes, we are able to evaluate a model’s detection performance without introducing content bias. This methodological design pro- motes a fairer and more reliable assessment, encouraging the detector to learn generalizable forgery artifacts rather than memorizing specific object- or scene-level patterns. Acknowledgment of LLM Usage. This manuscript has benefited from the assistance of a large language model, which was employed solely for grammar checking and language polishing. All scientific ideas, experimental designs, analyses, and conclusions are made by the authors. prompt:The video showcases the Alhambra in Granada, Spain, transitioning from warm golden sunset tones to deep violet hues as night falls. The palatial structures, set against the Sierra Nevada mountains and lined with cypress trees, shift from sunlit brilliance to dramatic nighttime illumination. A subtle zoom enhances the view, while the changing light casts a striking contrast between the fortress's golden glow and the darkening sky, creating a captivating visual transformation. realvideo generatedvideo Figure 7: Video Visualization from Magic Video Benchmark. 25 Published as a conference paper at ICLR 2026 prompt:A woman and a man engage in a friendly outdoor conversation amid wooden structures and greenery. The woman, wearing a purple headband and green tank top, sips her drink, signaling relaxation. Her expressions shift from savoring to engaging warmly, smiling and making eye contact. The man listens attentively, maintaining a steady demeanor. Both hold beverages, emphasizing the leisurely tone. Their uninterrupted dialogue features moments of humor and enjoyment in a serene setting. realvideo generatedvideo Figure 8: Video Visualization from Magic Video Benchmark. 26 Published as a conference paper at ICLR 2026 prompt:The video showcases billboards for Powerball ($470M) and Mega Millions ($999M) under a sunny sky, with a '3 News Now' banner highlighting a '$1 BILLION MEGA MILLIONS JACKPOT.' Vibrant designs and mentions of 'NEBRASKA POWERBALL POWERPLAY' add local context. A brief error misstates the Mega Millions jackpot as $9M before correcting it. The video ends with a wide shot of the billboards against a residential backdrop, emphasizing their public appeal. realvideo generatedvideo Figure 9: Video Visualization from Magic Video Benchmark. 27