Paper deep dive
The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics
Xiangbo Gao, Mingyang Wu, Siyuan Yang, Jiongze Yu, Pardis Taghavi, Fangzhou Lin, Zhengzhong Tu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 5:08:38 AM
Summary
The paper introduces 'Visual Chronometer', a predictor designed to measure the intrinsic Physical Frames Per Second (PhyFPS) of videos to address 'chronometric hallucination' in generative video models. By training on controlled temporal resampling, the model recovers the true temporal scale of motion, bypassing unreliable metadata. The authors establish two benchmarks, PhyFPS-Bench-Real and PhyFPS-Bench-Gen, to quantify temporal misalignment and instability in state-of-the-art video generators, demonstrating that PhyFPS-guided corrections improve perceived naturalness.
Entities (5)
Relation Signals (3)
Visual Chronometer â addresses â Chronometric Hallucination
confidence 95% ¡ we propose Visual Chronometer, a predictor designed to alleviate chronometric hallucination
PhyFPS-Bench-Gen â audits â Generative Video Models
confidence 95% ¡ PhyFPS-Bench-Gen, a benchmark designed to quantitatively audit the time-scale alignment of video generative models
Visual Chronometer â evaluatedby â PhyFPS-Bench-Real
confidence 90% ¡ We evaluate Visual Chronometer across multiple dimensions... we introduce PhyFPS-Bench-Real
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:While recent generative video models have achieved remarkable visual realism and are being explored as world models, true physical simulation requires mastering both space and time. Current models can produce visually smooth kinematics, yet they lack a reliable internal motion pulse to ground these motions in a consistent, real-world time scale. This temporal ambiguity stems from the common practice of indiscriminately training on videos with vastly different real-world speeds, forcing them into standardized frame rates. This leads to what we term chronometric hallucination: generated sequences exhibit ambiguous, unstable, and uncontrollable physical motion speeds. To address this, we propose Visual Chronometer, a predictor that recovers the Physical Frames Per Second (PhyFPS) directly from the visual dynamics of an input video. Trained via controlled temporal resampling, our method estimates the true temporal scale implied by the motion itself, bypassing unreliable metadata. To systematically quantify this issue, we establish two benchmarks, PhyFPS-Bench-Real and PhyFPS-Bench-Gen. Our evaluations reveal a harsh reality: state-of-the-art video generators suffer from severe PhyFPS misalignment and temporal instability. Finally, we demonstrate that applying PhyFPS corrections significantly improves the human-perceived naturalness of AI-generated videos. Our project page is this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2603.14375v1
- Canonical: https://arxiv.org/abs/2603.14375v1
Trouble viewing inline? Open PDF directly â
Full Text
49,765 characters extracted from source content.
Expand or collapse full text
The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics Xiangbo Gao, Mingyang Wu, Siyuan Yang, Jiongze Yu, Pardis Taghavi, Fangzhou Lin, Zhengzhong Tu Texas A&M University Abstract. While recent generative video models have achieved remarkable visual realism and are being explored as world models, true physical simulation requires mastering both space and time. Current models can produce visually smooth kinematics, yet they lack a reliable internal motion pulse to ground these motions in a consistent, real-world time scale. This temporal ambiguity stems from the common practice of indiscriminately training on videos with vastly different real-world speeds, forcing them into standardized frame rates. This leads to what we term chronometric hallucination: generated sequences exhibit ambiguous, unstable, and uncontrollable physical motion speeds. To address this, we propose Visual Chronometer, a predictor that recovers the Physical Frames Per Second (PhyFPS) directly from the visual dynamics of an input video. Trained via controlled temporal resampling, our method estimates the true temporal scale implied by the motion itself, bypassing unreliable metadata. To systematically quantify this issue, we establish two benchmarks,PhyFPS-Bench-Real andPhyFPS-Bench-Gen. Our evaluations reveal a harsh reality: state-of-the-art video generators suffer from severe PhyFPS misalignment and temporal instability. Finally, we demonstrate that applying PhyFPS corrections significantly improves the human-perceived naturalness of AI-generated videos. âNot only do we measure the movement by the time, but also the time by the movement, because they define each other.â â Aristotle, Physics Project Homepage: https://xiangbogaobarry.github.io/Visual_Chronometer/ Date: March 17, 2026 Contact: Xiangbo Gao (xiangbogaobarry@gmail.com), Zhengzhong Tu (tzz@tamu.edu) 1 Introduction While modern generative video models excel at spatial realismâproducing photorealistic textures, complex geometry, and coherent layouts [1,2,3,4,5]âan increasing number aspire to go further and act as physical world models [6,7]. However, faithfully simulating the physical world requires an intricate mastery of both space and time; physical motion is governed by a strict relationship between spatial displacement and elapsed time, yet todayâs video generation pipelines often lack a stable pulse of motion to track this. Consequently, while modern generators can produce visually fluid kinematics, these motions are rarely grounded in a consistent, real-world time scale. Much of this temporal ambiguity stems from the agnostic treatment of time during the training of modern video models [8]. Internet-scale video datasets are mixtures of varying capture and editing regimes, encompassing standard-rate footage, extreme slow-motion, and accelerated time-lapses. During training, models are typically blind to these inherent physical speeds; a time-lapse and a slow-motion video might be fed into the network identically. This lack of time-scale awareness severs the correspondence between a discrete frame step and the real-world time elapsed. As a result, models learn to generate plausible frame-to-frame transitions, but the underlying physical speed of the generated motion becomes ambiguous, unstable, and impossible to explicitly control. We refer to this prevalent failure mode as Chronometric Hallucination (see Figure 1). Aristotle once observed that ânot only do we measure the movement by the time, but also the time by the movement, because they define each other.â Operationalizing this ancient principle, we introduce Visual Chronometer, a predictor designed to alleviate chronometric hallucination by recovering this intrinsic motion arXiv:2603.14375v1 [cs.CV] 15 Mar 2026 A hummingbird hawk-moth hovering and darting between bright pink blossoms. A person stands on the bed and then lies down. Sora-2 Grok-Imgine-T2V ďźa) ďźb) Figure 1 Visualization of Chronometric Hallucination. Current video generators sometimes fail to ground their outputs in a consistent physical time scale, even when no speed-manipulating keywords (e.g., âslow motionâ) are prompted. (a) A hummingbird hawk-moth is rendered in extreme slow-motion, despite its naturally high wing-beat frequency. (b) A person falls onto a bed at a velocity significantly slower than standard gravity. These instances illustrate Chronometric Hallucination: a prevalent failure mode where generated motions exhibit an ambiguous, unstable, and uncontrollable physical time scale. pulse, formalized as Physical Frames Per Second (PhyFPS), directly from visual dynamics. We distinguish the inherent PhyFPS from the nominal metadata (meta FPS) by defining PhyFPS as the true frame rate that aligns with the real-world passage of time. Through controlled temporal resampling, we supervise the model to learn these motion-grounded dynamics, bypassing the often unreliable metadata. We evaluate Visual Chronometer across multiple dimensions. First, to validate the accuracy of our method, we introducePhyFPS-Bench-Real, comprising real-world videos where the true PhyFPS often diverges from the meta FPS due to complex speed variations. Second, we establishPhyFPS-Bench-Gento systematically audit state-of-the-art video generators along three complementary axes: (i) the alignment between meta FPS and actual PhyFPS, (i) intra-video stability (the consistency of PhyFPS across sliding windows within a single clip), and (i) inter-video stability across different outputs from the same model configuration. Our extensive measurements reveal a harsh reality: even strong generators exhibit substantial PhyFPS misalignment, alongside significant intra- and inter-video temporal jitter. Without a grounded physical time scale, these models fail to provide the reliable simulation necessary for true world modeling. Furthermore, we demonstrate that applying PhyFPS-guided post-corrections to generated videos substantially improves human-perceived naturalness, as validated by our user study. Finally, we evaluate strong Vision-Language Models (VLMs) as potential PhyFPS judges, finding them vastly unreliable for this specialized task, thereby underscoring the necessity of our dedicated Visual Chronometer. Our contributions are summarized as follows: â˘We identify and define the phenomenon of chronometric hallucination in modern video generators, and formalize Physical Frames Per Second (PhyFPS) as a temporal scale distinct from nominal meta FPS. ⢠We propose Visual Chronometer, a robust predictor that recovers PhyFPS directly from raw frames by learning motion-grounded dynamics through controlled temporal resampling. â˘We introducePhyFPS-Bench-Gento audit state-of-the-art generators, revealing severe time-scale misalign- ment in modern video generators. We further show that PhyFPS-guided post-correction significantly enhances human-perceived temporal naturalness. â˘ThroughPhyFPS-Bench-Real, we demonstrate our modelâs precision in predicting Physical FPS. Our analysis also reveals that the state-of-the-art VLMs are unreliable temporal judges, underscoring the necessity of a dedicated Visual Chronometer. 2 Related Works 2.1 Video Generation and the Quest for World Models Modern video generative models, spanning large-scale diffusion and autoregressive architectures, have achieved unprecedented perceptual quality and semantic coherence [9,1,2,10,11,12,13,14]. To capture dynamics, these systems employ sophisticated temporal modeling mechanisms, such as 3D spatiotemporal operators [15, 16], causal attention blocks [6,9], and temporal latent spaces [17,18]. As these architectures scale, they are 2 increasingly framed as âworld modelsâ capable of simulating physical environments [19,20,21,22,23]. However, while prior works focus heavily on optimizing frame-to-frame kinematic smoothness and spatial layout, the actual physical time scale of the depicted motion is rarely encoded or supervised [24,25]; instead, models rely entirely on the nominal frame rate (meta FPS) provided by the dataset container. Because these advanced generative mechanisms do not explicitly ground their temporal learning in real-world physics, they remain highly vulnerable to chronometric hallucinationâproducing motions that look perceptually smooth but lack a consistent physical speed. We argue that one cannot fix a physical flaw without first being able to measure it. Thus, we complement these generative advancements by developing the first dedicated tool to audit this structural blind spot. By explicitly defining and predicting the intrinsic Physical FPS (PhyFPS), we provide the necessary metric and benchmark to evaluate time-scale calibration in world models. 2.2 Visual Perception of Time and Dynamics Our methodology draws inspiration from a long-standing line of computer vision research aimed at understand- ing time and speed from visual cues. Early efforts in this domain focused on domain-specific heuristics, such as detecting slow-motion replays in sports broadcasts [26,27,28]. More recently, self-supervised approaches like SpeedNet [29] demonstrated that neural networks can discriminate between normal-rate and artificially sped-up clips. In a parallel vein, research on the âarrow of timeâ explores whether models can recognize the forward or backward directionality of video playback [30,31]. Furthermore, semantic hyperlapse and time-remapping techniques actively manipulate temporal sampling to summarize videos [32,33,34,35,36,37], proving that visual dynamics naturally dictate the perceived flow of time. However, these existing perception models typically frame time as a binary classification problem (e.g., faster vs. slower, forward vs. backward). They do not aim to recover a high-precision physical metric. In contrast, Visual Chronometer frames time-scale perception as an absolute continuous regression problem, directly predicting PhyFPS from frame sequences to audit generative models without relying on corrupted metadata. 2.3 Benchmarking Temporal and Physical Fidelity Evaluating video generation has traditionally been dominated by perceptual quality and semantic fidelity metrics. Standard protocols rely on frame-level similarity (PSNR [38], SSIM [39], LPIPS [40]), no-reference perceptual quality predictors for user-generated and variable-frame-rate videos such as RAPIQUE [41] and FAVER [42], and distribution-level feature matching, most notably the FrĂŠchet Video Distance (FVD) [43]. Recognizing the limitations of monolithic metrics, recent comprehensive suites like VBench [44,45,46] and WorldScore [47] have introduced multi-dimensional evaluations, including physics-adjacent axes such as temporal consistency and action alignment. Nevertheless, these benchmarks primarily evaluate whether the motion âlooks naturalâ rather than measuring the exact temporal speed governing the scene. Time-scale fidelityâspecifically, whether a video strictly adheres to a stable physical frame rate throughout its durationâ remains entirely unmeasured. Our introduced benchmarks,PhyFPS-Bench-RealandPhyFPS-Bench-Gen, fill this critical void. By shifting the evaluation paradigm from perceptual smoothness to chronometric measurement, we provide the first quantitative audit of intra-video and inter-video time-scale stability in generative world models. 3 Data Preparation 3.1 Data Collection To train Visual Chronometer to accurately predict the Physical Frames Per Second (PhyFPS), we require a training dataset with verified, ground-truth temporal labels. A model trained on data suffering from chronometric hallucination inherently cannot serve as a reliable temporal measurement tool. Therefore, we curate a dataset exclusively from video sources where the nominal metadata frame rate perfectly aligns with the real-world physical sampling rate (i.e., meta FPS = PhyFPS), strictly excluding videos with ambiguous post-hoc time-scale editing. We aggregate our high-fidelity source data from the following categories: â˘High-Frame-Rate Academic Datasets: We utilize high-speed benchmarks, including Adobe240 [48] and BVI-VFI [49] (up to 120 Hz), typically used for precise temporal analysis and frame interpolation. 3 Fast ShutterSlow ShutterRolling Shutter Figure 2 Physics-Grounded Temporal Augmentation. We synthesize diverse low-rate videos from high-frequency source data (240 FPS) to simulate real-world camera mechanics: Sharp Capture, Motion Blur, and Rolling Shutter. Data Distribution Figure 3 Dataset distribution across 18 target Physical Frame Rates. â˘Raw Broadcast Sequences: Uncompressed 4K YUV footage from UVG [50] (50/120 FPS) is included; its raw pipeline minimizes the risk of hidden temporal remapping. ⢠Sensor-Synchronized Autonomous Data: Datasets from NVIDIA and Honda [51] provide cross-sensor alignment (Camera/LiDAR/IMU), where strict synchronization guarantees physical time-scale integrity. â˘Physics-Grounded Human Motion: Human-centric sequences [52] are incorporated to leverage motion captured specifically for biomechanical and dynamic realism. ⢠Verified In-House Data: We supplement the public datasets with an internal collection captured under strictly controlled settings with verified frame-rate metadata. 3.2 Data Preprocessing and Augmentation To force the model to learn intrinsic visual dynamics rather than relying on semantic content priors, we expand our training distribution by synthetically generating a diverse array of PhyFPS variants from the source videos. We first temporally upsample all source videos to a high-frequency base rate of 240 FPS using a state-of-the-art frame interpolation model (RIFE) [53]. Let this high-rate video beI H at a frequencyF H = 240 FPS. For a target lower frame rateF L , we define the downsampling ratio asN=F H /F L . We then synthesize low-rate videos (I L ) using three distinct strategies (illustrated in Fig. 2), each designed to model specific real-world camera mechanics: (1) Sharp Capture (Fast Shutter): To simulate cameras operating with a very fast shutter speed (which minimizes motion blur), we uniformly subsample the high-rate sequence by settingI L k =I H âkNâ , whereI L k denotes the k-th frame of the synthesized low-rate video andâkNâis the corresponding discrete frame index in the high-rate source. This isolates pure spatial displacement over time, preserving sharp object boundaries but often resulting in the naturally aliased motion (stutter) typically seen in sports or action footage. (2) Motion Blur (Variable Exposure): Real-world cameras integrate light over an exposure window, resulting in motion blur that provides strong visual cues about object velocity. To mimic this exposure integration, we synthesize each low-rate frame by averaging a temporal window of high-rate frames:I L k = 1 M P Mâ1 i=0 I H âkNâ+i , whereMis the exposure window length. We simulate long, medium, and short effective exposures by setting M âN, N/2, N/4. (3) Synthetic Rolling Shutter Fast-moving objects captured by modern CMOS sensors frequently exhibit rolling shutter distortions [54] because sensor rows or columns are read sequentially rather than instantaneously. We simulate this intra-frame temporal distortion by partitioning the target frameâs spatial dimension (e.g., width W) into progressive bands. A pixel at columnxis sampled from the high-rate sequence at a progressively shifted time index:âkNâ+âM ¡ x/Wâ. By varying the readout durationM âN, N/2, N/4, we ensure the predictor is robust to these common spatiotemporal artifacts. Final Dataset Composition. As summarized in Fig. 3, alongside these synthetically augmented variants, we retain the original source videos at their native capture rates to preserve raw sensor statistics. e generate training data across 18 Physical Frame Rates, yielding a comprehensive dataset of 465,535 video clips, 4 uniformly standardized to a length of 128 frames to ensure balanced representation across different time scales. 4 Visual Chronometer 4.1 Model Architecture Backbone and Regression Head. We adopt VideoVAE+ [18] as the foundational video encoder to extract compact spatiotemporal latent representations. Given an input clip ofTframesV =I t T t=1 , the backbone produces a sequence of latent tokensZ = Enc(V). Instead of relying on conventional spatial pooling, we attach a lightweight, attention-based prediction head to aggregate temporal features into a clip-level representation. Specifically, we project the latent tokens into a hidden dimension and introduce a learnable query embedding that cross-attends to the token sequence. This query-based pooling mechanism effectively decouples the regression head from the input frame count, enabling Visual Chronometer to process videos of arbitrary lengths. Finally, a Multi-Layer Perceptron (MLP) maps the aggregated feature vector to a single scalarËsâR, which represents the predicted logarithmic frame rate, log(PhyFPS). We predict the logarithmic value rather than the absolute frequency to stabilize optimization across an exponentially wide range of time scales and to penalize relative, rather than absolute, errors. 4.2 Training Objective Let the ground-truth PhyFPS bey, with its log-space target defined ass=log y. The model outputs the prediction Ës = log Ëy. We optimize the model using a Mean Squared Error (MSE) in the logarithmic space: L log = 1 n n X i=1 (log y i â Ës i ) 2 ,(1) wherenis the batch size. Because the target PhyFPS values in our dataset are strictly positive (y i âĽ2), the logarithmic transformation is intrinsically well-defined. Therefore, we deliberately omit the standard offset term (+1) typically found in traditional Mean Squared Logarithmic Error (MSLE) formulations, allowing the loss to strictly reflect the true proportional scaling of time. 4.3 Model Training Details To train the Visual Chronometer, we extract clips from the dataset using a sliding window. During training, clips are sampled with a maximum temporal footprint ofT= 32 frames. To ensure robust performance across different deployment scenarios, we train two variants of the model targeting different operational regimes. The VC-Wide model is trained to predict across 18 distinct frame rates spanning from extreme slow-motion to high-speed capture:PhyFPSâ2,5,10,12,15,18,20,24,25,30,35,40,45,50,60,90,120,240. The VC-Common model focuses specifically on the most prevalent consumer and web video formats, narrowing the output space to PhyFPSâ12, 15, 18, 20, 24, 25, 30, 35, 40, 45, 50, 60. Both models are trained end-to-end, fine-tuning the VideoVAE+ backbone jointly with the attention-based prediction head. Optimization is performed using the Adam optimizer with a learning rate of 1Ă10 â5 for 125,000 iterations. We execute the training on a single computing node equipped with four NVIDIA RTX A6000 GPUs, utilizing a global batch size of 32. 5 Experiments In this section, we conduct three sets of experiments to validate the Visual Chronometer and demonstrate its utility in addressing chronometric hallucination, as well as enabling physics-grounded data preprocessing and video post-processing. First, we introducePhyFPS-Bench-Gento audit existing open- and closed-source video generative models by measuring their Meta-vs-PhyFPS alignment and temporal stability. Second, we build PhyFPS-Bench-Realto evaluate the prediction accuracy of our model against reliable ground-truth labels. Third, we compare our specialized predictor against strong Vision-Language Models (VLMs), demonstrating that general-purpose foundation models are not yet capable of reliable PhyFPS prediction. 5 5.1 Auditing Generative World Models PhyFPS-Bench-Gen. We introducePhyFPS-Bench-Gen, a benchmark designed to quantitatively audit the time- scale alignment of video generative models using our Visual Chronometer. We evaluate a diverse spectrum of leading generators. For open-source models, we assess the Wan series [1] (Wan2.1-1.3B, Wan2.1-14B, Wan2.2- 5B, Wan2.2-14B), the LTX series [10,2] (LTX-Video, LTX-2), the CogVideoX series [11] (CogVideoX-2B, CogVideoX-5B), HunyuanVideo [13], and the autoregressive model InfinityStar [55]. For closed-source models, we evaluate Veo-3.1-Fast [56], Sora-2 [57], Grok-Imagine-T2V [58], Kling-o3 [59], Seedance-1.0-Lite [60], and Seedance-1.5-Pro [60]. Benchmark Prompts. To ensure robust evaluation, we design 100 text-to-video prompts covering diverse content and motion patterns, strictly avoiding explicit speed-manipulation keywords (e.g.,slow motion, time-lapse,speed up). To guarantee that PhyFPS is observable, every prompt mandates at least one clearly dynamic instance, excluding purely static scenes. Prompt diversity is balanced across five axes: (i) primary entity (human, animal, vehicle, and nature), (i) motion type (articulated, rigid-body, fluid, and multi-agent), (i) camera behavior (static, pan, and tracking), (iv) environmental effects (rain, fire, and wind), and (v) scene context (indoor, urban, and nature). All models operate under default settings, extracting the nominal saved FPS (F meta ) from official documentation or output metadata. PhyFPS Estimation and Metrics. For all audits on generated videos, we employ the VC-Common predictor. For each videov â1, . . . , V, we extract C v overlapping clips ofT =32 frames with strides=4. Let Ë f v,c denote the predicted PhyFPS for clipc. The video-level PhyFPS ( Ě f v ) and the overall model-level PhyFPS ( Ë F) are computed as: Ě f v = 1 C v C v X c=1 Ë f v,c , Ë F = 1 V V X v=1 Ě f v .(2) We evaluate each generator along three critical dimensions. (1) Meta-vs-PhyFPS Alignment measures how well the nominal container rateF meta matches the predicted intrinsic speed. We report both the Avg. Error (FPS) and the Pct. Error (%): Avg. Error = 1 V V X v=1 Ě f v â F meta ,Pct. Error = 100 V V X v=1 Ě f v â F meta F meta .(3) (2) Inter-video Consistency and (3) Intra-video Consistency evaluate temporal stability across different prompts and within a single continuous video, respectively. Both utilize the coefficient of variation (CV): Inter CV = Std Ě f v V v=1 Mean Ě f v V v=1 , Intra CV = 1 V V X v=1 Std Ë f v,c C v c=1 Mean Ë f v,c C v c=1 .(4) Audit Results. Table 1 details the results of thePhyFPS-Bench-Genaudit. We observe a pervasive Meta- vs-PhyFPS mismatch across the majority of generators; despite outputs being stored at a fixed nominal meta FPS, the intrinsic visual speeds vary wildly. Notably, the Wan series models exhibit a relatively high adherence to their suggested meta FPS, achieving comparatively low average and percentage errors. To evaluate temporal consistency, we measure Inter CV and Intra CV, which represent the fluctuation of PhyFPS across different generated videos and the stability of PhyFPS across different time segments within a single video, respectively. The LTX-Video and LTX-2 models demonstrate strong performance on these stability metrics. This suggests their temporal representations are internally consistent, and the high absolute errors may primarily stem from inaccurate meta FPS metadata rather than structural chronometric hallucination. For instance, LTX-Videoâs outputs might simply need their meta FPS adjusted from 24 to around 46.5 to achieve high time-scale fidelity. Overall, closed-source models slightly outperform open-source counterparts in terms of absolute accuracy, maintaining average errors below 14 FPS and percentage errors under 60%. This indicates that commercial models may employ more carefully designed strategies for selecting meta FPS. However, despite these advantages in global alignment, the Intra and Inter CV scores of closed-source models are not significantly 6 Table 1 Quantitative Audit of Generative Models.PhyFPS-Bench-Genresults evaluating time-scale fidelity. Blue and red shaded cells indicate the best and second-best performance within each group (open vs. closed source). ModelMeta FPSPhyFPSAvg. Error âPct. Error(%)âIntra CV âInter CV â Open-sourced Models CogVideoX-2B2433.6412.46520.110.46 CogVideoX-5B2438.2617.96750.120.52 HunyuanVideo2435.8913.82580.120.36 Wan2.1-T2V-1.3B2426.28 7.54310.110.38 Wan2.1-T2V-14B2432.3710.87450.140.36 Wan2.2-T2V-A14B2431.5210.74450.120.38 Wan2.2-TI2V-5B2432.8111.63480.150.38 InfinityStar (5s)1634.4118.461150.110.38 InfinityStar (10s)1636.1520.191260.160.36 LTX-Video2446.5223.6799 0.100.33 LTX-22539.7715.70630.130.34 Closed-sourced Models Seedance-1.0-Lite2428.608.31350.150.37 Seedance-1.5-Pro2433.6910.67440.160.25 Sora-23036.218.40280.130.29 Grok-Imagine-T2V2436.9713.97580.160.28 Kling-o32430.049.1038 0.150.34 Veo-3.1-Fast2435.8313.62570.170.33 better than those of open-source models. This reveals that even heavily optimized, industry-scale generators still struggle with PhyFPS stability. It appears researchers predominantly prioritize visual fidelity and kinematic smoothness, inadvertently neglecting strict physical time-scale adherence. This lack of reliable temporal grounding poses a significant challenge for leveraging current video generative models as accurate world models. Finally, we observe a consistent trend where the predicted PhyFPS is generally higher than the assigned Meta FPS across almost all videos. According to the Visual Chronometer, most generated videos should be played back at a higher meta FPS, or directly at their intrinsic PhyFPS. This finding aligns with the widely recognized phenomenon that current generative models tend to produce âslow but smoothâ videos [8]. User Study: Perceptual Validation via Video Post-processing. To demonstrate the practical utility of our method and confirm that mathematical PhyFPS accuracy translates to improved human perception, we conduct a user study treating the predictor as a post-processing tool. Using VC-Common, we predict the clip-level PhyFPS for generated videos. We then present users with three variants of the same sequence. The first variant is the Original video, representing the untouched output directly from the generative model. The second variant, Pred, serves as a globally corrected version; here, we uniformly re-time the entire video to match its average predicted PhyFPS. The third variant, Pred Dyn, applies a dynamic local correction, where each distinct temporal segment within the video is independently re-timed based on its specific, clip-level PhyFPS prediction. Figure 4 Human Perceptual Preference on Temporal Naturalness. Bradley-Terry scores comparing the original generated videos against our post-processed variants. Both the global average correction (Pred) and dynamic local correction (Pred Dyn) are strongly preferred over the hallucinated original outputs, with 90% confidence intervals indicating statistical significance. We collected 1,490 pairwise comparisons from over 15 participants. Utilizing the Bradleyâ Terry model [61], we estimated the relative preference strength for each variant, comput- ing 90% confidence intervals via bootstrap- ping (Figure 4). The results reveal that both post-processed variants significantly out- perform the hallucinated original outputs (19.0%). Interestingly, the global correction (Pred, 44.2%) is preferred over the dynamic lo- cal correction (Pred Dyn, 36.9%). We hypoth- esize that while dynamic correction perfectly aligns local clips to their intrinsic PhyFPS, varying the playback frame rate within a sin- gle short sequence may introduce perceptual inconsistencies or jitter. Conversely, applying a constant, averaged Physical Frame Rate (Pred) remains visually smoother and more natural to human observers. Ultimately, these findings definitively highlight the value of physics-grounded post-processing. 7 5.2 Validating the Visual Chronometer To establish the reliability of our measurement tool, we evaluate its prediction accuracy on thePhyFPS-Bench-Real test set (comprising 4,000 verified clips partitioned from our dataset). Crucially, to ensure that the Visual Chronometer learns intrinsic physical time scales rather than overfitting to dataset-specific biases, we enforce a strict cross-source split; the training, validation, and test sets are derived from entirely disjoint video sources. Given the ground-truth PhyFPSy i and predicted PhyFPSËy i acrossntest samples, we report the Mean Absolute Error (MAE) and Mean Absolute Percentage Error (MAPE) to capture both absolute deviations and proportional accuracy: MAE = 1 n n X i=1 |y i â Ëy i | ,MAPE = 100 n n X i=1 |y i â Ëy i | y i .(5) Given the rapid advancements in Vision-Language Models (VLMs), it is tempting to deploy them as out-of-the- box evaluators for physical scene dynamics. To rigorously test this hypothesis, we establish a comprehensive baseline using state-of-the-art VLMs, including Gemini-3.1-Pro [62], Gemini-3-Flash [63], Seed-1.6 and Seed-1.6-Flash [60], as well as Qwen3.5+ and Qwen3.5-397B [64]. Table 2 Predictor Accuracy & VLM Baseline Compar- ison. Evaluating the Visual Chronometer (Ours) against state-of-the-art Vision-Language Models on PhyFPS-Bench-Real. The average ground-truth PhyFPS across the test set is 38.81. Blue and red shaded cells indicate the best and second-best performance. ModelAvg Pred MAE â MAPE(%)â Ours VC-Common39.203.469 VC-Wide45.48 7.7621 Video-based VLM Gemini-3.1-Pro 31.0021.6743 Gemini-3-Flash 26.6023.4047 Seed-1.629.6020.4041 Seed-1.6-Flash30.0020.0040 Qwen3.5+4.4645.5491 Qwen3.5-397B25.6024.4049 Image-based VLM Gemini-3.1-Pro 5.1544.8590 Gemini-3-Flash 1.7748.2396 Seed-1.66.3543.6587 Seed-1.6-Flash30.0020.0040 Qwen3.5+3.4846.5293 Qwen3.5-397B22.0327.9756 We evaluate these VLMs under two input paradigms (prompt details in the Appendix). First, we use a Video-based approach. Because modern VLMs typically subsample frames to manage context length, this inherent preprocessing disrupts temporal spacing, predictably degrading frame rate perception. To bypass this architectural bottleneck, we introduce an Image-based paradigm, unrolling the video into 128 discrete images fed sequentially to preserve the absolute frame count and temporal order. The results (Table 2) show that our Visual Chronome- ters (VC-Common and VC-Wide) achieve exceptionally low MAE and MAPE, with the narrower-range VC- Common predictably yielding the tightest margins. Qualitatively, Figure 5 confirms our modelâs ability to continuously and accurately track physical time scales across varying base rates. Conversely, all tested VLMs fail catastrophically at physical estimation for both video inputs and unrolled image sequences. Many suffer from severe mode collapse. For example, Seed-1.6-Flash degenerates to predicting exactly 30 FPS for all inputs regardless of the actual dynamics. These findings demonstrate that general-purpose foundation models lack a grounded internal motion pulse, reinforcing the necessity of our specialized architecture. 5.3 Ablation Studies To validate our core design choices, we conduct ablation studies on the VC-Common model, evaluating the impact of physics-grounded data augmentations and inference temporal context length. Impact of Temporal Augmentations. To verify the necessity of our physics-grounded augmentations (Fast Shutter, Motion Blur, and Synthetic Rolling Shutter), we train a naive baseline using only uniform temporal subsampling. Evaluated on the in-the-wild conditions ofPhyFPS-Bench-Real(Table 3), the baseline degrades significantly, with MAE increasing from 3.46 to 5.12. Without simulating exposure integration or sequential sensor readout during training, the naive model overfits to idealized spatial displacements and fails to disentangle physical speed from realistic motion artifacts. This confirms our augmentations are critical for learning robust, intrinsic visual dynamics. 8 PhyFPS: 60 PhyFPS: 24 PhyFPS: 12 Predicted PhyFPS Figure 5 Continuous PhyFPS Prediction on Real Dynamics. Qualitative results from our Visual Chronometer evaluating a single dynamic action (soccer ball juggling) captured at three distinct physical frame rates (60, 24, and 12 PhyFPS). The model not only accurately recovers the absolute time scale directly from visual cues but also maintains remarkable temporal stability across the entire sequence. Table 3 Ablation Study on Temporal Data Augmentations. Evaluated onPhyFPS-Bench-Realusing the VC-Common configuration. Augmentation StrategyMotion BlurRolling ShutterMAEâMAPE (%)â Naive Baselineâ5.1213 + Motion Blurââ4.8711 VC-Commonâ3.469 Impact of Temporal Context Length. Measuring physical speed computationally requires sufficient kinematic history. We evaluate the robustness of VC-Common across varying inference window lengths (patch sizes) T â8, 16, 32, 64, 128. As illustrated in Figure 6, the base model (trained on max 32 frames) expectedly struggles with ultra-short contexts (T= 8) due to insufficient visual evidence, optimizing atT= 32 (MAE= 3.46). Notably, it demonstrates strong length extrapolation, maintaining competitive accuracy atT= 64 and 128. Post-training the model on a maximum length of 128 frames further improves performance atT= 64 without degrading short-patch accuracy. However, a critical bottleneck emerges: increasing the inference patch size toT= 128 fails to outperform T= 64. This reveals an inherent trade-off in temporal modeling. While small patches lack sufficient receptive fields, extremely large patches (e.g.,T= 128, spanning the entire benchmark video) restrict evaluation to a single global inference pass. This loses the variance-reduction benefits of sliding-window ensembling and strips the model of its ability to capture fine-grained PhyFPS fluctuations within a single shot. Consequently, a mid-range patch size (T= 32 to 64) optimally balances kinematic context with local temporal granularity. 6 Discussion: Implications and Future Directions In this section, we contextualize our findings and explore future directions for temporal modeling in generative video through a question-and-answer format. 9 Figure 6 Ablation on Inference Context Length (T). Evaluating the VC-Common model across different inference patch sizes onPhyFPS-Bench-Real. We compare the base model (trained on max 32 frames) with a post-trained variant (max 128 frames) to analyze the trade-off between temporal receptive field and sliding-window granularity. Q1: IsstrictalignmentbetweenPhyFPSandmetaFPSalwaysdesirable? Inotherwords, ischronometrichallucination inherently problematic, given that intentional speed manipulation is a core creative tool in filmmaking? A: Dynamic retimingâsuch as deliberate slow-motion or time-lapseâis undeniably a vital creative tool. We do not argue that every generated video must strictly adhere to a 1Ăphysical time scale (i.e.,PhyFPS= meta FPS). Rather, the fundamental issue with chronometric hallucination lies in the absence of controllability. Currently, models hallucinate time scales arbitrarily; a user prompting for a âperson walkingâ might implicitly receive a sequence operating at 0.5Ăor 2Ăphysical speed without any explicit instruction. While variable speeds are essential for specific creative scenarios, the ability to stably generate a grounded, default 1Ăspeed is a prerequisite for true controllability. If generative video models are to evolve into reliable world models, they must possess a stable internal pulse. Only by first mastering baseline physical reality can a model faithfully execute deliberate NĂ speed manipulations upon request. Q2: How can future video generation pipelines resolve chronometric hallucination? A: To resolve this issue, future pipelines should treat time as an active, controllable condition. First, at the data curation level, training datasets should be rigorously relabeled with their true intrinsic PhyFPS. Our Visual Chronometer can serve as an automated, large-scale annotator to explicitly filter or condition the input distribution. Second, at the architectural level, models require temporal conditioning mechanisms that force the network to explicitly comprehend and disentangle the true pulse of varying physical frame rates during training. Finally, from an optimization standpoint, the Visual Chronometer has the potential to act as a specialized reward model. By providing direct, physics-grounded supervision signals during preference alignment (e.g., via RLHF or DPO), it can guide generative models to strictly adhere to desired temporal dynamics and structurally eliminate chronometric hallucination. 7 Conclusion In this work, we identify and formalize the phenomenon of chronometric hallucination in modern video generative models, where a reliance on arbitrary metadata containers leads to ambiguous and uncontrollable physical speeds. To address this issue, we propose the Visual Chronometer, a robust predictor trained via physics-grounded temporal resampling that accurately recovers the intrinsic Physical Frames Per Second (PhyFPS) directly from visual dynamics. Through our comprehensive benchmarks,PhyFPS-Bench-Genand PhyFPS-Bench-Real, we reveal a stark reality: state-of-the-art generators and vision-language models currently struggle to maintain a consistent internal pulse of motion. Nevertheless, by demonstrating that PhyFPS-guided dynamic retiming significantly improves the human-perceived temporal naturalness of AI-generated videos, we offer an immediate, practical mitigation. Ultimately, we hope this work inspires future generative world models to transition from passive metadata reliance to active, physics-grounded temporal conditioning. 10 References [1]T. Wan, A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang et al., âWan: Open and advanced large-scale video generative models,â arXiv preprint arXiv:2503.20314, 2025. [2] Y. HaCohen, B. Brazowski, N. Chiprut, Y. Bitterman, A. Kvochko, A. Berkowitz, D. Shalem, D. Lifschitz, D. Moshe, E. Porat et al., âLtx-2: Efficient joint audio-visual foundation model,â arXiv preprint arXiv:2601.03233, 2026. [3]Z. Jiang, Z. Han, C. Mao, J. Zhang, Y. Pan, and Y. Liu, âVace: All-in-one video creation and editing,â arXiv preprint arXiv:2503.07598, 2025. [4] R. Burgert, C. Herrmann, F. Cole, M. S. Ryoo, N. Wadhwa, A. Voynov, and N. Ruiz, âMotionv2v: Editing motion in a video,â arXiv preprint arXiv:2511.20640, 2025. [5] X. Gao, R. Li, X. Chen, Y. Wu, S. Feng, Q. Yin, and Z. Tu, âPisco: Precise video instance insertion with sparse control,â arXiv preprint arXiv:2602.08277, 2026. [6]A. Ali, J. Bai, M. Bala, Y. Balaji, A. Blakeman, T. Cai, J. Cao, T. Cao, E. Cha, Y.-W. Chao et al., âWorld simulation with video foundation models for physical ai,â arXiv preprint arXiv:2511.00062, 2025. [7]R. Team, Z. Gao, Q. Wang, Y. Zeng, J. Zhu, K. L. Cheng, Y. Li, H. Wang, Y. Xu, S. Ma et al., âAdvancing open-source world models,â arXiv preprint arXiv:2601.20540, 2026. [8] Z. Wu, A. Kag, I. Skorokhodov, W. Menapace, A. Mirzaei, I. Gilitschenski, S. Tulyakov, and A. Siarohin, âDensedpo: Fine-grained temporal preference optimization for video diffusion models,â arXiv preprint arXiv:2506.03517, 2025. [9]J. Liu, J. Han, B. Yan, H. Wu, F. Zhu, X. Wang, Y. Jiang, B. Peng, and Z. Yuan, âInfinitystar: Unified spacetime autoregressive modeling for visual generation,â 2025. [Online]. Available: https://arxiv.org/abs/2511.04675 [10]Y. HaCohen, N. Chiprut, B. Brazowski, D. Shalem, D. Moshe, E. Richardson, E. Levin, G. Shiran, N. Zabari, O. Gordon et al., âLtx-video: Realtime video latent diffusion,â arXiv preprint arXiv:2501.00103, 2024. [11]Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng et al., âCogvideox: Text-to-video diffusion models with an expert transformer,â arXiv preprint arXiv:2408.06072, 2024. [12] W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang, âCogvideo: Large-scale pretraining for text-to-video generation via transformers,â arXiv preprint arXiv:2205.15868, 2022. [13]W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang et al., âHunyuanvideo: A systematic framework for large video generative models,â arXiv preprint arXiv:2412.03603, 2024. [14]M. Elmoghany, L. Zhao, X. Shen, S. Mukherjee, Y. Zhou, G. Wu, V. D. Lai, S. Yoon, R. Rossi, A. Rashwan et al., âInfinitystory: Unlimited video generation with world consistency and character-aware shot transitions,â arXiv preprint arXiv:2603.03646, 2026. [15]D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, âLearning spatiotemporal features with 3d convolutional networks,â in Proceedings of the IEEE international conference on computer vision, 2015, p. 4489â4497. [16] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ĺ. Kaiser, and I. Polosukhin, âAttention is all you need,â Advances in neural information processing systems, vol. 30, 2017. [17] Z. Tong, Y. Song, J. Wang, and L. Wang, âVideomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training,â Advances in neural information processing systems, vol. 35, p. 10 078â10 093, 2022. [18]Y. Xing, Y. Fei, Y. He, J. Chen, J. Xie, X. Chi, and Q. Chen, âLarge motion video autoencoding with cross-modal video vae,â arXiv preprint arXiv:2412.17805, 2024. [19]B. Kang, Y. Yue, R. Lu, Z. Lin, Y. Zhao, K. Wang, G. Huang, and J. Feng, âHow far is video generation from world model: A physical law perspective,â arXiv preprint arXiv:2411.02385, 2024. [20]Y. Qin, Z. Shi, J. Yu, X. Wang, E. Zhou, L. Li, Z. Yin, X. Liu, L. Sheng, J. Shao et al., âWorldsimbench: Towards video generation models as world simulators,â arXiv preprint arXiv:2410.18072, 2024. [21]J. Ding, Y. Zhang, Y. Shang, Y. Zhang, Z. Zong, J. Feng, Y. Yuan, H. Su, N. Li, N. Sukiennik et al., âUnderstanding world or predicting future? a comprehensive survey of world models,â ACM Computing Surveys, vol. 58, no. 3, p. 1â38, 2025. [22]L. Wang, Z. Chen, Y. Du, D. Yan, W. Ge, G. Shen, X. Xu, L. Wu, M. Chen, T. Xu et al., âA mechanistic view on video generation as world models: State and dynamics,â arXiv preprint arXiv:2601.17067, 2026. 11 [23]Y. Wang, S. Xing, C. Can, R. Li, H. Hua, K. Tian, Z. Mo, X. Gao, K. Wu, S. Zhou et al., âGenerative ai for autonomous driving: Frontiers and opportunities,â arXiv preprint arXiv:2505.08854, 2025. [24]Y. Yuan, X. Wang, T. Wickremasinghe, Z. Nadir, B. Ma, and S. H. Chan, âNewtongen: Physics-consistent and controllable text-to-video generation via neural newtonian dynamics,â arXiv preprint arXiv:2509.21309, 2025. [25]Z. Gao, J. Mao, H.-X. Yu, H. Lou, E. Y.-T. Jia, J. Barbic, J. Wu, and Y. Wang, âSeeing the wind from a falling leaf,â arXiv preprint arXiv:2512.00762, 2025. [26] L. Wang, X. Liu, S. Lin, G. Xu, and H.-Y. Shum, âGeneric slow-motion replay detection in sports video,â in 2004 International Conference on Image Processing, 2004. ICIPâ04., vol. 3. IEEE, 2004, p. 1585â1588. [27] C.-M. Chen and L.-H. Chen, âA novel method for slow motion replay detection in broadcast basketball video,â Multimedia Tools and Applications, vol. 74, no. 21, p. 9573â9593, 2015. [28] V. Kiani and H. R. Pourreza, âAn effective slow-motion detection approach for compressed soccer videos,â International Scholarly Research Notices, vol. 2012, no. 1, p. 959508, 2012. [29]S. Benaim, A. Ephrat, O. Lang, I. Mosseri, W. T. Freeman, M. Rubinstein, M. Irani, and T. Dekel, âSpeednet: Learning the speediness in videos,â in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, p. 9922â9931. [30]L. C. Pickup, Z. Pan, D. Wei, Y. Shih, C. Zhang, A. Zisserman, B. Scholkopf, and W. T. Freeman, âSeeing the arrow of time,â in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, p. 2035â2042. [31] D. Wei, J. J. Lim, A. Zisserman, and W. T. Freeman, âLearning and using the arrow of time,â in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, p. 8052â8060. [32]E. P. Bennett and L. McMillan, âComputational time-lapse video,â in ACM SIGGRAPH 2007 papers, 2007, p. 102âes. [33]N. Petrovic, N. Jojic, and T. S. Huang, âAdaptive video fast forward,â Multimedia Tools and Applications, vol. 26, no. 3, p. 327â344, 2005. [34]F. Zhou, S. Bing Kang, and M. F. Cohen, âTime-mapping using space-time saliency,â in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, p. 3358â3365. [35]S. Lan, R. Panda, Q. Zhu, and A. K. Roy-Chowdhury, âFfnet: Video fast-forwarding via reinforcement learning,â in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, p. 6771â6780. [36]M. Silva, W. Ramos, J. Ferreira, F. Chamone, M. Campos, and E. R. Nascimento, âA weighted sparse sampling and smoothing frame transition approach for semantic fast-forward first-person videos,â in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, p. 2383â2392. [37]M. M. da Silva, âSemantic hyperlapse: a sparse coding based and multi-importance approach for first-person videos,â 2019. [38] B. Jähne, Digital image processing. Springer, 2005. [39] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, âImage quality assessment: from error visibility to structural similarity,â IEEE transactions on image processing, vol. 13, no. 4, p. 600â612, 2004. [40]R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, âThe unreasonable effectiveness of deep features as a perceptual metric,â in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, p. 586â595. [41] Z. Tu, X. Yu, Y. Wang, N. Birkbeck, B. Adsumilli, and A. C. Bovik, âRapique: Rapid and accurate video quality prediction of user generated content,â IEEE Open Journal of Signal Processing, vol. 2, p. 425â440, 2021. [42]Q. Zheng, Z. Tu, P. C. Madhusudana, X. Zeng, A. C. Bovik, and Y. Fan, âFaver: Blind quality prediction of variable frame rate videos,â Signal Processing: Image Communication, vol. 122, p. 117101, 2024. [43]I. Skorokhodov, S. Tulyakov, and M. Elhoseiny, âStylegan-v: A continuous video generator with the price, image quality and perks of stylegan2,â in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, p. 3626â3636. [44] Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit et al., âVbench: Comprehensive benchmark suite for video generative models,â in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, p. 21 807â21 818. [45]D. Zheng, Z. Huang, H. Liu, K. Zou, Y. He, F. Zhang, L. Gu, Y. Zhang, J. He, W.-S. Zheng et al., âVbench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness,â arXiv preprint arXiv:2503.21755, 2025. 12 [46]Z. Huang, F. Zhang, X. Xu, Y. He, J. Yu, Z. Dong, Q. Ma, N. Chanpaisit, C. Si, Y. Jiang et al., âVbench++: Comprehensive and versatile benchmark suite for video generative models,â IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. [47] H. Duan, H.-X. Yu, S. Chen, L. Fei-Fei, and J. Wu, âWorldscore: A unified evaluation benchmark for world generation,â in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, p. 27 713â 27 724. [48]S. Su, M. Delbracio, J. Wang, G. Sapiro, W. Heidrich, and O. Wang, âDeep video deblurring for hand-held cameras,â in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, p. 1279â1288. [49]D. Danier, F. Zhang, and D. R. Bull, âBvi-vfi: A video quality database for video frame interpolation,â IEEE Transactions on Image Processing, vol. 32, p. 6004â6019, 2023. [50]A. Mercat, M. Viitanen, and J. Vanne, âUvg dataset: 50/120fps 4k sequences for video codec analysis and development,â in Proceedings of the 11th ACM multimedia systems conference, 2020, p. 297â302. [51]V. Ramanishka, Y.-T. Chen, T. Misu, and K. Saenko, âToward driving scene understanding: A dataset for learning driver behavior and causal reasoning,â in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, p. 7699â7707. [52]D. Mehta, H. Rhodin, D. Casas, P. Fua, O. Sotnychenko, W. Xu, and C. Theobalt, âMonocular 3d human pose estimation in the wild using improved cnn supervision,â in 2017 international conference on 3D vision (3DV). IEEE, 2017, p. 506â516. [53]Z. Huang, T. Zhang, W. Heng, B. Shi, and S. Zhou, âReal-time intermediate flow estimation for video frame interpolation,â in Proceedings of the European Conference on Computer Vision (ECCV), 2022. [54] C.-K. Liang, Y.-C. Peng, and H. Chen, âRolling shutter distortion correction,â in Visual Communications and Image Processing 2005, vol. 5960. SPIE, 2005, p. 1315â1322. [55]J. Liu, J. Han, B. Yan, H. Wu, F. Zhu, X. Wang, Y. Jiang, B. Peng, and Z. Yuan, âInfinitystar: Unified spacetime autoregressive modeling for visual generation,â arXiv preprint arXiv:2511.04675, 2025. [56]DeepMind, âVeo 3 technical report,â DeepMind, Technical Report, 2025, accessed: 2026-02-18. [Online]. Available: https://storage.googleapis.com/deepmind-media/veo/Veo-3-Tech-Report.pdf [57] OpenAI, âSora: Creating video from text,â 2024. [58] xAI, âGrok Imagine â ai image & video generation by xai,â 2026. [59] Kling AI, âKling AI Omni / VIDEO O1 creative interface,â 2025. [60] ByteDance Seed Team, âByteDance Seed: Models and research,â https://seed.bytedance.com/, 2025, accessed: 2026-02-27. [61]R. A. Bradley and M. E. Terry, âRank analysis of incomplete block designs: I. the method of paired comparisons,â Biometrika, vol. 39, no. 3/4, p. 324â345, 1952. [62] Google DeepMind, âGemini 3.1 Pro â deepmind ai model,â 2025. [63] â, âGemini 3 Flash â deepmind ai model,â 2025. [64]A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv et al., âQwen3 technical report,â arXiv preprint arXiv:2505.09388, 2025. 13