Paper deep dive
RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement
Ziheng Jia, Jiaying Qian, Zicheng Zhang, Xiaorong Zhu, Lancheng Gao, Xiongkuo Min
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/12/2026, 3:26:04 AM
Summary
The paper introduces RAVEN-Eval, a rubric-guided automated evaluation framework for AI Video Generation Models (AIVGMs) using Large Multimodal Models (LMMs) as judges. It addresses the limitations of human evaluation and fixed-dimension metrics by curating 250 tasks (150 T2V, 100 I2V) and generating 4,500+ videos. The framework employs task-specific rubrics for pairwise preference judgments, establishing leaderboards for 20 AIVGMs and 13 LMM judges to ensure scalable, fine-grained, and trustworthy evaluation.
Entities (10)
Relation Signals (7)
RAVEN-Eval → uses → LMM-as-a-judge paradigm
confidence 95% · RAVEN-Eval, a rubric-guided automated evaluation framework for AIVGMs, built primarily on the LMM-as-a-judge paradigm.
RAVEN-Eval → employs → Rubric-Guided Pairwise Judgment
confidence 93% · At its core, RAVEN-Eval adopts rubric-guided automated LMM preference judgement
RAVEN-Eval → curates → 250 tasks
confidence 92% · RAVEN-Eval curates 150 text-to-video (T2V) tasks and 100 image-to-video (I2V) tasks
RAVEN-Eval → evaluates → 20 AIVGMs
confidence 90% · Finally, we evaluate 20 high-performance AIVGMs
RAVEN-Eval → introduces → Anchor-based Model Insertion
confidence 88% · It further introduces an anchor-based model insertion approach to reduce the evaluation cost of incorporating new models.
GPT-5.4 → usedas → Judge
confidence 85% · we perform LMM-as-a-judge pairwise preference judgement for each task using three LMMs: GPT-5.4
Seedance2 Fast → usedfor → Task Filtering
confidence 85% · we select Seedance2 Fast ... and LTX 2.3 Fast as preliminary solvers ... to assess prompt fulfillment
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video generation models~(AIVGMs) have become increasingly difficult to discern using conventional evaluation criteria, such as visual fidelity and semantic instruction following. Meanwhile, human evaluation now requires more expertise and sustained attention, substantially increasing annotation costs. This calls for automated evaluation that can reliably distinguish fine-grained differences among advanced AIVGMs with minimal human intervention. To address this challenge, we present RAVEN-Eval, a rubric-guided automated evaluation framework for AIVGMs, built primarily on the LMM-as-a-judge paradigm. Through an automatic task curation and quality-filtering pipeline, RAVEN-Eval curates 150 text-to-video~(T2V) tasks and 100 image-to-video~(I2V) tasks, and systematically collects more than 4,500 AIGVs. At its core, RAVEN-Eval adopts rubric-guided automated LMM preference judgement, in which LMM judges conduct pairwise comparisons according to task-specific rubrics. It further introduces an anchor-based model insertion approach to reduce the evaluation cost of incorporating new models. Finally, we evaluate 20 high-performance AIVGMs, as well as the judging capabilities of 13 LMM judges, and establish the RAVEN-Eval Leaderboards. Overall, RAVEN-Eval paves a scalable path for automatic and trustworthy evaluation of rapidly evolving AIVGMs.
Tags
Links
- Source: https://arxiv.org/abs/2608.09111v1
- Canonical: https://arxiv.org/abs/2608.09111v1
Trouble viewing inline? Open PDF directly →
Full Text
202,957 characters extracted from source content.
Expand or collapse full text
RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement Ziheng Jia 1 , Jiaying Qian 1 , Zicheng Zhang 2 , Xiaorong Zhu 1 , Lancheng Gao 1 , Xiongkuo Min 1 * 1 Shanghai Jiao Tong University 2 Shanghai Artificial Intelligence Laboratory jzhws1@sjtu.edu.cn Abstract AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video gen- eration models (AIVGMs) have become increasingly diffi- cult to discern using conventional evaluation criteria, such as visual fidelity and semantic instruction following. Mean- while, human evaluation now requires more expertise and sustained attention, substantially increasing annotation costs. This calls for automated evaluation that can reliably distin- guish fine-grained differences among advanced AIVGMs with minimal human intervention. To address this challenge, we present RAVEN-Eval, a rubric-guided automated evaluation framework for AIVGMs, built primarily on the LMM-as- a-judge paradigm. Through an automatic task curation and quality-filtering pipeline, RAVEN-Eval curates150text-to- video (T2V) tasks and100image-to-video (I2V) tasks, and systematically collects more than4, 500AIGVs. At its core, RAVEN-Eval adopts rubric-guided automated LMM pref- erence judgement, in which LMM judges conduct pairwise comparisons according to task-specific rubrics. It further intro- duces an anchor-based model insertion approach to reduce the evaluation cost of incorporating new models. Finally, we evaluate20high-performance AIVGMs, as well as the judging capabilities of13LMM judges, and establish the RAVEN- Eval Leaderboards. Overall, RAVEN-Eval paves a scalable path for automatic and trustworthy evaluation of rapidly evolv- ing AIVGMs. Introduction With the widespread commercialization of AI-generated videos (AIGVs), reliable and scalable evaluation of AI video generation models (AIVGMs) has become increasingly crit- ical since effective benchmarks guide model selection and characterize model strengths across diverse generation scenar- ios. However, the rapid progress of state-of-the-art (SOTA) AIVGMs is outpacing conventional evaluation paradigms, creating new demands for more discriminative assessment. Here, we identify two major challenges. First, widely used leaderboards such as (Artificial Analy- sis 2026) rely heavily on human preference annotations. As the number of AIGVs grows, the associated cost becomes in- creasingly substantial. Moreover, improvements in AIVGMs make quality differences more subtle, requiring finer-grained * Corresponding author. Motivation Guideline AIGVs rapidly evolve and commercialize Reliable evaluation becomes essential Challenge A Human preference-based Annotation Challenge B Models based on predefined dimensions or absolute scoring Challenge B Predefined dims. Visual Quality Motion Quality Consistency Aesthetic Naturalness ... Background Task-centric Benchmark T2V I2V Rubric-guided Pairwise LMM Judge LMM Judge RAVEN-Eval Leaderboards 250 Tasks 4500+ Videos 20 AIGV Models 13 LMM Judges Task- specific Rubric Guidance Costly,time-consuming,hard to scale Fine-grained differences remain hard to capture Attention Time Cost AIVGM Leaderboards LMM Judges Leaderboards Video Generation Evaluation LMM Preference Understanding Evaluation RAVEN-Eval A loofah being crushed by a hydraulic press. Multiple athletes completing a race in a stadium. Figure 1: To address the rising cost and limited fine-grained discrimination of existing AIVGM evaluation methods, we introduce RAVEN-Eval, a task-centric, rubric-guided LMM pairwise judgment framework for T2V and I2V tasks. perception. Annotators must therefore invest greater exper- tise, attention, and effort to make reliable decisions, further increasing cost and potentially reducing annotation consis- tency. Second, although semantic instruction following, basic physical consistency, and visual fidelity remain essential evaluation criteria, recent SOTA AIVGMs have become in- creasingly competitive along these dimensions. Consequently, frameworks based on fixed dimensions or handcrafted met- rics (Huang et al. 2024; Wang et al. 2025) are less effective at distinguishing advanced models. Meanwhile, commercial- ization raises user expectations for precise and professional generation, making the remaining capability gaps more evi- dent in fine-grained detail control and general intelligence under task-dependent and context-specific scenarios. Cap- turing these differences therefore requires task-specific, fine- grained preference judgments. This motivates our central question: How can an automatic evaluation framework with arXiv:2608.09111v1 [cs.AI] 10 Aug 2026 limited human involvement faithfully capture the capability differences among recent SOTA AIVGMs? To address this question, we propose RAVEN-Eval, a task-centricrubric-guidedautomatic evaluation framework that employs large multimodal model (LMM) for preference judgement, as illustrated in Fig. 1. Its feasibility is supported by two observations. First, current LMMs exhibit strong video understanding and can reliably identify fine-grained dif- ferences under effective rubric guidance. Compared with hu- man annotators, they are less susceptible to irrelevant visual content or contextual distractions, enabling more consistent attention to task-relevant details. Second, absolute scoring may fail to distinguish high-performing AIVGMs because of limited score granularity. Pairwise preference judgments of- fer a more suitable alternative, as relative comparison places lower perceptual and decision-making demands on the eval- uator than precise absolute scoring. Our main contributions are summarized as follows: •We construct the RAVEN-Eval Benchmark for fine- grained detail control and task-dependent model intelli- gence. Rather than relying on conventional predefined di- mensions, we build task-centric evaluation scenarios with context-specific requirements through an LMM-assisted generation and quality-filtering pipeline. The benchmark comprises250diverse T2V and I2V tasks and4,669high- quality generated videos. •We introduce rubric-guided pairwise preference judge- ment under the LMM-as-a-judge paradigm. Each task is equipped with a customized rubric that guides LMM judges to compare details and determine pairwise prefer- ences. Model capabilities are estimated through tie-aware optimization, while an anchor-based model insertion trick further enables efficient incorporation of new AIVGMs. •We establish two complementary RAVEN-Eval Leader- boards: the AIGV Leaderboards for ranking20re- cent open-source and proprietary generation models, and the Judge Leaderboards for benchmarking13 general-purpose LMM judges against human preferences. Their strong performance across diverse general LMMs demonstrates the training-free nature and strong model- harnessing capability of RAVEN-Eval. Related Works Existing AIVGM evaluation can be roughly categorized into human annotation-based platforms and automatic frame- works using trained models or predefined toolkits. Arena-style platforms such as Artificial Analysis (Artifi- cial Analysis 2026), OpenCompass (Cao et al. 2026), and Video Arena (Video Arena 2026) typically adopt human pair- wise comparison protocols and dynamic scheduling systems, such as Elo-style ranking (Dangauthier et al. 2007; Glick- man 1999), to continuously update model scores from human votes. Recent automatic evaluation approaches mainly rely on pretrained models and handcrafted metrics. EvalCrafter (Liu et al. 2024) and the VBench series (Huang et al. 2024) estab- lish benchmarks and toolkits based on predefined dimensions and corresponding automatic metrics. VBench++ (Huang et al. 2025) and VBench-2.0 (Zheng et al. 2025) further broaden the capability coverage and granularity of this dimension-based paradigm. VideoScore (He et al. 2024), Video-Bench (Han et al. 2025) explores using vision-language models as automatic evaluators. T2V-CompBench (Sun et al. 2025) evaluates compositional T2V generation through MLLM-, detection-, and tracking-based metrics tailored to different compositional categories, but still relies on abso- lute score assignment. VideoPhy (Bansal et al. 2025a) and VideoPhy-2 (Bansal et al. 2025b) focus on physical common- sense through tasks that test real-world consistency. Scorers such as AIGV-Assessor (Wang et al. 2025) and LOVE (Wang et al. 2026) investigate LMM-based AIGV assessment and mean opinion score (MOS) prediction. Despite their effectiveness, existing approaches still have limitations in evaluating SOTA AIVGMs as described above. Human evaluation requires substantial resources, while most automatic methods rely on predefined dimensions or absolute scoring, which may not sufficiently capture subtle capability differences among advanced models. These observations also provide compelling motivation for our work. The RAVEN-Eval Benchmark Task Curation To evaluate AIVGMs across diverse scenarios, we construct three task types: T2V generation, keyframe expansion task (KFT), and first–last frame completion task (FLT). Each is generated and filtered through a dedicated automatic pipeline, as illustrated in Fig. 2, with dataset statistics re- ported in Fig. 3. Examples are provided in Supp. Sec. A.6. T2V TasksAs discussed above, evaluating SOTA AIVGMs requires greater emphasis on fine-grained detail control and model intelligence. We therefore define three complemen- tary capability dimensions for prompt construction: scene complexity (S), dynamic interaction complexity (D), and expert-level knowledge requirement (K).Scontrols the difficulty of object and spatial-detail realization. S0 scenes contain only a few objects with simple layouts; S1 scenes contain multiple distinct objects with nontrivial composi- tional relationships; and S2 scenes contain groups of objects arranged in complex spatial structures.Dcontrols motion and interaction complexity. D0 tasks involve nearly static objects or limited motion with simple interactions, whereas D1 tasks require coordinated, highly dynamic, and complex interaction processes.Kintroduces domain-specific knowl- edge and reasoning requirements. K0 tasks depict everyday situations requiring no additional expertise, while K1 tasks require accurate representation of specialized phenomena, processes, or constraints, including physical experiments, chemical reactions, material transformations, precise text ren- dering, specialized photographic or aesthetic styles, and other knowledge-intensive scenarios. We additionally designate ap- proximately20%of the tasks as reasoning-based tasks to further assess general reasoning ability. CombiningS,D, and Kenables each task to jointly evaluate multiple advanced capabilities rather than a single predefined quality dimension. For each valid task, we use the proprietary LMM GPT- 5.5 (OpenAI 2026b) to generate a prompt skeleton defining A1 Task-specific Pairwise Rubric Construction T2V Task Construction Capability Dimensions Object Quantities S - Scene Complexity low/medium/high D - Dynamic Interaction K - Knowledge Requirement low / high low / high Prompt Skeleton Scene Objects Action Context Reasonig Clues Detail Enrichment Spatial Constraints Motion Details Interaction Process Expert Knowledge Two-stage Filtering Stage1: Easy tasks preliminary removal Stage2: Quality scrunity with two-level solvers 150 T2V Tasks A2 I2V Task Construction B1 Task Information Prompt Task Type Reference Frame(s) Reasoning Clues Two-level Importance Rubric Primary Criteria - Highest Priority In-context object details (T2V) Interaction and motion details Expected phenomenon / events (I2V) Physical or reasoning consistency Reference frame (s) adherence (I2V) Basic Criteria - Secondary Priority Non-primary object consistency Visual Fidelity Visual Attractiveness Naturalness LMM-as-judge Paradigm Priority-aware decision Rubric-guided LMM annotation Pairwise Preference Judgment B2 Video Generation 20 Models for T2V/KFT 14 Models for FLT 3 Generations per model- per task Best-of-three selection Reference consistency verification (I2V) capability C1 LMM Judges & Pair Input Video pair Task prompt Rubric Sampled frames Tie-aware Score Estimation Order-robust Pairwise Judging AIVGM Leaderboards Order-swap verification · Confidence-aware processing ·Three-judges aggregation Modified Davison Tie-aware MLE: Optimizes model scores while explicitly modeling tie probability ·Task-level score estimation · Leaderboards on different task types · Uncertainty-aware score process LMM Judges Candidate Pool Independent Ranking Estimation LMM Judge Ability Leaderboards Source Videos Quality Filtering Segments Mining Task Types · Timestamps samling · Segments extraction · Finegrained segments cropping 50 KFTs 50 FLTs 13 recent LMMs (open-sourced+proprietary) LMM-judge-specific ranking estimation · Fine-grained evaluation on LMM AIGV preference judging ability RAVEN-Eval AIGV Model Ranking Human Preference Annotation External Online Leaderboards High Ranking Consistency Fully connected Comparison Graph Reliablility-based Anchor Selection Ranking Reconstruction Ranking Reliability Validation Quality scrunity with different level solvers(presented in ) Task safety / reliability check by GPT · 50 representative tasks subset · 5 expert-level annotators · Trueskill based scheduling · Commonly used webside leaderboards (Artificial Analysis / Arena. AI) 20 models for T2V / KFT · Divide into 5 tiers · Top-K anchors per tier based on judge- agreement reliability Compare with: · Fully-connected graph optimized ranking · Human ranking · External leaderboards ranking New Model Selected Anchors Other Models ASMR Videos C Simulation/ Interaction Videos Detail Control FLT First-last-frame Completion Task KFT Keyframe Expansion Task I2V Reasoning Tasks To definite state To open-ended state ? Task LTX-2.3-Fast (low-level) Seedance2-Fast (high-level) Otherwise Semantic completion LMM Ranking Detail control Rank Model Elo Model A Model B Model C . . . LMM Judges I n p u t LMM A LMM B LMM C . . . Rank Model SRCC Conf. . . . A < B B < A = B Input 1 Input 2 AB BA Confident Preference Confident Tie Uncertain · Calculate alignment with human rankings A1 Extract shared leaderboards models GPT-5.4 Gemini-3.1-Flash-Lite Claude-4.6-Sonnet · Perform rubric-guided LMM annotation GPT-Series Gemini-Series Claude-Series Qwen-Series ... ... Measuring alignment with intra-dataset user preference Measuring alignment with generalized professional ranking 1 2 3 C2 D1 D2 RAVEN-Eval AIVGM Leaderboards RAVEN-Eval Judge Leaderboards Leaderboard Reliability Validation Anchor-based Model Insertion Task curation & Filtering Pairwise Rubric Construction & Video Generation RAVEN-Eval Leaderboards ReliabilityValidation & Anchor-based Model Insertion A B CD Primary Important Video A (Model A)Video B (Model B) + + + Prompt Prompt Clue ? Reasoning + ......... ......... ? Figure 2: The detailed visualization of the complete RAVEN-Eval construction pipeline. the core scene and interactions according to its assigned ca- pability dimensions, and then make task detail enrichment. For S1 and S2 tasks, we add fine-grained compositional and spatial constraints together with explicit object-count require- ments. For D1 tasks, we provide more detailed specifications of motion patterns and interaction processes. For K1 tasks, we incorporate domain-specific descriptions of the expected phenomena or processes. For reasoning-based tasks, how- ever, the prompt contains only clues, without explicitly stat- ing the expected visual outcome. The intended phenomenon or event, which is subsequently included in the task rubric, must be uniquely inferable from the provided clues using general reasoning and relevant domain knowledge. Following the task-generation scheme above, we construct a large candidate pool with balanced coverage across com- binations ofS,D, andK. We then apply a two-stage auto- matic filtering pipeline to retain T2V tasks with sufficient discriminative power. First, we remove all S0 and S1-D0 tasks. These low-difficulty combinations are generated during prompt generation to preserve a complete difficulty hierar- chy, but are unlikely to reveal meaningful differences among advanced AIVGMs and are therefore excluded from the final benchmark. In the second stage, inspired by the data-curation paradigm with two-level solvers in (Kulikov et al. 2026) and balancing capability with inference cost, we select Seedance2 Fast (Seedance Team et al. 2026) and LTX 2.3 Fast (LTX-2.3 2026) as preliminary solvers with substantially different ca- pability levels. The former is a high-performance proprietary model, whereas the latter is a smaller and relatively weaker open-weight model. Both solvers generate videos for every task retained after the first-stage filtering. We then use GPT- 5.4-mini (OpenAI 2026d) to assess prompt fulfillment on a five-point ordinal scale, jointly considering core semantic completion and realization of specified details. We retain tasks for which Seedance2 Fast scores at least3and LTX 2.3 Fast scores at most2. These tasks exhibit clear performance gaps between solvers with different prior capabilities and are therefore more informative for distinguishing AIVGMs. I2V Tasks We adopt a separate construction strategy for FLTs and KFTs. Since reference images impose stronger visual constraints than text alone, these tasks are better suited to evaluating whether models can generate well- specified, domain-relevant interaction and transforma- tion processes. We collect Creative Commons-licensed (C) videos, with a focus on autonomous sensory meridian response (ASMR) content (Barratt and Davis 2015). Such videos typ- T2V KFT FLT Figure 3: Statistical data of the RAVEN-Eval Benchmark. It comprises 250 tasks and 4, 669 videos in total. ically capture brief yet information-dense processes in spe- cialized settings, such as hydraulic presses crushing materials or red-hot iron balls penetrating different substances. Their reliance on physical and material behavior, together with durations well suited to mainstream AIVGMs, makes them valuable sources for challenging I2V tasks. We collect ap- proximately1, 000ASMR videos from YouTube and about 500additional videos featuring complex object interactions. In the final benchmark, the KFT set contains31tasks derived from ASMR videos and19from other sources, while the FLT set contains35ASMR-derived tasks and15other tasks. For each source video, we sample one frame every3seconds and provide the sequence to GPT-5.4-mini for segment mining. The model identifies intervals containing complete physical transformations or object interactions and estimates their tem- poral boundaries. We then uniformly sample16frames from each interval and use the same model to refine the segment boundaries. Segments with insufficient variation between the first and last frames are removed based on pooled pixel-vector similarity. For each retained segment, GPT-5.4-mini gener- ates an I2V prompt according to the task type. For FLTs, the prompt specifies a coherent transition between the provided first and last frames, with both endpoints strictly aligned to their references and the intermediate process reasonably inferred from the source segment. For KFTs, the prompt aligns the initial state with the provided keyframe, while the subsequent event is open-ended, only required to remain logically consistent. We also designate approximately20% of the I2V tasks as reasoning-based tasks, whose prompts provide only clues derived from the reference images without explicitly stating the expected outcome. These clues must uniquely imply a predictable phenomenon or event, requiring the AIVGM to infer and depict the intended content. Follow- ing the Stage-2 T2V filtering protocol, we use Seedance2 Fast, LTX 2.3 Fast, and GPT-5.4-mini to filter the I2V tasks. Finally, a GPT-assisted safety and reliability check removes potentially harmful or copyright-restricted content. Rubric GenerationAs discussed above, recent LMMs are suitable for fine-grained multimodal understanding tasks when provided with clear checklists. Therefore, construct- ing task-centric effective rubrics to accurately guide LMM judges to distinguish the quality differences between videos is critical. We also visualize this process in Fig. 2. For each task, we construct a two-level task-specific pref- erence evaluation rubric, consisting of primary basic crite- ria. The primary compliance criteria mainly emphasize how accurately the generated video reproduces the detailed requirements specified in the prompt. For T2V tasks, we use GPT-5.5 to extract the enriched details previously added to the prompts and combine them with the original prompt content to form itemized evaluation criteria. For I2V tasks, since the reference images provide richer semantic informa- tion and the prompt text mainly details the required events or phenomena, we directly reuse the prompt content as the pri- mary criteria. Furthermore, for all the reasoning-based tasks, since the prompts only contain clues, we supplement the pri- mary compliance criteria according to the task construction records, providing detailed requirements for the expected content presentation. The basic criteria are assigned lower priority than the primary criteria, and the same set is ap- plied uniformly to all tasks. They assess non-primary object consistency, visual fidelity, aesthetic appeal, and naturalness. Detailed definitions are provided in Supp. Sec. A.3. Benchmark Video GenerationWe demonstrate the video generation process in Fig 2. Since video generation models differ in their capable task ranges, we evaluate20models on both T2V tasks and KFTs, and14models on FLTs. For a clearer presentation, these models are reported together in the leaderboard shown in Tab. 1. For each task, we randomly select a configuration from the combinations of aspect ratio, resolution, and duration commonly supported by the models. If an individual model does not support the selected configu- ration, we use its closest available configuration instead. All videos are generated without audio. To reduce the effect of randomness in generation, each model generates3videos for each task. We use GPT-5.4-mini to rank the candidates based on task-semantic completion and detail control, and retain the highest-ranked video. For KFTs and FLTs, we further verify consistency with the reference images. Specifically, we compare the first frame of each KFT video, or the first and last frames of each FLT video, with the corresponding reference images using the similarity between pooled pixel-value vectors. Generated videos with a similarity below the threshold of 0.98 are excluded. RAVEN-Eval Leaderboards RAVEN-Eval AIVGM LeaderboardsAfter constructing the benchmark, we perform LMM-as-a-judge pairwise pref- erence judgement for each task using three LMMs: GPT- 5.4 (OpenAI 2026a), Claude Sonnet 4.6 (Anthropic 2026c), and Gemini 3.1 Flash-Lite (DeepMind 2026a). To avoid intra- dataset selection bias, these judges are chosen a priori based on the trade-off between their recognized multimodal un- derstanding capabilities and inference costs. The judgement input content includes8uniformly sampled frames from each video, the prompt, the task-specific rubric, and, for I2V tasks, the corresponding visual reference. Detailed settings are provided in Supp. Sec. A.5. Because some AIGV pairs remain difficult to distinguish even with task-specific rubrics, judges may output a “tie”. To control judgment stability, we adopt an order-robust pair- wise judging and confidence-aware processing protocol, under which each pair is evaluated twice with reversed pre- sentation orders. Both outputs are first mapped back to the corresponding model identities. A preference that remains consistent after reversal, or two “tie” outputs, is treated as a confident judgment; all other cases, including an order- inconsistent preference or a single “tie”, are treated as un- certain judgments and downweighted during capability esti- mation. For each task, all feasible model pairs are evaluated by all judges, forming a fully connected comparison graph that also supports the subsequent anchor-based insertion ex- periments. We then estimate task-level model capabilities using a modified Davidson tie-aware maximum likelihood estimation objective (Davidson 1970): arg max s i N i=1 ,γ 1 |P| X p= (i,j)∈P λ p 1 M M X m=1 h w (m) p,i>j logP (i≻ j)+ w (m) p,j>i logP(j≻i)+w (m) p,i∼j logP (i∼j) i − α N N X i=1 s 2 i −βγ 2 , s.t. N X i=1 s i = 0, θ i = exp (s i ), η = exp (γ), P (i≻ j) = θ i θ i + θ j + η p θ i θ j , P (j ≻ i) = θ j θ i + θ j + η p θ i θ j , P (i∼ j) = η p θ i θ j θ i + θ j + η p θ i θ j . (1) Here,Pdenotes the set of valid pairwise comparisons for a given task, wherep = (i,j)represents a comparison between two AIGV models.NandMdenote the numbers of AIVGMs to be evaluated and LMM judges, respectively. The variable s i ∈R is the latent capability score of model i, whileγ ∈Rcontrols the overall tie propensity throughθ i = exp (s i ) andη = exp (γ). Letc (m) p ∈0,1 indicate whether them-th judge provides a confident judgment for pairp. The pair-level reliability weight is defined as the proportion of confident judgments,λ p = 1 M P M m=1 c (m) p = 1−ℓ p /M, where ℓ p = P M m=1 (1−c (m) p )is the number of uncertain judgments. The outcome weightsw (m) p,i>j ,w (m) p,j>i , andw (m) p,i∼j encode the result provided by them-th judge. For a confident judgment, the weight of the resolved outcome is set to1, while the other two weights are set to0. For an uncertain judgment, we assignw (m) p,i>j =w (m) p,j>i = 1/2andw (m) p,i∼j = 0, representing equal uncertainty over the two preference directions. The coefficientsαandβregularize the latent capability scores and the tie parameter, respectively. The constraint P N i=1 s i = 0 resolves the location non-identifiability of the latent scores. We optimize the objective independently for each task for1,000iterations. For each task categoryc, a model’s final capability ̄s i,c is calculated as its mean latent score across all tasks in that category minus the standard deviation of these scores. The resulting capability scores are then linearly rescaled to an Elo-style range (Elo i,c = 1000 + 400 ln 10 ̄s i,c ) for intuitive presentation in Tab. 1. The process of building the leaderboard is also shown in Fig. 2. Reliability Verification To verify the reliability of the RAVEN-Eval AIVGM Leaderboards, we conduct human pref- erence experiments on50tasks, including30T2V tasks,10 KFTs, and10FLTs. T2V tasks are sampled to preserve the original distribution of theS,D, andKlevels, while I2V tasks approximately retain the original ratio of reasoning to non-reasoning tasks. We recruit5experts to provide pair- wise preference annotations. To reduce annotation cost, we adopt a TrueSkill-based adaptive scheduling strategy (Her- brich, Minka, and Graepel 2006). For each task, all annotators begin with the same cold-start pairs, with every model par- ticipating in4randomly sampled comparisons. We assess inter-annotator agreement by pooling the cold-start annota- tions across all tasks and computing nominal Krippendorff ’s α(Krippendorff 1970), treating the two directional prefer- ences and the tie as three categorical outcomes. The resulting αof0.714indicates acceptable agreement. After cold start, each annotator maintains an independent state and receives a separately scheduled sequence of comparisons. For each expert, the task-level ranking is updated using Eq. 1 after every10newly annotated pairs. Annotation terminates when all pairwise Spearman rank correlation coefficient (SRCC) values among the three most recent rankings exceed0.95, and the meanσacross models falls below0.3. For each task, we average the independently optimized capability scores from all annotators to obtain a consensus human score for every model. These scores are then averaged across the selected tasks to construct a human-reference ranking for each task type. Further details are provided in Supp. Sec. A.4. We compare the RAVEN-Eval AIVGM Leaderboards with the human rankings and two external online leader- boards, Artificial Analysis (A) (Artificial Analysis 2026) and Arena AI (Arena.) (Arena 2026). Since both external platforms provide only I2V leaderboards, we merge the KFT and FLT results into a unified I2V ranking. For each ex- ternal comparison, we retain only the models shared with our evaluation set (detailed in Supp.) and compute the corre- sponding SRCC. The overall validation process is illustrated in Fig. 2, and results are reported in the last row of Tab. 2 for readability. The strong agreement across multiple human- annotation-based references demonstrates the reliability of our evaluation framework. The Key Role of Rubrics To assess the effectiveness of rubric-guided pairwise evaluation, we conduct an ablation study using the same LMM judges. In no-rubric pairwise evaluation (NR-PW), the judges select the preferred video us- ing only the prompt and video, without access to the rubric. In rubric-guided absolute scoring (RG-AS), each video is inde- pendently assigned a continuous score from0to5under the original task-specific rubric, and model scores are averaged across tasks. Our rubric-guided pairwise evaluation (RG- PW) directly compares each video pair using the corre- sponding rubric. We additionally evaluate two representative trained, dimension-based absolute scorers, VideoScore (He et al. 2024) and LOVE (Wang et al. 2026), using their pub- licly released weights. For each task category, both models score every generated video along their predefined dimen- sions. Scores from each dimension are linearly normalized to a common scale across all evaluated videos and AIVGMs, av- eraged across dimensions for each video, and then averaged across videos to obtain the final score of each AIVGM. The CategoryT2VFLTKFT AIVGMRankScoreCI# VotesRankScoreCI# VotesRankScoreCI# Votes Seedance 21840.7 ±2184871890.5 ±10194111094.3 ±162847 G-I-V (Grok Imagine Video 2026)2791.5±198475–3978.9±162844 Kling 3.0 Pro (Kling 3.0 2026)3764.6 ±2384964817.5 ±2019354839.1 ±192832 HappyHorse V1.04762.6±288487–10690.1±272829 Kling 3.0 Omni5733.6 ±2184872854.7 ±1219358720.3 ±202826 Seedance 2 Fast6727.7±1585085799.5±1619475828.4±302835 Veo 3.1 Fast (Veo 3.1 2025)7680.0 ±13848412130.9 ±3519449706.8 ±172841 Veo 3.1 Lite8594.6±21849910363.3±30194412641.6±162823 PixVerse C1 (PixVerse C1 2026)9567.7 ±1685177749.7 ±14193821047.4 ±142832 PixVerse V6 (PixVerse V6 2026)10555.3±1784908609.5±35193211686.6±282838 Kling 3.0 Std11536.2 ±2584846786.7 ±1319446785.1 ±182832 Wan 2.7 (WanVideo 2026)12509.9±2384963848.4±1319357731.8±142832 PixVerse V5.5 (PixVerse V5.5 2025)13470.8 ±15849311274.0 ±28193516328.5 ±262832 Seedance 1.5 Pro (Chen et al. 2025)14393.9±1685299506.7±21193515337.9±292835 LTX 2.3 Fast15344.1 ±23852014-468.1 ±14194417-155.9 ±232841 Wan 2.616342.2±198484–13638.3±222826 LTX 2.3 Pro (LTX-2.3 2026)1713.3 ±2585381373.5 ±15193218-203.0 ±222838 Hailuo 2.3 (Hailuo2.3 2025)185.8±278520–14400.1±192826 LTX 2.0 Pro (HaCohen et al. 2026)19-27.6 ±238493–19-323.2 ±192832 LTX 2.0 Fast20-162.2±198514–20-336.8±222844 Table 1: The RAVEN-Eval AIVGM Leaderboards. # Votes denotes the number of effective LMM pairwise judgements involving the model. “CI” denotes the 95% confidence interval. [Per column, the highest value is shown in bold, and the second in italics.] CategoryT2VI2V SettingHuman A Arena. H. (KFT) H. (FLT) A Arena. NR-PW0.754 0.733 0.7670.8520.785 0.802 0.677 RG-AS0.6720.4550.6220.6330.5850.4320.548 VideoScore 0.723 0.611 0.5950.6540.542 0.577 0.536 LOVE0.7110.6430.7340.6250.5880.6070.575 Ours0.872 0.835 0.8100.9030.821 0.851 0.714 Table 2: Ablation results (SRCC) on the effectiveness of task- specific rubrics. The term “H.” is short for “Human”. resulting rankings are compared with our human-reference ranking and external leaderboards, as reported in Tab. 2. The results show that absolute scoring, including both rubric- based (RG-AS) and trained predefined-dimension evaluators, produces substantially weaker ranking consistency than our setting. Moreover, our setting clearly outperforms NR-PW, demonstrating the importance of task-specific rubric guid- ance. Anchor-Based New Model Insertion Although LMM evaluation substantially reduces costs, the number of pairwise comparisons still grows quadratically with the size of the model pool. Maintaining a fully connected comparison graph is therefore impractical for continuously updating the model pool. Inspired by (Zhu et al. 2024), we introduce an anchor-based insertion strategy (also shown in Fig. 2), in which newly added models are compared only with a subset of anchors from the existing model pool. Anchor Selection Since fully connected comparison graphs are available for all20AIVGMs on both T2V tasks and KFTs, we conduct the following experiments indepen- dently for these two task categories. In each trial, we ran- domly select15models to form a initial pool and treat the remaining5as newly inserted models. For each candidate CategoryT2VKFT SettingFull Human AAFull Human A K = 1,M = 1 (Rand.) 0.767 0.732 0.705 0.792 0.771 0.722 K = 1,M = 2 (Rand.)0.7920.7410.7300.8080.7820.731 K = 1,M = 3 (Rand.) 0.811 0.752 0.749 0.820 0.783 0.728 K = 1,M = 3 (Ours)0.8340.7700.7550.8510.7980.747 K = 2,M = 1 (Rand.) 0.818 0.789 0.768 0.856 0.801 0.772 K = 2,M = 2 (Rand.)0.8200.7850.7810.8640.8220.791 K = 2,M = 3 (Rand.) 0.832 0.791 0.785 0.878 0.828 0.801 K = 2,M = 3 (Ours)0.8540.8110.8140.9030.8570.835 Table 3: Ablation results (SRCC) under differentK,Mset- tings using either the proposed anchor-selection or random selection. “Full” uses the ranking from the complete compar- ison graph as the reference. “Rand.” denotes random pick. modeliin the initial pool, we construct four one-dimensional rank vectors for each task category. Each vector has the same length as the number of evaluated tasks, with each element recording the final rank of modelion one task. All ranks are derived exclusively from the fully connected compari- son graph of the selected15models. The vectorr (all) i is obtained using all three LMM judges (M = 3), while r (m) i 3 m=1 are obtained separately using only them-th judge (M = 1). We measure the agreement between the m-th judge and the joint evaluation of modeliasρ (m) i = SRCC (r (m) i ,r (all) i ). The mean agreement and cross-judge standard deviation are defined as ̄ρ i = 1 3 P 3 m=1 ρ (m) i and v i = Std (ρ (1) i ,ρ (2) i ,ρ (3) i ), respectively. We define the an- chor reliability score asq i = ̄ρ i − τv i , whereτcontrols the penalty for cross-judge inconsistency. A largerq i indicates stronger agreement with the joint evaluation and greater sta- bility across judges, making modelia more reliable anchor candidate in the current task space. CategoryT2VFLTKFT LMMRank SRCC KRCCConf.Rank SRCC KRCCConf.Rank SRCC KRCCConf. Claude Opus 4.6 (Anthropic 2026a)10.8530.68684.8%20.8670.72579.1%30.7890.64381.4% GPT-5.4 (OpenAI 2026a)20.8270.67378.9%10.8900.76580.1%10.8170.68772.3% GPT-5.6 Sol (OpenAI 2026c)30.8370.66068.5%50.8370.66070.2%40.7760.60672.3% Claude Opus 4.8 (Anthropic 2026b)40.8330.66085.3%30.8700.69983.2%60.6800.47279.5% Gemini 3.1 Pro (DeepMind 2026b)50.8080.58268.9%70.7790.62171.8%50.6830.47675.3% GPT-5.6 Terra60.7710.59578.5%40.8370.69980.4%100.4990.34576.8% Claude Sonnet 4.6 (Anthropic 2026c)70.7690.59566.2%80.7420.58265.7%20.8050.66768.4% Qwen3.7-Max (Qwen3.7 2026)80.7770.58264.3%60.8290.64758.5%90.5340.31548.2% Gemini 3.1 FL (DeepMind 2026a)90.7420.51752.2%100.6780.45134.7%70.6480.44549.3% GPT-5.4 Mini (OpenAI 2026d)100.7010.46248.6%120.5690.41235.3%120.3310.19437.4% Claude Haiku 4.5 (Anthropic 2025)110.6540.38640.5%90.6820.50338.7%80.6410.41543.8% GPT-5.6 Luna120.5920.29420.6%110.6080.45133.1%110.3850.21518.7% Qwen3.6-27B♠ (Qwen3.6-27B 2026)130.4230.1929.9%130.5350.36211.3%130.3470.17615.6% Table 4: The RAVEN-Eval Judge Leaderboards. “Conf.” denotes the proportion of pair judgments noted as confident for each LMM judge.♠ denotes an open-weight model. LMMs are ranked according to the sum of their SRCC and KRCC values. Model Insertion Simulation For each random split, we first reconstruct a leaderboard for the15initial models from their fully connected comparison graph, following the same procedure used for the RAVEN-Eval AIVGM Leaderboards. This process is performed independently for each task cate- gory. We then divide the15models into five consecutive tiers of three models according to the reconstructed ranking. Un- der our anchor-selection strategy, the top-Kmodels with the highest reliability scores are selected from each tier, yielding 5Kanchors in total. We evaluateK ∈1, 2, corresponding to 5 and 10 anchors, respectively. The remaining5models are treated as newly inserted mod- els. We retain the fully connected graph among the15initial models and add only comparisons between each inserted model and the selected anchors. Applying the optimization in Eq. 1 to this combined graph jointly estimates the capa- bility scores of all20models and reconstructs the expanded leaderboard. As a baseline, we repeat the same procedure using an equal number of randomly selected anchors from each tier. We further study the effect of the number of LMM judges usingM ∈1, 2, 3. More judges may provide more diverse evidence for distinguishing inserted models from the anchors, thereby improving ranking alignment. We evaluate different combinations ofK,M, and anchor- selection strategy. ForM= 1andM=2, we consider all three possible single-judge and two-judge combinations, respec- tively, and report their averaged results. Since our reliability- based anchor selection requires outputs from all three judges, theM= 1andM=2settings use random anchors, whereas M = 3includes both our strategy and the random baseline. For each setting, we compute the SRCC against the fully connected20-model ranking in Tab. 1, the human-annotated ranking, and the external A leaderboard. For comparison with the human-annotated ranking, anchors are selected us- ing the same50human-annotated tasks. Each configuration is evaluated over10independent15/5split trials, and the average results are reported in Tab. 3. When inserting5new models, the anchor-based strategy reduces the pairwise annotation cost by34.2%forK = 1 and21.1%forK = 2, while producing rankings close (with SRCC over0.8) to those inferred from the fully connected comparison graph. With a fixed number of anchors, increas- ingMis more likely to improve ranking reliability. Moreover, our reliability-based anchor selection clearly outperforms ran- dom selection across different settings. These results provide empirical support for the large-scale incorporation of future models using representative anchors. RAVEN-Eval Judge LeaderboardsEvaluating AIVGMs also provides a testbed for LMM judges, particularly their ability to recognize fine-grained quality differences among AIGVs. We therefore assess a range of recent open-weight and proprietary LMMs using the same inputs, evaluation pro- tocol, and score-estimation procedure as those used in the RAVEN-Eval AIVGM Leaderboards. For each judge, we inde- pendently derive an AIVGM ranking and measure its agree- ment with the human reference using SRCC and Kendall’s Rank Correlation Coefficient (KRCC). The resulting RAVEN- Eval Judge Leaderboards, reported in Tab. 4 and visualized in Fig. 2, benchmark fine-grained AIGV preference-judging capability. We observe that larger-size LMMs generally out- perform smaller ones, while stronger judges also tend to produce higher proportions of confident judgments. The Judge Leaderboards further demonstrate the potential of RAVEN-Eval for effective LMM judge model harness- ing. We observe that several mainstream proprietary LMMs achieve competitive performance, indicating that the frame- work is not overly sensitive to a particular judge choice. This flexibility provides a foundation for harnessing heteroge- neous LMMs according to user-specific requirements and for further building more adaptable and scalable evaluation systems. Conclusion We present RAVEN-Eval, an automated framework for eval- uating AIVGMs through rubric-guided LMM preference an- notations. The RAVEN-Eval Benchmark comprises250 T2V and I2V tasks and over4, 500AIGVs. Based on this, we establish the RAVEN-Eval Leaderboards, which compre- hensively record the capability of20SOTA AIVGMs and the evaluation capabilities of13LMM judges. We further intro- duce an anchor-based model insertion strategy to control the evaluation cost as the model pool continues to expand. Overall, RAVEN-Eval establishes an efficient and scalable paradigm for automatic AIVGM evaluation. References Anthropic. 2025. Claude Haiku 4.5 System Card. Anthropic. 2026a. Claude Opus 4.6 System Card. Anthropic. 2026b. Claude Opus 4.8 System Card. Anthropic. 2026c. Claude Sonnet 4.6 System Card. Arena. 2026. https://arena.ai/leaderboard. Artificial Analysis. 2026. https://artificialanalysis.ai/video/ leaderboard. Bansal, H.; Lin, Z.; Xie, T.; Zong, Z.; Yarom, M.; Bitton, Y.; Jiang, C.; Sun, Y.; Chang, K.-W.; and Grover, A. 2025a. Evaluating Physical Commonsense for Video Generation. In International Conference on Learning Representations. Bansal, H.; Peng, C.; Bitton, Y.; Goldenberg, R.; Grover, A.; and Chang, K.-W. 2025b. VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation. arXiv preprint arXiv:2503.06800. Barratt, E. L.; and Davis, N. J. 2015. Autonomous Sensory Meridian Response (ASMR): a flow-like mental state. PeerJ, 3: e851. Cao, M.; Chen, K.; Duan, H.; Fang, Y.; Fei, Z.; Gao, T.; Ge, J.; Li, M.; Liu, H.; Liu, J.; Liu, Y.; Lyu, C.; Lyu, H.; Ma, N.; Ma, Z.; Sun, Y.; Wu, Z.; Xiao, L.; Xiong, Z.; Xu, J.; Ye, H.; Yu, Z.; Yuan, Y.; Zhang, S.; Zhao, Y.; Zhou, F.; Zhou, P.; Zhu, D.; Zhu, L.; and Zhuo, J. 2026. A Universal Evaluation Platform for Large Language Models. arXiv preprint arXiv:2605.19276. Chen, S.; et al. 2025. Seedance 1.5 Pro: A Native Audio- Visual Joint Generation Foundation Model. arXiv preprint arXiv:2512.13507. Dangauthier, P.; Herbrich, R.; Minka, T.; and Graepel, T. 2007. Trueskill through time: Revisiting the history of chess. Advances in neural information processing systems, 20. Davidson, R. R. 1970. On Extending the Bradley-Terry Model to Accommodate Ties in Paired Comparison Exper- iments. Journal of the American Statistical Association, 65(329): 317–328. DeepMind. 2026a. Gemini 3.1 Flash-Lite Model Card. DeepMind. 2026b. Gemini 3.1 Pro Model Card. Glickman, M. E. 1999. Parameter estimation in large dy- namic paired comparison experiments. Journal of the Royal Statistical Society Series C: Applied Statistics, 48(3): 377– 394. Grok Imagine Video. 2026. https://docs.x.ai/developers/ models/grok-imagine-video. HaCohen, Y.; Brazowski, B.; Chiprut, N.; Bitterman, Y.; Kvochko, A.; Berkowitz, A.; Shalem, D.; Lifschitz, D.; Moshe, D.; Porat, E.; Richardson, E.; et al. 2026. LTX- 2: Efficient Joint Audio-Visual Foundation Model. arXiv preprint arXiv:2601.03233. Hailuo2.3. 2025. https://w.minimax.io/news/minimax- hailuo-23. Han, H.; Li, S.; Chen, J.; Yuan, Y.; Wu, Y.; Deng, Y.; Leong, C. T.; Du, H.; Fu, J.; Li, Y.; Zhang, J.; Zhang, C.; Li, L.- j.; and Ni, Y. 2025. Video-Bench: Human-Aligned Video Generation Benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 18858–18868. He, X.; Jiang, D.; Zhang, G.; Ku, M.; Soni, A.; Siu, S.; Chen, H.; Chandra, A.; Jiang, Z.; Arulraj, A.; Wang, K.; Do, Q. D.; Ni, Y.; Lyu, B.; Narsupalli, Y.; Fan, R.; Lyu, Z.; Lin, B. Y.; and Chen, W. 2024. VideoScore: Building Automatic Met- rics to Simulate Fine-Grained Human Feedback for Video Generation. In Proceedings of the 2024 Conference on Empir- ical Methods in Natural Language Processing, 2105–2123. Association for Computational Linguistics. Herbrich, R.; Minka, T.; and Graepel, T. 2006. TrueSkill: A Bayesian Skill Rating System. In Advances in Neural Information Processing Systems, volume 19. Huang, Z.; He, Y.; Yu, J.; Zhang, F.; Si, C.; Jiang, Y.; Zhang, Y.; Wu, T.; Jin, Q.; Chanpaisit, N.; Wang, Y.; Chen, X.; Wang, L.; Lin, D.; Qiao, Y.; and Liu, Z. 2024. VBench: Compre- hensive Benchmark Suite for Video Generative Models. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, 21807–21818. Huang, Z.; Zhang, F.; Xu, X.; He, Y.; Yu, J.; Dong, Z.; Ma, Q.; Chanpaisit, N.; Si, C.; Jiang, Y.; Wang, Y.; Chen, X.; Chen, Y.-C.; Wang, L.; Lin, D.; Qiao, Y.; and Liu, Z. 2025. VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models. IEEE Transactions on Pattern Analysis and Machine Intelligence. Kling 3.0. 2026. https://kling.ai/. Krippendorff, K. 1970. Estimating the reliability, systematic error and random error of interval data. Educational and psychological measurement, 30(1): 61–70. Kulikov, I.; Whitehouse, C.; Wu, T.; Nie, Y.; Saha, S.; Helenowski, E.; Yuan, W.; Golovneva, O.; Lanchantin, J.; Bachrach, Y.; et al. 2026. Autodata: An agentic data sci- entist to create high quality synthetic data. arXiv preprint arXiv:2606.25996. Liu, Y.; Cun, X.; Liu, X.; Wang, X.; Zhang, Y.; Chen, H.; Liu, Y.; Zeng, T.; Chan, R.; and Shan, Y. 2024. EvalCrafter: Bench- marking and Evaluating Large Video Generation Models. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, 22139–22149. LTX-2.3. 2026. https://ltx.io/model/ltx-2-3. OpenAI. 2026a. GPT-5.4 Thinking System Card. OpenAI. 2026b. GPT-5.5 System Card. https://openai.com/ index/gpt-5-5-system-card/. OpenAI. 2026c. GPT-5.6 System Card. OpenAI. 2026d.https://developers.openai.com/api/docs/ models/gpt-5.4-mini. PixVerse C1. 2026.https://pixverse.ai/en/blog/pixverse- introduces-c1-ai-video-model-for-film-production. PixVerse V5.5. 2025. https://app.pixverse.ai/. PixVerse V6. 2026.https://pixverse.ai/en/blog/pixverse- launches-v6-advancing-ai-video-generation. Qwen3.6-27B. 2026. https://huggingface.co/Qwen/Qwen3.6- 27B. Qwen3.7. 2026. https://qwen.ai/blog?id=qwen3.7. Seedance Team; et al. 2026. Seedance 2.0: Advancing Video Generation for World Complexity. arXiv preprint arXiv:2604.14148. Sun, K.; Huang, K.; Liu, X.; Wu, Y.; Xu, Z.; Li, Z.; and Liu, X. 2025. T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-Video Generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 8406–8416. Veo 3.1. 2025. https://deepmind.google/models/veo/. Video Arena. 2026. https://w.videoarena.tv/leaderboard. Wang, J.; Duan, H.; Jia, Z.; Zhao, Y.; Yang, W. Y.; Zhang, Z.; Chen, Z.; Wang, J.; Xing, Y.; Zhai, G.; and Min, X. 2026. LOVE: Benchmarking and Evaluating Text-to-Video Genera- tion and Video-to-Text Interpretation. International Confer- ence on Machine Learning. Wang, J.; Duan, H.; Zhai, G.; Wang, J.; and Min, X. 2025. AIGV-Assessor: Benchmarking and Evaluating the Percep- tual Quality of Text-to-Video Generation with LMM. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, 18869–18880. WanVideo. 2026. https://wan.video/. Zheng, D.; Huang, Z.; Liu, H.; Zou, K.; He, Y.; Zhang, F.; Gu, L.; Zhang, Y.; He, J.; Zheng, W.-S.; Qiao, Y.; and Liu, Z. 2025. VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness. arXiv preprint arXiv:2503.21755. Zhu, H.; Wu, H.; Li, Y.; Zhang, Z.; Chen, B.; Zhu, L.; Fang, Y.; Zhai, G.; Lin, W.; and Wang, S. 2024. Adaptive image quality assessment via teaching large multimodal model to compare. Advances in Neural Information Processing Sys- tems, 37: 32611–32629. A Supplementary Materials A.1 Information on the Evaluated AIVGMs This subsection provides supplementary information on the AIVGMs evaluated in RAVEN-Eval, including their open- source status, year of first public release, and official model or product pages. A model is regarded as open source when its official model weights are publicly available for down- load, even if its complete training code or training data are unavailable. 1.Seedance 2. Seedance 2 is a closed-source model first publicly released in 2026. Official Link: https://seed. bytedance.com/en/seedance20. 2.Kling 3.0 Pro. Kling 3.0 Pro is a closed-source model first publicly released in 2026. Official Link: https://home. klingai.com/. 3.Grok-Image-Video. Grok-Image-Video, officially pro- vided as Grok Imagine Video, is a closed-source model first publicly released in 2026. Official Link: https: //docs.x.ai/developers/models/grok-imagine-video. 4.HappyHorse V1.0. HappyHorse V1.0 is a closed-source model first publicly released in 2026. Official Link: https: //w.alibabacloud.com/blog/603068. 5.Kling 3.0 Omni. Kling 3.0 Omni is a closed-source model first publicly released in 2026. Official Link: https://home.klingai.com/. 6.Seedance 2 Fast. Seedance 2 Fast is a closed-source model first publicly released in 2026. Official Link: https://w.volcengine.com/product/dreamart. 7.PixVerse C1. PixVerse C1 is a closed-source model first publicly released in 2026. Official Link: https://pixverse.ai/en/blog/pixverse-introduces-c1- ai-video-model-for-film-production. 8.Kling 3.0 Std. Kling 3.0 Std is a closed-source model first publicly released in 2026. Official Link: https://home. klingai.com/. 9.WAN 2.7. WAN 2.7 is a closed-source model first publicly released in 2026. Official Link: https://w.alibabacloud.com/help/en/model- studio/video-generate-edit-model. 10.PixVerse V6. PixVerse V6 is a closed-source model first publicly released in 2026. Official Link: https://pixverse.ai/en/blog/pixverse-launches-v6- advancing-ai-video-generation. 11.Veo 3.1 Fast. Veo 3.1 Fast is a closed-source model first publicly released in 2025. Official Link: https://blog. google/innovation-and-ai/products/veo-updates-flow/. 12.Veo 3.1 Lite. Veo 3.1 Lite is a closed-source model first publicly released in 2026. Official Link: https://cloud.google.com/blog/products/ai-machine- learning/veo-3-1-lite-and-a-new-veo-upscaling- capability-on-vertex-ai. 13. LTX 2.3 Pro. LTX 2.3 Pro was first publicly released in 2026 and is regarded as open source because its official model weights are publicly available. Official Link: https: //ltx.io/blog/ltx-2-3-release. 14.WAN 2.6. WAN 2.6 is a closed-source model first publicly released in 2025. Official Link: https://w.alibabacloud.com/help/en/model- studio/video-generate-edit-model. 15.PixVerse V5.5. PixVerse V5.5 is a closed-source model first publicly released in 2025. Official Link: https://docs. platform.pixverse.ai/changelogs-906383m0. 16.LTX 2.3 Fast. LTX 2.3 Fast was first publicly released in 2026 and is regarded as open source because its official model weights are publicly available. Official Link: https: //ltx.io/blog/ltx-2-3-release. 17.Seedance 1.5 Pro. Seedance 1.5 Pro is a closed-source model first publicly released in 2025. Official Link: https: //seed.bytedance.com/en/seedance15. 18.Hailuo 2.3. Hailuo 2.3 is a closed-source model first publicly released in 2025. Official Link: https://w. minimax.io/news/minimax-hailuo-23. 19.LTX 2.0 Pro. LTX 2.0 Pro was first publicly released in 2025 and is regarded as open source because its official model weights are publicly available. Official Link: https://ltx.io/newsroom/ltx-2-foundation-ai-video- model-is-released. 20.LTX 2.0 Fast. LTX 2.0 Fast was first publicly released in 2025 and is regarded as open source because its official model weights are publicly available. Official Link: https://ltx.io/newsroom/ltx-2-foundation-ai-video- model-is-released. A.2 Information on the Evaluated LMMs This subsection provides supplementary information on the LMM judges evaluated in RAVEN-Eval, including their open- source status, year of first public release, and official model or product pages. Following the criterion used for the evaluated AIVGMs, a model is regarded as open source when its official model weights are publicly available for download, even if its complete training code or training data are unavailable. 1.Claude Opus 4.6. Claude Opus 4.6 is a closed-source model first publicly released in 2026. Official Link: https: //w.anthropic.com/news/claude-opus-4-6. 2.GPT-5.4. GPT-5.4 is a closed-source model first publicly released in 2026. Official Link: https://openai.com/index/ introducing-gpt-5-4/. 3. GPT-5.6 Sol. GPT-5.6 Sol is a closed-source model first publicly released in 2026 as the flagship member of the GPT-5.6 family. Official Link: https://openai.com/index/ gpt-5-6/. 4.Claude Opus 4.8. Claude Opus 4.8 is a closed-source model first publicly released in 2026. Official Link: https: //w.anthropic.com/news/claude-opus-4-8. 5.Gemini 3.1 Pro. Gemini 3.1 Pro is a closed-source model first publicly released in 2026. Official Link: https:// deepmind.google/models/model-cards/gemini-3-1-pro/. 6.GPT-5.6 Terra. GPT-5.6 Terra is a closed-source model first publicly released in 2026 as the balanced, lower- cost member of the GPT-5.6 family. Official Link: https: //openai.com/index/gpt-5-6/. 7.Claude Sonnet 4.6. Claude Sonnet 4.6 is a closed-source model first publicly released in 2026. Official Link: https: //w.anthropic.com/news/claude-sonnet-4-6. 8.Qwen3.7-Max. Qwen3.7-Max is a closed-source model first publicly released in 2026, with access provided through Qwen’s hosted services rather than download- able model weights. Official Link: https://qwen.ai/blog? id=qwen3.7. 9.Gemini 3.1 Flash-Lite. Gemini 3.1 Flash-Lite is a closed-source model first publicly released in 2026. Offi- cial Link: https://deepmind.google/models/model-cards/ gemini-3-1-flash-lite/. 10.GPT-5.4 Mini. GPT-5.4 Mini is a closed-source model first publicly released in 2026 as a smaller and more efficient member of the GPT-5.4 family. Offi- cial Link: https://openai.com/index/introducing-gpt-5- 4-mini-and-nano/. 11.Claude Haiku 4.5. Claude Haiku 4.5 is a closed-source model first publicly released in 2025. Official Link: https: //w.anthropic.com/news/claude-haiku-4-5. 12.GPT-5.6 Luna. GPT-5.6 Luna is a closed-source model first publicly released in 2026 as the fastest and most cost- efficient member of the GPT-5.6 family. Official Link: https://openai.com/index/gpt-5-6/. 13.Qwen3.6-27B. Qwen3.6-27B was first publicly released in 2026 and is regarded as open source because its official model weights are publicly available under the Apache 2.0 license. Official Link: https://huggingface.co/Qwen/ Qwen3.6-27B. A.3 Prompt Summary The Basic Criteria in Task-Specific Rubrics 1.Static non-target objects and background elements should remain consistent in both shape and quantity throughout the video. 2.Unless otherwise specified by the task instruction, the video should preserve practical realism as much as possible. It should follow real-world common sense, especially ensuring that any text appearing in the video is realistic, legible, and understandable. It should also follow real-world physical, material, and reaction laws. In particular, for videos involving material reactions or object interactions, dynamic objects undergoing reactions or interactions should exhibit logically coherent transformation processes. The total amount of material should be conserved as strictly as possible, without unrealistic material generation or abrupt changes. 3.For videos involving dynamic processes, the action or reaction process should be complete and reason- able. Unless otherwise specified by the task instruc- tion, all actions or reactions should appear natural and consistent with real-world logic. 4.After satisfying the above three criteria, which have higher priority, the generated video should have as much aesthetic value and visual appeal as possible. 5.If a primary criteria instruction conflicts with the above four criteria, such as a scene that explicitly does not require realistic physical behavior, the task- specific instruction should take priority. The above four criteria should then be treated as secondary evaluation standards. T2V LMM Judge Prompt T2V task: You are a professional AI text-to-video qual- ity evaluation expert. You are asked to perform pref- erence comparison for a series of AIGC video triples generated by different text-to-video models according to the given evaluation criteria. Each triple includes two AIGC videos generated by different models, the T2V prompt input for this task, and the specific com- parison criteria. The criteria contain two priority levels: primary considerations and basic criteria (lower prior- ity than the primary considerations). When judging preference, you should first strictly com- pare according to the primary considerations. If a pref- erence can be directly determined, use that preference as the output result. If the primary considerations are insufficient, jointly judge according to the basic crite- ria. The final output format is label only: better / worse / similar, indicating whether the first video is better than, worse than, or indistinguishable from the second video under the evaluation criteria. When determining the preference, you must: 1. Strictly base the evaluation result and related inference on the evaluation criteria and their priority order. Do not infer from knowledge outside these criteria. 2. Strictly base the evaluation result on the video pair content and the prompt content. Do not use any other modality or information. 3. Do not consider audio. Audio-related content is unnecessary and prohibited. Only visual content, the prompt, and the given evaluation criteria should be considered. This comparison only involves the following two gen- erated videos: - First video: generatedvideoa, model MODELA- Second video: generatedvideob, modelMODELB prompt:PROMPT RUBRIC:RUBRIC Strictly follow the guideline above. Do not output rea- soning, drafts, analysis, step-by-step explanations, self- correction, or statements about how you will answer. This run uses the labelonly output format. The final response must be exactly one of the following three labels: “better”, “worse”, or “similar”. Do not output any reason or extra text. KFT LMM Judge Prompt Keyframe extension task: You are a professional AI- generated video quality evaluation expert. You are asked to perform preference comparison for a series of AIGC video triples generated by different models according to the given evaluation criteria. Each triple includes two AIGC videos generated by different mod- els and the reference first frame used for the keyframe extension task (the first frame sampled from the ref- erence video is used as the input for the keyframe extension task). The task prompt input and the specific comparison criteria are also provided. The criteria con- tain two priority levels: primary considerations and basic criteria. When judging preference, you should first strictly com- pare according to the primary considerations. If a pref- erence can be directly determined, use that preference as the output result. If the primary considerations are insufficient, jointly judge according to the basic crite- ria. The final output format is label only: better / worse / similar, indicating whether the first video is better than, worse than, or indistinguishable from the second video under the evaluation criteria. When determining the preference, you must: 1. Strictly base the evaluation result and related inference on the evaluation criteria and their priority order. Do not infer from knowledge outside these criteria. 2. Strictly base the evaluation result on the video pair content and the prompt content. Do not use any other modality or information. 3. Do not consider audio. Audio-related content is unnecessary and prohibited. Only visual content, the prompt, and the given evaluation criteria should be considered. This comparison only involves the following two gen- erated videos: - First video: generatedvideoa, model MODELA- Second video: generatedvideob, modelMODELB prompt:PROMPT RUBRIC:RUBRIC Strictly follow the guideline above. Do not output rea- soning, drafts, analysis, step-by-step explanations, self- correction, or statements about how you will answer. This run uses the labelonly output format. The final response must be exactly one of the following three labels: “better”, “worse”, or “similar”. Do not output any reason or extra text. FLT LMM Judge Prompt First-last-frame completion task: You are a profes- sional AI-generated video quality evaluation expert. You are asked to perform preference comparison for a series of AIGC video triples generated by different models according to the given evaluation criteria. Each triple includes two AIGC videos generated by different models and the reference first and last frames used for the first-last-frame completion task (the first and last frames sampled from the reference video are used as the input for the first-last-frame completion task). The task prompt input and the specific comparison crite- ria are also provided. The criteria contain two priority levels: primary considerations and basic criteria. When judging preference, you should first strictly com- pare according to the primary considerations. If a pref- erence can be directly determined, use that preference as the output result. If the primary considerations are insufficient, jointly judge according to the basic crite- ria. The final output format is label only: better / worse / similar, indicating whether the first video is better than, worse than, or indistinguishable from the second video under the evaluation criteria. When determining the preference, you must: 1. Strictly base the evaluation result and related inference on the evaluation criteria and their priority order. Do not infer from knowledge outside these criteria. 2. Strictly base the evaluation result on the video pair content and the prompt content. Do not use any other modality or information. 3. Do not consider audio. Audio-related content is unnecessary and prohibited. Only visual content, the prompt, and the given evaluation criteria should be considered. This comparison only involves the following two gen- erated videos: - First video: generatedvideoa, model MODELA- Second video: generatedvideob, modelMODELB prompt:PROMPT RUBRIC:RUBRIC Strictly follow the guideline above. Do not output rea- soning, drafts, analysis, step-by-step explanations, self- correction, or statements about how you will answer. This run uses the labelonly output format. The final response must be exactly one of the following three labels: “better”, “worse”, or “similar”. Do not output any reason or extra text. Forward and Reversed Evaluation In forward evalua- tion,generatedvideoais used as the first video, and generatedvideob is used as the second video. In reversed evaluation, the input order of the two videos is swapped: the originalgeneratedvideobis used as the first video, and the originalgeneratedvideoais used as the second video. For both forward and reversed evaluation, the output label is always relative to the current input order: “better” means the first video is better, “worse” means the first video is worse, and “similar” means there is no clear quality difference between the two videos. Prompt for GPT-5.4-mini in the task quality filtering stage You are an expert evaluator for text-to-video generation tasks. Your goal is to assess how well a candidate video fulfills the given task on a five-point ordinal scale. You must evaluate the video independently. Do not compare it with videos generated by other models, and do not infer the identity or expected capability of the generation model. Input Task type: T2V TASK TYPE Generation prompt: GENERATION PROMPT Specified task details: DETAIL REQUIREMENTS Hidden expected outcome: HIDDEN EXPECTED OUTCOME Candidate video: VIDEO INPUT The hidden expected outcome is provided only for reasoning-based tasks. If it is marked as ”None”, evalu- ate the video solely according to the generation prompt and specified task details. Evaluation Objective Assess prompt fulfillment by jointly considering: 1. Core semantic completion - Whether the principal subjects, objects, scene, and event described in the prompt are present. - Whether the main action, interac- tion, transformation, or intended outcome is correctly realized. - Whether the generated content preserves the essential meaning of the task rather than merely depicting a related scene. 2. Realization of specified details - Whether explicitly required object quantities, attributes, spatial relations, and compositional constraints are satisfied. - Whether required actions, motion trajectories, interaction se- quences, and state transitions are visibly realized. - Whether domain-specific objects, operations, physical phenomena, material responses, or professional details are depicted correctly. - For reasoning-based tasks, whether the video presents the uniquely expected event or outcome implied by the clues. Core semantic completion has higher priority than secondary details. A visually attractive video that fails to realize the principal task should receive a low score. Conversely, minor visual defects should not dominate the judgment when the core task and most specified details are correctly completed. Only evaluate content that is visibly supported by the video. Do not assume that an unobserved action or event occurred outside the displayed duration. For dy- namic tasks, showing only the initial or final state is insufficient when the prompt explicitly requires an interaction or transformation process. Scoring Scale Score 5 – Excellent fulfillment The video correctly realizes the complete core task and nearly all speci- fied details. The main subjects, interactions, spatial relations, and expected phenomena are clearly and co- herently presented. Only negligible imperfections are present. Score 4 – Strong fulfillment The core task is correctly completed, and most important details are realized. Minor omissions, inaccuracies, or local inconsistencies are present, but they do not substantially affect the intended scene or event. Score 3 – Adequate fulfillment The principal seman- tic content and main task objective are recognizable and substantially completed. However, several speci- fied details are missing or inaccurate, or the process contains noticeable defects. The video remains a valid realization of the task. Score 2 – Weak fulfillment The video only partially re- alizes the task. The main event, interaction, or expected phenomenon is incomplete, incorrect, or insufficiently visible, and multiple important details are absent or contradicted. The video may depict a related scene but does not adequately complete the requested task. Score 1 – Failed fulfillment The video is unrelated to the prompt, omits the principal subjects or event, seriously contradicts the task, or is too corrupted or incomplete to evaluate as a valid realization. Decision Rules - If the principal task or expected event is absent or fundamentally incorrect, assign a score no higher than 2. - If the core task is completed but several secondary details are missing, distinguish between Scores 3 and 4 according to the importance and number of the miss- ing details. - Do not reduce the score solely because of general visual-quality defects unless they prevent recognition of the required content or violate an ex- plicit task requirement. - Do not reward additional content that is not requested if it does not improve fulfillment of the specified task. - Use an integer score only. Output Format Return only the following JSON object: ”score”: 1, ”core semantic completion”: ”complete — partial — failed”, ”detail realization”: ”high — mod- erate — low”, ”missing or incorrect requirements”: ”Briefly list each important missing or incorrect re- quirement” , ”rationale”: ”Provide a concise explana- tion grounded only in visible evidence from the video.” Prompt for T2V Prompt-Skeleton Construction You are an expert task designer for evaluating advanced AI video generation models. Given a specified scene-complexity levelS, dynamic-interaction levelD, domain-knowledge levelK, and reasoning-task indicatorR, your task is to construct a text-to-video prompt skeleton with a clear vi- sual objective and sufficient potential to distinguish models with different capabilities. At this stage, the prompt skeleton should specify only the core scene, principal subjects and objects, primary state or interaction, and necessary contex- tual information. Do not introduce excessive spatial, motion, procedural, or domain-specific details, as these requirements will be added during the subsequent detail-enrichment stage. Input Variables. The scene-complexity levelScontrols the number of subjects or objects and the complexity of their spatial organization.S0denotes a simple composition con- taining a single subject or only a few objects with straight- forward spatial relationships.S1denotes a scene containing multiple distinguishable subjects or objects with nontriv- ial spatial relationships, role assignments, or compositional requirements.S2denotes a scene containing numerous sub- jects or objects organized through complex spatial structures, group relationships, or multilayer compositions. The dynamic-interaction levelDcontrols the complexity of motion, interaction, and temporal evolution.D0denotes a static or nearly static scene that may contain only limited and simple motion, without a complex interaction process. D1denotes a clearly observable dynamic process involv- ing object interactions, coordinated actions, or multistage state transitions, for which temporal continuity and causal consistency are essential. The domain-knowledge levelKcontrols whether correct generation requires specialized knowledge.K0denotes an everyday or general-purpose scenario that does not depend on specialized disciplinary knowledge, professional oper- ations, or material mechanisms.K1denotes a task whose correct visual realization requires domain-specific knowl- edge, such as physical phenomena, chemical reactions, mate- rial transformations, professional sports movements, experi- mental procedures, text or design conventions, specialized photography, or stage performance. The reasoning-task indicatorRspecifies whether the ex- pected visual event is stated explicitly.R0denotes a non- reasoning task in which the required action, process, phe- nomenon, or final state may be described directly.R1de- notes a reasoning-based task in which the public generation prompt may contain only clues from which the expected event or outcome can be uniquely inferred using common- sense or relevant domain knowledge. For anR1task, the expected outcome must not be directly disclosed in the pub- lic prompt. Task Profiles for DifferentS–D–KCombinations. For S0–D0–K0, construct an everyday static scene with a sim- ple subject configuration and uncomplicated composition. ForS1–D0–K0, construct an everyday static scene contain- ing multiple subjects or objects, with emphasis on basic com- position and interpretable object relationships. ForS2–D0– K0, construct a complex everyday static scene containing many subjects or objects, emphasizing compositional plau- sibility, spatial naturalness, and subject consistency while imposing minimal professional-detail requirements. ForS0–D1–K0, construct an everyday dynamic process involving only a small number of subjects or objects and a simple composition. ForS1–D1–K0, construct an everyday dynamic interaction involving multiple subjects or objects, with emphasis on action coordination and plausible object in- teractions. ForS2–D1–K0, construct a complex and highly dynamic everyday scene involving multiple subjects, coordi- nated group motion, or concurrent object interactions, while avoiding unnecessary professional knowledge requirements. ForS0–D0–K1, construct a professional static scene con- taining a single subject or only a few objects. Representa- tive scenarios include standardized text presentation, pro- fessional product display, a specialized model pose, or a single-subject composition governed by explicit aesthetic conventions. The task should emphasize professional de- piction of the principal subject, overall visual style, and compositional quality. ForS1–D0–K1, construct a professional static scene con- taining several subjects or objects. Representative scenarios include arrangements of multiple products, multi-object ty- pography, posed groups of performers, or combinations of professional instruments. The task should emphasize limited static relationships among the principal objects, coherent composition, and accurate professional depiction. ForS2–D0–K1, construct a complex professional static scene containing numerous subjects or objects. Representa- tive scenarios include densely arranged professional prod- ucts, group poses, sculpture ensembles, or laboratory plat- forms containing multiple instruments. The task should em- phasize object-level consistency, structurally plausible ar- rangement, coherent group organization, and accurate pro- fessional details. ForS0–D1–K1, construct a professional dynamic scene involving a single principal object or a small number of objects in a simple composition. Representative scenarios include a hydraulic press deforming a material, a red-hot metal ball penetrating a substance, a glass bottle rolling down stairs, soap being squeezed or cut, or text being written dynamically. The task should emphasize continuity of the complete interaction process, plausible presentation of the professional setting, and completeness of the required action. ForS1–D1–K1, construct a moderately complex profes- sional dynamic scene involving multiple subjects or objects. Representative scenarios include chemical reactions, me- chanics experiments, air-pressure demonstrations, optical experiments, or multi-object material interactions. For sci- entific tasks, the task definition must specify the relevant substances, apparatus, experiment, or mechanism and the expected observable phenomenon with sufficient precision. ForS2–D1–K1, construct the most complex category of professional and highly dynamic tasks. Representative sce- narios include professionally filmed sprint races, coordinated multi-person dance performances, aerial formation maneu- vers, advanced group techniques such as leaf-like descending formations, or complex multi-object mechanics experiments. The task must specify the relevant professional action, ex- periment, or phenomenon sufficiently clearly to support ob- jective evaluation. Task-Design Requirements The task must be realizable within the short durations commonly supported by main- stream AI video generation models and should present a complete and observable state, interaction, or event within that duration. It must provide clear visual evidence for eval- uation and must not primarily depend on audio, dialogue, explanatory subtitles, or information unavailable within the video. The task should meaningfully distinguish advanced models in terms of subject consistency, composition, mo- tion, interaction, physical plausibility, material behavior, or professional knowledge. Avoid tasks that rely mainly on ab- stract concepts, subjective emotions, or outcomes that cannot be verified visually. Avoid real-person identity replication, copyrighted characters, brand-specific imitation, unsafe pro- cedural instructions, or other content unsuitable for a public benchmark. AK1task must contain genuine professional requirements that materially affect the correctness of the generated video. Merely placing an otherwise ordinary event in a laboratory, stadium, studio, or other professional-looking environment is insufficient to satisfyK1. AnR1task must provide clues that uniquely determine a visually observable outcome. If two or more substantially different outcomes remain equally plausible, the task is invalid and must be redesigned. Output Format. Return the result strictly using the follow- ing fields. Task Level: Specify the assigned values ofS,D,K, andR. Core Scene: Describe the environment and fundamental situation of the task in one concise sentence. Subjects and Objects: Identify the principal subjects, key objects, and their basic roles in the scene. Core State or Interaction: Describe the principal static state, action, transformation, or interaction that the video must present. Scene Context: Provide only the environmental, temporal, locational, or situational context required to understand the task. Do not add irrelevant decorative details. Reasoning Clues: Complete this field only whenR1is spec- ified. State the clues that may appear in the public generation prompt without revealing the expected outcome. For anR0 task, output “None.” Hidden Expected Outcome: Describe the event, phe- nomenon, process, or final state that constitutes the correct realization of the task. This field is intended only for subse- quent task filtering and rubric construction. It must not be copied directly into the public prompt for an R1 task. Core Evaluation Focus: State two to four principal capabil- ities assessed by the task, such as compositional consistency, motion continuity, interaction accuracy, physical plausibility, material behavior, or professional-detail realization. Prompt Skeleton: Integrate the core scene, principal sub- jects and objects, core state or interaction, and necessary context into one concise and natural text-to-video generation prompt. Do not include metadata or terms such asS,D,K, R, evaluation, benchmark, rubric, or model capability. For anR1task, include only the permitted reasoning clues and do not explicitly reveal the hidden expected outcome. Before returning the result, verify that the generated task matches the specifiedS,D,K, andRlevels; that its key requirements can be evaluated from visible video evidence; that it can be completed within a short video; and that any reasoning-based outcome is uniquely inferable without being explicitly disclosed. Prompt for T2V Prompt Detail Enrichment You are an expert in refining prompts for the evaluation of advanced AI video generation models. Your task is to enrich a given prompt skeleton according to its scene-complexity levelS, dynamic-interaction levelD, domain-knowledge levelK, and reasoning-task indicatorR, without altering its core semantics. The resulting prompt should contain ex- plicit, observable, and objectively assessable requirements while preserving the original scene, principal subjects, and in- tended event. You must also produce structured task records that can subsequently be used for quality filtering and rubric construction. The input task has already passed preliminary filtering and belongs to one of the following combinations:S1–D1–K0, S2–D0–K0,S2–D1–K0,S1–D1–K1,S2–D0–K1, or S2–D1–K1. Input. The input contains the assigned values ofS,D,K, andR, together with the core scene, subjects and objects, core state or interaction, scene context, reasoning clues, hid- den expected outcome, core evaluation focus, and the origi- nal prompt skeleton. General Enrichment Principles. Preserve the central scene, subjects, objects, and event defined in the original prompt skeleton. Do not replace the task with a different scenario merely to make it easier, harder, or more visually elaborate. Every added requirement must materially contribute to video- quality assessment and must be directly observable in the generated video. Avoid irrelevant decorative details and do not create artificial complexity through excessive adjectives or arbitrary visual constraints. The enriched prompt must remain natural, coherent, and compatible with the short durations commonly supported by mainstream video generation models. All required content should be realizable within one continuous and logically coherent visual event. Avoid tasks that require numerous scene transitions, extended narratives, or information ex- ternal to the video. The public generation prompt must not contain metadata or terms such as “evaluation,” “benchmark,” “rubric,” “model capability,” S1, D1, or K1. Scene-Complexity Enrichment. WhenS = 1, clearly spec- ify the roles of the principal subjects and objects and in- troduce observable spatial constraints, such as their relative positions, orientations, distances, foreground–background re- lationships, or compositional arrangement. The scene should exhibit a clear but not excessively crowded organization. For dynamic tasks, the identities and roles of all interact- ing subjects or objects must remain visually distinguishable throughout the event. WhenS = 2, establish a structured and visually interpretable organization involving multiple subjects or objects. This may include explicit foreground, middle-ground, and background layers, spatial groupings, relative positions, and visual hier- archies. For scenes involving occlusion, intersecting trajecto- ries, or visually similar entities, add requirements concerning identity preservation, spatial continuity, and stable group or- ganization. Scene complexity should arise from coherent multi-object composition rather than arbitrary accumulation of unrelated elements. Dynamic-Interaction Enrichment. WhenD = 0, preserve the static or nearly static nature of the task and do not in- troduce unnecessary complex motion. Enrich the prompt primarily through subject poses, object shapes, spatial ar- rangement, material appearance, and static consistency. Mi- nor movements may be included only when they improve visual naturalness without changing the fundamental static character of the task. WhenD = 1, specify the observable initial state, key in- termediate stages, and final state of the required action or interaction. Clarify the relevant motion directions, trajecto- ries, temporal order, contact relationships, changes in ve- locity, and causal responses among subjects or objects. The complete process should remain temporally continuous and visually coherent. The generated video should avoid abrupt appearances or disappearances, identity exchanges, unin- tended penetrations, implausible deformations, or discontin- uous state changes. For tasks involving collision, compres- sion, penetration, cutting, material transformation, chemical reaction, or coordinated group motion, describe the most important intermediate stages needed to evaluate whether the process has been correctly realized. Domain-Knowledge Enrichment. WhenK = 0, use ob- jects, actions, and scene logic that can be understood through everyday knowledge. Emphasize natural motion, plausible interaction, clear composition, and visual coherence. Do not introduce unnecessary scientific terminology, experimental conditions, or specialized professional conventions. WhenK = 1, identify the relevant professional or scientific domain and specify the key objects, operations, mechanisms, phenomena, or visual conventions that must be correctly depicted. Professional requirements must be translated into observable visual evidence rather than expressed through vague terms such as “professional,” “scientific,” or “accu- rate.” For physical or material processes, specify the relevant mate- rial properties, applied forces, motion or deformation mech- anisms, and expected observable responses. For chemical experiments, identify the relevant substances, operations, and visible outcomes, such as color changes, precipitation, bubbling, crystallization, phase transitions, or other exter- nally observable phenomena. For mechanics, air-pressure, or optical experiments, specify the apparatus, mode of op- eration, and expected visible result. For sports, dance, or professional performances, specify the action or technique, key poses, temporal sequence, coordination among partici- pants, and any essential professional filming requirements. For typography, product presentation, aesthetic composi- tion, or other professional static scenes, specify the relevant layout, text appearance, materials, poses, lighting, and com- positional conventions. Combination-Specific Requirements. For anS1–D1–K0 task, construct an everyday dynamic interaction involving multiple distinguishable subjects or objects. Emphasize ac- tion coordination, interaction order, contact relationships, motion continuity, and plausible causal responses without introducing unnecessary professional knowledge. For anS2–D0–K0task, construct a complex everyday static scene containing multiple subjects or objects. Emphasize spa- tial hierarchy, organized subject distribution, natural poses, object-level consistency, and overall compositional plausi- bility. For anS2–D1–K0task, construct a complex and highly dy- namic everyday scene involving multiple subjects, simulta- neous actions, or coordinated interactions. Emphasize group motion, concurrent object interactions, identity preservation, coherent trajectories, temporal continuity, and dynamically stable composition. For anS1–D1–K1task, construct a professional dynamic process involving several subjects or objects. Emphasize the professional operation, reaction, or interaction mechanism, its key intermediate stages, and the expected observable phenomenon. For anS2–D0–K1task, construct a complex professional static scene containing numerous subjects or objects. Empha- size the standardized arrangement of professional objects, subject poses, structural relationships, material depiction, visual hierarchy, and domain-appropriate composition. For anS2–D1–K1task, construct a complex, highly dy- namic, and knowledge-intensive professional scene involv- ing multiple subjects or objects. Emphasize professional ac- tions, group coordination, multi-object interactions, physical or material consistency, and the complete temporal evolution of the intended event. Reasoning-Task Processing. WhenR = 0, the final public prompt may explicitly describe the required action, process, phenomenon, and expected outcome. All principal task re- quirements may be directly stated. WhenR = 1, the final public prompt may contain only the initial conditions, contextual information, and necessary rea- soning clues. It must not explicitly state the hidden expected outcome, the name of the target phenomenon when that name directly reveals the answer, or the final event. Nevertheless, the clues must be sufficient for the expected outcome to be uniquely inferred through commonsense, physical principles, or relevant domain knowledge. The complete expected pro- cess and result must be retained in the internal task record and rubric information rather than disclosed in the generation prompt. If the available clues permit two or more substan- tially different but equally plausible outcomes, revise the clues until the intended outcome becomes uniquely infer- able. Output Format. Return the result strictly using the follow- ing fields. Task Level: Specify the assigned values ofS,D,K, andR. Final Video Generation Prompt: Produce one natural, com- plete, and directly usable text-to-video generation prompt. The prompt should be written in English and should integrate all necessary spatial, compositional, dynamic, interaction, and domain-specific requirements. For anR1task, include only the permitted clues and do not disclose the hidden ex- pected outcome. Added Spatial and Compositional Constraints: State the newly introduced requirements concerning the principal sub- jects and objects, their identities, relative positions, orien- tations, spatial layers, grouping relationships, and overall composition. Added Motion and Interaction Constraints: State the newly introduced requirements concerning action order, mo- tion trajectories, contact relationships, temporal continuity, causal responses, and state transitions. For aD0task, de- scribe the relevant static-state and consistency requirements instead. Added Domain-Specific Constraints: State the profes- sional objects, operations, mechanisms, phenomena, material properties, or visual conventions that must be accurately pre- sented. For aK0task, output “No additional domain-specific constraints.” Hidden Expected Process and Outcome: Describe the complete correct progression of the event and its intended final state. This information is intended for task filtering and rubric construction and must not be exposed in the public prompt for an R1 task. Likely Failure Modes: Identify the principal errors that would reveal insufficient model capability, such as missing subjects, incorrect object counts, incoherent composition, unstable identities, discontinuous actions, implausible inter- actions, unintended deformation, incorrect physical behavior, inaccurate professional phenomena, or omitted task-specific details. Candidate Rubric Criteria: Convert the task requirements into independently verifiable criteria. Distinguish the pri- mary criteria, which directly determine whether the core task has been correctly completed, from the basic criteria, which assess non-primary object consistency, visual fidelity, aesthetic quality, and naturalness. Self-Verification: Verify whether the final prompt conforms to the specifiedS,D,K, andRlevels; whether every princi- pal requirement can be assessed from visible evidence in the video; whether anR1prompt conceals the expected result while still allowing it to be uniquely inferred; whether the complete event can reasonably occur within a short video; and whether any irrelevant details have been introduced. If any condition is not satisfied, revise the task before returning the final result. Primary Criteria Generation Prompt You are an expert evaluator for AI video generation models. Your task is to construct the primary evalu- ation criteria for a task-specific pairwise preference rubric. The primary criteria should focus on whether the generated video accurately satisfies the core re- quirements of the task and should capture the most im- portant task-dependent factors that distinguish strong and weak generations. General Requirements: - Generate 3-6 concise and in- dependent primary criteria. - Each criterion must cor- respond to an observable property that can be verified from the generated video. - The criteria should pri- oritize task-specific requirements rather than generic video quality. - Do not include generic criteria such as visual fidelity, aesthetic appeal, naturalness, over- all realism, or image quality, as these belong to sec- ondary/basic criteria. - Avoid vague descriptions such as ”high quality”, ”professional”, or ”good consis- tency”. Instead, describe concrete visual evidence that should appear in the generated video. - The criteria should be suitable for pairwise comparison between two generated videos. The task type can be either Text-to-Video (T2V) or Image-to-Video (I2V). Follow the corresponding rules below. For Text-to-Video (T2V) Tasks: Input information includes: - Original task prompt. - Enriched prompt details generated during task con- struction, including object quantities, spatial con- straints, motion details, interaction processes, and domain-specific requirements. - Capability dimen- sions, including scene complexity (S), dynamic in- teraction complexity (D), and knowledge requirement (K). - Reasoning-task indicator. Generate primary criteria by extracting and reorganiz- ing the essential requirements from the original prompt and enriched details, especially the details. Specifically: - For scene-complexity requirements, evaluate whether important subjects and objects, their quantities, spatial relationships, composition structures, and scene organization are correctly real- ized. - For dynamic-interaction requirements, evaluate whether the intended actions, motion patterns, inter- action processes, temporal evolution, and causal rela- tionships are correctly completed. - For knowledge- intensive tasks, evaluate whether domain-specific phe- nomena, mechanisms, materials, operations, or pro- fessional details are accurately represented. - For reasoning-based tasks, the prompt only provides clues rather than explicitly describing the expected outcome. Generate criteria based on the intended phenomenon or event recorded in the task construction information, and evaluate whether the model correctly infers and presents the expected content. The criteria should mainly measure whether the gener- ated video fulfills the core task requirements and real- izes the most discriminative details among advanced AIVGMs. For Image-to-Video (I2V) Tasks: Input information includes: - Reference image(s). - Task prompt. - Task type (KFT or FLT). - Reasoning- task indicator. - For reasoning-based tasks, the ex- pected phenomenon or event recorded in the task con- struction information. For non-reasoning I2V tasks: - Directly reuse the task prompt as the basis of the primary criteria, since the prompt explicitly describes the required event, interac- tion, or transformation. - Organize the explicit require- ments in the prompt into concise and independently verifiable criteria. - Do not introduce additional re- quirements that are not specified by the prompt or ref- erence images. - Additionally consider the constraints imposed by the reference images: - For FLT tasks, evaluate whether the generated video preserves the provided first and last frames and produces a coher- ent transition process between them. - For KFT tasks, evaluate whether the generated video maintains con- sistency with the provided keyframe while generating a logically consistent subsequent event. For reasoning-based I2V tasks: - The prompt only pro- vides clues and does not explicitly reveal the expected outcome. - Do not directly reuse the prompt as the primary criteria. - Generate criteria according to the expected phenomenon or event recorded during task construction. - Describe the observable visual evidence required to determine whether the model correctly in- fers and realizes the intended event from the provided clues and reference images. - Ensure that the criteria evaluate both the inferred outcome and its consistency with the reference content. For all I2V tasks: - Prioritize task compliance, reference-image consistency, event or transformation realization, interaction correctness, and physical or log- ical consistency. - Do not include generic visual quality criteria, which belong to secondary/basic criteria. Output Format: Return only a JSON object: “primary-criteria”: [ “Criterion 1”, “Criterion 2”, “Cri- terion 3” ] Each criterion should be a single concise sentence describing an independently observable requirement. A.4 Human Experiment Details For each evaluation task, we conduct a pairwise human pref- erence study over the candidate generated videos. Given a tasktand its candidate model setM t , each comparison pair is denoted as(i,j), wherei,j ∈ M t andi ̸= j. For each pair, annotators watch the two videos side by side and choose one of three outcomes: the left video is better, the right video is better, or the two videos are similar. Cold-start stage Each task begins with a fixed cold-start stage before adaptive scheduling. We construct a coverage- balanced subset of comparison pairsB t from the full can- didate pair setP t . The cold-start pairs are selected with a round-robin strategy so that every model receives an initial number of comparisons. This prevents the adaptive scheduler from starting with completely unobserved ratings. In our implementation, each task uses four cold-start rounds. All annotations collected during the cold-start stage are replayed into the TrueSkill rating model and are therefore included when estimating the initial posterior distributions for adaptive scheduling. TrueSkill rating modelFor each modelm, we maintain a latent quality distribution: s m ∼N (μ m ,σ 2 m ), whereμ m represents the current estimated quality of the model, andσ m represents the uncertainty of this estimate. All models are initialized as μ m = 5, σ m = 5 6 . After each human comparison, the rating distributions of the two involved models are updated according to the TrueSkill posterior update rule. If one video is preferred, the winner’s mean rating is increased and the loser’s mean rat- ing is decreased, while the uncertainties of both models are reduced. If the two videos are judged to be similar, the com- parison is treated as a draw. The draw behavior is controlled by a draw probability parameter p draw . Adaptive pair schedulingAfter the cold-start stage, com- parison pairs are selected adaptively. For each unannotated pair (i,j), we compute the following acquisition score: A(i,j) = λ σ (σ i +σ j )+exp − |μ i − μ j | β +λ c C(i,j)−λ r R(i,j), whereμ i ,μ j andσ i ,σ j are the current TrueSkill mean and uncertainty values of the two models. The next pair is selected as (i ∗ ,j ∗ ) = argmax (i,j)∈P t t A(i,j), whereA t denotes the set of already annotated pairs for task t. In our implementation, the scheduling weights are set as λ σ = 1.2, λ c = 0.6, λ r = 0.8, with β = 25 6 . The first term, λ σ (σ i + σ j ), prioritizes model pairs whose ratings remain uncertain. The second term, exp − |μ i − μ j | β , prioritizes pairs whose current estimated qualities are close, since comparisons between similarly rated models are more informative for refining the final ranking. The coverage bonus C(i,j)encourages the scheduler to select under-compared models, while the repeat penaltyR(i,j)discourages repeat- edly showing the same model in consecutive comparisons. All previously annotated pairs are excluded from the candi- date set, preventing duplicate comparisons. Convergence criterion The adaptive process is checked after every 10 newly annotated adaptive pairs. At each check- point, we compare the current model ranking with the ranking from the previous checkpoint using Spearman’s rank correla- tion coefficient: ρ = SRCC (rank k , rank k−1 ). A task is considered converged when all three pairwise SRCC values among the most recent three rankings exceed 0.95, and the mean uncertainty satisfies: 1 |M t | X m∈M t σ m ≤ 0.3. Once a task satisfies the convergence condition, no additional pairs are scheduled for that task. In short, the adaptive scheduler favors pairs that are still uncertain and whose current TrueSkill means are close. This makes the annotation process focus on comparisons that are most useful for stabilizing the final model ranking. A.5 Additional Experiments and Details for Leaderboards Construction Due to space constraints in the main paper, we provide additional implementation details for the RAVEN-Eval Leaderboards below. AIVGM Leaderboards We estimate the95%confidence interval of each AIVGM score through stratified bootstrap resampling. Specifically, we perform200independent boot- strap rounds, each containing2,000model pairs sampled with replacement. The sampling budget is distributed approx- imately uniformly across all tasks to avoid overrepresenting particular tasks. To keep the key settings, for every sampled pair, we retain the complete set of judgments from all LMM judges rather than resampling individual judge outputs. We then rerun the same tie-aware optimization for each replicate to obtain a new score for every AIVGM. The2.5th and97.5th percentiles of the resulting score distribution are reported as the lower and upper confidence bounds, respectively. All other leaderboard construction details follow the main paper. As for external online references, we provide the exact shared AIVGMs with A and Arena. in Tab.6. Judge Leaderboards To improve the robustness of the reported SRCC and KRCC values in the RAVEN-Eval- Judge Leaderboards, we combine the estimate from the fully connected comparison graph (complete graph) with bootstrap-based estimates. For each LMM judge, we first optimize one ranking from the complete graph and construct another200rankings using the bootstrap protocol described above. We compute the correlation value (SRCC and KRCC) between each LMM-based ranking and the reference rank- ing(including human or external rankings), and first aver- age the corresponding correlations over the200bootstrap replicates. The final reported value (the bootstrap-stabilized SRCC / KRCC) is then obtained by averaging this bootstrap mean with the correlation value derived from the complete graph, assigning each component a weight of1/2. The cor- relation results in Tabs. 2 and 4 of the main paper follow this bootstrap-stabilized protocol, which reduces sensitivity to a single complete-graph optimization. For LMMs that provide an explicit reasoning-effort option, we enable reasoning at the lowest supported level; otherwise, we use the default inference configuration. The8sampled frames from each video are provided at their native resolu- tion. Additional LeaderboardsWe separately report another 6 AIVGM leaderboards forK1T2V tasks,S2T2V tasks,D1 T2V task,all reasoning-based T2V and I2V tasks, and the aggregated I2V tasks, together with an overall leaderboard aggregating the T2V, KFT, and FLT categories. If a task is incompatible with a particular model, it is excluded from the computation of that model’s final mean and variance. The resulting rankings are presented in Tab. 7. Optimization Settings and Ablations The optimization hyperparameters are fixed across task categories and sum- marized in Tab.5. We further evaluate several variants of the proposed score-estimation method. First, we remove explicit tie modeling. Both confidently identified ties and uncertain judgments are represented by assigning equal probability mass to the two directional out- comes. We retain the confidence-aware (dynamicλ p ) setting. Under this setting, the optimization reduces to a standard Bradley-Terry maximum likelihood formulation. Second, we remove confidence-aware weighting and con- sider two alternatives. In the full-trust (λ p = 1) setting, confi- dent and uncertain judgments contribute equally to optimiza- tion. In the confident-only setting, uncertain judgments are discarded, while confident judgments retain full weight. The results are reported in Tab.9. The standard Bradley-Terry formulation performs rela- tively worse than our modified Davidson tie-aware estimator, confirming the importance of explicitly modeling ties. As- signing full trust to uncertain judgments also causes a clear performance drop. By contrast, the confident-only variant performs close to, but slightly below, our confidence-aware setting. These observations suggest that uncertain outputs mainly reflect unstable or unresolved directional preferences rather than genuine judgments that two videos have similar quality. Treating them as fully reliable therefore introduces noise, whereas adaptive downweighting preserves limited useful evidence without allowing uncertain outcomes to dom- inate score estimation. Effects of the Number of LMM Judges We observe an apparent difference between the main-paper results: individ- ual LMM judges generally achieve stronger agreement with the human reference on FLTs than on KFTs in Tab. 4 in main paper, whereas the three-judge evaluation in Tab. 2 performs better on KFTs than on FLTs. This discrepancy may partly re- sult from the random composition of the10annotated KFTs and10annotated FLTs. More importantly, we assume that the performance gain from combining multiple judges may appears substantially larger than a single judge. To further examine the effects of the Number of LMM judges, we conduct an additional ablation using all50human- annotated tasks and construct a unified ranking without sepa- rating T2V and I2V tasks. For AIVGMs that do not support FLTs, unavailable FLT scores are omitted during aggregation. As in the main evaluation, the final score is computed from the mean task-level score with a standard-deviation penalty. We estimate rankings usingM = 1,M = 2, andM = 3 judges, evaluate every possible judge combination under each ensemble size, and report bootstrap-stabilized SRCC values against the human reference. The results are presented in Tab 10. We further select the top-7 LMM judges from the T2V Judge Leaderboard and evaluate ensemble sizes fromM = 1 toM = 7. For eachM, we enumerate all possible judge com- bination of sizeMand report their mean bootstrap-stabilized SRCC on all the50tasks with human reference. The re- sulting curve (Fig.4) shows a clear gain when increasing the ensemble size to three judges and a sustained upward trend as more capable judges are incorporated. These results indicate that, when the candidate judges are individually reli- able, increasing the ensemble size generally improves overall evaluation accuracy and reduces sensitivity to any particular judge combination. This finding also provides a plausible explanation for the stronger KFT performance observed under the three-judge setting in the main paper. More details about the anchor-based model inser- tion paradigm The experimental complexity reduction shown in the main paper (K = 1:34.2%,K = 2:21.1%) is derived from below. The theoretical reduction is based solely on the number of pairwise comparisons is31.6%for K = 1and18.4%forK = 2. Specifically, a fully connected graph over20models contains 20 2 = 190model pairs. After retaining the fully connected graph among the15existing models, the anchor-based strategy requires105+5×5 = 130 comparisons forK = 1and105 + 5× 10 = 155compar- isons forK = 2. The corresponding reductions are therefore (190− 130)/190 = 31.6%and(190− 155)/190 = 18.4%, respectively. In our implementation, we further account for the reduction in actual LMM invocation rounds. The anchor- based insertion procedure avoids additional LMM API call- ing rounds, which can be roughly treated as5pairwise- comparison-equivalent calls. After this implementation-level adjustment, the effective numbers of saved calls become (190− 130) + 5 = 65and(190− 155) + 5 = 40, yielding effective cost reductions of65/190 = 34.2%forK = 1and 40/190 = 21.1%forK = 2. Thus,31.6%and18.4%are the reductions derived purely from graph sparsification, whereas 34.2%and21.1%additionally reflect the saved LMM API calling rounds in the actual evaluation pipeline. As additional models are incorporated, the relative re- duction in pairwise evaluation cost is expected to become more pronounced, because the cost of fully connected evalu- ation grows quadratically with the model-pool size, whereas anchor-based insertion requires comparisons only against a limited anchor set. However, continued expansion of the model pool may weaken the reliability and representative- ness of the existing anchors, particularly when newly added models alter the capability distribution or introduce previ- ously uncovered generation characteristics. The anchor set should therefore be periodically refreshed by conducting full 1234567 M 0.850 0.860 0.870 Bootstrapping-stabilized SRCC 0.847 0.858 0.862 0.859 0.868 0.874 0.872 Figure 4: The trend of bootstrapping-stabilized SRCC as a function of M . pairwise evaluation on selected updates or by dynamically expanding the comparison graph to identify new reliable an- chors. Developing an adaptive anchor-maintenance strategy that balances evaluation efficiency against ranking reliability constitutes an important direction for future work. Computational Analysis Computational Analysis of the whole estimation process is shown in Tab.8. A.6 Benchmark Cases T2V Prompts and Task-Specific Rubrics Figures 5–10 present the complete prompts and task-specific rubrics of the 48 selected T2V benchmark cases. The cases are organized according to the six retained configurations of scene com- plexityS, dynamic interaction complexityD, and knowledge requirementK. Ten cases are additionally marked as rea- soning tasks. To avoid redundant text without omitting any evaluation criterion, the shared basic rubric is presented once at the top of each page and applies to all cases on that page. Each individual case retains its complete prompt, primary rubric, and additional task-specific criteria. T2V Case Examples We present representative T2V case examples for the six benchmark difficulty configurations and an additional reasoning-focused case in Figs. 11–17. Each case includes the complete English prompt, the task-specific primary rubric, and the outputs of six evaluated video gen- eration models. For every model, eight frames are sampled at fixed relative temporal positions and arranged in a2× 4 grid. The model selection varies across cases and collectively covers all 20 evaluated video generation models. I2V case Examples We provide representative case ex- amples for the two I2V task categories, i.e., first–last-frame transition (FLT) and keyframe continuation (KFT), as shown in Figs. 18–25. For each category, we include reasoning and non-reasoning cases collected from both YouTube and ASMR videos. Each case presents six different evaluated models, with eight uniformly sampled output frames arranged in a 2× 4grid for each model. For FLT, only the first and last input frames are displayed, whereas KFT shows only the starting keyframe. The model selection varies across cases and collectively covers all 20 evaluated video generation models. Prompt Word Clouds Figure 30 summarizes the lexical distributions of the prompts used in the three evaluation tasks. We aggregate all available prompts for T2V, FLT, and KFT, translate the original Chinese text into English using an of- fline machine-translation model, and then apply lowercasing, tokenization, stop-word removal, and basic singular-form nor- malization. The resulting corpora contain 150 T2V prompts, 50 FLT prompts, and 50 KFT prompts. Word size is propor- tional to corpus-level frequency after preprocessing. The T2V prompts emphasize motion, interaction, spatial relationships, stability, and material or physical properties, whereas the FLT and KFT prompts contain more concrete objects, materials, and state changes associated with temporally constrained video generation. Failure-Case Analysis We conduct a failure-case analy- sis for each task category. Since the RAVEN-Eval AIVGM Leaderboards exhibit strong correlations with established ex- ternal human-annotation-based leaderboards, we use the cor- responding category-level RAVEN-Eval ranking as a proxy reference. This choice is necessary because the external plat- forms do not provide separate rankings for KFT and FLT. T2V and KFT tasks evaluated by at least18AIVGMs and FLT tasks evaluated by at least14AIVGMs. For each eligible task, we compute the SRCC between its task-specific model ranking and the corresponding category-level ranking over their shared models. We then select the three tasks with the lowest SRCC in each category, yielding nine highly probable failure cases, as shown in Figs. 26,27,28. Although these tasks may contain task-specific biases that produce unusually divergent rankings, they remain informative for diagnosing systematic weaknesses in the evaluation framework. We fur- ther employ LMM-assisted qualitative analysis to identify their underlying failure modes. Our preliminary analysis reveals distinct patterns across task categories. For T2V, the primary failure mode arises from an imbalance between fine-grained requirement sat- isfaction and holistic scene quality. In several inconsistent cases, the evaluation places excessive emphasis on localized details, such as object counts or specific poses, while under- weighting the superior overall composition, coherence, and visual appeal produced by some strong models. This imbal- ance can lead to disagreement between the automated ranking and the proxy human-oriented reference. For KFT and FLT, the main difficulties involve highly specialized and com- plex scenarios, as well as the assessment of temporally continuous actions and fine-grained motion evolution. Be- cause the current judging pipeline relies primarily on sparsely sampled video frames, it has an inherent limitation in evaluat- ing motion continuity, action completeness, and intermediate state transitions, which can reduce ranking accuracy. These findings motivate two directions for future work: better bal- ancing holistic scene quality against local task-specific details in rubric design, and augmenting the primary LMM judge with complementary temporal-analysis models or external tools for tasks involving complex dynamics and specialized phenomena. Win-rate HeatmapsWe also provide the pairwise win-rate heatmaps for the T2V, KFT, and FLT tasks, computed from all 3 LMM judge results. It is shown in Fig.29. HyperparameterValueDefinition OptimizerL-BFGS-BQuasi-Newton optimizer used to minimize the Davidson negative log-likelihood. Learning rateN/ANo fixed learning rate is used; step sizes are selected automatically by L-BFGS-B line search. maxiter1, 000Maximum number of optimizer iterations allowed for each task-level fit. ftol10 −11 Convergence tolerance; optimization may stop when the relative improvement in objective value becomes sufficiently small. gtol10 −6 Gradient convergence tolerance; optimization may stop when the projected gradient norm is sufficiently small. Initialization of s i 0Initial skill score for every model in each task-level Davidson fit. Initialization of logη0Initial log tie parameter; equivalent to initializing η = exp(logη) = 1. α10 −3 L2 regularization coefficient on model skills, applied as α P i s 2 i . λ g 10.0An implementation-level gauge-fixing penalty corresponding to the the term P N i=1 s i = 0 in the main paper. β10 −4 L2 regularization coefficient on the log tie parameter, applied as β(logη) 2 . ηlearnedDavidson tie parameter controlling tie probability, parameterized as η = exp(logη) to ensure positivity. ̄stask meanMean task-level skill, ̄s = 1 N P i s i , used in the gauge regularization term. Bootstrap samples per round2000Number of model pairs sampled with replacement in each bootstrap round. Bootstrap rounds200Number of independent bootstrap leaderboard reconstructions used for confidence-interval estimation. CI level95%Percentile interval computed from the 2.5th and 97.5th percentiles of bootstrap Elo samples. Table 5: Hyperparameters used in the Davidson tie-aware Bradley–Terry optimization and leaderboard aggregation. Task set ReferenceShared modelsCount T2VAAseedance2, grokimagevideo, kling3pro, happyhorsev1, kling3omni, veo3.1fast, veo3.1lite, pixversev6, kling3std, wan2.6, pixversev5.5, seedance1.5pro, ltx23fast, wan2.7, ltx23pro, hailuo2.3, ltx2pro, ltx2fast18 T2VArena.AIseedance2, grokimagevideo, happyhorsev1, wan2.6, pixversev5.5, seedance1.5pro, wan2.7, hailuo2.3, pixversev6,veo3.1-lite, veo3.1-fast,kling3.0-pro,kling-3.0-omni,kling-3.0-std14 KFTAAgrok-image-video, seedance2, kling3.0-pro, wan2.7, kling-3.0-omni, kling-3.0-std, happyhorsev1, pixversev6, wan2.6, veo3.1-lite, seedance1.5-pro, hailuo2.3, veo3.1-fast, pixversev5.5, ltx23pro, ltx2pro, ltx2fast, ltx23fast18 KFTArena.AIseedance2, grokimagevideo, happyhorsev1, wan2.6, pixversev5.5, seedance1.5pro, wan2.7, hailuo2.3, pixversev6,veo3.1-lite, veo3.1-fast,kling3.0-pro,kling-3.0-omni,kling-3.0-std14 FLTAAseedance2, kling3.0-pro, wan2.7, kling-3.0-omni, kling-3.0-std, pixversev6, veo3.1-lite, seedance1.5-pro, veo3.1-fast, pixversev5.5, ltx23pro, ltx23fast12 FLTArena.AIseedance2, pixversev6, kling3.0-pro, wan2.7, seedance1.5-pro,veo3.1-lite, veo3.1-fast,kling-3.0-omni,pixversev5.5,kling-3.0-std10 Table 6: Overlapping models between our leaderboard and external reference leaderboards for alignment computation. CategoryReasoningS2D1K1I2V overallOverall AIVGMRankScoreRankScoreRankScoreRankScoreRankScoreRankScore Seedance 21999.72854.61852.61928.81978.91904.1 Seedance 2 Fast2930.76723.56691.610652.45809.14774.5 Grok-Imagine-Video3834.55759.12799.04803.32947.02857.1 Kling 3.0 Omni4752.78699.05693.22865.87787.05762.3 PixVerse C15741.39557.18600.012579.53840.67718.0 Kling 3.0 Pro6718.412485.04750.08697.44825.63799.6 Kling 3.0 Std7654.42072.411489.47721.78782.48677.2 Wan 2.78589.94812.615284.99655.06791.910573.8 Veo 3.1 Fast9565.23839.67673.714558.115323.113447.6 PixVerse V610533.010555.110531.411644.910646.09601.2 Veo 3.1 Lite11483.71966.49590.56764.412478.612524.7 HappyHorse V1.012414.37718.83768.85797.29690.16730.7 Wan 2.613413.614377.713468.513565.311638.311556.1 PixVerse V5.514270.615297.912469.618285.016294.215357.3 Seedance 1.5 Pro15257.319218.314361.916456.413421.714401.8 LTX 2.3 Pro1687.216284.818-31.415468.917-52.417-61.4 LTX 2.0 Pro1738.613437.717-29.317338.418-323.218-160.1 LTX 2.0 Fast1822.017259.920-117.81915.419-336.820-232.4 LTX 2.3 Fast19-10.311534.516276.43834.320-382.919-208.4 Hailuo 2.320-259.918220.019-62.820-169.914400.116128.7 Table 7: Sub-leaderboards display. Each leaderboard reports only rank and Elo-style score. ItemSetting / ValueDescription CPU2× Intel Xeon Platinum 855896 physical cores / 192 threads in total. Memory2.0 TiB RAMSystem memory available on the evaluation machine. Operating modeCPU-onlyPairwise likelihood optimization, Elo aggregation, and bootstrap sampling were run on CPU. Python3.11.10Runtime environment used for the benchmark. NumPy2.2.6Used for array operations, random bootstrap sampling, mean/std computation, and percentile CI computation. SciPy1.16.0Used for task-level Davidson optimization via scipy.optimize.minimize. OptimizerL-BFGS-BQuasi-Newton optimizer used for the Davidson tie-aware Bradley–Terry objective. Task set150 T2V + 50 KFT + 50 FLTCombined benchmark set, 250 task-level pairwise optimization problems in total. Bootstrap samples2000Number of resampled leaderboard estimates used for confidence interval estimation per round. Random seed20260631Seed used for bootstrap resampling. Task-level fits250 / 250 successfulAll task-level Davidson optimizations converged successfully in the measured run. Cached-score leaderboard latency8.15 sTime for loading fitted task scores, computing the Elo leaderboard, and running 2000 bootstrap samples. End-to-end optimization latency60.86 sTime for parsing pairwise JSONs, fitting all task-wise Davidson models, computing the Elo leaderboard, and running all bootstrap rounds samples. Table 8: Computational setup and measured latency for producing the combined T2V/KFT/FLT leaderboard. The cached-score latency starts from previously fitted task-level Davidson scores, while the end-to-end latency includes re-fitting all task-level Davidson models from pairwise JSON files. CategoryT2VI2V SettingHuman A Arena. H. (KFT) H. (FLT) A Arena. Standard BT-MLE 0.855 0.815 0.8010.8710.813 0.843 0.703 Confident-only0.867 0.825 0.8120.8940.817 0.846 0.712 Full-trust0.832 0.799 0.7870.8840.785 0.830 0.690 Ours0.872 0.835 0.8100.9030.821 0.851 0.714 Table 9: Ablation results (SRCC) under different optimization settings. The term “H.” is short for “Human.” The highest and second-highest values in each column are shown in bold and italics, respectively. LMM JudgesHuman Reference (50 Tasks) GPT-5.40.864 Gemini 3.1 FL0.817 Claude Sonnet 4.60.856 GPT-5.4 + Gemini 3.1 FL0.837 GPT-5.4 + Claude Sonnet 4.60.870 Claude Sonnet 4.6 + Gemini 3.1 FL0.848 All0.874 Table 10: Ablation results (SRCC) on the effect of the number and composition of LMM judges. The highest and second- highest values are shown in bold and italics, respectively. A.7 Limitations Although this work provides detailed algorithmic analyses and extensive validation of its key components, several lim- itations remain. Due to the substantial costs of human an- notation, proprietary high-performance LMM judges, and closed-source AIVGMs, we are unable to obtain human an- notations for all tasks or evaluate the full set of250tasks with a broader pool of capable LMM judges. The current AIVGM evaluation is also limited to20models. Consequently, the ev- idence for the large-scale scalability of RAVEN-Eval and its potential for LMM judge model harnessing still need further research. Extending the benchmark to more generation mod- els, judge ensembles, and human-validated tasks is therefore a central direction of our future work. S1–D1–K0 Moderate semantic complexity · Rich detail · No external knowledge 9 tasks SHARED BASIC RUBRIC —applies to every task on this page (1)All static objects and background objects must remain visually consistent throughout the video. (2)The video should appear realistic. In particular, when material reactions or object interactions are involved, and unless the prompt states otherwise, dynamic objects should follow real-world physical, material, and chemical laws and exhibit logically consistent changes. Material quantities should be conserved as strictly as possible. (3)For videos containing motion, each action should be complete and reasonable. Unless otherwise specified, all actions and reactions should remain continuous and natural, without obvious abrupt changes. (4)If a task-specific instruction conflicts with the three requirements above, such as a scene that intentionally disregards physical laws, the task-specific instruction takes precedence; the three shared requirements are secondary. 01. T2V_L1_D1_K0_0001 Prompt. One child throws a yellow ball to another child, who catches it and steps slightly backward. Use a fixed camera. Emphasize direct interaction between a small number of subjects, preserve natural continuity across the throw and catch, and make the action order and rhythm of both participants clear. Primary rubric. (1)The ball travels toward the second child along a continuous trajectory. (2)The second child successfully catches the ball and reacts by stepping slightly backward. (3)The full sequence of throwing, flight, and catching is complete and continuous. (4)The spatial relationship between the two children and the outcome of the interaction remain clearly visible. Additional criteria. (1)The visual treatment should emphasize direct interaction between a small number of subjects. (2)The yellow ball should serve as a visually salient interaction focus in the two-person throwing-and-catching sequence. 02. T2V_L1_D1_K0_0003 Prompt. Two cyclists approach each other from opposite directions on a path and pass one another. At the moment they meet, both make a slight handlebar adjustment. Keep the background simple and the camera fixed. Emphasize the path relationship and dynamic connection between the two subjects, with a clear action order and natural rhythm. Primary rubric. (1)The two cyclists approach from opposite directions. (2)Both cyclists maintain continuous motion before they meet. (3)Both make a slight handlebar adjustment at the moment of passing. (4)They successfully pass each other, and their path relationship remains clear. (5)Their motion trajectories remain continuous and stable throughout. Additional criteria. (1)The visual treatment should emphasize the path relationship and dynamic connection between the two subjects. (2)A simple background should foreground the subjects and preserve a restrained visual quality. 03. T2V_L1_D1_K0_0004 Prompt. One person opens an umbrella and then walks forward with another person. The second person moves under the umbrella and adjusts their pace. Use a fixed camera. Emphasize the interaction created when a small number of subjects share an action rhythm, and keep the action order and rhythmic transitions clear and natural. Primary rubric. (1)The first person completes the umbrella-opening action before walking. (2)The two people walk forward together. (3)The second person moves under the umbrella and adjusts their pace. (4)The shared-umbrella relationship remains clearly visible. Additional criteria. (1)The two participants' motion rhythms and positional changes should be natural and continuous. (2)The visual treatment should emphasize the interaction created by the subjects' shared action rhythm. (3)The background atmosphere, supporting details, and visual rhythm of the moving scene should remain consistent with the prompt. 04. T2V_L1_D1_K0_0007 Prompt. Two joggers approach in parallel on a path and exchange a relay baton. At the handoff, both slightly adjust the direction of their arms. Keep the background simple and the camera fixed. Emphasize their path relationship and dynamic connection, maintain reasonable changes in distance, and make the interaction outcome clearly visible. Actions should be continuous, directionally clear, and moderate in amplitude, without abrupt jumps. Use a stable medium shot with clear action paths and a readable interaction, focusing on two-subject coordination, credible contact, and motion continuity. Primary rubric. (1)The main scene shows two joggers exchanging a relay baton. (2)Two principal interacting subjects are present. (3)The action path is continuous, with clear start and end states. Additional criteria. (1)The distance or contact relationship between the subjects is reasonable. (2)The camera remains stable and the interacting subjects do not drift. (3)The visual treatment should show the subjects' path relationship and dynamic connection, preserve reasonable distance changes, and clearly present the outcome. 05. T2V_L1_D1_K0_0009 Prompt. A person leads a small dog around a yellow traffic cone. After the dog turns, the owner steps slightly backward. Use a fixed camera, emphasize direct interaction between the two subjects, preserve natural continuity in the guiding action, and keep the main motion path stable without sudden displacement of either subject. Actions should be continuous, directionally clear, and moderate in amplitude, without abrupt jumps. Use a stable medium shot with clear action paths and a readable interaction, focusing on coordination, credible contact, and motion continuity. Primary rubric. (1)The main scene shows an owner guiding a small dog around an obstacle. (2)Two principal interacting subjects are present. (3)The action path is continuous, with clear start and end states. Additional criteria. (1)The distance or contact relationship between the subjects is reasonable. (2)The camera remains stable and the interacting subjects do not drift. (3)The visual treatment should emphasize direct interaction between the two subjects. 06. T2V_L1_D1_K0_0010 Prompt. Two people jointly lift a small cardboard box and place it on a nearby table. Their hands, the box, and their body movements should remain continuous and natural. Use a fixed camera, emphasize dynamic interaction between a small number of subjects, and keep the main motion path stable without sudden displacement of the participants or box. Actions should be continuous, directionally clear, and moderate in amplitude, without abrupt jumps. Use a stable medium shot with clear action paths and a readable interaction, focusing on coordination, credible contact, and motion continuity. Primary rubric. (1)The main scene shows two people jointly moving a small cardboard box. (2)Two principal interacting subjects are present. (3)The action path is continuous, with clear start and end states. Additional criteria. (1)The distance or contact relationship between the subjects is reasonable. (2)The camera remains stable and the interacting subjects do not drift. (3)The visual treatment should emphasize dynamic interaction between a small number of subjects. 07. T2V_L1_D1_K0_0012 Prompt. One person steadies a short ladder and helps another person slowly step down from the final rung while the second person adjusts their footing. Use a fixed camera, emphasize the interaction created by a shared action rhythm, and keep the main motion path stable without sudden displacement. Actions should be continuous, directionally clear, and moderate in amplitude, without abrupt jumps. Use a stable medium shot with clear action paths and a readable interaction, focusing on coordination, credible contact, and motion continuity. Primary rubric. (1)The main scene shows one person steadying the ladder while another descends. (2)Two principal interacting subjects are present. (3)The action path is continuous, with clear start and end states. Additional criteria. (1)The distance or contact relationship between the subjects is reasonable. (2)The camera remains stable and the interacting subjects do not drift. (3)The visual treatment should emphasize the interaction created by the subjects' shared action rhythm. 08. T2V_L1_D1_K0_0013 Prompt. One person tilts a teapot to pour tea into another person's cup; the second person steadies the cup and then steps slightly backward. Use a fixed camera, emphasize direct interaction between the two subjects, preserve natural continuity before and after pouring, and fully show the key contact point and transfer moment. Actions should be continuous, directionally clear, and moderate in amplitude, without abrupt jumps. Use a stable medium shot with clear action paths and a readable interaction, focusing on coordination, credible contact, and motion continuity. Primary rubric. (1)The main scene shows one person pouring tea for another. (2)Two principal interacting subjects are present. (3)The action path is continuous, with clear start and end states. Additional criteria. (1)The distance or contact relationship between the subjects is reasonable. (2)The camera remains stable and the interacting subjects do not drift. (3)The visual treatment should emphasize direct interaction between the two subjects. 09. T2V_L1_D1_K0_0019 Prompt. Two people stand on opposite sides of a table and gently pull a tablecloth flat at the same time. During the pull, both slightly adjust the direction of their arms. Keep the background simple and the camera fixed. Emphasize their path relationship and dynamic connection, including natural reactions and subtle posture adjustments. Actions should be continuous, directionally clear, and moderate in amplitude, without abrupt jumps. Use a stable medium shot with clear action paths and a readable interaction, focusing on coordination, credible contact, and motion continuity. Primary rubric. (1)The main scene shows two people jointly pulling a tablecloth flat. (2)Two principal interacting subjects are present. (3)The action path is continuous, with clear start and end states. Additional criteria. (1)The distance or contact relationship between the subjects is reasonable. (2)The camera remains stable and the interacting subjects do not drift. (3)The visual treatment should emphasize the subjects' path relationship and dynamic connection, including natural reactions and subtle posture adjustments. RAVEN-Eval supplementary material · Full T2V prompts and rubrics · Page 1/6 Figure 5: Complete T2V benchmark cases forS=1,D=1, andK=0: moderate semantic complexity, rich detail, and no external knowledge requirement. S1–D1–K1 Moderate semantic complexity · Rich detail · Knowledge required 8 tasks SHARED BASIC RUBRIC —applies to every task on this page (1)All static objects and background objects must remain visually consistent throughout the video. (2)The video should appear realistic. In particular, when material reactions or object interactions are involved, and unless the prompt states otherwise, dynamic objects should follow real-world physical, material, and chemical laws and exhibit logically consistent changes. Material quantities should be conserved as strictly as possible. (3)For videos containing motion, each action should be complete and reasonable. Unless otherwise specified, all actions and reactions should remain continuous and natural, without obvious abrupt changes. (4)If a task-specific instruction conflicts with the three requirements above, such as a scene that intentionally disregards physical laws, the task-specific instruction takes precedence; the three shared requirements are secondary. 10. T2V_L1_D1_K1_0001 Prompt. Two investigators in long trench coats pass one another in a neon-lit alley on a rainy night. One raises a hand against the rain while the other looks back. Use a Blade Runner-style cyberpunk film-noir aesthetic with wet-ground reflections, cyan-purple neon, high-contrast shadows, layered mist, and a low camera angle. Motion should remain continuous and natural, with emphasis on cinematographic style and character-path relationships. Primary rubric. (1)Two investigators wearing long trench coats are visible. (2)They pass each other in a neon-lit alley on a rainy night with continuous motion. (3)The overall scene presents a Blade Runner-style cyberpunk film-noir aesthetic. Additional criteria. (1)Wet-ground reflections, mist, and high-contrast shadows are clearly visible. (2)The low camera angle and cyan-purple neon palette remain consistent. (3)The stylistic treatment includes wet-ground reflections, cyan-purple neon, high-contrast shadows, layered mist, and a low camera angle. (4)The visual treatment should convey cinematic photography and the characters' path relationship. 11. T2V_L1_D1_K1_0002 Prompt. Two modified off-road vehicles race side by side along a dusty section near a desert canyon. The lead vehicle makes a slight drift while the following vehicle avoids loose rocks. Use a Mad Max: Fury Road-style desert action- film aesthetic with high-contrast orange-yellow color, a handheld chase-camera feeling, airborne dust, exaggerated mechanical silhouettes, and a strong sense of speed. Emphasize the dynamic relationship between the vehicles and the cinematographic rhythm. Primary rubric. (1)Two modified off-road vehicles are visible. (2)The vehicles move side by side through a desert canyon or dusty road section. (3)The scene presents a Mad Max: Fury Road-style desert action-film aesthetic. Additional criteria. (1)The high-contrast orange-yellow palette and dust particles are prominent. (2)The changing front-to-back positions and sense of speed are clear. (3)The style includes a handheld chase-camera feeling, dust particles, exaggerated mechanical silhouettes, and strong speed. (4)The visual treatment should convey the dynamic relationship between the two vehicles and the cinematographic rhythm. 12. T2V_L1_D1_K1_0003 Prompt. One astronaut slowly walks toward a giant black monolith while a second astronaut stands beside a rover and observes. Use a 2001: A Space Odyssey-style austere science-fiction aesthetic with symmetrical composition, minimalist space, slow movement, cool white light, and extensive silent negative space. Motion should be subtle but continuous, emphasizing cinematic composition and a solemn atmosphere. Primary rubric. (1)The frame contains two astronauts and one giant black monolith. (2)One astronaut slowly approaches the monolith while the other observes beside the rover. (3)The scene presents an austere, minimalist science-fiction aesthetic in the style of 2001: A Space Odyssey. Additional criteria. (1)Symmetrical composition, cool white light, and extensive negative space are evident. (2)The movement remains slow and the spatial relationships remain stable. (3)The style includes symmetrical composition, minimalist space, slow movement, cool white light, and extensive silent negative space. (4)The visual treatment should convey cinematic composition and a solemn atmosphere. 13. T2V_L1_D1_K1_0005 Prompt. An explorer and an assistant slowly step backward in misty forest ruins as a huge shadow sweeps across the foreground; both look upward at the same time. Use a Jurassic Park-style Spielberg adventure-film aesthetic with a low-angle upward shot, warm flashlight illumination, layered mist, and staging that first builds suspense and then reveals a silhouette. Emphasize both characters' reactions and the adventure-film atmosphere. Primary rubric. (1)An explorer and an assistant are visible. (2)They step backward and look upward in misty forest ruins. (3)The scene presents a Jurassic Park-style adventure-film aesthetic. Additional criteria. (1)The low-angle shot, warm flashlight illumination, and layered mist are clearly visible. (2)A large shadow or silhouette creates suspense without making the scene confusing. (3)The style includes low-angle framing, warm flashlight light, layered mist, and suspense-first silhouette revelation. (4)The visual treatment should convey both characters' reactions and an adventure-film atmosphere. 14. T2V_L1_D1_K1_0006 [REASONING] Prompt. A hydraulic press slowly descends toward a rubber duck, with the press head, duck, and workbench visible together. The complete process should show the press head moving continuously downward, making contact with the target, and producing subsequent changes that follow the natural response of a material under force. The interaction path, final state, and professional scene logic should remain clear and explainable. Primary rubric. (1)After the press head contacts the rubber duck, the contact area should flatten first and deform visibly along the vertical force direction. (2)The duck should exhibit elastic-material behavior, such as body compression, lateral bulging, or slight displacement of the head or beak. (3)As the press continues downward, the deformation should deepen gradually; the duck should not suddenly shatter like a brittle material. (4)Contact among the press head, rubber duck, and workbench should remain stable, without interpenetration, floating, or an incorrect force direction. (5)The final state should show a physically reasonable elastic deformation caused by sustained compression. Additional criteria. None specified. 15. T2V_L1_D1_K1_0009 [REASONING] Prompt. A beam of light travels from a light source through a prism and onto a screen behind it. The frame contains at least the light source, prism, and screen. The light should interact continuously with these objects, and the resulting effect should follow common optical laws. Emphasize the professional dynamic relationship among a small number of objects while keeping the interaction path, final state, and scene logic clear and explainable. Primary rubric. (1)Before entering the prism, the light beam should follow a relatively concentrated propagation path. (2)After passing through the prism, the screen should show a positional change, displaced light spot, or dispersion effect caused by refraction. (3)If dispersion appears, the light on the screen should separate into a continuous or nearly continuous multicolored band rather than unrelated colors flashing randomly. (4)The light source, prism, optical path, and screen response should remain spatially aligned and causally connected. (5)The optical effect should remain stable with the light path and must not jump abruptly or appear unrelated to the prism's position. Additional criteria. None specified. 16. T2V_L1_D1_K1_0008 [REASONING] Prompt. A double-pendulum-ball experimental apparatus shows both balls and the support structure in the same frame. During the experiment, one ball is released and contacts the other. The interaction and subsequent changes should follow common mechanical laws. Emphasize professional interaction between a small number of objects while keeping the interaction path, final state, and professional scene logic clear and explainable. Primary rubric. (1)The released pendulum ball should swing along an arc toward the other ball, following gravity-driven pendulum motion. (2)A clear collision moment or contact position should be visible when the two balls meet. (3)After the collision, the struck ball should begin moving, while the released ball slows, pauses, or swings backward. (4)The changes in the two balls' motion should demonstrate momentum or energy transfer rather than unrelated simultaneous motion. (5)The balls, strings, and support should remain physically constrained: the balls must not detach, pass through each other, or suddenly change size. Additional criteria. None specified. 17. T2V_L1_D1_K1_0010 [REASONING] Prompt. A pneumatic experimental apparatus shows a main container, a pressure gauge, and a flexible component in the same frame. During the experiment, the internal state of the apparatus changes gradually and the components remain causally connected. Subsequent changes should follow common pressure behavior. Emphasize the interaction logic among a small number of professional objects while keeping the interaction path, final state, and professional scene logic clear and explainable. Primary rubric. (1)During pressurization or depressurization, the pressure-gauge needle, reading, or indicator state should change visibly with the internal pressure. (2)The flexible component should respond correspondingly by expanding, contracting, bulging, collapsing, or recovering its shape. (3)The pressure-gauge change and flexible-component response should be temporally correlated rather than contradictory or causally unrelated. (4)The main container, pressure gauge, and flexible component should maintain a clear physical connection or functional relationship. (5)The full process should form a continuous physical chain from pressure change to a visible mechanical or deformation response. Additional criteria. None specified. RAVEN-Eval supplementary material · Full T2V prompts and rubrics · Page 2/6 Figure 6: Complete T2V benchmark cases forS=1,D=1, andK=1: moderate semantic complexity, rich detail, and an external knowledge requirement. Reasoning cases are explicitly marked in the figure. S2–D0–K0 High semantic complexity · Standard detail · No external knowledge 9 tasks SHARED BASIC RUBRIC —applies to every task on this page (1)All static objects and background objects must remain visually consistent throughout the video. (2)The video should appear realistic. In particular, when material reactions or object interactions are involved, and unless the prompt states otherwise, dynamic objects should follow real-world physical, material, and chemical laws and exhibit logically consistent changes. Material quantities should be conserved as strictly as possible. (3)For videos containing motion, each action should be complete and reasonable. Unless otherwise specified, all actions and reactions should remain continuous and natural, without obvious abrupt changes. (4)If a task-specific instruction conflicts with the three requirements above, such as a scene that intentionally disregards physical laws, the task-specific instruction takes precedence; the three shared requirements are secondary. 18. T2V_L2_D0_K0_0001 Prompt. A crowded bookstore interior contains at least 6 people, 4 shelving areas, and more than 30 visible books. People remain reading or standing, with only very small page turns, eye movements, or posture adjustments, using a fixed wide-angle view. Requirements: (1) customers browse or read in the foreground, shelves and several customers occupy the middle ground, and shelves extend into the background to show depth; (2) natural occlusion is clear, such as foreground people partially blocking background shelves, while relative positions remain stable; (3) a slightly crooked stack of books and a partly exposed handwritten recommendation card are clearly visible; (4) subject orientations and reading directions vary naturally rather than forming neat rows; (5) use documentary photography with a clear light direction, coordinated color across depth layers, and strong but non-rigid visual order; (6) motion is limited to stillness or very small natural changes, without obvious physical interaction or large page-turning actions. Use a fixed camera, natural daylight or soft illumination, and deep depth of field, emphasizing spatial depth, visual order, and photographic aesthetics in a complex composition. Primary rubric. (1)The bookstore contains at least 6 people, 4 shelving areas, and more than 30 visible books; all quantity requirements must be met. Additional criteria. (1)Foreground customers browse or read, the middle ground mixes several customers with shelves, and background shelves extend into the distance to create clear depth. (2)Foreground-background occlusion, light direction, and overall color treatment remain realistic and consistent. (3)The photographic treatment includes stable composition, natural daylight or soft illumination, and deep depth of field. (4)The visual treatment should convey spatial depth, visual order, and photographic aesthetics in a complex composition. 19. T2V_L2_D0_K0_0002 Prompt. An indoor market contains at least 8 customers, 3 stalls, and 12 groups of goods. People remain mostly still, making only very small observational or positional adjustments. Use a fixed camera with clear foreground, middle-ground, and background layers. A slightly reflective price sign and the edge of a hanging shopping bag must appear. Subject distribution should be balanced, with at least one natural occlusion between foreground and background. Motion is restricted to stillness or subtle observation, breathing, page turning, or stance adjustment, with no obvious physical interaction. Use a fixed wide-angle or mildly compressed telephoto documentary style, emphasizing layered depth, natural occlusion, balanced distribution, unified color, clear lighting direction, and visual order in a complex scene. Primary rubric. (1)The indoor market contains at least 8 customers, 3 stalls, and 12 groups of goods with balanced subject distribution. Additional criteria. (1)The market retains a documentary photographic style with unified color and a clear lighting direction. (2)The wide-angle view supports readable multi-subject spatial depth. (3)The visual treatment should convey layered depth, natural occlusion, balanced subject distribution, unified color, clear lighting direction, and visual order. (4)Supporting details include a slightly reflective price sign and the edge of a hanging shopping bag. 20. T2V_L2_D0_K0_0004 Prompt. In an open-plan office, several people are distributed across different desks in working or conversational poses. Computers, papers, and cups are arranged naturally. Use a fixed camera and emphasize the realism of a static group scene. A horizontally placed pen and the slightly tilted corner of a sticky note must appear. Subject distribution should be balanced, with at least one natural foreground-background occlusion. Motion is restricted to stillness or subtle observation, breathing, page turning, or stance adjustment, without obvious physical interaction. Use a fixed wide-angle or mildly compressed telephoto documentary style, emphasizing layered depth, natural occlusion, balanced distribution, unified color, clear lighting direction, and visual order. Primary rubric. (1)A horizontally placed pen and the slightly tilted corner of a sticky note are clearly visible. Additional criteria. (1)Several people work or converse across the open office, while computers, papers, and cups are naturally distributed. (2)The office has realistic spatial layering, with unified color and a clear light direction. (3)The wide-angle view supports readable multi-subject spatial depth. (4)The visual treatment should convey the realism of a static group scene. 21. T2V_L2_D0_K0_0005 Prompt. A train-station waiting hall contains at least 7 passengers, 3 rows of seats, multiple suitcases, and 2 information screens. People wait or look down at tickets, making only very small page turns or posture adjustments. Use a fixed camera and emphasize a natural multi-subject static composition. A wheeled suitcase with an old luggage tag and a partly exposed paper ticket must appear. Maintain reasonable spacing among the main people and keep the aisle or table edge readable. Multiple distinct subjects or objects must be visible with clear spatial relationships. Use a fixed camera or slow stable shot, relatively deep depth of field, and clear foreground-middle-background layering, emphasizing multi-subject layout, occlusion, and complex-composition stability. Primary rubric. (1)At least 7 passengers and multiple groups of waiting-hall seats are visible. (2)Suitcases, information screens, and the paper-ticket element are clear. Additional criteria. (1)Aisle edges and seat arrangement remain readable. (2)Foreground, middle-ground, and background layers remain stable. (3)The photographic treatment includes stable composition or a slow stable shot, relatively deep depth of field, and clear depth layering. (4)The visual treatment should convey a natural multi-subject static composition. 22. T2V_L2_D0_K0_0009 Prompt. A living room during a move contains at least 6 people, 4 groups of cardboard boxes, and multiple pieces of furniture. People remain in packing or inventory-checking poses, with only very small label-applying movements. Use a fixed camera and emphasize a natural multi-subject static composition. A slightly crooked stack of boxes and a partly exposed handwritten inventory list must appear. Keep the distant background stable and avoid changes in local object counts. Multiple distinct subjects or objects must be visible with clear spatial relationships. Use a fixed camera or slow stable shot, relatively deep depth of field, and clear foreground-middle-background layering, emphasizing layout, occlusion, and complex-composition stability. Primary rubric. (1)At least 6 people and multiple box areas are visible. (2)Furniture, boxes, and the handwritten inventory list are clear. Additional criteria. (1)The distant background remains stable. (2)The photographic treatment includes stable composition or a slow stable shot, relatively deep depth of field, and clear depth layering. (3)The visual treatment should convey a natural multi-subject static composition. (4)Supporting details include a slightly crooked stack of boxes and a partly exposed handwritten inventory list. 23. T2V_L2_D0_K0_0011 Prompt. A flower shop contains at least 4 people arranged around a central worktable as they organize bouquets. The table holds varied flowers, ribbons, and scissors, while people make only small posture adjustments. Use a fixed camera and emphasize the plausibility of a complex static scene. A fallen leaf beside the flower stems and small water droplets on a glass vase must be visible. Keep the distant background stable and avoid changes in local object counts. Multiple distinct subjects or objects must remain visible with clear spatial relationships. Use a fixed camera or slow stable shot, relatively deep depth of field, and clear depth layering, emphasizing layout, occlusion, and complex-composition stability. Primary rubric. (1)At least 4 people are positioned around the flower-shop worktable. (2)Flowers, ribbons, scissors, and the vase are clearly visible. Additional criteria. (1)The number and position of background flower racks remain stable. (2)The photographic treatment includes stable composition or a slow stable shot, relatively deep depth of field, and clear depth layering. (3)The visual treatment should convey the plausibility of a complex static scene. (4)Multiple distinct subjects or objects remain visible with clear spatial relationships. 24. T2V_L2_D0_K0_0012 Prompt. In a bicycle-repair workshop, several people are distributed across different work stands in inspection or conversational poses. Tires, tools, and parts boxes are naturally arranged. Use a fixed camera and emphasize the realism of a static group scene. A horizontally placed wrench and a slightly tilted corner of a repair label must appear. Keep the distant background stable and avoid changes in local object counts. Multiple distinct subjects or objects must remain visible with clear spatial relationships. Use a fixed camera or slow stable shot, relatively deep depth of field, and clear depth layering, emphasizing layout, occlusion, and complex-composition stability. Primary rubric. (1)Several people are visible inside the bicycle-repair workshop. (2)Tires, tools, parts boxes, and work stands are clear. Additional criteria. (1)The number of background objects remains stable. (2)The photographic treatment includes stable composition or a slow stable shot, relatively deep depth of field, and clear depth layering. (3)The visual treatment should convey the realism of a static group scene. (4)Supporting details include a horizontally placed wrench and a slightly tilted corner of a repair label. 25. T2V_L2_D0_K0_0014 Prompt. A bus interior contains at least 8 passengers, 3 rows of seats, and 12 visible handrails or advertisements. People remain mostly still, making only small observational or positional adjustments. Use a fixed camera with clear foreground, middle-ground, and background layers. A slightly reflective transit card and the edge of a hanging backpack must appear. Standing and seated positions should remain clearly layered and the complex composition easy to read. Multiple distinct subjects or objects must remain visible with clear spatial relationships. Use a fixed camera or slow stable shot, relatively deep depth of field, and clear depth layering, emphasizing multi-subject layout, occlusion, and composition stability. Primary rubric. (1)At least 8 passengers are visible inside the bus. (2)Seats, handrails, advertisements, the transit card, and the backpack are visible. Additional criteria. (1)Standing and seated layers are clear. (2)Spatial relationships inside the bus remain stable. (3)The photographic treatment includes stable composition or a slow stable shot, relatively deep depth of field, and clear depth layering. (4)The visual treatment should convey multi-subject layout, occlusion, and complex-composition stability. 26. T2V_L2_D0_K0_0020 Prompt. In a music rehearsal room, several people are distributed among different music stands in tuning or conversational poses. Instruments, sheet music, and water cups are naturally arranged. Use a fixed camera and emphasize the realism of a static group scene. A horizontally placed pencil and a slightly tilted metronome must appear. Subject orientations and attention points should vary naturally rather than forming a mechanical arrangement. Multiple distinct subjects or objects must remain visible with clear spatial relationships. Use a fixed camera or slow stable shot, relatively deep depth of field, and clear depth layering, emphasizing layout, occlusion, and complex- composition stability. Primary rubric. (1)Several people are visible inside the music rehearsal room. (2)Instruments, music stands, sheet music, and water cups are clearly distributed. Additional criteria. (1)Orientations and attention points are naturally varied. (2)The photographic treatment includes stable composition or a slow stable shot, relatively deep depth of field, and clear depth layering. (3)The visual treatment should convey the realism of a static group scene. (4)Supporting details include a horizontally placed pencil and a slightly tilted metronome. RAVEN-Eval supplementary material · Full T2V prompts and rubrics · Page 3/6 Figure 7: Complete T2V benchmark cases forS=2,D=0, andK=0: high semantic complexity, standard detail, and no external knowledge requirement. S2–D0–K1 High semantic complexity · Standard detail · Knowledge required 8 tasks SHARED BASIC RUBRIC —applies to every task on this page (1)All static objects and background objects must remain visually consistent throughout the video. (2)The video should appear realistic. In particular, when material reactions or object interactions are involved, and unless the prompt states otherwise, dynamic objects should follow real-world physical, material, and chemical laws and exhibit logically consistent changes. Material quantities should be conserved as strictly as possible. (3)For videos containing motion, each action should be complete and reasonable. Unless otherwise specified, all actions and reactions should remain continuous and natural, without obvious abrupt changes. (4)If a task-specific instruction conflicts with the three requirements above, such as a scene that intentionally disregards physical laws, the task-specific instruction takes precedence; the three shared requirements are secondary. 27. T2V_L2_D0_K1_0001 Prompt. A laboratory tabletop contains at least 6 labeled containers, 5 pieces of glassware, 3 instruments, and 4 loose tools. At least four labels remain readable throughout the clip and must read “Acid A,” “Base B,” “pH 7 Buffer,” and “Sample C.” A pipette gun partly pinned beneath a record sheet and a small water mark at the bottom of a glass vessel must appear. All key labels, signs, and structural edges remain clear and stable throughout. Motion is limited to stillness or very small environmental changes, with no experimental reaction or obvious physical interaction. Use a fixed camera, professional laboratory or industrial documentary photography, even but directional lighting, deep depth of field, layered depth, and strict visual order, emphasizing object arrangement, label legibility, material detail, and overall aesthetics. Primary rubric. (1)At least 6 labeled containers, 5 pieces of glassware, 3 instruments, and 4 tools are visible. (2)At least four labels—“Acid A,” “Base B,” “pH 7 Buffer,” and “Sample C”—are clearly readable. Additional criteria. (1)Object layout remains stable and clearly layered. (2)The complex laboratory-table composition is reasonable. (3)The visual treatment should convey professional object arrangement, label legibility, material detail, and overall aesthetics. (4)Supporting details include a pipette gun partly pinned beneath a record sheet and a small water mark at the bottom of a glass vessel. 28. T2V_L2_D0_K1_0002 Prompt. An industrial workbench contains at least 2 control panels, 8 buttons or knobs, 3 warning signs, and 2 screens. Visible text must include “HIGH PRESSURE,” “EMERGENCY STOP,” and “Temp 68 C.” A slightly coiled cable and a small maintenance label attached to a panel edge must appear. All key labels, signs, and structural edges remain clear and stable throughout. Motion is limited to stillness or very small environmental changes, with no experimental reaction or obvious physical interaction. Use a fixed camera, professional laboratory or industrial documentary photography, even but directional lighting, deep depth of field, layered depth, and strict visual order, emphasizing object arrangement, label legibility, material detail, and overall aesthetics. Primary rubric. (1)Two control panels, at least 8 buttons or knobs, 3 warning signs, and 2 screens are visible. (2)“HIGH PRESSURE,” “EMERGENCY STOP,” and “Temp 68 C” are clearly readable. (3)The coiled cable and maintenance label are present. (4)The panel layout remains stable. Additional criteria. (1)The visual treatment should convey professional object arrangement, label legibility, material detail, and overall aesthetics. (2)Supporting details include a slightly coiled cable and a small maintenance label attached to the panel edge. (3)Deep depth of field and layered depth strengthen visual order. (4)Professional laboratory or industrial documentary photography reinforces object arrangement and material detail. 29. T2V_L2_D0_K1_0004 Prompt. A set of specimens or models with exhibit labels is neatly arranged on a display platform. At least 8 main objects and multiple signs are present, with sign text fixed as “Exhibit 07,” “Gallery B,” and “Specimen 12.” The static scene emphasizes shape consistency and compositional logic under dense text labeling. Slight wear on one base edge and a partly obscured numbered card must appear. All key labels, signs, and structural edges remain clear and stable throughout. Motion is limited to stillness or very small environmental changes, with no experimental reaction or obvious physical interaction. Use a fixed camera, professional laboratory or industrial documentary photography, even but directional lighting, deep depth of field, layered depth, and strict visual order. Primary rubric. (1)At least 8 main objects and multiple signs are visible on the display platform. (2)“Exhibit 07,” “Gallery B,” and “Specimen 12” are displayed accurately. (3)The relationship between each sign and its corresponding exhibit is clear. Additional criteria. (1)The display composition is neat and stable. (2)The visual treatment should maintain shape consistency and compositional logic under dense text labeling. (3)Supporting details include slight wear on a base edge and a partly obscured numbered card. (4)Deep depth of field and layered depth strengthen visual order. 30. T2V_L2_D0_K1_0005 Prompt. A professional experimental platform contains multiple numbered devices, measuring tools, and instruction labels in a complex composition. Visible text must include “Unit 04,” “Step 2,” and “Valve B.” A horizontally placed marker and a thin wire draped over a tray edge must appear. Emphasize text legibility in a complex professional static scene and keep all key labels, signs, and structural edges clear and stable throughout. Motion is limited to stillness or very small environmental changes, with no experimental reaction or obvious physical interaction. Use a fixed camera, professional laboratory or industrial documentary photography, even but directional lighting, deep depth of field, layered depth, and strict visual order. Primary rubric. (1)Multiple devices, measuring tools, and instruction labels appear together. (2)“Unit 04,” “Step 2,” and “Valve B” are clearly readable. (3)The horizontal marker and thin wire draped over the tray edge are present. Additional criteria. (1)A professional experimental-platform appearance is maintained. (2)The visual treatment should emphasize text legibility in a complex professional static scene. (3)Supporting details include a horizontally placed marker and a thin wire draped over a tray edge. (4)Deep depth of field and layered depth strengthen visual order. 31. T2V_L2_D0_K1_0006 Prompt. A radiology control console contains at least 4 image-preview screens, 5 examination forms, 3 controllers, and 4 loose marker notes. At least four text labels remain readable throughout and must read “Scan A,” “Dose Log,” “Patient ID,” and “Series 03.” A stylus partly pinned beneath a form and a small reflection point on a screen edge must appear. Spatial order, material consistency, text legibility, and compositional balance among the many objects should remain professional. Primary rubric. (1)The radiology console includes screens, forms, controllers, and marker notes. (2)“Scan A,” “Dose Log,” “Patient ID,” and “Series 03” are readable. Additional criteria. (1)Spatial order and compositional balance remain professional. (2)Text legibility remains stable. (3)Supporting details include a stylus partly pinned beneath a form and a small reflection point on a screen edge. (4)Professional visual quality and clear visual order are maintained. 32. T2V_L2_D0_K1_0007 Prompt. An aerospace-parts inspection bench contains at least 2 parts trays, 8 small fasteners or gauges, 3 warning signs, and 2 inspection drawings. Visible text must include “TORQUE LIMIT,” “INSPECTED,” and “Panel 68C.” A slightly coiled antistatic wrist strap and a small maintenance label attached to a drawing edge must appear. Spatial order, material consistency, text legibility, and compositional balance among the many objects should remain professional. Primary rubric. (1)The inspection bench contains parts trays, fasteners, gauges, and drawings. (2)“TORQUE LIMIT,” “INSPECTED,” and “Panel 68C” are readable. Additional criteria. (1)Metallic and engineering materials remain consistent. (2)The complex static composition remains stable. (3)Supporting details include a slightly coiled antistatic wrist strap and a small maintenance label attached to a drawing edge. (4)Professional visual quality and clear visual order are maintained. 33. T2V_L2_D0_K1_0008 Prompt. A geological-sample cataloging table contains at least 4 rock trays, 5 sample boxes, 5 labeled objects, and 2 record sheets. Sampling has not begun, but the layout clearly suggests preparation for cataloging. Visible text must include “Basalt Core” and “Layer 1.” A slight crease at the corner of a record sheet and a small amount of rock- powder residue beside a sample box must appear. Spatial order, material consistency, text legibility, and compositional balance should remain professional. Primary rubric. (1)Rock trays, sample boxes, labeled objects, and record sheets are visible. (2)“Basalt Core” and “Layer 1” are clear. Additional criteria. (1)The spatial arrangement of objects appears professional. (2)Relationships between text labels and samples are clear. (3)Supporting details include a slight crease on a record-sheet corner and rock-powder residue beside a sample box. (4)Professional visual quality and clear visual order are maintained. 34. T2V_L2_D0_K1_0009 Prompt. At an electronics quality-inspection station, multiple circuit boards and test fixtures are neatly arranged on an antistatic mat. At least 8 main objects and multiple signs are present, with sign text fixed as “QC 07,” “Board B,” and “Reject 12.” The static scene emphasizes shape consistency and compositional logic under dense text labeling. Slight wear on one fixture edge and a partly obscured numbered card must appear. Spatial order, material consistency, text legibility, and compositional balance should remain professional. Primary rubric. (1)The electronics inspection station contains circuit boards, test fixtures, and signs. (2)“QC 07,” “Board B,” and “Reject 12” are clear. Additional criteria. (1)The composition remains logical despite dense text. (2)The visual treatment should maintain shape consistency and compositional logic under dense text labeling. (3)Supporting details include slight wear on a fixture edge and a partly obscured numbered card. (4)Professional visual quality and clear visual order are maintained. RAVEN-Eval supplementary material · Full T2V prompts and rubrics · Page 4/6 Figure 8: Complete T2V benchmark cases forS=2,D=0, andK=1: high semantic complexity, standard detail, and an external knowledge requirement. S2–D1–K0 High semantic complexity · Rich detail · No external knowledge 8 tasks SHARED BASIC RUBRIC —applies to every task on this page (1)All static objects and background objects must remain visually consistent throughout the video. (2)The video should appear realistic. In particular, when material reactions or object interactions are involved, and unless the prompt states otherwise, dynamic objects should follow real-world physical, material, and chemical laws and exhibit logically consistent changes. Material quantities should be conserved as strictly as possible. (3)For videos containing motion, each action should be complete and reasonable. Unless otherwise specified, all actions and reactions should remain continuous and natural, without obvious abrupt changes. (4)If a task-specific instruction conflicts with the three requirements above, such as a scene that intentionally disregards physical laws, the task-specific instruction takes precedence; the three shared requirements are secondary. 35. T2V_L2_D1_K0_0001 Prompt. A busy street simultaneously contains at least 10 pedestrians, 3 bicycles, 4 cars, and 5 storefront façades. Pedestrians, bicycles, and cars are all moving. Use a fixed wide-angle street view and emphasize complex dynamic relationships among many subjects. A slightly swaying roadside sign and a small bag placed beside a seat must appear. Motion directions in the foreground, middle ground, and background should all be clearly readable. Primary rubric. (1)At least 10 pedestrians, 3 bicycles, 4 cars, and 5 storefront façades are visible simultaneously. (2)Pedestrians, bicycles, and cars remain in continuous motion, with readable directions across all depth layers. (3)Different subject types maintain reasonable spatial distribution without abnormal overlap or disappearance. Additional criteria. (1)The wide-angle view supports readable multi-subject spatial depth. (2)The visual treatment should convey complex dynamic relationships among many subjects. (3)Supporting details include a slightly swaying roadside sign and a small bag beside a seat, while motion directions remain readable across depth layers. 36. T2V_L2_D1_K0_0002 Prompt. A market passage contains at least 9 pedestrians, 2 handcarts, 6 groups of boxes, and 4 stalls or storefronts. Moving subjects appear in the foreground, middle ground, and background with natural overall rhythm. A half-open cardboard box and a slightly curled paper sign at a stall edge must appear. Motion directions across all depth layers should be clearly readable. Primary rubric. (1)At least 9 pedestrians, 2 handcarts, 6 groups of boxes, and 4 stalls or storefronts are visible simultaneously. (2)Continuously moving subjects appear across all depth layers with clearly readable directions. (3)The half-open box and curled paper sign are clearly visible. Additional criteria. (1)Pedestrian and handcart motion is continuous and natural, preserving plausible traffic relationships. (2)The overall rhythm is stable and fully conveys the market's dynamic layering. (3)Supporting details include the half-open box and curled paper sign, while motion directions remain readable across depth layers. (4)The background atmosphere, supporting details, and visual rhythm remain consistent with the prompt. 37. T2V_L2_D1_K0_0003 Prompt. A playground contains at least 8 children, 1 ball, 4 playground structures, and 1 set of bleachers. The children run, pass the ball, and cross one another's paths. Use a fixed camera and emphasize a dynamic multi-subject scene. A water bottle at the edge of the field and a short chalk line on the ground must appear. Motion directions in the foreground, middle ground, and background should all be clearly readable. Primary rubric. (1)At least 8 children, 1 ball, 4 playground structures, and 1 bleacher area are visible simultaneously. (2)The children run, pass the ball, and cross paths, with the ball-transfer process continuously visible. (3)The water bottle and chalk line are clearly visible. (4)Moving subjects appear across all depth layers with clearly readable directions. Additional criteria. (1)Multi-subject interaction is continuous, natural, and consistent with playground activity. (2)The visual treatment should convey a dynamic multi-subject scene. (3)Supporting details include a water bottle and short chalk line, while motion directions remain readable across depth layers. 38. T2V_L2_D1_K0_0004 Prompt. A port loading area simultaneously contains at least 10 workers, 2 forklifts, 3 handcarts, and the edges of 5 container groups. Workers, forklifts, and handcarts are moving, with at least one natural crossing or occlusion among moving subjects. Multiple subjects must form readable front-back and left-right relationships without random drift or identity confusion. Use relatively deep depth of field and a stable medium-long shot or slow tracking shot, emphasizing multi-subject dynamics, occlusion handling, and overall motion order. Primary rubric. (1)Workers, forklifts, handcarts, and containers are visible in the port loading area. (2)At least 10 workers and multiple moving vehicles are visible. (3)Moving subjects exhibit natural crossing or occlusion. Additional criteria. (1)Front-back and left-right relationships remain clear. (2)The photographic treatment uses relatively deep depth of field and a stable medium-long or slow tracking shot. (3)The visual treatment should convey multi-subject dynamics, occlusion handling, and overall motion order. (4)Deep depth of field and layered depth strengthen visual order. 39. T2V_L2_D1_K0_0005 Prompt. An airport terminal contains at least 9 travelers, 2 luggage carts, 6 groups of suitcases, and 4 counter or gate areas. Moving subjects appear in the foreground, middle ground, and background with natural rhythm. A half-open boarding bag and a slightly curled baggage tag at a counter edge must appear. At least one natural crossing or occlusion occurs among moving subjects. Multiple subjects form readable front-back and left-right relationships without random drift or identity confusion. Use relatively deep depth of field and a stable medium-long or slow tracking shot, emphasizing multi-subject dynamics, occlusion handling, and overall motion order. Primary rubric. (1)Travelers, luggage carts, suitcases, and counters are visible in the airport terminal. (2)Moving subjects appear in the foreground, middle ground, and background. (3)The half-open boarding bag and baggage tag are clear. Additional criteria. (1)Moving subjects exhibit natural crossing. (2)Spatial layering and subject identities remain stable. (3)The photographic treatment uses relatively deep depth of field and a stable medium-long or slow tracking shot. (4)The visual treatment should convey multi-subject dynamics, occlusion handling, and overall motion order. 40. T2V_L2_D1_K0_0006 Prompt. An indoor basketball gym contains at least 8 players, 2 basketballs, 4 training cones, and 1 courtside seating area. Players run, pass, or defend in different directions under a stable wide-angle camera. A strip of tape on the floor and a slightly swaying towel beside the seats must appear. At least one natural crossing or occlusion occurs among moving subjects. Multiple subjects form readable front-back and left-right relationships without random drift or identity confusion. Use relatively deep depth of field and a stable medium-long or slow tracking shot, emphasizing multi-subject dynamics, occlusion handling, and overall motion order. Primary rubric. (1)At least 8 players are moving inside the basketball gym. (2)Basketballs, training cones, and the courtside seating area are visible. (3)Running, passing, or defensive actions remain continuous. Additional criteria. (1)The tape mark and towel serve as supporting references. (2)Occlusion is natural and subject identities do not become confused. (3)The photographic treatment uses relatively deep depth of field and a stable medium-long or slow tracking shot. (4)The visual treatment should convey multi-subject dynamics, occlusion handling, and overall motion order. 41. T2V_L2_D1_K0_0009 Prompt. An orchestral rehearsal hall contains at least 8 performers, multiple instruments, 4 music-stand areas, and 1 observation-seating area. Following the conductor's cues, performers make subtle but synchronized body and arm movements. Use a stable wide-angle camera. A pencil beside a music stand and one slightly fluttering sheet of music in the front seating row must appear. Main motion paths remain continuous without abrupt jumps or broken rhythm. Multiple subjects form readable front-back and left-right relationships without random drift or identity confusion. Use relatively deep depth of field and a stable medium-long or slow tracking shot, emphasizing multi- subject dynamics, occlusion handling, and overall motion order. Primary rubric. (1)At least 8 performers and multiple instruments are visible. (2)Subtle synchronized movement occurs under the conductor's cues. (3)The music stands, pencil, and slightly fluttering sheet music are clear. Additional criteria. (1)The wide-angle composition remains stable. (2)Group identities and spatial relationships do not become confused. (3)The photographic treatment uses relatively deep depth of field and a stable medium-long or slow tracking shot. (4)The visual treatment should convey multi-subject dynamics, occlusion handling, and overall motion order. 42. T2V_L2_D1_K0_0010 Prompt. A warehouse sorting area simultaneously contains at least 10 workers, 3 handcarts, 4 conveyor segments, and 5 shelving façades. Workers, handcarts, and parcels are moving, while the temporal order and path relationships among moving subjects remain clear. Multiple subjects form readable front-back and left-right relationships without random drift or identity confusion. Use relatively deep depth of field and a stable medium-long or slow tracking shot, emphasizing multi-subject dynamics, occlusion handling, and overall motion order. Primary rubric. (1)Workers, handcarts, conveyors, and shelves are visible in the warehouse sorting area. (2)At least 10 workers participate in the motion. (3)Path relationships among parcels and people are clear. Additional criteria. (1)The temporal order remains traceable. (2)Overall motion order remains stable. (3)The photographic treatment uses relatively deep depth of field and a stable medium-long or slow tracking shot. (4)The visual treatment should convey multi-subject dynamics, occlusion handling, and overall motion order. RAVEN-Eval supplementary material · Full T2V prompts and rubrics · Page 5/6 Figure 9: Complete T2V benchmark cases forS=2,D=1, andK=0: high semantic complexity, rich detail, and no external knowledge requirement. S2–D1–K1 High semantic complexity · Rich detail · Knowledge required 6 tasks SHARED BASIC RUBRIC —applies to every task on this page (1)All static objects and background objects must remain visually consistent throughout the video. (2)The video should appear realistic. In particular, when material reactions or object interactions are involved, and unless the prompt states otherwise, dynamic objects should follow real-world physical, material, and chemical laws and exhibit logically consistent changes. Material quantities should be conserved as strictly as possible. (3)For videos containing motion, each action should be complete and reasonable. Unless otherwise specified, all actions and reactions should remain continuous and natural, without obvious abrupt changes. (4)If a task-specific instruction conflicts with the three requirements above, such as a scene that intentionally disregards physical laws, the task-specific instruction takes precedence; the three shared requirements are secondary. 43. T2V_L2_D1_K1_0001 Prompt. A professional wide-angle recording of a 100-meter sprint contains at least 8 athletes, a clearly visible track, and a timing environment. The race is shown continuously from start to finish, with athletes' relative positions gradually changing as the event progresses. The composition is complex and professional. A small towel beside a starting block and a slightly swaying handheld sign in the front row of the stands must appear. Transitions between key stages remain clear and supporting reference objects stay visible throughout the complex motion. Primary rubric. (1)The start, acceleration, mid-race running, and finish stages appear in sequence. (2)The small towel beside the starting block remains identifiable. (3)The handheld sign in the front row of the stands sways slightly. (4)The changing positions of multiple athletes remain continuously traceable. Additional criteria. (1)The wide-angle view supports readable multi-subject spatial depth. (2)Supporting details include the small towel beside the starting block and the slightly swaying handheld sign. (3)Professional visual quality and clear visual order are maintained. 44. T2V_L2_D1_K1_0002 Prompt. A multi-person dance performance on a professional stage contains at least 8 dancers. The dancers complete a continuous unified choreography with observable changes in group positions. The camera preserves the full stage space and group interaction. A floor light at the side of the stage remains partly visible, and a fabric accessory at one dancer's waist sways slightly. Transitions between key stages remain clear and supporting reference objects stay visible throughout the complex motion. Primary rubric. (1)At least 8 dancers participate simultaneously. (2)Formation changes are clear and complete. (3)The floor light at the side of the stage remains visible. (4)Group movement is synchronized and spatial relationships remain stable. Additional criteria. (1)The specified dancer's waist accessory sways naturally. (2)Supporting details include the partly visible floor light and the slightly swaying fabric accessory. (3)Professional visual quality and clear visual order are maintained. 45. T2V_L2_D1_K1_0003 Prompt. In a complex multi-object mechanics experiment, multiple devices, measuring tools, and target objects are visible together. The apparatus is triggered in a specified manner, after which the objects continue interacting and produce subsequent changes. A measurement scale sheet attached to a support and a spare weight block at a corner of the workbench must appear. Transitions between key stages remain clear and supporting reference objects stay visible throughout the complex motion. Primary rubric. (1)Multiple devices, measuring tools, and target objects are visible together. (2)The trigger condition, main phenomenon, and final state are clearly presented. (3)The measurement scale sheet on the support remains visible. (4)The spare weight block remains at the corner of the workbench. (5)Interactions among the devices remain continuously traceable. Additional criteria. (1)Supporting details include the measurement scale sheet on the support and the spare weight block at the workbench corner. (2)The background atmosphere, supporting details, and visual rhythm remain consistent with the prompt. 46. T2V_L2_D1_K1_0004 Prompt. In an aerial formation or complex stunt scene, multiple high-speed subjects execute continuous flight maneuvers simultaneously. Their relative positions change continuously and gradually form new spatial relationships. A small flag in the ground observation area and a faint, short-lived smoke trail behind one aircraft must appear. Transitions between key stages remain clear and supporting reference objects stay visible throughout the complex motion. Primary rubric. (1)Multiple flying subjects complete a formation change. (2)Motion stages and relative-position changes are clearly presented. (3)The small flag in the ground observation area remains visible. (4)At least one aircraft leaves a brief faint smoke trail. (5)The flying subjects gradually form a new spatial relationship. Additional criteria. (1)Supporting details include the small flag in the observation area and the faint smoke trail behind one aircraft. (2)The background atmosphere, supporting details, and visual rhythm remain consistent with the prompt. 47. T2V_L2_D1_K1_0005 Prompt. On a complex chemical experimental platform, materials from multiple containers are added to a main container in a specified sequence. The states of the materials, addition order, and operation process remain clearly visible, and subsequent changes follow common real-world physical or chemical laws. A record card marked by a small liquid droplet and a small amount of residue on the end of a stirring rod must appear. Transitions between key stages remain clear and supporting reference objects stay visible throughout the complex motion. Primary rubric. (1)Multiple materials are added or mixed sequentially. (2)The mixing process remains continuously visible. (3)An observable change appears inside the container. (4)The change occurs after mixing begins. (5)The full process has a clear temporal order. Additional criteria. (1)Supporting details include the droplet-marked record card and residue on the end of the stirring rod. (2)The droplet and stirring-rod residue provide experimental reference details. 48. T2V_L2_D1_K1_0010 Prompt. On a robotic assembly line, multiple robotic arms sequentially grasp, rotate, and place parts in a specified order. Part states, transfer order, and the operating process remain clearly visible, and subsequent changes follow common real-world mechanical-assembly logic. A calibration card attached to the workstation edge and a small amount of lubricant residue on a robotic-arm end effector must appear. Relative-position changes among people or devices remain fully traceable, while the scene preserves a stable, credible experimental environment and professional operating procedure. Primary rubric. (1)The scene is a robotic assembly line. (2)Multiple robotic arms grasp, rotate, and place parts in sequence. (3)The transfer order and relative-position changes remain fully traceable. Additional criteria. (1)The calibration card and lubricant residue serve as supporting references. (2)The mechanical-assembly logic remains stable and credible. (3)Supporting details include the calibration card and lubricant residue; relative-position changes remain fully traceable in a stable, credible environment and professional workflow. (4)Professional visual quality and clear visual order are maintained. RAVEN-Eval supplementary material · Full T2V prompts and rubrics · Page 6/6 Figure 10: Complete T2V benchmark cases forS=2,D=1, andK=1: high semantic complexity, rich detail, and an external knowledge requirement. T2V Case Example Moderate semantics · Rich detail · No external knowledge · six model generations S1-D1-K0 PROMPT One child throws a yellow ball to another child, who catches it and steps slightly backward. Use a fixed camera. Emphasize direct interaction between a small number of subjects, preserve natural continuity across the throw and catch, and make the action order and rhythm of both participants clear. TASK-SPECIFIC PRIMARY RUBRIC 1. The ball travels toward the second child along a continuous trajectory. 2. The second child successfully catches the ball and reacts by stepping slightly backward. 3. The full sequence of throwing, flight, and catching is complete and continuous. 4. The spatial relationship between the two children and the outcome of the interaction remain clearly visible. Seedance 2 6.0 s Kling 3.0 Pro 6.0 s WAN 2.7 6.0 s Veo 3.1 Fast 6.0 s PixVerse C1 6.0 s LTX 2.3 Pro 6.1 s RAVEN-Eval supplementary material · T2V · S1-D1-K0 · frames sampled at fixed relative positions · 1/7 Figure 11: A representative T2V case under the S1-D1-K0 configuration. The case evaluates a continuous ball-throwing and catching interaction between two children without requiring external knowledge. T2V Case Example Moderate semantics · Rich detail · External knowledge · six model generations S1-D1-K1 PROMPT One astronaut slowly walks toward a giant black monolith while a second astronaut stands beside a rover and observes. Use a 2001: A Space Odyssey- style austere science-fiction aesthetic with symmetrical composition, minimalist space, slow movement, cool white light, and extensive silent negative space. Motion should be subtle but continuous, emphasizing cinematic composition and a solemn atmosphere. TASK-SPECIFIC PRIMARY RUBRIC 1. The frame contains two astronauts and one giant black monolith. 2. One astronaut slowly approaches the monolith while the other observes beside the rover. 3. The scene presents an austere, minimalist science-fiction aesthetic in the style of 2001: A Space Odyssey. Seedance 2 Fast 6.0 s Kling 3.0 Omni 6.0 s Grok-Image-Video 6.0 s WAN 2.6 6.0 s Veo 3.1 Lite 6.0 s LTX 2.3 Fast 6.1 s RAVEN-Eval supplementary material · T2V · S1-D1-K1 · frames sampled at fixed relative positions · 2/7 Figure 12: A representative T2V case under the S1-D1-K1 configuration. The case requires a restrained science-fiction composi- tion involving two astronauts and a monolith while following the specified cinematic conventions. T2V Case Example High semantics · Standard detail · No external knowledge · six model generations S2-D0-K0 PROMPT A train-station waiting hall contains at least 7 passengers, 3 rows of seats, multiple suitcases, and 2 information screens. People wait or look down at tickets, making only very small page turns or posture adjustments. Use a fixed camera and emphasize a natural multi-subject static composition. A wheeled suitcase with an old luggage tag and a partly exposed paper ticket must appear. Maintain reasonable spacing among the main people and keep the aisle or table edge readable. Multiple distinct subjects or objects must be visible with clear spatial relationships. Use a fixed camera or slow stable shot, relatively deep depth of field, and clear foreground-middle-background layering, emphasizing multi-subject layout, occlusion, and complex- composition stability. TASK-SPECIFIC PRIMARY RUBRIC 1. At least 7 passengers and multiple groups of waiting-hall seats are visible. 2. Suitcases, information screens, and the paper-ticket element are clear. Seedance 1.5 Pro 6.0 s Kling 3.0 Std 6.0 s HappyHorse V1.0 6.1 s PixVerse V6 6.0 s Hailuo 2.3 5.9 s LTX 2.0 Pro 6.1 s RAVEN-Eval supplementary material · T2V · S2-D0-K0 · frames sampled at fixed relative positions · 3/7 Figure 13: A representative T2V case under the S2-D0-K0 configuration. The case evaluates the stability of a crowded train- station waiting hall containing multiple passengers, seats, suitcases, and information screens. T2V Case Example High semantics · Standard detail · External knowledge · six model generations S2-D0-K1 PROMPT A laboratory tabletop contains at least 6 labeled containers, 5 pieces of glassware, 3 instruments, and 4 loose tools. At least four labels remain readable throughout the clip and must read “Acid A,” “Base B,” “pH 7 Buffer,” and “Sample C.” A pipette gun partly pinned beneath a record sheet and a small water mark at the bottom of a glass vessel must appear. All key labels, signs, and structural edges remain clear and stable throughout. Motion is limited to stillness or very small environmental changes, with no experimental reaction or obvious physical interaction. Use a fixed camera, professional laboratory or industrial documentary photography, even but directional lighting, deep depth of field, layered depth, and strict visual order, emphasizing object arrangement, label legibility, material detail, and overall aesthetics. TASK-SPECIFIC PRIMARY RUBRIC 1. At least 6 labeled containers, 5 pieces of glassware, 3 instruments, and 4 tools are visible. 2. At least four labels—“Acid A,” “Base B,” “pH 7 Buffer,” and “Sample C”—are clearly readable. Seedance 2 6.0 s Kling 3.0 Pro 6.0 s WAN 2.7 6.0 s PixVerse V5.5 5.0 s Veo 3.1 Lite 6.0 s LTX 2.0 Fast 6.1 s RAVEN-Eval supplementary material · T2V · S2-D0-K1 · frames sampled at fixed relative positions · 4/7 Figure 14: A representative T2V case under the S2-D0-K1 configuration. The case evaluates object-count compliance, text legibility, material consistency, and spatial organization on a densely populated laboratory tabletop. T2V Case Example High semantics · Rich detail · No external knowledge · six model generations S2-D1-K0 PROMPT A playground contains at least 8 children, 1 ball, 4 playground structures, and 1 set of bleachers. The children run, pass the ball, and cross one another's paths. Use a fixed camera and emphasize a dynamic multi-subject scene. A water bottle at the edge of the field and a short chalk line on the ground must appear. Motion directions in the foreground, middle ground, and background should all be clearly readable. TASK-SPECIFIC PRIMARY RUBRIC 1. At least 8 children, 1 ball, 4 playground structures, and 1 bleacher area are visible simultaneously. 2. The children run, pass the ball, and cross paths, with the ball-transfer process continuously visible. 3. The water bottle and chalk line are clearly visible. 4. Moving subjects appear across all depth layers with clearly readable directions. Seedance 2 Fast 6.0 s Kling 3.0 Omni 6.0 s Grok-Image-Video 6.0 s WAN 2.6 6.0 s Veo 3.1 Fast 6.0 s PixVerse V6 6.0 s RAVEN-Eval supplementary material · T2V · S2-D1-K0 · frames sampled at fixed relative positions · 5/7 Figure 15: A representative T2V case under the S2-D1-K0 configuration. The case evaluates multi-subject motion, ball passing, path crossings, and readable movement across foreground, middle-ground, and background regions. T2V Case Example High semantics · Rich detail · External knowledge · six model generations S2-D1-K1 PROMPT A multi-person dance performance on a professional stage contains at least 8 dancers. The dancers complete a continuous unified choreography with observable changes in group positions. The camera preserves the full stage space and group interaction. A floor light at the side of the stage remains partly visible, and a fabric accessory at one dancer's waist sways slightly. Transitions between key stages remain clear and supporting reference objects stay visible throughout the complex motion. TASK-SPECIFIC PRIMARY RUBRIC 1. At least 8 dancers participate simultaneously. 2. Formation changes are clear and complete. 3. The floor light at the side of the stage remains visible. 4. Group movement is synchronized and spatial relationships remain stable. Seedance 1.5 Pro 6.0 s Kling 3.0 Std 6.0 s HappyHorse V1.0 6.1 s WAN 2.7 6.0 s Hailuo 2.3 5.9 s LTX 2.3 Pro 6.1 s RAVEN-Eval supplementary material · T2V · S2-D1-K1 · frames sampled at fixed relative positions · 6/7 Figure 16: A representative T2V case under the S2-D1-K1 configuration. The case requires at least eight dancers to perform synchronized choreography with coherent formation changes and stable spatial relationships. T2V Case Example Physical outcome inference · Material response · six model generations S1-D1-K1Reasoning PROMPT A hydraulic press slowly descends toward a rubber duck, with the press head, duck, and workbench visible together. The complete process should show the press head moving continuously downward, making contact with the target, and producing subsequent changes that follow the natural response of a material under force. The interaction path, final state, and professional scene logic should remain clear and explainable. TASK-SPECIFIC PRIMARY RUBRIC 1. The press head continuously approaches the target object. 2. The press head makes clear contact with the target. 3. The target visibly deforms after force is applied. 4. The final state is consistent with the applied-force process. Seedance 2 6.0 s Kling 3.0 Pro 6.0 s Grok-Image-Video 6.0 s WAN 2.6 6.0 s PixVerse C1 6.0 s LTX 2.3 Fast 6.1 s RAVEN-Eval supplementary material · T2V · Reasoning · frames sampled at fixed relative positions · 7/7 Figure 17: A representative reasoning-focused T2V case. The models must infer and generate the material response of a rubber duck as a hydraulic press descends, makes contact, and applies force. FLT Case Example Reference sequence and six model generations ReasoningYouTube Source PROMPT A large brick is suddenly thrown into the still surface of a sewage pool. Complete the entire process. TASK-SPECIFIC PRIMARY RUBRIC 1. A large brick must be thrown into the initially still sewage pool. 2. A large splash should appear immediately, followed by several clearly visible concentric ripples. MODEL INPUT · JSON-defined first and last frames JSON t = 12.706 / 15.581 s FIRST FRAMELAST FRAME Seedance 2 4.0 s Kling 3.0 Pro 3.0 s WAN 2.7 3.0 s Veo 3.1 Fast 8.0 s PixVerse C1 5.0 s LTX 2.3 Pro 6.1 s RAVEN-Eval supplementary material · FLT · Reasoning · YouTube · generated frames: first + 6 uniformly spaced intermediate + last· 1/8 Figure 18: A reasoning FLT case collected from YouTube. The first and last frames are provided as model inputs. Six model outputs are shown using eight uniformly sampled frames per model. FLT Case Example Reference sequence and six model generations ReasoningASMR Source PROMPT Show the complete process of a hydraulic press compressing several white steamed buns: the buns first flatten, and then thin white strips are extruded through holes in the press plate. TASK-SPECIFIC PRIMARY RUBRIC 1. The full sequence must show the press moving downward, the buns flattening, and thin white strips being forced through the holes in the press plate. MODEL INPUT · JSON-defined first and last frames JSON t = 25.452 / 28.444 s FIRST FRAMELAST FRAME Seedance 2 Fast 4.0 s Kling 3.0 Omni 3.0 s Grok-Image-Video 3.0 s WAN 2.6 4.0 s Veo 3.1 Lite 8.0 s LTX 2.3 Fast 6.1 s RAVEN-Eval supplementary material · FLT · Reasoning · ASMR · generated frames: first + 6 uniformly spaced intermediate + last · 2/8 Figure 19: A reasoning FLT case collected from an ASMR video. The case requires the models to infer and generate the physical transition between the given first and last frames. FLT Case Example Reference sequence and six model generations Non-reasoningYouTube Source PROMPT Keep the top-down viewpoint identical to the first frame, steady and free of motion blur. As the boat moves, white wake continuously forms beneath it and spreads outward with clearly discernible texture. TASK-SPECIFIC PRIMARY RUBRIC 1. The camera must remain in the same stable top-down view as the first frame, without motion blur. 2. The boat's motion must continuously generate a spreading wake with clear white-water texture. MODEL INPUT · JSON-defined first and last frames JSON t = 59.570 / 61.990 s FIRST FRAMELAST FRAME Seedance 1.5 Pro 4.0 s Kling 3.0 Std 3.0 s WAN 2.7 2.0 s PixVerse V6 5.0 s Veo 3.1 Fast 8.0 s LTX 2.3 Pro 6.1 s RAVEN-Eval supplementary material · FLT · Non-reasoning · YouTube · generated frames: first + 6 uniformly spaced intermediate + last · 3/8 Figure 20: A non-reasoning FLT case collected from YouTube. The case primarily evaluates visual continuity, camera consistency, and the temporal development of the boat wake. FLT Case Example Reference sequence and six model generations Non-reasoningASMR Source PROMPT Show two hands kneading the petal-shaped soap into fragments, including the detailed breakage caused by the applied hand pressure. TASK-SPECIFIC PRIMARY RUBRIC 1. The two hands must knead the petal-shaped soap into fragments, and the crushing details produced by the applied force must be visible. MODEL INPUT · JSON-defined first and last frames JSON t = 281.890 / 285.324 s FIRST FRAMELAST FRAME Seedance 2 4.0 s Kling 3.0 Pro 3.0 s WAN 2.6 4.0 s PixVerse V5.5 5.0 s Veo 3.1 Lite 8.0 s LTX 2.3 Fast 6.1 s RAVEN-Eval supplementary material · FLT · Non-reasoning · ASMR · generated frames: first + 6 uniformly spaced intermediate + last · 4/8 Figure 21: A non-reasoning FLT case collected from an ASMR video. The case evaluates whether the models faithfully complete the soap-crushing process between the input frames. KFT Case Example Reference sequence and six model generations ReasoningYouTube Source PROMPT Using the communicating-vessels principle, infer the final equilibrium water levels in the large bucket and the small bottle, and generate the complete water-level transition. Keep the liquid color and container shapes consistent throughout. TASK-SPECIFIC PRIMARY RUBRIC 1. The water level in the small bottle should rise slowly until it visually matches the level in the large bucket. 2. The liquid color and the shapes of both containers must remain consistent throughout. MODEL INPUT · starting keyframe 0.7 s reference KEYFRAME HappyHorse V1.0 6.0 s Seedance 2 6.0 s Kling 3.0 Pro 6.0 s WAN 2.7 6.0 s Veo 3.1 Fast 6.0 s PixVerse V6 6.0 s RAVEN-Eval supplementary material · KFT · Reasoning · YouTube · frames sampled at fixed relative positions · 5/8 Figure 22: A reasoning KFT case collected from YouTube. Starting from the given keyframe, the models must infer the water- level transition according to the communicating-vessels principle. KFT Case Example Reference sequence and six model generations ReasoningASMR Source PROMPT A red-hot iron ball falls freely from above six balloons. Predict and show the complete subsequent phenomenon. TASK-SPECIFIC PRIMARY RUBRIC 1. Because of its high temperature, the falling iron ball should burst all six balloons from top to bottom almost instantaneously. 2. The six bursts should be extremely fast but still continuous, with their order discernible. MODEL INPUT · starting keyframe 0.7 s reference KEYFRAME Seedance 2 Fast 6.0 s Kling 3.0 Omni 6.0 s Grok-Image-Video 6.0 s HappyHorse V1.0 6.0 s Hailuo 2.3 5.9 s LTX 2.3 Fast 6.1 s RAVEN-Eval supplementary material · KFT · Reasoning · ASMR · frames sampled at fixed relative positions · 6/8 Figure 23: A reasoning KFT case collected from an ASMR video. The models are required to predict the rapid sequential bursting of six balloons caused by a falling red-hot iron ball. KFT Case Example Reference sequence and six model generations Non-reasoningYouTube Source PROMPT At the start, only a heavier Shiba Inu lies quietly above the stairs. A slimmer Shiba Inu then appears from the upper-left corner of the stairway and finally stops to the left of the first dog. TASK-SPECIFIC PRIMARY RUBRIC 1. The second, slimmer dog must enter continuously and naturally from the upper-left stair corner and stop to the left of the first dog, without teleportation. MODEL INPUT · starting keyframe 3.4 s reference KEYFRAME Seedance 1.5 Pro 6.0 s Kling 3.0 Std 6.0 s WAN 2.6 6.0 s Veo 3.1 Lite 6.0 s PixVerse C1 6.0 s LTX 2.0 Pro 6.1 s RAVEN-Eval supplementary material · KFT · Non-reasoning · YouTube · frames sampled at fixed relative positions · 7/8 Figure 24: A non-reasoning KFT case collected from YouTube. The case evaluates whether the second Shiba Inu enters naturally and stops at the specified position without temporal discontinuities. KFT Case Example Reference sequence and six model generations Non-reasoningASMR Source PROMPT Show multiple marbles rolling along the tracks. New marbles should continuously emerge from the end of each track, so the number of marbles on every track increases, while each track's marble color remains unchanged. TASK-SPECIFIC PRIMARY RUBRIC 1. Marbles must roll along all tracks as new marbles continuously emerge and the count on each track increases. 2. The marble color assigned to each track must remain consistent from beginning to end. MODEL INPUT · starting keyframe 1.4 s reference KEYFRAME Seedance 2 6.0 s Kling 3.0 Pro 6.0 s WAN 2.7 6.0 s PixVerse V5.5 5.0 s Hailuo 2.3 5.9 s LTX 2.0 Fast 6.1 s RAVEN-Eval supplementary material · KFT · Non-reasoning · ASMR · frames sampled at fixed relative positions · 8/8 Figure 25: A non-reasoning KFT case collected from an ASMR video. The case evaluates continuous marble motion, increasing marble counts, and color consistency across the generated sequence. Figure 26: T2V failure cases. Figure 27: KFT failure cases. Figure 28: FLT failure cases. Figure 29: Pairwise win-rate heatmaps for the T2V, KFT, and FLT tasks, computed from all available LMM judge results. Each cell shows the win rate of the row model against the column model, where a merged pairwise sample is counted only when the same judge provides valid forward and reversed labels. Ties and inconsistent forward–reverse judgments are counted as0.5wins for both models. Diagonal cells are left blank. For FLT, missing off-diagonal pairs are filled by interpolation for visualization only. (a) T2V (n = 150)(b) FLT (n = 50)(c) KFT (n = 50) Figure 30: Prompt word clouds for the three evaluation tasks: (a) text-to-video generation (T2V), (b) first–last-frame-to-video generation (FLT), and (c) keyframe-to-video generation (KFT). Word size indicates frequency in the corresponding translated and preprocessed prompt corpus.