Paper deep dive
VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation
Longteng Jiang, DanDan Zheng, Qianqian Qiao, Heng Huang, Huaye Wang, Yihang Bo, Bao Peng, Jingdong Chen, Jun Zhou, Xin Jin
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/14/2026, 2:01:09 AM
Summary
VGA-Bench is a comprehensive, unified evaluation framework for AIGC-based video generation models. It introduces a three-tier taxonomy covering Aesthetic Quality, Aesthetic Tagging, and Generation Quality, supported by a large-scale dataset of 60,000 videos generated from 1,016 diverse prompts. The framework includes three specialized neural assessorsâVAQA-Net, VTag-Net, and VGQA-Netâto provide automated, fine-grained, and scalable evaluation of video generation models, addressing limitations in existing benchmarks like V-Bench.
Entities (5)
Relation Signals (4)
VGA-Bench â includesassessor â VAQA-Net
confidence 100% ¡ we develop three dedicated multi-task neural assessors: VAQA-Net for aesthetic quality prediction
VGA-Bench â includesassessor â VTag-Net
confidence 100% ¡ VTag-Net for automatic aesthetic tagging
VGA-Bench â includesassessor â VGQA-Net
confidence 100% ¡ VGQA-Net for generation and basic quality attributes
VGA-Bench â builtupon â V-Bench
confidence 90% ¡ Building upon V-Bench, we refine the taxonomy
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The rapid advancement of AIGC-based video generation has underscored the critical need for comprehensive evaluation frameworks that go beyond traditional generation quality metrics to encompass aesthetic appeal. However, existing benchmarks remain largely focused on technical fidelity, leaving a significant gap in holistic assessment-particularly with respect to perceptual and artistic qualities. To address this limitation, we introduce VGA-Bench, a unified benchmark for joint evaluation of video generation quality and aesthetic quality. VGA-Bench is built upon a principled three-tier taxonomy: Aesthetic Quality, Aesthetic Tagging, and Generation Quality, each decomposed into multiple fine-grained sub-dimensions to enable systematic assessment. Guided by this taxonomy, we design 1,016 diverse prompts and generate a large-scale dataset of over 60,000 videos using 12 video generation models, ensuring broad coverage across content, style, and artifacts. To enable scalable and automated evaluation, we annotate a subset of the dataset via human labeling and develop three dedicated multi-task neural assessors: VAQA-Net for aesthetic quality prediction, VTag-Net for automatic aesthetic tagging, and VGQA-Net for generation and basic quality attributes. Extensive experiments demonstrate that our models achieve reliable alignment with human judgments, offering both accuracy and efficiency. We release VGA-Bench as a public benchmark to foster research in AIGC evaluation, with applications in content moderation, model debugging, and generative model optimization.
Tags
Links
- Source: https://arxiv.org/abs/2604.10127v1
- Canonical: https://arxiv.org/abs/2604.10127v1
Trouble viewing inline? Open PDF directly â
Full Text
58,881 characters extracted from source content.
Expand or collapse full text
VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Longteng Jiang 1 DanDan Zheng 1 Qianqian Qiao 1 Heng Huang 1 Huaye Wang 1 Yihang Bo 2 Bao Peng 2 Jingdong Chen 1, â JUN ZHOU 1 Xin Jin 3, â 1 Ant Group 2 Beijing Film Academy 3 State Key Laboratory of General Artificial Intelligence, BIGAI Abstract The rapid advancement of AIGC-based video genera- tion has underscored the critical need for comprehensive evaluation frameworks that go beyond traditional genera- tion quality metrics to encompass aesthetic appeal. How- ever, existing benchmarks remain largely focused on tech- nical fidelity, leaving a significant gap in holistic assess- mentâparticularly with respect to perceptual and artistic qualities. To address this limitation, we introduce VGA- Bench, a unified benchmark for joint evaluation of video generation quality and aesthetic quality. VGA-Bench is built upon a principled three-tier taxon- omy: Aesthetic Quality, Aesthetic Tagging, and Generation Quality, each decomposed into multiple fine-grained sub- dimensions to enable systematic assessment. Guided by this taxonomy, we design 1,016 diverse prompts and gen- erate a large-scale dataset of over 60,000 videos using 12 video generation models, ensuring broad coverage across content, style, and artifacts. To enable scalable and automated evaluation, we an- notate a subset of the dataset via human labeling and de- velop three dedicated multi-task neural assessors: VAQA- Net for aesthetic quality prediction, VTag-Net for automatic aesthetic tagging, and VGQA-Net for generation and basic quality attributes. Extensive experiments demonstrate that our models achieve reliable alignment with human judg- ments, offering both accuracy and efficiency. We release VGA-Bench as a public benchmark to foster research in AIGC evaluation, with applications in content moderation, model debugging, and generative model optimization. 1. Introduction In recent years, Artificial Intelligence Generated Content (AIGC) technologies, particularly in the realm of video gen- eration [2, 3, 14, 17, 23, 29], have seen rapid advancements. â Corresponding authors. Leveraging progress in diffusion models [2, 13, 30, 43], transformers [1, 22, 28], and large-scale vision-language pretraining [6, 8, 9, 37], current video generation mod- els [2, 10, 11, 18, 21, 24, 32, 33, 35, 36, 38, 40, 42] can now produce highly coherent, temporally stable, and visually ap- pealing video sequences from text prompts. These capabil- ities hold significant potential for applications in digital art, film production, and virtual reality. However, as these gen- erative models become increasingly sophisticated, the need for a comprehensive, reliable, and interpretable evaluation framework becomes more pressing. Traditional metrics such as FVD [34], CLIP Score [12], or their upgraded versions [20] primarily focus on technical fidelityâmeasuring temporal consistency, prompt align- ment, or image distortion levelsâbut often fail to capture higher-level perceptual qualities, especially aesthetic ex- pressiveness that critically influences visual content. Al- though recent studies have attempted to address this gap, most existing benchmarks still suffer from limited coverage and coarse-grained assessment. Among them, V-Bench [15] represents one of the first systematic efforts to evaluate AIGC videos across multiple dimensions, marking an important step towards standard- ized evaluation. However, it simplifies âvideo aestheticsâ into a single score metric and heavily relies on external scor- ing models (e.g., MUSIQ [16], DINO [5]), resulting in in- sufficient granularity, significant bias, and weak controlla- bility. As shown in Figure 1, to overcome the aforementioned limitations, this paper introduces VGA-Bench: a unified and fine-grained evaluation benchmark for AIGC-generated videos, aiming to enable joint assessment of generation quality, aesthetic quality, and visual formal elements (tags). Our main contributions are as follows: ⢠A detailed and systematic three-dimensional evaluation framework: Building upon V-Bench, we refine the taxon- omy by proposing a comprehensive structure encompass- ing three core dimensionsâgeneration quality, aesthetic quality, and visual formal elements. Each dimension is arXiv:2604.10127v1 [cs.CV] 11 Apr 2026 Generated Videos Human Annotation Prompt Suite Composition+Shot Size The distant mountains and the nearby stream alternate, creating a balanced composition with clear focal points. The variation in scenery enhances the natural awe-inspiring effect. LLM Aesthetic QualityTen-point scale Aesthetic TaggingClassification Generation Quality Hierarchical classification Evaluation Dimensions General Character -specific Composition Shot Size Expression Costume ... ... Composition Types Number of Light Sources Light Source Position ... Video Clarity Video Noise-Free ... Video-Text Consistency Realism& Plausibility Basic Quality Character Generation Quality Fluid Motion Realism ... Character-Text Object-Text ... Quality Total:10 dimensions Tagging Total:11 dimensions Model Training and Evaluation Prompt Video Annotation Generation VGQA -NetVAQA -NetVTag-Net Aesthetic score Total:31 dimensions Generated level Tag classification Aesthetic Total:21 dimensions Figure 1. Overview of VGA-Bench. We propose a unified benchmark and multi-model framework for video aesthetic and generation quality assessment, comprising a Prompt Suite with design guidelines, a large-scale generated video dataset, a subset of human-annotated data, and three trained evaluation modelsâVAQA-Net, VTag-Net, and VGQA-Netâfor assessing video aesthetic quality, aesthetic tags, and generation quality, respectively. further decomposed into well-defined sub-attributes (e.g., composition, color harmony, lighting usage, motion aes- thetics), enabling fine-grained, interpretable, and holistic evaluation. ⢠A diverse prompt suite and large-scale test dataset: We design 1,016 diverse prompts based on the evaluation framework, covering various scenes, actions, styles, and challenging scenarios. Using 12 state-of-the-art video generation models, we generate a total of 60,000 videos, constructing the largest integrated testing platform to date and supporting fair cross-model comparisons. ⢠Three dedicated multi-task automated evaluators: Based on professional human annotations, we train three spe- cialized neural assessorsâVGQA-Net (for generation quality prediction), VAQA-Net (for aesthetic quality as- sessment), and VTag-Net (for automatic aesthetic tag- ging)âeliminating reliance on external scoring models and enabling end-to-end, consistent, and scalable auto- mated evaluation. ⢠Comprehensive empirical analysis of mainstream mod- els with full open-source commitment: We conduct a systematic evaluation of 12 cutting-edge models using VGA-Bench, revealing their strengths and weaknesses across different dimensions. Upon publication, we will fully release: (1) the complete benchmark suite (includ- ing taxonomy, prompt templates, and annotation data); (2) public API interfaces for all evaluation models; (3) the entire generated video datasetâensuring reproducibility and broad accessibility for the research community. We believe that VGA-Bench serves not only as a rigor- ous evaluation platform but also as a key infrastructure for advancing the next generation of video generation systems with enhanced aesthetic intelligence and artistic controlla- bility. 2. Related Work 2.1. Video Generative Models In recent years, driven by the rapid advancement of deep generative models, text-to-video (T2V) generation technol- ogy [2, 3, 14, 17, 23, 29] has achieved significant break- throughs. Generative architectures represented by diffusion models [2, 13, 19, 30, 43] are now capable of producing high-resolution, temporally coherent, and creatively rich dynamic content from natural language descriptions. These technologies not only show broad application prospects in film production, advertising design, game development, and virtual reality, but are also increasingly integrated into so- cial media creation and personalized content generation pipelines, becoming a core component of the AIGC ecosys- tem. Previously, mainstream T2V model architectures were dominated by U-Net-based designs [2, 10, 36], but their limitationsâsuch as difficulties in modeling long-range de- pendencies and poor scalabilityâsoon became apparent. With the emergence of Sora [21], pure Transformer-based architectures exemplified by DiT [19, 24, 40] have rapidly gained prominence due to their unparalleled global model- ing capability and excellent scalability, and are now becom- ing the dominant paradigm and future direction for high- Total DimensionsAesthetic DimensionsEvaluated ModelsPrompts VBench [15]1614âź1600 VBench2.0 [44]1824âź1600 T2V-CompBench [31]70231400 ChronoMagic-Bench [41]40131649 StoryEval [39]8011423 VGA-Bench(ours)5221121016 Table 1. Comparison of existing evaluation methods for text-to-video generative models end, large-scale video generation models. However, as generation capabilities improve, user ex- pectations have evolved beyond basic technical correct- ness (e.g., absence of artifacts, plausible motion) to in- creasingly emphasize artistic expressiveness and aesthetic qualityâsuch as whether the composition is visually pleas- ing, lighting is skillfully employed, color harmony is well- balanced, or character expressions are graceful. At the same time, the fidelity with which generated content re- flects key visual elements described in the prompt (e.g., âa cyberpunk cityscape at nightâ or âa slow-motion dance un- der soft backlightingâ)âi.e., consistency in visual formal elementsâhas become a crucial metric for assessing model controllability and semantic understanding. In this paper, we evaluate a series of text-to-video mod- els released over the past three years, including both offi- cially open-sourced and commercial models. This compre- hensive evaluation ensures diversity in T2V approaches and provides insightful analysis into their capabilities. 2.2. Evaluation of Video Generative Models Despite continuous performance improvements in video generation models, the scientific and fair evaluation of their comprehensive capabilities remains a key challenge in research. Early assessment methods primarily relied on human ratings or simple technical metrics [12, 20, 34], which are insufficient for quantifying complex human per- ceptual experiences. In recent years, with the develop- ment of the AIGC ecosystem, a series of specialized eval- uation benchmarks for text-to-video generation have been proposed, driving the evolution of assessment frameworks from single metrics toward multi-dimensional and auto- mated paradigms. Among them, V-Bench [15] is the first comprehensive benchmark for video generation, decomposing evaluation into multiple sub-tasksâincluding visual quality, prompt alignment, and motion plausibilityâand incorporating hu- man annotations for holistic scoring.Its successor, V- Bench2 [44], extends the original framework by introduc- ing additional dimensions such as generated duration, frame rate, and style diversity, while also incorporating more generative models and test samples to enhance evaluation breadth and representativeness. Subsequent benchmarks such as ChronoMagic-Bench [41], T2V-CompBench [31], and StoryEval [39] have also evaluated model performance from multiple perspectives, covering aspects like tempo- ral coherence, compositional fidelity, and narrative consis- tency. Table 1 presents a comparative overview of key data characteristics between our work and these existing bench- marks. However, existing benchmarks still suffer from several limitations: ⢠Over-simplified aesthetic evaluation: Most benchmarks treat âaesthetic qualityâ as a single holistic metric, lack- ing fine-grained modeling of specific aesthetic elements such as composition, color harmony, lighting, and visual rhythm. ⢠Reliance on external models: Many metrics depend on pre-trained image or video understanding models for indi- rect inference, which may introduce bias and fail to align with genuine human perception. To address these shortcomings, this paper proposes VGA-Bench, which introduces systematic improvements at three levels: evaluation dimensions, data construction, and model design. We not only refine the sub-dimensions of aesthetic quality but also develop dedicated multi-task eval- uation models specifically designed for video aesthetics and generation quality, enabling more comprehensive and accu- rate assessment of generated videos. 3. VGA-Bench Suite 3.1. Evaluation Dimension Suite 3.1.1. Aesthetic Quality Video aesthetics refers to the perceptual appeal and artis- tic expressiveness conveyed through visual formal ele- mentsâsuch as composition, color, lighting, and mo- tionâin artificially generated dynamic content. Our aes- thetic quality dimensions are adapted from the VADB dataset [26], and specifically include the following ten di- mensions: Overall Score, Composition, Shot Size, Light- ing, Visual Tone, Color, Depth of Field, Expression, Cos- tume, and Makeup. The definitions of these dimensions in real-world videos are thoroughly described in the original dataset paper and thus will not be repeated here. Instead, we focus on their manifestation and interpretability within the context of generated videos. Composition (Com): Generated videos often suffer from âfloating compositionâ or âvisual center offsetâ due to a lack of spatial logic, yet they can achieve surreal arrange- ments that are difficult to realize in real-world filming. Shot Size (S): In real videos, shot selection is con- strained by physical camera setups, whereas generated videos allow free perspective switchingâbut sometimes lack narrative coherence in âcinematic language.â Lighting (Lig): Real-world lighting appears natural and physically plausible, while generated videos may exhibit âuniform illuminationâ or ânon-physical light sources,â leading to stylized yet distorted appearances. Visual Tone (VT): Generated videos demonstrate more consistent tone control, but tend toward âtemplate-like emo- tional expressionâ and lack the subtle transitions present in real lighting dynamics. Color (Col): Colors in real videos are rich and in- fluenced by environmental conditions, whereas generated videos often adopt an âidealizedâ palette, frequently ex- hibiting stylistic biases such as âover-saturationâ or âlow contrast.â Depth of Field (DoF): In real footage, depth of field dy- namically changes with focus; in contrast, generated videos often feature âstatic blurâ effects, lacking the dynamic per- ception of spatial layers. Expression (Exp): Real performances contain micro- expressions and emotional fluctuations, while expressions in generated videos are often âmechanicalâ or âstiff,â failing to capture complex psychological states. Costume (Cos): Costumes in real videos are grounded in cultural and historical context, whereas generated videos frequently produce âstyle mismatchesâ or âinappropriate at- tireâ due to inconsistent semantic reasoning. Makeup (Mak): Real makeup emphasizes detail fidelity and skin-tone harmony, while generated videos often ex- hibit âtexture discontinuitiesâ or âproportional distortionsâ in virtual makeup rendering. We define these aesthetic quality dimensions to guide generative models toward the high-level aesthetic standards observed in real-world videos, enabling systematic eval- uation of their alignment with human perception in as- pects such as composition, lighting, and color. Through fine-grained aesthetic assessment, we aim to examine the modelâs understanding and reconstruction ability regarding advanced visual aesthetics, thereby promoting AIGC sys- tems to more deeply grasp and generate high-quality con- tent that aligns with human aesthetic preferences. 3.1.2. Aesthetic Tagging Aesthetic video tags are structured annotations of identifi- able and quantifiable visual aesthetic features in a video, used to describe artistic expression elements such as com- position style, lighting application, and color properties. Similarly, our aesthetic video tags are adapted from the VADB dataset [26] and supplemented by established pho- tographic theory [4, 7, 25]. We select the following 11 aes- thetic tags: Composition Types, Number of Light Sources, Light Source Position, Light Quality, Light Color, Shot Type, Depth of Field, Saturation, Brightness, Color Tem- perature, and Contrast. Definitions for each dimension are provided below: Composition Types (CT): Refers to the spatial arrange- ment of the main subject and visual elements within the frame, influencing visual balance and narrative guidance. Includes: Rule of Thirds Composition, Symmetrical Com- position, Asymmetrical Composition, Centered Composi- tion, Framing Composition, Leading Lines Composition. Number of Light Sources (NoLS): The number of pri- mary illumination sources in the scene, affecting depth per- ception, atmosphere, and spatial layering. Includes: Single Light Source, Dual Light Sources, Multiple Light Sources. Light Source Position (LSP): The direction of the light relative to the subject, shaping contours, volume, and emo- tional tone. Includes: Back Light, Front-Side Light, Side Light, Bottom Light, Top Light, Front Light, Back-Side Light. Light Quality (LQ): The hardness or softness of lightâsoft light is diffused and even, hard light is sharp and directionalâdirectly influencing mood, texture rendering, and visual texture expression. Includes: Hard Light, Soft Light, Diffused Light. Light Color (LC): The chromatic property of the light source, used to convey emotion, indicate time of day, or cre- ate stylized atmospheres. Includes: White (Neutral) Light, Warm Light, Cool Light, Colored Light. Shot Type (ST): The distance relationship between the camera and the subject, determining information density and psychological engagement with the viewer. Includes: Wide Shot, Full Shot, Medium Shot, Close-Up, Extreme Close-Up. Depth of Field (DoF): The range of spatial area that ap- pears in focus; shallow depth of field emphasizes the sub- ject, while deep depth of field reveals environmental con- textâserving as a key tool for directing visual attention. Includes: Shallow DOF, Deep DOF. Saturation (Sat):The intensity or purity of col- orsâhigh saturation appears vivid and striking, low satu- ration conveys subtlety and restraintâimpacting visual im- pact and emotional expression. Includes: High, Medium, Low. Brightness (Bri): The overall luminance level of the image, affecting readability, mood, and perceived spatial depth. Includes: Bright, Medium, Dark. Color Temperature (Col): The warmth or coolness of the lighting, a critical factor in establishing emotional tone and temporal cues (e.g., dawn vs. dusk). Includes: Cool, Medium, Warm. Contrast (Con): The difference between the brightest and darkest regions in the imageâhigh contrast enhances dramatic tension, while low contrast creates a soft, harmo- nious feel. Includes: High, Medium, Low. We define these aesthetic video tags to construct an inter- pretable and reproducible visual aesthetic language system, enabling evaluation to move beyond subjective judgments such as âwhether it looks good,â toward concrete analysis of where the visual appeal lies and why it is aesthetically effective. Through standardized tag annotation, we can ef- fectively measure a modelâs understanding of photographic aesthetic principles, and provide training signals and op- timization objectives for future generation of high-quality videos that better align with human aesthetic preferences. 3.1.3. Generation Quality Our generation quality assessment further refines the frame- work of V-Bench [15] by categorizing it into three broad categories comprising a total of 31 sub-dimensions. Video- Text Consistency measures the semantic alignment between the generated content and the input prompt; Reality & Plausibility evaluates the credibility of scenes, actions, and physical dynamics with respect to real-world laws; Basic Quality focuses on the intrinsic visual clarity and technical stability of the video itself. The Video-Text Consistency dimension includes: Character-Text Consistency (1), Action-Text Consistency (2), Scene-Text Consistency (3), Object Position-Text Consistency (4), Object Attribute-Text Consistency (5), Object-Text Consistency (6), Video Content-Text Con- sistency (7), Video Speed-Text Consistency (8), Video Style-Text Consistency (9),Camera Movement-Text Consistency (10), Unrealistic Description Imaginative Presentation (11). The Realism & Plausibility dimension includes: Rigid Body Collision Realism (12), Action Realism (13), Scene Realism (14), Weather Representation Realism (15), Time Period Representation Realism (16), Gaseous Motion Re- alism (17), Fluid Motion Realism (18), Gradual Change Motion Realism (19), Object Motion Trajectory Realism (20), Object Realism (21), Character Generation Quality (22), Textual Attribute Representation Realism (23), Video Lighting and Shadow Realism (24), Moving Scene Reason- ableness (25), Overall Realism (26). The Basic Quality dimension includes: Abnormal Light- ing Detection (27), Video Noise-Free (28), Video Clarity (29), Static Content Non-distortion (30), Static Content Sta- bility (31). The definitions of all dimensions are summarized in Ap- pendix. We define these three categories of generation quality and their sub-dimensions to systematically evaluate AIGC videos in terms of semantic understanding, physical com- monsense, and visual fidelity. This ensures that models not only âunderstandâ the input prompts but also generate con- tent that is logically coherent and visually natural. Through fine-grained decomposition, our framework provides clear optimization directions for model improvement, and pro- motes the evolution of generative systems toward greater realism, controllability, and practical usability. 3.2. Prompt Suite 3.2.1. Prompt Design Prompt design is a critical component in text-to-video eval- uation.A clear and precise prompt can effectively re- duce stochastic interference during generation, enabling the model to focus on user intent and thus more faithfully reflect its semantic understanding and content generation capabil- ities. To this end, our core design principle is: the targeted aesthetic or generation quality dimension must be explicitly specified in the prompt, ensuring that the model can per- ceive and respond to the intended attribute. For example, a video should only be used for composition assessment if the prompt explicitly includes descriptions such as âcomposed using the rule of thirdsâ; otherwise, the corresponding di- mension should not be included in the evaluation. Building upon this, we further emphasize prompt di- versity: prompts should vary in length, cover both single- dimension and multi-dimensional scenarios, and span a wide range of themes and scenes to enhance the representa- tiveness and robustness of the test set. Based on these principles, we construct a systematic Prompt Suite containing 1,016 carefully designed prompts, distributed as follows: 200 for aesthetic quality dimensions, 220 for aesthetic tag dimensions, and 596 for generation quality dimensions. Each evaluation dimension is covered by at least 50 prompts, ensuring statistical validity. Further- more, to accommodate different testing requirements, we provide two lightweight subsets: one with 508 prompts and another with 127 prompts. All versions maintain balanced dimension distribution and diverse prompt lengths, and sup- port combinations of 1 to 5 dimensions per prompt, facili- tating flexible fine-grained analysis and efficient lightweight evaluation. 3.2.2. Use of LLMs For the aesthetic quality dimensions, we select high-scoring real video comments from the VADB dataset [26] in the corresponding dimensions and extract descriptive sentences that emphasize specific aesthetic attributes (e.g., âbalanced compositionâ, âsoft lightingâ). These human aesthetic feed- backs are used as input to guide the LLM in generating prompts with similar expressive styles and semantic focus. Aesthetic Quality ModelďźCogvideoX PromptďźThe market stalls are colorful, with a well-arranged layout that features strong yet harmonious color contrasts, providing a visually pleasing experience. DimentionsďźColor + Composition Aesthetic Tagging ModelďźLTXVideo PromptďźA diver swims in the deep sea with side lighting from a flashlight, creating a hard beam of light from a single source. The scene is framed through a submarine window and shot at a distance, with a cool color temperature. DimentionsďźFraming Composition + Single Light Source + Side Light + cool color temperature Generation Quality - Basic Quality ModelďźMochi PromptďźA long shot of the study, with bookshelves that remain stable and undistorted, static without any drifting, and the image is sharp. DimentionsďźVideo Clarity + Static Content Non-distortion + Static Content Stability Generation Quality ModelďźLatte PromptďźOn a clear and cloudless day, the parking lot is full of cars. Keywords: scene + weather DimentionsďźScene-Text Consistency + Scene Realism + Weather Representation Realism Figure 2. Prompts for the three core dimensions, their correspond- ing sub-dimensions, and example generated videos. For the aesthetic tag dimensions, we directly feed the categorical labels of each sub-dimension (e.g., Back Light, Shallow DOF, High Saturation) into the LLM, instructing it to generate natural language descriptions that explicitly include the given keyword while maintaining semantic co- herence. For the generation quality dimensions, we first sum- marize each sub-dimension into one or more representa- tive keywordsâfor example, âObjectâ for both Object-Text Consistency and Object Realism, and âGaseous Motionâ for Gaseous Motion Realismâensuring that each keyword covers one or two core attributes. Subsequently, these key- words are used to prompt the LLM to generate text in- puts that precisely elicit the target characteristics. Notably, for the Basic Quality sub-dimensions, we set the keywords as single adjectives and directly incorporate them into the prompt as modifiersâfor instance, âVideo Clarityâ is real- ized in the prompt as a directive such as âgenerate a clear videoâ. Concrete examples are illustrated in Figure 2. 3.3. Human Annotation To train the multi-task evaluation models of this benchmarkâVGQA-Net, VAQA-Net, and VTag-Netâwe adopt an âexpert-led + crowd-assistedâ annotation paradigm. First, domain experts from the film and video industry perform exemplary annotations on a subset of samples based on predefined dimension definitions and rating guidelines.Subsequently, a crowdsourced team completes the labeling of the remaining data by following these exemplars.Finally, experts conduct batch-wise Aesthetic Quality Overall Scoreďź4.8 Compositionďź5.6 Lightingďź4.4 Aesthetic Tagging Number of Light Sourcesďź Single Light Source Brightnessďź Bright Shot Typeďź Close-Up Generation Quality Static Content Stabilityďź 2: Relatively unstable, with noticeable changes in content, but the overall content can still be recognized as consistent. Abnormal LightingDetectionďź 2: Obvious lighting issues, such as overexposure or abnormal highlight placement, which begin to affect the viewing experience. Light Qualityďź Soft Light Depth of Fieldďź Shallow DOF Color Temperatureďź Warm Overall Scoreďź4.2 Compositionďź4.4 Colorďź4 Makeupďź3.6 Time Period Representation Realismďź 2: Fairly realistic, but with some flaws. Character Generation Quality ďź 4: Good quality, with only minor, barely noticeable anomalies. Figure 3. Examples of human annotations for the three core di- mensions. sampling audits; if annotation errors are identified, the entire batch is rejected and re-labeled to ensure consistent quality. All annotations strictly follow the prompt design principles: raters score only those dimensions explicitly mentioned in the prompt, avoiding subjective inference on unmentioned attributes. For the aesthetic quality and aesthetic tag dimensions, we adopt the standardized scoring guidelines from the VADB dataset [26], with rating boundaries defined through both textual descriptions and example videos. Each aes- thetic sub-dimension is scored on a 0â10 scale, and the final score is computed as the average of three independent an- notators. For aesthetic tags, treated as a multi-label classi- fication task, each sample is independently labeled by three annotators, and labels are retained only if at least two agree (âmajority votingâ). For each evaluation dimension under generation qual- ity, we design specific assessment questions paired with structured response options. These options represent dis- tinct levels of qualityâfunctioning effectively as ordinal scoresâtailored to the semantic meaning of the respective dimension. Example (Object-Text Alignment): This dimension in- cludes four response options: ⢠-1 (Invalid Question): A universal option present in most dimensions. Annotators select this to discard a sample when: The prompt for a âconsistencyâ dimension lacks a specified target (e.g., object, scene). The prompt for a ârealismâ dimension intentionally describes an unrealistic scenario. ⢠1 (Completely Inconsistent): The object exhibits no align- ment with the text description. ⢠2 (Partially Consistent): The object exhibits characteris- â /âĄ/⢠Prompt Video Encoder Clip Text Encoder M L P Result Linear ⢠â /⥠⢠Figure 4.Architecture of (1)VAQA-Net, (2)VTag-Net, and (3)VGQA-Net. The video encoders in VAQA-Net and VTag-Net are those trained in the first stage of the VADB dataset [26] us- ing a dual-text encoder with dynamic fusion module for language comments and aesthetic tags; these encoders are frozen during the training phase in this work. Compared to the former two models, VGQA-Net includes an additional CLIP [27] branch before the input MLP. tics of the described target. ⢠3 (Fully Consistent): The object perfectly matches the text description. Likewise, the results for generation quality are deter- mined using the âmajority votingâ principle. 4. Experiments and Results 4.1. VGA Evaluation Network The network architectures of VAQA-Net, VTag-Net, and VGQA-Net are illustrated in Figure 4. Despite significant progress in visual fidelity of gener- ated videos, they still lag far behind human-created con- tent in core aspects of aesthetic intelligenceâsuch as inten- tional artistry, emotional authenticity, and cultural context embedding. Real-world videos, especially professionally produced films and documentaries, exhibit deliberate artis- tic decisions in composition, lighting, and narrative pac- ing, reflecting a deep understanding of human perception and cultural norms. These qualities make real-world video data an indispensable resource for training models to recog- nize and reason about âmeaningful beauty,â going beyond merely capturing superficial visual patterns. Therefore, we initialize VAQA-Net and VTag-Net with the video encoder pre-trained in the first stage of the VADB dataset [26], along with its associated training data and pa- rameters from real video scoring and tagging tasks, en- abling the models to inherit aesthetic understanding ac- quired from professional cinematography. In the second stage, we fine-tune the models on an extended dataset that includes 1,300 generated videos from 12 mainstream gen- erative models, each paired with high-quality human anno- tations. Evaluation is conducted on a separate set of 400 generated videos: for the aesthetic tag task, standard accu- Table 2. 5-Class Accuracy of VAQA-Net Dim.Acc.â %Dim.Acc.â % Overall76.9Com73.6 S72.4Lig69.5 VT67.9Col71.1 DoF69.5Exp74.8 Cos77.4Mak71.8 Table 3. Accuracy of VTag-Net (Top-2 Predicted Tags Match Ground Truth) Dim.Acc.â %Dim.Acc.â % CT45NoLS77 LSP66LQ74 LC80ST57 DoF82Sat89 Bri82Col71 Con90 Table 4. Accuracy of VGQA-Net Num.Acc.â %Num.Acc.â %Num.Acc.â % 186.01282.32383.0 278.41391.02482.1 378.91492.62580.7 476.31581.62685.6 572.41687.82789.9 674.81784.42881.2 774.31871.72985.6 875.81970.13076.5 976.72071.23181.8 1081.42183.5 1167.32280.6 racy (Acc) is used as the metric; for aesthetic scoring, the 0â10 scale is discretized into five levels, and five-class ac- curacy is computed. Results are presented in Table 2 and Table 3. In contrast, our VGQA-Net is fully focused on gener- ated videos. To comprehensively evaluate its cross-model generalization capability, we select two representative gen- erative models from each year over three years, resulting in six models in total: HunyuanVideo [18], LTXVideo [11], Mochi [33], Latte-1 [24], CogVideoX [40], and Show- 1 [42], covering 12,000 generated videos. The model is trained on videos produced by three of these models and tested on videos from the remaining three, ensuring no over- lap between training and test sets in terms of model prove- nance. Accuracy (Acc) is used as the evaluation metric, and results are presented in Table 4. 4.2. VGA-Bench Evaluation Results We evaluate all generative models on the trained VAQA- Net, VTag-Net, and VGQA-Net to assess their performance Aesthetic Quality ModelďźLatte PromptďźA lightning striking atop of eiffel tower, dark clouds in the sky. Light scoreďź2.8 Color scoreďź2 ModelďźMochi PromptďźA lightning striking atop of eiffel tower, dark clouds in the sky. Light scoreďź4 Color scoreďź4 Generation Quality ModelďźLtxvideo PromptďźHeavy downpour Weather Representation Realismďź 1: Completely unrealistic. ModelďźLatte PromptďźHeavy downpour Weather Representation Realismďź 3: Highly realistic. Figure 5. Comparison of generated videos from different models using the same dimension-aligned prompt. Aesthetic scores and generation quality levels are derived from human annotations. across various aesthetic and quality dimensions. To ensure a fair and unbiased ranking, all generated videos used for evaluation are held out from the model training process, pre- venting data leakage from influencing the results. The eval- uated models are ordered by release date in ascending order, including: Stable Video Diffusion(SVD) [2], AnimateDiff- v2 [10], LaVie [38], Show-1 [42], ModelScope [36], CogVideoX [40], Latte-1 [24], Mochi [33], LTXVideo [11], HunyuanVideo [18], Wan2.1 [35], and Sora2 [21]. For the aesthetic quality dimension, models are ranked by average score, with higher scores indicating better per- formance. For the aesthetic tag dimension, classification ac- curacy is used as the metric, and models with higher mean accuracy rank higher. For the generation quality dimension, models are ranked by average level rating, where a higher average level indicates better generation quality and thus a better (higher) rank. The results are obtained by normalizing and averaging across all sub-dimensions within the three main dimensions, as shown in Table 5. For more detailed results of each model across all sub-dimensions, please refer to the Appendix. 4.3. User Study We conducted a user study in the form of a questionnaire. Since most general users have not received professional training, and it is practically infeasible to train every par- ticipant in a large-scale survey, we only invited non-expert users to perform ranking evaluations on the outputs of 12 generative models across two dimensions: aesthetic quality and generation quality. In the experiment, we randomly selected five prompt sets from the Prompt Suite for aesthetic quality and another ModelAes. ScoreTag Cla.Gen. Level SVD [2]0.200.320.55 AnimateDiff [10]0.360.300.49 LaVie [38]0.340.310.47 Show-1 [42]0.290.320.28 ModelScope [36]0.310.300.49 CogVideoX [40]0.410.390.55 Latte-1 [24]0.350.310.45 Mochi [33]0.210.340.54 LTXVideo [11]0.220.340.47 Hunyuan [18]0.450.360.55 Wan2.1 [35]0.46 0.380.53 Sora2 [21]0.500.180.54 Table 5. Performance comparison of state-of-the-art text-to-video generation models on aesthetic score (Aes. Score), tag classifica- tion accuracy (Tag Cla.), and generation level (Gen. Level) met- rics. Recall@1Recall@3Recall@5 Aes.0.100.700.80 Gen.0.500.800.83 Table 6. Comparison between human and model-based rankings on aesthetic and generation quality dimensions. Recall@5, Re- call@3, and Recall@1 are reported based on 40 user surveys. five from the Prompt Suite for generation quality. For each prompt, we collected the corresponding videos generated by 12 different models, forming comparative sequences. Par- ticipants were asked to rank the videos according to two subjective yet representative criteria: âTo what extent does the video reflect the beauty described in the text?â and âHow accurately does the video depict the textual content?â As a reference, we normalize the evaluation scores across all sub-dimensions and compute their average to de- rive an overall ranking of the models. We then compare the overlap between human rankings and our model-based rankings using Recall@5, Recall@3, and Recall@1. The results from 40 collected questionnaires are summarized in Table 6. 5. Conclusion We propose VGA-Bench, a fine-grained AIGC video eval- uation benchmark comprising 52 sub-dimensions, 1,016 prompts, and over 60,000 annotated videos. Through our dedicated evaluators (VAQA-Net, VTag-Net, and VGQA- Net), this work delivers human-aligned insights into state- of-the-art models and systematically integrates artistic prin- ciples into the evaluation pipeline. Marking a paradigm shift from âhow realâ to âhow beautiful,â VGA-Bench not only quantifies key elements like composition, color, and lighting, but also paves the way for models to achieve gen- uine perceptual aesthetics and expressiveness. VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation Supplementary Material Figure 6. Radar chart comparing the performance of various video generation models across three evaluation dimensions. The concentric circular grids represent score levels, ranging from 0.1 at the center to 0.9 outward in 0.1 intervals. Higher values indicate better performance. Each model is represented by a closed polygon with distinct colors and line styles for easy comparison. 6. Dimension of Generation Quality The meanings of the dimensions of generation quality are summarized in Table 9. 7. Overview of Generative Model Performance The performance comparison of various generative models across the three core dimensions is shown in Figure 6. 8. Comprehensive Evaluation Results of Gen- erative Models across Sub-Dimensions in VGA-Bench 8.1. Performance Comparison of Generative Mod- els on Aesthetic Quality Dimensions Table 7 presents the scoring results of 12 mainstream gener- ative models across sub-attributes under the aesthetic qual- ity dimension. All scores are generated by VAQA-Net through automated evaluation, reflecting the visual appeal and artistic expressiveness of the videos produced by these models. 8.2. Comparison of Aesthetic Tag Prediction Capa- bilities Table 8 reports aesthetic tag prediction accuracies using VTag-Net, evaluating the modelsâ capabilities in under- standing and generating complex aesthetic semantics. 8.3. Performance Comparison of Generative Mod- els on the Generation Quality Dimension Due to the extensive sub-dimensions within generation quality, we conduct automated annotations using VGQA- Net and present the performance comparisons across three separate tables: Table 10 details 11 metrics on video-text consistency and spatio-temporal alignment; Table 11 as- sesses 14 realism metrics concerning physical laws and real-world commonsense; and Table 12 evaluates 6 tech- nical dimensions reflecting basic low-level visual fidelity. Table 7. Performance Comparison of Generative Models on Aesthetic Quality Dimensions ModelOverallComSSLigVTColDoFExpCosMak Show-1 [42]0.2900.4150.3830.3670.3570.3970.3300.2940.3170.294 Latte-1 [24]0.3450.4720.4350.4200.4100.4450.3930.3130.3400.296 LaVie [38]0.3410.4680.4350.4210.4070.4420.3960.3090.3400.293 AnimateDiff [10]0.3560.4750.4460.4270.4150.4600.4000.3170.3640.310 ModelScope [36]0.3120.4450.4130.3970.3810.4300.3770.2870.3340.266 Wan2.1 [35]0.459 0.5520.5230.5230.5210.5210.5030.4360.4590.412 SVD [2]0.2040.3050.2600.2290.2330.2770.1910.2070.1510.164 CogVideoX [40]0.4050.4800.4540.4370.4310.4530.3990.3890.4200.354 LTXVideo [11]0.2140.3130.2750.2380.2410.2850.2080.1890.1420.125 Mochi [33]0.2110.3060.2690.2400.2340.2830.2030.1980.1470.147 Hunyuan [18]0.4520.5310.5050.5070.5040.5120.4850.455 0.4650.430 Sora2 [21]0.5040.5590.5320.5430.5640.5600.5400.5200.5320.452 Table 8. Comparison of Aesthetic Tag Prediction Capabilities ModelCTNoLSLSPLQLCSTDoFSatBriColCon Show-1[42]0.170.470.190.460.170.320.540.290.420.480.35 Latte-1[24]0.160.470.190.460.120.250.600.290.350.460.35 LaVie[38]0.170.470.190.460.110.250.660.290.410.390.35 AnimateDiff[10]0.18 0.470.190.460.060.230.660.290.370.390.35 ModelScope[36]0.170.470.190.460.010.230.600.290.430.400.35 Wan2.1[35]0.180.600.190.570.230.340.680.380.470.580.36 SVD[2]0.190.500.240.460.280.220.520.290.250.460.37 CogVideoX[40]0.180.710.190.510.330.380.620.290.480.630.37 LTXVideo[11]0.150.520.270.460.360.200.570.290.250.640.36 Mochi[33]0.170.490.240.460.440.200.550.290.250.690.35 Hunyuan[18]0.160.660.210.480.230.400.680.270.450.470.35 Sora2[21]0.130.320.140.290.190.320.270.110.150.190.10 Table 9. Number and explanation of different assessment dimensions TypeNum.Assessment DimensionDescription Video-Text Consistency 1Character-Text Consistency Whether specific characters in the video match the text descrip- tion (e.g., Elon Musk should appear as the correct individual). 2Action-Text Consistency Whether actions in the video match the text description (e.g., running, jumping), focusing solely on the action regardless of the subject. 3Scene-Text Consistency Whether scenes in the video match the described settings (e.g., hospital, school), including identifiable scene elements. 4Object Position-Text Consistency Object positions refer to relative placement based on camera orientation (e.g., if âa motorcycle is to the left of a bus,â they should appear on corresponding sides of the video frame). 5Object Attribute-Text Consistency Object attributes include descriptive features like color, shape, and texture. 6Object-Text Consistency Whether objects in the video can be correctly identified as those mentioned in the text. 7Video Content-Text Consistency Overall alignment where every textual description should be ac- curately generated. 8Video Speed-Text Consistency Whether video speed matches textual descriptions (current sam- ples only include slow-motion). 9Video Style-Text Consistency Whether artistic styles mentioned in text (e.g., Van Gogh, Pi- casso) are recognizable in the video. 10Camera Movement-Text Consistency Whether camera movements described in text (e.g., pan left, tilt right) are properly executed. 11Unrealistic Description Imaginative Presentation When text describes unrealistic scenarios (e.g., âan astronaut riding a horse in spaceâ), whether the video presentation aligns with imaginative expectations. Realism & Plausibility 12Rigid Body Collision Realism Whether rigid body collisions in videos appear physically plau- sible. 13Action RealismWhether actions could realistically be performed. 14Scene Realism Whether scenes appear sufficiently realistic when no special style is specified in text. 15Weather Representation RealismWhether weather conditions appear realistic. 16Time Period Representation RealismWhether time-period representations appear authentic. 17Gaseous Motion Realism Whether gas dynamics (smoke, vapor) appear physically accu- rate. 18Fluid Motion RealismWhether fluid movements appear physically plausible. 19Gradual Change Motion Realism Whether gradual transformations (balloon inflation, plant growth) appear physically accurate. 20Object Motion Trajectory Realism Whether object movement paths follow physically plausible dy- namics. 21Object RealismWhether objects appear sufficiently realistic. 22Character Generation QualityWhether human characters appear sufficiently realistic. 23Textual Attribute Representation Realism Whether object attributes (color, shape, texture) match real- world appearances. 24Video Lighting and SGQAow RealismWhether lighting and sGQAows appear physically accurate. 25Moving Scene Reasonableness Whether scene transitions during camera movements maintain proper perspective. 26Overall RealismWhether the entire video looks realistic overall. Basic Quality 27Abnormal Lighting Detection Videos should avoid lighting artifacts (overexposure, abnormal flares). 28Video Noise-FreeVideos should exhibit no noticeable noise artifacts. 29Video ClarityWhether video resolution is sufficiently sharp. 30Static Content Non-distortion Stationary objects shouldnât distort abnormally during camera movement. 31Static Content Stability Stationary objects shouldnât distort abnormally over time (tem- poral consistency). Table 10. Evaluation Results on Video-Text Consistency Model1234567891011 Show-1[42]0.0370.3640.3620.3460.5050.4660.1960.0270.0300.4330.250 Latte-1[24]0.0520.2850.7700.5650.8750.6720.5980.0000.0340.5000.288 LaVie[38]0.0760.2120.8520.5120.8870.6820.5980.0070.0640.6120.340 AnimateDiff[10]0.0240.1680.7830.5000.9770.7700.5470.0000.0140.8610.400 ModelScope[36]0.0210.2080.7540.4640.9640.8160.5670.0140.0370.806 0.480 Wan2.1[35]0.0260.8400.6720.629 0.9000.7950.6280.2950.1840.5950.448 SVD[2]0.0660.9580.8930.5720.9340.9170.5910.0710.0500.6550.410 CogVideoX[40]0.0550.9360.8630.5120.9290.8880.5990.0450.0650.7200.419 LTXVideo[11]0.0320.5610.5780.5420.7810.6400.5380.0380.0430.5330.303 Mochi[33]0.0390.8970.8030.7010.9510.8130.6790.0390.0430.6320.387 Hunyuan[18]0.0560.8560.7990.5920.970 0.8090.6000.0340.0410.6200.350 Sora2[21]0.0240.8420.7900.5320.9340.8510.6030.0860.0940.7150.425 Table 11. Evaluation Results on Realism & Plausibility Model121314151617181920212223242526 Show-1[42]0.0280.2670.2940.3770.2060.3440.3770.6090.1270.3510.2370.3890.2840.2820.227 Latte-1[24]0.0380.4350.5110.6400.3880.3500.4880.8250.1500.4280.5130.6000.4130.4630.598 LaVie[38]0.0350.4020.5290.6180.3450.4150.4900.8300.1450.4080.5300.6070.4400.5330.598 AnimateDiff[10]0.0250.4900.5350.6440.4150.3900.5000.7600.1450.4660.6110.6770.4850.5560.547 ModelScope[36]0.0700.4840.5220.6760.4050.3750.4900.8700.1800.5080.5740.6600.4900.5740.567 Wan2.1[35]0.1630.5170.5210.5390.4810.5750.5000.9380.2100.5270.5210.5720.4940.5330.628 SVD[2]0.0250.4960.5860.6720.4450.6550.5000.6600.1650.605 0.6520.6400.4400.6710.591 CogVideoX[40]0.0300.5120.5510.6770.4610.5970.4700.8750.1390.5900.6310.6020.4200.6410.599 LTXVideo[11]0.0350.4650.5200.6050.3910.4710.4440.8640.1310.4970.5200.5950.4240.5090.623 Mochi[33]0.0470.5060.5400.6650.4150.5000.5600.937 0.1740.5350.6530.6310.4660.5780.786 Hunyuan[18]0.0350.5260.5630.7840.4280.5260.5470.8640.1720.5930.7140.7000.4380.6160.694 Sora2[21]0.0850.5200.5530.6600.4500.5850.4900.7800.1900.6200.6040.6940.4500.6680.603 Table 12. Evaluation Results on Basic Visual Quality Model2728293031 Show-1[42]0.3690.0970.2120.2570.285 Latte-1[24]0.7660.1580.3650.4470.742 LaVie[38]0.8610.2070.4520.5000.717 AnimateDiff[10]0.8430.2140.5260.5250.680 ModelScope[36]0.6330.2320.4540.4820.664 SVD[2]0.898 0.2920.6500.5110.733 CogVideoX[40]0.8550.3040.592 0.5150.748 LTXVideo[11]0.8230.2380.4760.5180.808 Mochi[33]0.8240.2230.5000.5110.805 Hunyuan[18]0.8760.2530.5560.5780.797 Sora2[21]0.9060.2150.5400.5180.720 Acknowledgments This work was supported by the Ant Group Research Fund, the National Natural Science Foundation of China under Grant No. 62072014, and the Opening Project of the State Key Laboratory of General Artificial Intelligence, BIGAI/Peking University, Beijing, China (Project No. SKLAGI2025OP01). References [1] Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lu Ë ci Ě c, and Cordelia Schmid. Vivit: A video vision transformer. In Proceedings of the IEEE/CVF inter- national conference on computer vision, pages 6836â6846, 2021. 1 [2] Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023. 1, 2, 8, 10, 12 [3] Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dock- horn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with la- tent diffusion models. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, pages 22563â22575, 2023. 1, 2 [4] Blain Brown. Cinematography: theory and practice: im- age making for cinematographers and directors. Routledge, 2016. 4 [5] Mathilde Caron, Hugo Touvron, Ishan Misra, Herv Ě e J Ě egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9650â9660, 2021. 1 [6] Fei-Long Chen, Du-Zhen Zhang, Ming-Lun Han, Xiu-Yi Chen, Jing Shi, Shuang Xu, and Bo Xu. Vlp: A survey on vision-language pre-training. Machine Intelligence Re- search, 20(1):38â56, 2023. 1 [7] Maya Deren. Cinematography: the creative use of reality. Daedalus, 89(1):150â167, 1960. 4 [8] Zi-Yi Dou, Aishwarya Kamath, Zhe Gan, Pengchuan Zhang, Jianfeng Wang, Linjie Li, Zicheng Liu, Ce Liu, Yann Le- Cun, Nanyun Peng, et al. Coarse-to-fine vision-language pre-training with fusion in the backbone. Advances in neural information processing systems, 35:32942â32956, 2022. 1 [9] Zhe Gan, Linjie Li, Chunyuan Li, Lijuan Wang, Zicheng Liu, Jianfeng Gao, et al. Vision-language pre-training: Basics, re- cent advances, and future trends. Foundations and TrendsÂŽ in Computer Graphics and Vision, 14(3â4):163â352, 2022. 1 [10] Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text- to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725, 2023. 1, 2, 8, 10, 12 [11] Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, Dudu Moshe, Eitan Richardson, Eran Levin, Guy Shiran, Nir Zabari, Ori Gordon, et al. Ltx-video: Realtime video latent diffusion.arXiv preprint arXiv:2501.00103, 2024. 1, 7, 8, 10, 12 [12] Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation met- ric for image captioning. In Proceedings of the 2021 confer- ence on empirical methods in natural language processing, pages 7514â7528, 2021. 1, 3 [13] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. Advances in neural information processing systems, 33:6840â6851, 2020. 1, 2 [14] Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion mod- els. arXiv preprint arXiv:2210.02303, 2022. 1, 2 [15] Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. Vbench: Comprehensive bench- mark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21807â21818, 2024. 1, 3, 5 [16] Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. Musiq: Multi-scale image quality transformer. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5148â5157, 2021. 1 [17] Levon Khachatryan, Andranik Movsisyan, Vahram Tade- vosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Text2video-zero: Text- to-image diffusion models are zero-shot video generators. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15954â15964, 2023. 1, 2 [18] Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603, 2024. 1, 7, 8, 10, 12 [19] Junlong Li, Yiheng Xu, Tengchao Lv, Lei Cui, Cha Zhang, and Furu Wei. Dit: Self-supervised pre-training for docu- ment image transformer. In Proceedings of the 30th ACM international conference on multimedia, pages 3530â3539, 2022. 2 [20] Yuanxin Liu, Lei Li, Shuhuai Ren, Rundong Gao, Shicheng Li, Sishuo Chen, Xu Sun, and Lu Hou. Fetv: A bench- mark for fine-grained evaluation of open-domain text-to- video generation. Advances in Neural Information Process- ing Systems, 36:62352â62387, 2023. 1, 3 [21] Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jian- feng Gao, et al. Sora: A review on background, technology, limitations, and opportunities of large vision models. arXiv preprint arXiv:2402.17177, 2024. 1, 2, 8, 10, 12 [22] Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3202â3211, 2022. 1 [23] Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan.Videofusion: Decomposed diffusion mod- els for high-quality video generation.arXiv preprint arXiv:2303.08320, 2023. 1, 2 [24] Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Zi- wei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. Latte: Latent diffusion transformer for video generation. arXiv preprint arXiv:2401.03048, 2024. 1, 2, 7, 8, 10, 12 [25] Mustafa Yousry Matbouly. Quantifying the unquantifiable: the color of cinematic lighting and its effect on audienceâs impressions towards the appearance of film characters. Cur- rent Psychology, 41(6):3694â3715, 2022. 4 [26] Qianqian Qiao, DanDan Zheng, Yihang Bo, Bao Peng, Heng Huang, Longteng Jiang, Huaye Wang, Jingdong Chen, Jun Zhou, and Xin Jin. Vadb: A large-scale video aesthetic database with professional and multi-dimensional annota- tions. arXiv preprint arXiv:2510.25238, 2025. 3, 4, 5, 6, 7 [27] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, pages 8748â8763. PmLR, 2021. 7 [28] Javier Selva, Anders S Johansen, Sergio Escalera, Kamal Nasrollahi, Thomas B Moeslund, and Albert Clap Ě es. Video transformers: A survey. IEEE Transactions on Pattern Anal- ysis and Machine Intelligence, 45(11):12922â12943, 2023. 1 [29] Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022. 1, 2 [30] Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equa- tions. arXiv preprint arXiv:2011.13456, 2020. 1, 2 [31] Kaiyue Sun, Kaiyi Huang, Xian Liu, Yue Wu, Zihan Xu, Zhenguo Li, and Xihui Liu. T2v-compbench: A comprehen- sive benchmark for compositional text-to-video generation. In Proceedings of the Computer Vision and Pattern Recogni- tion Conference, pages 8406â8416, 2025. 3 [32] Shixiang Tang, Yizhou Wang, Lu Chen, Yuan Wang, Sida Peng, Dan Xu, and Wanli Ouyang. Human-centric founda- tion models: Perception, generation and agentic modeling. arXiv preprint arXiv:2502.08556, 2025. 1 [33] Genmo Team.Mochi 1. https://github.com/ genmoai/models, 2024. 1, 7, 8, 10, 12 [34] Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Rapha Ě el Marinier, Marcin Michalski, and Sylvain Gelly. Fvd: A new metric for video generation. 2019. 1, 3 [35] Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video gen- erative models. arXiv preprint arXiv:2503.20314, 2025. 1, 8, 10, 12 [36] Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571, 2023. 1, 2, 8, 10, 12 [37] Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhil- iang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mo- hammed, Saksham Singhal, Subhojit Som, et al. Image as a foreign language: Beit pretraining for vision and vision- language tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19175â 19186, 2023. 1 [38] Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, et al. Lavie: High-quality video generation with cascaded latent diffusion models. International Journal of Computer Vision, 133(5):3059â3078, 2025. 1, 8, 10, 12 [39] Yiping Wang, Xuehai He, Kuan Wang, Luyao Ma, Jianwei Yang, Shuohang Wang, Simon Shaolei Du, and Yelong Shen. Is your world simulator a good story presenter? a consecu- tive events-based benchmark for future long video genera- tion. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 13629â13638, 2025. 3 [40] Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiao- han Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 1, 2, 7, 8, 10, 12 [41] Shenghai Yuan, Jinfa Huang, Yongqi Xu, Yaoyang Liu, Shaofeng Zhang, Yujun Shi, Rui-Jie Zhu, Xinhua Cheng, Jiebo Luo, and Li Yuan. Chronomagic-bench: A bench- mark for metamorphic evaluation of text-to-time-lapse video generation. Advances in Neural Information Processing Sys- tems, 37:21236â21270, 2024. 3 [42] David Junhao Zhang, Jay Zhangjie Wu, Jia-Wei Liu, Rui Zhao, Lingmin Ran, Yuchao Gu, Difei Gao, and Mike Zheng Shou. Show-1: Marrying pixel and latent diffusion models for text-to-video generation. International Journal of Com- puter Vision, 133(4):1879â1893, 2025. 1, 7, 8, 10, 12 [43] Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, pages 3836â3847, 2023. 1, 2 [44] Dian Zheng, Ziqi Huang, Hongbo Liu, Kai Zou, Yinan He, Fan Zhang, Lulu Gu, Yuanhan Zhang, Jingwen He, Wei- Shi Zheng, et al. Vbench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness. arXiv preprint arXiv:2503.21755, 2025. 3