Paper deep dive
TDMM-LM: Bridging Facial Understanding and Animation via Language Models
Luchuan Song, Pinxin Liu, Haiyang Liu, Zhenchao Jin, Yolo Yunlong Tang, Zichong Xu, Susan Liang, Jing Bi, Jason J Corso, Chenliang Xu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/22/2026, 5:04:16 AM
Summary
TDMM-LM introduces Open3DFaceVid, a large-scale synthetic dataset of 80 hours of facial videos paired with 3DMM parameters and text descriptions. The framework utilizes a geometry VQ-VAE to discretize facial motion into tokens, enabling a unified language model to perform bidirectional tasks: Motion2Language (interpreting facial motion into natural language) and Language2Motion (generating 3D facial animation from text prompts).
Entities (5)
Relation Signals (4)
TDMM-LM → utilizes → Open3DFaceVid
confidence 100% · Building on this dataset, we probe language models for bidirectional competence
Geometry VQ-VAE → discretizes → 3DMM
confidence 95% · We train a geometry VQ-VAE that discretizes 3DMM sequences into a perceptually coherent codebook
TDMM-LM → performs → Motion2Language
confidence 95% · probe language models for bidirectional competence over facial motion via two complementary tasks: (1) Motion2Language
TDMM-LM → performs → Language2Motion
confidence 95% · probe language models for bidirectional competence over facial motion via two complementary tasks: (2) Language2Motion
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Text-guided human body animation has advanced rapidly, yet facial animation lags due to the scarcity of well-annotated, text-paired facial corpora. To close this gap, we leverage foundation generative models to synthesize a large, balanced corpus of facial behavior. We design prompts suite covering emotions and head motions, generate about 80 hours of facial videos with multiple generators, and fit per-frame 3D facial parameters, yielding large-scale (prompt and parameter) pairs for training. Building on this dataset, we probe language models for bidirectional competence over facial motion via two complementary tasks: (1) Motion2Language: given a sequence of 3D facial parameters, the model produces natural-language descriptions capturing content, style, and dynamics; and (2) Language2Motion: given a prompt, the model synthesizes the corresponding sequence of 3D facial parameters via quantized motion tokens for downstream animation. Extensive experiments show that in this setting language models can both interpret and synthesize facial motion with strong generalization. To best of our knowledge, this is the first work to cast facial-parameter modeling as a language problem, establishing a unified path for text-conditioned facial animation and motion understanding.
Tags
Links
- Source: https://arxiv.org/abs/2603.16936v1
- Canonical: https://arxiv.org/abs/2603.16936v1
Trouble viewing inline? Open PDF directly →
Full Text
52,680 characters extracted from source content.
Expand or collapse full text
TDMM-LM: Bridging Facial Understanding and Animation via Language Models † Luchuan Song 1 , * Pinxin Liu 1 , Haiyang Liu 2 , Zhenchao Jin, Yolo Yunlong Tang 1 , Zichong Xu 1 , Susan Liang 1 , Jing Bi 1 , Jason J Corso 3,4 , Chenliang Xu 1 1 University of Rochester 2 University of Tokyo 3 University of Michigan 4 Voxel51 lsong11@ur.rochester.edu,pliu23, yunlong.tang, sliang22@ur.rochester.edu, haiyangliu1997@gmail.com, jjcorso@eecs.umich.edu, chenliang.xu@rochester.edu Open3DFaceVid User: Please describe expressions and movements of the 3D facial sequence in the video. Agent: The person is speaking in a slightly happy tone, keeping the head almost still. User: Please describe expressions and movements of the 3D facial sequence in the video. Agent: A clear look of surprise on the face, and the head moved up and down. User: The person is speaking with the expression of strong anger. User: The person is speaking with slight happiness on his face. User: A person is speaking with smile on the face. And his head is nodding. Figure 1. Overview of the proposed Open3DFaceVid dataset and 3D facial understanding/animation pipeline. The left panel visualizes the Open3DFaceVid corpus, which covers a wide range of identities, emotions, and speaking styles generated via text-to-video (T2V) models. The right panel illustrates our interactive 3D facial interface: given a 3DMM sequence, the user prompts the agent to describe expressions and head motion in natural language, and the agent returns fine-grained, parameter-based interpretations. In the reverse direction, the agent is able to condition on user prompts to generate new 3DMM trajectories with controllable emotion and pose. Please refer to https://songluchuan.github.io/TDMM-LM/ for visualization results and datasets. Abstract Text-guided human body animation has advanced rapidly, yet facial animation lags due to the scarcity of well- annotated, text-paired facial corpora. To close this gap, we leverage foundation generative models to synthesize a large, balanced corpus of facial behavior. We design prompts suite covering emotions and head motions, gener- ate about 80 hours of facial videos with multiple generators, and fit per-frame 3D facial parameters, yielding large-scale (prompt and parameter) pairs for training. Building on this dataset, we probe language models for bidirectional com- petence over facial motion via two complementary tasks: (1) Motion2Language: given a sequence of 3D facial pa- † Project Leader. * Equal contribution. rameters, the model produces natural-language descrip- tions capturing content, style, and dynamics; and (2) Lan- guage2Motion: given a prompt, the model synthesizes the corresponding sequence of 3D facial parameters via quan- tized motion tokens for downstream animation. Extensive experiments show that in this setting language models can both interpret and synthesize facial motion with strong gen- eralization. To best of our knowledge, this is the first work to cast facial-parameter modeling as a language problem, establishing a unified path for text-conditioned facial ani- mation and motion understanding. 1. Introduction Multimodal large language models (MLLMs) [3, 15, 17, 20, 24, 33, 35, 44, 45, 47, 61, 62, 64, 67, 70] have sig- 1 arXiv:2603.16936v1 [cs.CV] 14 Mar 2026 22.4% 37.6% 37.9% .6% .7% .8% (b) T2V ContributionsProportion(c) Emotion Words (d) Prompts Words (a) Control Categories 57.1% 10.23%10.20% 7.55% 5.94% 8.98% Figure 2. The analysis of the Open3DFaceVid dataset. We summarize the control categories induced by prompts and their corresponding video counts, broken down by underlying T2V backbones. We further visualize the vocabulary with word clouds, separately for emotion- related terms and for full-text prompts, to highlight the diversity and saliency of affective descriptors. nificantly advanced visual understanding through joint rea- soning over images, audio, and language. Modern sys- tems [2, 19, 23, 51, 57] achieve state-of-the-art results on a wide range of perception tasks [7, 9, 25, 31, 34, 59]. Re- cent domain-specialized MLLMs [14, 37, 48, 60, 69] fur- ther demonstrate strong potential for human-centered visual reasoning. However, fine-grained facial-behavior understanding re- mains fundamentally limited. Token inefficiency is a major bottleneck: existing V-LLMs must process every frame as a large set of image tokens. To reduce cost, models down- sample video or sparsely sample keyframes. While accept- able for coarse actions, this is detrimental for facial mo- tion, where expressions unfold over only a few frames and temporal continuity is crucial. Micro-expressions such as brow raises, lip twitches, or brief smirks often disappear when frames are dropped; even retained frames yield hun- dreds of image tokens per second, restricting temporal con- text and forcing models to ignore subtle dynamics. The result is a systematic bias toward static, neutral facial be- havior. Emotion imbalance in existing training corpora amplifies this issue. Large-scale MLLMs [51, 60, 69] are mostly trained on in-the-wild videos (e.g., YouTube, Tik- Tok, VoxCeleb [5, 38]), which overwhelmingly feature neu- tral, frontal talking heads. High-intensity or atypical expres- sions (e.g. pouting, frowning, laughing, smirking) are rare. Coupled with temporal sparsity from downsampling, mod- els learn a narrow, low-variance facial prior and struggle to perceive or generate expressive motion. These limitations raise a central question: “Can fine- grained facial emotion and movement be represented using low-dimensional geometric signals (e.g., 3D facial param- eters) that preserve subtle temporal variation while avoid- ing redundant visual tokens?” A major obstacle is the lack of class-balanced, richly annotated facial-motion data. Ex- isting datasets [6, 13, 16, 30, 56] rely on studio capture and provide coarse, one-hot emotion labels, lacking the di- versity and open-form descriptions needed for language- driven modeling. To overcome this limitation, we intro- duce Open3DFaceVid, a scalable synthetic pipeline that generates diverse facial videos using multiple T2V mod- els [55, 58]. We curate a lexicon of∼200 emotion and facial-action descriptors (e.g., grin, pout, smirk, squint), uniformly sample all categories to ensure balanced cover- age, and extract 3DMM parameters [8] for every clip. This yields a large corpus of paired videos, text descriptions, and high-quality 3D facial-motion trajectories. Dataset charac- teristics are shown in Fig. 3. With this dataset providing the necessary coverage and annotation richness, we explore a modeling paradigm that bypasses image tokens entirely. We train a geometry VQ- VAE [53] that discretizes 3DMM sequences into a percep- tually coherent codebook, producing compact geometry to- kens that replace image tokens in the MLLM input space. This structured representation preserves subtle dynamics, removes visual redundancy, and drastically reduces token consumption. Finetuning a LLM on paired motion–text samples enables it to directly reason over 3D facial motion. Built on this alignment interface, our framework sup- ports two complementary directions using the same LLM: (1) Motion2Language: interpreting geometry-token se- quences to produce natural-language descriptions of emo- tion, intensity, micro-expressions, and head motion; (2) Language2Motion:generating expressive 3D facial- motion trajectories directly from user prompts, enabling fine-grained, text-driven facial animation. Extensive exper- iments validate the effectiveness of our dataset, geometry- token representation, and unified modeling framework. We demonstrate strong performance across facial-motion un- derstanding and language-driven animation. In summary, our contributions include: • We present Open3DFaceVid, an∼80-hour synthetic facial-motion corpus generated using multiple foundation T2V models. It provides the largest collection to date of paired Text2Face annotations with diverse emotions, sub- jects, and intensity variations. • We extend the LLM to the Motion2Language setting, enabling natural-language interpretation of 3D facial mo- tion. Given a facial-motion token sequence, the model produces compact, expressive descriptions of emotions, 2 micro-expressions, and head movements. • We further introduce a Language2Motion framework that generates 3D facial motion from natural-language prompts. By conditioning an autoregressive geometry decoder on word-level LLM embeddings, users can pre- cisely control fine-grained facial dynamics through text. 2. Related Work Text-to-Motion Generation Text–motion alignment un- derpins much of modern motion understanding [28, 29, 54], with models such as MotionCLIP [52] and TMR [40] map- ping language descriptions to dynamic motion sequences. Building on this foundation, recent work has shifted to- ward text-to-motion generation: autoregressive models [11, 12, 32, 43, 65, 66] tokenize motion and decode it in a language-like space. Supported by large-scale body motion datasets [11, 26, 41, 42], these approaches enable natural and expressive full-body movement generation. In contrast, progress in facial animation has lagged behind, largely due to the absence of emotion-rich, text-aligned corpora. Text-to-Video Generation Text-to-video generation has made significant progress in recent years. Early diffusion- based [10, 21, 68, 72, 73] frameworks factorize the gener- ation via intermediate image conditioning, improving sta- bility and quality [46]. More recent works target human- centric video or long shot sequences: Identity-Preserving T2V ensures consistent human identity while generating high-fidelity video [63], and Text-to-Multi-Shot Video Gen- eration addresses transitions and scene continuity [18]. Commercial systems such as Sora2 [39]/Veo3 [58] have fur- ther demonstrated text-driven open-world video generation. 3. Open3DFaceVid We build our dataset by programmatically prompting a fam- ily of Text-to-Video (T2V) models [49] to synthesize fa- cial videos that can be reconstructed in our 3D setting. The prompt space covers a large attribute lexicon, including sub- ject descriptors (e.g., male, female, young lady) and affec- tive cues (e.g., joy, anger, smile), to encourage broad cover- age over identities and emotions. To avoid overfitting to the biases of a single generator, we instantiate the pipeline with multiple underlying T2V models (e.g., Wan-2.1/2.2 [55], Open-Sora [71], HuMo [4], Veos [58]), which expands the range of styles and mitigates model-specific artifacts. Prompts Construction The prompt design factorizes into facial appearance and video dynamics. On the facial side, we sample a rich space of affective descriptions, varying both emotion type and intensity, and augment them with frequent micro-expressions such as blinking, pouting, and related facial actions to increase behavioral diversity. On the video side, we employ a small set of templated prompts that constrain camera motion, encouraging stable framing. For example, instructions like “The frame remains steady, with the head and shoulders centered” are applied to keep the subject fixed in view and reduce cinematic jitter. We try to maintain the balance between the number of facial emo- tion categories and the number of subject categories in the prompts. Programmatically generated templates are then re- fined with the ChatGPT-4.1 [1], yielding natural, conversa- tional descriptions that better resemble daily language while preserving the intended control condition. Human-Centric T2V Synthesis We apply the polished prompts as control condition to synthesize a large collec- tion of facial videos. To avoid collapsing the motion dis- tribution onto the prior of any single generator, we instan- tiate the pipeline with a suite of base T2V models. For each backbone, the facial-attribute portion of the prompt is kept standardized, while only the video-related clauses are lightly adapted to match model-specific preferences in framing and cinematography. Our corpus contains roughly 60K synthetic clips, each about 4–6 seconds. To narrow the gap to real data, we augment this set with about 10K wild clips and derive prompt-style annotations with a vi- sion–language model (Gemini [51]). Geometry Facial Estimation We estimate facial motion with 3D Morphable Model (3DMM) [8] parameters, fitted to monocular videos in our corpus to recover per-clip iden- tity, expression, and head pose. Specifically, we adopt the FLAME [22] model to regress facial parameters. 4. Language-Motion Alignment Base on Open3DFaceVid, we learn a shared space be- tween 3D facial motion (3DMM) and natural language, en- abling both interpretation and synthesis of expressive fa- cial behavior. The alignment pipeline has three key com- ponents: (1) the Geometry VQ-VAE operates on facial ge- ometry rather than continuous 3DMM parameters, map- ping sequences into the discrete token space; (2) an LLM- based motion interpreter that directly decodes geometry- token sequences into natural-language descriptions, sup- porting Motion2Language understanding; (3) a shared lan- guage–motion transformer that conditions geometry-token prediction on word-level embeddings, enabling controllable Language2Motion synthesis. Geometry VQ-VAE Shown in Fig. 4, we quantize facial dynamics into a discrete latent space. Rather than dis- cretizing on 3DMM parametersm which can map similar expressions to different coefficient patterns. We operate di- rectly on reconstructed facial geometry, ensuring that visu- ally similar expressions are encoded with consistent tokens. This design mitigates the many-to-one ambiguity in which multiple expression codes map to nearly facial geometry. Motion2Language To translate 3D facial motion into nat- ural language, we utilize geometry tokenizer to tokenize fa- 3 Prompt: The young man has an earnest and slightly concernedexpression. His eyebrows are raised and furrowedas he speaks. Prompt: The man's face is contorted with intense anger. His brows are furrowed, and his eyes are wide. He repeatedly opens his mouth. Prompt: The man has a wide, enthusiastic smile, with his mouth open and eyes crinkled, conveying genuine happiness. He bobs his head. Wan2.1 Wan2.2 HuMo Prompt: The man issurprised. His eyebrows are raised high, and mouth is open in an "O" shape. His head bobs slightly forward and back. Prompt: The young girl has a proudexpression, with mouth open as she speaks and eyes focused. Her head moves in a natural motion. Figure 3. Dataset overview. Top two rows: starting from a fixed text prompt, we vary the random seed and emphasize different prompt keywords to modulate facial identity and video attributes, showcasing subjects across different genders. Bottom three rows: we recover FLAME facial parameters and pair the resulting trajectories with the corresponding prompt, forming Text–3DMM dataset. Codebooks 1 4 9 72 14 972 ℒ 1 ( mesh ) Figure 4. Geometry-aware facial tokenization learning. We quan- tize facial expression codes into a discrete codebook and enforce reconstruction in mesh space. The input facial expression codes are mapped to code indices, decoded back to FLAME meshes (bot- tom), and supervised with anL 1 loss on vertex positions cial geometry, where discrete tokens are fed as symbolic ob- servations to the LLM. Instead of relying on image inputs, the LLM directly consumes sequences of geometry tokens, producing rich free-form descriptions of emotions, intensi- ties, micro-expressions (e.g., blinking, pouting), and head dynamics. Each training instance provides a geometry- token sequence and an accompanying textual description. Geometry EncoderText Tokenizer Large Language Model Text Tokenizer ...... ...... ...... ... ... ......... Agent: The person has a happinessexpression as speaking. Thehead is still. Z 1 Z 2 Z i Z i+1 Z i+2 Z 4096 T 1 T i T i+1 T i+3 T i+2 T n-1 T n T * 1 T * 2 T * 3 T * i+1 T * i T * n-1 T * n [UserInteraction]: Summarize the face in one short sentence; rely only on the codes: Create a minimal description of the face encoded by these tokens: Describe the person's facial appearance in one brief sentence: From the token stream, describe the face video motion: Figure 5. Motion2Language. Geometry sequences are encoded into discrete facial tokens by the geometry encoder and fed, to- gether with text tokens from the user prompt, into a LLM. Condi- tioned only on these geometry tokens, the agent generates natural- language descriptions of expression/head motion, enabling inter- active question to answering about 3D facial behavior. To improve linguistic coverage without altering semantics, we expand each annotation with multiple paraphrases pro- duced by an auxiliary LLM. As shown in Fig. 5, the LLM is instruction-tuned to map geometry-token observations to natural-language responses conditioned on task templates. 4 [UserInteraction]: The person has a warmand engaging facial expression. He offers a gentle, closed-lip smile. The man has a warm and gentle smile, with his eyes crinkling kindly at the corners. The person has a serious and concerned expression. Her brow is slightly furrowed, T5 Text Encoder Geometry Encoder AutoregressionTransformer ...... ...... Z 1 Z 2 Z i Z i+1 Z i+2 Z 4096 Geometry Decoder ...... ...... Z 1 Z 2 Z i Z i+1 Z i+2 Z 4096 ...... ... T 1 T i T i+1 T i+3 T i+2 T n-1 T n Token Masks Figure 6. Language2Motion.The user provides a natural- language description of the desired facial behavior (top left). The text tokenizer converts the prompt into word-level tokens, while a paired 3D facial sequence is encoded into discrete geometry to- kens by the geometry encoder. The autoregressive transformer predicts future geometry tokens conditioned on the text prefix. Methods Cor E ↑ Cor M ↑ Cor I ↑USER E ↑ USER M ↑ [GPT-4][Human] HumanOmni [50]1.841.171.092.041.00 Gemini2.5 VLM [56]2.452.913.512.883.41 Ours4.02 3.35 3.634.293.79 Table 1. Quantitative evaluation of Motion2Language. Mean 1–5 correctness scores from GPT-4 and human raters for emotion (Cor E , USER E ), motion (Cor M , USER M ) and intensity (Cor I ). We bold the best. Our geometry-token model consistently outper- forms HumanOmni and Gemini-2.5 VLM across the metrics. Language2Motion Compared to full-body scenarios, text- driven facial motion generation remains considerably more challenging. Available datasets are smaller, facial actions are subtler, and conventional text-to-motion pipelines com- press the entire prompt into a global embedding, losing the token-level cues that dictate fine-grained muscular dynam- ics. To leverage the linguistic granularity, we introduce a word-level language prefix that injects pretrained LLM em- beddings into an autoregressive facial-motion transformer. Importantly, this prefix is processed by the text encoder which takes the user’s text prompt, produces token-level embeddings, and conditions the motion decoder without modifying the vocabulary of the language model. As illus- trated in Fig. 6, this prefix-based fusion preserves the struc- ture of the prompt and permits individual words to steer lo- calized facial movements, enabling controllable and seman- tically aligned 3D facial-motion generation. 5. Experiments 5.1. Dataset Details The Open3DFaceVid has a total duration of 81 hours and contains 57.2K video clips ranging from 4 to 6 seconds. Methods Cor E ↑ Cor M ↑ Cor I ↑USER E ↑ USER M ↑ [GPT-4][Human] HumanOmni [69]3.661.591.423.821.27 Gemini2.5 VLM [51]4.213.17 3.924.173.82 Ours4.02 3.353.634.293.79 Table 2. Quantitative evaluation of Motion2Language with nat- ural images input. We employ the same correctness metric as in Table 1, while replacing geometry-image inputs with natural facial images to better adapt the evaluation setting to VLMs. Methods L 2 ↓FD↓Tok↑Cor E ↑USER↑ [Parameters][Language] T2M-X [27]0.47147.590.6712.213.40 T2M-GPT [65]0.22637.040.8953.573.91 Ours 0.219 31.75 0.9204.133.95 Table 3. Quantitative evaluation of Language2Motion. Compar- ison with T2M-X and T2M-GPT on parameter-space measures (L 2 /FD/Tok) and language-based metrics (Cor E /USER).L 2 /FD assess expression/pose fidelity, Tok is token-level accuracy. The videos in datasets are standardized to 25 FPS and the resolution is resized 618× 360. We employ 32 H200 GPUs to synthesize videos in T2V setting using Wan2.2 (5B and 14B), HuMo (17B), Open-Sora (11B) and Wan2.1 (14B), consuming roughly 400 GPU Hours in total. We obtain Veo2/3 samples via its API interface. Due to the high per- clip generation cost, we include a modest number of Veo2/3 generated videos in our corpus. For face estimation, we use the FLAME template to regress per-frame expression, and head rotation (yaw/roll/pitch) parameters within videos. We present the model implementation details in the Appendix. 5.2. Baselines Given the challenges of constructing text–motion datasets, there exist few directly comparable prior methods. We therefore adopt the most relevant techniques as baselines across both Motion2Language and Language2Motion. Motion2Language.We evaluate against strong vi- sual–language models applied directly to video frames. (1) HumanOmni [69], a human-centric VLM designed for holistic audio–visual reasoning. We treat it as a vision–only model by sampling video frames from dataset and prompt- ing it to describe head motion and emotion. During infer- ence, it operates entirely in pixel space. (2) Gemini VLM, a general-purpose multimodal model [51]. We uniformly sample video frames and provide an instruction prompt ask- ing for descriptions of head pose and affective state. Language2Motion.We draw baselines from text-to- human-motion generation. (3) T2M-X [27], a transformer- based text-to-motion model that represents motion as dis- crete tokens and autoregressively predicts pose sequences; we adapt its facial-motion branch as a baseline. (4) T2M- GPT [65], a GPT-style generative model trained on discrete motion codes. It directly models the distribution of tok- enized motion trajectories conditioned on text descriptions. We follow their discrete VQ-token formulation and train the 5 User:Pleasegivemeacompact,fluentdescriptionofthegeometryfacehere(onlyfocusonemotion andheadpose). Ground-Truth:Ayounggirlspeakswithaslightboredomexpression.Herheadremainsrelatively still,withonlyslight,naturalmovementsasshetalks. Gemini-2.5Pro:Thevideodisplaysa3Dgeometricmodelofahumanfacethatisanimatedtoopen andcloseitsmouth,asifspeakingorsinging. HumanOmni:Smile. OurMotion2Language:Thepersonspeaksasthebrowsfurrowinaworriedexpression,theheadis almoststill. User:Pleasesummarizethefacialfeatures(emotionandheadmotion)withinonesentencefromthe geometrysequence. Ground-Truth:Thewoman’sexpressionisoneofintenseangeranddisbelief.Herbrowisdeeply furrowedasshespeaks.Herheadmovesinsharp,punctuatedmotions,shakingslightly. Gemini-2.5Pro:Thegeometricfaceanimatesthroughaspeakingorsingingmotionbyopeningand closingitsmouthwhileperformingasubtleside-to-sideheadturn. HumanOmni:Openmouthwide. OurMotion2Language:Thepersonspeaksinkeepingtheamplifiedangrylook,theheadisslight shakinginspeaking. User:Pleasewriteabriefnatural-languagedescriptionofthefaceemotionandheadmotionfromthe geometrysequence. Ground-Truth:Thewoman'sfacialexpressionissurpriseoremphasis.Herheadmovessubtly, noddingandturningtopunctuateherwords. Gemini-2.5Pro:The3Dmodelofthefaceanimatesthroughaspeech-likemotion,openingand closingitsmouth.Simultaneously,theheadperformsasubtleside-to-siderotation,turningslightly fromitsrighttowardsitsleft. HumanOmni:Smile. OurMotion2Language:Thepersonrecitesabriefline,withaverystrongjoyfuldemeanor.The headmotionisslightnodding. Figure 7. Qualitative comparison for Motion2Language. For rep- resentative clips, we show the input frames/geometry. The Gem- ini/HumanOmni operate on rendered geometry, while our model relies on 3D facial tokens. Our approach produces more accurate descriptions of emotion/motion, closely matching GT annotations. facial-motion pathway on our dataset for fair comparison. 5.3. Quantitative and Qualitative Evaluations Motion2Language. We validate: (1) geometry-token se- quences preserve sufficient expressive information for accu- rate semantic interpretation, and (2) the compact tokeniza- tion provides higher efficiency than image-token VLMs. To quantify these aspects, we measure correctness in ex- pression (Cor E ), motion (Cor M ), and intensity (Cor I ) us- ing GPT-4 as an automatic judge on the 2K test clips, and collect human ratings (USER E , USER M ) on 300 samples. Moreover, to ensure a fair comparison with the baseline methods, we use natural facial videos/images as inputs, as shown in Tab. 2. We observe that commercial models such as Gemini achieve leading performance in the natural- image setting. Nevertheless, it is important to note that our method uses only a single token as input for each image, which is substantially fewer than the number of input to- The woman has a bright and friendly expression, smiling widely. Her head is mostly stationary. The man begins with a serious and focused expression. Then leans forward surprise expression. T2M - X LM - Listener Ours Ours LM - Listener T2M - X Figure 8. Qualitative comparison for Language2Motion. Given prompts, we visualize generated 3D facial motion from different method. Our model produces expressions and head poses that more faithfully follow the described affect Train-Corpus Cor E ↑ Cor M ↑ Cor I ↑USER E ↑ USER M ↑ [GPT-4][Human] MEAD [56]3.031.293.592.921.18 YouTube [Gemini] 2.743.923.763.143.65 Open3DFaceVid3.973.44 3.854.053.75 Table 4.Ablation on training corpus for Motion2Language. Identical architectures are trained on MEAD, YouTube (Gemini- annotated) and Open3DFaceVid. The metrics defined in Sec. 5.3. kens required by Gemini. Fig. 7 compares model outputs under identical prompts. HumanOmni and Gemini-2.5 Pro, both trained primarily on natural images, struggle to interpret temporal dynamics of facial motions, resulting in correctness scores close to 1. This reflects a domain gap rather than a failure of the task itself. In contrast, our model consistently recovers the cor- rect emotion and motion semantics from 3DMM-derived tokens, demonstrating that structured geometric representa- tions retain the key information needed for facial-behavior reasoning. Furthermore, our method encodes each frame with a single geometry token instead of 300–500 visual to- kens, yielding a more efficient and temporally responsive understanding pipeline. Language2Motion. We test: (1) geometry-token genera- 6 (1) Open3DFaceVid(2) MEAD(3) YouTube Happy Angry Surprise Sad Fear Disgust Neutral Excited Cry Interest Worry Proud Shy Happy Angry Sad Fear Contempt Disgust Surprise Neutral Calm Neutral Happy Sad Laugh Joyful Cool Disappoint Depressed Flat Smile Angry Total Labels = 187 Total Labels = 8 Total Labels = 37 ... ... Figure 9. Emotion coverage across datasets. Top-row: emotion-label distributions for (1) Open3DFaceVid, (2) MEAD [56], and (3) YouTube-derived clips, showing the number and balance of emotion categories (187/8/37 labels, respectively). Bottom-row: 2D projections of the corresponding facial (expression+pose) video embeddings, where Open3DFaceVid exhibits rich, well-separated clusters, while MEAD and YouTube provide coarser and less diverse affective coverage. Each point in the t-SNE [36] plots corresponds to the facial- expression features of a single video clip. Due to color variety limitations, we are unable to display all categories. Train-Corpus L 2 ↓FD↓Tok↑Cor E ↑USER↑ [Parameters][Language] MEAD [56]0.217 32.140.8923.562.88 YouTube [Gemini]0.25933.470.9103.173.59 Open3DFaceVid0.22632.48 0.9153.903.92 Table 5. Ablation on training corpus for Language2Motion. Iden- tical models are trained on MEAD, YouTube (Gemini-annotated), and Open3DFaceVid. Training on Open3DFaceVid yields the best trade-off between geometric fidelity and text–motion alignment. tion conditioned on language can faithfully follow the in- tended facial semantics, and (2) the predicted 3DMM tra- jectories exhibit realistic temporal dynamics despite using a compact discrete representation. To validate these as- pects, we adopt quantitative and human-centric metrics. Text–motion consistency is assessed through human ratings (USER) on 300 samples. Low-level reconstruction fidelity is measured usingL 2 distance on expression and pose. Mo- tion realism is evaluated via Fr ́ echet Distance (FD), com- paring feature distributions of generated and ground-truth sequences. Token prediction accuracy (Tok) reflects dis- crete code fidelity within our VQ token space. Finally, we compute expression correctness (Cor E ) using the trained Motion2Language model as an external semantic evaluator, providing an independent measure of alignment. As shown in Tab. 3, our method achieves consistently higher scores across all metrics, supporting both claims: the model generates semantically aligned facial motion and maintains high-quality temporal structure. T2M-GPT per- forms reasonably well due to its autoregressive architecture but falls short on semantic metrics, highlighting the benefit of conditioning on a stronger LLM backbone. Qualitative The young man on continues speaking, with his lips pushed forward in a lingering pout. Open3DFaceVid YouTube MEAD Figure 10. The qualitative ablation on training corpus for Lan- guage2Motion. We visualize facial motion generated by models trained on ablation datasets. The MEAD- and YouTube-trained models miss the “pout” behavior, whereas the our model produces clear lip protrusion (red arrows), matching the description. examples in Fig. 8 show that our model responds sensitively to nuanced instructions (e.g., transitioning from “focused” to “surprised”) and that higher-capacity variants capture expressive affect (e.g., “joyful”) more distinctly than com- peting methods. 5.4. Ablation Study Dataset Ablation. To understand the value of our synthetic corpus, we compare Open3DFaceVid against two repre- sentative alternatives: (1) the MEAD [56] studio-captured 7 User:Pleasegiveabriefnatural-languagedescriptionofthefaceemotionandheadmotionfromthe inputgeometrysequence. Ground-Truth:Theman'sexpressiontransitionsfromworriedandconcernedtooneofsudden shockanddisbelief.Hiseyebrows,initiallyfurrowed,andhismouthopensinan"O"ofsurprise. TrainonMEAD:Thepersonisspeakingwithaslightfear.Theheadisalmoststill. TrainonYouTube:Thepersonspeakswithanengagedandexpressiveface.Theeyesarewide beforebrieflyclosingandreopeningashetalks.Theheadremainsrelativelystillandforward-facing. TrainonOpen3DFaceVid:Thepersonspeaksfromsubtleworriedtosurprised.Theheadismove fromlefttoright. Figure 11. Qualitative ablation on corpus for Motion2Language. Given user query, we show responses from models trained on MEAD, YouTube and Open3DFaceVid. The MEAD- and YouTube-trained variants produce generic or partially correct de- scriptions, while the Open3DFaceVid model most accurately cap- tures the subtle worry-to-surprise transition and head motion. Training Iterations Tokens Accuracy Motion2LanguageLanguage2Motion Qwen3sLlamas Figure 12. The scaling behavior of Motion2Language and Lan- guage2Motion. Left: token accuracy over training iterations for Motion2Language with Qwen3 backbones of different sizes (.6B- 32B). Right: token accuracy for Language2Motion with LLaMA- based backbones (0.7B-3B). dataset with one-hot emotion labels, converted into textual prompts; and (2) an in-the-wild YouTube set automatically annotated by Gemini. These two sources reflect the two dominant data paradigms in facial-behavior research, con- trolled lab capture and unconstrained internet videos, both of which are known to lack the fine-grained expressive cov- erage required for language-driven modeling. Evaluation Protocol. We follow a TMR-style [40] proto- col to assess whether each dataset provides semantically reliable text–motion pairs. For every dataset, we train a CLIP-like text–motion encoder on its paired annotations and project all samples into a shared embedding space. If a dataset contains high-quality, well-aligned annotations and exhibits sufficient expression diversity, its embeddings should form well-separated clusters corresponding to dif- ferent facial categories. Conversely, datasets with limited variation or noisy annotations should collapse into overlap- ping or ambiguous regions. The t-SNE visualizations in Fig. 9 confirm this behav- ior. Open3DFaceVid displays clear and balanced clusters across 187 labels, reflecting both the richness of our curated lexicon and the reliability of the T2V generation pipeline in following prompts. MEAD, despite being high-quality The person speaks without hesitation, paired with a high-energy happy expression. The person speaks without hesitation, paired with a high-energy angryexpression. The person speaks without hesitation, paired with a lowhappy expression. Figure 13. Language-controlled expressive ablation. We modify one keyword in prompt and let the Language2Motion model gen- erate the corresponding facial motion. The two happy prompts produce different intensity level, while the angry prompt yields clearly distinct mouth shapes, demonstrating fine-grained text con- trol over emotional style. capture, offers only eight coarse labels and limited expres- sive variation, leading to tighter, less informative clusters. YouTube data, even after Gemini annotation, remains heav- ily skewed toward neutral expressions and exhibits consid- erable overlap, highlighting the difficulty of recovering sub- tle facial behaviors from uncontrolled videos. Downstream comparisons.We next train both Mo- tion2Language and Language2Motion models on each dataset. For Motion2Language, the broader coverage of Open3DFaceVid yields noticeably richer semantic under- standing (Fig. 11), even though its head-motion range is slightly narrower than the YouTube variant (reflected by Cor M in Tab. 4). For Language2Motion, models trained on our dataset produce clearer expression articulation, the lin- gering pout in Fig. 10 for instance, demonstrating stronger text–motion alignment.Although MEAD occasionally shows lower feature-space distances due to its simpler la- bel space, models trained on Open3DFaceVid achieve the highest USER scores (Tab. 5), indicating superior percep- tual quality and semantic fidelity. Overall, these ablations show that neither controlled studio capture nor in-the-wild videos offer the diversity or balance needed for expressive facial-motion learning. Our synthetic corpus provides significantly stronger annota- tion quality, more comprehensive expressive coverage, and more reliable text–motion alignment, benefiting both Mo- tion2Language and Language2Motion tasks. Additional analysis of dataset bias and synthetic-data effectiveness is 8 provided in the Appendix. Scaling Behavior. We present the scaling behavior by plotting token accuracy for Motion2Language and Lan- guage2Motion in Fig. 12. For Motion2Language, perfor- mance saturates around 4B–8B parameters, and larger back- bones offer limited gains. In contrast, Language2Motion benefits more strongly from scale, with larger models yield- ing higher token-generation accuracy. Language2Motion Keywords. As shown in Fig. 13, we obtain facial motions by changing only a single word in the prompt. The edit can adjust style intensity (“high-energy”- “low”) or swap the affective category itself (“happy”- “angry”). These results demonstrate both the model nu- anced language understanding and its ability to translate fine-grained prompts into controllable facial motion. 6. Conclusion We study a previously underexplored direction, aligning language with 3D facial motion. We first analyze a key bottleneck, the scarcity of facial datasets with paired text and address it by synthesizing a large-scale corpus of fa- cial videos with T2V models. Building on this resource, we study two complementary settings: Motion2Language and Language2Motion. We hope our dataset and bidirectional framework will serve as a foundation for future research in 3D facial motion understanding and animation. References [1] Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ah- mad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 3 [2] Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923, 2025. 2 [3] Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Zhenyu Tang, Li Yuan, et al. Sharegpt4video: Improving video understand- ing and generation with better captions. Advances in Neural Information Processing Systems, 37:19472–19495, 2024. 1 [4] Liyang Chen, Tianxiang Ma, Jiawei Liu, Bingchuan Li, Zhuowei Chen, Lijie Liu, Xu He, Gen Li, Qian He, and Zhiyong Wu.Humo: Human-centric video generation via collaborative multi-modal conditioning. arXiv preprint arXiv:2509.08519, 2025. 3 [5] Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. Voxceleb2: Deep speaker recognition.arXiv preprint arXiv:1806.05622, 2018. 2 [6] Andrzej Czyzewski, Bozena Kostek, Piotr Bratoszewski, Jozef Kotus, and Marcin Szykulski. An audio-visual cor- pus for multimodal automatic speech recognition. Journal of Intelligent Information Systems, 49(2):167–192, 2017. 2 [7] Ali Diba, Mohsen Fayyaz, Vivek Sharma, Manohar Paluri, J ̈ urgen Gall, Rainer Stiefelhagen, and Luc Van Gool. Large scale holistic video understanding. In European Conference on Computer Vision, pages 593–610. Springer, 2020. 2 [8] Bernhard Egger, William AP Smith, Ayush Tewari, Stefanie Wuhrer, Michael Zollhoefer, Thabo Beeler, Florian Bernard, Timo Bolkart, Adam Kortylewski, Sami Romdhani, et al. 3d morphable face models—past, present, and future. ACM Transactions on Graphics (ToG), 39(5):1–38, 2020. 2, 3 [9] Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6202–6211, 2019. 2 [10] Daiheng Gao, Shilin Lu, Shaw Walters, Wenbo Zhou, Ji- aming Chu, Jie Zhang, Bang Zhang, Mengxi Jia, Jian Zhao, Zhaoxin Fan, et al. Eraseanything: Enabling con- cept erasure in rectified flow transformers. arXiv preprint arXiv:2412.20413, 2024. 3 [11] Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating diverse and natural 3d human motions from text. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5152–5161, 2022. 3 [12] Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang, and Li Cheng. MoMask: Generative masked mod- eling of 3D human motions. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. 3 [13] Naomi Harte and Eoin Gillen. Tcd-timit: An audio-visual corpus of continuous speech. IEEE Transactions on Multi- media, 17(5):603–615, 2015. 2 [14] Yining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng, Yilun Du, Zhenfang Chen, and Chuang Gan. 3d-llm: In- jecting the 3d world into large language models. Advances in Neural Information Processing Systems, 36:20482–20494, 2023. 2 [15] Bin Huang, Xin Wang, Hong Chen, Zihan Song, and Wenwu Zhu. Vtimellm: Empower llm to grasp video moments. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 14271–14280, 2024. 1 [16] Chao Huang, Yapeng Tian, Anurag Kumar, and Chenliang Xu. Egocentric audio-visual object localization. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22910–22921, 2023. 2 [17] Chao Huang, Zeliang Zhang, Jiang Liu, Ximeng Sun, Jialian Wu, Xiaodong Yu, Ze Wang, Chenliang Xu, Emad Barsoum, and Zicheng Liu. Directional reasoning injection for fine- tuning mllms. arXiv preprint arXiv:2510.15050, 2025. 1 [18] Ozgur Kara, Krishna Kumar Singh, Feng Liu, Duygu Cey- lan, James M. Rehg, and Tobias Hinz. Shotadapter: Text- to-multi-shot video generation with diffusion models, 2025. 3 [19] Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603, 2024. 2 [20] Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, 9 and Jianfeng Gao. Llava-med: Training a large language- and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems, 36:28541–28564, 2023. 1 [21] Leyang Li, Shilin Lu, Yan Ren, and Adams Wai-Kin Kong.Set you straight:Auto-steering denoising tra- jectories to sidestep unwanted concepts.arXiv preprint arXiv:2504.12782, 2025. 3 [22] Tianye Li, Timo Bolkart, Michael J Black, Hao Li, and Javier Romero. Learning a model of facial shape and expression from 4d scans. ACM Trans. Graph., 36(6):194–1, 2017. 3 [23] Bin Lin, Zhenyu Tang, Yang Ye, Jiaxi Cui, Bin Zhu, Peng Jin, Jinfa Huang, Junwu Zhang, Yatian Pang, Munan Ning, et al.Moe-llava: Mixture of experts for large vision- language models. arXiv preprint arXiv:2401.15947, 2024. 2 [24] Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual repre- sentation by alignment before projection. In Proceedings of the 2024 Conference on Empirical Methods in Natural Lan- guage Processing, pages 5971–5984, 2024. 1 [25] Ji Lin, Chuang Gan, and Song Han. Tsm: Temporal shift module for efficient video understanding. In Proceedings of the IEEE/CVF international conference on computer vision, pages 7083–7093, 2019. 2 [26] Jing Lin, Ailing Zeng, Shunlin Lu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, and Lei Zhang.Motion-x: A large-scale 3d expressive whole-body human motion dataset. Advances in Neural Information Processing Systems, 36: 25268–25280, 2023. 3 [27] Mingdian Liu, Yilin Liu, Gurunandan Krishnan, Karl S Bayer, and Bing Zhou. T2m-x: Learning expressive text- to-motion generation from partially annotated data, 2024. 5 [28] Pinxin Liu, Haiyang Liu, Luchuan Song, and Chenliang Xu. Intentional Gesture: Deliver Your Intentions with Gestures for Speech. arXiv preprint arXiv:2505.15197, 2025. 3 [29] Pinxin Liu, Pengfei Zhang, Hyeongwoo Kim, Pablo Gar- rido, Ari Shapiro, and Kyle Olszewski. Contextual Ges- ture: Co-Speech Gesture Video Generation through Context- aware Gesture Representation. In ACM International Con- ference on Multimedia, 2025. 3 [30] Steven R Livingstone and Frank A Russo.The ryer- son audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english.PloS one, 13(5): e0196391, 2018. 2 [31] Shilin Lu, Xinghong Hu, Chengyou Wang, Lu Chen, Shulu Han, and Yuejia Han. Copy-move image forgery detection based on evolving circular domains coverage. Multimedia Tools and Applications, 81(26):37847–37872, 2022. 2 [32] Shunlin Lu, Ling-Hao Chen, Ailing Zeng, Jing Lin, Ruimao Zhang, Lei Zhang, and Heung-Yeung Shum. Humantomato: Text-aligned whole-body motion generation. arXiv preprint arXiv:2310.12978, 2023. 3 [33] Shilin Lu, Zilan Wang, Leyang Li, Yanzhu Liu, and Adams Wai-Kin Kong.Mace: Mass concept erasure in diffu- sion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6430– 6440, 2024. 1 [34] Shilin Lu, Zihan Zhou, Jiayou Lu, Yuanzhi Zhu, and Adams Wai-Kin Kong. Robust watermarking using generative pri- ors against image editing: From benchmarking to advances. arXiv preprint arXiv:2410.18775, 2024. 2 [35] Shilin Lu, Zhuming Lian, Zihan Zhou, Shaocong Zhang, Chen Zhao, and Adams Wai-Kin Kong. Does flux already know how to perform physically plausible image composi- tion? arXiv preprint arXiv:2509.21278, 2025. 1 [36] Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9 (Nov):2579–2605, 2008. 7 [37] Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Khan. Video-chatgpt: Towards detailed video un- derstanding via large vision and language models. In Pro- ceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12585–12602, 2024. 2 [38] Arsha Nagrani, Joon Son Chung, Weidi Xie, and Andrew Zisserman. Voxceleb: Large-scale speaker verification in the wild. Computer Speech & Language, 60:101027, 2020. 2 [39] OpenAI.Video generation models as world simulators, 2024. 3 [40] Mathis Petrovich, Michael J Black, and G ̈ ul Varol. TMR: Text-to-motion retrieval using contrastive 3D human motion synthesis. In IEEE/CVF International Conference on Com- puter Vision, pages 9488–9497, 2023. 3, 8 [41] Matthias Plappert, Christian Mandery, and Tamim Asfour. The kit motion-language dataset. Big data, 4(4):236–252, 2016. 3 [42] Abhinanda R Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez, and Michael J Black. Babel: Bodies, action and behavior with english la- bels. In Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition, pages 722–731, 2021. 3 [43] Jose Ribeiro-Gomes, Tianhui Cai, Zolt ́ an ́ A Milacski, Chen Wu, Aayush Prakash, Shingo Takagi, Amaury Aubel, Daeil Kim, Alexandre Bernardino, and Fernando De La Torre. Mo- tionGPT: Human Motion Synthesis with Improved Diver- sity and Realism via GPT-3 Prompting. In IEEE/CVF Win- ter Conference on Applications of Computer Vision, pages 5070–5080, 2024. 3 [44] Yunhang Shen, Chaoyou Fu, Peixian Chen, Mengdan Zhang, Ke Li, Xing Sun, Yunsheng Wu, Shaohui Lin, and Rongrong Ji. Aligning and prompting everything all at once for univer- sal visual perception. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 13193–13203, 2024. 1 [45] Fangxun Shu, Lei Zhang, Hao Jiang, and Cihang Xie. Audio- visual llm for video understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4246–4255, 2025. 1 [46] Uriel Singer, Amit Zohar, Yuval Kirstain, Shelly Sheynin, Adam Polyak, Devi Parikh, and Yaniv Taigman. Video edit- ing via factorized diffusion distillation. In European Con- 10 ference on Computer Vision, pages 450–466. Springer, 2024. 3 [47] Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, Yuxuan Wang, and Chao Zhang. video-salmonn: Speech-enhanced audio-visual large language models. arXiv preprint arXiv:2406.15704, 2024. 1 [48] Licai Sun, Zheng Lian, Bin Liu, and Jianhua Tao. Hic- mae: Hierarchical contrastive masked autoencoder for self- supervised audio-visual emotion recognition. Information Fusion, 108:102382, 2024. 2 [49] Rui Sun, Yumin Zhang, Tejal Shah, Jiahao Sun, Shuoying Zhang, Wenqi Li, Haoran Duan, Bo Wei, and Rajiv Ranjan. From sora what we can see: A survey of text-to-video gener- ation. arXiv preprint arXiv:2405.10674, 2024. 3 [50] Zhiyao Sun, Tian Lv, Sheng Ye, Matthieu Lin, Jenny Sheng, Yu-Hui Wen, Minjing Yu, and Yong-jin Liu. Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models.ACM Transactions on Graphics (TOG), 43(4):1–9, 2024. 5 [51] Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean- Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023. 2, 3, 5 [52] Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. Motionclip: Exposing human motion generation to clip space. In European Conference on Com- puter Vision, pages 358–374. Springer, 2022. 3 [53] Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information pro- cessing systems, 30, 2017. 2 [54] Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Repre- sentation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2019. 3 [55] Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video gen- erative models. arXiv preprint arXiv:2503.20314, 2025. 2, 3 [56] Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. Mead: A large-scale audio-visual dataset for emotional talking-face generation. In European conference on com- puter vision, pages 700–717. Springer, 2020. 2, 5, 6, 7 [57] Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024. 2 [58] Thadd ̈ aus Wiedemer, Yuxuan Li, Paul Vicol, Shixiang Shane Gu, Nick Matarese, Kevin Swersky, Been Kim, Priyank Jaini, and Robert Geirhos. Video models are zero-shot learn- ers and reasoners. arXiv preprint arXiv:2509.20328, 2025. 2, 3 [59] Chao-Yuan Wu and Philipp Krahenbuhl. Towards long-form video understanding. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 1884–1894, 2021. 2 [60] Qize Yang, Detao Bai, Yi-Xing Peng, and Xihan Wei. Omni- emotion: Extending video mllm with detailed face and audio modeling for multimodal emotion analysis. arXiv preprint arXiv:2501.09502, 2025. 2 [61] Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen.A survey on multimodal large language models. National Science Review, 11(12): nwae403, 2024. 1 [62] Xinlei Yu, Zhangquan Chen, Yudong Zhang, Shilin Lu, Ruolin Shen, Jiangning Zhang, Xiaobin Hu, Yanwei Fu, and Shuicheng Yan. Visual document understanding and ques- tion answering: A multi-agent collaboration framework with test-time scaling. arXiv preprint arXiv:2508.03404, 2025. 1 [63] Shenghai Yuan, Jinfa Huang, Xianyi He, Yunyang Ge, Yu- jun Shi, Liuhan Chen, Jiebo Luo, and Li Yuan. Identity- preserving text-to-video generation by frequency decompo- sition. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 12978– 12988, 2025. 3 [64] Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M Ni, and Heung-Yeung Shum. Dino: Detr with improved denoising anchor boxes for end-to-end object detection. arXiv preprint arXiv:2203.03605, 2022. 1 [65] Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Shaoli Huang, Yong Zhang, Hongwei Zhao, Hongtao Lu, and Xi Shen. T2M-GPT: Generating Human Motion from Textual Descriptions with Discrete Representations. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 3, 5 [66] Pengfei Zhang, Pinxin Liu, Pablo Garrido, Hyeongwoo Kim, and Bindita Chaudhuri. KinMo: Kinematic-aware Human Motion Understanding and Generation. In IEEE/CVF Inter- national Conference on Computer Vision, 2025. 3 [67] Xiaoman Zhang, Chaoyi Wu, Ziheng Zhao, Weixiong Lin, Ya Zhang, Yanfeng Wang, and Weidi Xie. Pmc-vqa: Vi- sual instruction tuning for medical visual question answer- ing, 2024. URL https://arxiv. org/abs/2305.10415. 1 [68] Chen Zhao, Jiawei Chen, Hongyu Li, Zhuoliang Kang, Shilin Lu, Xiaoming Wei, Kai Zhang, Jian Yang, and Ying Tai.Luve: Latent-cascaded ultra-high-resolution video generation with dual frequency experts.arXiv preprint arXiv:2602.11564, 2026. 3 [69] Jiaxing Zhao, Qize Yang, Yixing Peng, Detao Bai, Shimin Yao, Boyuan Sun, Xiang Chen, Shenghao Fu, Xihan Wei, Liefeng Bo, et al. Humanomni: A large vision-speech lan- guage model for human-centric video understanding. arXiv preprint arXiv:2501.15111, 2025. 2, 5 [70] Duo Zheng, Shijia Huang, and Liwei Wang. Video-3d llm: Learning position-aware video representation for 3d scene understanding. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 8995–9006, 2025. 1 [71] Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. Open-sora: Democratizing efficient video production for all. arXiv preprint arXiv:2412.20404, 2024. 3 [72] Zihan Zhou, Shilin Lu, Shuli Leng, Shaocong Zhang, Zhum- ing Lian, Xinlei Yu, and Adams Wai-Kin Kong. Dragflow: 11 Unleashing dit priors with region based supervision for drag editing. arXiv preprint arXiv:2510.02253, 2025. 3 [73] Yuanzhi Zhu, Ruiqing Wang, Shilin Lu, Junnan Li, Han- shu Yan, and Kai Zhang.Oftsr: One-step flow for im- age super-resolution with tunable fidelity-realism trade-offs. arXiv preprint arXiv:2412.09465, 2024. 3 12