Paper deep dive
MoRight: Motion Control Done Right
Shaowei Liu, Xuanchi Ren, Tianchang Shen, Huan Ling, Saurabh Gupta, Shenlong Wang, Sanja Fidler, Jun Gao
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 93%
Last extracted: 4/10/2026, 4:03:08 AM
Summary
MoRight is a unified video generation framework that enables disentangled control of camera and object motion while modeling motion causality. By using a dual-stream architecture with temporal cross-view attention, it separates canonical object motion from target camera viewpoints. It further decomposes motion into active (user-driven) and passive (consequence) components, allowing for forward reasoning of scene dynamics and inverse reasoning of plausible actions.
Entities (5)
Relation Signals (3)
MoRight → uses → DiT
confidence 95% · We build upon a DiT-based latent video diffusion models.
MoRight → integrates → ViPE
confidence 90% · we estimate per-frame depth maps, camera poses, and intrinsics using ViPE
MoRight → integrates → SAM2
confidence 90% · segment the first frame with SAM2
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Generating motion-controlled videos--where user-specified actions drive physically plausible scene dynamics under freely chosen viewpoints--demands two capabilities: (1) disentangled motion control, allowing users to separately control the object motion and adjust camera viewpoint; and (2) motion causality, ensuring that user-driven actions trigger coherent reactions from other objects rather than merely displacing pixels. Existing methods fall short on both fronts: they entangle camera and object motion into a single tracking signal and treat motion as kinematic displacement without modeling causal relationships between object motion. We introduce MoRight, a unified framework that addresses both limitations through disentangled motion modeling. Object motion is specified in a canonical static-view and transferred to an arbitrary target camera viewpoint via temporal cross-view attention, enabling disentangled camera and object control. We further decompose motion into active (user-driven) and passive (consequence) components, training the model to learn motion causality from data. At inference, users can either supply active motion and MoRight predicts consequences (forward reasoning), or specify desired passive outcomes and MoRight recovers plausible driving actions (inverse reasoning), all while freely adjusting the camera viewpoint. Experiments on three benchmarks demonstrate state-of-the-art performance in generation quality, motion controllability, and interaction awareness.
Tags
Links
- Source: https://arxiv.org/abs/2604.07348v1
- Canonical: https://arxiv.org/abs/2604.07348v1
Trouble viewing inline? Open PDF directly →
Full Text
73,965 characters extracted from source content.
Expand or collapse full text
2026-4-9 MoRight: Motion Control Done Right Shaowei Liu 1,2* Xuanchi Ren 1 Tianchang Shen 1 Huan Ling 1 Saurabh Gupta 2 Shenlong Wang 2 Sanja Fidler 1 Jun Gao 1 1 NVIDIA 2 University of Illinois Urbana-Champaign *Work done during an internship at NVIDIA. https://research.nvidia.com/labs/sil/projects/moright forward reasoning inverse reasoning Motion Causality Reasoning zoom - in orbit - down Zoom - out Disentangled Camera-object Motion Control FramesFrames Figure 1: Given a single input image, our method enables controllable interactive motion generation with motion causality reasoning. Left: Users can provide active motion (e.g. action of hand) to drive scene dynamics (forward reasoning) or specify desired passive outcomes (e.g. trajectory of teapot) and recover plausible driving actions (inverse reasoning). Right: The model further enables disentangled control of object motion and camera viewpoint, allowing users to explore the scene with custom viewpoints and motions. Abstract Generating motion-controlled videos—where user-specified actions drive physically plausible scene dynamics under freely chosen viewpoints—demands two capabilities: (1) disentangled motion control, allowing users to separately control the object motion and adjust camera viewpoint; and (2) motion causality, ensuring that user-driven actions trigger coherent reactions from other objects rather than merely displacing pixels. Existing methods fall short on both fronts: they entangle camera and object motion into a single tracking signal and treat motion as kinematic displacement without modeling causal relationships between object motion. We introduce MoRight, a unified framework that addresses both limitations through disentangled motion modeling. Object motion is specified in a canonical static-view and transferred to an arbitrary target camera viewpoint via temporal cross-view attention, enabling disentangled camera and object control. We further decompose motion into active (user-driven) and passive (consequence) components, training the model to learn motion causality from data. At inference, users can either supply active motion and MoRight predicts consequences (forward reasoning), or specify desired passive outcomes and MoRight recovers plausible driving actions (inverse reasoning), all while freely adjusting the camera viewpoint. Experiments on three benchmarks demonstrate state-of-the-art performance in generation quality, motion controllability, and interaction awareness. Keywords: Video Generation; Disentangled Motion Control; Causal Motion Reasoning © 2026 NVIDIA. All rights reserved. arXiv:2604.07348v1 [cs.CV] 8 Apr 2026 MoRight: Motion Control Done Right 1. Introduction Humans interact with the physical world as active agents: we move our viewpoint, manipulate objects, and reason about how actions lead to consequences. Yet existing video generation models lack this unified capability [7,12,34,67,83]. Bridging this gap is essential for applications that demand interactive visual reasoning, from embodied AI agents [6,72,80] that must anticipate action outcomes [23,27,29], to world models [8,32,35,82,85] that simulate physical interactions, and to immersive content creation [11,22,44, 49,50,77,89] where users freely navigate and manipulate scenes. A desirable video generation system should therefore offer joint controllability over both camera and object motion—generating visually coherent frames under arbitrary viewpoint changes while producing causally consistent scene dynamics driven by user-specified actions. Existing controllable video generation methods [10,15,24,68,70,75] work as renderers: given displace- ments for all pixels, they generate a visually realistic video that adheres to the displacements. These approaches have two key limitations in practice. First, they entangle camera and object motion, making joint control ambiguous because viewpoint changes alter pixel trajectories. Second, they interpret user-specified motion as simple kinematic displacement and largely ignore the consequences of the given trajectories. Models, therefore, focus on following trajectories rather than reasoning about causal relationships between objects. In reality, actions cause consequences—pushing a cup may cause it to slide and collide with other objects, while lifting a teapot may cause water to pour. Specifying the effects of all the objects’ motion through motion prompts is often impractical. Without modeling these action–consequence relationships, motion-controlled generation cannot capture the causal structure of real-world interactions. To overcome these challenges, we introduce MoRight, a unified framework for video generation with disentangled camera–object motion control and motion causality reasoning as shown in Fig. 1. Given a reference image, user-specified motion trajectories, and target viewpoints, MoRight generates videos where objects follow the desired motion and the scene is rendered from the specified cameras. For disentangled camera-object motion control, our key insight is that specifying the motion of objects under camera changes is inherently difficult. We therefore introduce a dual-stream motion formulation. The first branch models and generates object motion in the source image plane under a canonical static viewpoint, allowing users to easily specify dynamic trajectories. The second branch represents the target camera motion and transfers object dynamics from the canonical branch via temporal cross-view attention. This cross-view motion transfer enables independent control of camera and object motion while maintaining coherent scene dynamics. For motion causality reasoning, we achieve this by decomposing object motion into two categories during training: active motion, representing user-driven actions, and passive motion, representing their causal outcomes. By conditioning on either active or passive motion signals, the video model learns to generate all the scene dynamics, capturing action–consequence relationships within the scene. This yields two complementary capabilities: forward reasoning of scene evolution from user actions, and inverse reasoning of plausible actions that produce a desired outcome. We evaluate MoRight on three benchmarks covering diverse interaction scenarios. Results show that MoRight outperforms existing methods in generation quality, motion controllability, and interaction awareness, validating the effectiveness of disentangled motion control and causal motion reasoning. In summary, our contributions are threefold. (1) We propose a disentangled framework for camera and object motion control, enabling users to draw motion trajectories directly in the image plane while freely adjusting viewpoints to generate coherent videos. (2) We empower video generation models with motion reasoning capability by modeling action–consequence relationships, allowing user-driven motions to meaningfully interact with the environment and produce consistent scene dynamics. (3) We demonstrate that MoRight supports both forward and inverse reasoning: given active motion inputs, it predicts future scene evolution; given desired passive outcomes, it recovers plausible actions that achieve them. Together, these advances establish a unified 2 MoRight: Motion Control Done Right framework for controllable and reasoning-aware video generation. 2. Related Work 2.1. Motion-Controlled Video Generation Controllable video generation has been studied across a spectrum of motion granularity, from coarse region- level signals such as bounding boxes [54,70,79], sparse keypoint tracks [41,75,77,86], optical flow fields [38,42,62,63,91], and dense per-pixel trajectories [15,24]. Despite steady progress, trajectory-based methods share two fundamental limitations. First, because trajectories are defined in pixel space, they inevitably entangle object and camera motion: any viewpoint change alters all trajectories, making joint control ill-posed without explicit foreground–background decomposition. Second, producing physically plausible motion typically demands carefully crafted trajectories from dedicated motion-generation models [11,49,50,53,55,62] or laborious manual annotation. MoRight overcomes both issues by disentangling camera and object motion by design, and by accepting lightweight inputs—simple strokes or sparse tracklets—that the model completes into coherent, interaction-aware dynamics. 2.2. Camera–Object Motion Disentanglement Separating camera motion from scene dynamics [36,40,43,84,90] is a long-standing challenge in controllable generation [24,65,73,74] and video understanding [37,51,78]. Existing methods [14,26,62,91] attempt to decouple the two by treating them as independent control signals, yet they typically rely on privileged information such as per-frame depth [14,26], 3D object trajectories [17,30,39], or foreground–background segmentation masks [46,56,62,91], and pre-warp all signals to their anticipated future locations. These assumptions implicitly require the full video sequence or 3D motion to be known in advance, severely limiting applicability in image-to-video settings where only a single reference frame is available. MoRight instead introduces a canonical static-view branch for object dynamics and transfers them to arbitrary target viewpoints via cross-view attention, eliminating the need for explicit 3D supervision or pre-computed scene decomposition. 2.3. Interaction and Causal Reasoning in Video Generation A complementary line of work seeks to move beyond kinematic animation toward causally grounded video generation. One family of methods incorporates external physics engines [11,44,50,55] or conditions on explicit action representations such as force vectors [25] to model specific phenomena (e.g., fluid flow, rigid-body collisions). While effective in constrained domains, these approaches are tailored to particular interaction types and require a simulation module in the loop, limiting their generality. Another family delegates causal reasoning to vision–language or large language models [45,57,76,81], which first predict outcomes in text and then transfer them to the video generator through intermediate representations such as flow fields [53,55], edge maps [53], or depth [53,87]. This two-stage pipeline is prone to error propagation: imprecise textual predictions are further degraded during cross-modal conversion, yielding spatially inaccurate dynamics. MoRight sidesteps both limitations by learning latent action–response structure directly from video data. Decomposing motion into active (user-driven) and passive (consequence) components allows the model to perform cause–effect reasoning at the pixel level within a single forward pass, enabling both forward prediction of scene consequences and inverse inference of plausible underlying actions. 3. Approach Given a single image퐼, we aim to generate a video of푇framesx∈ R 푇×퐻×푊×3 that follows the user-defined motion of an object represented as pixel trajectories풯, and the specified camera motion sequence퐶 푖 푇 푖=1 . 3 MoRight: Motion Control Done Right Consistent Dynamics Camera Tr a j Input Image Motion Tr a j <latexit sha1_base64="bWD2JGyhBJFZ4+he1/mEF+qp61U=">AAAB8nicbVDLSgMxFL1TX7W+qi7dBIvgqsxIUZdFEVxWsA+YDiWTZtrQTDIkGaEM/Qw3LhRx69e482/MtLPQ1gOBwzn3knNPmHCmjet+O6W19Y3NrfJ2ZWd3b/+genjU0TJVhLaJ5FL1QqwpZ4K2DTOc9hJFcRxy2g0nt7nffaJKMykezTShQYxHgkWMYGMlvx9jMyaYZ3ezQbXm1t050CrxClKDAq1B9as/lCSNqTCEY619z01MkGFlGOF0VumnmiaYTPCI+pYKHFMdZPPIM3RmlSGKpLJPGDRXf29kONZ6God2Mo+ol71c/M/zUxNdBxkTSWqoIIuPopQjI1F+PxoyRYnhU0swUcxmRWSMFSbGtlSxJXjLJ6+SzkXdu6w3Hhq15k1RRxlO4BTOwYMraMI9tKANBCQ8wyu8OcZ5cd6dj8VoySl2juEPnM8feVCRZA==</latexit> E <latexit sha1_base64="bWD2JGyhBJFZ4+he1/mEF+qp61U=">AAAB8nicbVDLSgMxFL1TX7W+qi7dBIvgqsxIUZdFEVxWsA+YDiWTZtrQTDIkGaEM/Qw3LhRx69e482/MtLPQ1gOBwzn3knNPmHCmjet+O6W19Y3NrfJ2ZWd3b/+genjU0TJVhLaJ5FL1QqwpZ4K2DTOc9hJFcRxy2g0nt7nffaJKMykezTShQYxHgkWMYGMlvx9jMyaYZ3ezQbXm1t050CrxClKDAq1B9as/lCSNqTCEY619z01MkGFlGOF0VumnmiaYTPCI+pYKHFMdZPPIM3RmlSGKpLJPGDRXf29kONZ6God2Mo+ol71c/M/zUxNdBxkTSWqoIIuPopQjI1F+PxoyRYnhU0swUcxmRWSMFSbGtlSxJXjLJ6+SzkXdu6w3Hhq15k1RRxlO4BTOwYMraMI9tKANBCQ8wyu8OcZ5cd6dj8VoySl2juEPnM8feVCRZA==</latexit> E <latexit sha1_base64="bWD2JGyhBJFZ4+he1/mEF+qp61U=">AAAB8nicbVDLSgMxFL1TX7W+qi7dBIvgqsxIUZdFEVxWsA+YDiWTZtrQTDIkGaEM/Qw3LhRx69e482/MtLPQ1gOBwzn3knNPmHCmjet+O6W19Y3NrfJ2ZWd3b/+genjU0TJVhLaJ5FL1QqwpZ4K2DTOc9hJFcRxy2g0nt7nffaJKMykezTShQYxHgkWMYGMlvx9jMyaYZ3ezQbXm1t050CrxClKDAq1B9as/lCSNqTCEY619z01MkGFlGOF0VumnmiaYTPCI+pYKHFMdZPPIM3RmlSGKpLJPGDRXf29kONZ6God2Mo+ol71c/M/zUxNdBxkTSWqoIIuPopQjI1F+PxoyRYnhU0swUcxmRWSMFSbGtlSxJXjLJ6+SzkXdu6w3Hhq15k1RRxlO4BTOwYMraMI9tKANBCQ8wyu8OcZ5cd6dj8VoySl2juEPnM8feVCRZA==</latexit> E ! ! Motion-Condition Camera-Condition Source View Ta r g e t View Active Motion Realistic Motion Causality Cross-View Self-Attention Weight Shared Consistent Motion under Ta r g e t V i e w Condition EncodingDual-stream generation2-view outputs ❄ ❄ ! Figure 2: Model architecture. Our model adopts a dual-stream architecture with shared weights to disentangle object motion from camera motion. The canonical stream encodes motion trajectories using a track encoder and learns motion in a fixed canonical view. The target stream encodes camera pose signals through a camera encoder. The resulting motion and camera conditions are injected into every attention block of the network. Cross-view self-attention connects the two streams, transferring motion learned in the canonical view to the target view and enabling disentangled camera–object motion generation. The generated video should be able to model the causality of motion and produce a coherent dynamics within the scene. To achieve this, we first present our approach for disentangled camera and object motion control in Sec. 3.1 and Fig. 2, and then introduce motion causality modeling in Sec. 3.2. Sec. 3.3 describes our training data curation pipeline, and training details along with the inference pipeline are provided in Sec. 3.4. 3.1. Disentangled Camera-Object Motion Control Most existing approaches adopt pixel-wise trajectories as motion control signals. However, such representations inherently entangle object motion with viewpoint changes. The video model must implicitly reason how an object moves, and the camera changes, without explicit geometric cues, when conditioned on these signals. Our key insight is that object motion is intrinsically unambiguous when expressed in a canonical camera. Building on this observation, we decouple the motion signal from the viewpoint transformation through a dual-stream generation framework [1,2]. The first stream synthesizes a canonical video in a static camera, where object motion can be directly and faithfully controlled. The second stream generates the target video with both camera and object motion. The two streams can interact with each other through self-attention as shown in Fig. 2. By jointly denoising both streams, the model learns to transfer motion cues from canonical space to arbitrary camera poses, where the canonical stream serves as an anchor for motion control. This design resolves motion–camera entanglement at generation time while naturally supporting heterogeneous supervision—including motion-only, camera-only, and fully coupled data—as detailed in Sec. 3.3. Preliminaries. We build upon a DiT-based [58] latent video diffusion models [16,67]. A pretrained VAE encoder first encodes the video into a latent spacez 0 =ℰ(x)∈ R 푇× 퐻× 푊×푑 . The diffusion model is trained in this latent space via flow matching [48]. Specifically, we first sample a noise from a Gaussian distribution: 휖∼풩(0,I), and form z 푡 = (1−푡)z 0 + 푡휖 for 푡∈ [0, 1]. The DiT풢 휃 is trained to regress the velocity: ℒ = E z 0 ,푡,휖 [︁ ‖풢 휃 (z 푡 ,푡,c)− (휖− z 0 )‖ 2 ]︁ ,(1) wherecdenotes conditioning signals (e.g., text). At inference, an ODE solver [92] integrates the learned 4 MoRight: Motion Control Done Right velocity from a noise to a clean latent ˆ z 0 , from which a decoder reconstructs the original video ˆ x =풟( ˆ z 0 ). Dual-stream generation. The user first provides the object motion휏 can 푖 푇 푖=1 in the canonical frame. The first stream generates the videos only with object motion, and the second stream generates the videos with both object and camera motion. Concretely, the two streams receive their individual conditions: c can = ︀ 퐼, 퐶 1 , 휏 can 푖 ︀ , c tar = ︀ 퐼, 퐶 푖 , ∅ ︀ , where퐶 1 denotes the camera in the first image and is an identity matrix, and∅denotes empty object motion. In the following, we assume we have the paired training data: one ground truth videox can with object motion only, and the corresponding videox tar with both object and camera motion. Extensions to other training data are described in Sec. 3.3 and Sec. 3.4. We add independent noise with the same timestep푡to each stream and obtainz can 푡 andz tar 푡 . We then concatenate two latents along the temporal dimension and jointly denoise them. We slightly modify the positional embedding to indicate the difference between the two streams, with details in the supplement. In this way, we reuse the same DiT weights for two streams, and the only difference is their input and conditioning. The two streams naturally exchange information in the self-attention layers of each transformer block. During inference, the two streams are jointly denoised, and we provide the output from the target stream x =풟( ˆ z tar 0 ) to the user, and the canonical stream serves as a “virtual” anchor. Condition injection and motion transfer. We inject camera and motion conditions into the latents of DiT at every transformer block. Specifically: Camera encoding. We follow Gen3C [88] and warp the first image퐼using the corresponding camera pose and estimated depth [47]. We then encode the warped frames via encoderℰfrom VAE and obtain the latent z cam ∈ R 푇× 퐻× 푊×푑 . For the canonical stream, we use the identity matrix for warping. Motion encoding. Following [24], we build a per-pixel trajectory map where pixels along one trajectory share the same temporal-correspondence embedding. We then encode it via a lightweight encoder to obtain e trk ∈ R 푇× 퐻× 푊×푑 . For the target stream, we simply set e trk = 0 since the condition is empty. Condition injection. The camera and motion encodings are fused via learned linear projections and added into the latent feature at each transformer block. Let f denote the feature of one block: f 푖 ← f 푖 + 푊 cam z 푖,cam + 푊 trk e 푖,trk , 푖∈can, tar.(2) The features from the two streams are then concatenated and passed through the self-attention layer: [︀ f can ; f tar ]︀ := SelfAttn (︀[︀ f can ; f tar ]︀)︀ ,(3) allowing target-view tokens to attend to motion-conditioned canonical tokens and vice versa, implicitly exchang- ing the motion information in latent space. This inject-then-synchronize repeats at every block, progressively transferring motion across views, as shown in Fig. 2. 3.2. Motion Causality Modeling Figure 3: Active vs. passive motion. The ac- tive object (hand) initiates the action, while the passive object (cloth) responds. Disentangling the camera from motion alone is insufficient for realistic interactions: when a hand pushes a cup, the cup must slide; when a ball strikes a stack of blocks, the blocks must scatter. We term this as motion causality: the ability to reason plausible consequences from the given actions, and vice versa. To model this, we decompose the motion tracks of all fore- ground objects풯 =휏 푖 푇 푖=1 into two complementary compo- nents as shown in Fig. 3:휏 푖 =휏 act 푖 ∪휏 pas 푖 ,where휏 act 푖 captures 5 MoRight: Motion Control Done Right PoseEstDepthEstTracking Reprojection Active-passive segmentation Filtering ... ... Stage-1 Motion Extraction & Canonicalization Stage-2 Motion Decomposition Stage-3 Paired Multi-view synthesis Figure 4: Data curation pipeline. Foundation models [30,36,59] extract depth, camera poses, and tracks from raw videos. A VLM [3] segments tracks into active/passive regions. We further optionally use a video-to-video model [21] to generate paired videos with the same object motion but different camera motions. the active (causal) motion—the intentional action applied to the scene (e.g., a hand pushing)—and휏 pas 푖 captures the passive (consequential) motion—the reaction from other objects (e.g., the pushed object sliding). Reasoning such causality is critical in applications such as embodied AI [9, 19, 28]. The central mechanism to enable the causality modeling is through motion dropout. Specifically, during training, we randomly drop out one motion component from the input, and supervise the model on the full video containing both active and passive motion: ̃ 휏 푖 := ⎧ ⎨ ⎩ 휏 act 푖 , 휉 < 푝, 휏 pas 푖 , otherwise, (4) where휉is sampled from a uniform distribution풰(0, 1),푝is the dropout probability and ̃ 휏 푖 is used as tracking condition to video model. In our model training, we do not distinguish active motion or passive motion when feeding them into the model, and only rely on the model’s capability to reason the dropped component to generate plausible videos. During training, this asymmetric supervision encourages the video models to internalize the causal relationship between actions and their consequences, rather than simply replaying the provided trajectories. At inference, this learned causality enables two complementary applications: forward reasoning (action →reaction), where users specify an action and the model generates the resulting consequences; and inverse reasoning (reaction→action), where users prescribe a desired outcome and the model synthesizes a plausible action that drives it. We demonstrate these capabilities in Sec. 4.4. 3.3. Training Data Curation Curating data to train our dual-stream model is challenging since most real-world videos are single-view and always entangle camera and object motion, while our model requires paired videos depicting the same dynamics under different viewpoints. We first describe our data annotation pipeline to extract the pixel trajectories, camera poses, and active/passive motion, and provide details of our data curation for training afterwards. The overview of pipeline is shown in Fig. 4. Motion extraction and canonicalization. Given a videox, we estimate per-frame depth maps퐷 푖 , cam- era poses퐶 푖 , and intrinsics퐾using ViPE [36], and extract dense pixel trajectories풯 = 휏 푖 푇 푖=1 with 6 MoRight: Motion Control Done Right AllTracker [30]. Each trajectory is unprojected to 3D and reprojected into the first frame: 휏 can 푖 = 휋 (︀ 퐾, 퐶 0 퐶 −1 푢 휋 −1 (퐾,휏 푖 , 퐷 푖 ) )︀ ,(5) where휋 −1 lifts 2D points to 3D using depth and intrinsics, and휋projects onto the image plane of퐼 0 . We assume constant intrinsics across the video. Active and passive motion decomposition. For the given video, we prompt a vision-language model (Qwen3 [3]) to identify the active and passive objects, then segment the first frame with SAM2 [59], yielding masks푀 act and푀 pas , for active and passive objects, respectively. Trajectories are assigned to each component by mask membership, producing휏 can,act and휏 can,pas for the motion dropout training in Eq. (4). We also generate per-video captions describing each motion component and only provide the caption of one component during training to prevent information leakage. Synthetic data generation for paired two-view videos. We leverage a synthetic data generation pipeline to generate paired two-view videos for training. Specifically, we first curate the videos whose cameras are static by checking the displacement of camera poses estimated from ViPE [36]. We then synthesize corresponding moving-camera videos using a camera-control video-to-video model [20], providing supervision for the second stream. To increase camera diversity, we further augment the data with basic camera operations (e.g., orbit, pan, zoom) as well as dynamic camera trajectories extracted from real videos. Single-view real-world data for mixed-training. The generated paired videos inevitably contain visual artifacts. In our dual-stream mode, the abundant single-view real-world data can be leveraged for mixed- training. First, for the videos with static cameras (object motion only), we duplicate the video and treat it as the target video. In this case, the video model learns to exchange the motion condition from the first stream (with motion condition) into the target stream (without motion condition). Second, for the videos that exhibit both camera and object motion, we feed the condition into the video model as described in Sec. 3.1 and only supervise the second stream, leaving the loss at the first stream being zero. These two mixed-training strategies expose video models to real-world data with diverse camera and object motion, increasing the robustness and generalizability to various camera and motion configurations, while mitigating artifacts from synthetic data. Rendered Graphics Data. We further incorporate synthetic data from SyncCamMaster [2] to expose our model to more camera diversity. 3.4. Training and Inference We train the DiT using the flow matching loss from Eq. (1), and apply two complementary dropout strategies to encourage the model learn the motion causality. Multi-granularity motion dropout. Per Eq. (4), we randomly retain either휏 act or휏 pas to encourage causal reasoning. We further obtain multi-granularity trajectories by averaging per-pixel trajectories within each patch. During training, we randomly select the granularity, enabling the model to capture both fine-grained pixel control and object-level manipulation. Occlusion and track dropout. For the obtained trajectory, we randomly mask a subset of it to simulate occlusion and tracking failures that can happen during inference, improving robustness to missing or unreliable tracks. Inference. At test time, users first specify motion by drawing sparse trajectories (simple curves or strokes on the first image) to indicate the desired direction and magnitude of movement, along with an optional text prompt and target camera poses퐶 푖 푇 푖=1 . We further perform occlusion-aware masking by approximating visibility ordering from the first-frame depth. The model then jointly denoises both streams, with the second stream being presented to the user. 7 MoRight: Motion Control Done Right Table 1: Controllable video generation on DynPose-100K [61] and Cooking. We compare models with camera and object motion control. Tracking-based methods (MP, ATI, WanMove) require privileged fore- ground/background tracks, while MoRight uses only first-frame reprojected trajectories and camera poses. Despite weaker inputs, MoRight achieves comparable visual quality and more accurate motion control. All methods use the Wan2.1-14B backbone; * denotes models reimplemented and trained by us. Best and second- best results are marked in bold and underline. DynPose-100KCooking MethodFull info PSNR↑ SSIM↑Rot↓Trans↓EPE↓ PSNR↑ SSIM↑Rot↓Trans↓EPE↓ Wan2.1 [67] ×11.23 0.435--–14.23 0.527--– Gen3c* [60] ×12.450.5075.464.09-15.37 0.6131.9710.03- MP* [24]✓11.72 0.4556.766.047.5615.68 0.5642.5012.244.25 ATI [68]✓13.18 0.4935.626.548.43 15.93 0.5824.2516.945.87 WanMove [15]✓13.91 0.5214.123.568.05 16.420.5892.9313.275.47 Ours×12.30 0.4574.554.617.64 16.44 0.5942.1610.114.27 4. Experiments 4.1. Implementation Details We build upon the pretrained Wan2.1-14B [67] and fine-tune only the camera encoder, trajectory encoder, and self-attention layers. The trajectory embedding dim is 64, and the camera encoder uses 32 channels. We train the model with 15K iterations on 64 GPUs with a global batch size of 16, using AdamW [52] with a learning rate of3× 10 −5 and weight decay0.001. Trajectory dropout is set to 0.1 and text-conditioning dropout to 0.2. Following the data curation pipeline in Sec. 3.3, we build our training data from large-scale public video datasets, including Panda-70M [13] and Wild-SDG-1M [36]. From these sources, we collect 76K static-view videos, from which we synthesize 43K paired dynamic-view videos using camera-controlled video-to-video generation, and 3.4K synthetic interaction videos from SyncMaster [2]. All videos are processed at 480p resolution. At inference, we sample 35 diffusion steps; generating one video takes approximately 15 minutes on a single A100 GPU. More implementation details are presented in Sec. A. 4.2. Experiment Settings Evaluation metrics. We evaluate our model across four different aspects: Video quality: PSNR and SSIM against reference videos, and FID [33] and FVD [66] for distribution-level similarity. Camera accuracy: rotation and translation errors [1,2,31] between reference poses and poses estimated from generated videos using ViPE [36]; we report median errors across frames to mitigate estimation noise. Motion accuracy: end-point error (EPE) [15,24], theℓ 2 distance between ground-truth object tracks and predicted tracks extracted with AllTracker; we report the median EPE to reduce the impact of outlier tracks. Motion realism: Physical Commonsense (PC) and Semantic Adherence (SA) from VideoPhy [4], both 5-point scores normalized to[0, 1]. All evaluations are conducted at 480p resolution. Evaluation Datasets. We evaluate on three datasets spanning diverse interaction scenarios. DynPose-100K [61] is an in-the-wild dataset with highly dynamic camera motion; we manually select 50 videos exhibiting strong viewpoint changes and clear object interactions. WISA [69] is a large-scale physical-dynamics dataset; we select 50 videos from categories including collision, deformation, elasticity, liquid, and rigid-body motion. We further collect 50 real-world cooking videos, featuring complex hand-object interactions. 8 MoRight: Motion Control Done Right forward reasoning inverse reasoning zoom - in orbit - down orbit - right Figure 5: Disentangled camera–object control. MoRight enables independent control of object motion and camera viewpoint. Rows 1-3 fix the camera and vary object motion (rows 1-2: forward reasoning; row 3: inverse reasoning), while rows 4-6 fix object motion and vary camera motion. 4.3. Disentangled Camera-Object Motion Control Existing works on controllable video generation typically focus on either camera or object motion in isolation. Motion-conditioned baselines [15,24,68] typically receive privileged signals of both foreground and background tracks of all the pixel trajectories for control, while our method only uses reprojected trajectories defined on the canonical frame, without access to future-frame pixel trajectories. We compare our method with state-of-the-art baselines to evaluate the quality in motion control and camera control. While challenging, this evaluation allows us to demonstrate our model’s unique ability to generate faithful motion from disentangled controls—a task that is fundamentally more difficult than baselines. Baselines and setup. We compare with several controllable video generation methods. Wan2.1 [67] is our base model without motion control. Gen3C [60] only supports camera control. Recent state-of-the-art motion-conditioned models (Motion Prompting (MP) [24], ATI [68] and WanMove [15] all take dense pixel tracks as input. For a fair comparison, we retrain Gen3C and MP in our setup, and all methods share the same Wan2.1-14B backbone. Evaluation is conducted on DynPose-100K [61] and Cooking dataset. Motion accuracy is evaluated in future-frame pixel space via EPE for all methods. Results. Quantitative results are provided in Tab. 1, with controllable results in Fig. 5, qualitative comparisons in Fig. 6. On DynPose-100K [61], WanMove [15] achieves the best overall numbers. Our method is slightly behind 9 MoRight: Motion Control Done Right ATI WanMove MoRight ATI WanMove MoRight ATI WanMove MoRight Figure 6: Qualitative comparison of ATI [68], WanMove [15], and MoRight on interactive motion generation with camera control. All methods use the same input image. ATI and WanMove rely on pixel-aligned per-frame tracks (top-left), which entangle camera and object motion and require privileged future tracks. In contrast, MoRight uses only reprojected first-frame tracks. The first two rows show active motion reasoning, and the third row shows passive motion reasoning. Our model disentangles camera and object control and produces more coherent interactions. 10 MoRight: Motion Control Done Right Table 2: Interactive motion generation on WISA [69] and Cooking. We compare motion-conditioned video generation models on WISA and Cooking, evaluating video quality (FID, FVD) and motion realism (PC, SA). Prior methods require detailed motion captions with full interaction descriptions, while MoRight uses only a single active motion description yet achieves comparable quality with stronger physical commonsense reasoning. All methods use the Wan2.1-14B backbone for fair comparison; * denotes models reimplemented and trained by us. Best and second-best results are marked in bold and underline. WISACooking MethodFull info FID↓ FVD↓PC↑SA↑ FID↓ FVD↓PC↑SA↑ MP* [24]✓57.29975.940.750.8243.49759.530.870.89 ATI [68]✓69.80 990.820.750.83 55.80 881.940.850.90 WanMove [15]✓61.34 1088.230.730.83 53.51 882.900.840.87 Ours×52.95 876.030.760.8239.94 730.460.880.89 Table 3: Ablation of motion controllability and reasoning on the Cooking benchmark. We ablate archi- tectural choices, causal reasoning, hybrid training, and different motion input granularities. Our full model achieves the best overall performance across photometric quality, controllability, and motion realism, while remaining robust to different motion granularities and input conditions (active and passive). PhotometricControllabilityMotion SettingFID↓ FVD↓ PSNR↑ SSIM↑Rot↓Trans↓EPE↓PC↑SA↑ cascaded41.74 728.8015.98 0.5692.6911.505.050.870.89 w/o fixed view51.17 997.83 14.15 0.5153.3614.5714.300.870.89 w/o reasoning44.04 784.19 15.55 0.5622.8812.495.050.870.88 w/o mixed training 41.94 808.96 16.29 0.5832.2212.804.090.870.89 ours (coarse)39.83 725.88 16.45 0.5942.2110.984.370.880.88 ours (passive)44.20 838.67 15.99 0.5882.2111.047.270.870.88 ours (active)39.94 730.46 16.440.5942.1610.114.270.880.89 under highly dynamic camera motion, where errors in camera pose estimation and trajectory reprojection can degrade the input control signals. Nevertheless, MoRight attains comparable controllability to methods that rely on privileged future-frame tracking information and achieves the best EPE for object motion accuracy. We observe that ATI [68] and WanMove [15], which couple camera and object motion in a single tracking signal, tend to favor the dominant motion mode in highly dynamic settings—sometimes sacrificing camera accuracy or object tracking fidelity. On the Cooking benchmark, our method achieves the best overall performance in both visual quality and motion control accuracy. More controllable generation results are shown in Fig. 14, and qualitative comparisons are provided in Fig. 15. 4.4. Motion Causality Modeling Setup. We evaluate the causality of modeling motion on WISA [69] and Cooking, measuring both generation quality (FID, FVD) and motion realism (PC, SA). We compare with MP* [24], ATI [68], and WanMove [15], all following the same input protocol as Sec. 4.3. Every model receives active motion representing the user-specified 11 MoRight: Motion Control Done Right active active active passive passive passive Figure 7: Causal interaction reasoning. In the first 3 rows, we provide active motion (e.g., hand movement) as input, and the model infers the resulting passive motion (e.g., cloth movement). In the last 3 rows, we provide passive motion (e.g., ball movement), and the model infers the corresponding active motion (e.g., human movement). action (e.g., a hand pushing an object); the goal is to generate plausible interaction outcomes. Baseline methods use their original prompts containing both motion descriptions and expected consequences. Our model receives only the active motion description, without specifying passive outcomes, and must infer the resulting interactions. Results. Quantitative results are reported in Tab. 2. MoRight achieves the highest PC score on WISA, indicating strong physical commonsense, and the best video quality (FID, FVD) on both datasets. For SA, we rewrite the input prompt to remove passive motion descriptions and avoid information leakage; consequently, our score is slightly lower than methods that use full prompts containing both actions and outcomes, yet remains comparable. This confirms that MoRight’s generations stay semantically aligned with intended outcomes while demonstrating genuine causal motion reasoning rather than relying on prompt-supplied answers. Qualitative examples of both reasoning modes—forward (action→reaction) and inverse (reaction→action)—are shown in Fig. 7. Our model generates plausible reactions when providing the active motion, and can reason the meaningful active motion when providing passive motion. More qualitative visualizations are shown in Fig. 13. 12 MoRight: Motion Control Done Right ATIWanMoveoursNone 0 10 20 30 40 50 60 Preference (%) 18.8% 25.0% 53.5% 2.7% Controllability ATIWanMoveoursNone 18.2% 25.7% 54.6% 1.5% Motion Realism ATIWanMoveoursNone 17.4% 23.1% 55.9% 3.6% Photorealism User Preference by Metric ATIWanMoveoursNone Figure 8: Human perceptual evaluation. From 330 responses by 11 participants, our method is preferred across controllability, motion realism, and photorealism, outperforming ATI [68] and WanMove [15], which rely on privileged 3D tracks but lack interaction reasoning. 4.5. Human Perceptual Evaluation In addition to objective metrics, we conduct a human perceptual study to evaluate generation quality. ATI [68] and WanMove [15] rely on pixel-aligned per-frame tracks projected from privileged 3D trajectories, including both foreground/background and full interaction (active and passive) motion. In contrast, our method uses only first-frame active trajectories, requiring the model to infer interactions without privileged information. We randomly sample 30 examples from the combined test datasets. For each example, videos from different methods are presented in randomized order to avoid positional bias. Participants evaluate results based on three criteria: Controllability (alignment with input object and camera motion), Motion Realism (physical plausibility of interactions), and Photorealism (visual quality). For each criterion, participants select the best result, with ties and a None option allowed. After filtering unreliable submissions, we collect responses from 11 participants, yielding 330 evaluations per criterion (the evaluation interface is shown in Fig. 12). As shown in Fig. 8, our method is preferred in the majority of cases, achieving 53.5%, 54.6%, and 55.9% for controllability, motion realism, and photorealism, respectively. This outperforms ATI [68] (18.8%, 18.2%, 17.4%) and WanMove [15] (25.0%, 25.7%, 23.1%). Despite access to privileged 3D trajectories, baseline methods lack explicit interaction reasoning and entangle camera and object motion, leading to inferior performance. In contrast, our disentangled formulation enables more controllable and realistic video generation. 4.6. Ablation Studies We ablate model design, training strategies, and input conditions on the Cooking dataset in Tab. 3. Model design and training. Cascaded pipeline (row 1): a naive solution for disentangling camera-object motion is to first generate motion-controlled video under a static camera, followed by a Gen3C-style camera controller to move the camera. However, this approach introduces error accumulation between two stages, yielding larger control errors. W/o fixed-view branch (row 2): we only train with dynamic camera views and jointly encode the reprojected tracks and camera embeddings, removing the canonical-view anchor. The model struggles to disentangle camera and object motion, resulting in significantly worse camera and tracking accuracy. W/o motion reasoning (row 3): we disable active/passive decomposition during training. This approach increases FID/FVD and reduces PC, indicating degraded interaction quality. W/o mixed supervision (row 4): We only train the model on paired data. It slightly degrades camera accuracy, as the paired subset contains limited camera motion diversity. 13 MoRight: Motion Control Done Right Ours GT Ours GT Ours GT Ours GT Figure 9: Limitation analysis. Input tracks are overlaid on the first frame as in previous figures. (1) Incorrect interaction reasoning may lead to implausible outcomes (two kabobs merging). (2) Unnatural motion can occur when input tracks become temporally sparse due to occlusion (hand example). (3) Physically unrealistic dynamics may appear, such as objects disappearing during motion (soccer ball). (4) Hallucinated content may emerge in later frames (extra hand). 14 MoRight: Motion Control Done Right Input conditions. We vary the motion input configuration to evaluate the robustness of our models. Specifically, we evaluate with coarse segment-level trajectories vs. fine-grained pixel tracks, as well as active vs. passive motion inputs. Performance remains stable across all settings, confirming that MoRight flexibly handles different motion granularities and types while maintaining strong controllability and causal reasoning capability. 4.7. Limitation Analysis Despite promising results, our method still exhibits several limitations, as illustrated in Fig. 9. First, the model may produce incorrect interaction reasoning, leading to implausible outcomes such as two kabobs merging into a single object. Second, unnatural motion can occur when the input trajectories become temporally sparse due to occlusion, making it difficult for the model to reliably infer the intended motion (e.g., the hand example). Third, the generated motion may violate physical consistency, such as objects disappearing during interaction (soccer example). Fourth, the model may occasionally hallucinate new content in later frames, such as an extra hand appearing during generation. In addition, our method has difficulty modeling very complex or fast camera motion (e.g., drastic egomotion). Our camera control is designed for common smooth camera trajectories, and when the input camera motion changes drastically, the predicted interaction dynamics may degrade. 5. Conclusion We present MoRight, a unified framework for controllable and interaction-aware video generation. MoRight addresses two key limitations of prior motion-controlled methods: (1) entangled camera and object motion, resolved through a dual-stream design that enables independent control of object trajectories and camera viewpoints; and (2) limited causal reasoning, addressed by decomposing motion into active (user-driven) and passive (consequence) components to learn action–response dynamics. At inference, MoRight supports both forward prediction—generating scene outcomes from active motion—and inverse reasoning—inferring actions from desired passive results. Experiments on DynPose-100K, WISA, and Cooking show strong performance in generation quality, motion control, and interaction awareness, establishing MoRight as a step toward more interactive and physically grounded video generation. Acknowledgement We would like to thank Jiahui Huang, Zian Wang, Xiao Fu, and Chen-Hsuan Lin for their help and support with video data processing, data generation, and infrastructure. 15 2026-4-9 Appendix A. Implementation Details A.1. Network Architecture Our model builds on the Wan2.1 I2V-14B [67]. We first encode the two-view videos and then concatenate the tokens along the temporal dimension before feeding them into the model. The two streams share the same spatial RoPE [64] embeddings but use different temporal indices. The object tracking condition, represented as a trajectory map, is encoded by a lightweight temporal encoder with RMSNorm, SiLU [18], and two3× 1× 1 Conv3D layers that downsample the temporal dimension by4×to match the Wan latent resolution. For camera motion control, we follow Gen3C [60] by warping the first frame with the camera trajectory and encoding it with the VAE, producing features in the same latent space and resolution. Both camera and tracking features are linearly projected to the Wan hidden dimension (5120) and added to the video tokens before the self-attention layer of each Wan2.1 transformer block. During training, we only train the lightweight temporal encoder and the self-attention layers together with the camera and tracking encoders in each block, and freeze other parts of the network. A.2. Training Data Curation In training data curation, we need to identify active and passive object and its motion in a given video. We first identify those objects by querying Qwen3-VL [3] and use SAM2 [59] for video object segmentation. The system prompt to Qwen3-VL [3] is shown in Fig. 10. We further use Qwen3 to rewrite video captions by decomposing object motion into active or passive descriptions, ensuring that each rewritten caption contains only one type of motion. During training, the original caption and the rewritten caption are randomly sampled with equal probability, encouraging the model to infer plausible interaction consequences. To generate paired multi-view data, we select videos from our collected Internet videos with nearly static cameras using the camera poses provided by ViPE [36], requiring a maximum rotation of 0.5 ∘ and translation of 5 m. System: You are a video understanding specialist. Follow all rules exactly and output only valid JSON. • active_dynamic: objects that move by their own power or actuation (e.g., person, hand, animal, robot, drone, car, bus, train, boat, ship, airplane, motorcycle). • passive_dynamic : objects that move only because of other objects or forces. If a person/hand is visible manipulating an item, that item is passive_dynamic. User: You are given a video. Identify all objects that actually move: • active_dynamic: objects moving by their own power (or actuation). • passive_dynamic: objects that move only because other objects or natural forces cause them to move. Figure 10: Prompt used for active and passive object identification for Qwen3 [3] in data curation pipeline. A.3. Training During training, we apply several data augmentations to improve robustness. For each sample, we randomly simplify the input trajectories with probability 0.5, where tracks are averaged per object such that all pixels of the same object share a single trajectory. To encourage the model to reason about motion causality, we randomly provide active or passive motion tracks with probabilities 0.8 and 0.2, respectively. We further apply motion dropout by randomly dropping visible tracks with probability 0.2 to simulate occlusion and tracking errors commonly observed in off-the-shelf trackers at inference time. In addition, we randomly truncate tracks after a sampled middle frame to simulate partial observations of motion. During training, we randomly sample © 2026 NVIDIA. All rights reserved. Appendix Figure 11: Interactive demo interface. Our system enables users to control both object and camera motion from a single image. Users draw trajectories on the first frame to specify object motion (active or passive), either by moving a selected region using keypoint trajectories or by defining fine-grained motion paths for detailed control. between 500 and 2000 tracks per iteration. At inference time, we fix the number of input tracks to 1500 for all experiments to ensure consistent evaluation. Finally, since multi-view supervision is critical for learning camera–object disentanglement, we control the sampling ratio of multi-view and single-view training samples, ensuring that single-view data is sampled at a lower rate to prevent the model from overfitting to single-view motion patterns. A.4. Inference At inference time, users can freely select objects in the first frame and specify their motion trajectories. Motion control can be provided either in a coarse manner, where the entire object moves with a shared trajectory, or in a fine-grained manner using sparse point tracks. To facilitate interaction-driven editing, we also provide several simple motion primitives for hand interactions, such as push, pull (along a specified direction), and reach (toward a target location). For passive motion control, users can define arbitrary 2D trajectories; Fig. 13 illustrates an example using straight-line trajectories with different directions for inverse reasoning. To enable flexible control, we implement an interactive GUI as shown in Fig. 11. Starting from a single input image, users draw motion trajectories directly on the first frame while specifying camera motion independently through a sequence of camera poses (the first frame is treated as the identity pose). The interface supports trajectory visualization across time steps and occlusion checking using the first-frame depth estimated by MoGe [71], enabling intuitive editing of object dynamics and camera viewpoints during generation. A.5. Evaluation The Cooking Benchmark is constructed from real-world cooking videos collected from YouTube that contain rich hand–object interactions. These scenes involve diverse manipulation behaviors such as pushing, cutting, and picking, making them a suitable testbed for evaluating interactive motion reasoning. The benchmark contains 50 video clips covering a variety of kitchen environments and object interactions. For motion quality evaluation, we adopt Physical Commonsense (PC) and Semantic Adherence (SA), 17 Appendix Figure 12: Human perceptual evaluation interface. Given an input image, object trajectories, and a target camera motion, participants evaluate generated videos under three criteria: Controllability (matching object tracks and camera motion), Motion Realism (physically plausible interactions and scene responses), and Photorealism (overall visual quality). For each criterion, participants select the best video (multiple selections allowed for ties, or None if none satisfy it). The study contains 30 randomly selected video sets with shuffled candidate order. from [4,5]. PC measures whether the generated video follows real-world physical behaviors. SA evaluates whether the generated video remains semantically consistent with the input text prompt. For SA, we use the original caption of each video as the evaluation prompt. We use the automatic evaluation rater provided to score the generated videos and report the resulting normalized scores across different methods. For human perceptual evaluation, the interface layout is shown in Fig. 12. B. More Qualitative Results B.1. Interactive Motion Generation We demonstrate the causal reasoning ability of our model in Fig. 13. We generate videos and select the first-stream generation (static view) to visualize the interaction dynamics. The input motion tracks are overlaid on the generated frames, where colored tracks indicate either user actions (active motion) or passive trajectories. Our model supports both forward and inverse reasoning. In forward reasoning (first two samples), the model predicts plausible scene consequences given the specified active motion. In inverse reasoning (last sample), the model infers feasible driving actions that could lead to the observed passive outcomes. These examples highlight the model’s ability to reason about causal interactions between actions and objects. 18 Appendix Figure 13: Causal interaction reasoning. Input tracks are shown in color and overlaid on the generated static reference-view video. The tracks represent user actions (active) or passive trajectories. Given these inputs, our model either predicts plausible consequences (forward reasoning) or recovers feasible driving actions that produce the desired outcomes (inverse reasoning, last row). 19 Appendix Orbit-left Zoom-in Zoom-out Orbit-left Zoom-in Zoom-out Orbit-left Zoom-in Zoom-out Figure 14: Additional controllable generation-1. Object motion trajectories are overlaid on the input image. For each video, we show different camera and object motion control. Each group shares the same object motion but uses different camera motions. Minor variations under the same object motion but different camera motions arise from the stochastic nature of interaction generation. 20 Appendix ATI WanMove MoRight ATI WanMove MoRight ATI WanMove MoRight Figure 15: Additional qualitative comparison with ATI [68], WanMove [15], and MoRight. ATI and WanMove rely on privileged 3D trajectories (with depth) projected to pixel-aligned per-frame tracks and take full interaction trajectories (active and passive) as input. In contrast, MoRight uses only first-frame active tracks without privileged information and infers plausible interactions. 21 Appendix B.2. Disentangled Camera-Object Control We present additional disentangled controllable generation results in Fig. 14. Object motion trajectories are overlaid on the input image. We demonstrate three different object motions and three different camera motions (orbit-left, zoom-in, and zoom-out), resulting in 9 generated videos in 3 group. Each group shares the same object motion while varying the camera viewpoint, highlighting the model’s ability to maintain consistent object dynamics under different camera controls. Minor variations under the same object motion but different camera motions arise from the stochastic nature of interaction generation. B.3. Qualitative Comparison We provide additional qualitative comparisons with ATI [68] and WanMove [15] in Fig. 15. Both baselines rely on privileged 3D trajectories (with depth) projected to pixel-aligned per-frame tracks and take full interaction trajectories (active and passive) as input. In contrast, our method only requires 2D motion trajectories on the first frame, while camera motion is introduced in the second stream of our dual-stream architecture. Despite using weaker inputs, our approach achieves stronger controllability and produces more coherent interactions while maintaining disentangled camera–object motion. 22 Appendix References [1]J. Bai, M. Xia, X. Fu, X. Wang, L. Mu, J. Cao, Z. Liu, H. Hu, X. Bai, P. Wan, et al. Recammaster: Camera-controlled generative rendering from a single video. arXiv preprint arXiv:2503.11647, 2025. 4, 8 [2]J. Bai, M. Xia, X. Wang, Z. Yuan, X. Fu, Z. Liu, H. Hu, P. Wan, and D. Zhang. Syncammaster: Synchronizing multi-camera video generation from diverse viewpoints. Proc. ICLR, 2025. 4, 7, 8 [3]S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, W. Ge, Z. Guo, Q. Huang, J. Huang, F. Huang, B. Hui, S. Jiang, Z. Li, M. Li, M. Li, K. Li, Z. Lin, J. Lin, X. Liu, J. Liu, C. Liu, Y. Liu, D. Liu, S. Liu, D. Lu, R. Luo, C. Lv, R. Men, L. Meng, X. Ren, X. Ren, S. Song, Y. Sun, J. Tang, J. Tu, J. Wan, P. Wang, P. Wang, Q. Wang, Y. Wang, T. Xie, Y. Xu, H. Xu, J. Xu, Z. Yang, M. Yang, J. Yang, A. Yang, B. Yu, F. Zhang, H. Zhang, X. Zhang, B. Zheng, H. Zhong, J. Zhou, F. Zhou, J. Zhou, Y. Zhu, and K. Zhu. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631, 2025. 6, 7, 16 [4]H. Bansal, Z. Lin, T. Xie, Z. Zong, M. Yarom, Y. Bitton, C. Jiang, Y. Sun, K.-W. Chang, and A. Grover. Videophy: Evaluating physical commonsense for video generation. arXiv preprint arXiv:2406.03520, 2024. 8, 18 [5]H. Bansal, C. Peng, Y. Bitton, R. Goldenberg, A. Grover, and K.-W. Chang. Videophy-2: A challenging action-centric physical commonsense evaluation in video generation. arXiv preprint arXiv:2503.06800, 2025. 18 [6]A. Bar, G. Zhou, D. Tran, T. Darrell, and Y. LeCun. Navigation world models. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 15791–15801, 2025. 2 [7]A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22563–22575, 2023. 2 [8]T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. Ng, R. Wang, and A. Ramesh. Video generation models as world simulators. OpenAI technical reports, 2024. 2 [9]Q. Bu, J. Cai, L. Chen, X. Cui, Y. Ding, S. Feng, S. Gao, X. He, X. Hu, X. Huang, et al. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems. arXiv preprint arXiv:2503.06669, 2025. 6 [10] R. Burgert, Y. Xu, W. Xian, O. Pilarski, P. Clausen, M. He, L. Ma, Y. Deng, L. Li, M. Mousavi, et al. Go-with-the-flow: Motion-controllable video diffusion models using real-time warped noise. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 13–23, 2025. 2 [11]B. Chen, H. Jiang, S. Liu, S. Gupta, Y. Li, H. Zhao, and S. Wang. Physgen3d: Crafting a miniature interactive world from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6178–6189, 2025. 2, 3 [12]H. Chen, Y. Zhang, X. Cun, M. Xia, X. Wang, C. Weng, and Y. Shan. Videocrafter2: Overcoming data limitations for high-quality video diffusion models. arXiv preprint arXiv:2401.09047, 2024. 2 [13]T.-S. Chen, A. Siarohin, W. Menapace, E. Deyneka, H.-w. Chao, B. E. Jeon, Y. Fang, H.-Y. Lee, J. Ren, M.-H. Yang, et al. Panda-70m: Captioning 70m videos with multiple cross-modality teachers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13320–13331, 2024. 8 [14]Y. Chen, Y. Men, Y. Yao, M. Cui, and L. Bo. Perception-as-control: Fine-grained controllable image animation with 3d-aware motion representation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14380–14389, 2025. 3 [15]R. Chu, Y. He, Z. Chen, S. Zhang, X. Xu, B. Xia, D. Wang, H. Yi, X. Liu, H. Zhao, et al. Wan-move: Motion-controllable video generation via latent trajectory guidance. arXiv preprint arXiv:2512.08765, 2025. 2, 3, 8, 9, 10, 11, 13, 21, 22 [16] T. Cosmos. Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575, 2025. 4 23 Appendix [17]C. Doersch, Y. Yang, M. Vecerik, D. Gokay, A. Gupta, Y. Aytar, J. Carreira, and A. Zisserman. Tapir: Tracking any point with per-frame initialization and temporal refinement. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10061–10072, 2023. 3 [18]S. Elfwing, E. Uchibe, and K. Doya. Sigmoid-weighted linear units for neural network function approximation in reinforcement learning. Neural networks, 107:3–11, 2018. 16 [19]C. Finn and S. Levine. Deep visual foresight for planning robot motion. In 2017 IEEE international conference on robotics and automation (ICRA), pages 2786–2793. IEEE, 2017. 6 [20] X. Fu, S. Tang, M. Shi, X. Liu, J. Gu, M.-Y. Liu, D. Lin, and C.-H. Lin. Plenoptic video generation. arXiv preprint arXiv:2601.05239, 2025. 7 [21]X. Fu, S. Tang, M. Shi, X. Liu, J. Gu, M.-Y. Liu, D. Lin, and C.-H. Lin. Plenoptic video generation. arXiv preprint arXiv:2601.05239, 2026. 6 [22]Q. Gao, Q. Xu, Z. Cao, B. Mildenhall, W. Ma, L. Chen, D. Tang, and U. Neumann. Gaussianflow: Splatting Gaussian dynamics for 4D content creation. arXiv preprint arXiv:2403.12365, 2024. 2 [23]S. Gao, W. Liang, K. Zheng, A. Malik, S. Ye, S. Yu, W.-C. Tseng, Y. Dong, K. Mo, C.-H. Lin, Q. Ma, S. Nah, L. Magne, J. Xiang, Y. Xie, R. Zheng, D. Niu, Y. L. Tan, K. Zentner, G. Kurian, S. Indupuru, P. Jannaty, J. Gu, J. Zhang, J. Malik, P. Abbeel, M.-Y. Liu, Y. Zhu, J. Jang, and L. J. Fan. Dreamdojo: A generalist robot world model from large-scale human videos. arXiv preprint arXiv:2602.06949, 2026. 2 [24] D. Geng, C. Herrmann, J. Hur, F. Cole, S. Zhang, T. Pfaff, T. Lopez-Guevara, C. Doersch, Y. Aytar, M. Rubinstein, C. Sun, O. Wang, A. Owens, and D. Sun. Motion prompting: Controlling video generation with motion trajectories. arXiv preprint arXiv:2412.02700, 2024. 2, 3, 5, 8, 9, 11 [25]N. Gillman, C. Herrmann, M. Freeman, D. Aggarwal, E. Luo, D. Sun, and C. Sun. Force prompting: Video generation models can learn and generalize physics-based control signals. arXiv preprint arXiv:2505.19386, 2025. 3 [26]Z. Gu, R. Yan, J. Lu, P. Li, Z. Dou, C. Si, Z. Dong, Q. Liu, C. Lin, Z. Liu, et al. Diffusion as shader: 3d-aware video diffusion for versatile video generation control. In SIGGRAPH, 2025. 3 [27]D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi. Dream to control: Learning behaviors by latent imagination. arXiv preprint arXiv:1912.01603, 2019. 2 [28]D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson. Learning latent dynamics for planning from pixels. In International conference on machine learning, pages 2555–2565. PMLR, 2019. 6 [29]D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap. Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104, 2023. 2 [30]A. W. Harley, Y. You, X. Sun, Y. Zheng, N. Raghuraman, Y. Gu, S. Liang, W.-H. Chu, A. Dave, S. You, et al. Alltracker: Efficient dense point tracking at high resolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5253–5262, 2025. 3, 6, 7 [31]H. He, Y. Xu, Y. Guo, G. Wetzstein, B. Dai, H. Li, and C. Yang. Cameractrl: Enabling camera control for text-to-video generation. arXiv preprint arXiv:2404.02101, 2024. 8 [32] X. He, C. Peng, Z. Liu, B. Wang, Y. Zhang, Q. Cui, F. Kang, B. Jiang, M. An, Y. Ren, et al. Matrix-game 2.0: An open-source real-time and streaming interactive world model. arXiv preprint arXiv:2508.13009, 2025. 2 [33] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Proc. NeurIPS, 2017. 8 [34] W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868, 2022. 2 [35]Y. Hong, Y. Mei, C. Ge, Y. Xu, Y. Zhou, S. Bi, Y. Hold-Geoffroy, M. Roberts, M. Fisher, E. Shechtman, et al. Relic: Interactive video world model with long-horizon memory. arXiv preprint arXiv:2512.04040, 2025. 2 24 Appendix [36]J. Huang, Q. Zhou, H. Rabeti, A. Korovko, H. Ling, X. Ren, T. Shen, J. Gao, D. Slepichev, C.-H. Lin, et al. Vipe: Video pose engine for 3d geometric perception. arXiv preprint arXiv:2508.10934, 2025. 3, 6, 7, 8, 16 [37]N. Huang, W. Zheng, C. Xu, K. Keutzer, S. Zhang, A. Kanazawa, and Q. Wang. Segment any motion in videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3406–3416, 2025. 3 [38]W. Jin, Q. Dai, C. Luo, S.-H. Baek, and S. Cho. Flovd: Optical flow meets video diffusion model for enhanced camera-controlled video synthesis. In Proc. CVPR, 2025. 3 [39]N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht. Cotracker: It is better to track together. In Proc. ECCV, 2024. 3 [40]J. Kopf, X. Rong, and J.-B. Huang. Robust consistent video depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1611–1621, 2021. 3 [41]Q. Li, Z. Xing, R. Wang, H. Zhang, Q. Dai, and Z. Wu. Magicmotion: Controllable video generation with dense-to-sparse trajectory guidance. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12112–12123, 2025. 3 [42]Z. Li, S. Niklaus, N. Snavely, and O. Wang. Neural scene flow fields for space-time view synthesis of dynamic scenes. In CVPR, 2021. 3 [43] Z. Li, R. Tucker, F. Cole, Q. Wang, L. Jin, V. Ye, A. Kanazawa, A. Holynski, and N. Snavely. Megasam: Accurate, fast and robust structure and motion from casual dynamic videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10486–10496, 2025. 3 [44] Z. Li, H.-X. Yu, W. Liu, Y. Yang, C. Herrmann, G. Wetzstein, and J. Wu. Wonderplay: Dynamic 3d scene generation from a single image and actions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9080–9090, 2025. 2, 3 [45]L. Lian, B. Shi, A. Yala, T. Darrell, and B. Li. Llm-grounded video diffusion models. arXiv preprint arXiv:2309.17444, 2023. 3 [46]F. Liang, B. Wu, J. Wang, L. Yu, K. Li, Y. Zhao, I. Misra, J.-B. Huang, P. Zhang, P. Vajda, et al. Flowvid: Taming imperfect optical flows for consistent video-to-video synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8207–8216, 2024. 3 [47]H. Lin, S. Chen, J. Liew, D. Y. Chen, Z. Li, G. Shi, J. Feng, and B. Kang. Depth anything 3: Recovering the visual space from any views. arXiv preprint arXiv:2511.10647, 2025. 5 [48]Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2022. 4 [49]S. Liu, C. Guo, B. Zhou, and J. Wang. Ponimator: Unfolding interactive pose for versatile human-human interaction animation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12068–12077, 2025. 2, 3 [50]S. Liu, Z. Ren, S. Gupta, and S. Wang. Physgen: Rigid-body physics-grounded image-to-video generation. In European Conference on Computer Vision, pages 360–378. Springer, 2024. 2, 3 [51] S. Liu, D. Y. Yao, S. Gupta, and S. Wang. Visual sync: Multi-camera synchronization via cross-view object motion. arXiv preprint arXiv:2512.02017, 2025. 3 [52] I. Loshchilov and F. Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 8 [53]J. Lv, Y. Huang, M. Yan, J. Huang, J. Liu, Y. Liu, Y. Wen, X. Chen, and S. Chen. Gpt4motion: Scripting physical motions in text-to-video generation via blender-oriented gpt planning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1430–1440, 2024. 3 [54]W.-D. K. Ma, J. P. Lewis, and W. B. Kleijn. Trailblazer: Trajectory control for diffusion-based video generation. arXiv preprint arXiv:2401.00896, 2023. 3 25 Appendix [55]A. Montanaro, L. Savant Aira, E. Aiello, D. Valsesia, and E. Magli. Motioncraft: Physics-based zero-shot video generation. Advances in Neural Information Processing Systems, 37:123155–123181, 2024. 3 [56]M. Niu, X. Cun, X. Wang, Y. Zhang, Y. Shan, and Y. Zheng. Mofa-video: Controllable image animation via generative motion field adaptions in frozen image-to-video diffusion model. In European conference on computer vision, pages 111–128. Springer, 2024. 3 [57]C. Pan, B. Yaman, T. Nesti, A. Mallik, A. G. Allievi, S. Velipasalar, and L. Ren. Vlp: Vision language planning for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14760–14769, 2024. 3 [58]W. Peebles and S. Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4195–4205, 2023. 4 [59]N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C.-Y. Wu, R. Girshick, P. Dollár, and C. Feichtenhofer. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714, 2024. 6, 7, 16 [60]X. Ren, T. Shen, J. Huang, H. Ling, Y. Lu, M. Nimier-David, T. Müller, A. Keller, S. Fidler, and J. Gao. Gen3c: 3d- informed world-consistent video generation with precise camera control. In CVPR, pages 6121–6132, 2025. 8, 9, 16 [61] C. Rockwell, J. Tung, T.-Y. Lin, M.-Y. Liu, D. F. Fouhey, and C.-H. Lin. Dynamic camera poses and where to find them. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12444–12455, 2025. 8, 9 [62]X. Shi, Z. Huang, F.-Y. Wang, W. Bian, D. Li, Y. Zhang, M. Zhang, K. C. Cheung, S. See, H. Qin, et al. Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling. In ACM SIGGRAPH 2024 Conference Papers, pages 1–11, 2024. 3 [63]A. Siarohin, S. Lathuilière, S. Tulyakov, E. Ricci, and N. Sebe. First order motion model for image animation. In NeurIPS, 2019. 3 [64]J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024. 16 [65]S. Tulyakov, M.-Y. Liu, X. Yang, and J. Kautz. Mocogan: Decomposing motion and content for video generation. In CVPR, 2018. 3 [66]T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717, 2018. 8 [67]T. Wan, A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, et al. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314, 2025. 2, 4, 8, 9, 16 [68]A. Wang, H. Huang, Z. Fang, Y. Yang, and C. Ma. Ati: Any trajectory instruction for controllable video generation. arXiv preprint, arXiv:2505.22944, 2025. 2, 8, 9, 10, 11, 13, 21, 22 [69]J. Wang, A. Ma, K. Cao, J. Zheng, Z. Zhang, J. Feng, S. Liu, Y. Ma, B. Cheng, D. Leng, et al. Wisa: World simulator assistant for physics-aware text-to-video generation. arXiv preprint arXiv:2503.08153, 2025. 8, 11 [70] J. Wang, Y. Zhang, J. Zou, Y. Zeng, G. Wei, L. Yuan, and H. Li. Boximator: Generating rich and controllable motions for video synthesis. arXiv preprint arXiv:2402.01566, 2024. 2, 3 [71]R. Wang, S. Xu, C. Dai, J. Xiang, Y. Deng, X. Tong, and J. Yang. Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision, 2024. 17 [72] X. Wang, Z. Zhu, G. Huang, X. Chen, J. Zhu, and J. Lu. Drivedreamer: Towards real-world-drive world models for autonomous driving. In European conference on computer vision, pages 55–72. Springer, 2024. 2 [73]Y. Wang, P. Bilinski, F. Bremond, and A. Dantcheva. G3AN: Disentangling appearance and motion for video generation. In CVPR, 2020. 3 26 Appendix [74]Y. Wang, F. Bremond, and A. Dantcheva. Inmodegan: Interpretable motion decomposition generative adversarial network for video generation. arXiv preprint arXiv:2101.03049, 2021. 3 [75]Z. Wang, Z. Yuan, X. Wang, T. Chen, M. Xia, P. Luo, and Y. Shan. Motionctrl: A unified and flexible motion controller for video generation. In SIGGRAPH, 2024. 2, 3 [76]T.-H. Wu, L. Lian, J. E. Gonzalez, B. Li, and T. Darrell. Self-correcting llm-controlled diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6327–6336, 2024. 3 [77]W. Wu, Z. Li, Y. Gu, R. Zhao, Y. He, D. J. Zhang, M. Z. Shou, Y. Li, T. Gao, and D. Zhang. Draganything: Motion control for anything using entity representation. In Proc. ECCV, 2024. 2, 3 [78]X. Wu, D. Paschalidou, J. Gao, A. Torralba, L. Leal-Taixé, O. Russakovsky, S. Fidler, and J. Lorraine. Where is motion from? scalable motion attribution for video generation models. In 1st Workshop on Reliable and Interactive World Model in Computer Vision Non Archival, 2026. 3 [79]J. Xing, L. Mai, C. Ham, J. Huang, A. Mahapatra, C.-W. Fu, T.-T. Wong, and F. Liu. Motioncanvas: Cinematic shot design with controllable image-to-video generation. In SIGGRAPH, 2025. 3 [80]M. Yang, Y. Du, K. Ghasemipour, J. Tompson, D. Schuurmans, and P. Abbeel. Learning interactive real-world simulators. arXiv preprint arXiv:2310.06114, 1(2):6, 2023. 2 [81]X. Yang, B. Li, Y. Zhang, Z. Yin, L. Bai, L. Ma, Z. Wang, J. Cai, T.-T. Wong, H. Lu, et al. Vlipp: Towards physically plausible video generation with vision and language informed physical prior. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12360–12370, 2025. 3 [82]Z. Yang, W. Ge, Y. Li, J. Chen, H. Li, M. An, F. Kang, H. Xue, B. Xu, Y. Yin, et al. Matrix-3d: Omnidirectional explorable 3d world generation. arXiv preprint arXiv:2508.08086, 2025. 2 [83]Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 2 [84]D. Y. Yao, A. J. Zhai, and S. Wang. Uni4d: Unifying visual foundation models for 4d modeling from a single video. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 1116–1126, 2025. 3 [85]S. Ye, Y. Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y. L. Tan, C. Zhu, J. Xiang, et al. World action models are zero-shot policies. arXiv preprint arXiv:2602.15922, 2026. 2 [86]S. Yin, C. Wu, J. Liang, J. Shi, H. Li, G. Ming, and N. Duan. Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory. arXiv preprint arXiv:2308.08089, 2023. 3 [87]L. Zhang, A. Rao, and M. Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, pages 3836–3847, 2023. 3 [88]Q. Zhang, C. Wang, A. Siarohin, P. Zhuang, Y. Xu, C. Yang, D. Lin, B. Zhou, S. Tulyakov, and H.-Y. Lee. SceneWiz3D: Towards text-guided 3D scene composition. In Proc. CVPR, 2024. 5 [89]T. Zhang, H.-X. Yu, R. Wu, B. Y. Feng, C. Zheng, N. Snavely, J. Wu, and W. T. Freeman. Physdreamer: Physics-based interaction with 3d objects via video generation. In Proc. ECCV, 2024. 2 [90]Z. Zhang, F. Cole, Z. Li, M. Rubinstein, N. Snavely, and W. T. Freeman. Structure and motion from casual videos. In European Conference on Computer Vision, pages 20–37. Springer, 2022. 3 [91]Z. Zhang, F. Long, Z. Qiu, Y. Pan, W. Liu, T. Yao, and T. Mei. Motionpro: A precise motion controller for image-to-video generation. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 27957–27967, 2025. 3 [92]W. Zhao, L. Bai, Y. Rao, J. Zhou, and J. Lu. Unipc: A unified predictor-corrector framework for fast sampling of diffusion models. Advances in Neural Information Processing Systems, 36:49842–49869, 2023. 4 27