Paper deep dive
Controlling Motion Transfer in Diffusion Transformers via Attention Heads
Sunyoung Jung, Jiwoo Park, Yoonseok Choi, Kyobin Choo, Ming-Hsuan Yang, Seong Jae Hwang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/18/2026, 1:59:03 PM
Summary
The paper introduces HALO, a head-aware controllable motion transfer framework for video Diffusion Transformers (DiTs). It identifies distinct attention heads specialized for motion (temporal heads with high displacement map alignment) and structure (low-entropy heads with diagonal attention patterns). HALO leverages these findings to refine motion cues via semantic correspondence guidance and preserve structure through selective feature injection, achieving accurate motion transfer without parameter updates.
Entities (8)
Relation Signals (6)
HALO → enables → Motion Transfer
confidence 95% · This head-level control not only enables accurate motion transfer but also provides an interpretable foundation for controllable video generation with DiTs.
HALO → uses → Motion-Specific Heads
confidence 95% · HALO constructs inter-frame displacement maps using only motion-specific heads to guide motion optimization.
HALO → uses → Structure-Specialized Heads
confidence 95% · We leverage our second insight that low-entropy attention heads primarily encode structural information... we introduce selective structural feature injection.
Motion-Specific Heads → encodes → Motion Cues
confidence 90% · A subset of heads exhibits strong patch correspondences across frames, capturing temporally coherent motion cues.
Structure-Specialized Heads → encodes → Spatial Structure
confidence 90% · Another subset of heads encodes structural information... yielding attention features with concentrated structural information.
Semantic Correspondence Refinement → refines → Displacement Map
confidence 85% · We introduce semantic guidance modules that align motion cues with semantic similarities... This refinement produces semantically aligned displacement maps.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challenging due to limited understanding of motion and structure representations within DiTs. We analyze video DiTs at the attention-head level and identify distinct heads specialized for motion and spatial structure. Based on this insight, we propose a head-aware controllable motion transfer framework that requires no parameter updates. Our method refines motion cues from motion-specialized heads via semantic correspondence guidance and preserves structure through selective feature injection. This head-level control not only enables accurate motion transfer but also provides an interpretable foundation for controllable video generation with DiTs.
Tags
Links
- Source: https://arxiv.org/abs/2607.11081v1
- Canonical: https://arxiv.org/abs/2607.11081v1
Trouble viewing inline? Open PDF directly →
Full Text
77,458 characters extracted from source content.
Expand or collapse full text
Controlling Motion Transfer in Diffusion Transformers via Attention Heads Sunyoung Jung 1* , Jiwoo Park 1,2* , Yoonseok Choi 1 , Kyobin Choo 1 , Ming-Hsuan Yang 3 , and Seong Jae Hwang 1† 1 Yonsei University 2 LG Electronics 3 University of California, Merced Project page: https://sunyj-hxppy.github.io/halo GWTF A red Porschedrives beside a lakeat sunrise. Benchmark Reference Ours Reference In-the-wildMovie Scene Ours GWTF Two stormtroopersriding mechanical steedsacross a concrete field. Fig. 1: Overview. We present HALO, a head-aware controllable motion transfer framework for video Diffusion Transformers, which identifies motion- and structure- specialized attention heads within the model. Leveraging these findings, HALO gener- ates videos that follow the target prompt while remaining motion- and structure-aligned with reference videos, achieving accurate motion transfer. Abstract. Diffusion Transformers (DiTs) have advanced video gener- ation with high-quality, temporally coherent results. However, extend- ing them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challenging due to lim- ited understanding of motion and structure representations within DiTs. We analyze video DiTs at the attention-head level and identify distinct heads specialized for motion and spatial structure. Based on this insight, we propose a head-aware controllable motion transfer framework that requires no parameter updates. Our method refines motion cues from motion-specialized heads via semantic correspondence guidance and pre- serves structure through selective feature injection. This head-level con- trol not only enables accurate motion transfer but also provides an in- terpretable foundation for controllable video generation with DiTs. Keywords: Motion Transfer· Diffusion Transformers· Attention Heads 1 Introduction Motion transfer in video generation synthesizes a video that follows the motion of a reference video while adhering to a target prompt. The primary objectives * Equal contribution. † Corresponding author. arXiv:2607.11081v1 [cs.CV] 13 Jul 2026 2S. Jung et al. are (1) motion fidelity, ensuring temporal adherence to the reference motion, and (2) structural alignment, maintaining spatial layout of the reference [14,48]. Achieving these goals requires modeling of spatio-temporal dependencies, an area in which recent video Diffusion Transformers (DiTs) [25, 40, 46, 47] have shown strong capability. Given their ability to capture spatial structure and temporal dynamics, DiTs have become a natural choice for motion transfer [7,14,35]. Existing DiT-based motion transfer approaches, such as noise warping [7], rotary positional embedding manipulation [14], and cross-frame attention op- timization [35], offer varying degrees of motion controllability. However, these methods focus on manipulating motion representations without an understand- ing of how motion and structure are encoded within DiTs. This lack of under- standing is due to the distributed functionality of DiTs, which makes their inter- nal mechanisms challenging to analyze [3,8,12]. Consequently, generated videos often exhibit seemingly plausible motion yet inaccurate trajectories or misaligned object structures relative to the reference video. As illustrated in Fig. 1, Go-with- the-flow (GWTF) [7], a state-of-the-art method, exhibits motion deviations, in- cluding a car moving straight instead of turning and two stormtroopers with misaligned positions compared to the reference. Therefore, we conduct a detailed analysis of video DiTs, focusing on the attention heads to answer the fundamental question: How are motion and struc- tural cues internally encoded in video DiTs? To identify the heads specialized in modeling motion, we introduce the first head-level analysis based on displace- ment maps. The displacement map [35] encodes motion as patch-wise coordinate differences between frames, making it effective for analyzing motion properties of heads. For structural cues, we compute the visual token attention-map entropy, which represents the uncertainty in patch dependency. By utilizing the entropy, we select heads that reliably capture structural information. Our analysis identifies two distinct attention head subsets within video DiTs: (1) Motion-Specific Heads. A subset of heads exhibits strong patch correspon- dences across frames, capturing temporally coherent motion cues in the dis- placement maps. These heads present clear displacement maps that accurately reflect the motion, demonstrating their superior capability to represent coherent motion flow. (2) Structure-Specialized Heads. Another subset of heads encodes structural information. The attention maps of the heads show low entropy, char- acterized by sharply diagonal patterns. This attention pattern yields attention features with concentrated structural information, confirming that attention map entropy serves as an indicator of the structural content. Building on these findings, we propose HALO, a head-driven semantic-structu ral motion transfer framework that simultaneously enforces motion fidelity and structural alignment. By leveraging the properties of attention heads in video DiTs, our approach first constructs inter-frame displacement maps using only motion-specific heads to guide motion optimization. These displacement maps capture the motion dynamics but lack semantic information. This often leads to transferring motion into semantically irrelevant regions, causing inconsistent or reversed object motion, as shown in Fig. 2(a). To overcome this issue, we Controlling Motion Transfer3 The motorcycleis driving down the highway.A lionwalking on a sandy pathin a wildlife reserve. HALO Displacement - only Optimization Reference (a)(b) Fig. 2: Limitations of displacement-based motion transfer. Comparison between HALO and displacement-only optimization [35]. (a) Lack of semantic alignment causes motion errors. (b) Missing structural preservation leads to spatial misalignment. HALO ensures consistent motion and spatial fidelity. introduce semantic guidance modules that align motion cues with semantic sim- ilarities derived from diffusion features, known to encode rich object-level seman- tics [38,41]. This refinement produces semantically aligned displacement maps, enabling temporally consistent and semantically coherent motion transfer. While the refined displacement map provides motion flow information, it remains insufficient for preserving the spatial layout of the reference video. Con- sequently, the generated results often exhibit structural misalignment with the reference, as illustrated in Fig. 2(b). To overcome this limitation, we leverage our second insight that low-entropy attention heads primarily encode struc- tural information. Specifically, we introduce selective structural feature injec- tion by incorporating attention features from the reference video through these low-entropy heads. This head selection reinforces structural information while suppressing noisy and diffuse cues, yielding motion transfer that is structurally aligned with the reference. The main contributions of this work are: – We present a head-level functional analysis of video DiTs, revealing motion-specific and structure-specialized attention heads and validating them as control primitives for controllable video generation. – We propose HALO, a head-aware controllable motion transfer framework that enhances motion fidelity through semantic-aware displacement op- timization and preserves spatial consistency via selective feature injec- tion from structurally informative heads. – We perform extensive evaluations on standard motion transfer benchmarks and our Movie Scene Dataset, demonstrating that it achieves better motion coherence and structural alignment. 4S. Jung et al. 2 Related Work Text-to-Video Generation. Early Text-to-Video (T2V) generation was driven by U-Net-based diffusion models [9,23,29], which extended image diffusion with temporal attention or 3D convolutions to capture spatio-temporal dynamics [4, 6,15]. More recently, DiTs have emerged as the dominant paradigm [10,16,25,40, 47], significantly improving temporal consistency and visual quality. Motivated by these advances, we build our framework upon a video DiT backbone. Motion Transfer. Motion transfer aims to generate a target video by extracting and reflecting only the motion information from a reference video [13,33,37,48]. The main challenge of this task lies in effectively decoupling the motion and appearance information from the reference [21,45]. Early studies have addressed this challenge by extracting motion embedding from the temporal modules of U-Net [27,42,50]. Recently, the strong generative capability of DiTs has motivated their adoption in the motion transfer field. For instance, GWTF [7] explicitly warps noise to manipulate the motion, while RoPECraft [14] handles motion by warping the RoPE embeddings using optical flow. Furthermore, DiTFlow [35] derives displacement from cross-frame attention to guide patch movement. While these prior works have explored various ways to handle motion information in DiTs, there has been a lack of direct analysis of how motion and structure are encoded within DiTs’ internal representations. To bridge this analytical gap, our work analyzes DiT internal representa- tions to identify motion-specific and structure-specialized heads. By leveraging these heads, we achieve training-free motion transfer and extend our approach to controllable video generation. Attention in Video DiTs. U-Net-based video diffusion models [9, 23] explic- itly separate spatial and temporal attention, enabling motion extraction from temporal modules [21,27]. In contrast, DiTs [40,47] adopt unified attention that jointly models spatial and temporal information. While this design enhances expressiveness, it obscures the distinction between motion and structure, com- plicating motion-specific control [28,35]. In contrast, we posit that this motion information is independently encoded within the unified structure of attention in the DiTs. Consequently, we focus on analyzing this unified attention structure at the level of individual attention heads to identify the distinct encoding of motion and structure in terms of motion transfer. 3 Analyzing Attention Heads of Video DiTs We focus on understanding how motion and structural information are repre- sented within video DiTs [47]. Here, we present a head-level analysis of attention heads to uncover their distinct functional roles in modeling motion and spatial structure, offering a mechanistic perspective overlooked in prior studies [1,24,44]. Analysis 1: Motion-Specific Heads. We aim to identify the motion-specific properties encoded within individual attention heads. Building on SparseGen [44], which identifies temporal and spatial head patterns for efficient video generation Controlling Motion Transfer5 Fig. 3: Head Configuration Comparison. (a) Attention maps show distinct patterns: temporal heads capture cross-frame diag- onals, while spatial heads maintain intra- frame locality. (b) Quantitative evaluation using directional alignment and correla- tion shows that temporal heads align more closely with reference motion. (c) Displace- ment maps further confirm temporal heads more accurately capture motion flow. Attention Map Reference 3.34 / 0.44 Feature Map Entropy Distribution 6.66 / 0.537.89 / 0.77 L20 H8L20 H15L20 H6 0.35 Attention Entropy / Feature Entropy 0.40.50.450.3 ℋ hidden ℋ attn 2.0-4.0 4.0-6.0 6.0-8.0 8.0-10.0 Fig. 4: Relationship between attention maps and structural cues in hidden fea- tures. Heads with lower attention en- tropy show lower feature entropy, indi- cating stronger structural fidelity. Ex- ample: L20 H8 denotes the 8th atten- tion head in the 20th layer. (see Fig. 3(a)), we first classify heads based on their similarity to pattern masks (e.g., cross-frame diagonal for temporal, sparse localized patterns for spatial). Based on this, we construct head-specific displacement maps from cross-frame at- tention [31,35]. Specifically, cross-frame attention is computed using the queries Q and keys K derived from the latent features z f ∈R head×N×d as: M h f,f ′ = softmax Q h f · K h⊤ f ′ √ d ! , (1) where h denotes the attention head, N = H×W is the number of spatial tokens, and d is the feature dimension. Here, f and f ′ represent the frame indices. We then derive a head-specific displacement map D h f,f ′ ∈R H×W×2 , which captures the patch-wise motion between frames (f,f ′ ), defined as: I h f,f ′ = argmax f ′ M h f,f ′ , D h f,f ′ (i) =g I h f,f ′ (i) −g(i),(2) whereg(·) converts a 1D patch index into 2D spatial coordinates and i denotes the patch index. Using displacement maps, we analyze cross-frame motion patterns across at- tention heads to quantify motion capture effectiveness, as shown in Fig. 3(c). Temporal heads capture motion variations more accurately than spatial or com- bined heads, showing stronger alignment with object motion (blue) and back- ground movement (red). To validate these observations, we measure the simi- larity between predicted displacement maps and ground-truth optical flow [39] 6S. Jung et al. through two quantitative metrics: (1) Directional Alignment (DA), calculating cosine similarity between motion directions, and (2) Correlation (Corr), using Pearson correlation for global pattern consistency. As illustrated in Fig. 3(b), temporal heads consistently achieve higher DA and Corr scores. These results confirm that while spatial heads attend to intra- frame patch relationships, temporal heads specialize in cross-frame dependencies, effectively encoding motion-related information within video DiTs. Analysis 2: Structure-Specialized Heads. To identify attention heads that encode structural information, we analyze the visual token attention maps A h corresponding to each head. Our analysis reveals distinct patterns across differ- ent heads: Heads exhibiting well-defined diagonal alignment patterns with low entropy (e.g., L20 H8) stand in contrast to those with irregular or diffuse atten- tion distributions with high entropy (e.g., L20 H6), as shown in Fig. 4. Since the entropy serves to quantify the uncertainty in capturing the relationships between visual tokens, low entropy implies the presence of structural information. To validate this, we compute the attention entropy [32] and the spatial en- tropy of corresponding attention features [5,22,34]. Given a head h with attention map A h , the attention entropy H h attn is defined as H h attn = − P N k=1 A h k logA h k , where lower values indicate a more concentrated attention distribution. To cap- ture spatial properties in attention features, we compute the spatial entropy of PCA-transformed feature embeddings. As illustrated in Fig. 4, lower entropy heads exhibit clear diagonal atten- tion patterns and yield feature maps maintaining structural layouts. Fig. 4 also provides quantitative evidence of a strong correlation between lower attention entropyH attn and lower hidden feature spatial entropyH hidden . Therefore, these diagonal attention patterns indicate strong spatial patch correlations, validating that such heads are specialized in encoding structural information. 4 Method We introduce HALO, a head-aware controllable motion transfer framework that jointly achieves motion fidelity and structural alignment with a reference video (see Fig. 5(a)). Building on Sec. 3, we propose (i) semantic-aware motion opti- mization (Sec. 4.1) and (i) structure-guided feature injection (Sec. 4.2), illus- trated in Fig. 5(b)–(c). 4.1 Semantic-Aware Motion Guidance For motion transfer, we construct cross-frame displacement maps within DiTs. Our analysis (Analysis 1, Sec. 3) shows that motion-specific heads M capture motion more effectively. We therefore aggregate cross-frame attentionM h f,f ′ for h∈M, identified via temporal pattern masks to extract a displacement map. While cross-frame attention provides patch-level correlation cues, it can overem- phasize visually similar yet semantically unrelated regions, producing mismatched displacements and degraded motion alignment (e.g., camera-following move- ments that disregard objects and incorrect object trajectories; see Supp. Sec. B). Controlling Motion Transfer7 Fig. 5: Overview of HALO. (a) From a reference video, we extract displacement maps and head features. Displacements from motion-specific heads guide motion by opti- mizing the latent representation, while selected head features from the reference are injected to preserve structure during generation. (b) To enhance motion guidance, se- mantic correspondence derived from diffusion features refines the displacement map, ensuring semantically aligned motion flow. (c) For structural guidance, we inject value features from heads selected via entropy analysis, targeting heads that encode essential spatial information. To mitigate this, we introduce a semantic-aware motion guidance that refines displacements using semantic similarities from semantically rich diffusion fea- tures [17,38,41], as shown in Fig. 5(b). Semantic Correspondence Refinement. To refine the reference displace- ment map, we first select the top-k attention candidates from cross-frame at- tention rather than relying on a single maximum, as illustrated in Fig. 6(a). We then construct semantic correspondence C using pairwise cosine similarity over diffusion features φ: C f,f ′ = φ f · φ f ′ ∥φ f ∥ 2 ∥φ f ′ ∥ 2 ,(3) For each patch index, we compute the distances between the most similar se- mantic indices I cor f,f ′ ∈R N between frames and the top-k attention candidates. The nearest candidate is then selected as the reference best-match coordinate I ref f,f ′ . The reference inter-frame displacement map D ref is then constructed from I ref f,f ′ following Eq. 2 (first row in Fig. 7). Motion Optimization. Given D ref , we optimize z T to align generated motion with the reference. We apply Semantic Reweighting (SRW) to adjust cross-frame attention via a correspondence-guided bias derived fromC (Fig. 6(b)). Using I cor , 8S. Jung et al. Fig. 6: Details of Semantic Correspondence Refinement (SCR) and Semantic Reweight- ing (SRW). (a) I ref f,f ′ is obtained by choosing, among top-k attention candidates, the patch closest to the semantic best match I cor f,f ′ . (b) SRW manipulates target cross-frame attention by adding a correspondence-based bias at I cor f,f ′ . we form the bias matrix B f,f ′ : B f,f ′ (i,j) = ( β, if j = I cor f,f ′ (i), 0, otherwise, (4) where β controls bias strength. The refined cross-frame attention scores are ̃ E f,f ′ = E f,f ′ +B f,f ′ , and refined attention is ̃ M f,f ′ = softmax( ̃ E f,f ′ ). We then estimate displacement as an expectation over coordinate differences, ∆x f,f ′ = X i,j (x j − x i ) ̃ M f,f ′ (i,j), ∆y f,f ′ = X i,j (y j − y i ) ̃ M f,f ′ (i,j), (5) and stack them in D t = [∆x f,f ′ ,∆y f,f ′ ] (second row in Fig. 7). Finally, to optimize the latent for motion alignment, we minimize the semantic motion loss L SM , defined as the L2 distance between the reference displacement map D ref and the target displacement map D t . Incorporating semantic cues into dis- placement estimation unifies motion alignment with semantic coherence, yielding consistent motion transfer. Notably, this strategy represents the first attempt to incorporate semantic correspondences within motion representations, facilitating high-fidelity, semantically consistent video generation. 4.2 Selective Structural Head Injection Although displacement maps capture directional motion, they lack sufficient intra-frame structure, which can cause spatial misalignment with the reference (Fig. 2, right). To address this, we propose a structural guidance mechanism for preserving the spatial integrity of the reference. We inject reference value fea- tures during generation. Naïvely injecting all head features, however, introduces artifacts such as noise and identity leakage [2] (see Supp. Sec. C.4, Supp. Fig. 6). Controlling Motion Transfer9 SCR Reference Displacement Video SRW Target The person is riding a motorbike on a muddy course. Fig. 7: Effect of SCR and SRW on displacement maps. The refinements improve ro- bustness to fine-grained object motion and better preserve object shape. We therefore select structurally informative heads via entropy-based anal- ysis (Analysis 2, Sec. 3). Low-entropy heads exhibit strong diagonal attention and produce spatially coherent features; we inject their value features into the corresponding heads, supplying structure-aware guidance without noise or leak- age (Fig. 5(c)). This selective injection complements motion optimization and enforces spatial alignment with the reference. 5 Experiments Implementation Details. Following prior works, we evaluate our approach using a standard motion transfer benchmark and settings [14, 36, 48]. Across all experiments, we adopt the video diffusion transformer CogVideoX [47] as the base model, using 50 denoising steps, 12 optimization steps T opt , and a guidance scale of 7, consistent with common practice. For additional results with the video DiT model, Wan [40], please refer to Supp. Sec. D. The bias strength β for semantic reweighting is fixed to 0.1, and the entropy range τ for selective structure specialized-head injection is set to τ < 7, based on 7 being the median entropy value observed in our analysis (validated in Section 5.3). Hyperparameter configurations (e.g., β, T opt , τ, top-k) and HALO optimization algorithm are detailed in Supp. Sec. C.1 and C.2, respectively. Baselines. We compare against representative motion-transfer approaches, in- cluding U-Net-based methods (MOFT [45], ConMO [13], and MotionClone [27]) and recent DiT-based methods (DiTFlow [35], RoPECraft [14], and GWTF [7]). Movie Scene Dataset. Motion transfer has emerged as a pivotal task in video production and digital cinematography, where it is increasingly demanded to replicate sophisticated film-style shot compositions and intricate post-production effects [30]. However, existing benchmarks [27,36] are largely derived from generic internet videos (e.g., DAVIS/WebVid), which do not fully reflect these profes- sional requirements. Therefore, we curate a Movie Scene Dataset specifically for the motion transfer task (see Supp. Sec. E for details). Complementing existing benchmarks, this dataset provides a production-oriented evaluation setting to assess whether methods can achieve faithful motion transfer under real-world cinematic conditions. 10S. Jung et al. Fig. 8: Qualitative comparison between U-Net- and DiT-based baselines and HALO. 5.1 Main Results Qualitative Results. Fig. 8 presents qualitative comparisons with state-of-the- art models. Baseline methods often exhibit camera drift, generation failures, or identity leakage (e.g., generating a motorcycle instead of a horse). They also pro- duce semantically inconsistent motion, such as misaligned orientations between correlated objects (e.g., car vs. airplane). In contrast, our method produces se- mantically coherent motion aligned with the reference while preserving subject identity and structural integrity. Furthermore, in Fig. 9, HALO exhibits remark- able robustness on the Movie Scene Dataset. Even when guided by complex cin- ematic prompts, HALO faithfully preserves the target’s identity integrity while following the reference dynamics. Quantitative Results. We evaluate performance using CLIP score (CLIP) [18] for text–video alignment, Temporal Consistency (TC) [49] for frame coherence, and Motion Fidelity (MF) [48] and FTD [14] for reference-motion alignment. As shown in Table 1, HALO achieves notable improvements in CLIP, MF, and FTD, demonstrating that head-level analysis within DiTs enables motion transfer with strong motion fidelity and structural alignment. Since motion transfer requires following the reference motion while adher- ing to the target text prompt, CLIP and MF must be jointly evaluated. Prior Controlling Motion Transfer11 Hike: 0, 18, 29 Horse-jumplow: 0, 18, 39 Bus: 0, 25, 48 ReferenceHALOMotionCloneDiTFlowGWTF A small Minions walks past one another in a loose line, moving steadily across the green space. A SpongeBob-style adventurer rides a bouncing anemone-creature across a smooth open ground. RoPECraft Fig. 9: Qualitative results in the movie scene dataset. works exhibit a trade-off between these metrics, as shown in Table 1. In con- trast, our method achieves balanced performance, maintaining high CLIP and MF scores. This indicates that our head-analysis-based formulation transfers reference motion effectively while suppressing irrelevant semantic information. Although HALO reports a slightly lower TC than DiTFlow and GWTF in Ta- ble 1, TC tends to favor more static outputs. Further analysis of the CLIP–MF balance and the relationship between motion dynamics and temporal consistency is provided in Supp. Sec. F.2. We also report the runtime and peak GPU memory in Supp. Sec. F.3, showing that HALO incurs only marginal overhead relative to the displacement optimization framework. User Study. To ensure an assessment aligned with human preference, we con- duct a user study to evaluate HALO across three criteria: editing accuracy, temporal consistency, and motion accuracy. Twenty participants rated each re- sult on a 5-point Likert scale (1: lowest, 5: highest), with scores subsequently normalized to a 0–100 range. As reported in Table 4, HALO achieves the high- est mean score across all criteria, demonstrating a consistent preference among participants. 5.2 Ablation Study We perform ablation studies under conditions identical to the main experiments, using displacement optimization [35] as the baseline. Effect of Semantic-Aware Motion Guidance. To evaluate the role of seman- tic correspondence refinement in displacement computation, we directly compute displacement maps from raw cross-frame attention. As shown in Fig. 10 (green 12S. Jung et al. Table 1: Quantitative results with state-of-the-art motion transfer methods. Best and second results are represented with bold and underlined . ModelCLIP ↑ TC ↑ MF ↑ FTD ↓ U-Net-based MoFT [45]NeurIPS’24 30.9 85.8 34.8 23.0 ConMO [13]CVPR’2529.8 85.3 52.0 17.4 MotionClone [27] ICLR’2530.4 78.6 55.6 19.8 DiT-based DiTFlow [35]CVPR’2531.0 89.5 59.6 23.0 RoPECraft [14] NeurIPS’25 30.3 85.9 58.2 19.6 GWTF [7]CVPR’2531.6 88.262.521.6 HALO31.787.566.219.4 Table 2: Quantitative results in the Movie Scene dataset (meth- ods requiring video masks are ex- cluded). ModelCLIP ↑ TC ↑ MF ↑ MotionClone [27] 27.1 74.5 47.7 DiTFlow [35]29.7 90.4 48.4 RoPECraft [14]28.3 88.2 49.1 GWTF [7]30.287.6 46.6 HALO30.588.952.6 Table 3: Ablation results of proposed components on performance. “Semantic” denotes the semantic- aware motion guidance, and “Injection” denotes the selective structural head injection. Exp.# Semantic Injection CLIP ↑ TC ↑ MF ↑ FTD ↓ 131.0 89.5 59.6 23.0 2✓30.6 88.3 61.8 20.5 3✓30.9 88.5 61.3 20.9 4 (HALO)✓31.787.566.219.4 Table 4: User preference study on motion transfer. The results represent the preference rate across three dimensions. ModelEdit Acc. TC Motion Acc. DiTFlow84.1 83.169.0 RoPECraft 72.4 75.677.6 GWTF84.0 77.172.8 HALO86.587.893.3 Exp.#3 The penguin is walking in a zoo enclosure, possibly towards its habitat pool. ReferenceHALOExp.#1Exp.#2 Fig. 10: Qualitative results of our method across ablation studies. Exp.# corresponds to Table 3. box), the absence of semantic refinement causes motion to the background rather than the actual moving object. This demonstrates that cross-frame attention alone captures visually similar but semantically unrelated regions, leading to motion misalignment. Semantic correspondence refinement is therefore crucial for achieving semantically consistent motion transfer. Effect of Selective Structure-Specialized Head Injection. Removing se- lective structure-specialized head injection significantly degrades motion fidelity (Exp.#3 in Table 3). As shown in Fig. 10 (yellow box), this omission results in structural distortions and duplicated objects (e.g., two penguins). These find- ings indicate that injecting low-entropy structural features helps maintain spatial consistency, making explicit structural guidance essential for accurate motion transfer. Controlling Motion Transfer13 Table 5: Performance of motion-transfer methods across different motion difficulty levels categorized based on the complexity of motion trajectories. Model HardMediumEasy CLIP ↑ MF ↑ CLIP ↑ MF ↑ CLIP ↑ MF ↑ MotionClone 26.8 39.0 24.8 37.0 27.8 39.0 ConMo30.7 51.0 31.0 54.0 28.3 64.0 DiTFlow32.1 54.0 32.5 67.0 32.1 66.0 RoPECraft 27.1 55.7 24.7 69.9 24.1 59.5 GWTF33.6 56.0 32.7 69.0 30.8 74.0 HALO33.259.233.872.331.475.5 Table 6: Analysis of head config- uration and selective injection en- tropy threshold. (a) Different attention heads. (b) Validation of selective struc- tural head injection entropy range τ. (a) Attention Head Head MF ↑ FTD ↓ All63.1 21.6 Spatial 53.6 22.7 Motion66.219.4 (b) Entropy Range τ MF ↑ FTD ↓ 7 < 54.1 23.5 < 766.219.4 5.3 Additional Validations Complex Motion. Robust motion transfer remains challenging under rapid movements and long-range trajectories. To evaluate the robustness of HALO in such settings, we utilize a benchmark stratified into three difficulty levels (Easy/Medium/Hard) based on trajectory complexity. We assess performance using prompt alignment (CLIP) and motion fidelity (MF), while omitting MoFT due to its lack of competitiveness in these scenarios. As summarized in Table 5, HALO consistently achieves superior performance across all levels with minimal performance decay, indicating high stability. This robustness is further evidenced qualitatively in Fig. 11(a); our method handles extreme cases effectively, such as occlusions in a motorcycle jump and the high- velocity dynamics of boxing. Overall, these results validate that our approach remains stable across a wide spectrum of complex motions. Validation of Motion-Specific Heads. To verify motion-head selection, Ta- ble 6a compares motion-specific, spatial, and all-head configurations. Spatial heads alone yield significantly lower MF and FTD, indicating a limited motion- related signal for motion transfer. Although aggregating all heads offers a slight improvement, it still underperforms motion-specific heads. These results suggest that incorporating spatial heads can dilute relevant motion cues; in contrast, iso- lating motion-specific heads better captures dynamic correlations essential for motion transfer. Qualitative comparisons are included in Supp. Sec. C.5. To demonstrate the broader applicability of motion-specific heads as a motion- control mechanism, we conduct additional experiments focusing on motion dy- namics. Building on Layer Perturbation Guidance [31], we extend this approach to our motion-specific head by introducing head perturbation during inference. Evaluation using VBench metrics [19] on 100 video samples reveals that inte- grating this method with CogVideoX [47] improves image quality by +14%, aesthetics by +8%, and motion dynamics by +15%. These results suggest that our motion-head analysis provides a lightweight yet effective mechanism to guide motion in DiT-based architectures. Validation of Structure-Specialized Heads. Table 6b presents results under different entropy thresholds for structure-specialized head selection. Based on 14S. Jung et al. Fig. 11: Qualitative results on complex motion and video editing. (a) Generated sam- ples illustrating the HALO’s capability in handling complex motion dynamics. (b) Application to video editing using the structure-specialized head. our analysis showing a median entropy of 7, we adopt τ=7 as the threshold. Injecting features from high-entropy heads (τ>7) yields lower MF and FTD scores, indicating that such heads contain weaker, more diffuse structural cues. This validates the relationship between attention-map entropy and structural information in the context of motion transfer. Additional experiments over a wider entropy range and qualitative examples are provided in Supp. Sec. C.1. To further assess the applicability of structure-specialized heads, we extend their utility to video editing, a task where preserving structural integrity is cru- cial. Specifically, we perform the attention value injection within these selected structure-specialized heads. As shown in Fig. 11(b), our approach successfully modifies subject identity (e.g., rhino→lion/bear) while maintaining the orig- inal layout and structural details. These results demonstrate that structure- specialized heads serve as a precise mechanism for structure-preserving video editing, extending their effectiveness beyond motion transfer. 6 Conclusion In this work, we demonstrate for the first time that video DiTs contain dis- tinct motion-specific and structure-specialized heads responsible for temporal dynamics and spatial organization. Based on this insight, we present HALO, a training-free motion transfer framework that leverages these heads for high motion fidelity and structural alignment. Our method integrates semantic cor- respondences with motion representations and proposes a head-wise injection to enable motion-consistent, structure-aligned transfer. Extensive experiments val- idate that HALO outperforms prior approaches, highlighting the effectiveness of head-level analysis. Furthermore, by introducing our Movie Scene Dataset, we provide a new direction for future research in practical motion transfer. Future Work. Our findings reveal that video DiTs inherently disentangle mo- tion and structure across attention heads, opening new directions for controllable video generation. Future research may explore explicit control over motion di- rection and intensity for fine-grained editing. Controlling Motion Transfer15 Acknowledgments This work was supported in part by the IITP RS-2024-00457882 (AI Research Hub Project), IITP 2020-I201361, NRF RS-2024-00345806, and NRF RS-2023- 002620, while the authors are affiliated with the Department of Artificial In- telligence (S.J., J.P, S.J.H.) and the Department of Computer Science (K.C.). This work was supported by the AI Seoul Tech Research Support Program of the Seoul Future Foundation. We would like to express our sincere gratitude to RASCA FX for providing access to the movie dataset used in this work. We also thank the RASCA FX team for their valuable support and collaboration throughout this project. 16S. Jung et al. Controlling Motion Transfer in Diffusion Transformers via Attention Heads Supplementary Material Sunyoung Jung 1∗ Jiwoo Park 1,2∗ Yoonseok Choi 1 Kyobin Choo 1 Ming-Hsuan Yang 3 Seong Jae Hwang 1 1 Yonsei University 2 LG Electronics 3 University of California, Merced Contents A Code & Presentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .16 B Details of Problems in Previous Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 C Details of HALO . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .18 D Application to the Video DiT Model: Wan . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 E Movie Scene Dataset . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26 F Additional Experiment Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 G Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33 A Code & Presentation The code for HALO will be released to the public. In addition, we have prepared a presentation video that introduces HALO in a clear and accessible manner, il- lustrating qualitative results and representative video demonstrations. This video is included as the file presentation.mp4. B Details of Problems in Previous Methods This section discusses the limitations of existing DiT-based motion transfer ap- proaches [7, 13, 14], which largely arise from the lack of understanding of the functional roles of attention heads. We further examine challenges observed in cross-frame attention-based methods [35]. Although displacement maps can cap- ture motion dynamics, they do not encode semantic correspondence, often re- sulting in motion being transferred to irrelevant regions and causing inconsistent or even reversed object motion. Results of DiT-based Methods. We analyze the outputs of prior DiT-based motion transfer approaches and identify several characteristic failure patterns. As illustrated in Fig.12(a), existing methods often produce misaligned or dis- torted object structures: although the peacock roughly follows the reference mo- tion trajectory, it suffers from severe morphological inconsistency. Furthermore, Fig.12(b) reveals identity leakage from the reference video. In motion transfer, it Controlling Motion Transfer17 Fig. 12: Limitations of existing DiT-based motion transfer methods. (a) The gener- ated subject exhibits structural distortion and misalignment, failing to preserve the spatial layout of the reference. (b) Identity cues from the reference video leak into the generated subject, causing identity corrup- tion and appearance inconsistency across frames. Fig. 13: Problems of cross-frame attention-based motion transfer methods. (a) The generated video reproduces the reference camera motion, but the object fails to follow the reference object’s trajectory, leading to degraded motion fidelity. (b) The generated video omits the moving object, and its motion is incorrectly mapped onto the background. is essential to preserve the semantics specified by the target prompt while align- ing with the reference motion. However, previous methods inherit appearance cues from the reference, indicating insufficient disentanglement between content and motion representations. In contrast, HALO achieves accurate motion trans- fer, faithfully following the reference motion and adhering to the intended target semantics. Results of Cross-Frame Attention-based methods. Cross-frame attention- based motion transfer approaches [35] suffer from a fundamental limitation: be- cause they derive correspondence solely from query-key similarity, the resulting displacement maps tend to reflect coarse attention activation patterns rather than semantic alignment. As shown in Fig.13(a), even when the global camera motion correctly mirrors the reference (e.g., a right-to-left sweep), the generated object often follows an incoherent spatial trajectory, causing the layout to diverge from the reference. This discrepancy arises when spurious attention activations dominate displacement estimation in semantically irrelevant regions. Fig.13(b) illustrates an additional failure mode: the foreground object disappears while its motion is incorrectly transferred to the background. In this case, the intended motion of the foreground subject is improperly transferred to the background, leading to degraded motion alignment. These shortcomings highlight the necessity of integrating semantic correspon- dence refinement within the motion transfer pipeline. Therefore, HALO achieves precise and coherent synthesis by grounding motion guidance in semantic align- ment rather than coarse feature similarity. To the best of our knowledge, HALO 18S. Jung et al. Table 7: Optimization steps T opt results. T opt CLIP ↑MF ↑FTD ↓ 1030.950.126.2 1231.766.219.4 1530.556.022.7 Table 8: Bias strength β analysis results. βCLIP ↑MF ↑FTD ↓ 0.131.766.219.4 0.330.753.623.4 0.530.753.923.7 Table 9: Hyperparameter sensitivity. 4 5 6 7 8 9 MF 67.0 67.2 67.1 72.3 70.2 69.1 (a) τ sensitivity 0.01 0.05 0.1 MF 66.6 68.1 69.8 (b) β sensitivity 3 45 MF 75.4 76.1 73.1 (c) top-k sensitivity represents the first attempt to explicitly bridge the gap between high-level se- mantic correspondence and low-level attention-based motion representations. C Details of HALO We provide the hyperparameter configurations in Sec. C.1. Our extensive exper- iments demonstrate that HALO is highly robust to hyperparameter variations, maintaining consistency across diverse videos and model architectures without the need for specific tuning. The full algorithmic procedure for HALO is provided in Sec. C.2. C.1 Hyperparameters Optimization iteration. We set the optimization steps to T opt =12. Increasing T opt improves quantitative performance, especially MF, but yields diminishing returns after 12 iterations (see Table 7). We therefore choose T opt =12 to balance accuracy and efficiency. Entropy range. Beyond the broad range discussed in the main text, Table 9(a) provides a fine-grained entropy range τ analysis on the curated dataset. MF degrades when the value deviates from 7, indicating that an entropy threshold of 7 optimally distinguishes structural information. Bias strength. Semantic bias strength β scales the correspondence-guided bias on displacement maps in Semantic Reweighting (SRW). Table 8 shows that a small bias strength yields the best overall performance, whereas larger values lead to degraded results. These results indicate that an excessive semantic bias can suppress motion dynamics in the displacement maps, thereby compromising the motion fidelity. Based on this observation, we set β=0.1 in all experiments. A finer-grained sweep on the curated dataset further shows stable performance within the low-bias range, with β=0.1 achieving the best result, as shown in Table 9(b). Top-k Selection. To construct the reference displacement map, Controlling Motion Transfer19 ReferenceHALO (k=4)k=10k=20 A zebra walking on a gravel path in a safari park. Fig. 14: Qualitative results with corresponding reference displacement maps D ref in hyperparameter Top-k analysis. Table 10: Hyperparameter Top-k analysis. Top-k CLIP ↑ MF ↑ FTD ↓ 431.766.219.4 1031.245.226.5 2030.839.630.3 we extract the top-4 candidates from cross-frame attention and then select the final match by comparing their distances to the semantically best- aligned location. To justify the choice of k=4, we additionally evaluate larger candidates (k=10, k=20). As shown in Table 10, increasing k consistently degrades all metrics, indicat- ing that using the top 4 candidates yields the most stable and reliable displace- ment estimation. This performance drop arises from inaccurate displacement predictions caused by irrelevant cross-frame attention candidates. As illustrated in Fig.14, while HALO with k=4 produces clean, object-aligned displacement maps, larger k settings (k=10, 20) generate noisy and spatially inconsistent maps. These inaccuracies propagate directly into motion transfer, ultimately leading to weaker motion alignment. Additionally, evaluating the performance sensitivity within a narrower top-k range (3, 4, and 5) yields consistent results (Table 9(c)), demonstrating low variance and robustness to this specific setting. C.2 Algorithm Algorithm 1 outlines the inference pipeline of HALO, encompassing both the motion-optimization and the structural feature-injection steps. Notably, this en- tire procedure is training-free, requiring no additional parameters or model fine- tuning. C.3 Metric Details for Motion-Specific Heads To verify that our head-selection strategy isolates motion-specific behavior, we compare displacement fields from temporal, spatial, and all-head configurations against ground-truth optical flow (RAFT [39]) using Directional Alignment (DA) and Correlation (Corr). Both metrics operate on displacement maps of shape (T,H,W, 2), where T is the number of frame pairs used for displacement estima- tion, (H,W) is the latent spatial resolution, and the final dimension corresponds to the (x,y) flow components. 20S. Jung et al. Algorithm 1 HALO inference process Input: Reference video x ref , text prompt P, DiT model ε θ , noise scheduler N, encoder E decoder G Output: Generated video x 0 Note: Semantic extractor DIFT, Attention head h T ,S : Motion Heads, Structure Heads 1: Extract correspondence map: C ←− DIFT ▷cosine similarity 2: Compute attention: Q,K←− ε θ (z ref ,∅, 0) ▷z ref =E(x ref ) 3: Head Classification for h∈T▷Temporal or Spatial head 4: for each (i,j) where i,j ∈ [1,F] do 5: if h∈T then 6:Calculate cross-frame attention A i,j 7:Construct displacement matrix D i,j 8: end if 9: end for 10: Construct reference displacement: D ref ←− SCR(D i,j ,C i,j ) 11: Initialize z T ∼N(0,I) 12: for t = T to 0 do 13: if t > T opt then 14:for optimization step k = 0 to N opt do 15:Head Classification for h∈T 16:Entropy Calculation for h∈S 17:z ′ t ←−N(z ref ,t) ▷Noise Forward 18:if h∈S then▷Structure Heads 19:Store Value: V ′ ←− ε θ (z ′ t ,P,t) 20:Inject Value: V ′ ←− V 21:end if 22:if h∈T then ▷Motion Heads 23:for each (i,j) where i,j ∈ [1,F] do 24:Calculate cross-frame attention A i,j 25:Reweight ̃ A i,j : ̃ A i,j ←− SRW(A i,j ,C i,j ) 26:Construct displacement matrix D i,j 27:end for 28:L SM ←−∥D ref −D t ∥ 2 2 29:Update z t by minimizing L SM 30:end if 31:end for 32: end if 33: z t−1 = f(z t ,ε θ (z t ,P,t)) 34: end for 35: Return x 0 =G(z 0 ) Directional Alignment (DA). DA evaluates how well the predicted displace- ment directions align with the reference optical flow, independent of magnitude. Given the reference flow f ref and generated displacement f gen , DA is defined as: Controlling Motion Transfer21 Fig. 15: (a) Attention Entropy Distribution: Distribution of attention-map entropy values across samples. (b) Layer-wise Head Entropy Consistency: Distribution of head- level entropy values across layers, showing consistent patterns across different samples. Similarity Metric (SSIM): A heatmap is used to visualize pairwise similarity between samples. DA = mean D f ref ∥f ref ∥ , f gen ∥f gen ∥ E , (1) where ∥·∥ denotes the L2 norm and ⟨·,·⟩ is the dot product. The computation proceeds as follows: 1) Normalize f ref and f gen at each pixel. 2) Compute per- pixel cosine similarity between normalized vectors. 3) Average the directional similarity across all spatial locations and time steps. Correlation (Corr). Corr measures the linear relationship between the pre- dicted and reference flows, capturing both directional and magnitude consis- tency. It is computed using the Pearson correlation coefficient: Corr = Cov(f ref ,,f gen ) σ ref ,σ gen , (2) where Cov denotes covariance and σ represents the standard deviation. To com- pute Corr, we flatten each displacement field into a one-dimensional array of length T × H × W × 2 and apply the standard Pearson correlation formula. Whereas DA focuses solely on directional agreement, Corr reflects overall linear consistency, enabling complementary assessment of both direction- and magnitude-level motion alignment. This multi-dimensional analysis serves as a robust proof of concept for Motion-Specific head, proving that HALO achieves superior motion-aware guidance through a well-validated selection process. C.4 Entropy Validation in Structure-Specialized Heads Entropy Consistency Across Samples. We investigate the consistency of head entropy distributions across diverse video inputs. Fig. 15(a) shows the entropy patterns remain consistent across different samples, particularly within layers 14 to 27, which are utilized for structural feature injection. To quantify 22S. Jung et al. Fig. 16: Entropy distribution across all layers, computed by averaging the entropy of all attention heads within each layer. The orange dashed line denotes the median entropy value across layers, which is 6.78. A monkey is riding a bicycle on a trail next to a mural-covered wall in an urban park. ReferenceHALO(b)(a) Fig. 17: Qualitative results under different entropy thresholds. (a) Injecting features only from heads with entropy greater than 7. (b) Injecting features from all heads. this consistency, we compute similarity metrics using heat maps that visualize entropy values across the entire benchmark (Fig. 15(b)). The resulting SSIM score of 0.988 confirms a highly stable and invariant entropy pattern regardless of the input. This high degree of similarity indicates that the structural information encoded by the low-entropy heads is inherently consistent. Entropy Threshold. For structural-head selection, we set the entropy thresh- old to τ < 7. This choice is motivated by our empirical observation that the median entropy of attention maps across layers is approximately 7, as shown in Fig.16. Intuitively, lower-entropy heads exhibit more concentrated attention dis- tributions, which aligns with spatially localized, structure-preserving behavior. To derive this distribution, we compute the entropy of attention maps across all 90 samples, 42 layers, and 50 denoising steps. Fig.16 reports the mean entropy aggregated over all samples, layers, and steps, yielding a median value of roughly 7. This provides a principled basis for setting the threshold at τ < 7. Controlling Motion Transfer23 Fig. 18: Qualitative results and reference displacement maps D ref across various at- tention head configurations. We evaluate various entropy ranges to confirm that our selected entropy range, τ, consistently identifies structural-specialized heads. Fig. 17(a), high- entropy heads (τ > 7) produce diffuse attention patterns, resulting in layouts misaligned with the reference. Meanwhile, injecting all heads (b) causes visual artifacts due to the accumulation of disparate spatial features. These results highlight that selective structural-head injection is essential for providing clean structural guidance, allowing HALO to maintain a coherent spatial layout with- out artifacts. C.5 Validation of Motion-Specific Heads Fig. 18 illustrates how the head analysis translates to the motion transfer results. Motion-Specific Head demonstrates results that faithfully track the motion of the reference video while effectively capturing the movement of the car in the refer- ence displacement map. In contrast, All-Head and Spatial Head fail to properly capture the car in the displacement map, leading to misaligned motion in the generated videos. Consequently, the results validate that Motion-Specific Head plays a crucial role in effectively processing motion information. D Application to the Video DiT Model: Wan To demonstrate the generalizability of HALO, we extend it to Wan [40] video DiT model. Notably, we utilize the identical hyperparameter configurations as those used for CogVideoX [47], achieving consistent performance without any model-specific tuning. Unlike CogVideoX [47], Wan employs a dedicated self- attention module that exclusively processes visual tokens. Accordingly, our anal- ysis is conducted on this specialized attention mechanism. Head Analysis in Wan. Using the same procedure as in the main model, we identify both Motion-Specific and Structure-Specialized heads within Wan. To analyze motion behavior, we compute displacement maps separately for the tem- poral and spatial heads. As shown in Fig. 19, temporal heads more accurately 24S. Jung et al. Fig. 19: Displacement-field comparison between temporal and spatial heads in Wan. Temporal heads better capture reference motion, showing clearer separation of back- ground and object motion. Fig. 20: Quantitative comparison of each head in Wan using Directional Alignment (DA) and Correlation (Corr) with ground- truth optical flow. Fig. 21: Relationship between attention- map entropy and spatial entropy in Wan, demonstrating that low-entropy heads correspond to structure-specialized heads. capture the reference motion than spatial heads, exhibiting clear distinctions between background motion (red) and object motion (blue). We further com- pute Correlation (Corr) and Directional Alignment (DA) between RAFT-based optical flow [39] and the displacement maps of each head. Fig. 20 shows that tem- poral heads consistently achieve higher Corr and DA scores than spatial heads, confirming that temporal heads in Wan correspond to motion-specific heads. We also compute attention-map entropy [32] and compare it with the spa- tial entropy [5, 34] derived from corresponding attention features. As shown in Fig. 21, attention-map entropy is strongly correlated with spatial entropy, con- sistent with the findings in the main paper. This confirms that low-entropy heads in Wan encode structural information. Together, these results demonstrate that our analysis methodology generalizes effectively across different video DiT ar- chitectures. Experimental Results. We conduct experiments on the same dataset [36] Controlling Motion Transfer25 Fig. 22: Motion transfer results of HALO applied to Wan. Fig. 23: Overview of our movie scene dataset: video samples and corresponding prompts. Table 11: Quantitative results of Wan. ModelCLIP ↑ TC ↑ MF ↑ FTD ↓ WAN + HALO 30.2 90.1 71.6 18.3 used for the primary evaluations in the main paper. For faster experimen- tation, we generate 21-frame videos using 25-frame reference sequences. Table 11 shows that when applied to Wan, HALO produces videos that accurately reflect both the target prompt and the reference motion, achieving comparable motion fidelity to the results demon- strated in the main model. Qualitative motion transfer examples are shown in Fig. 22. Cross-Architecture Robustness. While the main paper primarily reports results using CogVideoX [47], HALO is not tied to this specific architecture. To examine whether our entropy-based head selection depends on a particu- lar backbone, we compare attention-entropy statistics between CogVideoX [47] and Wan [40]. We observe highly aligned timestep-wise entropy trends (r=0.97) and strong agreement in mean and median entropy (0.92/0.96), indicating that low-entropy structural heads emerge consistently across both video DiTs. This 26S. Jung et al. supports entropy-based head selection as a backbone-robust criterion rather than architecture-specific tuning. E Movie Scene Dataset This section discusses practical applications of motion transfer in Sec. E.1, pro- vides the details of the movie-scene dataset in Sec. E.2, and presents additional results of both our method and baseline models in Sec. E.3. E.1 Application Potential of Motion Transfer Existing visual effects (VFX) and computer graphics (CG) productions involve time-consuming and labor-intensive pipelines, requiring specialized skills across stages, such as modeling, animation, and rendering [11]. With recent advances in video generation accelerating content creation [26, 30], motion transfer has emerged as a practical, controllable video generation with direct utility in film- making and game production. The practical demand for such flexible content creation necessitates specialized datasets tailored to these domains. To address this need, we introduce a movie scene dataset that serves as a foundational resource for advancing and benchmarking motion transfer in professional pro- duction contexts. E.2 Movie Scene Dataset Details To evaluate our method in more realistic scenarios, we construct a movie scene dataset reflecting real-world video production environments, as shown in Fig. 23. While existing benchmarks [27,36] already utilize real videos, we incorporate this dataset to further assess the generalizability and robustness of our approach in production-level settings. The dataset consists of 20 videos, each paired with five different text prompts, totaling 100 samples, and the prompts describe both the primary objects and their backgrounds. We generate prompts using GPT-4o [20] to cover diverse cinematic scenarios. Moreover, the dataset involves complex mo- tion patterns that go beyond simple movements, such as sword fighting, horse- back riding, and walking in crowds. Such variety provides a robust testbed for verifying the model’s ability to faithfully transfer motion in dynamic real-world environments. E.3 Movie Scene Qualitative Results We present more qualitative comparisons on the movie scene dataset in Fig. 24, evaluating multiple motion transfer methods. HALO achieves the most consis- tent motion transfer, exhibiting high motion fidelity and coherent structural alignment. ConMO [13] is excluded from this comparison because it requires an input mask for generation, which is not available for our movie scene dataset. Controlling Motion Transfer27 MOI_sc: 8, 14, 24 MOI_sc 2: 0, 18, 32 ReferenceHALOMotionCloneDiTFlowGWTF Twovintage cars cruise smoothly across a wide sandy field in warm late-afternoon. A stormtrooper pulling a Santa sleigh through a snowy night city, neon reflections and drifting snow. Toy Story’s Woody gallops forward on horseback across a warm sunlit terrain with natural dust. RoPECraft Fig. 24: Qualitative comparison across diverse motion transfer tasks with the movie scene dataset. HALO is evaluated against U-Net- and DiT-based baselines across mul- tiple subjects and motions. F Additional Experiment Results F.1 More Qualitative Results We provide additional qualitative results in Fig. 25 across both the motion trans- fer benchmark (top) and our movie-scene dataset (bottom). Extended qualita- tive results in Fig. 26 further show the robustness of HALO across varying motion scales, from routine activities (e.g., walking, bus motion) to challenging high-dynamic scenes (e.g., horse jumping). The results highlight our method’s superior ability to maintain motion fidelity even as the complexity of the scene increases. 28S. Jung et al. Fig. 25: Qualitative results of HALO on benchmark and movie-scene datasets. The top rows illustrate performance on the motion transfer benchmark, while the bottom rows show generalization to movie scene videos. Controlling Motion Transfer29 Fig. 26: Qualitative comparison across diverse motion transfer tasks. HALO is evalu- ated against U-Net- and DiT-based baselines across multiple subjects and motions. F.2 Quantitative Trade-off and metric Analysis CLIP vs. MF trade-off Given that motion transfer necessitates a delicate bal- ance between preserving target semantics and following reference motion, CLIP and MF serve as complementary metrics that should be interpreted jointly. While baseline methods exhibit a pronounced CLIP–MF trade-off–consistent with prior studies [36,48]–HALO achieves a superior balance, attaining high scores in both metrics simultaneously (Fig. 27). These results suggest that our head-analysis approach effectively transfers motion while reducing irrelevant semantic inter- ference. TC vs. Motion Dynamics analysis We further analyze temporal consis- tency by measuring the average optical-flow magnitude across generated videos to quantify motion dynamics. As shown in Table 12, methods with higher au- tomated TC scores can exhibit lower motion dynamics, suggesting that TC may favor relatively static outputs. Therefore, TC should be interpreted jointly with reference-motion alignment metrics such as MF. Although HALO shows a slightly lower TC score than DiTFlow [35] and GWTF [7], it achieves sub- stantially higher MF in the main quantitative results. Moreover, the user study 30S. Jung et al. Fig. 27: Objective comparison across motion-transfer methods in terms of Motion Fidelity (MF) and CLIP Score (CLIP). Baseline methods exhibit a trade-off between MF and CLIP, whereas HALO (Ours) achieves balanced performance. ranks HALO highest in perceived temporal consistency, indicating that the lower automated TC score does not necessarily imply weaker temporal coherence. Table 12: Temporal consistency and motion dynamics analysis. HALO GWTF DiTFlow Motion Dynamics ↑5.794.815.19 Temporal Consistency ↑ 87.588.289.5 F.3 Computational cost Table 13: Runtime and Memory Re- port. Base +Sem. +Inj. HALO Time(s)1076 1088 1115 1125 Memory(GB) 38.1 38.2 38.2 38.2 Table 14: Comparison with Baselines DiTFlow RoPECraft GWTF HALO Time(s)10768187 456+α 1125 Memory(GB) 38.135.7 14.8+α 38.2 We report the runtime and memory efficiency of HALO on a single NVIDIA RTX A6000. As shown in Table 13, incorporating semantic correspondence and head-wise injection incurs only a minor runtime overhead; compared to the base model, HALO adds just 49s in inference time and 0.1GB in peak memory. We further compare HALO’s computational cost against other DiT-based meth- ods in Table 14. A direct comparison with GWTF is challenging because its training requirements introduce significant additional overhead α beyond simple inference. HALO demonstrates computational parity with leading training-free Controlling Motion Transfer31 Fig. 28: Failure cases of our method. In these examples, the output closely follows the reference motion but exhibits unnatural articulated movements (e.g., a penguin walk- ing like a bear), which occurs when semantic correspondence over-aligns structurally incompatible articulations. methods. Specifically, our approach remains highly competitive with DiTFlow and RopeCraft, achieving faithful motion transfer without sacrificing practical inference scalability. F.4 Limitations In some cases, the generated motion closely follows the reference yet becomes unnatural (e.g., a penguin walking like a bear), as illustrated in Fig. 28. This failure arises when semantic correspondence over-aligns structurally incompat- ible articulations. Because our semantic correspondence module refines motion representations via fine-grained correspondences, it can occasionally enforce near one-to-one part alignment, leading to implausible kinematics. Interestingly, these cases can exhibit higher motion fidelity scores because they more strictly track the reference dynamics. This reveals an inherent trade-off between reference- motion following and kinematic plausibility, which is not fully captured by ex- isting evaluations. Fine-Grained Local Motion We further analyze HALO under fine-grained lo- cal motion, such as facial expressions. As shown in Fig. 29, HALO preserves the overall head pose and coarse facial motion of the reference, showing comparable motion alignment to existing methods. However, subtle non-rigid deformations, such as detailed mouth shapes and fine facial expressions, remain challenging. This limitation arises because HALO represents motion using patch-level dis- placements, which are effective for object-level and medium-scale motion but less precise for very local deformation. F.5 Evaluation of Structure-Specialized Heads for Video Editing As shown in the main paper, we further extend our method to video editing. For quantitative evaluation, we report CLIP score (CLIP), Temporal Consis- tency (TC), Motion Fidelity (MF), masked PSNR (m.P), and masked LPIPS (m.L) [43]. To compute m.P and m.L, we use the provided segmentation masks and evaluate reconstruction quality on the background region, which is expected 32S. Jung et al. A mime artist screaming silently on a black stage. Reference Ours Fig. 29: Fine-grained local motion transfer. HALO preserves coarse facial motion and head pose, but subtle non-rigid deformations, such as detailed facial expressions, remain challenging due to the patch-level displacement representation. Reference cow → rhino cow→ tiger Reference flamingo → crane flamingo → goose Fig. 30: Application to video editing using the structure-specialized heads. to remain unchanged from the reference. As shown in Table 15, HALO achieves strong editing quality while preserving background structure, supporting the effectiveness of structure-specialized heads beyond motion transfer. Additional qualitative results are provided in Fig. 30. Table 15: Video editing results. CLIP ↑ TC ↑ MF ↑ m.P ↑ m.L ↓ HALO 29.3 93.1 81.9 17.20.4 F.6 Structural Consistency in Long Videos To validate structural consistency in long videos, we evaluate HALO on 112- frame sequences. This setting is substantially longer than the standard gen- eration length used in the main experiments, allowing us to examine whether the selected structure-specialized heads remain stable over extended temporal horizons. As shown in Fig. 31, low-entropy heads preserve consistent structural attention patterns across the sequence, without noticeable drift or degradation Controlling Motion Transfer33 A truck is driving through the mountains. Reference Ours Features Fig. 31: Structural consistency in long-video generation. We evaluate HALO on 112- frame sequences and visualize the behavior of selected low-entropy heads. The selected heads maintain stable structural attention patterns over extended temporal horizons, supporting long-sequence structural preservation. over time. This indicates that entropy-based head selection is not limited to short clips but provides stable structural cues for long-horizon motion transfer. F.7 Scalability of SCR SCR operates entirely in the latent space, making its computational overhead moderate even at higher output resolutions. As shown in Table 16, increasing the resolution from 720p to 1080p increases the SCR runtime only from 13.2s to 19.1s, while the total inference time increases from 1125s to 1404s. The memory usage of SCR also remains nearly unchanged across the two resolutions. These results indicate that SCR does not introduce a major bottleneck and remains practical for high-resolution motion transfer. Table 16: Scalability analysis of SCR at higher resolutions. HALO 720p SCR HALO 1080p SCR Time (s)112513.2140419.1 Mem. (GB)38.218.238.218.4 G User Study For human evaluation, we conduct a comprehensive user study involving 20 participants with expertise in computer vision. The evaluation is performed on a set of 10 representative video samples. Participants are asked to assess the results based on three key criteria: editing accuracy, temporal consistency, and motion fidelity. The specific evaluation metrics are defined as follows: 34S. Jung et al. – Edit Accuracy: How well does the video content align with the provided text prompt? (e.g., whether the characters, actions, and attributes described in the prompt are accurately reflected in the visual output). – Temporal Consistency: Do the motion, appearance, and background el- ements transition naturally throughout the video? (e.g., checking for the absence of flickers or abrupt changes in character identity or background details). – Motion Fidelity: How closely does the overall motion in the generated video resemble the reference motion? (e.g., whether the pose, flow of action, speed, and timing align with the source motion). Controlling Motion Transfer35 References 1. Ahn, D., Kang, J., Lee, S., Kim, M., Min, J., Jang, W., Lee, S., Paul, S., Hong, S., Kim, S.: Fine-grained perturbation guidance via attention head selection. arXiv e-prints p. arXiv–2506 (2025) 2. Atzmon, Y., Gal, R., Tewel, Y., Kasten, Y., Chechik, G.: Motion by queries: Identity-motion trade-offs in text-to-video generation. arXiv preprint arXiv:2412.07750 (2024) 3. Avrahami, O., Patashnik, O., Fried, O., Nemchinov, E., Aberman, K., Lischinski, D., Cohen-Or, D.: Stable flow: Vital layers for training-free image editing. In: CVPR. p. 7877–7888 (2025) 4. Bar-Tal, O., Chefer, H., Tov, O., Herrmann, C., Paiss, R., Zada, S., Ephrat, A., Hur, J., Liu, G., Raj, A., et al.: Lumiere: A space-time diffusion model for video generation. In: SIGGRAPH Asia 2024 Conference Papers. p. 1–11 (2024) 5. Batty, M.: Spatial entropy. Geographical analysis 6(1), 1–31 (1974) 6. Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al.: Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127 (2023) 7. Burgert, R., Xu, Y., Xian, W., Pilarski, O., Clausen, P., He, M., Ma, L., Deng, Y., Li, L., Mousavi, M., et al.: Go-with-the-flow: Motion-controllable video diffusion models using real-time warped noise. In: CVPR. p. 13–23 (2025) 8. Cai, M., Cun, X., Li, X., Liu, W., Zhang, Z., Zhang, Y., Shan, Y., Yue, X.: Ditctrl: Exploring attention control in multi-modal diffusion transformer for tuning-free multi-prompt longer video generation. In: CVPR. p. 7763–7772 (2025) 9. Chen, H., Zhang, Y., Cun, X., Xia, M., Wang, X., Weng, C., Shan, Y.: Videocrafter2: Overcoming data limitations for high-quality video diffusion models. In: CVPR. p. 7310–7320 (2024) 10. Chen, S., Xu, M., Ren, J., Cong, Y., He, S., Xie, Y., Sinha, A., Luo, P., Xiang, T., Perez-Rua, J.M.: Gentron: Diffusion transformers for image and video generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 6441–6451 (2024) 11. Du, T., Wu, K., Ma, P., Wah, S., Spielberg, A., Rus, D., Matusik, W.: Diffpd: Differentiable projective dynamics. ACM Transactions on Graphics (ToG) 41(2), 1–21 (2021) 12. Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al.: Scaling rectified flow transformers for high-resolution image synthesis. In: ICML (2024) 13. Gao, J., Yin, Z., Hua, C., Peng, Y., Liang, K., Ma, Z., Guo, J., Liu, Y.: Conmo: Con- trollable motion disentanglement and recomposition for zero-shot motion transfer. In: CVPR. p. 7191–7200 (2025) 14. Gokmen, A.B., Ekin, Y., Bilecen, B.B., Dundar, A.: Ropecraft: Training-free mo- tion transfer with trajectory-guided rope optimization on diffusion transformers. arXiv preprint arXiv:2505.13344 (2025) 15. Guo, Y., Yang, C., Rao, A., Liang, Z., Wang, Y., Qiao, Y., Agrawala, M., Lin, D., Dai, B.: Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. In: ICLR (2024), https://openreview.net/forum?id= Fx2SbBgcte 16. HaCohen, Y., Chiprut, N., Brazowski, B., Shalem, D., Moshe, D., Richardson, E., Levin, E., Shiran, G., Zabari, N., Gordon, O., et al.: Ltx-video: Realtime video latent diffusion. arXiv preprint arXiv:2501.00103 (2024) 36S. Jung et al. 17. Hedlin, E., Sharma, G., Mahajan, S., Isack, H., Kar, A., Tagliasacchi, A., Yi, K.M.: Unsupervised semantic correspondence using stable diffusion. NeurIPS 36, 8266– 8279 (2023) 18. Hessel, J., Holtzman, A., Forbes, M., Le Bras, R., Choi, Y.: Clipscore: A reference- free evaluation metric for image captioning. In: Proceedings of the 2021 conference on empirical methods in natural language processing. p. 7514–7528 (2021) 19. Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., Wang, Y., Chen, X., Wang, L., Lin, D., Qiao, Y., Liu, Z.: VBench: Comprehensive benchmark suite for video generative models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2024) 20. Hurst, A., Lerer, A., Goucher, A.P., Perelman, A., Ramesh, A., Clark, A., Os- trow, A., Welihinda, A., Hayes, A., Radford, A., et al.: Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024) 21. Jeong, H., Park, G.Y., Ye, J.C.: Vmc: Video motion customization using temporal attention adaption for text-to-video diffusion models. In: CVPR (2023) 22. Kang, S., Kim, J., Kim, J., Hwang, S.J.: Your large vision-language model only needs a few attention heads for visual grounding. In: CVPR. p. 9339–9350 (2025) 23. Khachatryan, L., Movsisyan, A., Tadevosyan, V., Henschel, R., Wang, Z., Navasardyan, S., Shi, H.: Text2video-zero: Text-to-image diffusion models are zero- shot video generators. In: ICCV. p. 15954–15964 (2023) 24. Kim, C., Shin, H., Hong, E., Yoon, H., Arnab, A., Seo, P.H., Hong, S., Kim, S.: Seg4diff: Unveiling open-vocabulary segmentation in text-to-image diffusion transformers. arXiv preprint arXiv:2509.18096 (2025) 25. Kong, W., Tian, Q., Zhang, Z., Min, R., Dai, Z., Zhou, J., Xiong, J., Li, X., Wu, B., Zhang, J., et al.: Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603 (2024) 26. Li, B., Zhang, Y., Wang, Q., Ma, L., Shi, X., Wang, X., Wan, P., Yin, Z., Zhuge, Y., Lu, H., Jia, X.: Vfxmaster: Unlocking dynamic visual effect generation via in-context learning (2025), https://arxiv.org/abs/2510.25772 27. Ling, P., Bu, J., Zhang, P., Dong, X., Zang, Y., Wu, T., Chen, H., Wang, J., Jin, Y.: Motionclone: Training-free motion cloning for controllable video generation. arXiv preprint arXiv:2406.05338 (2024) 28. Liu, B., Wang, C., Su, T., Ten, H., Huang, J., Guo, K., Jia, K.: Understanding attention mechanism in video diffusion models. arXiv preprint arXiv:2504.12027 (2025) 29. Liu, Y., Zhang, K., Li, Y., Yan, Z., Gao, C., Chen, R., Yuan, Z., Huang, Y., Sun, H., Gao, J., et al.: Sora: A review on background, technology, limitations, and opportunities of large vision models. arXiv preprint arXiv:2402.17177 (2024) 30. Ma, Y., Feng, K., Hu, Z., Wang, X., Wang, Y., Zheng, M., He, X., Zhu, C., Liu, H., He, Y., et al.: Controllable video generation: A survey. arXiv preprint arXiv:2507.16869 (2025) 31. Nam, J., Son, S., Chung, D., Kim, J., Jin, S., Hur, J., Kim, S.: Emer- gent temporal correspondences from video diffusion transformers. arXiv preprint arXiv:2506.17220 (2025) 32. Pardyl, A., Rypeść, G., Kurzejamski, G., Zieliński, B., Trzciński, T.: Active vi- sual exploration based on attention-map entropy. arXiv preprint arXiv:2303.06457 (2023) 33. Park, G.Y., Jeong, H., Lee, S.W., Ye, J.C.: Spectral motion alignment for video motion transfer using diffusion models. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 39, p. 6398–6405 (2025) Controlling Motion Transfer37 34. Peruzzo, E., Sangineto, E., Liu, Y., De Nadai, M., Bi, W., Lepri, B., Sebe, N.: Spatial entropy as an inductive bias for vision transformers. Machine Learning 113(9), 6945–6975 (2024) 35. Pondaven, A., Siarohin, A., Tulyakov, S., Torr, P., Pizzati, F.: Video motion trans- fer with diffusion transformers. In: CVPR. p. 22911–22921 (2025) 36. Shi, Q., Wu, J., Bai, J., Zhang, J., Qi, L., Tong, Y., Li, X.: Decouple and track: Benchmarking and improving video diffusion transformers for motion transfer. In: ICCV. p. 10995–11005 (2025) 37. Siarohin, A., Lathuilière, S., Tulyakov, S., Ricci, E., Sebe, N.: Animating arbitrary objects via deep motion transfer. In: CVPR. p. 2377–2386 (2019) 38. Tang, L., Jia, M., Wang, Q., Phoo, C.P., Hariharan, B.: Emergent correspondence from image diffusion. NeurIPS 36, 1363–1389 (2023) 39. Teed, Z., Deng, J.: Raft: Recurrent all-pairs field transforms for optical flow. In: ECCV. p. 402–419. Springer (2020) 40. Wan, T., Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.W., Chen, D., Yu, F., Zhao, H., Yang, J., Zeng, J., Wang, J., Zhang, J., Zhou, J., Wang, J., Chen, J., Zhu, K., Zhao, K., Yan, K., Huang, L., Feng, M., Zhang, N., Li, P., Wu, P., Chu, R., Feng, R., Zhang, S., Sun, S., Fang, T., Wang, T., Gui, T., Weng, T., Shen, T., Lin, W., Wang, W., Wang, W., Zhou, W., Wang, W., Shen, W., Yu, W., Shi, X., Huang, X., Xu, X., Kou, Y., Lv, Y., Li, Y., Liu, Y., Wang, Y., Zhang, Y., Huang, Y., Li, Y., Wu, Y., Liu, Y., Pan, Y., Zheng, Y., Hong, Y., Shi, Y., Feng, Y., Jiang, Z., Han, Z., Wu, Z.F., Liu, Z.: Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314 (2025) 41. Wang, J., Ma, Y., Guo, J., Xiao, Y., Huang, G., Li, X.: Cove: Unleashing the diffusion feature correspondence for consistent video editing. NeurIPS 37, 96541– 96565 (2024) 42. Wang, L., Mai, Z., Shen, G., Liang, Y., Tao, X., Wan, P., Zhang, D., Li, Y., Chen, Y.C.: Motion inversion for video customization. In: Proceedings of the Spe- cial Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers. p. 1–12 (2025) 43. Wang, Y., Wang, L., Ma, Z., Hu, Q., Xu, K., Guo, Y.: Videodirector: Precise video editing via text-to-video models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 2589–2598 (2025) 44. Xi, H., Yang, S., Zhao, Y., Xu, C., Li, M., Li, X., Lin, Y., Cai, H., Zhang, J., Li, D., et al.: Sparse videogen: Accelerating video diffusion transformers with spatial- temporal sparsity. arXiv preprint arXiv:2502.01776 (2025) 45. Xiao, Z., Zhou, Y., Yang, S., Pan, X.: Video diffusion models are training-free motion interpreter and controller. NeurIPS 37, 76115–76138 (2024) 46. Xing, J., Xia, M., Liu, Y., Zhang, Y., Zhang, Y., He, Y., Liu, H., Chen, H., Cun, X., Wang, X., et al.: Make-your-video: Customized video generation using textual and structural guidance. IEEE Transactions on Visualization and Computer Graphics 31(2), 1526–1541 (2024) 47. Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., Yin, D., Yuxuan.Zhang, Wang, W., Cheng, Y., Xu, B., Gu, X., Dong, Y., Tang, J.: Cogvideox: Text-to-video diffusion models with an expert transformer. In: ICLR (2025), https://openreview.net/forum?id=LQzN6TRFg9 48. Yatim, D., Fridman, R., Bar-Tal, O., Kasten, Y., Dekel, T.: Space-time diffusion features for zero-shot text-driven motion transfer. In: CVPR. p. 8466–8476 (2024) 49. Zhang, H., Li, F., Liu, S., Zhang, L., Su, H., Zhu, J., Ni, L., Shum, H.Y.: Dino: Detr with improved denoising anchor boxes for end-to-end object detection. In: The Eleventh International Conference on Learning Representations (2023) 38S. Jung et al. 50. Zhao, R., Gu, Y., Wu, J.Z., Zhang, D.J., Liu, J.W., Wu, W., Keppo, J., Shou, M.Z.: Motiondirector: Motion customization of text-to-video diffusion models. In: ECCV. p. 273–290. Springer (2024)