Paper deep dive
RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation
Feng Jiang, Yang Chen, Kyle Xu, Yuchen Liu, Haifeng Wang, Zhenhao Shen, Jasper Lu, Shengze Huang, Yuanfei Wang, Chen Xie, Ruihai Wu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/26/2026, 10:48:17 PM
Summary
RoboWM-Bench is a manipulation-centric benchmark designed to evaluate the physical executability of video world models in robotic manipulation. Unlike existing benchmarks that focus on visual realism, RoboWM-Bench assesses whether predicted behaviors (from both human-hand and robotic videos) can be translated into executable action sequences and successfully performed in a high-fidelity 'real-to-sim' environment. The benchmark utilizes a unified pipeline involving human-centric retargeting (using 3D hand pose estimation) and robot-centric inverse dynamics modeling (IDM) to validate predicted interactions across diverse tasks, including rigid, articulated, and deformable object manipulation.
Entities (6)
Relation Signals (4)
RoboWM-Bench → builton → LeHome
confidence 100% · RoboWM-Bench is built upon the LeHome simulation framework [30]
Inverse Dynamics Model → convertsto → Embodied Action Sequences
confidence 100% · The predicted behaviors are then converted into embodied action sequences through inverse dynamics modeling
Real-to-Sim → enables → Reproducible Evaluation
confidence 100% · To enable fair and accessible evaluation across real-world scenarios, we adopt a real-to-simulation (real-to-sim) framework
RoboWM-Bench → evaluates → Video World Models
confidence 100% · RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent advances in large-scale video world models have enabled increasingly realistic future prediction, raising the prospect of leveraging imagined videos for robot learning. However, visual realism does not imply physical plausibility, and behaviors inferred from generated videos may violate dynamics and fail when executed by embodied agents. Existing benchmarks begin to incorporate notions of physical plausibility, but they largely remain perception- or diagnostic-oriented and do not systematically evaluate whether predicted behaviors can be translated into executable actions that complete the intended task. To address this gap, we introduce RoboWM-Bench, a manipulation-centric benchmark for embodiment-grounded evaluation of video world models. RoboWM-Bench converts generated behaviors from both human-hand and robotic manipulation videos into embodied action sequences and validates them through robotic execution. The benchmark spans diverse manipulation scenarios and establishes a unified protocol for consistent and reproducible evaluation. Using RoboWM-Bench, we evaluate state-of-the-art video world models and find that reliably generating physically executable behaviors remains an open challenge. Common failure modes include errors in spatial reasoning, unstable contact prediction, and non-physical deformations. While finetuning on manipulation data yields improvements, physical inconsistencies still persist, suggesting opportunities for more physically grounded video generation for robots.
Tags
Links
- Source: https://arxiv.org/abs/2604.19092v1
- Canonical: https://arxiv.org/abs/2604.19092v1
Trouble viewing inline? Open PDF directly →
Full Text
70,642 characters extracted from source content.
Expand or collapse full text
RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation Feng Jiang 1,∗ , Yang Chen 1,∗ , Kyle Xu 1,∗ , Yuchen Liu 1 , Haifeng Wang 1 , Zhenhao Shen 1 , Jasper Lu 1 , Shengze Huang 2 , Yuanfei Wang 1 , Chen Xie 3 , and Ruihai Wu 1,† 1 Peking University 2 Tsinghua University 3 Lightwheel wuruihai@pku.edu.cn Webpage: https://robowm-bench.github.io/RoboWM-Bench/ * Equal contribution,† Corresponding author Abstract. Recent advances in large-scale video world models have en- abled increasingly realistic future prediction, raising the prospect of lever- aging imagined videos for robot learning. However, visual realism does not imply physical plausibility, and behaviors inferred from generated videos may violate dynamics and fail when executed by embodied agents. Existing benchmarks begin to incorporate notions of physical plausibil- ity, but they largely remain perception- or diagnostic-oriented and do not systematically evaluate whether predicted behaviors can be trans- lated into executable actions that complete the intended task. To ad- dress this gap, we introduce RoboWM-Bench, a manipulation-centric benchmark for embodiment-grounded evaluation of video world models. RoboWM-Bench converts generated behaviors from both human-hand and robotic manipulation videos into embodied action sequences and validates them through robotic execution. The benchmark spans diverse manipulation scenarios and establishes a unified protocol for consistent and reproducible evaluation. Using RoboWM-Bench, we evaluate state- of-the-art video world models and find that reliably generating physi- cally executable behaviors remains an open challenge. Common failure modes include errors in spatial reasoning, unstable contact prediction, and non-physical deformations. While finetuning on manipulation data yields improvements, physical inconsistencies still persist, suggesting op- portunities for more physically grounded video generation for robots. Keywords: World models· Robotic manipulation· Benchmarking 1 Introduction Recent advances in large-scale video generation and prediction have led to in- creasingly realistic video world models capable of synthesizing temporally co- herent and visually plausible futures [8, 16, 17, 49], opening new opportunities arXiv:2604.19092v1 [cs.RO] 21 Apr 2026 VEO, WAN, Cosmos ... Predicted Human Video Predicted Robot Video Retargeting Real2Sim IDM Retargeted Human Actions IDM Robot Actions execute execute RoboWM-Bench Initial Image Task: Throw the trash in the can... Initial Image World Models (b) Tasks (a) Framework Input Video World Models Physical Embodied Validation Task: Pick up the cup and pour the water... (c) Performance Fold towel Stack cups Put banana on plate Fold clothes Make hamburger Discard trash Put banana in drawer Push tape Cut sausage Turn faucet Put stapler in box Robot (Real2Sim) Human Robot Robot (Sim) Human Put cup on plate Fig. 1: Overview of RoboWM-Bench. RoboWM-Bench is a manipulation-centric benchmark for evaluating video world models under embodied execution. (a) Given an initial scene observation and task description, world models generate manipulation videos with human hands or robot arms. The predicted behaviors are converted into embodied action sequences and validated in simulation through real-to-sim scene re- construction. (b) RoboWM-Bench spans a diverse suite of manipulation tasks with varying interaction dynamics, object properties, and temporal horizons. (c) Perfor- mance of state-of-the-art video world models on RoboWM-Bench. for robot learning from predicted videos. However, despite their impressive vi- sual fidelity, such models may generate behaviors that violate physical consis- tency when grounded in embodiment and real-world dynamics. Motivated by this challenge, a growing body of work has begun to tailor world models to robotic manipulation scenarios, seeking to better capture active robot–environment in- teractions and embodiment-aware dynamics [1, 2, 9, 12, 21, 46]. As both visual realism and physical consistency continue to improve, imagined manipulation videos are increasingly viewed as a scalable source of supervision for robot learn- ing [9, 23, 45]. In this context, reliable evaluation of predicted video quality is essential to accurately assess world model capability, and to ensure that down- stream robotic policies learned from them are physically grounded and robust. Existing evaluation protocols for video world models primarily emphasize vi- sual fidelity, semantic consistency, and temporal coherence [15,22,24,32,33,59]. However, visual realism alone does not guarantee physical validity. More recent benchmarks have introduced metrics aimed at capturing physical plausibility, revealing that state-of-the-art models often fail to maintain coherent physical dynamics despite strong perceptual quality [43,60,61]. While these efforts repre- sent important progress, they remain largely perception- or diagnostic-oriented. To reliably assess predicted videos and to support scalable robot learning, eval- uation should extend beyond perceptual plausibility to consider whether the behaviors implied by predicted videos can be faithfully executed by embodied agents and successfully accomplish the intended tasks. A recent study [14] takes a further step by attempting to assess physical executability through real-world robot validation. Although this approach moves closer to embodiment-grounded evaluation, the comprehensive evaluation requires broader coverage and greater task diversity. Moreover, reliance on real-world testing makes large-scale, repro- ducible benchmarking challenging. In this paper, we introduce RoboWM-Bench, a systematic and repro- ducible manipulation-centric benchmark for embodiment-grounded evaluation of video world models. RoboWM-Bench operationalizes physical executability as a principled and measurable evaluation criterion, grounding assessment in verifiable control execution rather than perceptual judgment alone. Specifically, the benchmark evaluates whether interactions predicted in generated videos can be translated into executable action sequences and successfully performed in a physically grounded environment. As illustrated in Figure 1, given initial observations and task descriptions, video world models generate future manipulation videos involving either human hands or robot arms. The predicted behaviors are then converted into embodied action sequences through inverse dynamics modeling [5,23] for robotic videos, or pose tracking and retargeting [27, 28, 39] for human demonstrations. To enable fair and accessible evaluation across real-world scenarios, we adopt a real-to- simulation (real-to-sim) framework in which scenes and interaction dynamics are reconstructed in simulation to match their real-world counterparts [11,52]. The extracted actions are then executed within the reconstructed environments using visually and physically high-fidelity simulation, enabling standardized and repro- ducible validation of physical executability. RoboWM-Bench provides hierarchi- cal evaluation with diverse complexity, incorporating both step-level executabil- ity metrics and final task-level success rates, which together enable fine-grained diagnostic analysis as well as holistic performance assessment. RoboWM-Bench spans a broad spectrum of manipulation scenarios, including diverse object dy- namics, short- and long-horizon tasks, and both single-arm and bimanual inter- actions. Through this unified protocol, RoboWM-Bench enables consistent and reproducible comparison of video world models under embodiment constraints. We conduct extensive experiments using RoboWM-Bench to evaluate state- of-the-art video world models under embodied execution. The results suggest a noticeable gap between visual realism and physical executability. Execution success declines as task complexity increases, particularly for long-horizon in- teractions and deformable-object manipulation. Qualitative analysis further re- veals common inconsistencies in generated videos, including unrealistic object deformation and inaccurate contact prediction, which may lead to dynamically infeasible actions during execution. While fine-tuning on manipulation-specific data improves executability, cer- tain physical inconsistencies remain. These observations suggest that ensuring physically consistent behavior under embodied interaction remains a challenging problem for current video world models, and highlight promising opportunities for future research toward more physically grounded and embodiment-aware world modeling. To summarize, our main contributions are as follows: – We introduce RoboWM-Bench, a manipulation-centric benchmark for embodiment-grounded evaluation of video world models, which assesses whether predicted videos can be translated into executable actions and successfully performed in physically grounded environments. – We develop a unified evaluation pipeline that converts predicted videos into embodied actions, and validates them through execution in real-to-sim re- constructed environments with visually and physically high-fidelity engines, enabling standardized and reproducible evaluation. – We conduct extensive experiments across diverse manipulation tasks to eval- uate the embodied executability of state-of-the-art video world models, re- vealing their challenges and potential for physically grounded robot learning. 2 Related Work 2.1 World Models for Robotics Recent advances in large-scale video generation have renewed interest in world models as predictive models of physical dynamics [1, 2, 19, 20, 47, 50, 53]. Mod- els such as Sora [8], Veo [17], Wan [49], and Seedance [16] demonstrate strong visual realism and temporal coherence, suggesting that internet-scale training can capture rich spatiotemporal priors. However, while these models can gen- erate visually plausible manipulation videos, it remains unclear whether they preserve physical consistency and interaction dynamics in a manner that sup- ports executable control. Motivated by this challenge, a growing body of work develops robotics-oriented world models that explicitly target embodied con- trol [3,57,62]. DreamGen [23] finetunes video world models to better learn the robot’s physical constraints and movement capabilities. Large Video Planner (LVP) [9] investigates video-conditioned planning, leveraging predicted visual rollouts as intermediate representations for downstream control. WoW [12] em- phasizes physically grounded intuition through large-scale embodied interaction data. EnerVerse [21] proposes an embodied future-space formulation for manip- ulation reasoning, while GigaWorld-0 [46] frames world models as scalable data engines for embodied AI. As world models evolve from perceptual generators into embodied simulators and planning engines, evaluation must extend beyond visual fidelity to verify physical consistency and control feasibility. 2.2 Learning Robotic Actions from Video Learning robotic control from video has attracted increasing attention as large- scale human video data and visual world models become widely available [6,18, 36,41,42,44,55]. R3M [37] and VIP [35] focus on learning visual representations for manipulation by pretraining vision encoders on diverse video data to improve downstream policy learning. Another line of work explores latent action model- ing, aiming to recover implicit control signals from videos without explicit robot action annotations, including latent action pretraining [58] and domain adapta- tion methods [40]. Other approaches explicitly model inverse dynamics to extract executable actions from predicted or observed videos [13, 23, 48]. For example, DreamGen [23] combines video world models with inverse dynamics to generate neural trajectories for policy learning. Complementary to these, human-to-robot retargeting frameworks such as Phantom [28] and Masquerade [27] leverage in- the-wild human videos to train policies without direct robot data. More recently, large-scale vision-language-action (VLA) models demonstrate emergent human- to-robot transfer capabilities through joint pretraining on multimodal data [25]. In parallel, large-scale world models further enable robot learning from world- model-generated videos [7,23,34,45]. As such generated videos increasingly serve as training or planning signals for robot policies, systematic evaluation of their physical validity and control compatibility becomes essential. 2.3 Evaluation and Benchmarks of World Models Existing benchmarks for video world models primarily emphasize perceptual realism and generative quality [15, 32]. Comprehensive evaluation suites such as VBench [22], EvalCrafter [33], and T2VEval [24] assess visual fidelity, tem- poral consistency, text–video alignment, and motion coherence. More recently, physical-AI-oriented benchmarks [29,43,59] such as PAI-Bench [61] and VBench- 2.0 [60] have begun to evaluate models from the perspective of physical reason- ing and embodied intelligence. While these efforts provide systematic and repro- ducible assessment of generative models, they largely focus on perceptual plausi- bility or high-level physical understanding, without explicitly evaluating whether predicted manipulation dynamics are executable under embodied robotic con- straints. Several robotics-oriented works have incorporated real-robot validation. For example, LVP [9] evaluates video-conditioned planning through real-world manipulation tasks, and “Wow, wo, val!” [14] introduces embodied evaluation protocols inspired by Turing-test-style assessments. These efforts represent im- portant steps toward physically grounded evaluation. However, their coverage and task diversity remain limited, and real-robot experiments are often costly and difficult to reproduce, making large-scale and accessible evaluation challeng- ing. In contrast, RoboWM-Bench establishes a manipulation-centric benchmark Real-to-Sim Pipeline � � � � �풉�捠,� � � �풏�,� 푮 풓 풂 풔 - 풂 풏 � � �-,� � � 풓�풏,� � � �풏㵀,� Human-Hand Retargeting Joint Space Gripper Pose Marble Predicted VideosEmbodied Validation SAM 3D Real World Real-to-Sim Engine Simulation Platform Inverse Dynamics Model World Models Fig. 2: Pipeline of RoboWM-Bench. Given an initial scene observation, the corre- sponding real-world scene is reconstructed in simulation through a real-to-sim pipeline, enabling consistent and reproducible evaluation. Predicted videos are then converted into executable robot actions through two pathways: human-centric retargeting, which estimates 3D hand poses and retargets them to robot end-effector actions, and robot- centric inverse dynamics, which predicts joint-space actions via an inverse dynamics model (IDM). The resulting actions are executed in simulation and evaluated using step-level checkers and final task success rates to measure embodied executability. with broader and standardized task suites, along with a unified and reproducible evaluation protocol spanning both simulation and real-scenario settings, enabling systematic assessment of the physical executability and action-level consistency of predicted manipulation videos. 3 RoboWM-Bench 3.1 Benchmark Overview We introduce RoboWM-Bench, a benchmark for evaluating video world models through embodiment-grounded validation. RoboWM-Bench treats physical exe- cutability as a principled and measurable criterion for generated videos, and sup- ports evaluation on both human-hand and robotic manipulation scenarios. Given an initial scene observation and a task description, a world model predicts a fu- ture manipulation video. The embodied behaviors depicted in the video are then converted into executable action sequences through video-to-action mappings (Section 3.2). To enable consistent and reproducible evaluation, the correspond- ing real-world scenario is reconstructed in simulation via a real-to-sim pipeline, where the extracted actions are executed (Section 3.3). RoboWM-Bench defines a suite of manipulation tasks spanning diverse object properties and interac- tion regimes (Section 3.4). Executability is determined by whether the executed actions accomplish the intended manipulation objective (Section 3.5). 3.2 Embodied Video-to-Action Execution RoboWM-Bench evaluates both human-hand and robotic manipulation videos, enabling comprehensive assessment across different interaction modalities. While current video world models often exhibit stronger performance in predicting human-hand interactions, robotic manipulation videos are more directly aligned with downstream robot policy learning and control execution. We next describe how actions are extracted from both human-hand and robotic videos. Human-Centric Retargeting. Inspired by prior work [27, 28], we estimate human hand poses from videos and retarget them to robot end-effector poses. We first reconstruct 3D hand poses with HaMeR [39], and use the recovered 3D keypoints as the basis for motion retargeting. The gripper target position is defined as the midpoint between the thumb and index fingertips. For orientation estimation, we observe that existing formulations [28] may yield unstable or tilted end-effector poses, as they rely on global finger configurations rather than contact-relevant geometry. To improve stability, after fitting a plane through the thumb and index finger keypoints, we project the thumb and index fingertips onto this plane, and define thex-axis as the line connecting their projections, while thez-axis remains the plane normal, consistent with previous formulations. The resulting end-effector pose better preserves the human–object interaction geometry. For gripper opening, rather than using the thumb–index distance as in prior work, we empirically adopt the minimum distance between the thumb tip and all other fingertips, which accounts for cases where the index fingertip is not the primary contact point. Finally, we apply trajectory smoothing and temporal denoising to stabilize the retargeted motion signals, following previous work [27,28]. Robot-Centric Execution. For robotic manipulation videos, we recover ac- tion sequences from predicted frames using an inverse dynamics model (IDM) [5, 23]. We adopt the IDM architecture from [23], which takes two consecutive image frames as input and predicts the intermediate action chunk in joint space. To pretrain the IDM, we collect large-scale simulation data by executing di- verse trajectories with a Franka arm in a physics simulator, recording paired RGB observations and corresponding joint-space actions. Compared with real- world recordings, simulation trajectories can be generated at higher temporal resolution, providing smoother motion supervision during pretraining and facil- itating stable inverse-dynamics learning. To mitigate the visual sim-to-real gap when applying the IDM to real-world scenarios and the model-generated videos, we adopt a background-masking strategy during simulation pretraining. Specif- ically, inspired by [46], background regions in simulation videos are removed, retaining only the robot arm to minimize domain discrepancies between simu- lated and real observations. Finally, we collect a small amount of real-world data from a physical Franka arm to finetune the IDM without background masking. This two-stage strategy, consisting of simulation pretraining followed by real-world finetuning, improves data efficiency and enhances sim-to-real generalization, enabling reliable action extraction for embodiment-grounded evaluation. 3.3 High-Fidelity Real-to-Sim Framework To ensure both accessibility and reproducibility, RoboWM-Bench conducts eval- uation entirely in open-source simulation environments. The benchmark includes both purely simulated tasks and real-to-sim tasks reconstructed from real-world scenes. This design enables faithful reproduction of predicted interactions with- out requiring physical robotic platforms or real-world environments, which are often costly and difficult to replicate. RoboWM-Bench is built upon the LeHome simulation framework [30], which provides physically realistic object dynamics and visually consistent scene ren- dering. These properties are critical for faithfully reconstructing real-world scenes in simulation and ensuring that reproduced interactions yield outcomes consis- tent with those observed in the real environment. For real-to-sim reconstruction, we adopt a modular pipeline to replicate real- world scenes within a high-fidelity simulation environment [30]. Inspired by re- cent work [52], the background scenes are reconstructed using 4D Gaussian rep- resentations to preserve visual realism and spatial consistency. For interactive ob- jects, rigid geometries are obtained via 3D segmentation and reconstruction [11], while articulated and deformable object pairs between real and simulated do- mains are acquired following [30]. Objects’ initial poses relative to the camera are estimated using pose estimation models [26,51], ensuring accurate initialization within the simulation environment, and the real-world camera pose is calibrated using FEEPE [54] with averaged results across multiple runs. This real-to-sim framework preserves physical structure and spatial config- uration while ensuring reproducibility and accessibility. Moreover, RoboWM- Bench is inherently scalable, allowing new tasks to be added through simulation- native assets or reconstructed real-world scenes within the same modular pipeline. 3.4 Manipulation Task Suite with Diverse Complexity RoboWM-Bench comprises a suite of manipulation tasks with varying levels of complexity, designed to systematically evaluate the embodied reasoning and physical consistency of video world models. Built on the LeHome simulation en- gine [30], the benchmark enables reliable reproduction of a wide range of manip- ulation behaviors and object interactions within a physically realistic simulation. RoboWM-Bench includes diverse object types and interaction regimes. The task suite begins with basic rigid-object manipulation tasks, such as object pickup and trash disposal, which primarily assess contact precision and spa- tial reasoning. It further incorporates articulated-object interactions, including drawer opening and faucet rotation, requiring models to reason over kinematic constraints and structured motion. To evaluate the modeling of non-rigid dynam- ics, we introduce deformable-object manipulation tasks such as towel folding. Beyond relatively short-horizon tasks, RoboWM-Bench features long-horizon compositional tasks, such as assembling a hamburger, which require multi-stage planning and temporal consistency. Additionally, the benchmark includes bi- manual manipulation tasks, such as object handover and collaborative towel folding, which introduce coordination constraints between two hands. Together, this structured task design enables comprehensive evaluation across varying ob- ject properties, interaction dynamics, and temporal horizons. 3.5 Evaluation of Embodied Executability We define embodied executability as whether predicted behaviors can be trans- lated into dynamically feasible action sequences that accomplish the intended task. Our evaluation protocol consists of both step-level verification and final task-level success assessment. Concretely, for each task, we predefine a set of key action nodes corresponding to semantically meaningful and task-structured interaction stages, such as contact events (e.g., grasping) or designated moments when the end-effector is required to reach a stable configuration (e.g., lifting). During execution, we conduct step-level verification to evaluate whether the pre- dicted behavior satisfies the required interaction and dynamical constraints at each key node. A trajectory is considered task-level successful only if all key nodes pass the step-level checks and the task objective is ultimately achieved. This hierarchical protocol enables fine-grained diagnosis of failure modes while maintaining a clear and measurable definition of overall task completion. 4 Experiments We conduct extensive experiments with state-of-the-art video world models across diverse manipulation tasks (Section 4.1) to systematically evaluate their embod- ied executability and validate RoboWM-Bench as a reliable evaluation frame- work. Specifically: (1) We quantify the embodied executability of current video world models, providing a detailed analysis of their performance, key limitations, and potential for generating physically executable behaviors (Section 4.2). (2) We examine the limitations of existing benchmarks and show that RoboWM- Bench offers a more principled, embodiment-grounded assessment of executabil- ity (Section 4.3). (3) We further validate the robustness of RoboWM-Bench by analyzing the consistency and soundness of its action extraction module and simulation-based execution pipeline (Section 4.4). 4.1 Environment Setup Tasks and Environments. We evaluate video world models on a diverse suite of manipulation tasks spanning both human-hand and robotic embodiments. For robotic evaluation, a Franka arm is positioned on a table with a fixed camera capturing the scene. Objects are placed on the tabletop with randomized Table 1: Embodied execution success rates (%) on RoboWM-Bench for both human- hand (top two sections) and robotic manipulation tasks (bottom two sections), reported at both the task and step levels. Human (Task Level) MethodPick Object Push Button Put on Plate Pour Water Stack Cups Open Drawer Put in Drawer Fold Towel Cosmos23%40%15%0%10%10%10%0% Wan 2.257%80%55%60%40%0%20%0% Wan 2.6 83%100%70%80%80%80%80%40% Veo 3.173%100%30%60%20%20%60%0% LVP 70%40%70%40%20%80%40%20% Human (Step Level) MethodPut on PlatePut in Drawer contactliftplacecontactliftabove drawer in drawer close drawer Cosmos90%20%15%80%20%20%20%10% Wan 2.2100%60%55%100%60%60%40%20% Wan 2.6100%75%70%100%80%80%80%80% Veo 3.1 100%70%30%100%70%70%60%60% LVP100%75%70%100%70%60%50%40% Robot (Task Level) MethodClose Drawer Pick Object Push Object Push Button Put on Plate Discard Trash Pull Object Put in Drawer Cosmos0%10%10%10%10%0%0%0% Wan 2.230%10%0%0%0%0%0%0% Wan 2.650%20%40%40%20%10%0%0% Veo 3.120%20%10%20%10%0%0%0% Cosmos-FT90%50%50%60%40%30%40%20% Robot (Step Level) MethodPut on PlatePut in Drawer contactliftplacecontactliftabove drawer in drawer close drawer Cosmos30%10%10%10%0%0%0%0% Wan 2.220%0%0%0%0%0%0%0% Wan 2.640%20%20%30%0%0%0%0% Veo 3.140%10%10%30%0%0%0%0% Cosmos-FT60%40%40%60%20%20%20%20% initial poses. We include both purely simulated scenarios and real-world scenar- ios, with the latter reconstructed into simulation via the real-to-sim pipeline for unified and controlled evaluation. For human-hand evaluation, a real human hand is initially placed above the table, and objects are placed with randomized initial poses. We focus exclusively on real-world scenarios, as simulating a real human hand within the physics engine would not faithfully capture human manipulation dynamics and thus does not provide meaningful evaluation. The evaluation environment is built on LeHome [30,38], which provides high- fidelity rendering and supports deformable object simulation. For each task, we run 10 episodes with different object initializations and report the average accuracy. To ensure reproducibility, environment initialization, task descriptions, random seeds, and evaluation protocols are standardized across all models. Baselines. We evaluate multiple SOTA video world models spanning both general-purpose and interaction-oriented systems. Among general-purpose mod- els, we include closed-source systems, Veo3.1 [17] and Wan2.6 [49], as well as open-source counterparts, Wan2.2 [49] and Cosmos-Predict2.5 [2]. For interaction- oriented models, we include LVP [9], which is specifically trained to capture complex human interactive behaviors, enabling evaluation in human-centric ma- nipulation scenarios. To further investigate the potential of video world models for embodied manipulation, we introduce Cosmos-Finetune, a variant of Cosmos fine-tuned on our collected real-world manipulation dataset (50 trajectories per task). 4.2 Embodied Executability of Video World Models Table 1 reports the execution success rates. Results for purely simulated robotic tasks are provided in Section A of the supplementary material. The experimental results reveal several consistent trends, which we analyze below. First, human-hand videos achieve higher execution success rates than robotic manipulation videos. This discrepancy is likely attributable to biases in pretraining data, as large-scale video datasets predominantly contain human interactions, whereas robotic data remains relatively scarce. Additionally, we observe that in generated videos, human hands typically maintain stable ge- ometry during interaction, while robotic manipulators are more likely to exhibit structural distortions, which can lead to execution failures. Second, task difficulty and interaction complexity significantly af- fect execution success. As tasks progress from short-horizon interactions (e.g., Push Button, Pick Object) to longer-horizon tasks (e.g., Put in Drawer), success rates decrease due to the accumulation of errors across multiple steps. Among the human-hand tasks, Fold Towel is the most challenging, suggesting that in- teractions with deformable objects remain particularly challenging for current video world models. Third, fine-tuning on robotic manipulation data significantly im- proves embodied executability. Cosmos-Finetune achieves substantially higher success rates than its pretrained counterpart, indicating that even limited task- specific data (50 trajectories per task) can enhance the generation of dynami- cally consistent and executable behaviors on robotic tasks. Fine-tuning helps the model better capture the structural priors of robotic arms, reducing deforma- tion artifacts and improving joint articulation. However, its 3D spatial reasoning remains limited, often resulting in inaccurate object localization and grasping failures. These findings highlight the potential for further improving the embod- ied capabilities of video world models. Beyond aggregated metrics, qualitative results are shown in Figure 3. Among the evaluated models, Wan2.6 achieves the strongest performance on RoboWM- Bench, and we visualize representative success and failure cases. Notably, even strong world models may generate visually plausible yet physically inconsistent interactions. For example, in Put on Plate, the predicted video shows the fingers merely touching the object without forming a stable grasp, yet H u m a n - H a n d R o b o t x Stack Cups x Push Object x Discard Trash Put in Drawer x x Put on Plate x Open Drawer Fig. 3: Qualitative execution results on RoboWM-Bench. For each task, pre- dicted videos (left) are converted into robot actions and executed in simulation (right). the object is lifted. Such interactions are physically implausible and fail during execution. A similar issue occurs in Open Drawer, the predicted motion resembles closing the drawer without establishing a proper grasp, while the simulated ex- ecution correctly reflects the physical outcome, where the gripper instead closes the drawer. For robotic manipulation videos, in addition to unrealistic contact behaviors, the predicted robot structure is also more prone to geometric distor- tions, further reducing execution reliability. 4.3 Perceptual Plausibility vs. Embodied Executability We compare execution accuracy in RoboWM-Bench with the domain scores in PAI-Bench, a commonly used metric for evaluating the perceptual plausibility of generated videos. As shown in Figure 4, the same predicted videos obtain near- saturated scores on PAI-Bench across different world models, whereas RoboWM- Bench evaluates whether the predicted behaviors are physically executable, re- sulting in more discriminative outcomes. This discrepancy arises because some actions may appear visually plausible (Figure 3) yet remain physically infeasible, an issue that perceptual domain scores may not capture but becomes evident under embodied execution. These results highlight the complementary role of RoboWM-Bench and show that embodiment-grounded evaluation provides a more direct measure of physical executability. 4.4 Robustness of RoboWM-Bench To assess the robustness of RoboWM-Bench, we evaluate two key components of the evaluation pipeline: (1) the accuracy of action extraction from videos, includ- The banana is held between the thumb and index finger, so it won't fall. Fig. 4: Comparison between PAI-Bench and RoboWM-Bench. Table 2: Accuracy of action extraction methods. Human MethodPick Object Stack Cups Pour Water Open Drawer Fold Towel Put on Plate Put in Drawer Average Retargeting100%90%90%100%100%100%100%97.1% Robot MethodPick Object Pull Object Push Object Discard Trash Close Drawer Put on Plate Put in Drawer Average IDM Real 70%70%80%70%90%70%50%71.4% IDM Sim+Real 100%90%100%90%100%100%90%95.7% ing pose tracking and retargeting for human-hand videos, and inverse dynamics modeling (IDM) for robotic manipulation videos; (2) the fidelity of reconstructed simulation environments with respect to their corresponding real-world scenes. Action Extraction Accuracy. To evaluate the accuracy of action extraction from videos, we collect a set of real-world manipulation trajectories that suc- cessfully accomplish the target tasks. The corresponding videos are processed using our action extraction pipeline, and the extracted actions are executed in simulation to verify whether they reproduce the original task outcomes. As shown in Table 2, for human-hand videos, the pose tracking and retarget- ing pipeline achieves near-perfect execution success, indicating high reliability. Failure in Stack Cups and Pour Water arises from slight discrepancies between the robot gripper’s contact locations on the cup and those of human fingertips. As cylindrical objects impose strict constraints on grasp contact locations, even small deviations can lead to unstable grasps and cause execution failure. SuccessConsistencyFailureConsistencyConsistency Table Fig. 5: Real-to-sim consistency evaluation. Identical manipulation trajectories are executed in real-world scenes and reconstructed simulation environments, yielding consistent success and failure outcomes. For robotic videos, we compare two IDM training strategies: IDM Real , trained directly on real-world data (50 trajectories per task), and IDM Sim+Real , a two- stage approach consisting of simulation pretraining followed by real-world fine- tuning. The latter significantly improves execution success rates, demonstrating that simulation pretraining provides useful motion priors and leads to more stable inverse-dynamics prediction. Nevertheless, a small number of tasks still exhibit non-perfect success rates. This is primarily caused by minor prediction errors in the IDM outputs. Although these errors have negligible impacts in most tasks, they may cause failure when the object is grasped with intentionally shallow contacts, where even slight deviations can compromise grasp stability. Simulation Reconstruction Fidelity. We also evaluate the consistency be- tween reconstructed simulation environments and their real-world counterparts. Ideally, executing the same manipulation trajectory in both domains should yield identical outcomes. As the reconstruction pipeline is identical for human-hand and robotic tasks, we use robotic tasks as representative examples. To assess this consistency, for each task, we collect 10 successful and 10 failed real-world manipulation trajectories and replay the same actions in the corre- sponding reconstructed simulation environments. Quantitative and qualitative results are shown in Figure 5. Execution outcomes remain consistent across do- mains, with successes and failures faithfully reproduced in simulation. These results indicate that the reconstructed simulation environments faithfully pre- serve the physical structure and interaction dynamics of real scenes, enabling reliable and reproducible evaluation without requiring the original physical se- tups. Additional visualizations are provided in the supplementary video. 5 Conclusion In this work, we introduce RoboWM-Bench, a systematic and reproducible benchmark for evaluating video world models through embodiment-grounded robotic manipulation. By operationalizing physical executability as a measur- able criterion, RoboWM-Bench assesses whether predicted behaviors can be translated into dynamically feasible actions and accomplish the intended tasks. The benchmark supports both human-hand and robotic videos and spans di- verse tasks involving rigid, articulated, deformable, and long-horizon interac- tions. Experiments with state-of-the-art video world models reveal that em- bodied executability degrades significantly as task complexity increases, par- ticularly for long-horizon manipulations and deformable-object interactions. In addition, human-hand videos generally achieve higher execution success than robotic videos, reflecting biases in large-scale training data. While fine-tuning on manipulation-specific data improves executability, notable physical inconsis- tencies remain. In addition, we validate the robustness of RoboWM-Bench by evaluating the accuracy of action extraction and the fidelity of simulation recon- struction. Overall, RoboWM-Bench provides a principled evaluation protocol that bridges video generation and robotic control, offering a standardized foun- dation for advancing physically consistent world models for robotic applications. References 1. Agarwal, N., Ali, A., Bala, M., Balaji, Y., Barker, E., Cai, T., Chattopadhyay, P., Chen, Y., Cui, Y., Ding, Y., et al.: Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575 (2025) 2. Ali, A., Bai, J., Bala, M., Balaji, Y., Blakeman, A., Cai, T., Cao, J., Cao, T., Cha, E., Chao, Y.W., et al.: World simulation with video foundation models for physical ai. arXiv preprint arXiv:2511.00062 (2025) 3. Assran, M., Bardes, A., Fan, D., Garrido, Q., Howes, R., Muckley, M., Rizvi, A., Roberts, C., Sinha, K., Zholus, A., et al.: V-jepa 2: Self-supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985 (2025) 4. Bai, S., Cai, Y., Chen, R., Chen, K., Chen, X., Cheng, Z., Deng, L., Ding, W., Gao, C., Ge, C., Ge, W., Guo, Z., Huang, Q., Huang, J., Huang, F., Hui, B., Jiang, S., Li, Z., Li, M., Li, M., Li, K., Lin, Z., Lin, J., Liu, X., Liu, J., Liu, C., Liu, Y., Liu, D., Liu, S., Lu, D., Luo, R., Lv, C., Men, R., Meng, L., Ren, X., Ren, X., Song, S., Sun, Y., Tang, J., Tu, J., Wan, J., Wang, P., Wang, P., Wang, Q., Wang, Y., Xie, T., Xu, Y., Xu, H., Xu, J., Yang, Z., Yang, M., Yang, J., Yang, A., Yu, B., Zhang, F., Zhang, H., Zhang, X., Zheng, B., Zhong, H., Zhou, J., Zhou, F., Zhou, J., Zhu, Y., Zhu, K.: Qwen3-vl technical report. arXiv preprint arXiv:2511.21631 (2025) 5. Baker, B., Akkaya, I., Zhokov, P., Huizinga, J., Tang, J., Ecoffet, A., Houghton, B., Sampedro, R., Clune, J.: Video pretraining (vpt): Learning to act by watching unlabeled online videos. Advances in Neural Information Processing Systems 35, 24639–24654 (2022) 6. Bharadhwaj, H., Dwibedi, D., Gupta, A., Tulsiani, S., Doersch, C., Xiao, T., Shah, D., Xia, F., Sadigh, D., Kirmani, S.: Gen2act: Human video generation in novel sce- narios enables generalizable robot manipulation. arXiv preprint arXiv:2409.16283 (2024) 7. Bjorck, J., Castañeda, F., Cherniadev, N., Da, X., Ding, R., Fan, L., Fang, Y., Fox, D., Hu, F., Huang, S., et al.: Gr00t n1: An open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734 (2025) 8. Brooks, T., Peebles, W., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Ng, C., Wang, R., Ramesh, A., et al.: Video generation models as world simulators. https://openai.com/research/video-generation- models-as-world-simulators (2024), openAI Technical Report 9. Chen, B., Zhang, T., Geng, H., Song, K., Zhang, C., Li, P., Freeman, W.T., Malik, J., Abbeel, P., Tedrake, R., et al.: Large video planner enables generalizable robot control. arXiv preprint arXiv:2512.15840 (2025) 10. Chen, S., Guo, H., Zhu, S., Zhang, F., Huang, Z., Feng, J., Kang, B.: Video depth anything: Consistent depth estimation for super-long videos. In: Proceedings of the Computer Vision and Pattern Recognition Conference. p. 22831–22840 (2025) 11. Chen, X., Chu, F.J., Gleize, P., Liang, K.J., Sax, A., Tang, H., Wang, W., Guo, M., Hardin, T., Li, X., et al.: Sam 3d: 3dfy anything in images. arXiv preprint arXiv:2511.16624 (2025) 12. Chi, X., Jia, P., Fan, C.K., Ju, X., Mi, W., Zhang, K., Qin, Z., Tian, W., Ge, K., Li, H., et al.: Wow: Towards a world omniscient world model through embodied interaction. arXiv preprint arXiv:2509.22642 (2025) 13. Du, Y., Yang, S., Dai, B., Dai, H., Nachum, O., Tenenbaum, J., Schuurmans, D., Abbeel, P.: Learning universal policies via text-guided video generation. Advances in neural information processing systems 36, 9156–9172 (2023) 14. Fan, C.K., Chi, X., Ju, X., Li, H., Bao, Y., Wang, Y.K., Chen, L., Jiang, Z., Ge, K., Li, Y., et al.: Wow, wo, val! a comprehensive embodied world model evaluation turing test. arXiv preprint arXiv:2601.04137 (2026) 15. Feng, W., Li, J., Saxon, M., Fu, T.j., Chen, W., Wang, W.Y.: Tc-bench: Bench- marking temporal compositionality in text-to-video and image-to-video generation. arXiv preprint arXiv:2406.08656 (2024) 16. Gao, Y., Guo, H., Hoang, T., Huang, W., Jiang, L., Kong, F., Li, H., Li, J., Li, L., Li, X., et al.: Seedance 1.0: Exploring the boundaries of video generation models. arXiv preprint arXiv:2506.09113 (2025) 17. Google DeepMind: Veo: a text-to-video generation system (veo-3 technical re- port). Tech. Rep. Veo-3-Tech-Report, Google DeepMind (2025), https://storage. googleapis.com/deepmind-media/veo/Veo-3-Tech-Report.pdf, technical Re- port 18. Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Ham- burger, J., Jiang, H., Liu, M., Liu, X., et al.: Ego4d: Around the world in 3,000 hours of egocentric video. In: Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition. p. 18995–19012 (2022) 19. HaCohen, Y., Chiprut, N., Brazowski, B., Shalem, D., Moshe, D., Richardson, E., Levin, E., Shiran, G., Zabari, N., Gordon, O., et al.: Ltx-video: Realtime video latent diffusion. arXiv preprint arXiv:2501.00103 (2024) 20. Hong, W., Ding, M., Zheng, W., Liu, X., Tang, J.: Cogvideo: Large-scale pretrain- ing for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868 (2022) 21. Huang, S., Chen, L., Zhou, P., Chen, S., Jiang, Z., Hu, Y., Liao, Y., Gao, P., Li, H., Yao, M., et al.: Enerverse: Envisioning embodied future space for robotics manipulation (2025). arXiv preprint arXiv:2501.01895 (2025) 22. Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., et al.: Vbench: Comprehensive benchmark suite for video gener- ative models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 21807–21818 (2024) 23. Jang, J., Ye, S., Lin, Z., Xiang, J., Bjorck, J., Fang, Y., Hu, F., Huang, S., Kundalia, K., Lin, Y.C., et al.: Dreamgen: Unlocking generalization in robot learning through video world models. arXiv preprint arXiv:2505.12705 (2025) 24. Ji, P., Xiao, C., Tai, H., Huo, M.: T2vbench: Benchmarking temporal dynamics for text-to-video generation. In: Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition. p. 5325–5335 (2024) 25. Kareer, S., Pertsch, K., Darpinian, J., Hoffman, J., Xu, D., Levine, S., Finn, C., Nair, S.: Emergence of human to robot transfer in vision-language-action models. arXiv preprint arXiv:2512.22414 (2025) 26. Labbé, Y., Manuelli, L., Mousavian, A., Tyree, S., Birchfield, S., Tremblay, J., Carpentier, J., Aubry, M., Fox, D., Sivic, J.: Megapose: 6d pose estimation of novel objects via render & compare. arXiv preprint arXiv:2212.06870 (2022) 27. Lepert, M., Fang, J., Bohg, J.: Masquerade: Learning from in-the-wild human videos using data-editing. arXiv preprint arXiv:2508.09976 (2025) 28. Lepert, M., Fang, J., Bohg, J.: Phantom: Training robots without robots using only human videos. URL https://arxiv. org/abs/2503.00779 2 (2025) 29. Li, D., Fang, Y., Chen, Y., Yang, S., Cao, S., Wong, J., Luo, M., Wang, X., Yin, H., Gonzalez, J.E., et al.: Worldmodelbench: Judging video generation models as world models. arXiv preprint arXiv:2502.20694 (2025) 30. Li, Z., Yang, J., Xu, J., Xie, S., Chen, T., Wang, Y., Shen, Z., Shen, Y., Zheng, Y., Li, W., et al.: Lehome: A simulation environment for deformable object ma- nipulation in household scenarios. In: IROS 2025-5th Workshop on RObotic MA- nipulation of Deformable Objects: holistic approaches and challenges forward 31. Li, Z., Tucker, R., Cole, F., Wang, Q., Jin, L., Ye, V., Kanazawa, A., Holynski, A., Snavely, N.: Megasam: Accurate, fast and robust structure and motion from casual dynamic videos. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 10486–10496 (2025) 32. Ling, X., Zhu, C., Wu, M., Li, H., Feng, X., Yang, C., Hao, A., Zhu, J., Wu, J., Chu, X.: Vmbench: A benchmark for perception-aligned video motion generation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 13087–13098 (2025) 33. Liu, Y., Cun, X., Liu, X., Wang, X., Zhang, Y., Chen, H., Liu, Y., Zeng, T., Chan, R., Shan, Y.: Evalcrafter: Benchmarking and evaluating large video generation models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 22139–22149 (2024) 34. Luo, Y., Du, Y.: Grounding video models to actions through goal conditioned exploration. arXiv preprint arXiv:2411.07223 (2024) 35. Ma, Y.J., Sodhani, S., Jayaraman, D., Bastani, O., Kumar, V., Zhang, A.: Vip: Towards universal visual reward and representation via value-implicit pre-training. arXiv preprint arXiv:2210.00030 (2022) 36. McCarthy, R., Tan, D.C., Schmidt, D., Acero, F., Herr, N., Du, Y., Thuruthel, T.G., Li, Z.: Towards generalist robot learning from internet video: A survey. Jour- nal of Artificial Intelligence Research 83 (2025) 37. Nair, S., Rajeswaran, A., Kumar, V., Finn, C., Gupta, A.: R3m: A universal visual representation for robot manipulation. arXiv preprint arXiv:2203.12601 (2022) 38. NVIDIA Corporation: Nvidia isaac sim: High-fidelity simulation for robotics. https://developer.nvidia.com/isaac-sim (2023), accessed: 2026-03-03 39. Pavlakos, G., Shan, D., Radosavovic, I., Kanazawa, A., Fouhey, D., Malik, J.: Reconstructing hands in 3d with transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 9826–9836 (2024) 40. Punamiya, R., Patel, D., Aphiwetsa, P., Kuppili, P., Zhu, L.Y., Kareer, S., Hoff- man, J., Xu, D.: Egobridge: Domain adaptation for generalizable imitation from egocentric human data. In: Human to Robot: Workshop on Sensorizing, Modeling, and Learning from Humans (2025) 41. Qin, Y., Wu, Y.H., Liu, S., Jiang, H., Yang, R., Fu, Y., Wang, X.: Dexmv: Imitation learning for dexterous manipulation from human videos. In: European Conference on Computer Vision. p. 570–587. Springer (2022) 42. Radosavovic, I., Xiao, T., James, S., Abbeel, P., Malik, J., Darrell, T.: Real-world robot learning with masked visual pre-training. In: Conference on Robot Learning. p. 416–426. PMLR (2023) 43. Shang, Y., Li, Z., Ma, Y., Su, W., Jin, X., Wang, Z., Jin, L., Zhang, X., Tang, Y., Su, H., et al.: Worldarena: A unified benchmark for evaluating perception and functional utility of embodied world models. arXiv preprint arXiv:2602.08971 (2026) 44. Team, G.R., Choromanski, K., Devin, C., Du, Y., Dwibedi, D., Gao, R., Jindal, A., Kipf, T., Kirmani, S., Leal, I., et al.: Evaluating gemini robotics policies in a veo world simulator. arXiv preprint arXiv:2512.10675 (2025) 45. Team, G., Ye, A., Wang, B., Ni, C., Huang, G., Zhao, G., Li, H., Li, J., Zhu, J., Feng, L., et al.: Gigabrain-0: A world model-powered vision-language-action model. arXiv preprint arXiv:2510.19430 (2025) 46. Team, G., Ye, A., Wang, B., Ni, C., Huang, G., Zhao, G., Li, H., Zhu, J., Li, K., Xu, M., et al.: Gigaworld-0: World models as data engine to empower embodied ai. arXiv preprint arXiv:2511.19861 (2025) 47. Team, K., Chen, J., Ding, Y., Fang, Z., Gai, K., Gao, Y., He, K., Hua, J., Jiang, B., Lao, M., et al.: Klingavatar 2.0 technical report. arXiv preprint arXiv:2512.13313 (2025) 48. Tian, Y., Yang, S., Zeng, J., Wang, P., Lin, D., Dong, H., Pang, J.: Predictive inverse dynamics models are scalable learners for robotic manipulation. arXiv preprint arXiv:2412.15109 (2024) 49. Wan, T., Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.W., Chen, D., Yu, F., Zhao, H., Yang, J., et al.: Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314 (2025) 50. Wang, Y., Chen, X., Ma, X., Zhou, S., Huang, Z., Wang, Y., Yang, C., He, Y., Yu, J., Yang, P., et al.: Lavie: High-quality video generation with cascaded latent diffu- sion models. International Journal of Computer Vision 133(5), 3059–3078 (2025) 51. Wen, B., Yang, W., Kautz, J., Birchfield, S.: Foundationpose: Unified 6d pose esti- mation and tracking of novel objects. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 17868–17879 (2024) 52. World Labs: Marble: A Multimodal World Model (2025), https://w.worldlabs. ai/blog/marble-world-model, accessed: 2026-02 53. Wu, B., Zou, C., Li, C., Huang, D., Yang, F., Tan, H., Peng, J., Wu, J., Xiong, J., Jiang, J., et al.: Hunyuanvideo 1.5 technical report. arXiv preprint arXiv:2511.18870 (2025) 54. Wu, T., Zhang, J., Liang, S., Han, Z., Dong, H.: Foundation feature-driven online end-effector pose estimation: A marker-free and learning-free approach. In: 2025 IEEE International Conference on Robotics and Automation (ICRA). p. 1921– 1928. IEEE (2025) 55. Xiong, H., Li, Q., Chen, Y.C., Bharadhwaj, H., Sinha, S., Garg, A.: Learning by watching: Physical imitation of manipulation skills from human videos. In: 2021 IEEE/RSJ international conference on intelligent robots and systems (iros). p. 7827–7834. IEEE (2021) 56. Yang, L., Kang, B., Huang, Z., Xu, X., Feng, J., Zhao, H.: Depth anything: Un- leashing the power of large-scale unlabeled data. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 10371–10381 (2024) 57. Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Schuurmans, D., Abbeel, P.: Learning interactive real-world simulators. arXiv preprint arXiv:2310.06114 1(2), 6 (2023) 58. Ye, S., Jang, J., Jeon, B., Joo, S., Yang, J., Peng, B., Mandlekar, A., Tan, R., Chao, Y.W., Lin, B.Y., et al.: Latent action pretraining from videos. arXiv preprint arXiv:2410.11758 (2024) 59. Yue, H., Huang, S., Liao, Y., Chen, S., Zhou, P., Chen, L., Yao, M., Ren, G.: Ewmbench: Evaluating scene, motion, and semantic quality in embodied world models. arXiv preprint arXiv:2505.09694 (2025) 60. Zheng, D., Huang, Z., Liu, H., Zou, K., He, Y., Zhang, F., Gu, L., Zhang, Y., He, J., Zheng, W.S., et al.: Vbench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness. arXiv preprint arXiv:2503.21755 (2025) 61. Zhou, F., Huang, J., Li, J., Ramanan, D., Shi, H.: Pai-bench: A comprehensive benchmark for physical ai. arXiv preprint arXiv:2512.01989 (2025) 62. Zhou, S., Du, Y., Chen, J., Li, Y., Yeung, D.Y., Gan, C.: Robodreamer: Learning compositional world models for robot imagination. arXiv preprint arXiv:2404.12377 (2024) Supplementary Material A Additional Results for Purely Simulated Robotic Tasks RoboWM-Bench also includes a set of robotic manipulation tasks evaluated en- tirely in simulation environments. The task setup follows the same protocol described in the main paper. For each task, video world models generate future manipulation behaviors condi- tioned on the simulation observations and the corresponding task descriptions. The predicted behaviors are then converted into executable action sequences and executed in simulation to assess task completion. Table 3 reports both task-level and step-level execution success rates for these purely simulated tasks. The overall trends are consistent with those observed in the real-to-sim evaluation. As task complexity increases, the success rates of most models decrease, particularly for tasks requiring long-horizon reasoning or precise contact interactions. Table 3: Task-level and step-level embodied execution success rates (%) on RoboWM- Bench for robotic manipulation tasks evaluated entirely in simulation. MethodRobot (Task Level) Close Drawer Push Button Cut Sausage Turn Off Faucet Assemble Burger Fold Clothes Cosmos0%10%10%0%0%0% Wan 2.210%10%10%0%0%0% Wan 2.630%20%40%0%0%0% Veo10%20%20%0%0%0% Robot (Step Level) MethodTurn Off FaucetAssemble BurgerFold Clothes contactrot. > 20 ◦ turn offcontactliftplaceL. sleeve R. sleeve Cosmos10%0%0%10%0%0%0%0% Wan 2.210%0%0%10%0%0%0%0% Wan 2.640%10%0%30%0%0%0%0% Veo20%0%0%10%0%0%0%0% B Additional Results on Human Tasks We additionally include human-hand manipulation tasks involving bimanual co- ordination. The detailed execution success rates are reported in Table 4. The results show that bimanual tasks are generally more challenging, as success- ful execution requires coordinated interactions between both hands to maintain physically feasible manipulation. Table 4: Task-level and step-level embodied execution success rates (%) on the addi- tional human manipulation tasks involving bimanual coordination. Task LevelStep Level MethodCookgrasp spatula grasp pan lift spatula lift panspatula– pan contact place spatula place pan Cosmos10%30%20%20%20%20%10%20% Wan 2.2 30%70%60%40%30%30%30%30% Wan 2.650%90%70%70%70%60%50%60% Veo 40%80%70%70%50%60%40%50% LVP30%80%70%60%50%40%40%30% Task LevelStep Level MethodLift Large Boxgrasp left grasp right liftplace Cosmos10%20%20%10%10% Wan 2.230%70%60%30%30% Wan 2.660%90%90%70%60% Veo50%80%80%50%50% LVP30%70%70%30%30% C Comparison with PAI-Bench Quality Scores We additionally compare the average quality scores in PAI-Bench [61] with the execution accuracy in RoboWM-Bench, as shown in Figure 6. Specifically, Fig- ure 6 presents two scatter plots corresponding to human-hand tasks and robotic tasks, respectively. In both plots, the horizontal axis denotes the average quality score in PAI-Bench, while the vertical axis represents the execution accuracy measured in RoboWM-Bench. Each point corresponds to a specific world model baseline (indicated by color) on a particular manipulation task, reporting its average quality score and execution accuracy. As observed in the scatter plots, most points cluster along a vertical line around an average quality score of approximately 0.78, suggesting that the PAI- Bench quality scores are relatively consistent across different tasks and models. In contrast, the execution accuracy measured by RoboWM-Bench exhibits sub- stantially greater variation. This discrepancy suggests that visual plausibility does not necessarily imply physical correctness, and the embodiment-grounded evaluation in RoboWM-Bench provides a more sensitive and informative mea- sure of physical executability. Following the definitions and reporting protocol of PAI-Bench, we addition- ally present detailed domain and quality scores in tabular form. Note that all scores are computed using videos generated for the manipulation tasks in RoboWM-Bench, rather than the broader video sources used in PAI-Bench. Specifically, the detailed domain scores are reported in Table 5 for human-hand tasks and Table 6 for robotic tasks, while the detailed quality scores are reported in Table 7 and Table 8, respectively. Fig. 6: Comparison between the average quality scores in PAI-Bench with the execution accuracy in RoboWM-Bench. The left scatter plot shows human-hand tasks, and the right scatter plot shows robotic tasks. Table 5: PAI-Bench Domain Score on Human-Hand Tasks. ModelsDomain ScoreAvg. Pick Object Push Button Put on Plate Pour Water Stack Cups Open Drawer Put in Drawer Fold Towel Cosmos961009272781001008690.5 Wan 2.2921001009410010010010098.3 Wan 2.6100921001001001001008697.3 Veo10082100100100929210095.8 LVP1007810078100100867687.0 Table 6: PAI-Bench Domain Score on Robotic Real Tasks. ModelsDomain ScoresAvg. Close Drawer Pick Object Push Object Push Button Put on Plate Discard Trash Pull Object Put in Drawer Cosmos8410010010010010010010098.0 Wan 2.21001009210094100988696.3 Wan 2.610010010010010010010010099.5 Veo100100100901001001008096.3 Cosmos-FT 1001001001001009210010099.0 Table 7: PAI-Bench Quality Score on Human-Hand Tasks. Metrics follow the PAI- Bench protocol, including Subject Consistency (SC), Background Consistency (BC), Motion Smoothness (MS), Aesthetic Quality (AQ), Imaging Quality (IQ), Overall Con- sistency (OC), I2V Subject (IS), I2V Background (IB). ModelsQuality ScoreAvg. SCBCMSAQIQOCISIB Cosmos95.094.999.545.874.925.698.398.579.1 Wan 2.2 95.594.999.141.774.526.298.698.578.6 Wan 2.696.995.699.240.575.925.898.497.978.8 Veo 96.095.599.640.975.426.398.598.678.9 LVP96.395.999.436.474.425.698.198.078.0 Table 8: PAI-Bench Quality Score on Robotic Real Tasks. ModelsQuality ScoreAvg. SCBCMSAQIQOCISIB Cosmos94.992.999.444.374.522.294.093.577.0 Wan 2.296.594.699.543.973.523.294.693.277.4 Wan 2.696.996.299.445.177.123.394.392.878.1 Veo95.192.599.644.470.123.695.094.876.9 Cosmos-FT 94.694.299.544.364.922.898.298.577.1 D More Visualizations In this section, we provide additional visualization results for qualitative analysis. D.1 Predicted Manipulation Videos and Embodied Execution Figure 7 provides additional qualitative examples of the predicted manipulation videos and their corresponding embodied executions. These results complement the examples shown in Figure 3 of the main paper. D.2 Real-to-Sim Scene Consistency Figure 8 provides additional qualitative examples of identical manipulation tra- jectories executed in real-world scenes and their reconstructed simulation envi- ronments, complementing the examples shown in Figure 5 of the main paper. E Discussion of Depth in Human-Hand Tracking In our experiments, we empirically found that Phantom [28] provides the most reliable performance for pose tracking and retargeting in human-hand manip- ulation videos. Therefore, as described in Section 3.2 of the main paper, our pipeline builds upon Phantom with several adaptations. Human-Hand Robot Sim x Stack Cups Put in Drawer Push Button Put on Plate Push Button Cut Sausage Robot Real Discard Trash Pick Banana Turn off Faucet Assemble Burger Pick Banana Cook Fig. 7: Additional qualitative results on RoboWM-Bench. For each task, pre- dicted videos (left) are converted into robot actions and executed in simulation (right). Real Sim Real Sim Success ConsistencyFailure Consistency Fig. 8: Additional qualitative results on real-to-sim consistency evaluation. Identical manipulation trajectories are executed in real-world scenes (left) and recon- structed simulation environments (right). Consistent outcomes demonstrate the fidelity of the reconstruction pipeline. However, Phantom assumes access to ground-truth depth information, whereas in our setting such depth is not available for the videos predicted by world mod- els. In practice, only the depth of the first frame can be obtained, as it is captured by the camera in the real-world or simulation environment. Recent works have explored leveraging depth estimation for video under- standing [9,31]. In our setting, we investigate whether estimated depth can im- prove human-hand tracking. To this end, we estimate depth from RGB frames using Video Depth Anything [10, 56] and evaluate two strategies. First, we di- rectly use the absolute depth predicted by the model. However, as shown in Figure 9(a), the predicted depth shows a large discrepancy from the ground- truth. Second, we estimate relative depth for subsequent frames and align it with the ground-truth depth of the first frame to recover absolute depth values. As illustrated in Figure 9(b), the recovered depth remains imperfect. Empirically, we observe that incorporating these estimated depths does not improve downstream pose tracking and retargeting performance. Therefore, depth information is not used in the final pipeline. F Prompt Details for World Model Video Generation In this section, we provide details on how instruction prompts are generated for world model video generation. (a) Predicted Absolute Depth Comparison (b) Predicted Relative Depth Comparison Ground TruthPredictedError Map Ground Truth Predicted (Aligned) Error Map Mean Error: 0.0512m Mean Error: 0.7645m Fig. 9: Visualization of depth estimation results. (a) The predicted absolute depth shows a large discrepancy from the ground-truth. (b) Aligning relative depth with the first-frame ground-truth depth improves consistency, although non-negligible errors still remain. F.1 Human-Hand Tasks To generate instruction prompts for video world models on human tasks, we leverage the Qwen3-vl-flash [4] model to generate concise task-specific descrip- tions conditioned on the initial observation image and high-level task instructions (e.g., Pick up the stapler on the table, and stay still.). The resulting descriptions generated by Qwen are summarized in Table 9 and Table 10. To further encourage physical consistency of the synthesized videos, we ap- pend a standardized set of constraints to the generated instruction: "The entire human hand must remain fully visible in the frame at all times, with no cropping, no fingers cut off, and no part of the hand outside the camera view. The cam- era must remain completely static with no movement, no panning, no tilting, no zooming, and no change in viewpoint during the entire video. No extra or unnec- essary motions are allowed, the hand must not perform any additional gestures such as turning to show the palm, posing, rotating unnecessarily. All other objects in the scene that are not being manipulated must remain completely stationary, with no position shift, no rotation, and no change in placement throughout the entire video." Empirically, we observe that including these constraints improves the quality of videos generated by the world models. Table 9: Overview of human-hand manipulation tasks and their corresponding in- struction prompts. Task NameInput ImageTask Prompt Pick ObjectThe hand lowers slowly toward the stapler and gently grasps it using the thumb and index finger. The stapler is lifted slowly from the table surface without any bouncing or sliding. Push ButtonThe hand moves downward toward the button. The fingers make contact with the yellow surface. The hand applies pressure causing the button to compress. The hand lifts up and rest in the air. Put on PlateThe hand slowly extend downward to gently grasp the banana using the thumb and index finger and then slowly lifts the banana straight up. The banana is then slowly placed on the center of the plate. Pour WaterThe hand grasps the plastic cup using the thumb and the index finger. It slowly lifts the cup from the table. It tilts the cup to pour water into the paper cup. It releases the plastic cup. Stack CupsThe hand grasps the left edge of the left cup using the thumb and the index fingers and slowly lifts it, then moves it above the right cup and lowers it to stack inside. Open DrawerThe hand grasps the left edge of the transparent drawer using the thumb and the index fingers then pulls it leftward causing the drawer to slide open. Put in DrawerThe hand slowly moves to grasp the banana using the thumb and the index finger, lifts it and places it inside the clear drawer. The hand pushes the drawer into the container until it is fully closed. Fold TowelThe hand grasps the left edge using the thumb and the index fingers and lifts it. The hand folds the towel over to the right edge. The hand aligns the edges and presses down to secure the fold. Table 10: Overview of human-hand manipulation tasks and their corresponding in- struction prompts (continued). Task NameInput ImageTask Prompt CookThe person uses their left hand to pick up the wooden spatula and places it into the frying pan, while their right hand holds the handle of the pan and lift it up. They move the spatula around inside the pan briefly as if stirring or scraping, then lift the spatula out and place the wooden spatula and the frying pan back on the table in their original position. Lift Large BoxThe person grasps both the left and right edges of the box using the thumb and the index fingers, lifting it vertically off the table. F.2 Robotic Tasks To generate instruction prompts for video world models on robotic manipulation tasks, we also leverage the Qwen3-vl-flash [4] model to generate concise task- specific descriptions conditioned on the initial observation image and high-level task instructions (e.g., close drawer). The resulting descriptions generated by Qwen are summarized in Table 11 and Table 12. To further encourage physical consistency of the synthesized videos, we ap- pend a standardized set of constraints to the generated instruction: "Throughout the entire video, the gripper undergoes no structural deformation, only the open- ing angle of its jaws changes; the rotational movements of the robotic arm joints strictly adhere to its inherent mechanical structure; and the camera perspective remains completely unchanged." Empirically, we observe that including these constraints improves the quality of videos generated by the world models. Table 11: Overview of real-to-sim robotic manipulation tasks and their corresponding instruction prompts. Task NameInput ImageTask Prompt Close DrawerThe robotic arm closes its gripper and pushes the drawer back to its closed position. Pick ObjectThe robotic arm picks up the white cup from the table surface. Push ObjectThe robotic arm closes its gripper and pushes the green cube away from the base of the arm for a short distance. Push ButtonThe robotic arm closes its gripper, then uses the gripper to press the yellow button on the tabletop. Put on PlateThe robotic arm picks up the white cup from the table and moves its gripper above the plate, then releases the cup placing it on the plate. Discard TrashThe robotic arm picks up the green cube from the table and moves its gripper above the trash bin, then releases the gripper to drop the green cube into the trash bin. Pull ObjectThe robotic arm closes its gripper and pulls the green cube toward the robotic arm base for a short distance. Put in DrawerThe robotic arm picks up the yellow banana, moves it above the open drawer, releases the gripper so the banana falls inside, then closes the gripper to push the drawer back to its closed position. Table 12: Overview of purely simulated robotic manipulation tasks and their corre- sponding instruction prompts. Task NameInput ImageTask Prompt Close DrawerThe robotic arm closes its gripper and pushes the drawer back to its closed position. Push ButtonThe robotic arm closes its gripper, then uses the gripper to press the yellow button on the tabletop. Cut SausageThe robotic arm uses a knife to cut the sausage on the table. Turn Off Faucet The robotic arm grasps the lever handle on the left side of the faucet, then rotates it inward to turn off the water. Assemble Burger The robotic arm first pushes the meat patty from the left side of the cutting board to its edge, then picks it up and places it on top of the cheese in the plate located on the right side of the cutting board. Fold ClothesThe left arm grasps the cuff of the left sleeve, while the right arm grasps the cuff of the right sleeve. Both arms then lift and fold the sleeves inward toward the center of the shirt: the left arm places the left cuff onto the left side of the shirt’s body, and the right arm places the right cuff onto the right side of the body, aligning them neatly along the torso. G Implementation Details of PAI-Bench Domain Score To facilitate evaluation within the PAI-Bench framework, we design targeted VQA suites for human-hand tasks and robotic tasks, respectively. Within each suite, only the first question is task-dependent and varies across scenarios, while the remaining four questions are identical across all tasks. G.1 VQA Pairs for Human-Hand Tasks Question 1: Does the human successfully complete the task: The task descrip- tion, such as the human hand grasping and lifting the banana from the table surface. options: (A) yes (B) no (C) unclear ground truth: A Question 2: Does the human hand make physical contact with the target ob- ject? options: (A) yes (B) no (C) unclear ground truth: A Question 3: Does the human hand maintain anatomically plausible hand struc- ture throughout the video (no impossible bends/broken fingers/extra joints)? options: (A) yes (B) no (C) unclear ground truth: A Question 4: Do all task-relevant objects maintain their structural integrity without undergoing physically implausible deformations? options: (A) yes (B) no (C) unclear ground truth: A Question 5: Do any of the task-relevant objects exhibit physically implausible motions, such as sudden spatial displacement? options: (A) yes (B) no (C) unclear ground truth: A G.2 VQA Pairs for Robotic Tasks Question 1: Does the robot successfully complete the task: The task descrip- tion, such as the robotic arm picks up the small yellow cube from the tabletop. options: (A) yes (B) no (C) unclear ground truth: A Question 2: Does the robot gripper/hand make physical contact with the target object? options: (A) yes (B) no (C) unclear ground truth: A Question 3: Does the robotic system (arm and gripper) maintain structural integrity and a physically plausible configuration throughout the video, with no deformation of rigid links or gripper, and realistic joint rotations? options: (A) yes (B) no (C) unclear ground truth: A Question 4: Do all task-relevant objects maintain their structural integrity without undergoing physically implausible deformations? options: (A) yes (B) no (C) unclear ground truth: A Question 5: Do any of the task-relevant objects exhibit physically implausible motions, such as sudden spatial displacement? options: (A) yes (B) no (C) unclear ground truth: A