Paper deep dive
DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning
Yang Zhou, Xiaofeng Wang, Hao Shao, Letian Wang, Guosheng Zhao, Jiangnan Shao, Jiagang Zhu, Tingdong Yu, Zheng Zhu, Guan Huang, Steven L. Waslander
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/2/2026, 11:56:44 PM
Summary
DriveDreamer-Policy is a unified world-action model for autonomous driving that integrates depth generation, future video generation, and motion planning. By utilizing a large language model (LLM) to process multi-view images and instructions, it generates geometry-aware world embeddings that guide lightweight generative experts. This approach improves planning robustness and future prediction quality by explicitly modeling 3D geometric structure, achieving state-of-the-art performance on Navsim v1 and v2 benchmarks.
Entities (4)
Relation Signals (3)
DriveDreamer-Policy â evaluatedon â NAVSIM
confidence 100% · Experiments on the Navsim v1 and v2 benchmarks demonstrate that DriveDreamer-Policy achieves strong performance
DriveDreamer-Policy â usesarchitecture â Qwen3-VL-2B
confidence 95% · For the large language model, we use Qwen3-VL-2B
DriveDreamer-Policy â generates â Depth Map
confidence 90% · we propose DriveDreamer-Policy, a unified driving world-action model that jointly generates 1) a depth-based 3D geometric representation
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recently, world-action models (WAM) have emerged to bridge vision-language-action (VLA) models and world models, unifying their reasoning and instruction-following capabilities and spatio-temporal world modeling. However, existing WAM approaches often focus on modeling 2D appearance or latent representations, with limited geometric grounding-an essential element for embodied systems operating in the physical world. We present DriveDreamer-Policy, a unified driving world-action model that integrates depth generation, future video generation, and motion planning within a single modular architecture. The model employs a large language model to process language instructions, multi-view images, and actions, followed by three lightweight generators that produce depth, future video, and actions. By learning a geometry-aware world representation and using it to guide both future prediction and planning within a unified framework, the proposed model produces more coherent imagined futures and more informed driving actions, while maintaining modularity and controllable latency. Experiments on the Navsim v1 and v2 benchmarks demonstrate that DriveDreamer-Policy achieves strong performance on both closed-loop planning and world generation tasks. In particular, our model reaches 89.2 PDMS on Navsim v1 and 88.7 EPDMS on Navsim v2, outperforming existing world-model-based approaches while producing higher-quality future video and depth predictions. Ablation studies further show that explicit depth learning provides complementary benefits to video imagination and improves planning robustness.
Tags
Links
- Source: https://arxiv.org/abs/2604.01765v1
- Canonical: https://arxiv.org/abs/2604.01765v1
Trouble viewing inline? Open PDF directly â
Full Text
58,967 characters extracted from source content.
Expand or collapse full text
2026-4-3 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning Yang Zhou 1,2 Xiaofeng Wang 1 Hao Shao 3 Letian Wang 2 Guosheng Zhao 1 Jiangnan Shao 1 Jiagang Zhu 1 Tingdong Yu 1 Zheng Zhu 1,* Guan Huang 1 Steven L. Waslander 2 1 GigaAI 2 University of Toronto 3 CUHK MMLab * Corresponding Author Project Website: https://drivedreamer-policy.github.io/ Abstract Recently, worldâaction models (WAM) have emerged to bridge visionâlanguageâaction (VLA) models and world models, unifying their reasoning and instruction-following capabilities and spatio-temporal world modeling. However, existing WAM approaches often focus on modeling 2D appearance or latent representations, with limited geometric grounding-an essential element for embodied systems operating in the physical world. We present DriveDreamer-Policy, a unified driving worldâaction model that integrates depth generation, future video generation, and motion planning within a single modular architecture. The model employs a large language model to process language instructions, multi-view images, and actions, followed by three lightweight generators that produce depth, future video, and actions. By learning a geometry-aware world representation and using it to guide both future prediction and planning within a unified framework, the proposed model produces more coherent imagined futures and more informed driving actions, while maintaining modularity and controllable latency. Experiments on the Navsim v1 and v2 benchmarks demonstrate that DriveDreamer-Policy achieves strong performance on both closed-loop planning and world generation tasks. In particular, our model reaches 89.2 PDMS on Navsim v1 and 88.7 EPDMS on Navsim v2, outperforming existing world-model-based approaches while producing higher-quality future video and depth predictions. Ablation studies further show that explicit depth learning provides complementary benefits to video imagination and improves planning robustness. 1. Introduction Autonomous driving systems have undergone a paradigm shift from handcrafted modular stacks (Thrun et al., 2006) to end-to-end learning systems (Bojarski et al., 2016; Hu et al., 2023), and more recently to vision-language-action (VLA) models built on large language models (LLMs) that offer richer common-sense knowledge, instruction following, and reasoning capabilities (Zhou et al., 2025; Kirby et al., 2026; Zhou et al., 2025; Li et al., 2025). Despite these advances, most VLA planners primarily optimize action outputs and do not explicitly model how the future world may evolve under alternative actions, which limits interpretability and can reduce reliability in rare or safety-critical situations where forward-looking reasoning about occlusions and hidden hazards is essential. In parallel, driving world models have emerged as a powerful direction for learning spatio-temporal traffic dynamics directly from large-scale sensor logs (Gao et al., 2024; Russell et al., 2025; Hassan et al., 2024; Mousakhan et al., 2025; Li et al., 2025; Bartoccioni et al., 2025; Zhou et al., 2026; © 2026 GigaAI. All rights reserved. arXiv:2604.01765v1 [cs.CV] 2 Apr 2026 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning Perception Vision-based Planner Lidar e.g., Transfuser, UniAD Image Scene Description VLA Planner e.g., AutoVLA, DriveVLA-W0 Image Action World Model Action e.g., Vista, GEM Action LanguageImage World-Action Model (3D+2D) Action Action Depth ImageLanguage DriveDreamer-Policy (Ours) World-Action Model (2D) Action Video e.g., PWM, UniUGP Image ActionVideoAction Language Scene Description Video Figure 1: Comparison of our DriveDreamer-Policy with existing models. Items with dashed lines are optional. Vision-based and VLA planners directly map observations (and optional inputs) to actions without explicitly predicting the future world. World models generate future observations but often rely on external action signals. Recent worldâaction models unify future world generation and planning, but typically operate on image/video representations. DriveDreamer-Policy extends this line by explicitly generating depth alongside video and actions, enabling geometry-grounded imagination and planning within a unified model. Zhao et al., 2026; Ye et al., 2026). By forecasting future observations (e.g., videos, BEV features, or other state representations), they enable scalable simulation, rare-event evaluation, and controllable synthetic data generation. Motivated by the complementary strengths of these two lines, recent world-action models aim to unify future world generation with motion prediction or planning (Zhang et al., 2025; Xia et al., 2025; Lu et al., 2025; Zhao et al., 2025), thereby bridging imagination and decision-making in a single framework. However, in many existing world-action models, the world component is still implemented as image/video prediction or latent rollouts without explicit geometric grounding, and the benefits of imagined futures for action prediction can be constrained by representation mismatch or tightly coupled module designs. As a result, the generated world may be visually plausible but not maximally informative for planning, and the planner may not consistently benefit from structured safety cues such as geometric layout and free-space constraints. A key observation motivating this work (Fig. 1) is that autonomous driving is fundamentally a 4D physical process: 3D geometry evolves over time. Consequently, an actionable world model should not only synthesize appearance, but also preserve geometric structure that is essential for occlusion reasoning, distance estimation, and physically consistent motion. Depth-centric modeling is particularly attractive here: depth is compact, directly tied to geometry, and can serve as an explicit scaffold that constrains future image/video generation and informs planning decisions. Moreover, recent progress in depth foundation models (Yang et al., 2024; Lin et al., 2025; Piccinelli et al., 2024, 2026; Xu et al., 2025) suggests that high-fidelity depth can be produced off- the-shelf, without collecting extra data or training a depth estimator from scratch. These developments suggest an opportunity to more effectively drive world-action models: explicitly generating depth representations and studying how this benefits both future video generation and motion planning within a unified architecture. To this end, we propose DriveDreamer-Policy, a unified driving world-action model that jointly generates 1) a depth-based 3D geometric representation of the current scene, 2) action-conditioned future videos, and 3) future trajectories for planning. The system is built on a large language model for perception and reasoning, producing a compact set of world embeddings and action embeddings. These embeddings serve as conditions for the multimodal generator: a pixel-space depth generator, a latent-space video generator, and an action generator. Importantly, we impose a structured causal attention mask across query groups in a depthâvideoâaction manner: video queries may consume depth context, and action queries may consume both depth and video context. This yields a simple, single-pass information flow, while enabling video imagination to benefit from 3D understanding and allow planning to leverage both 3D structure and predicted future-world context. Our contributions are threefold. 1) We introduce DriveDreamer-Policy, a unified, modular world-action architecture for autonomous driving that combines an LLM with generative experts connected through a fixed- size query interface, enabling practical compute control. 2) We incorporate an explicit 3D depth generation module and utilize a causal 3Dâ2Dâ1D conditioning pathway, allowing geometry to directly scaffold future 2 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning video generation and motion planning and supporting multimodal data generation/simulation. (3) We conduct comprehensive experiments and analyses on Navsim (Dauner et al., 2024; Cao et al., 2025), evaluating both planning performance and world generation quality, achieving state-of-the-art compared to existing world-based models. For instance, our model achieves an EPDMS score of 88.7 (+2.6 over the previous method) and an FVD of 53.59 (-32.36 over the previous method), improving both planning performance and video generation quality. We also perform targeted ablations to quantify how depth and future imagination jointly contribute to planning robustness. 2. Related Works 2.1. Driving World-Action Models Driving-focused generative models use sensor data, such as images, to generate future videos, thereby enabling scalable data synthesis and simulation (Hassan et al., 2024; Mousakhan et al., 2025; Bartoccioni et al., 2025; Wang et al., 2024; Zhao et al., 2025; Agarwal et al., 2025; Liang et al., 2025; NVIDIA et al., 2025; Ni et al., 2025; Lu et al., 2025; Zhao et al., 2025; Team et al., 2025). Recently, driving world action models that unify generation and planning have emerged as an active frontier. Epona (Zhang et al., 2025) introduces an autoregressive diffusion world model that decouples causal temporal latents from per-step diffusion generation to support long-horizon video rollout and trajectory planning. ReSim (Yang et al., 2025) trains a diffusion-transformer world simulator on real logs plus simulator non-expert behaviors to improve action-following reliability and adds Video2Reward for reward estimation. DriveVLA-W0 (Li et al., 2025) combats the VLA supervision deficit by adding future-image world modeling and using a lightweight MoE action expert for lower latency. PWM (Zhao et al., 2025) treats a unified autoregressive transformer as a Policy World Model that performs action-free future forecasting and collaborative state-action prediction to benefit planning. DriveLaW (Xia et al., 2025) unifies planning and generation by feeding video-generator latents into a diffusion trajectory planner to align imagined futures with control. OmniNWM (Li et al., 2025) jointly generates panoramic RGB, semantics, depth, and 3D occupancy, conditions trajectories via Plucker ray-maps, and derives intrinsic occupancy-based dense rewards. UniPGT (Lu et al., 2025) unifies understanding, video generation, and trajectory planning by integrating a pretrained VLM with a video generator through hybrid experts. Similar to existing world-action models, DriveDreamer-Policy uses an LLM to model driving-world knowledge, serving as the perception module. To incorporate multimodal generators, fixed-size latent queries are used as cross-attention keys, enabling generative experts to jointly predict depth, video, and action for the first time. 2.2. Driving Vision-Language-Action Models The use of VLMs in autonomous driving has gradually evolved from scene interpretation to direct action genera- tion. Early works such as DriveGPT4 (Xu et al., 2024) primarily used LLMs/VLMs to generate scene descriptions and high-level maneuver suggestions, thereby improving interpretability. To bridge language understanding and low-level control, a line of modular language-to-action frameworks (Zhou et al., 2025; Arai et al., 2025) introduced multi-stage pipelines that pass intermediate textual commands between modules; however, these non-differentiable interfaces hinder end-to-end optimization by blocking gradient backpropagation across perception, reasoning, and control. More recent end-to-end VLA methods (Yang et al., 2025; Zeng et al., 2025; Team et al., 2025) adopt unified architectures that directly map sensor observations to trajectories. Within this paradigm, DriveMoE (Yang et al., 2025) introduces a Mixture-of-Experts (MoE) design with an action decoder attached after a VLM, ReCogDrive (Li et al., 2025) couples a VLM with a diffusion planner trained by imitation and reinforcement learning to better align semantic reasoning and control, and AutoVLA (Zhou et al., 2025) discretizes trajectories into action primitives so that a single autoregressive model can jointly learn adaptive reasoning and planning. DriveVLA-W0 (Li et al., 2025) scales driving VLA learning by adding dense world modeling (future image prediction) as supervision and introducing a lightweight MoE action expert to reduce inference cost and improve latency. Our method is similar to this end-to-end VLA direction and further extends it with a unified world-action modeling design. DriveDreamer-Policy follows driving VLA models but replaces 3 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning image-token world prediction with a diffusion video generation head, enabling controllable action-conditioned rollouts. We further add a depth-based 3D representation to resolve missing-geometry grounding. 3. Methodology We propose DriveDreamer-Policy, a unified driving worldâaction model that couples a large language model with lightweight generative experts to jointly support 1) 3D world representation, 2) 2D world generation, and 3) motion planning. We introduce world understanding with large language models and world-action prediction using generative experts in Sec. 3.1 and Sec. 3.2, then provide training details in Sec. 3.3. 3.1. Preliminaries Flow Matching. We use conditional flow matching (Lipman et al., 2023) as the training principle for our generative experts when predicting continuous targets. Letí„â R í denote a target variable, and letídenote conditioning information. Flow matching learns a time-dependent velocity fieldíŁ í (í„ íĄ ,íĄ|í)that transports a simple base distribution toward the data distribution along a predefined path. We adopt the standard linear interpolation path between a data sample í„ 0 ⌠í data and a noise sample í„ 1 ⌠í noise : í„ íĄ = (1â íĄ)í„ 0 + íĄí„ 1 , íĄâŒí°(0, 1),(1) with the target velocity given by Ìí„ íĄ = í„ 1 â í„ 0 . The objective is a regression loss on the velocity: â FM = E í„ 0 ,í„ 1 ,íĄ [ïž âíŁ í (í„ íĄ ,íĄ|í)â (í„ 1 â í„ 0 )â 2 2 ]ïž .(2) At inference, we sampleí„ 1 ⌠í noise and integrate the induced ODE backward fromíĄ=1toíĄ=0to obtain a sample consistent with the conditioning í. 3.2. DriveDreamer-Policy Our overall pipeline is illustrated in Fig. 2. Multi-view images, language instructions, and the action are first encoded as tokens and processed by an LLM, together with a compact set of learned world and action queries. The resulting world embeddings and action embeddings serve as a geometry-aware interface that conditions three modular experts: a depth generator, a video generator, and an action generator. The LLM is responsible for multimodal understanding and producing compact state representations, while the experts generate modality-specific outputs (depth, video, and action, respectively), all mediated by a fixed-size query bottleneck. This design is motivated by the complementary strengths of the two components: the LLM provides stable semantics and strong contextual reasoning, whereas generative experts better capture multi-modality and uncertainty in long-horizon prediction. As a result, the model can operate in multiple modes: planning-only (enabling only the action expert), imagination-enabled planning (running action plus depth/video generation when needed), or full generation for offline simulation and data synthesis. 3.2.1. World Understanding Input Processing. At each decision step, the model takes as input a natural-language instruction and syn- chronized multi-view RGB observations. We also provide the current action as context to the LLM, thereby contributing to both world modeling and planning. We tokenize the inputs into three streams. First, the instruction is converted into standard text tokens using the LLMâs tokenizer. Second, each camera view is encoded by the vision encoder into a sequence of visual patch tokens. Third, the action context is embedded into a set of action tokens using a lightweight action encoder. Finally, we append three fixed-size groups of learnable query tokensâdepth queries, video queries, and action queriesâin that order. This design yields a stable, compact interface: the backbone always consumes the same set of query slots, and downstream heads can read out the corresponding query embeddings to produce depth maps, future videos, and future actions. 4 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning Large Language Model Multi-view Images WorldQueries âkeep straightâ Action Encoder Vision Encoder Te x t To ke n i z e r Action ActionQueries World Embeddings GT Diffusion Transformer Training Only Add Noise Image Depth Generator Current Depth Concat Diffusion Transformer Noised Latents Image VideoGenerator Future Video VAE VAE Clip Model Concat Diffusion Transformer Noise Future Action ActionGenerator Action Embeddings Depth Queries Video Queries Figure 2: Overview of our DriveDreamer-Policy pipeline. The large language model takes the language instruction, multi-view images and current action, along with a set of learnable queries as inputs to reason and generate world and action embeddings. The generated embeddings are then passed into our three generative expert models as cross-attention conditions to generate depth, future images, and future action. Embeddings Generation. All tokens are processed by the LLM to produce contextualized hidden states. We impose a causal ordering across query groups: depth queries come first, video queries can attend to depth context, and action queries can attend to both depth and video context. Concretely, within the same time step the attention pattern satisfies depth queriesâvideo queriesâaction queries. This structured mask provides a clean, single-pass information flow without extra synchronization or iterative refinement across branches. 3.2.2. World and Action Prediction Depth Generator. We generate a monocular depth map as an explicit 3D scaffold for the world action model. Depth is not only a geometric output for visualization: it provides a compact representation that is directly useful for downstream video imagination (e.g., occlusions and object boundaries) and action planning (e.g., free space and distance-to-collision cues). Our model predicts depth with a generative objective rather than a purely deterministic regression head, thereby better capturing the inherent ambiguity of monocular depth and preserving sharp depth discontinuities. Generating depth directly in pixel-space is practical here because depth has much lower dimensionality than RGB video and preserves boundary fidelity without requiring an additional learned codec. As shown in the upper-left of Fig. 2, our depth generator is a pixel-space diffusion transformer trained with a standard flow-matching objective. During training, we sample a continuous flow time and corrupt the ground-truth depth with noise; the denoiser takes as input the concatenation of the noisy depth and the corresponding RGB image, and predicts the denoising update. To ground pixel-space generation in global scene semantics, we condition the depth denoiser on the LLM world depth embeddings through cross-attention: the depth-query world embedding acts as a compact global representation (keys/values) that guides the diffusion transformer to maintain global structural consistency while recovering fine-grained geometric details. This makes depth a queryable modality in DriveDreamer-Policy: the predicted depth can be generated on demand with depth embeddings, and the depth embeddings serve as an upstream geometric feature that later query groups (video/action) can attend to. Video Generator. For future video generation, we employ a text-image-to-video diffusion transformer (Peebles and Xie, 2023; Wan et al., 2025) (see the upper-middle of Fig. 2). Given the current RGB images, we first encode them into a compact latent representation using a VAE and initialize a sequence of noisy video latents 5 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning for the target horizon. Instead of conditioning the diffusion model on text embedding (as in standard text-to- video pipelines), we condition it on the LLM world video embeddings produced by the video query tokens. These world embedding tokens summarize language intent, multi-view perception, and action context, and extrinsically incorporate upstream geometric cues from the depth queries. The video denoiser attends to the world embeddings at each transformer block via cross-attention. To preserve appearance, identity, and camera content, we additionally extract a lightweight visual condition from the current image frame using a CLIP (Radford et al., 2021) model and inject it into the denoiser as an explicit conditioning signal, concatenated with the world video embeddings as shown in our pipeline. This design keeps the generator tightly grounded in the current scene and commanded maneuver, while enabling controllable, action-aware video generation. Action Generator. As presented in the upper-right of Fig. 2, the action generator is implemented as a standalone diffusion transformer that maps a noise trajectory to a feasible future action sequence. It is conditioned on the action embedding produced by the LLM from the action query tokens, which aggregates instruction semantics, multi-view observations, and upstream geometric and imagination cues. This conditioning is injected via cross-attention, thereby keeping the action head lightweight while still leveraging rich scene context. Since the action generator does not depend on explicit depth and video generation, it can be activated independently for planning, while implicitly benefiting from predicted future world context. We parameterize each trajectory state by position and heading using a continuous representation (Zhou et al., 2024), namely (í„,íŠ, cosí, siní), which avoids angular wrap-around and encourages smooth turn dynamics. 3.3. Training Details Depth Normalization. We normalize depth to a stable range before training the depth generator. Given a depth map, we first apply a log transform and then compute per-map percentiles to normalize to the range [-0.5, 0.5]. During inference, we invert the transform to recover metric or relative depth as needed. Model Initialization and Adaptation. For the large language model, we use Qwen3-VL-2B (Bai et al., 2025) to process and understand multimodal inputs. For the depth generator, we initialize our model from PPD (Xu et al., 2025). For the video generator, we initialize the model from Wan-2.1-T2V-1.3B (Wan et al., 2025) and adapt it to the image-to-video task. For both depth and video generation, we fine-tune at a spatial resolution of 144Ă 256 to reduce computational and memory costs. The video training horizon is 9 frames. Training Objective and Optimization. We train all components in a single stage with a joint multi-task loss: â = í í â í + í íŁ â íŁ + í í â í ,(3) whereâ í is the loss for depth prediction,â íŁ is the loss for video prediction, andâ í is the loss for trajectory prediction. We useí í = 0.1and set the remaining hyperparameters to 1.0 by default. The depth label used in training is obtained from an off-the-shelf depth foundation model, Depth Anything 3 (DA3) (Lin et al., 2025). 4. Experiments 4.1. Experimental Setup Datasets and Planning Metrics. We train and evaluate our method on the Navsim benchmark (Dauner et al., 2024; Cao et al., 2025), which is derived from real-world driving logs and provides synchronized surround-view sensory inputs for end-to-end planning evaluation. Following the standard Navsim protocol, we train on the navtrainsplit and evaluate on thenavtestsplit, which contains 100k and 12k data samples sampled at 2Hz, respectively. Navsim evaluates closed-loop planning performance using the predictive driver model score (PDMS) on v1 and the extended PDMS (EPDMS) on v2. PDMS aggregates multiple safety and quality terms, including no-at-fault collision, drivable-area compliance, time-to-collision, ego progress, and comfort; EPDMS further includes direction and traffic-light compliance, as well as lane-keeping and comfort. For Navsim-v2, we 6 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning Table 1: Comparison with state-of-the-art methods on the thenavtestof Navsim v1 (Dauner et al., 2024) benchmark. The best result outlines as bold. â*â indicates results with imitation learning. Methods fall into three categories: vision-based E2E models, VLA-based models, and world-model-based models. MethodsVenueSensorsNCâ DACâ TTCâCâEPâPDMSâ Humanâ100.0 100.0 100.0 99.9 87.594.8 Vision-Based End-to-End Methods TransFuser (Chitta et al., 2022)TPAMIâ233ĂC + L97.7 92.8 92.8 100.0 79.284.0 UniAD (Hu et al., 2023)CVPRâ236ĂC97.8 91.9 92.9 100.0 78.883.4 PARA-Drive (Weng et al., 2024)CVPRâ246ĂC97.9 92.4 93.0 99.8 79.384.0 DiffusionDrive (Liao et al., 2025) CVPRâ253ĂC + L98.2 96.2 94.7 100.0 82.288.1 Vision-Language-Action Methods AutoVLA (Zhou et al., 2025)NeurIPSâ253ĂC98.4 95.6 98.0 99.9 81.989.1 Recogdrive * (Li et al., 2025)ICLRâ263ĂC98.1 94.7 94.2 100.0 80.986.5 DriveVLA-W0 (Li et al., 2025)ICLRâ26 1ĂC98.7 96.2 95.5 100.0 82.288.4 World-Model-Based Methods LAW (Li et al., 2025)ICLRâ25 1ĂC96.4 95.4 88.7 99.9 81.784.6 DrivingGPT (Chen et al., 2025)ICCVâ25 1ĂC98.9 90.7 94.9 95.6 79.782.4 WoTE (Li et al., 2025)ICCVâ25 3ĂC + L98.5 96.8 94.4 99.9 81.988.3 Epona (Zhang et al., 2025)ICCVâ25 3ĂC97.9 95.1 93.8 99.9 80.486.2 FSDrive (Zeng et al., 2025)NeurIPSâ25 3ĂC98.2 93.8 93.3 99.9 80.185.1 PWM (Zhao et al., 2025)NeurIPSâ25 1ĂC98.6 95.9 95.4 100.0 81.888.1 DriveDreamer-Policy (Ours)â3ĂC98.497.195.1100.083.589.2 Table 2: Comparison with state-of-the-art methods on the navtest of Navsim v2 (Cao et al., 2025). MethodsVenueNCâ DACâ DDCâ TLCâ EPâ TTCâ LKâ HCâ ECâEPDMSâ Vision-Based End-to-End Methods TransFuser (Chitta et al., 2022)TPAMIâ2396.9 89.9 97.8 99.7 87.1 95.4 92.7 98.3 87.276.7 DiffusionDrive (Liao et al., 2025) CVPRâ25 98.2 95.9 99.4 99.8 87.5 97.3 96.8 98.3 87.784.5 Drivesuprim (Yao et al., 2025)AAAIâ26 97.5 96.5 99.4 99.6 88.4 96.6 95.5 98.3 77.083.1 ARTEMIS (Feng et al., 2026)RALâ26 98.3 95.1 98.6 99.8 81.5 97.4 96.5 98.3 89.183.1 Vision-Language-Action Methods DriveVLA-W0 (Li et al., 2025)ICLRâ2698.5 99.1 98.0 99.7 86.4 98.1 93.2 97.9 58.986.1 World-Model-Based Methods DriveDreamer-Policy (Ours)â98.497.199.599.987.997.797.698.379.488.7 Table 3: World generation performance on Navsim. For video generation, we compare against existing generative world-model methods trained and evaluated on Navsim. For depth generation, we report comparisons among our model variants. Our method achieves higher quality for depth and video. (a) Video performance comparison. MethodsVenueLPIPSâ PSNRâFVDâ PWM (Zhao et al., 2025)NeurIPSâ250.2321.5785.95 DriveDreamer-Policy (Ours)â0.2021.0553.59 (b) Depth performance comparison. MethodsVenueAbsRelâíż 1 âíż 2 âíż 3 â PPD (Xu et al., 2025)NeurIPSâ2518.580.4 94.0 97.2 PPD-FintunedNeurIPSâ259.391.4 98.3 99.5 DriveDreamer-Policy (Ours)â8.192.898.699.5 align with the common practice in recent methods (Li et al., 2025; Liao et al., 2025) and evaluate EPDMS on the navtest split for fair and convenient comparison. World Generation Metrics. In addition to planning, we evaluate our generative experts. Video is evaluated on Navim using the recorded future RGB frames as ground truth. Depth is evaluated using dense depth targets provided by DA3, which are also used for training. We report Absolute Relative Error (AbsRel) to quantify relative depth differences and threshold accuracy (íż) to measure the proportion of accurate predictions within a specified relative error. Higheríżvalues and lower AbsRel values indicate better depth estimation performance. For video evaluation, we follow (Zhao et al., 2025) and report perceptual quality and temporal consistency of predicted future frames using Learned Perceptual Image Patch Similarity (LPIPS) (Zhang et al., 2018), Peak Signal-to-Noise Ratio (PSNR) (Huynh-Thu and Ghanbari, 2008) and FrĂ©chet Video Distance (FVD) (Unterthiner et al., 2019). Baselines. We compare against strong Navsim baselines spanning three families. 1) Classical vision-based end- 7 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning Table 4: Ablations on World Learning for Planning. All world learning strategies effectively improve planning performance compared to training from scratch. Strategy World RepresentationPlanning Metrics DepthVideo NCâ DACâ TTCâCâEPâPDMSâ Without World Learning %%98.0 96.3 94.4 100.0 82.588.0 With World Learning "%98.1 96.7 94.9 100.0 82.888.5 %"98.1 97.0 95.0 100.0 83.188.9 ""98.4 97.1 95.1 100.0 83.589.2 Table 5: Ablations on Depth Learning for Video Generation. Using depth as a prior in joint learning improves video generation accuracy. Strategy World RepresentationVideo Metrics Depth LPIPSâ PSNRâ FVDâ Without Depth Learning %0.2219.89 65.82 With Depth Learning "0.20 21.05 53.59 Table 6: Ablations on Number of Queries. More query tokens provide higher-capacity slots for storing relevant context, thereby enhancing both generation and planning. Number of QueriesMetrics Depth Video Action DepthVideoAction AbsRelâ íż 1 â íż 2 â íż 3 âLPIPSâ PSNRâ FVDâNCâ DACâ TTCâCâEPâPDMSâ 323249.790.2 97.9 99.40.2020.67 57.9798.2 97.0 95.0 100.0 83.288.9 646488.192.8 98.6 99.50.20 21.05 53.5998.4 97.1 95.1 100.0 83.589.2 to-end planners that utilize vision models and map sensor inputs to trajectories, including TransFuser (Chitta et al., 2022), UniAD (Hu et al., 2023) and DiffusionDrive (Liao et al., 2025). 2) Vision-language-action planners that use large language models and predict trajectories as tokens or by a diffusion-based expert, including approaches: DriveVLA-W0 (Li et al., 2025), AutoVLA (Zhou et al., 2025) and Recogdrive (Li et al., 2025). 3) World-model-based planners that integrate foresight into planning, including LaW (Li et al., 2025), DrivingGPT (Xu et al., 2024), WoTE (Li et al., 2025), Epona (Zhang et al., 2025), FSDrive (Zeng et al., 2025) and PWM (Zhao et al., 2025). All baselines are reported under their official Navsim performance. Implementation Details. We implement the action encoder as a 2-layer MLP with layer normalization (Ba et al., 2016). We train DriveDreamer-Policy in a single stage for 100k steps with a batch size of 32 on 8 NVIDIA H20 GPUs, using AdamW optimizer (Loshchilov and Hutter, 2019) with a learning rate of1Ă10 â5 . Unless otherwise stated, all experiments use the same query configuration (64 depth-query tokens, 64 video-query tokens, and 8 action-query tokens). We use Navsim training data, without additional datasets or extra pre-training beyond the initialized backbones. 4.2. Quantitative Results Planning Performance Comparison with Other Methods. As shown in Table 1 and Table 2, we first report planning performance compared with other methods, on thenavtestset of Navsim v1 using PDMS and Navsim v2 using EPDMS. For a fair and meaningful comparison, we consider only methods that were published at the time of this paperâs submission. Specifically, we categorize comparisons against state-of-the-art vision-based end-to-end planners, vision-language-action planners, and world-model-based models. Our DriveDreamer- Policy outperforms all considered methods and achieves PDMS scores of 89.2 and 88.7 on Navsim v1 and v2, respectively. Moreover, DriveDreamer-Policy achieves strong performance on specific planning-critic sub-scores, 8 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning Keep Straight Planning Human Pred Left ViewFront ViewRight View Depth Video First Frame Video Last Frame Keep Straight Planning Human Pred Left ViewFront ViewRight View Depth Video First Frame Video Last Frame Figure 3: Visualization Results of our method. We show the generated depth, video, and actions, respectively. Depth is truncated to below 80 meters for better visualization. Our generation results remain spatially stable, and the planning performs well compared with human trajectories (e.g., aligns with human trajectories (top) and slows down more effectively than human trajectories (bottom)). indicating our methodâs control of planning trajectories (e.g., a DAC of 97.1 and an EP of 83.5 on Navsim v1, and a DDC of 99.5 and an LK of 97.6 on Navsim v2). World Performance Comparison with Other Methods. We also compare with other generative world-based methods. For video quality, we compare with the existing method PWM (Zhao et al., 2025). Since PWM supports only single-view generation, we evaluate the single-view (front) quality for a fair and convenient comparison. As presented in Table 3(a), our method shows a larger improvement than PWM by a substantial margin (e.g., improvement of 32.36 on FVD). For depth accuracy, we compare against PPD, because our depth generator is initialized from it. To the best of our knowledge, there is no widely adopted Navsim benchmark that reports directly comparable results for depth prediction. Specifically, we evaluate two variants: (i) zero-shot PPD on Navsim and (i) fine-tune PPD on Navsim. As shown in Table 3(b), DriveDreamer-Policy achieves lower depth error. We attribute the improvement to LLM conditioning: the depth denoiser is guided by world depth embeddings, resolving locally ambiguous regions and producing more consistent geometry than image cues alone. Ablations on World Learning for Planning. To isolate the contribution of each modality, we evaluate four 9 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning Keep Straight Action-Only Human Pred Keep Straight Depth-Action Human Pred Keep Straight Video-Action Human Pred Keep Straight Depth-Video-Action Human Pred Turn Left Action-Only Human Pred Turn Left Depth-Action Human Pred Turn Left Video-Action Human Pred Turn Left Depth-Video-Action Human Pred Turn Left Action-Only Human Pred Turn Left Depth-Action Human Pred Turn Left Video-Action Human Pred Turn Left Depth-Video-Action Human Pred Figure 4: Visualization of world learning for planning. Columns compare Action-Only, Depth-Action, Video- Action, and Depth-Video-Action variants. Green denotes the human (expert) trajectory and red denotes the predicted trajectory. The three rows correspond to (top) avoiding potential collision by a slower trajectory, (middle) correcting an initially wrong maneuver, and (bottom) aligning more closely with the human trajectory. Depth and video provide complementary world cues that improve safety margins and trajectory consistency. variants under identical training budgets: 1) action-only: the simplest VLA version without any world modeling, 2) depth+action: generates depth and future action, 3) video+action: generates video and future action, and 4) depth+video+action: our full model. We report improvements in PDMS for each variant. The results are shown in Table 4. All world training strategies effectively improve planning performance compared to training from scratch. However, single-world training (depth or video) yields smaller improvement, likely due to incomplete world information. Joint world training, leveraging both depth geometry and video temporal evolution, achieves the largest performance gains by enabling the model to learn more generalizable and robust features. Ablations on Depth Learning for Video Generation. The impact of depth learning on future video generation is evaluated by comparing two variants under the same training data and compute budget: 1) video-only, where the video generator is conditioned on the backbone features without depth joint learning; 2) depth+video, where depth is trained jointly and the video queries are causally conditioned on depth queries. Video quality is reported in Table 5. Joint learning with depth improves video generation accuracy, indicating that depth provides an effective 3D scaffold for coherent future prediction. 10 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning 4.3. Ablation Studies Ablations on Number of Queries. We ablate the query budget to study the trade-off between accuracy. Specifically, we compare the default setting (64 depth + 64 video + 8 action query tokens) against smaller budgets (32 depth + 32 video + 4 action query tokens). The results, presented in Table 6, indicate that increasing the query budget generally improves both planning and world-generation performance, as more query tokens provide higher-capacity slots to store geometry, appearance, and action-relevant context. 4.4. Qualitative Results Visualizations of World generation and Motion Planning. We visualize representative Navsim scenarios in Fig. 3 by overlaying predicted trajectories on BEV renderings and by showing the corresponding generated depth maps and future frames. These examples highlight how depth-conditioned imagination reduces common failure modes (e.g., short-horizon collision risk, off-road drift) and improves interpretability by exposing the modelâs predicted geometry and future scene evolution. Visualizations of World Learning for Planning. We present additional visualization results to demonstrate the effectiveness of world learning in benefiting planning. Fig. 4 presents qualitative comparisons that demonstrate how learned world representations improve planning. Each row visualizes a representative Navsim scenario, and each column corresponds to an ablation variant: Action-Only, Depth-Action, Video-Action, and Depth- Video-Action. Across the three cases, adding world learning consistently produces safer and more human-like trajectories. In the top-row example, action-only planning tends to drift toward conflicting traffic, whereas depth- and video-conditioned variants maintain a clearer safety margin. In the middle row example, world learning helps recover the correct maneuver early, reducing lateral deviation and preventing late, abrupt turns. In the bottom-row example, the full model yields trajectories that more closely match the expert path, suggesting that geometry cues (depth) and future appearance/dynamics cues (video) provide complementary guidance for action prediction. 5. Conclusion We present DriveDreamer-Policy, a unified driving worldâaction model that jointly performs depth generation, future video imagination, and motion planning within a single framework. The model combines a large language model with modular generative experts connected through a compact query interface, enabling flexible operating modes for both planning and world generation. By introducing depth as an explicit geometric scaffold and organizing information flow in a depthâvideoâaction manner, the model allows planning to leverage both scene geometry and predicted future dynamics. Experiments on Navsim demonstrate strong performance across planning and world generation tasks, and ablations show that depth and video provide complementary cues that improve planning robustness. 11 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning References [1]Sebastian Thrun, Mike Montemerlo, Hendrik Dahlkamp, David Stavens, Andrei Aron, James Diebel, Philip Fong, John Gale, Morgan Halpenny, Gabriel Hoffmann, Kenny Lau, Celia Oakley, Mark Palatucci, Vaughan Pratt, Pascal Stang, Sven Strohband, Cedric Dupont, Lars-Erik Jendrossek, Christian Koelen, Charles Markey, Carlo Rummel, Joe van Niekerk, Eric Jensen, Philippe Alessandrini, Gary Bradski, Bob Davies, Scott Ettinger, Adrian Kaehler, Ara Nefian, and Pamela Mahoney. Stanley: The robot that won the darpa grand challenge: Research articles. J. Robot. Syst., 23(9):661â692, September 2006. ISSN 0741-2223. 1 [2]Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D. Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba. End to end learning for self-driving cars, 2016. URL https://arxiv.org/abs/1604.07316. 1 [3] Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, Lewei Lu, Xiaosong Jia, Qiang Liu, Jifeng Dai, Yu Qiao, and Hongyang Li. Planning-oriented autonomous driving, 2023. URL https://arxiv.org/abs/2212.10156. 1, 7, 8 [4]Xingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma, Volker Tresp, and Alois Knoll. Opendrivevla: Towards end-to-end autonomous driving with large vision language action model, 2025. URLhttps: //arxiv.org/abs/2503.23463. 1, 3 [5] Ellington Kirby, Alexandre Boulch, Yihong Xu, Yuan Yin, Gilles Puy, Ăloi Zablocki, Andrei Bursuc, Spyros Gidaris, Renaud Marlet, Florent Bartoccioni, Anh-Quan Cao, Nermin Samet, Tuan-Hung VU, and Matthieu Cord. Driving on registers, 2026. URL https://arxiv.org/abs/2601.05083. 1 [6]Zewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang, Zhiyu Huang, Bolei Zhou, and Jiaqi Ma. Au- tovla: A vision-language-action model for end-to-end autonomous driving with adaptive reasoning and reinforcement fine-tuning, 2025. URL https://arxiv.org/abs/2506.13757. 1, 3, 7, 8 [7]Yingyan Li, Shuyao Shang, Weisong Liu, Bing Zhan, Haochen Wang, Yuqi Wang, Yuntao Chen, Xiaoman Wang, Yasong An, Chufeng Tang, Lu Hou, Lue Fan, and Zhaoxiang Zhang. Drivevla-w0: World models amplify data scaling law in autonomous driving, 2025. URLhttps://arxiv.org/abs/2510.12796. 1, 3, 7, 8 [8] Shenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta, Yihang Qiu, Andreas Geiger, Jun Zhang, and Hongyang Li. Vista: A generalizable driving world model with high fidelity and versatile controllability. Advances in Neural Information Processing Systems, 37:91560â91596, 2024. 1 [9]Lloyd Russell, Anthony Hu, Lorenzo Bertoni, George Fedoseev, Jamie Shotton, Elahe Arani, and Gianluca Corrado. Gaia-2: A controllable multi-view generative world model for autonomous driving, 2025. URL https://arxiv.org/abs/2503.20523. 1 [10] Mariam Hassan, Sebastian Stapf, Ahmad Rahimi, Pedro M B Rezende, Yasaman Haghighi, David BrĂŒgge- mann, Isinsu Katircioglu, Lin Zhang, Xiaoran Chen, Suman Saha, Marco Cannici, Elie Aljalbout, Botao Ye, Xi Wang, Aram Davtyan, Mathieu Salzmann, Davide Scaramuzza, Marc Pollefeys, Paolo Favaro, and Alexandre Alahi. Gem: A generalizable ego-vision multimodal world model for fine-grained ego-motion, object dynamics, and scene composition control, 2024. URLhttps://arxiv.org/abs/2412.11198. 1, 3 [11]Arian Mousakhan, Sudhanshu Mittal, Silvio Galesso, Karim Farid, and Thomas Brox. Orbis: Overcoming challenges of long-horizon prediction in driving world models, 2025. URLhttps://arxiv.org/abs/ 2507.13162. 1, 3 [12] Bohan Li, Zhuang Ma, Dalong Du, Baorui Peng, Zhujin Liang, Zhenqiang Liu, Chao Ma, Yueming Jin, Hao Zhao, Wenjun Zeng, and Xin Jin. Omninwm: Omniscient driving navigation world models, 2025. URL https://arxiv.org/abs/2510.18313. 1, 3 12 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning [13]Florent Bartoccioni, Elias Ramzi, Victor Besnier, Shashanka Venkataramanan, Tuan-Hung Vu, Yihong Xu, Loick Chambon, Spyros Gidaris, Serkan Odabas, David Hurych, Renaud Marlet, Alexandre Boulch, Mickael Chen, Ăloi Zablocki, Andrei Bursuc, Eduardo Valle, and Matthieu Cord. Vavim and vavam: Autonomous driving through video generative modeling, 2025. URL https://arxiv.org/abs/2502.15672. 1, 3 [14]Yang Zhou, Hao Shao, Letian Wang, Zhuofan Zong, Hongsheng Li, and Steven L Waslander. Drivinggen: A comprehensive benchmark for generative video world models in autonomous driving. arXiv preprint arXiv:2601.01528, 2026. 1 [15] Guosheng Zhao, Yaozeng Wang, Xiaofeng Wang, Zheng Zhu, Tingdong Yu, Guan Huang, Yongchen Zai, Ji Jiao, Changliang Xue, Xiaole Wang, et al. Unidrivedreamer: A single-stage multimodal world model for autonomous driving. arXiv preprint arXiv:2602.02002, 2026. 2 [16]Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang, Guosheng Zhao, Hao Li, Hengtao Li, Jie Li, Jindi Lv, Jingyu Liu, et al. Gigaworld-policy: An efficient action-centered worldâaction model. arXiv preprint arXiv:2603.17240, 2026. 2 [17] Kaiwen Zhang, Zhenyu Tang, Xiaotao Hu, Xingang Pan, Xiaoyang Guo, Yuan Liu, Jingwei Huang, Li Yuan, Qian Zhang, Xiao-Xiao Long, Xun Cao, and Wei Yin. Epona: Autoregressive diffusion world model for autonomous driving. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025. 2, 3, 7, 8 [18]Tianze Xia, Yongkang Li, Lijun Zhou, Jingfeng Yao, Kaixin Xiong, Haiyang Sun, Bing Wang, Kun Ma, Guang Chen, Hangjun Ye, Wenyu Liu, and Xinggang Wang. Drivelaw:unifying planning and video generation in a latent driving world, 2025. URL https://arxiv.org/abs/2512.23421. 2, 3 [19] Hao Lu, Ziyang Liu, Guangfeng Jiang, Yuanfei Luo, Sheng Chen, Yangang Zhang, and Ying-Cong Chen. Uniugp: Unifying understanding, generation, and planing for end-to-end autonomous driving, 2025. URL https://arxiv.org/abs/2512.09864. 2, 3 [20]Zhida Zhao, Talas Fu, Yifan Wang, Lijun Wang, and Huchuan Lu. From forecasting to planning: Policy world model for collaborative state-action prediction, 2025. URLhttps://arxiv.org/abs/2510.19654. 2, 3, 7, 8, 9 [21]Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything v2, 2024. URL https://arxiv.org/abs/2406.09414. 2 [22]Haotong Lin, Sili Chen, Junhao Liew, Donny Y. Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth anything 3: Recovering the visual space from any views, 2025. URLhttps://arxiv.org/ abs/2511.10647. 2, 6 [23] Luigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis, Mattia Segu, Siyuan Li, Luc Van Gool, and Fisher Yu. Unidepth: Universal monocular metric depth estimation, 2024. URLhttps://arxiv.org/abs/2403. 18913. 2 [24] Luigi Piccinelli, Christos Sakaridis, Yung-Hsu Yang, Mattia Segu, Siyuan Li, Wim Abbeloos, and Luc Van Gool. Unidepthv2: Universal monocular metric depth estimation made simpler. IEEE Transactions on Pattern Analysis and Machine Intelligence, 48(3):2354â2367, March 2026. ISSN 1939-3539. doi: 10.1109/tpami.2025.3628473. URL http://dx.doi.org/10.1109/TPAMI.2025.3628473. 2 [25]Gangwei Xu, Haotong Lin, Hongcheng Luo, Xianqi Wang, Jingfeng Yao, Lianghui Zhu, Yuechuan Pu, Cheng Chi, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye, Sida Peng, and Xin Yang. Pixel-perfect depth with semantics-prompted diffusion transformers, 2025. URLhttps://arxiv.org/abs/2510.07316. 2, 6, 7 13 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning [26]Daniel Dauner, Marcel Hallgarten, Tianyu Li, Xinshuo Weng, Zhiyu Huang, Zetong Yang, Hongyang Li, Igor Gilitschenski, Boris Ivanovic, Marco Pavone, Andreas Geiger, and Kashyap Chitta. Navsim: Data- driven non-reactive autonomous vehicle simulation and benchmarking, 2024. URLhttps://arxiv.org/ abs/2406.15349. 3, 6, 7 [27] Wei Cao, Marcel Hallgarten, Tianyu Li, Daniel Dauner, Xunjiang Gu, Caojun Wang, Yakov Miron, Marco Aiello, Hongyang Li, Igor Gilitschenski, Boris Ivanovic, Marco Pavone, Andreas Geiger, and Kashyap Chitta. Pseudo-simulation for autonomous driving, 2025. URL https://arxiv.org/abs/2506.04218. 3, 6, 7 [28] Xiaofeng Wang, Zheng Zhu, Guan Huang, Xinze Chen, Jiagang Zhu, and Jiwen Lu. Drivedreamer: Towards real-world-drive world models for autonomous driving. In European conference on computer vision, pages 55â72. Springer, 2024. 3 [29]Guosheng Zhao, Xiaofeng Wang, Zheng Zhu, Xinze Chen, Guan Huang, Xiaoyi Bao, and Xingang Wang. Drivedreamer-2: Llm-enhanced world models for diverse driving video generation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 10412â10420, 2025. 3 [30]Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575, 2025. 3 [31]Dingkang Liang, Dingyuan Zhang, Xin Zhou, Sifan Tu, Tianrui Feng, Xiaofan Li, Yumeng Zhang, Mingyang Du, Xiao Tan, and Xiang Bai. Seeing the future, perceiving the future: A unified driving world model for future generation and perception, 2025. URL https://arxiv.org/abs/2503.13587. 3 [32]NVIDIA, :, Arslan Ali, Junjie Bai, Maciej Bala, Yogesh Balaji, Aaron Blakeman, Tiffany Cai, Jiaxin Cao, Tianshi Cao, Elizabeth Cha, Yu-Wei Chao, Prithvijit Chattopadhyay, Mike Chen, Yongxin Chen, Yu Chen, Shuai Cheng, Yin Cui, Jenna Diamond, Yifan Ding, Jiaojiao Fan, Linxi Fan, Liang Feng, Francesco Ferroni, Sanja Fidler, Xiao Fu, Ruiyuan Gao, Yunhao Ge, Jinwei Gu, Aryaman Gupta, Siddharth Gururani, Imad El Hanafi, Ali Hassani, Zekun Hao, Jacob Huffman, Joel Jang, Pooya Jannaty, Jan Kautz, Grace Lam, Xuan Li, Zhaoshuo Li, Maosheng Liao, Chen-Hsuan Lin, Tsung-Yi Lin, Yen-Chen Lin, Huan Ling, Ming-Yu Liu, Xian Liu, Yifan Lu, Alice Luo, Qianli Ma, Hanzi Mao, Kaichun Mo, Seungjun Nah, Yashraj Narang, Abhijeet Panaskar, Lindsey Pavao, Trung Pham, Morteza Ramezanali, Fitsum Reda, Scott Reed, Xuanchi Ren, Haonan Shao, Yue Shen, Stella Shi, Shuran Song, Bartosz Stefaniak, Shangkun Sun, Shitao Tang, Sameena Tasmeen, Lyne Tchapmi, Wei-Cheng Tseng, Jibin Varghese, Andrew Z. Wang, Hao Wang, Haoxiang Wang, Heng Wang, Ting-Chun Wang, Fangyin Wei, Jiashu Xu, Dinghao Yang, Xiaodong Yang, Haotian Ye, Seonghyeon Ye, Xiaohui Zeng, Jing Zhang, Qinsheng Zhang, Kaiwen Zheng, Andrew Zhu, and Yuke Zhu. World simulation with video foundation models for physical ai, 2025. URL https://arxiv.org/abs/2511.00062. 3 [33] Chaojun Ni, Guosheng Zhao, Xiaofeng Wang, Zheng Zhu, Wenkang Qin, Guan Huang, Chen Liu, Yuyin Chen, Yida Wang, Xueyang Zhang, et al. Recondreamer: Crafting world models for driving scene reconstruction via online restoration. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 1559â1569, 2025. 3 [34]Yuhang Lu, Yichen Yao, Jiadong Tu, Jiangnan Shao, Yuexin Ma, and Xinge Zhu. Can lvlms obtain a driverâs license? a benchmark towards reliable agi for autonomous driving. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 5838â5846, 2025. 3 [35] Guosheng Zhao, Chaojun Ni, Xiaofeng Wang, Zheng Zhu, Xueyang Zhang, Yida Wang, Guan Huang, Xinze Chen, Boyuan Wang, Youyi Zhang, et al. Drivedreamer4d: World models are effective data machines for 4d driving scene representation. In Proceedings of the computer vision and pattern recognition conference, pages 12015â12026, 2025. 3 14 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning [36]GigaWorld Team, Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang, Guosheng Zhao, Haoyun Li, Jiagang Zhu, Kerui Li, Mengyuan Xu, et al. Gigaworld-0: World models as data engine to empower embodied ai. arXiv preprint arXiv:2511.19861, 2025. 3 [37]Jiazhi Yang, Kashyap Chitta, Shenyuan Gao, Long Chen, Yuqian Shao, Xiaosong Jia, Hongyang Li, Andreas Geiger, Xiangyu Yue, and Li Chen. Resim: Reliable world simulation for autonomous driving, 2025. URL https://arxiv.org/abs/2506.09981. 3 [38]Zhenhua Xu, Yujia Zhang, Enze Xie, Zhen Zhao, Yong Guo, Kwan-Yee K. Wong, Zhenguo Li, and Hengshuang Zhao. Drivegpt4: Interpretable end-to-end autonomous driving via large language model. IEEE Robotics and Automation Letters, 9(10):8186â8193, 2024. doi: 10.1109/LRA.2024.3440097. 3, 8 [39]Hidehisa Arai, Keita Miwa, Kento Sasaki, Kohei Watanabe, Yu Yamaguchi, Shunsuke Aoki, and Issei Yamamoto. Covla: Comprehensive vision-language-action dataset for autonomous driving. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 1933â1943, 2025. doi: 10.1109/WACV61041.2025.00195. 3 [40] Zhenjie Yang, Yilin Chai, Xiaosong Jia, Qifeng Li, Yuqian Shao, Xuekai Zhu, Haisheng Su, and Junchi Yan. Drivemoe: Mixture-of-experts for vision-language-action model in end-to-end autonomous driving, 2025. URL https://arxiv.org/abs/2505.16278. 3 [41]Shuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu, Yifan Bai, Zheng Pan, Mu Xu, Xing Wei, and Ning Guo. Futuresightdrive: Thinking visually with spatio-temporal cot for autonomous driving, 2025. URL https://arxiv.org/abs/2505.17685. 3, 7, 8 [42]GigaBrain Team, Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang, Guosheng Zhao, Haoyun Li, Jie Li, Jiagang Zhu, Lv Feng, et al. Gigabrain-0: A world model-powered vision-language-action model. arXiv preprint arXiv:2510.19430, 2025. 3 [43]Yongkang Li, Kaixin Xiong, Xiangyu Guo, Fang Li, Sixu Yan, Gangwei Xu, Lijun Zhou, Long Chen, Haiyang Sun, Bing Wang, Kun Ma, Guang Chen, Hangjun Ye, Wenyu Liu, and Xinggang Wang. Recogdrive: A reinforced cognitive framework for end-to-end autonomous driving, 2025. URLhttps://arxiv.org/ abs/2506.08052. 3, 7, 8 [44] Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling, 2023. URL https://arxiv.org/abs/2210.02747. 4 [45]William Peebles and Saining Xie. Scalable diffusion models with transformers, 2023. URLhttps: //arxiv.org/abs/2212.09748. 5 [46] Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pandeng Li, Pingyu Wu, Ruihang Chu, Ruili Feng, Shiwei Zhang, Siyang Sun, Tao Fang, Tianxing Wang, Tianyi Gui, Tingyu Weng, Tong Shen, Wei Lin, Wei Wang, Wei Wang, Wenmeng Zhou, Wente Wang, Wenting Shen, Wenyuan Yu, Xianzhong Shi, Xiaoming Huang, Xin Xu, Yan Kou, Yangyu Lv, Yifei Li, Yijing Liu, Yiming Wang, Yingya Zhang, Yitong Huang, Yong Li, You Wu, Yu Liu, Yulin Pan, Yun Zheng, Yuntao Hong, Yupeng Shi, Yutong Feng, Zeyinzi Jiang, Zhen Han, Zhi-Fan Wu, and Ziyu Liu. Wan: Open and advanced large-scale video generative models, 2025. URL https://arxiv.org/abs/2503.20314. 5, 6 [47]Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021. URL https://arxiv.org/abs/2103.00020. 6 15 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning [48]Yang Zhou, Hao Shao, Letian Wang, Steven L Waslander, Hongsheng Li, and Yu Liu. Smartrefine: A scenario-adaptive refinement framework for efficient motion prediction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 15281â15290, 2024. 6 [49]Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan Liu, Dunjie Lu, Ruilin Luo, Chenxu Lv, Rui Men, Lingchen Meng, Xuancheng Ren, Xingzhang Ren, Sibo Song, Yuchong Sun, Jun Tang, Jianhong Tu, Jianqiang Wan, Peng Wang, Pengfei Wang, Qiuyue Wang, Yuxuan Wang, Tianbao Xie, Yiheng Xu, Haiyang Xu, Jin Xu, Zhibo Yang, Mingkun Yang, Jianxin Yang, An Yang, Bowen Yu, Fei Zhang, Hang Zhang, Xi Zhang, Bo Zheng, Humen Zhong, Jingren Zhou, Fan Zhou, Jing Zhou, Yuanzhi Zhu, and Ke Zhu. Qwen3-vl technical report, 2025. URL https://arxiv.org/abs/2511.21631. 6 [50]Bencheng Liao, Shaoyu Chen, Haoran Yin, Bo Jiang, Cheng Wang, Sixu Yan, Xinbang Zhang, Xiangyu Li, Ying Zhang, Qian Zhang, and Xinggang Wang. Diffusiondrive: Truncated diffusion model for end-to-end autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12037â12047, June 2025. 7 [51] Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. 7 [52]Q Huynh-Thu and M Ghanbari. Scope of validity of PSNR in image/video quality assessment. Electron. Lett., 44(13):800, 2008. 7 [53]Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. Towards accurate generative models of video: A new metric & challenges, 2019. URL https://arxiv.org/abs/1812.01717. 7 [54]Kashyap Chitta, Aditya Prakash, Bernhard Jaeger, Zehao Yu, Katrin Renz, and Andreas Geiger. Transfuser: Imitation with transformer-based sensor fusion for autonomous driving, 2022. URLhttps://arxiv.org/ abs/2205.15997. 7, 8 [55] Xinshuo Weng, Boris Ivanovic, Yan Wang, Yue Wang, and Marco Pavone. Para-drive: Parallelized architecture for real-time autonomous driving. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15449â15458, 2024. doi: 10.1109/CVPR52733.2024.01463. 7 [56]Bencheng Liao, Shaoyu Chen, Haoran Yin, Bo Jiang, Cheng Wang, Sixu Yan, Xinbang Zhang, Xiangyu Li, Ying Zhang, Qian Zhang, and Xinggang Wang. Diffusiondrive: Truncated diffusion model for end-to-end autonomous driving, 2025. URL https://arxiv.org/abs/2411.15139. 7, 8 [57]Yingyan Li, Lue Fan, Jiawei He, Yuqi Wang, Yuntao Chen, Zhaoxiang Zhang, and Tieniu Tan. Enhancing end-to-end autonomous driving with latent world model, 2025. URLhttps://arxiv.org/abs/2406. 08481. 7, 8 [58]Yuntao Chen, Yuqi Wang, and Zhaoxiang Zhang. Drivinggpt: Unifying driving world modeling and planning with multi-modal autoregressive transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 26890â26900, October 2025. 7 [59]Yingyan Li, Yuqi Wang, Yang Liu, Jiawei He, Lue Fan, and Zhaoxiang Zhang. End-to-end driving with online trajectory evaluation via bev world model. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 27137â27146, October 2025. 7, 8 16 DriveDreamer-Policy: A Geometry-Grounded WorldâAction Model for Unified Generation and Planning [60]Wenhao Yao, Zhenxin Li, Shiyi Lan, Zi Wang, Xinglong Sun, Jose M. Alvarez, and Zuxuan Wu. Drivesuprim: Towards precise trajectory selection for end-to-end planning, 2025. URLhttps://arxiv.org/abs/2506. 06659. 7 [61]Renju Feng, Ning Xi, Duanfeng Chu, Rukang Wang, Zejian Deng, Anzheng Wang, Liping Lu, Jinxiang Wang, and Yanjun Huang. Artemis: Autoregressive end-to-end trajectory planning with mixture of experts for autonomous driving. IEEE Robotics and Automation Letters, 11(1):226â233, 2026. doi: 10.1109/LRA.2025.3632616. 7 [62] Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization, 2016. URLhttps: //arxiv.org/abs/1607.06450. 8 [63] Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization, 2019. URLhttps://arxiv. org/abs/1711.05101. 8 17