Paper deep dive
A Wireless World Model for AI-Native 6G Networks
Ziqi Chen, Yi Ren, Yixuan Huang, Qi Sun, Nan Li, Yuhong Huang, Chih-Lin I, Yifan Li, Liang Xia
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/27/2026, 1:34:11 AM
Summary
The paper introduces the Wireless World Model (WWM), a multi-modal foundation framework for 6G networks that integrates 3D geometry, user trajectories, and channel state information (CSI) to predict wireless channel evolution. By utilizing a Joint Embedding Predictive Architecture (JEPA) and a Multi-Modal Mixture-of-Experts (MMoE) Transformer, WWM overcomes the generalization limitations of traditional data-driven models, achieving state-of-the-art performance in channel prediction, compression, beam management, and localization across both simulated and real-world environments.
Entities (5)
Relation Signals (3)
Wireless World Model â utilizes â Joint Embedding Predictive Architecture
confidence 100% ¡ the model adopts a Joint Embedding Predictive Architecture (JEPA)
Wireless World Model â supports â CSI temporal prediction
confidence 95% ¡ Across the five key downstream tasks supported by WWM... including CSI prediction
Wireless World Model â trainedon â Sionna RT
confidence 90% ¡ Pre-trained on a massive ray-traced multi-modal dataset... integrating high-fidelity Sionna RT ray tracing simulations
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Integrating AI into the physical layer is a cornerstone of 6G networks. However, current data-driven approaches struggle to generalize across dynamic environments because they lack an intrinsic understanding of electromagnetic wave propagation. We introduce the Wireless World Model (WWM), a multi-modal foundation framework predicting the spatiotemporal evolution of wireless channels by internalizing the causal relationship between 3D geometry and signal dynamics. Pre-trained on a massive ray-traced multi-modal dataset, WWM overcomes the data authenticity gap, further validated under real-world measurement data. Using a joint-embedding predictive architecture with a multi-modal mixture-of-experts Transformer, WWM fuses channel state information, 3D point clouds, and user trajectories into a unified representation. Across the five key downstream tasks supported by WWM, it achieves remarkable performance in seen environments, unseen generalization scenarios, and real-world measurements, consistently outperforming SOTA uni-modal foundation models and task-specific models. This paves the way for physics-aware 6G intelligence that adapts to the physical world.
Tags
Links
- Source: https://arxiv.org/abs/2603.25216v1
- Canonical: https://arxiv.org/abs/2603.25216v1
Trouble viewing inline? Open PDF directly â
Full Text
85,890 characters extracted from source content.
Expand or collapse full text
A Wireless World Model for AI-Native 6G Networks Ziqi Chen 1*â , Yi Ren 1â , Yixuan Huang 1â , Qi Sun 1â , Nan Li 1* , Yuhong Huang 1* , Chih-Lin I 1 , Yifan Li 1 , Liang Xia 1 1* China Mobile Research Institute, No. 32 Xuan Wu Men West Street, Beijing, 100053, China. *Corresponding author(s). E-mail(s): chenziqiyjy@chinamobile.com; linan@chinamobile.com; huangyuhong@chinamobile.com; Contributing authors: renyiyjy@chinamobile.com; huangyixuan@chinamobile.com; sunqiyjy@chinamobile.com; icl@chinamobile.com; liyifan 059052@q.com; xialiang@chinamobile.com; â These authors contributed equally to this work. Abstract Integrating AI into the physical layer is a cornerstone of 6G networks. However, current data-driven approaches struggle to generalize across dynamic environ- ments because they lack an intrinsic understanding of electromagnetic wave propagation. We introduce the Wireless World Model (WWM), a multi-modal foundation framework predicting the spatiotemporal evolution of wireless chan- nels by internalizing the causal relationship between 3D geometry and signal dynamics. Pre-trained on a massive ray-traced multi-modal dataset, WWM overcomes the data authenticity gap, further validated under real-world measure- ment data. Using a joint-embedding predictive architecture with a multi-modal mixture-of-experts Transformer, WWM fuses channel state information, 3D point clouds, and user trajectories into a unified representation. Across the five key downstream tasks supported by WWM, it achieves remarkable per- formance in seen environments, unseen generalization scenarios, and real-world measurements, consistently outperforming SOTA uni-modal foundation models and task-specific models. This paves the way for physics-aware 6G intelligence that adapts to the physical world. Keywords: 6G, air interface AI, wireless foundation model, joint embedding predictive architecture, multi-modal learning 1 arXiv:2603.25216v1 [cs.NI] 26 Mar 2026 1 Introduction Sixth-generation (6G) wireless networks are envisioned as the core information infras- tructure for the AI era, necessitating a quantum leap in performance and new capabilities like integrated AI and communication [1, 2]. Central to this evolution is the enhancement of spectral efficiency at the air interfaceâthe most fundamental and historically challenging goal of every mobile generation. The progression from 2G to 5G achieved this primarily by increasing bandwidth and deploying large-scale Mutiple Input and Multiple Output (MIMO) antenna systems, guided by Shannonâs informa- tion theory [3]. However, this traditional scaling approach has reached a bottleneck. Further expanding to extremely large-massive MIMO systems introduces prohibitive overhead for acquiring precise Channel State Information (CSI) and intractable com- putational complexity for precoding [4, 5]. Furthermore, a persistent gap to the theoretical Shannon limit exists, widened by hardware non-idealities and complex network interference that conventional signal processing models cannot adequately address [6]. AI presents a transformative solution, leveraging its advanced capabilities in feature extraction and complex problem-solving to break these barriers [7]. This pro- pels the vision of an AI-native 6G air interface, where AI is fundamentally integrated into core physical layer functions [8â10]. Initial efforts toward an AI-native air inter- face relied on task-specific models. However, this approach is fundamentally limited by poor generalization to dynamic environments, and a disjointed design that creates significant computational and management burdens at the base station (BS). In recent years, foundation models have demonstrated superior performance and changed the landscape of AI [11â14]. This has catalyzed a paradigm shift toward a more scalable wireless AI solution: the Wireless Foundation Model (WFM). A WFM is a large-scale model, pre-trained on diverse wireless data via self-supervision to improve general- ization capability, and designed to adapt to a wide range of downstream tasks with minimal fine-tuning. Early explorations into WFMs initially focused on adapting pre-trained Large Language Models (LLMs) to wireless domain through fine-tuning or prompt engineering[15â21]. While this demonstrated the potential of large-scale AI model architectures, the approach often yielded physically inconsistent predictions, as the models lacked an intrinsic understanding of electromagnetic wave propagation. This critical flaw spurred a necessary shift toward the current generation of AI-native architectures, such as WiFo [22, 23] and WirelessGPT [24], which are pre-trained from the ground up on vast, domain-specific corpora. Despite this progress, the field remains constrained by two limitations that form a âgeneralization ceiling.â First is a data authenticity gap; most models are trained on statistical channel data that fails to capture real-world complexity, a problem addressable by integrating high-fidelity, physics-based data like ray tracing [25]. Second is a modality gap, as existing solutions typically process a single data type (e.g., CSI), neglecting complementary information from 3D environment, which are crucial for true environmental understanding. To address these fundamental challenges, we propose the Wireless World Model (WWM), marking a shift from data-driven pattern matching to environment-aware cognitive intelligence. Conceptually, a World Model generates internal predictive rep- resentation of the environment that enables system to internalize physical laws and 2 predict future states [26â28]. While a WFM excels at capturing statistical correlations within signal distributions, the WWM specifically aims to learn the âphysics of the wireless worldââthe causal mapping between 3D spatial geometry and electromag- netic propagation. By learning the causal mapping from environmental semantics to channel characteristics, the WWM develops a âworld-centricâ latent representation that captures how physical objects and spatial dynamics dictate wave behavior. By perceiving the environment as a structured physical entity, the WWM exhibits three distinct features: environmental grounding, which links signal variations to physical objects; predictive consistency, ensuring channel estimations align with EM theory; and multi-task versatility, supporting diverse downstream functions through a shared understanding of the underlying physical space and electromagnetic pattern. In this work, we present the first WWM implementation that leverages 3D point cloud physical information to bridge the âgeneralization ceilingâ. The WWM intro- duces three fundamental innovations. First, we construct a massive hybrid multi-modal dataset of approximately 800 thousand samples by integrating high-fidelity Sionna RT ray tracing simulations [29] with real-world 6G prototype BS field trials, ensuring that the model learns from both physically consistent and realistic electromagnetic environments. Second, the model adopts a Joint Embedding Predictive Architecture (JEPA) [30, 31], which shifts the learning objective from signal reconstruction to pre- diction in a semantic feature space. This compels the model to internalize the intrinsic causal laws of wave propagation rather than merely fitting data correlations. Third, a Multi-Modal Mixture of Experts (MMoE) structure [32] is employed, enabling effi- cient fusion of multi-modal data including channel signal, point cloud and trajectory data within a unified Transformer backbone. This holistic architecture enables a âone- model-for-allâ paradigm, where a single pre-trained core can simultaneously perform high-precision channel prediction, channel compression and feedback, beam manage- ment, and user localization through lightweight, task-specific heads. As the first model to systematically incorporate 3D geometric priors into a wireless predictive frame- work, our WWM provides a crucial reference for future exploration in AI-native 6G and the development of more sophisticated wireless autonomous agents. 2 Results 2.1 Wireless world model understands Electromagnetic propagation via pretraining We introduce the WWM, a world model for wireless communications that not only reconstructs signals but also predicts the coherent spatiotemporal evolution of elec- tromagnetic environments in a manner consistent with physical laws. By extensive pre-training on a massive dataset, WWM learns generalizable features that enable effective adaptation to diverse downstream tasks, even with limited task-specific training data. Figure 1 depicts the overall WWM framework. A hybrid large-scale multi-modal wireless dataset is constructed through ray tracing simulation by Sionna RT and field measurement with China Mobile 6G prototype system. By incorporating MMoE model architecture within JEPA pre-training framework, WWM fuses hetero- geneous dataâCSI, 3D point clouds, and user trajectoriesâinto a unified semantic 3 OBSERVED MODALITIES UNIFIED REPRESENTATION DOWNSTREAM TASKS a Wireless Data Abstraction Layer b Wireless World Model c Downstream Tasks Discretized Space Latent Representation ⪠Ray-tracing Simulation Field Measurements Wireless World Model with Modality-Aware Mixture-of-Experts Transformer (MMoE) CSI Embeddings Point Clouds Embeddings UE Trajectory Embeddings Channel Point Cloud Trajectory Unified Wireless World Model Enabled Tasks Trajectory Channel State Information Point Cloud Trajectory Sionna RT 6G Testbed Prototype Point Cloud Channel State Information Large-Scale Multimodal Dataset Uniform physical-layer representation across simulation and measurement H-FNN PC-FNN P-FNN Multi-Strategies Masking CSITemporal Prediction Beam Prediction CSICompression & Feedback User Localization ... Conv3D Embedder Point-BERT Embedder MLP Embedder ... + concatenate + concatenate ... Shared Self-Attention M ... ... M M M M M Strategy â Strategy ⥠Strategy ⢠... ... ... M M M M ... ... ... ... CSI Frequency-Domain Prediction Fig. 1 The workflow of WWM. a, Multi-modal data source for pre-training and evaluation. The ray tracing simulation is performed on Sionna RT, which consisted of channel State Information (CSI), 3D Point Clouds and User Equipment (UE) Trajectory. The Field Measurement is collected outdoor from China Mobile 6G prototype system. b, Pre-training model architecture and pre-training tasks. The WWM is a pre-trained on JEPA, involving an encoder-predictor architecture. Both encoder and predictor are Multi-modal Mixture of Expert Transformer model, trained on 3 self-supervised mask tasks. c, Downstream tasks for validation based on WWM embedding. The representational capabilities of WWM is verified on 4 downstream tasks with simulated data, and its real-world generalization ability is evaluated on CSI frequency-domain prediction based field measurement. representation. This unified approach enables a single pre-trained model to support multiple network optimization downstream tasks, including CSI prediction, channel compression and feedback, beam prediction, and user localization, without requiring separate, task-specific Algorithms or AI models. To enable wireless world modeling, we constructed a large-scale hybrid multi- modal wireless dataset (Figure 2). The dataset integrates timeâfrequencyâspace CSI, scenario-level 3D point clouds and synchronized UE trajectories collected across five representative urban environments: Munich, Paris, Beijing CBD, the Forbidden City and Wall Street. Using physics-based ray tracing [29], we generated more than 700,000 channel samples under multiple user mobility regimes (5, 30 and 60 km/h), providing diverse spatioâtemporal wireless observations (Fig. 2a). To assess real-world applicability, we further collected uplink CSI measurements from a 6G prototype system deployed at the China Mobile International Information Port in Beijing (Fig. 2b). This real-world dataset introduces hardware impairments and environmental noise, enabling evaluation of the modelâs ability to transfer from simula- tion to practical wireless environments. Detailed simulation parameters, measurement configurations and dataset composition are provided in Extended Data Tables 1â4. 4 Fig. 2 Large-scale multi-modal wireless dataset spanning diverse simulated and real- world environments. Across these environments, we collect multi-modal data including scenario 3D point clouds, user trajectories and time-synchronized CSI for each sample. a, Representative real- world photographs of the selected urban environments used for ray tracing simulation. From top to bottom: Place de lâ Ě Etoile (Paris), Forbidden City (Beijing), Munich urban district (Germany), central business district (Beijing), and Wall Street (New York). b, Corresponding 3D scenario mod- els constructed from geographic data and ground signal coverage maps generated in Sionna RT. c, Extracted 3D point clouds of corresponding 3D scenario models. d, Photograph of the real-world outdoor measurement environment used for channel data acquisition, with the base station (BS) location indicated. e, Satellite view of the measurement site, where the yellow cross marks the BS position and the green trapezoid indicates the UE trajectory. f, 3D point clouds reconstructed from the measurement environment. g, BS hardware of the 6G prototype system used for real-world mea- surements. h, UE device used for outdoor channel data acquisition. As illustrated in Fig. 3a, WWM adopts a JEPA pre-training framework, which fun- damentally differs from traditional masked autoencoders [33]. Instead of reconstructing raw data, WWM predicts wireless channel semantic representations in a latent space, forcing the model to understand and forecast electromagnetic evolution abstractly like a world model. Before entering the Transformer backbone, each modality is processed by a dedicated embedder tailored to its physical characteristics. Central to this archi- tecture is a MMoE Transformer (Fig. 3b). This design enables the seamless fusion of heterogeneous modalitiesâCSI, 3D point clouds, and user trajectoriesâinto a unified latent space, allowing a single pre-trained backbone to support diverse downstream tasks. The intelligence of WWM emerges from its self-supervised pre-training strategy. We used three complementary masking tasks (Fig. 3c) during pre-training. First, fine- grained CSI masking encourages the reconstruction of local multipath components. Second, coarse CSI masking forces the model to infer global channel structures from environmental context. Third, trajectory masking requires the model to deduce user motion solely from channel evolution. By alternating between these tasks, WWM learns the mapping between electromagnetic signal evolution and physical user motion. 5 The effectiveness of this pre-training is evident in the modelâs ability to recon- struct missing information from context. We pre-trained WWM based on simulation data across 4 cities (Munich, Paris, Beijing CBD, Beijing Forbidden City) as detailed in Extended Data Table 3, while the simulated data of fifth city (Wall street) and of additional velocities in CBD is reserved for generalization testing, as detailed in Extended Data Table 4.) Fig. 4a visualizes the reconstruction of a masked CSI sam- ple of 16 timesteps. Even when significant time-frequency blocks are masked, WWM accurately restores the channel structure by reasoning from the unmasked CSI blocks, 3D geometry and trajectory cues. We utilized t-distributed stochastic neighbour embedding (t-SNE) [34] to visualize the WWM encoderâs final-layer CSI embeddings with different data label. The outcome, as illustrated in Fig. 4b, revealed that the embeddings produced by WWM organizes samples into meaningful clusters based on unique characteristics, demonstrating that it has successfully internalized a structured representation of the wireless physical environment without explicit supervision. 2.2 WWM augments RAN downstream tasks We evaluated the performance of the pre-trained WWM across four downstream tasks, benchmarking it against SOTA task-specific models and representative WFMs, specif- ically LWM [35] and WiFo [22]. To ensure a rigorous comparison, WFMs were assessed using either official checkpoints or checkpoints pre-trained on the same multi-modal dataset (extended data table 3) if training code is provided. We kept backbone frozen for all foundation models, where task-specific knowledge was captured by training lightweight output heads. In contrast, task-specific SOTA baselines were trained full- shot from scratch using the full labeled dataset for each respective task. As detailed in the following sections, WWM consistently achieved SOTA performance across all four tasks. This demonstrates the superior adaptation and transferability of WWMâs latent representations. Furthermore, ablation experiments reveal that reverting WWM to a single-modal configurationâby pre-training without 3D point clouds and user trajectory priorsâobserves a obvious performance degradation. This underscores the necessity of multi-modal environmental grounding for robust wireless representation. Implementation details and specific task analyses are provided in the respective results sections as followed and further detailed in Methods. CSI temporal prediction: To assess whether WWM captures channel evolution beyond statistical correlation, we evaluated its performance on CSI temporal predic- tion against SOTA baselines WiFo [22] and LSTM [36], As shown in Fig. 5a. WWM encoder and predictor is configured to predict future CSI based on 14 history CSI timesteps in latent space, while a decoder is trained to recover CSI from latent space representation. For in-pattern urban environments (CBD, Etoile, Forbidden City and Munich), WWM achieves consistently high Squared Generalized Cosine Similarity (SGCS) scores of 0.80â0.96, outperforming best baselines by an average of 0.12 (Fig. 5b). Crucially, WWM breaks the generalization ceiling: in the completely unseen âWall Streetâ environment, labeled as Gen-City, it sustains an SGCS of 0.92âa relative gain of 56% over LSTM (0.59) and 21% over WiFo (0.76). These results suggest that with the aid of multi-modal data, forecasting in the latent space, rather than directly in the raw channel domain, leads to more robust channel predictions under seen and unseen 6 Shared Multi-head Self-Attention UE Position x N Conv3D Embedder H-FFNPC-FFN P-FFN Point-net Embedder MLP Embedder Switch Modality Experts Channel Matrix Point Cloud Predictor \ Remove masked tokens Concaten ate mask tokens . . . Masks Positions . . . Encoder . . . . . . . . . . . \ Remove mask tokens . . Loss Embedded Tokens EMA Encoder Channel Matrix Point Clouds UE Trajectories EMA-based online updating Wireless World Model â Fine-grained CSI masking âĄCoarse CSI masking â˘Trajectory masking Channel Matrix Point Clouds UE Trajectories JEPA Style Pre-training Process Multi-Modal Mixture of Experts (MMoE) Pre-training Masking Strategies a b c + + + ++ + Fig. 3 The model and pre-training process. a, WWM employs a Joint Embedding Predictive Architecture (JEPA) to infer masked multi-modal features in latent space. An online encoder processes visible tokens while a predictor estimates masked embeddings, supervised by an Exponential Moving Average (EMA) based momentum encoder to ensure representation stability. b, Multi-modal Mixture of Experts (MMoE). Heterogeneous inputsâCSI, 3D point clouds, and trajectoriesâare tokenized via domain-specific embedders (Conv3D, Point-net, and MLP). Within each Transformer block, shared self-attention performs global cross-modal reasoning, followed by modality-specific experts (H-FFN, PC-FFN, P-FFN) to preserve physical inductive biases. c, Pre-training masking strategies. Three complementary strategies supervise the model: Fine-grained and coarse CSI masking to extract multi- scale spatio-temporal propagation features. Trajectory masking to capture kinematic dynamics and their interaction with the electromagnetic environment. propagation environments. A detailed quantitative comparison across models and test scenarios is provided in Extended Data Table 5. CSI compression and feedback: In massive MIMO systems, minimizing feed- back overhead while ensuring high CSI acquisition accuracy is critical for accurate precoding and thus wireless transmission capacity. Building on the output of WWM, a pair of highly efficient deep learning-based compressor and de-compressor networks can be trained to significantly reduce the feedback payload while maintaining CSI fidelity, as demonstrated in Fig. 5c. Experimental results across various compression ratios (from 1/1024 to 1/128) demonstrate that the WWM-based compressor consistently achieves superior performance compared with baseline methods such as QCR-NET [37] and CR-NET [38], as shown in Fig. 5d. For in-pattern urban environments (CBD, 7 (a) Original(b) Masked (c) Reconstructed a b Embedding Distributions (b) Cities (c) LOS/NLOS (a) Token Level (d) Base Stations (e) Noise Levels Masks and Reconstruction Heatmaps LOS NLOS CBD Etoile Forbidden City Munich BS-0 BS-1 Normal 0dB Each Color Represents a Sample Fig. 4 Pre-training results. a, The visualizations of CSI reconstruction. 1st graph: Original 16 timestep CSI sample. 2nd graph: Masked CSI sample used as input to the WWM models for fine- grained masking strategy. 3rd graph: Reconstructed CSI sample using masked input. b, The t-SNE [34] maps show the encoderâs final-layer embeddings across five sampling schemesârandomly sampled token-level embeddings, samples grouped by city, samples grouped by LOS/NLOS conditions, samples grouped by BS and samples grouped by noise levels. Etoile, Forbidden City, and Munich), WWM maintains strong reconstruction fidelity with SGCS scores ranging from 0.65 to 0.95 across varying compression ratios, out- performing the most competitive baseline (QCR-NET) by an average absolute SGCS gain of 0.13. Crucially, WWM exhibits robust generalization capabilities under shift- ing conditions: in the completely unseen city generalization scenario, it achieves an average relative gain of 21% over QCR-NET(0.58) and 62% over CR-NET(0.43) across all evaluated compression ratios. Notably, even at the extreme compression ratio of 1/1024, it secures an SGCS of 0.62 in unseen cities compared to QCR-NETâs 0.48. Similarly, for velocity generalization, WWM consistently surpasses QCR-NET by an average relative gain of 9%. These outcomes highlight the effectiveness of the WWM architecture not only in preserving CSI with high accuracy but also in generalizing robustly to unseen propagation environments and mobility conditions, underscoring its practical viability for real-world massive MIMO deployments. A detailed quantita- tive comparison across compression rates and test scenarios is provided in Extended Data Table 6. Beam prediction: Sub-6 GHz signals propagate through the same physical envi- ronment as upper-6G (U6G) signals and therefore their dominant propagation paths are shaped by the same geometry. Leveraging this property, the BS can use Sub-6 GHz CSI to predict the most suitable U6G beam direction, thereby avoiding an exhaustive 8 c d Base Station UE Reference Signal v Compressed CSI (Bitstream) 101100101101... Compression head De-compression head Reconstructed CSI Wireless channel UE WWM encoder Reference Signal HistoryCSI Base Station ... a b f g h UE Base Station Regression head (í,í) i j UE Reference Signal Base Station WWM predictor Decoder head Wireless channel WWM encoder PredictedCSI WWM encoder MeasuredCSI ... Coordinate Wireless channel WWM encoder PartialCSI ... WWM predictor Decoder head PredictedCSI e U6G AAU Sub6G AAU Classification head Beam index WWM encoder Wireless channel Reference Signal UE Base Station Wireless channel Reference Signal 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 0.40 PRB4PRB5PRB6PRB7Avg. NMSE WWMWiFoCMixer2 0.70 0.75 0.80 0.85 0.90 0.95 1.00 PRB4PRB5PRB6PRB7Avg. SGCS 0.40 0.50 0.60 0.70 0.80 0.90 1.00 CBDEtoileForbidden CityMunichGen-VelocityGen-City SGCS CSI Temporal Prediction Performance Comparison WWMWifoLSTM 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Accuracy Top-1 Accuracy Comparison WWMLWMSPMM 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Ratio Beam Gain Ratio Comparison 0.75 0.76 0.78 0.89 0.60 0.68 0.73 0.80 0.30 0.46 0.60 0.68 0.20 0.40 0.60 0.80 1.00 1/10241/5121/2561/128 SGCS In-pattern WWMQCR-NETCR-NET 0.62 0.62 0.65 0.90 0.48 0.54 0.58 0.70 0.26 0.36 0.52 0.58 1/10241/5121/2561/128 Gen-City 0.65 0.66 0.66 0.75 0.59 0.61 0.62 0.67 0.38 0.53 0.62 0.63 1/10241/5121/2561/128 Gen-Velocity CBD Etoile Forbidden City Munich Gen-Velocity Error (m) Error (m) Error (m) Error (m) Error (m) CDF (%) CSI Compression and Feedback Performance Comparison Beam Prediction Performance Comparison User Localization Performance Comparison CSI frequency-domain Prediction Performance Comparison WWMdeep-CNN MeasuredCSI MeasuredCSI Fig. 5 Downstream task and results comparison a, The CSI temporal prediction task details. b, CSI temporal prediction performance comparison with SOTA WFM Wifo and SOTA task-specific model LSTM. c, CSI compression and feedback task details. d, CSI compression and feedback performance comparison with SOTA task-specific models QCR-NET and CR-NET across various compression rate. e, Beam prediction task details. f, Beam prediction across frequency bands performance comparison with SOTA WFM LWM and SOTA task-specific model SPMM. g, User localization task details. h, User localization performance comparison with SOTA task-specific model deep-CNN. i, CSI frequency-domain prediction in real-world measurement task details. j, CSI frequency-domain prediction performance comparison with SOTA WFM Wifo and SOTA task- specific model C-Mixer. 9 pilot-based beam search and reducing both air-interface overhead and energy consump- tion (Fig. 5e). We evaluate beam prediction using top-1 classification accuracy and beam gain ratio, where beam gain ratio is the channel gain achieved by predicted beam divided by that of theoretical best beam. We compare WWM against two baselines: (i) LWM [35], a SOTA uni-modal WFM and (i) SPMM [39], a SOTA task-specific. As shown in Fig. 5f, WWM achieves an averae Top-1 accuracy of 94.0%âa relative increase of 7.3% over LWM (87.6%) and 13.6% over SPMM (82.6%), demonstrating stronger feature extraction capability. It also demonstrates a very strong beam gain ratio (99%), suggesting the predicted beam can always achieve near optimal channel gain. A detailed quantitative comparison across test scenarios is provided in Extended Data Table 7. User localization: Leveraging its multi-modal understanding of the environment, WWM enables high-precision user localization by treating wireless CSI as a distinctive âfingerprintâ (Fig. 5g). We quantify localization performance using the cumulative distribution of 2D absolute error in meters, as shown in Fig. 5h. WWM achieved an average localization error of 2.2 meters, compared with a conventional CNN-based regression baseline model which scored at 4.1 meters, WWM shows a reduction of 46% in terms of average localization error across all tested scenarios, highlighting its ability to extract position-relevant features directly from raw CSI and 3D point clouds. A detailed quantitative comparison across test scenarios is provided in Extended Data Table 8. Ablation Study: To quantify the contribution of multi-modal fusion to the WWM, we conducted ablation experiment in which the model was pre-trained using only CSI data, with 3D point cloud and user trajectory modalities masked. This uni-modal variant preserves the same JEPA pre-training scheme and model architec- ture (despite removing 2 un-used modality-specific expert FFNs), differing only in the absence of auxiliary modalities. When evaluated on the CSI temporal prediction task, the uni-modal model exhibited consistently lower prediction SGCS (a decrease of 6.0% on average across all scenarios, from 0.886 to 0.833). These results indicate that geometric context and motion cues provide essential complementary information that cannot be recovered from CSI alone. Detailed quantitative comparisons between the uni-modal and multi-modal variants are provided in the Extended Data Table 9. To validate the deployment feasibility of WWM in RAN hardware, we quantified the end-to-end inference latency of WWM. For the frozen WWM backbone (including encoder and predictor) paired with decoder heads, the average single-sample end-to- end inference latency is 8.5 ms on an NVIDIA RTX 4090 24GB GPU under mixed- precision bfloat16 inference. 2.3 WWMâs generalization capability in Real-world measurement In time-division duplex (TDD) systems, uplink and downlink channels are reciprocal within the coherence time, allowing BS precoding and multi-user coordination to rely on uplink channel observations. Sounding reference signals (SRS) provide uplink CSI at the BS; inferring unmeasured frequency resources from partial SRS observations reduces measurement overhead while maintaining channel awareness for precoding, 10 scheduling and link adaptation. Beyond its direct system-level relevance, predicting CSI in frequency-domain serves as a critical test of model transferability from large- scale simulated CSI pre-training to real-world measurements. Specifically, WWM is first pre-trained on simulated CSI data (Extended Data Table 3), and then fine-tuned using a comparatively small real-world measured CSI dataset collected from the 6G prototype system. Given CSI measured on four SRS Physical Resource Blocks (PRBs) in frequency domain, the model predicts the CSI on adjacent four PRBs within the same time window (Fig. 5i). As illustrated in Fig. 6, the modelâs CSI frequency-domain prediction outputs are visualized, showing strong agreement with the ground truth across real, imaginary, magnitude, and phase representations. Fig. 6 Visualization of CSI frequency-domain prediction based on SRS measurement. Comparison between ground-truth and predicted channel tensors across real, imaginary, magnitude, and phase components for a representative sample. The horizontal axis corresponds to the joint UE antenna and subband dimension, and the vertical axis corresponds to base-station antennas. a, b, Ground-truth real and imaginary parts of the channel tensor. c, d, Corresponding predicted real and imaginary parts produced by the model. e, f, Squared error maps for the real and imaginary components, respectively, showing low reconstruction error across most spatialâfrequency elements. g, h, Ground-truth channel magnitudes and phase. i, j, Corresponding predicted channel magnitudes and phase. k, l, Magnitude and Phase error map, indicating small deviations concentrated in limited regions, confirming accurate reconstruction of signal magnitudes and phase. Despite the substantial gap between simulated and measured channelâarising from hardware impairments, measurement noise, and mismatched feature statisticsâWWM rapidly adapts to real environments and outperforms the baseline WFM method WiFo[22] and task-specific model C-Mixer [40] (Extended Data Table 10). As shown 11 in Fig. 5j, it achieves improvements of 15.9% and 27.6% in NMSE, and 3.1% and 4.0% in SGCS, respectively. These results indicate that the model learns transferable representations of wireless dynamics that extend beyond the specific characteristics of ray tracing simulations, and that limited real measurements (13.2% of simulated pre-training data volume) are sufficient to specialize the model to real-world systems. This observation suggests a practical pathway for deploying foundation models in future 6G networks, where large-scale simulation can be leveraged for pre-training, and lightweight fine-tuning on sparse real data enables fast adaptation to new environments and system configurations. 3 Discussion Our work establishes the WWM as a key step toward transitioning wireless AI from fragmented, task-specific âexpert modelsâ to a unified and scalable foundation model. We show that the limitations of existing specialized modelsâincluding limited robust- ness, narrow cognitive scope due to single-modality inputsâcan be overcome by learning a universal representation of the wireless environment. Our primary contribu- tion is a multi-modal world model tailored to AI-Native 6G Network integrated with a MMoE model architecture. This architecture enables fusion of heterogeneous data sourcesâCSI, 3D point clouds, and user trajectoriesâinto a coherent latent space. In addition, we introduce a large, high-fidelity hybrid dataset of more than 700,000 sam- ples that bridges ray tracing simulations and real-world 6G prototype measurements, setting a new benchmark for data diversity and realism. WWM demonstrates strong performance across tasks, which we attribute to large- scale pre-training on diverse scenarios. This pre-training enables WWM to capture complex, non-linear dependencies in the channel that smaller, capacity-limited models fail to resolve, consistent with empirical scaling laws for large models. Furthermore, the multi-modal pre-training paradigm of the WWM transcends the limitations of conventional uni-modal models by establishing a coherent physical mapping between environmental geometry and electromagnetic propagation. By integrating 3D point clouds and user trajectories during self-supervised masked modeling, the WWM inter- nalizes a âphysical priorâ that characterizes how spatial dynamics dictate channel evolution. Our ablation analysis highlights the criticality of this multi-modal integra- tion: for channel prediction tasks, the exclusion of geometric and mobility priors leads to a sharp degradation in prediction accuracy. The WWM architecture represents a paradigm shift from the conventional âsiloedâ AI framework, where discrete, task-specific models are optimized for individual func- tions such as CSI prediction, beam management, and localization. By unifying these heterogeneous tasks within a single representational backbone, the WWM achieves remarkable cross-site generalization. While the initial pre-training of such a founda- tion model involves computational overhead, our empirical results confirm this barrier is low: full pre-training can be completed on a single consumer-grade GPU in 87 hours, with the cost further amortized across the vast network of sites and diverse downstream tasks. This achieves a long term computational equilibrium: the total lifecycle expen- diture (pre-training plus fine-tuning) for a centralized WWM becomes significantly 12 more efficient than the cumulative burden of training, deploying, and maintaining a large number of fragmented task-specific models across a massive RAN deployment. By consolidating the operational management from thousands of site-specific entities into a single âworld-awareâ engine, the WWM significantly mitigates the complex- ity of AI-native 6G networks, providing a scalable and maintainable trajectory for automated network evolution. Despite these advances, challenges remain. First, the current WWM relies on high-fidelity 3D point clouds, which may not be available in real-time for real deploy- ments. Future work may explore inferring geometry directly from sparse channel measurements. Second, while we validated the model on a 6G prototype system, scal- ing to massive multi-user interference scenarios requires further investigation. Third, the computational cost and inference latency can be further reduced through model light-weighting technologies, including knowledge distillation, model pruning, sparse attention mechanism, and mixed-precision computing. Looking forward, WWM estab- lishes the groundwork for autonomous wireless networks. By equipping networks with a predictive âworld modelâ, we move closer to the ultimate vision of 6G: a self-evolving, physics-aware infrastructure that optimizes itself in harmony with the physical reality it serves. 4 Methods 4.1 Large-scale multi-modal dataset construction To establish a comprehensive and realistic data foundation for the wwm, we con- structed a large-scale multi-modal dataset integrating urban geometric information, wireless channel state information (CSI), and user trajectories, alongside a comple- mentary real-world dataset to validate generalization under empirical conditions. Five representative urban environments were selected to capture diverse city structure: the Forbidden City (Beijing), the central business district (Beijing), Place de lâ Ě Etoile (Paris), an urban district of Munich and Wall Street (New York). Raw geographic data were obtained from OpenStreetMap, which provides building footprints, road layouts and associated metadata at city scale [41]. These map data were converted into three- dimensional urban scenes through a preprocessing pipeline in which building footprints were extruded using available height information or standardized urban assumptions when height data were unavailable. The resulting geometries were imported into Blender to refine scene structure and assign electromagnetic surface materials. Since OpenStreetMap-derived meshes may contain non-manifold edges, self-intersections and open surfaces, all scenario elements were further processed using Mitsuba [42] to generate closed manifold triangular meshes. This procedure resolves non-manifold geometry, unifies surface normals and removes degenerate triangles, ensuring geomet- ric consistency for ray-based electromagnetic simulation. The sanitized scenes were exported in a ray-tracing-compatible format that serves as a unified geometric backend for both wireless channel simulation and environment point-cloud generation. Wireless channels were generated using the physics-based ray tracing framework provided by Sionna [29]. For each scenario, multiple base-station placements were sim- ulated, with CSI computed along continuous UE trajectories spanning several city 13 blocks. The UE trajectories follow straight-line paths at constant speeds typical of normal pedestrians and vehicles. At regularly sampled time instants along each tra- jectory, frequency-selective MIMO channels were computed based on line-of-sight and specular reflection propagation paths determined by the urban geometry and material properties. Each channel realization is represented as a complex-valued tensor in the form C N UE ĂN r ĂN t ĂTĂF , where N UE denotes the number of UE samples, N r and N t are the numbers of receive and transmit antenna ports, T is the number of timesteps per sample and F is the number of subcarriers. The antenna dimensions N t and N r are determined by the antenna array size and polarization configuration at the base station and UE. Simulation parameters and their correspondence to these variables are summarized in Extended Data Table 1. To facilitate storage and model training, the complex CSI tensors were transformed into a real-valued representation by separating real and imaginary components and grouping subcarriers into non-overlapping subbands. The resulting representation has the form R N UE Ă2ĂTĂN t ĂN Ⲡr , where the factor of two corresponds to real and imaginary components and N Ⲡr = N r ĂN sb , with N sb denoting the number of subbands obtained by aggregating consecutive 12 subcarriers. This representation preserves frequency selectivity at the subband level while reducing the dimensionality of the channel tensor. In parallel with channel simulation, explicit geometric representations of the physical environment were constructed as 3D point clouds. Using the same scene descriptions employed for ray tracing, points were sampled from the surfaces of all triangular meshes, producing global-view point clouds represented as R N PC Ă3 that encode the static geometry of each urban environment, where N PC denotes the number of point clouds in the environment. All modalitiesâincluding CSI, 3D point clouds and User trajectoriesâare expressed in a unified global coordinate system and temporally synchronized. The final pre-training dataset contains 24 simulated subsets spanning four urban environments (Forbidden City, Beijing CBD, Munich and Place de lâ Ě Etoile), multiple base-station deployments and UE speeds of 5, 30 and 60 km/h. To evaluate generalization, the Wall Street datasets (5, 30 and 60 km/h) were excluded from pre-training and used exclusively for scenario-level testing. In addition, Beijing CBD datasets collected at previously unseen speeds of 40 and 70 km/h were reserved for speed-level generaliza- tion evaluation. Detailed dataset composition and splits are summarized in Extended Data Tables 3 and 4. To evaluate model performance under empirical wireless conditions, we collected a real-world wireless dataset containing uplink CSI measurements derived from SRS cap- tured using a 6G prototype system developed by the China Mobile Research Institute. The prototype platform integrates advanced transmission technologies and supports high-fidelity wireless experimentation and data acquisition. We collected outdoor wire- less dataset at the China Mobile International Information Port in Beijing using a carrier frequency of 6.6 GHz and a bandwidth of 400 MHz. The resulting dataset pro- vides realistic channel observations that include hardware impairments, environmental noise and non-ideal propagation effects. The 3D point clouds and user trajectories were constructed from on-site geospatial measurements. Detailed system parameters are summarized in Extended Data Table 2. 14 4.2 Model architecture The WWM is implemented as a multi-modal JEPA. The model comprises three main components: an online encoder f θ , a predictor g Ď and a target encoder f Ě Î¸ whose param- eters are maintained as an exponential moving average (EMA) of the online encoder. The online encoder operates on partially observed inputs and produces context embed- dings, the predictor uses these embeddings together with learnable mask tokens to infer the representations of masked regions, and the target encoder provides slowly varying target embeddings for the same inputs without masking. All supervision is applied in the latent space, and the model is never asked to reconstruct raw samples. JEPA-style multi-modal prediction For each training sample, the raw input consists of three synchronized modalities: a CSI tensor x CSI , a 3D point cloud tensor x PC and a user trajectory vector x POS in a BS-centered coordinate frame. After modality-specific embedding (further explained in 4.2.1), the resulting token sets X CSI , X PC , X POS are concatenated into a unified sequence X 0 = Concat X CSI , X PC , X POS â R NĂD . A masking module then selects two disjoint index sets over this sequence, I full =I enc âI pred ,(1) which specify visible (context) tokensI enc and masked (prediction) tokensI enc indices, respectively. The online encoder f θ receives only the visible tokens and produces context embeddings z ctx = f θ (X;I enc ),(2) In parallel, the target encoder f Ě Î¸ processes the complete, unmasked input and outputs target embeddings h = f Ě Î¸ (X;I full ),(3) from which the targets at the masked positions h I pred are extracted. The predictor g Ď then takes the context embeddings together with a set of learnable mask tokens m j â R D jâI pred , each associated with a masked position in I pred , and produces predicted embeddings Ë h I pred = g Ď z ctx , m j jâI pred ; I pred ,(4) A latent-space loss (here an â 1 distance) is computed between Ë h I pred and h I pred , and gradients are applied to θ and Ď only; the target parameters Ě Î¸ are updated via EMA. This scheme encourages the model to learn stable, semantically meaningful embed- dings that capture the physical structure linking channel responses, geometry and motion. 4.2.1 Multi-modal input embedding and unified token space To enable joint processing of heterogeneous modalities, all inputs are mapped to a shared embedding space with dimension D. Each modality is first converted into a sequence of tokens using a modality-specific encoder, and these tokens are then projected into the common latent dimension. 15 Channel (CSI) embedding The raw CSI tensor x CSI â R C in ĂTĂHĂW (where C in = 2 corresponds to the real and imaginary components) is treated as a 3D spatiotemporal volume over time, frequency, and antenna (spatial) indices. A 3D convolution with kernel size and stride both equal to (T p ,H p ,W p ) serves as the patchification operator: it partitions the volume into non-overlapping tubes of size (T p ,H p ,W p ) and simultaneously projects each tube into a D-dimensional embedding vector, yielding N CSI = T T p Ă H H p Ă W W p CSI tokens X CSI i â R D N CSI i=1 . To preserve the spatiotemporal ordering, a fixed 3D sinusoidal positional encoding is added to each token. Point-cloud embedding The raw 3D point cloud x PC â R N PC Ă3 is encoded using a discrete-variational tok- enizer inspired by point-cloud auto-encoding methods. Farthest-point sampling selects a fixed number of centers, and local neighborhoods around these centers are con- structed by nearest-neighbor grouping. A lightweight PointNet-style encoder from Point-BERT [43] extracts a feature vector for each neighborhood, which is then passed through a learned code book via GumbelâSoftmax quantization and refined by a shal- low geometric network. The result is a set of N PC patch-level point cloud tokens in the shared latent space, x PC â X PC j â R D N PC j=1 . This procedure preserves local geometry while compressing the raw point cloud into a compact, fixed-size sequence. Trajectory embedding The raw user trajectory vector x POS =p t â R 3 T pos t=1 is represented as a sequence of positions over time. A small projection network, implemented as a multi-layer percep- tron with non-linear activations, maps each position to a D-dimensional embedding, and a temporal positional encoding is added: x POS âX POS k â R D N POS k=1 . Unified token sequence After modality-specific embedding, the three token sequences are concatenated along the sequence dimension in a fixed order, X 0 = Concat X CSI , X PC , X POS â R NĂD ,(5) where N = N CSI +N PC +N POS and X CSI , X PC , X POS denote all the tokens for each modality. The model keeps track of the segment lengths (N CSI ,N PC ,N POS ) for use in the modality-aware expert layers. In our implementation, the numbers of tokens for CSI, point cloud, and trajectory are N CSI = 512, N PC = 256, and N POS = 16, respectively. 4.2.2 Shared cross-modal attention and modality-specific experts The unified token sequence X 0 is processed by a stack of L e Transformer blocks. Within each block, layer-normalized tokens first pass through a shared multi-head self-attention (MHSA) layer that operates on all modalities jointly, enabling cross- modal information exchange. The output is then layer-normalized, split by modality 16 and routed to three parallel feed-forward sub-networks, each specialized to CSI, point cloud, or trajectory tokens respectively. The expert outputs are concatenated back and added as a residual. This two-stage process repeats for L e layers, yielding the final representation X L e . Shared self-attention Given the input X â â R NĂD at layer â, a standard multi-head self-attention (MHSA) layer with pre-normalization and residual connection is applied: X Ⲡâ = X â + MHSA LN(X â ) .(6) Because X â contains tokens from all three modalities, the attention mechanism learns to exchange information across CSI, geometry and motion, for example by allowing channel tokens to attend to nearby building structures or trajectory tokens. Modality-specific feed-forward experts Rather than a single shared feed-forward layer, 3 feed-forward sub-network, one for each modality, in each Transformer block is implemented as three parallel experts. After a second normalization, the sequence X Ⲡâ is partitioned into CSI, 3D point cloud and user trajectory segments according to the stored lengths: X CSI â , X PC â , X POS â = Split ( LN(X Ⲡâ );N CSI ,N PC ,N POS ) .(7) Each segment is then passed through its own two-layer feed-forward expert, Ě X CSI â = f CSI X CSI â , Ě X PC â = f PC X PC â , Ě X POS â = f POS X POS â .(8) where each feed-forward expert f is a position-wise non-linear mapping with separate parameters. The updated segments are concatenated back to their original order and combined with a residual connection: X â+1 = X Ⲡâ + Concat Ě X CSI â , Ě X PC â , Ě X POS â .(9) This âshared-attention plus modality-expertâ design can be viewed as a multi-modal mixture-of-experts architecture with deterministic routing based on modality identity. It allows global context modeling to be shared across modalities, while maintaining specialized pathways tuned to the statistics and physical constraints of CSI, 3D point cloud, and use trajectory. 4.2.3 Encoder and predictor configurations Both the online encoder f θ and the predictor g Ď are built from the shared-attention plus modality-expert Transformer blocks described in the previous section, and share the same ViT-Small hyper-parameters (D=384, L e =L p =12). Despite this architec- tural symmetry, the two networks serve distinct roles: the encoder operates directly on 17 the embedded input tokens at the visible positions I enc and must learn rich, general- purpose representations of the observed context; the predictor, by contrast, receives the encoderâs output together with learnable mask tokens at positions I pred and is tasked with inferring the latent content of the masked regions. Both the encoder and predictor share an identical architecture, consisting of 12 MMoE Transformer layers with an embedding dimension of 384, 6 attention heads, and a head dimension of 64. 4.3 Pre-training details The WWM was pre-trained in a self-supervised manner using JEPA described above. In this setting, the model is presented with partially observed multi-modal inputs and is trained to predict the latent representations of the masked regions in a shared embed- ding space, rather than reconstructing raw measurements. This approach encourages the encoderâpredictor pair to capture the underlying physical structure of the wireless environment. 4.3.1 Data preprocessing Raw CSI tensors acquired from either simulation or the 6G prototype system are preprocessed into a numerically stable representation to facilitate large-scale self-supervised training. Samples containing zero-valued elements are first discarded. The remaining CSI values span a wide dynamic range, which can hinder training stability. To compress this range while preserving sign information, we apply a signed log transform: Ě H = sign(H)¡ log 1 +|H|/Îľ ,(10) where H denotes the CSI tensor and Îľ = 10 â7 is a small scaling constant that amplifies near-zero magnitudes before the logarithm, ensuring fine-grained distinctions among weak signal components are retained. Finally, meanâvariance standardization is applied to rescale the dataset to zero mean and unit standard deviation. This pipeline converts heterogeneous raw inputs into clean, standardized tensors with controlled dynamic range, which is critical for stable JEPA pre-training across diverse cities and user speeds. 4.3.2 Masking strategies for CSI and trajectories To expose the model to complementary forms of partial observability, three masking configurations were interleaved during pre-training. All configurations operate on the unified token sequence obtained by concatenating CSI, point-cloud and trajectory tokens, but place the emphasis on different aspects of the wireless scenario. Fine-grained CSI masking: In the first configuration, the CSI volume is parti- tioned into a 3D grid of spatiotemporal patches along time, frequency and antenna (or spatial) dimensions. For each sample, several relatively small 3D blocks are sampled at random within this grid, and all CSI patches inside these blocks are designated as masked. Concretely, we sample 8 blocks per clip, each with a temporal extent covering 50% of the tubeletized time axis and a spatial extent corresponding to approximately 18 15% of the CSI patch grid in both spatial dimensions. The remaining CSI patches, together with all point-cloud and trajectory tokens, form the visible context. This configuration yields a moderate masking ratio over CSI and encourages the model to reconstruct fine-scale multipath structure when sufficient local context is available. Coarse CSI masking: In the second configuration, fewer but substantially larger 3D blocks are masked in the CSI grid, producing sizeable âholesâ in timeâfrequencyâspace that must be inferred from the remaining CSI context. Here we mask 2 blocks per clip, each covering 50% of the tubeletized time axis and roughly 70% of the patch grid in each spatial dimension. Point-cloud and trajectory tokens remain fully visible. Compared with the fine-grained configuration, this setting places more emphasis on long-range dependencies and global consistency. Trajectory masking: The third configuration targets user trajectory inference. In this setting, all CSI patches and point-cloud tokens remain fully visible. Instead, the entire trajectory token sequence is masked. The model thus receives a complete description of the channel evolution and environment, but no explicit user coordinates, and is required to reconstruct the latent embeddings associated with the trajectory. This configuration encourages the model to internalize the inverse relationship from channel and geometry back to user motion patterns. Across a mini-batch, the three configurations are sampled with equal probabil- ity, and the total loss is obtained by averaging the latent-space prediction losses from each configuration. This multi-task JEPA objective ensures that the same pre- trained model is simultaneously optimized for both channel completion and trajectory inference under different visibility patterns. 4.3.3 Pre-training configurations The WWM model, including encoder and predictor, was pre-trained on a simulated large-scale multi-modal dataset (Extended Data Table 3), with each sample containing CSI tokens of 16 time steps, 3D point cloud tokens centered around user location and synchronized user trajectory vector. The model was trained with a global batch size of 128 using the AdamW optimizer and the L1 loss function. Following a cosine learning-rate schedule, the learning rate was linearly warmed up from 1.0Ă 10 â5 to a peak value of 2.0 Ă 10 â5 over the first 2 epochs, and then decayed by a cosine scheduler to a final learning rate of 1.0Ă 10 â5 by the end of training. Weight decay was fixed at 0.04 throughout. The target encoder was updated with a momentum of 0.9925 (applied to both the encoder and predictor branches), providing a slowly evolving target network that stabilizes training in the JEPA setting. Pre-training was run for 16 epochs with a mixed-precision bfloat16 data type. Notably, the full pre- training workflow on the complete multi-modal dataset was completed in 87 hours on a single NVIDIA RTX 4090 24GB consumer-grade GPU, demonstrating an extremely low training computational barrier compared to large-scale foundation models in other domains. 19 5 Data and model availability Once the paper is accepted, all datasets and model checkpoints used in this study will be publicly available via https://zenodo.org/communities/wwm/. 6 Code availability Once the paper is accepted, the model pre-training, downstream task training and testing code will be publicly available via GitHub at https://github.com/Wireless- World-Model/WWM-V1. 20 Extended data Extended Data Table 1 Table 1 Extended Data Table 1 â Simulation Dataset Configuration SymbolDescriptionValue f c Center frequency2.6 GHz âfSubcarrier spacing15 kHz FNumber of subcarriers96 N sb Number of subbands8 F sb Subcarriers per subband12 âtTemporal sampling interval5 ms TTemporal samples per sample16 N t BS antenna ports32 = 4 horizontalĂ 4 verticalĂ 2 polarization N r UE antenna ports4 = 2 horizontalĂ 1 verticalĂ 2 polarization N Ⲡr Effective receiveâfrequency dimension32 = N r Ă N sb d ant Antenna element spacing0.5 wavelength N scen Urban scenarios5 Extended Data Table 2 Table 2 Extended Data Table 2 â Real Measured Dataset Configuration SymbolDescriptionValue f c Center frequency6.6 GHz âfSubcarrier spacing120 kHz N sb Number of subcarriers3333 N RB Number of PRBs264 N RBsample Number of PRBs per sample8 âtTemporal sampling interval10 ms TTemporal samples per sequence16 N t BS antenna ports32 = 16 horizontalĂ 1 verticalĂ 2 polarization N r UE antenna ports4 = 4 horizontalĂ 1 verticalĂ 1 polarization N Ⲡr Effective receiveâfrequency dimension32 = N r Ă N RBsample d ant Antenna element spacing0.5 wavelength 21 Extended Data Table 3 Table 3 Extended Data Table 3 â Summary of simulation scenarios and dataset parameters. Scenario ID CityBS Position UE SpeedData Volume 1 Munich BS0 5 km/h 2048 (Trajectories) Ă 16 (Time steps) Ă 7 (runs) per scenario 230 km/h 360 km/h 4 BS1 5 km/h 530 km/h 660 km/h 7 BS2 5 km/h 830 km/h 960 km/h 10 Etoile BS0 5 km/h 1130 km/h 1260 km/h 13 BS1 5 km/h 1430 km/h 1560 km/h 16 Beijing Forbidden City BS05 km/h 17BS15 km/h 18BS25 km/h 19 Beijing CBD BS0 5 km/h 2030 km/h 2160 km/h 22 BS1 5 km/h 2330 km/h 60 km/h 22 Extended Data Table 4 Table 4 Extended Data Table 4 â Summary of generalization scenarios and dataset parameters. This portion of the dataset is specifically designed to evaluate the modelâs robustness to unseen velocities and urban environments. Scenario ID Generalization Type BS Position UE SpeedData Volume 1 Velocity Generalization (Beijing CBD) BS0 40 km/h 2048 (Trajectories)Ă16 (Time steps) per scenario 270 km/h 3 BS1 40 km/h 470 km/h 5 City Generalization (Wall Street) BS0 5 km/h 630 km/h 760 km/h Extended Data Table 5 Table 5 Extended Data Table 5 â CSI temporal prediction performance of WWM compared to other baselines across in-distribution and generalization scenarios. A Metric of SGCS of predicted CSI is listed Scenario WWMWiFoLSTM T=15T=16Avg.T=15T=16Avg.T=15T=16Avg. CBD0.7960.8070.802â0.7360.6970.6740.685 Etoile0.9480.9410.944â0.8440.8960.8900.893 Forbidden City0.9590.9570.958â0.8630.8610.8590.860 Munich0.9230.9130.918â0.7430.8790.8820.881 Velocity Generalization0.7670.7860.776â0.6770.6940.6710.682 City Generalization0.9260.9060.916â0.7630.5900.5850.588 23 Extended Data Table 6 Table 6 Extended Data Table 6 â Comprehensive comparison of CSI compression and feedback performance (SGCS) across varying compression ratios and scenarios. Performance is evaluated for WWM and baselines (QCR-NET, CR-NET) under in-distribution and generalization regimes. ModelScenario1/1024 1/512 1/256 1/128 WWMCBD0.6516 0.6666 0.6610 0.7522 Etoile0.8027 0.8027 0.8460 0.9556 Forbidden City0.7901 0.8014 0.8244 0.9240 Munich0.7704 0.7652 0.7995 0.9467 Velocity Generalization 0.6468 0.6623 0.6598 0.7454 City Generalization0.6214 0.6187 0.6535 0.8961 QCR-NET CBD0.58720.61090.61720.6652 Etoile0.60840.75100.81180.8862 Forbidden City0.39980.48470.53670.6220 Munich0.66510.7454 0.8150 0.8940 Velocity Generalization0.59350.61220.62190.6693 City Generalization0.48280.54020.57810.7016 CR-NETCBD0.37320.52170.61420.6273 Etoile0.29410.43510.60540.7361 Forbidden City0.19240.31580.43210.4744 Munich0.28930.48010.64320.7438 Velocity Generalization0.37550.52970.62140.6342 City Generalization0.25650.36440.52080.5753 24 Extended Data Table 7 Table 7 Extended Data Table 7 â beam prediction performance of WWM compared to other baselines across in-distribution and generalization scenarios. The Metrics of Top-1 accuracy of predicted DFT codeword and their respective beam gain ratio are used Scenario WWMLWMSPMM Top-1 AccSE ratioTop-1 AccSE ratioTop-1 AccSE ratio CBD0.9831.0000.9270.9950.9020.987 Etoile0.9170.9800.8360.9440.7780.931 Forbidden City0.9060.9810.8540.9700.8280.963 Munich0.9150.9840.8360.9560.7670.926 Velocity Generalization0.9780.9990.9260.9950.8570.972 Extended Data Table 8 Table 8 Extended Data Table 8 â User localization performance of WWM compared to other baselines across in-distribution and generalization scenarios. A Metric of mean average error of 2D distance is used ScenarioWWMdeep-CNN CBD1.2127432.2889 Etoile2.9829865.2257 Forbidden City2.8682116.0771 Munich2.7184764.8768 Velocity Generalization1.2269491.9978 Extended Data Table 9 Table 9 Extended Data Table 9 â Ablation study comparing the full multimodal WWM with its unimodal variant across scenarios, reporting per-timestep SGCS results (T =15, T =16) and their average. Scenario WWMWWM-Unimodal T=15T=16Avg.T=15T=16Avg. CBD0.7960.8070.8020.6870.6870.687 Etoile0.9480.9410.9440.9350.9220.928 Forbidden City0.9590.9570.9580.9230.9120.918 Munich0.9230.9130.9180.9010.8790.890 Velocity Generalization0.7670.7860.7760.6530.6620.657 City Generalization0.9260.9060.9160.9270.9040.915 25 Extended Data Table 10 Table 10 Extended Data Table 10 â CSI frequency-domain prediction performance of WWM compared to other baselines by metric of NMSE and SGCS. Scenario WWMWiFoC-Mixer NMSESGCSNMSESGCSNMSESGCS PRB40.2120.9320.2520.9220.2960.918 PRB50.2180.9180.2610.8950.3080.892 PRB60.2600.9140.3090.8820.3490.875 PRB70.2430.9020.2870.8580.3340.843 Avg.0.2330.9170.2770.8890.3220.882 26 References [1] Liu, G., Huang, Y., Li, N., et al.: Vision, requirements and network architecture of 6G mobile network beyond 2030. China Communications 17(9), 92â104 (2020) https://doi.org/10.23919/JCC.2020.09.008 [2] Wang, C.-X., You, X., Gao, X., et al.: On the Road to 6G: Visions, Requirements, Key Technologies, and Testbeds. IEEE Communications Surveys & Tutorials 25(2), 905â974 (2023) https://doi.org/10.1109/COMST.2023.3249835 [3] Shannon, C.E.: A mathematical theory of communication. The Bell System Tech- nical Journal 27(3), 379â423 (1948) https://doi.org/10.1002/j.1538-7305.1948. tb01338.x [4] Wang, Z., Zhang, J., Du, H., Niyato, D., Cui, S., Ai, B., Debbah, M., Letaief, K.B., Poor, H.V.: A tutorial on extremely large-scale mimo for 6g: Fundamentals, signal processing, and applications. IEEE Communications Surveys & Tutorials 26(3), 1560â1605 (2024) https://doi.org/10.1109/COMST.2023.3349276 [5] Ziao, Q.: A review of codebooks for csi feedback in 5g new radio and beyond. China Communications 22(2), 112â127 (2025) https://doi.org/10.23919/JCC.ja. 2023-0117 [6] Yu, B., Qian, C., Lin, P., Lee, J., Li, Q., Park, S., Kim, S., Yoon, C., Hu, S., Liu, L.: Light-weight ai enabled non-linearity compensation leveraging high order modulations. IEEE Transactions on Communications 72(1), 539â552 (2024) https://doi.org/10.1109/TCOMM.2023.3321735 [7] Shi, Y., Lian, L., Shi, Y., et al.: Machine learning for large-scale optimization in 6g wireless networks. IEEE Communications Surveys & Tutorials 25(4), 2088â2132 (2023) https://doi.org/10.1109/COMST.2023.3300664 [8] Farhadi, H., Banerjee, B., Berkvens, R., et al.: 6G AI-Driven Air Interface â Hexa-X-I View. IEEE Communications Magazine 63(10), 118â125 (2025) https: //doi.org/10.1109/MCOM.001.2400394 . Accessed 2026-01-22 [9] Hoydis, J., Aoudia, F.A., Valcarce, A., et al.: Toward a 6G AI-Native Air Inter- face. IEEE Communications Magazine 59(5), 76â81 (2021) https://doi.org/10. 1109/MCOM.001.2001187 . Accessed 2026-01-22 [10] Zheng, X., Xiao, H., Jin, S., et al.: Ai-native 6g physical layer with cross-module optimization and cooperative control agents. IEEE Journal on Selected Areas in Communications, 1â1 (2026) https://doi.org/10.1109/JSAC.2026.3652936 [11] Moor, M., Banerjee, O., Abad, Z.S.H., Krumholz, H.M., Leskovec, J., Topol, E.J., Rajpurkar, P.: Foundation models for generalist medical artificial intelligence. Nature 616(7956), 259â265 (2023) 27 [12] Abramson, J., Adler, J., Dunger, J., Evans, R., Green, T., Pritzel, A., Ron- neberger, O., Willmore, L., Ballard, A.J., Bambrick, J., et al.: Accurate structure prediction of biomolecular interactions with alphafold 3. Nature 630(8016), 493â500 (2024) [13] He, Y., Fang, P., Shan, Y., Pan, Y., Wei, Y., Chen, Y., Chen, Y., Liu, Y., Zeng, Z., Zhou, Z., et al.: Generalized biological foundation model with unified nucleic acid and protein language. Nature Machine Intelligence, 1â12 (2025) [14] Binz, M., Akata, E., Bethge, M., Br Ěandle, F., Callaway, F., Coda-Forno, J., Dayan, P., Demircan, C., Eckstein, M.K., Ě Eltet Ěo, N., et al.: A foundation model to predict and capture human cognition. Nature, 1â8 (2025) [15] Shao, J., Tong, J., Wu, Q., Guo, W., Li, Z., Lin, Z., Zhang, J.: Wirelessllm: Empowering large language models towards wireless intelligence. Journal of Com- munications and Information Networks 9(2), 99â112 (2024) https://doi.org/10. 23919/JCIN.2024.10582827 [16] Liu, B., Liu, X., Gao, S., et al.: Llm4cp: Adapting large language models for channel prediction. Journal of Communications and Information Networks 9(2), 113â125 (2024) https://doi.org/10.23919/JCIN.2024.10582829 [17] Cui, Y., Guo, J., Wen, C.-K., et al.: Exploring the Potential of Large Language Models for Massive MIMO CSI Feedback (2025) https://doi.org/10.48550/arXiv. 2501.10630 [18] Zheng, T., Dai, L.: Large Language Model Enabled Multi-Task Physical Layer Network (2025) https://doi.org/10.48550/arXiv.2412.20772 . arXiv:2412.20772 [cs]. Accessed 2025-04-19 [19] Wen, Y., Chen, X., Zhang, M., et al.: ICWLM: A Multi-Task Wireless Large Model via In-Context Learning. arXiv (2025). https://doi.org/10.48550/arXiv. 2507.18167 . http://arxiv.org/abs/2507.18167 [20] Noh, H., Shim, B., Yang, H.J.: Adaptive resource allocation optimization using large language models in dynamic wireless environments. IEEE Transactions on Vehicular Technology 74(10), 16630â16635 (2025) https://doi.org/10.1109/TVT. 2025.3572440 [21] Zhang, C., Zhang, H., Qiao, J., Li, Z., Alouini, M.-S.: TIDES: Traffic Intelli- gence with DeepSeek Enhanced Spatial Temporal Prediction. IEEE Journal on Selected Areas in Communications, 1â1 (2025) https://doi.org/10.1109/JSAC. 2025.3643397 [22] Liu, B., Gao, S., Liu, X., et al.: Wifo: Wireless Foundation Model for Channel Prediction. SCIENCE CHINA Information Sciences 68(6) (2025) https://doi. org/10.1007/s11432-025-4349-0 28 [23] Cheng, X., Liu, B., Liu, X., Liu, E., Huang, Z.: Foundation model empow- ered synesthesia of machines (som): Ai-native intelligent multi-modal sensing- communication integration. IEEE Transactions on Network Science and Engi- neering 13, 762â782 (2026) https://doi.org/10.1109/TNSE.2025.3587238 [24] Yang, T., Zhang, P., Zheng, M., et al.: WirelessGPT: A Generative Pre-Trained Multi-Task Learning Framework for Wireless Communication. IEEE Network 39(5), 58â65 (2025) https://doi.org/10.1109/MNET.2025.3579496 . Accessed 2026-01-31 [25] He, D., Ai, B., Guan, K., et al.: The design and applications of high-performance ray-tracing simulation platform for 5g and beyond wireless communications: A tutorial. IEEE Communications Surveys & Tutorials 21(1), 10â27 (2019) https: //doi.org/10.1109/COMST.2018.2865724 [26] Ha, D., Schmidhuber, J.: Recurrent world models facilitate policy evolution. In: Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., Garnett, R. (eds.) Advances in Neural Information Processing Systems, vol. 31, p. 2451â 2463. Curran Associates, Inc., Red Hook, NY, USA (2018) [27] LeCun, Y.: A path towards autonomous machine intelligence. OpenReview 1(1), 1â64 (2022) [28] Ren, X., Lu, Y., Cao, T., Gao, R., Huang, S., Sabour, A., Shen, T., Pfaff, T., Wu, J.Z., Chen, R., et al.: Cosmos-drive-dreams: Scalable synthetic driving data generation with world foundation models. arXiv preprint arXiv:2506.09042 (2025) [29] Hoydis, J., et al.: Sionna: An open-source library for next-generation physical layer research. IEEE Journal on Selected Areas in Communications (2022) [30] Bardes, A., Garrido, Q., Ponce, J., Chen, X., Rabbat, M., LeCun, Y., Assran, M., Ballas, N.: Revisiting Feature Prediction for Learning Visual Representations from Video (2024). https://arxiv.org/abs/2404.08471 [31] Assran, M., Bardes, A., Fan, D., Garrido, Q., Howes, R., Muckley, M., Rizvi, A., Roberts, C., Sinha, K., Zholus, A., et al.: V-jepa 2: Self-supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985 (2025) [32] Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O.K., Singhal, S., Som, S., et al.: Image as a foreign language: Beit pretraining for vision and vision-language tasks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 19175â19186 (2023) [33] He, K., Chen, X., Xie, S., Li, Y., Doll Ěar, P., Girshick, R.: Masked autoencoders 29 are scalable vision learners. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 16000â16009 (2022) [34] Maaten, L., Hinton, G.: Visualizing data using t-sne. Journal of machine learning research 9(11) (2008) [35] Alikhani, S., Charan, G., Alkhateeb, A.: Large wireless model (lwm): A founda- tion model for wireless channels. arXiv preprint arXiv:2411.08872 (2024) [36] Graves, A.: Long short-term memory. Supervised sequence labelling with recur- rent neural networks, 37â45 (2012) [37] Zhang, X., Lu, Z., Zeng, R., Wang, J.: Quantization adaptor for bit-level deep learning-based massive mimo csi feedback. IEEE Transactions on Vehicular Technology 73(4), 5443â5453 (2023) [38] Lu, Z., Wang, J., Song, J.: Multi-resolution csi feedback with deep learning in massive mimo system. In: ICC 2020-2020 IEEE International Conference on Communications (ICC), p. 1â6 (2020). IEEE [39] Alrabeiah, M., Alkhateeb, A.: Deep learning for mmwave beam and blockage pre- diction using sub-6 ghz channels. IEEE Transactions on Communications 68(9), 5504â5518 (2020) [40] Chen, Z., Zhang, Z., Yang, Z., Liu, L.: Channel mapping based on inter- leaved learning with complex-domain mlp-mixer. IEEE Wireless Communications Letters 13(5), 1369â1373 (2024) [41] OpenStreetMap Contributors: OpenStreetMap. https://w.openstreetmap.org. Accessed 2024 (2024) [42] Nimier-David, M., Vicini, D., Zeltner, T., Jakob, W.: Mitsuba 2: A retargetable forward and inverse renderer. ACM Transactions on Graphics (2019) [43] Yu, X., Tang, L., Rao, Y., Huang, T., Zhou, J., Lu, J.: Point-bert: Pre-training 3d point cloud transformers with masked point modeling. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 19313â 19322 (2022) [44] Gu, J., Wang, Z., Kuen, J., Ma, L., Shahroudy, A., Shuai, B., Liu, T., Wang, X., Wang, G., Cai, J., et al.: Recent advances in convolutional neural networks. Pattern recognition 77, 354â377 (2018) 30 Supplementary Note Supplementary Note 1 Channel Prediction Task Details We evaluate the WWM as a sequence model for short-horizon channel prediction. Here we formulate channel prediction as a downstream task on top of the pre-trained WWM: given history CSI together with the corresponding environment 3D point cloud and user trajectory, the model is asked to infer the CSI at a set of future time steps. In this setting, the WWM encoder and predictor operate exactly as in the pre-training stage, but their parameters are kept frozen. A multi-time step CSI sample, along with the associated point cloud and user trajectory, is fed into the encoder. Only the first 14 time steps in the sample are treated as visible context, while the channel tokens corresponding to subsequent T pred = 2 time steps within the sample are designated as masked positions. The predictor receives the context embeddings and a set of learnable mask tokens at these future positions, and produces latent tokens that represent the predicted CSI in the embedding space. These predicted tokens, restricted to the CSI modality, are then passed to a dedicated channel decoder that maps them back to the complex CSI tensor on the original timeâfrequencyâspace format. The channel decoder is implemented as a compact transformer-based network spe- cialized for CSI reconstruction. It consists of a stack of Transformer blocks operating purely on CSI tokens, followed by a projection head that reshapes the output into a tensor of shape (2,T pred ,H,W ). Here, the first dimension corresponds to real and imaginary parts; T pred denotes the number of future channel time steps being pre- dicted, H corresponds to the number of base-station antennas (32 in our dataset), and W denotes the joint frequencyâantenna dimension defined by the product of user- side antenna elements and subband groups (here 4Ă 8). This structure allows the decoder to reconstruct the full complex CSI on timeâfrequencyâspace format from the predicted latent tokens. In our implementation, the decoder is a 6-layer transformer trained with AdamW and a learning rate of 2Ă 10 â5 . The decoder is trained with a complex-valued reconstruction loss that explicitly separates magnitude and phase components of the channel, while also incorporating a structural similarity objective. Given predicted and ground-truth channels in the form Ë Y, Y â R 2ĂT pred ĂHĂW (real and imaginary parts separated in the first dimension), we form complex tensors Ë H, H â C T pred ĂHĂW and their magnitudes and phases of each element. The CSI loss combines a raw mean-squared error in the realâimaginary plane, a magnitude loss, a phase-consistency term, and an SGCS regularization term: L CSI = 1 N Ë Yâ Y 2 2 |z raw MSE +Îą 1 N | Ë H|â|H| 2 2 |z magnitude loss +β 1â 1 N X n cos( Ë Ď n â Ď n ) | z phase-consistency loss +Îł 1 T pred T pred X t=1 L SGCS,t |z average SGCS loss . (11) Here N = 2Ă T Ă H Ă W denotes the number of elements in Ë Y, and Ë Ď n and Ď n represent the predicted and true phases at element n ⤠N . In implementation, L SGCS = 1â SGCS is computed on the final predicted time step by reshaping the reconstructed channel into subcarrier groups and measuring the similarity between the 31 dominant spatial singular vectors of the predicted and ground-truth CSI. Specifically, for each subcarrier k, we reshape the channel matrix as H k â C N ue ĂN bs and perform singular value decomposition (SVD): H k = U k ÎŁ k V H k ,(12) where the first right singular vector v k (corresponding to the largest singular value) captures the dominant spatial direction. Let Ë v k and v k denote the dominant right singular vectors of the predicted and ground-truth CSI, respectively. The SGCS of a single CSI time step is computed as the average normalized inner product across all K subcarriers: SGCS = 1 K K X k=1 Ë v H k v k âĽ Ë v k ⼠2 âĽv k ⼠2 .(13) In our training setup, the CSI-loss weights are progressively adjusted across epochs to balance stable amplitude reconstruction, phase alignment, and SGCS optimization: epochs 1â10 use Îą = 1.0 and β = 0.2; epochs 11â15 use Îą = 1.0 and β = 0.5; epochs 16â20 use Îą = 1.0 and β = 1.0; and epochs 21â25 use Îą = 1.0, β = 0.2, and Îł = 1.0. During evaluation, we report SGCS to quantify the structural similarity between the reconstructed and ground-truth CSI across the antenna array and frequency bands. To assess performance and generalization, we compare the WWM against two baseline models: Wifo[22]âa foundation model for wireless channel prediction, and Long short- term memory (LSTM)[36]. Supplementary Note 2 Channel Compression and Feedback Task Details We evaluate the WWM as a learned channel compression and feedback module. Here, channel compression is formulated as a downstream task operating on the latent rep- resentations extracted by the pre-trained WWM. Out of a 16-timestep CSI sample (processed as 8 temporal tubelets), only the 4th tubelet (corresponding to time steps 7 and 8) is kept unmasked and fed into WWM encoder (together with the corresponding 3D point cloud and User trajectory), which produces N = 64 continuous CSI latent tokens Z csi â R NĂD in the embedding space, each with a dimension of D = 384. A dedicated lightweight compression head, implemented as CRNetTokens, is attached on top of these tokens. The compression and quantization process can be formulated as: Z comp =Q Îź,b ( f reduce (Z csi ) ) (14) where f reduce (¡) applies a dimensionality reduction layer that shrinks the embedding dimension by a configurable reduction factor r (e.g., r = 96), followed by a token- mixing stage. The function Q Îź,b (¡) represents a Îź-law scalar quantizer that maps the continuous vectors to 2 b discrete levels (e.g., b = 4 bits) and is optimized via a straight-through estimator during training. A decompressor f expand (¡) subsequently expands the discrete feedback payload back to the original token dimension D. These reconstructed tokens are then passed to 32 a channel decoderDâa Vision Transformer jointly fine-tuned with the compressorâto recover the complex channel tensor the original timeâfrequencyâspace format: Ë Y =D ( f expand (Z comp ) ) (15) The compressor and decoder are trained end-to-end using the AdamW optimizer with a learning rate of 1Ă 10 â4 and a weight decay of 0.01. The training objective is a complex-valued CSI reconstruction loss that combines a raw mean-squared error, a magnitude loss (weight Îą = 1.0), and a phase-consistency term (weight β = 0.5). We evaluate the WWM-based compressor against strong neural baselines (CR- NET [38] and QCR-NET[37]) under matched compression budgets. Generalization is assessed across three scenarios: in-distribution (seen cities and velocities), velocity gen- eralization (unseen speeds in seen environments), and city generalization (completely unseen urban layouts). By compressing in the semantically rich multimodal latent space rather than the raw channel domain, the WWM maintains high reconstruction fidelity and exhibits strong robustness to unseen speeds and propagation conditions. Supplementary Note 3 Beam prediction task details The beam prediction task adopts the same urban scenarios and user velocity ranges as those used for pre-training and other downstream tasks. For each user trajectory, CSI is collected at two distinct central frequency simultaneously: a Sub-6GHz band (2.6GHz) and an upper 6GHz (U6G) band (6.62505GHz). At every sampled UE position, the Sub-6GHz CSI tensor X sub-6 is used as input, together with the corresponding 3D point cloud and user trajectory. From the concurrent U6G channel, the optimal Precoding matrix indicator (PIM) index b â is determined as the beam that maximizes the received power over a predefined Type I Single-Panel Codebook of K beams. Thus the task learns the mapping: f beam : X sub-6 ââ b â â1, 2,...,K, where f beam denotes the prediction model and X sub-6 â R 2ĂTĂN t ĂN Ⲡr is the pro- cessed low-frequency CSI tensor. The mapping is learned using the pre-trained WWM as a frozen feature extractor, followed by a task-specific 1-layer attentive classifier. The attentive classifier takes the encoded feature tokens and outputs a probability distribution over the K beams, trained with the cross-entropy loss: L beam =â 1 N N X i=1 K X k=1 y i,k log(p i,k ), where N is the batch size, y i,k is the one-hot encoding of the ground-truth beam index for the i-th sample, and p i,k is the predicted probability for the k-th beam. Throughout training, only the attentive classifier parameters are updated, while the WWM backbone remains fixed. To assess performance and generalization, we compare the WWM-based predictor against two baseline models: the LWM[35]âa general-purpose foundation model for wireless channelsâand the Sub-6-Preds-mmWave (SPMM)[39]âa deep neural net- work specifically designed for cross-band beam prediction. We measure the Top-1 33 classification accuracy, as well as the achieved beam gain relative to the optimal beam. The beam gain ratio for a set of M test samples is computed as: R BG = 1 M M X j=1 H U 6G w (j) b H U 6G w (j) b â , where H U 6G is the last timestep CSI of j-th sample in U6G band, w (j) b is the PMI corresponding to predicted beam index for the j-th sample, and w (j) b â is the PMI corresponding to the theoretical optimal beam. This metric directly reflects how closely the modelâs beam selections approach the ideal link performance. Supplementary Note 4 User localization task details We further evaluate the WWMâs capability for high-precision user localization. In this downstream task, the model is required to infer the 2D geographical coordinates (x,y) of a UE based on its CSI as well as the corresponding 3D cloud point. Similar to the channel prediction task, the pre-trained WWM backbone remains frozen to leverage its internalized physical representations. The explicit user trajectory tokens are intention- ally omitted from WWM input, compeling the model to derive positional information solely from the interaction between CSI patterns and the geometric structure of the environment. The resulting latent tokens from the WWM encoder, which encapsulate joint EM-geometric features, are then fed into a dedicated attentive regression head. The attentive regression head is designed to regress the sequence of latent embed- dings into a 2D coordinate Ë p = ( Ë x, Ë y). The localization head is trained to minimize the Euclidean distance between the predicted coordinates and the ground-truth posi- tion p = (x,y) corresponding to the final position of the UE trajectory. The objective function is defined by the mean squared error (MSE) loss. we quantify performance using the CDF of the absolute localization error. We compare the WWM-based localization framework against a CNN-based baseline[44]. Specifically, the baseline consists of a one-layer temporal fusion module, followed by a ResNet-18 backbone (comprising 17 convolutional layers and 1 fully connected layer), and a final linear regression layer that directly outputs the 2D coor- dinates. The baseline model is trained from scratch on the same labeled dataset and takes raw CSI tensors as input. Both models are evaluated across the same urban lay- outs used in the general benchmark to assess their robustness under environmental shifts. Our results indicate that, by extracting high-level semantic features from the joint EM-geometric space, the WWM-based approach significantly outperforms the conventional regression baseline. Supplementary Note 5 CSI frequency-domain prediction based on SRS measurement task details We introduce an CSI frequency-domain prediction task to evaluate the modelâs capa- bility to reduce uplink measurement overhead under realistic deployment conditions. The data are collected from a 6G prototype system in an outdoor slow-mobility sce- nario. In practical 6G systems, sounding reference signals (SRS) are transmitted over 34 multiple physical resource blocks (PRBs) in frequency domain for uplink channel esti- mation at the base station. Dense PRB-level measurements, however, incur substantial signaling overhead. This task aims to infer CSI on on unmeasured SRS PRBs from partially observed PRBs within the same time window. The prototype system operates at 6.6 GHz with 400 MHz bandwidth and 120 kHz subcarrier spacing. Each PRB contains 12 subcarriers (1.44 MHz bandwidth). The raw measurements comprise 264 PRBs, which are partitioned into 33 samples of 8 con- secutive PRBs. For each sample, the first four PRBs serve as visible context, and the remaining four PRBs are designated as prediction targets. Each sample corresponds to a continuous UE trajectory containing 16 consecutive CSI time steps, which are jointly modeled to capture temporal channel evolution. The task is formulated as learning the mapping: f CSI : X vis â X mask , where X vis and X mask denote the CSI tensors of the observed and masked PRBs, respectively. CSI frequency-domain prediction is implemented as a downstream task on top of the pre-trained WWM. The encoder and predictor retain the same architecture as in pre-training. In practice, training on the measured dataset is conducted in two stages. First, the encoder and predictor are jointly fine-tuned to adapt the pre-trained representations to real-world data. Then, the encoder is frozen, and the predictor is further optimized together with a lightweight Transformer-based channel decoder. A full 16-timestep trajectory sample, together with its 3D point cloud and user trajec- tory tokens, is fed into the encoder. Within each time step, only the first four PRBs are provided as input, while the remaining four PRBs are withheld as prediction targets. The predictor produces latent representations corresponding to the full CSI matrix, rather than only masked positions. These representations are passed to the channel decoder, which operates solely on CSI tokens and outputs a complete CSI esti- mate. A final projection head reshapes the output into a tensor of size (2,T,H,W ), where the first dimension corresponds to real and imaginary parts; T = 16 is the number of temporal samples; H = 32 is the number of base-station antennas; and W denotes the joint frequencyâantenna dimension (4 UE antennas Ă 8 PRBs). The model is trained using mean squared error between the predicted and ground-truth CSI over the full frequency band, and performance is evaluated with NMSE and SGCS. By reconstructing unobserved PRBs from partial observations, this task reflects the modelâs ability to exploit frequency-domain channel correlations in real-world scenar- ios and provides a practical mechanism for reducing SRS overhead in operational 6G systems. To assess performance, we compare the WWM against baseline foundation model WiFo[22] and task-specific model C-Mixer [40]. To ensure a rigorous and fair comparison, WiFo adopts the same training strategy as WWMâinitial pre-training on simulated dataset followed by specialized fine-tuning on field-measured SRS data. In contrast, the task-specific C-Mixer is trained from scratch directly on the measured SRS dataset. 35