Paper deep dive
Skillful Kilometer-Scale Regional Weather Forecasting via Global and Regional Coupling
Weiqi Chen, Wenwei Wang, Qilong Yuan, Lefei Shen, Bingqing Peng, Jiawei Chen, Bo Wu, Liang Sun
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/31/2026, 2:22:38 AM
Summary
The paper introduces ScaleMixer, a global-regional coupling framework for kilometer-scale weather forecasting. It integrates a pretrained Transformer-based global model with a high-resolution regional network using a bidirectional coupling module that adaptively identifies meteorologically critical regions to resolve multiscale interactions, outperforming existing NWP and AI baselines.
Entities (4)
Relation Signals (3)
Global Model → trainedon → ERA5
confidence 98% · The global model M global is pretrained on ERA5 reanalysis
ScaleMixer → couples → Global Model
confidence 95% · synergistically couples a pretrained Transformer-based global model with a high-resolution regional network via a novel bidirectional coupling module, ScaleMixer.
ScaleMixer → enables → Cross-scale feature interaction
confidence 92% · enables cross-scale feature interaction through dedicated attention mechanisms.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Data-driven weather models have advanced global medium-range forecasting, yet high-resolution regional prediction remains challenging due to unresolved multiscale interactions between large-scale dynamics and small-scale processes such as terrain-induced circulations and coastal effects. This paper presents a global-regional coupling framework for kilometer-scale regional weather forecasting that synergistically couples a pretrained Transformer-based global model with a high-resolution regional network via a novel bidirectional coupling module, ScaleMixer. ScaleMixer dynamically identifies meteorologically critical regions through adaptive key-position sampling and enables cross-scale feature interaction through dedicated attention mechanisms. The framework produces forecasts at $0.05^\circ$ ($\sim 5 \mathrm{km}$ ) and 1-hour resolution over China, significantly outperforming operational NWP and AI baselines on both gridded reanalysis data and real-time weather station observations. It exhibits exceptional skill in capturing fine-grained phenomena such as orographic wind patterns and Foehn warming, demonstrating effective global-scale coherence with high-resolution fidelity. The code is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2603.28173v1
- Canonical: https://arxiv.org/abs/2603.28173v1
Trouble viewing inline? Open PDF directly →
Full Text
68,196 characters extracted from source content.
Expand or collapse full text
Skillful Kilometer-Scale Regional Weather Forecasting via Global and Regional Coupling Weiqi Chen DAMO Academy, Alibaba Group Hangzhou, China Wenwei Wang DAMO Academy, Alibaba Group Hangzhou, China Qilong Yuan Northwest Polytechnical University Xi’an, China Lefei Shen Zhejiang University Hangzhou, China Bingqing Peng DAMO Academy, Alibaba Group Hangzhou, China Jiawei Chen Zhejiang University Hangzhou, China Bo Wu Institute of Atmospheric Physics, Chinese Academy of Sciences Beijing, China Liang Sun DAMO Academy, Alibaba Group Hangzhou, China Abstract Data-driven weather models have advanced global medium-range forecasting, yet high-resolution regional prediction remains chal- lenging due to unresolved multiscale interactions between large- scale dynamics and small-scale processes such as terrain-induced circulations and coastal effects. This paper presents a global-regional coupling framework for kilometer-scale regional weather fore- casting that synergistically couples a pretrained Transformer-based global model with a high-resolution regional network via a novel bidirectional coupling module, ScaleMixer. ScaleMixer dynami- cally identifies meteorologically critical regions through adaptive key-position sampling and enables cross-scale feature interaction through dedicated attention mechanisms. The framework produces forecasts at 0.05 ◦ (∼5km) and 1-hour resolution over China, signif- icantly outperforming operational NWP and AI baselines on both gridded reanalysis data and real-time weather station observations. It exhibits exceptional skill in capturing fine-grained phenomena such as orographic wind patterns and Foehn warming, demon- strating effective global-scale coherence with high-resolution fi- delity. The code is available at https://anonymous.4open.science/r/ ScaleMixer-6B66. Keywords Regional weather forecasting, Downscaling, Deep Neural networks ACM Reference Format: Weiqi Chen, Wenwei Wang, Qilong Yuan, Lefei Shen, Bingqing Peng, Ji- awei Chen, Bo Wu, and Liang Sun. 2018. Skillful Kilometer-Scale Regional Weather Forecasting via Global and Regional Coupling. In . ACM, New York, NY, USA, 24 pages. https://doi.org/X.X Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. Conference’17, Washington, DC, USA © 2018 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-X-X/2018/06 https://doi.org/X.X 1 Introduction Accurate weather forecasting is essential for disaster mitigation, agriculture, transportation, and energy management [6]. Tradi- tional numerical weather prediction (NWP) systems solve the gov- erning equations of atmospheric dynamics involving mass con- tinuity, momentum conservation, and thermodynamics, and pa- rameterize subgrid-scale processes such as turbulence and cloud microphysics [3,13]. Although NWP models provide physically consistent forecasts and remain operational standards, their compu- tational demands and sensitivity to parameterization schemes limit the skill in resolving kilometer-scale weather phenomena governed by multiscale interactions. Recent data-driven AI models, particularly Transformer-based architectures trained on global reanalysis data such as ERA5, have achieved remarkable success in medium-range forecasting at syn- optic scales at resolution of 0.25 ◦ and coarser. However, high- resolution operational regional forecasting (e.g., 0.05 ◦ , or∼5km) remains a significant challenge. Kilometer-scale weather is gov- erned by complex multiscale interactions: large-scale circulations modulate local processes such as topographic flows, coastal breezes, and convective systems, while fine-scale features also feedback to broader dynamics. A prime example is the Hengduan Mountains, where large-scale dynamics including the Indian Monsoon, East Asian Monsoon, and Tibetan Plateau climate, interact with extreme terrain gradients. These terrain gradients, which exceed 3,000 m within 100km, drive localized wind accelerations, sharp tempera- ture contrasts, and convective processes that are poorly captured by coarse global models or isolated regional models [30]. Such intricate multiscale interactions challenge conventional models, necessitat- ing forecasting models that reconcile global-scale coherence with high-resolution fidelity. Recent studies have begun to explore data-driven regional weather forecasting and downscaling, typically treating global forecasts as static inputs [20,21,23,26,31]. However, these decoupled meth- ods neglect dynamic cross-scale interactions and suffer from tem- poral misalignment between low-frequency global forecasts (e.g., 6-hourly) and high-resolution regional observations (e.g., hourly). In summary, to make accurate high-resolution regional weather arXiv:2603.28173v1 [cs.LG] 30 Mar 2026 Conference’17, July 2017, Washington, DC, USAWeiqi Chen, Wenwei Wang, Qilong Yuan, Lefei Shen, Bingqing Peng, Jiawei Chen, Bo Wu, and Liang Sun forecasting requires addressing two key challenges: (1) a mecha- nism to dynamically identify regions where cross-scale interactions are active, and (2) a bidirectional coupling framework that ensures spatial-temporal consistency across scales. To address the aforementioned challenges, we propose a novel global–regional coupling framework for high-resolution re- gional weather prediction. Our approach seamlessly integrates a pretrained global Transformer model, which provides synoptic- scale (large scale) context, with a regional refinement model operat- ing at 0.05 ◦ resolution. Central to this architecture is ScaleMixer, a module that adaptively identifies key spatial regions exhibiting strong multiscale interactions and enable bidirectional feature en- coding between global and regional tokens. This allows the model to prioritize meteorologically critical areas such as typhoon bound- aries and mountain ridges, and maintain global coherence while resolving fine-grained regional dynamics. The main contributions of this work are summarized as follows: • A global–regional coupling framework for 0.05 ◦ and 1- hour forecasting by integrating a pretrained global model for synoptic-scale context with a high-resolution regional model •The ScaleMixer module for dynamic identification of cross- scale interaction regions and bidirectional feature fusion. •Experiments on both hindcast and operational settings show our model’s superiority against operational NWP and lead- ing AI baselines. Case studies over complex terrain in China further demonstrates the model’s notable skill in capturing orographic wind effects and Foehn warming. •Similar to other AI-based models, the model is highly effi- cient during inference. It takes less than 3 minutes for 48- hour forecasting on a single GPU, while numerical models like IFS-HRES typically requires about 1 hour on massive CPU clusters for a similar forecasting window. 2 Related Work Numerical Weather Forecasting. As the predominant paradigm, NWP systems typically formulate the atmospheric physical laws through PDEs and then solve them using numerical simulations. Representative examples include earth system models (ESMs) [13] and the operational Integrated Forecast System (IFS) of European Centre for Medium-Range Weather Forecasts (ECMWF) [3]. By integrating physics laws, NWP approaches have enjoyed remark- able success with great accuracy, stability, and interpretability. IFS- HRES [10] is a world-leading high-resolution deterministic NWP system (0.1 ◦ ) and serves as a benchmark for operational forecasting and research. However, NWP models are sensitive to initial con- ditions, prone to errors in parameterization, and computationally expensive [15]. These limitations hinder their ability to accurately resolve kilometer-scale weather driven by complex multiscale in- teractions. Deep Learning for Global Weather Forecasting. Recent progress in deep learning models for global weather forecasting has been trans- formative. They predominantly employ two architectural paradigms: Transformer-based models [2,4,5,22] and Graph Neural Net- work (GNN)-based architectures [14,16,25]. These models demon- strate computational efficiency and competitive skill in predict- ing synoptic-scale weather patterns. However, they fail to capture finer-grained mesoscale weather dynamics due to limited resolution (0.25 ◦ or coarser). Deep Learning for Regional Weather Forecasting and Downscal- ing. Recently, regional weather models have been developed for fine-scale forecasting and downscaling over regions of interest. For instance, CorrDiff [20] combines U-Net and diffusion models to correct and downscale the global forecasts to improve local pre- dictions. Machine learning limited area models [1] further take into account boundary conditions, because the evolution of atmo- spheric states relies on both the internal dynamics and the exter- nal forcing. A sequence of GNN layers is applied to capture both large-scale circulations and local microscale processes. Other lim- ited area modeling methods also employ GNN architectures with stretched-grid [11] and nested-grid [21] to make low-resolution global and high-resolution regional weather forecasts simultane- ously. They model cross-scale interactions through grid deforma- tion and nesting, with a static graph structure; however, these rigid, geometry-driven interactions limit the model’s ability to efficiently capture highly dynamic and non-local coupling processes. In con- trast, our model employs a bidirectional coupling module to learn the content-dependent cross-scale interactions between global and regional tokens adaptively. 3 Methodology Accurate regional weather forecasting requires seamless integration of large-scale atmospheric dynamics with localized, high-resolution features. As nearly all AI-based global models are trained on the EAR5 dataset, we assume a pretrained Vision Transformer (ViT)- based global weather forecasting model, denotedM global , which operates on low-resolution (0.25 ◦ ) global reanalysis data U 푡 0 ∈ R 퐻×푊×퐶 . At time푡 0 , the model generates 6-hour-ahead global pre- dictions capturing synoptic-scale dynamics: ˆ U 푡 0 +6H =M global (U 푡 0 ),(1) where퐻 ×푊 ×퐶represents the stacked weather state with multi- ple levels of upper air and surface variables, in which latitude and longitude are divided into퐻and푊grids for each variable. Con- currently, high-resolution regional analysis data풖 푡 0 ∈R ℎ×푤×푉 reg provides critical surface variables (wind components푈,푉, temper- ature푇, specific humidity푄, pressure푃, radiation fluxes푆푅퐷, and total cloud cover푇퐶) within a region of interest at 1hour temporal resolution and 0.05 ◦ spatial resolution. Problem Formulation. As the fundamental challenge lies in ef- fectively coupling multiscale information: coarse-grained global features fromM global and fine-resolution regional features, we for- malize the task as developing a hybrid global-regional weather fore- casting frameworkM global−regional that extendsM global through the integration of large-scale atmospheric dynamics and small-scale weather effects. ˆ U 푡 0 +6H ; ˆ 풖 푡 0 +푖H 6 푖=1 =M global−regional U 푡 0 ; 풖 푡 0 +푖H 0 푖=−5 ,(2) Skillful Kilometer-Scale Regional Weather Forecasting via Global and Regional CouplingConference’17, July 2017, Washington, DC, USA where ˆ 풖 푡 0 +푖H 0 푖=−5 denotes the temporally aligned regional analy- sis data with 1-hour intervals. This formulation establishes a prin- cipled framework for generating high-fidelity regional forecasts by systematically bridging global-scale dynamics with localized meteorological processes with deep learning architectures. Model Overview. We propose a multiscale weather forecasting framework that dynamically integrates global and regional-scale atmospheric dynamics to resolve high-resolution mesoscale fea- tures in the target region. As shown in Figure 1, the framework comprises two Transformer-based sub-models with shared archi- tectural principles: (1) a global model for synoptic-scale dynamics, and (2) a regional model for mesoscale processes. The ScaleMixer module enables bidirectional coupling between the global and re- gional models via adaptive key position identification and encod- ing to preserve cross-scale meteorological consistency. The global model, pretrained on ERA5 reanalysis data [12], remains fixed dur- ing regional optimization, while the regional model and ScaleMixer module are trained from scratch. 3.1 Pretrained Global Weather Model Our global modelM global prioritizes architectural simplicity, flexi- bility, and scalability, and implements a vision Transformer (ViT) architecture [7]. Without loss of generality, our framework can work with any ViT-based global weather forecasting model. In the following discussion, we will limit the discussion on our in- house developed ViT-based global model, comprising three core components: Patch Embedding and Tokenization: A 2-D convolutional layer partitions the multivariate input atmospheric state U 푡 0 ∈ R 퐻×푊×퐶 into non-overlapping spatial patches of size푃 × 푃. This generates token representations S∈R 푁×푑 , where푁= (퐻/푃)× (푊/푃) and 푑 is the embedding dimension. Transformer Encoder: A stack of푀Transformer encoder layers processes the sequence S through multi-head self-attention and feed-forward networks [7,29], enabling global information interaction across spatial scales. Prediction Head: A deconvolution block upscales the processed sequence back to the original spatial resolution퐻 ×푊, produc- ing a 6-hour ahead deterministic global forecast U 푡 0 +6H of the full atmospheric state. The global modelM global is pretrained on ERA5 reanalysis [12] using weighed mean absolute error (MAE) as the loss function (detailed in Section 3.4). The dataset includes five pressure level variables (13 vertical levels each): geopotential (푧), specific humid- ity (푞), wind components (푢,푣), and temperature (t), and multiple surface variables, e.g., 2-meter temperature (t2m), 10-meter wind (u10, v10), and mean sea level pressure (msl), surface pressure (sp), etc. (detailed in Section B). 3.2 Modifications in Regional Weather Model The regional modelM regional inherits the Transformer architecture from the global model but introduces necessary modifications: (1) modified patch embedding layer to incorporate fine-grained to- pography and temporal encodings, (2) enhanced prediction head with adaptive layer normalization (AdaLN) [24] to amplify the high-frequency signal for hourly temporal alignment, and (3) fewer Transformer encoder layers (푘 ≪ 푀) to reduce computational overhead while preserving regional meteorological fidelity. Patch Embedding: In addition to the input 풖 푡 0 +푖H 0 푖=−5 ∈ R ℎ×푤×푉 reg ×6 , the block also needs to process the static topography, land-sea mask, and dynamic hourly temporal information. Regional analyses are tokenized across 6 time steps using a shared patch embedding layer, with topography, land-sea masks, and temporal embeddings (hour-of-day, day-of-year) added via MLP. To ensure ge- ographic consistency with global patches, we set patch size푝=5×푃, generating regional tokens s∈R 푛×푑 , where 푛=(ℎ/푝)×(푤/푝). Transformer Encoders: The regional model employs푘encoder layers (푘 ≪ 푀, where푀= 푘 × 퐿) to achieve computational effi- ciency in regional optimization. Each cross-scale coupling block comprises퐿global encoder layers, 1 regional encoder layer, and 1 ScaleMixer module. Prediction Head: To generate 6-hour forecasts at hourly in- tervals, 6 dedicated prediction heads produce lead time-specific outputs (Δ푡=1Hto6H). Temporal alignment is enforced via AdaLN[24], where scale and shift parameters훾,훽are derived from Fourier embeddings ofΔ푡 : FourierEmbed(Δ푡)= [ cos(2휋푎 푖 Δ푡 +푏 푖 ), sin(2휋푎 푖 Δ푡 +푏 푖 ) ] (3) for 0≤ 푖< 푑/2, 훾,훽= MLP ( FourierEmbed(Δ푡) ) ,(4) where푎 푖 and푏 푖 are learnable Fourier embedding parameters. This formulation ensures high-frequency signal amplification for re- gional forecasting. Moreover, regional prediction heads take the concatenation of regional tokens and spatially-aligned global to- kens as input to make full use of multi-scale information. 3.3 ScaleMixer: Bidirectional Global and Regional Scale Coupling Accurate high-resolution regional prediction requires resolving multiscale atmospheric processes–from synoptic-scale forcings to mesoscale circulations–while maintaining global dynamical con- sistency. To this end, we introduce ScaleMixer, a differentiable coupling mechanism that explicitly models interactions between the global foundation model and the regional refinement model. As illustrated in Figure 1 (right), ScaleMixer enables bidirectional feature fusion by adaptively identifying meteorologically critical re- gions and performing token-level encoding, effectively prioritizing areas with strong cross-scale interactions. Adaptive key position identification. To capture spatial regions exhibiting strong multiscale interactions, we implement a dynamics- aware sampling module that identifies critical spatial positions from global token embeddings S. Spatial dynamics are extracted via a con- volutional network, followed by softmax-normalized importance scores Pr∈R 푁 (푁 is the number of global tokens): Pr= Softmax ( Conv(S) ) ,(5) whereConv(·)consists of a convolutional layer followed by a linear projection. We then select top-푚 salient positions: c= arg top-푚(Pr),h= Pr[c] ⊙ S[c],(6) Conference’17, July 2017, Washington, DC, USAWeiqi Chen, Wenwei Wang, Qilong Yuan, Lefei Shen, Bingqing Peng, Jiawei Chen, Bo Wu, and Liang Sun Prediction Head ERA5 pretrained ScaleMixer Transformer Encoder Transformer Encoder Transformer Encoder Transformer Encoder Transformer Encoder <latexit sha1_base64="52LXA0equLXkC5hUWoU18ImqmxM=">AAAB73icbVA9SwNBEJ3zM8avqKXNYhCswl2QaBmwsbCIYD4gOcLeZi9Zsrd37s4J4cifsLFQxNa/Y+e/cZNcoYkPBh7vzTAzL0ikMOi6387a+sbm1nZhp7i7t39wWDo6bpk41Yw3WSxj3Qmo4VIo3kSBkncSzWkUSN4Oxjczv/3EtRGxesBJwv2IDpUIBaNopU4PRcQNueuXym7FnYOsEi8nZcjR6Je+eoOYpRFXyCQ1puu5CfoZ1SiY5NNiLzU8oWxMh7xrqaJ2jZ/N752Sc6sMSBhrWwrJXP09kdHImEkU2M6I4sgsezPxP6+bYnjtZ0IlKXLFFovCVBKMyex5MhCaM5QTSyjTwt5K2IhqytBGVLQheMsvr5JWteLVKrX7y3K9msdRgFM4gwvw4ArqcAsNaAIDCc/wCm/Oo/PivDsfi9Y1J585gT9wPn8Ap6aPrw==</latexit> →L <latexit sha1_base64="52LXA0equLXkC5hUWoU18ImqmxM=">AAAB73icbVA9SwNBEJ3zM8avqKXNYhCswl2QaBmwsbCIYD4gOcLeZi9Zsrd37s4J4cifsLFQxNa/Y+e/cZNcoYkPBh7vzTAzL0ikMOi6387a+sbm1nZhp7i7t39wWDo6bpk41Yw3WSxj3Qmo4VIo3kSBkncSzWkUSN4Oxjczv/3EtRGxesBJwv2IDpUIBaNopU4PRcQNueuXym7FnYOsEi8nZcjR6Je+eoOYpRFXyCQ1puu5CfoZ1SiY5NNiLzU8oWxMh7xrqaJ2jZ/N752Sc6sMSBhrWwrJXP09kdHImEkU2M6I4sgsezPxP6+bYnjtZ0IlKXLFFovCVBKMyex5MhCaM5QTSyjTwt5K2IhqytBGVLQheMsvr5JWteLVKrX7y3K9msdRgFM4gwvw4ArqcAsNaAIDCc/wCm/Oo/PivDsfi9Y1J585gT9wPn8Ap6aPrw==</latexit> →L Patch EmbeddingPatch Embedding Prediction Head Coupling Block <latexit sha1_base64="5YliE42eXDzgekC4/NXV49CCS1w=">AAAB73icbVBNS8NAEJ3Ur1q/qh69LBbBU0mKVI8FLx4r2A9oQ9lsN+3SzSbuToQS+ie8eFDEq3/Hm//GbZuDtj4YeLw3w8y8IJHCoOt+O4WNza3tneJuaW//4PCofHzSNnGqGW+xWMa6G1DDpVC8hQIl7yaa0yiQvBNMbud+54lrI2L1gNOE+xEdKREKRtFK3T6KiBsyGZQrbtVdgKwTLycVyNEclL/6w5ilEVfIJDWm57kJ+hnVKJjks1I/NTyhbEJHvGeponaNny3unZELqwxJGGtbCslC/T2R0ciYaRTYzoji2Kx6c/E/r5dieONnQiUpcsWWi8JUEozJ/HkyFJozlFNLKNPC3krYmGrK0EZUsiF4qy+vk3at6tWr9furSqOWx1GEMziHS/DgGhpwB01oAQMJz/AKb86j8+K8Ox/L1oKTz5zCHzifP9aij84=</latexit> →k <latexit sha1_base64="5YliE42eXDzgekC4/NXV49CCS1w=">AAAB73icbVBNS8NAEJ3Ur1q/qh69LBbBU0mKVI8FLx4r2A9oQ9lsN+3SzSbuToQS+ie8eFDEq3/Hm//GbZuDtj4YeLw3w8y8IJHCoOt+O4WNza3tneJuaW//4PCofHzSNnGqGW+xWMa6G1DDpVC8hQIl7yaa0yiQvBNMbud+54lrI2L1gNOE+xEdKREKRtFK3T6KiBsyGZQrbtVdgKwTLycVyNEclL/6w5ilEVfIJDWm57kJ+hnVKJjks1I/NTyhbEJHvGeponaNny3unZELqwxJGGtbCslC/T2R0ciYaRTYzoji2Kx6c/E/r5dieONnQiUpcsWWi8JUEozJ/HkyFJozlFNLKNPC3krYmGrK0EZUsiF4qy+vk3at6tWr9furSqOWx1GEMziHS/DgGhpwB01oAQMJz/AKb86j8+K8Ox/L1oKTz5zCHzifP9aij84=</latexit> →k ... <latexit sha1_base64="4cRwnCzN6I+oVG68vCMDFrPHCb4=">AAAB+XicbVBNS8NAFNzUr1q/oh69LBbBU0mKVI8FLx4rmLbQhrDZbtqlm03YfSmU0H/ixYMiXv0n3vw3btoctHVgYZh5jzc7YSq4Bsf5tipb2zu7e9X92sHh0fGJfXrW1UmmKPNoIhLVD4lmgkvmAQfB+qliJA4F64XT+8LvzZjSPJFPME+ZH5Ox5BGnBIwU2PYwJjAJo9xbBDkEziKw607DWQJvErckdVSiE9hfw1FCs5hJoIJoPXCdFPycKOBUsEVtmGmWEjolYzYwVJKYaT9fJl/gK6OMcJQo8yTgpfp7Iyex1vM4NJNFTr3uFeJ/3iCD6M7PuUwzYJKuDkWZwJDgogY84opREHNDCFXcZMV0QhShYMqqmRLc9S9vkm6z4bYarcebertZ1lFFF+gSXSMX3aI2ekAd5CGKZugZvaI3K7derHfrYzVascqdc/QH1ucPy+iTuw==</latexit> U t 0 <latexit sha1_base64="0tEETaCWRvLqf1FNenwa/DjUTdw=">AAACKXicbVDLSsNAFJ34tr6iLt0EiyCIJSla3QgFNy4rWBU6MUwmk3Zw8mDmRihDfseNv+JGQVG3/oiTtgtfB4Y5nHMv994T5oIrcN13a2p6ZnZufmGxtrS8srpmr29cqqyQlHVpJjJ5HRLFBE9ZFzgIdp1LRpJQsKvw9rTyr+6YVDxLL2CYMz8h/ZTHnBIwUmC3sWAxYI0HBDQOMxGpYWI+XZRloCFw9zhOCAxkos/KEkveHwA2Dj/ZPyxvtFsGdt1tuCM4f4k3IXU0QSewn3GU0SJhKVBBlOp5bg6+JhI4Fays4UKxnNBb0mc9Q1OSMOXr0aWls2OUyIkzaV4Kzkj93qFJoqoDTGW1tfrtVeJ/Xq+A+NjXPM0LYCkdD4oL4UDmVLE5EZeMghgaQqjkZleHDogkFEy4NROC9/vkv+Sy2fBajdb5Qb3dnMSxgLbQNtpFHjpCbXSGOqiLKLpHj+gFvVoP1pP1Zn2MS6esSc8m+gHr8wsLYqjg</latexit> ˆ u t 0 +iH 0 i=→5 : Regional patch embeddings <latexit sha1_base64="8DLjom7jD/mQvVYCz35dTOzcsbE=">AAAB9XicbVDLSgMxFL3js9ZX1aWbYBFclRmR6rLoxmUF+4B2LJlMpg3NJEOSUcrQ/3DjQhG3/os7/8ZMOwttPRByOOdecnKChDNtXPfbWVldW9/YLG2Vt3d29/YrB4dtLVNFaItILlU3wJpyJmjLMMNpN1EUxwGnnWB8k/udR6o0k+LeTBLqx3goWMQINlZ66AeSh3oS2yvT00Gl6tbcGdAy8QpShQLNQeWrH0qSxlQYwrHWPc9NjJ9hZRjhdFrup5ommIzxkPYsFTim2s9mqafo1CohiqSyRxg0U39vZDjWeTQ7GWMz0oteLv7n9VITXfkZE0lqqCDzh6KUIyNRXgEKmaLE8IklmChmsyIywgoTY4sq2xK8xS8vk/Z5zavX6ncX1cZ1UUcJjuEEzsCDS2jALTShBQQUPMMrvDlPzovz7nzMR1ecYucI/sD5/AFPRZMP</latexit> s : Global patch embeddings <latexit sha1_base64="gDyuhtg4kYcYfkwBE8o5K5PDz6A=">AAAB8XicbVDLSgMxFL1TX7W+qi7dBIvgqsyIVJdFNy4r2ge2pWTSO21oJjMkGaEM/Qs3LhRx69+482/MtLPQ6oHA4Zx7ybnHjwXXxnW/nMLK6tr6RnGztLW9s7tX3j9o6ShRDJssEpHq+FSj4BKbhhuBnVghDX2BbX9ynfntR1SaR/LeTGPsh3QkecAZNVZ66IXUjP0gvZsNyhW36s5B/hIvJxXI0RiUP3vDiCUhSsME1brrubHpp1QZzgTOSr1EY0zZhI6wa6mkIep+Ok88IydWGZIgUvZJQ+bqz42UhlpPQ99OZgn1speJ/3ndxASX/ZTLODEo2eKjIBHERCQ7nwy5QmbE1BLKFLdZCRtTRZmxJZVsCd7yyX9J66zq1aq12/NK/SqvowhHcAyn4MEF1OEGGtAEBhKe4AVeHe08O2/O+2K04OQ7h/ALzsc3yJKRAg==</latexit> S Key position Identification Position embeddings Global-to-position Attention Position-to-regional Attention QKV QKV Concat & Projection Orography Time embedding Adaptive Layer Normalization Forecast lead time <latexit sha1_base64="VoVc7PUzLpk3wX/8aq2Zs5mRutI=">AAAB73icbVBNS8NAEJ34WetX1aOXxSJ4KolI9VjUg8cK9gPaUDbbTbt0s4m7E6GE/gkvHhTx6t/x5r9x2+agrQ8GHu/NMDMvSKQw6Lrfzsrq2vrGZmGruL2zu7dfOjhsmjjVjDdYLGPdDqjhUijeQIGStxPNaRRI3gpGN1O/9cS1EbF6wHHC/YgOlAgFo2ildveWS6QEe6WyW3FnIMvEy0kZctR7pa9uP2ZpxBUySY3peG6CfkY1Cib5pNhNDU8oG9EB71iqaMSNn83unZBTq/RJGGtbCslM/T2R0ciYcRTYzoji0Cx6U/E/r5NieOVnQiUpcsXmi8JUEozJ9HnSF5ozlGNLKNPC3krYkGrK0EZUtCF4iy8vk+Z5xatWqvcX5dp1HkcBjuEEzsCDS6jBHdShAQwkPMMrvDmPzovz7nzMW1ecfOYI/sD5/AGTB4+v</latexit> !t Global Forecast Regional Forecast Pretrained and frozen Train from scratch Target region in global field Figure 1: Left: The Architecture of Global-Regional Weather Forecasting Model: Synoptic-scale context (M global ) drives mesoscale regional refinement (M regional ) via ScaleMixer, ensuring cross-scale coupling and consistency. Right: ScaleMixer Module: Bidirectional Cross-Scale Coupling via Key Position Identification and encoding. Key components include (1) key position identification, (2) coupling regional dynamics with global context via global-to-position and position-to-regional attention, and (3) global token adaptation incorporating regional features. with⊙denoting element-wise product, c=c 푖 푚 푖=1 ∈R 푚×2 (p 푖 ∈ [ 0 :퐻/푃 −1]×[0 :푊/푃 −1]) representing the coordinates of푚 selected tokens, and h∈R 푚×푑 their corresponding embeddings. Regional features alignment with global context. To effectively bridge the scale gap between global context and regional features, we design a two-stage cross-attention mechanism operating on identified key positions. Directly correlating all global and regional tokens is computationally expensive and may weaken localized meteorological features. Instead, we first condense global informa- tion into a sparse set of dynamically identified key positions, then propagate these enriched features to regional tokens. Global-to-Position Attention first aggregates global context into the key positions. Using the concatenated token embeddings and coordinates of key positions h||c∈R 푚×(푑+2) as queries, and the global tokens S as keys and values, we compute: Glo-to-Pos(h||c, S, S)= Softmax (W 푄 · h||c)(W 퐾 S) ⊤ √ 푑 W 푉 S, (7) h global ||c ′ = h||c+ Glo-to-Pos(h||c, S, S),(8) where W 푄 , W 퐾 and W 푉 are linear projections. To better model the dynamics of key positions, the key representations are further refined by incorporating regional features via bilinear interpolation at the updated coordinates c ′ h ′ = MLP Proj Bilinear(s, c ′ )||h global .(9) Position-to-regional attention subsequently integrates the globally informed key features into regional tokens s: Pos-to-Reg(s, h ′ ||c ′ , h ′ ||c ′ )= Softmax W ′ 푄 s(W ′ 퐾 · h ′ ||c ′ ) ⊤ √ 푑 ! W ′ 푉 · h ′ ||c ′ , (10) s ′ = s+ Pos-to-Reg(s, c ′ ||p ′ , c ′ ||p ′ ). (11) with distinct learnable projections W ′ 푄 , W ′ 퐾 , and W ′ 푉 . This two-step attention mechanism ensures that synoptic-scale dynamics are ef- fectively integrated into high-resolution regional features, enabling globally consistent and locally accurate weather prediction. Global token adaptation with regional feedback. To enable large- scale dynamics to adapt to regional details, global tokens spatially aligned with regional tokens (S aligned ) are updated via token-wise concatenation and an adapter MLP: S ′ aligned = Concat S aligned , s ′ ∈R 푛×2푑 ,(12) S ′ aligned = S aligned + MLP Adapter S ′ aligned .(13) These adapted tokens S ′ aligned replace their counterparts in the global token sequence, allowing regional fine-scale information to recur- sively influence the global context in subsequent encoder layers. 3.4 Model Optimization The optimization schedule follows a three-stage training protocol: (1) global model pretraining, (2) regional model one-step training (6-hours ahead), and (3) regional model autoregressive roll-out fine-tuning (12∼ 48-hours ahead). Skillful Kilometer-Scale Regional Weather Forecasting via Global and Regional CouplingConference’17, July 2017, Washington, DC, USA Training objective. For global model pretraining, we employ the weighted mean absolute error (MAE) across multivariate atmo- spheric states. Decomposing the weather state U 푡 into surface-level variables and upper-air atmospheric variables, ˆ U 푡 = ( ˆ S 푡 , ˆ A 푡 )and U 푡 =(S 푡 , A 푡 ), the loss can be written as: L( ˆ U, U 푡 )= 1 푉 푆 +푉 퐴 " 푉 푆 ∑︁ 푘=1 푤 푆 푘 퐻 ×푊 퐻 ∑︁ 푖=1 푊 ∑︁ 푗=1 | ˆ S 푡 푖,푗,푘 − S 푡 푖,푗,푘 | ! + 푉 퐴 ∑︁ 푘=1 1 퐻 ×푊 × 푃 푃 ∑︁ 푝=1 푤 퐴 푐,푘 퐻 ∑︁ 푖=1 푊 ∑︁ 푗=1 | ˆ A 푡 푖,푗,푝,푘 − A 푡 푖,푗,푝,푘 | !# , (14) where푉 퐴 and푉 퐾 are numbers of upper-air and surface varibles, 푃is the number of pressure levels,푤 푆 푘 is the weight associated with surface-level variable푘, and푤 퐴 푘,푐 is the weight associated with atmospheric variable 푘 at pressure level 푝. During both one-step training and roll-out fine-tuning of re- gional model, we directly using MAE as the objective: L( ˆ 풖,풖 푡 )= 1 푉 reg 푉 reg ∑︁ 푘=1 1 ℎ×푤 ℎ ∑︁ 푖=1 푤 ∑︁ 푗=1 | ˆ 풖 푡 푖,푗,푘 −풖 푡 푖,푗,푘 | ! ,(15) where푉 reg is the number of variables in regional analyses. Roll-out Fine-tuning. The model makes forecasts for the next 6 hours in one step, and longer forecasts are obtained by rolling auto-regressively, which may suffer from error accumulations. To enhance the multi-step forecasting accuracy, we adopt a rolling-out fine-tuning strategy for 48 hours. It is performed by predicting ˆ U 푡 0 +6(푛+1)+6H and ˆ 풖 푡 0 +6(푛+1)+푖H 6 푖=1 with the predictions from the previous step ˆ U 푡 0 +6푛+6H and ˆ 풖 푡 0 +6푛+푖H 6 푖=1 as input, for푛=0,1, ... recursively, and optimize the loss function of 3.4 and 3.4 over the 48 time spans. Implementation and Training Details. The global Transformer encoder comprises 24 layers (푀=24), while the regional encoder and ScaleMixer modules each contain 4 layers (푘=4). The model employs a hidden dimension of 1536 and identifies푚=64 key positions for cross-scale interaction in each ScaleMixer module. The framework contains 1.07 billion parameters, with the global model M global accounting for 736 million. Full implementation details are summarized in Section A. The global model was pretrained for 150,000 steps on 32×NVIDIA A800 GPUs using the AdamW optimizer [19] with a per-GPU batch size of 1. A cosine learning rate schedule was applied with linear warmup over 1,000 steps, decaying from 7×10 −4 to 1×10 −7 . Regional model training followed identical hyperparameters over 80, 000 it- erations on 8×A800 GPUs, withM global parameters frozen. During regional roll-out fine-tuning, the model was trained for 100,000 steps at a fixed learning rate of 1×10 −6 . The one-time training takes approximately 20 days on 8 NVIDIA A800 GPUs, and the inference stage is quite efficient, taking less than 3 minutes for 48-hour forecasting on a single GPU. 4 Experiments To resolve high-impact meteorological phenomena such as con- vective storms and boundary layer dynamics, weather prediction (a) Latitude-weighted RMSE for 7 surface variables (2024/10–2024/12 hindcast period) (b) Latitude-weighted RMSE for 7 surface variables (2025/01–2025/04 operational period) Figure 2: ScaleMixer demonstrates superior deterministic forecasting skill compared to IFS-HRES at 0.05° resolu- tion. Seven surface variables (T2M, U10, V10, Q, P, TCC, and SSRD) are evaluated using latitude-weighted RMSE (lower values indicate superior performance). (a) Hind- cast results show ScaleMixer outperforms IFS-HRES across all variables during 2024/10–2024/12. (b) Operational fore- casts confirm ScaleMixer maintains superiority performance (2025/01–2025/04). systems require high-resolution spatial-temporal modeling capabil- ities. We evaluate ScaleMixer through two complementary experi- mental paradigms: (1) hindcast for verification using reanalysis data, and (2) operational forecast to assess predictive skill under dynamically evolving initial conditions consistent with production environment management system. 4.1 Datasets Global Reanalysis (ERA5). The European Centre for Medium- Range Weather Forecasts (ECMWF) ERA5 reanalysis provides 0.25° horizontal resolution (1440× 720 latitude-longitude grid) atmo- spheric states with 37 hybrid pressure levels. Spanning 1979–2015, this dataset serves as the primary training source for the global model (M global ), with 2016 reserved for validation. ERA5’s spa- tiotemporal continuity and multivariate fidelity make it a standard for data-driven weather modeling [12]. Similar to most AI-based weather forecasting foundation models, we select 5 atmospheric variables (u-component of wind speed, v-component of wind speed, temperature, specific humidity, geopotential) at 13 pressure levels (50 hPa, 100 hPa, 150 hPa, 200 hPa, 250 hPa, 300 hPa, 400 hPa, 500 hPa, 600 hPa, 700 hPa, 850 hPa, 925 hPa, 1000 hPa), 6 surface vari- ables and 6 static variables from the raw ERA5 dataset, and use z-score normalization for each variable. All variables are inputs of M 푔푙표푏푎푙 , and all variables except the static ones are used as model outputs. Global Operational Analysis. Operational analysis utilize the ini- tial conditions from ECMWF’s High-Resolution Deterministic Pre- diction (HRES) system, which assimilates observations through Conference’17, July 2017, Washington, DC, USAWeiqi Chen, Wenwei Wang, Qilong Yuan, Lefei Shen, Bingqing Peng, Jiawei Chen, Bo Wu, and Liang Sun Table 1: Average RMSE and ACC acrossΔ푡=1∼48 hours lead time regional weather hindcast at 0.05 ◦ resolution. The best results are bolded. Variable Latitude-weighted RMSELatitude-weighted ACC IFS-HRES Baguan M global M regional OneForecast LAM ScaleMixerΔRMSEIFS-HRES M global M regional OneForecast LAM ScaleMixerΔACC T2M1.8151.9281.9912.4521.5711.742 1.382 ↓23.86%0.8620.8810.8450.8970.901 0.921 ↑6.84% U101.9341.9561.9672.6312.2511.889 1.644 ↓14.99% 0.7210.7530.7160.7440.742 0.793 ↑9.99% V101.9281.9371.9702.8862.3161.905 1.617 ↓16.13%0.7230.7510.7130.7320.737 0.785 ↑8.57% Q0.8070.6110.8111.2430.7140.676 0.559 ↓30.73%0.7740.7680.7710.7510.781 0.812 ↑4.91% P13.2722.9710.274.2412.5452.141.874 ↓85.88%0.8870.9020.8940.8990.912 0.920 ↑3.72% TCC38.8331.83(N/A)43.7137.7434.13 28.76 ↓25.93%0.563(N/A)0.6170.6210.679 0.721 ↑28.06% SSRD51.2541.18(N/A)68.4252.5642.41 33.26 ↓34.04%0.824(N/A)0.8340.8380.850 0.887 ↑7.64% (1)ΔRMSE andΔACC donote RMSE and ACC improvement of ScaleMixer compared to IFS-HRES. (2)M global denotes standalone global model, andM regional denotes uncoupled regional model. (3) Results of IFS-HRES (0.1°) and Baguan, and M global (0.25°) are corrected and downscaled to target grid (0.05°) using a pretrained bias-correction and downscaling model (based on a ViT backbone trained on ERA5 and CLDAS data) for comparison. (4) T2M: 2m temperature; U10/V10: 10m wind components; Q: Specific humidity; P: Surface Pressure; TCC: Total cloud cover; SSRD: Radiation flux (surface solar radiation downward). 2024/10/30 22 UTC 2024/10/31 8 UTC CLDASScaleMixerIFS-HRESCLDASScaleMixerIFS-HRES 2m temperature (°C)10m wind speed (m/s) Wind Ridge Leeward slope Widward slope Figure 3: Left: Temporal evolution of 10m wind speed predictions initialized at 2024/10/30 12 UTC over the Hengduan Moun- tains (25.0–35.0°N, 95.0–105.0°E), China. Black arrows represent wind flow fields. ScaleMixer resolves enhanced resolution of orographic wind heterogeneity (peaking >10 m/s at crests and <2 m/s in valleys). Right: Corresponding temperature fields. Foehn effects are illustrated in the picture, characterized by 4–8°C leeward warming relative to windward slopes through adiabatic compression processes. ScaleMixer captures fine-grained temperature gradients, contrasting with IFS-HRES exhibiting spatial smoothing forecasts. 4D-variational data assimilation [27]. The 0.1° analysis fields (inter- polated to ERA5 resolution, 0.25°) provide dynamically real-time initial conditions for ScaleMixer’s operational deployment during 2025/01–2025/04. Regional Analysis (CLDAS). The China Meteorological Adminis- tration’s Land Data Assimilation System (CLDAS) is a near-realtime regional reanalysis product. It contains 7 critical surface varialbes with 1 hour intervals, covering East Asia (0–65°N, 60–160°E) area at a spatial resolution 0f 0.01°. To reduce computational complex- icity and maintain detailed information, we interpolate the data to 0.05° latitude-longitude gridding. The full dataset covers from 2020 onward. We use data from 2022/01-2024/09 for training the global-regional model (M global−regional ), with two independent eval- uation periods defined as: Hindcast evaluation (ERA5 input): 2024/10–2024/12 and Operational evaluation (operational anal- ysis input): 2025/01–2025/04. All raw variables are also z-score normalized before fed into the model. More details of datasets and experimental settings can be found in Section B. 4.2 Evaluation Metrics and Baselines Evaluation metrics. To measure the performance of regional weather forecasting, we evaluate all methods using latitude-weighted root mean squared error (RMSE) and latitude-weighted anomaly correlation coefficient (ACC). More details of metrics can be found in Section C. Baselines. We comprehensively evaluate ScaleMixer against sev- eral strong baselines: (1) our internal global model (M global ); (2) a standalone regional model initialized from CLDAS data without global coupling (M regional ); (3) an AI-based global forecast model Skillful Kilometer-Scale Regional Weather Forecasting via Global and Regional CouplingConference’17, July 2017, Washington, DC, USA Baguan [22], which demonstrates superior performance among a set of state-of-the-art data-driven weather models on Weather- Bench 1 and provides more comprehensive surface meteorological variables than other global models like Pangu-Weather [2] and GraphCast [17] (detailed comparisons can be found in Section F); (4) the operational high-resolution NWP system IFS-HRES from ECMWF [9], which serves as a gold-standard reference. The res- olutions of AI-based global forecasts and IFS-HRES are 0.25 ◦ and 0.1 ◦ , respectively, and there may exist systematic bias between their forecast values and CLDAS. For a fair comparison, we employ a downscaling and correction model to map the original forecast values to the target 0.05 ◦ grids. The downscaling and correction model is trained on ERA5 and CLDAS, using Swin Transformer [18] as the backbone, containing 4 layers block and a hidden dimension of 192; (5) OneForecast [11], which introduces a Neural Nested Grid method that typically passes boundary feature maps between grids of different resolutions via direct interpolation and concatenation; and (6) Limited Area Model (LAM), which is built uponM 푟푒푔푖표푛푎푙 (Standalone Regional Model) but enhanced to take both regional initial conditions and external global forecasts as input. 4.3 Skillful Regional Weather Forecasting at 0.05 ◦ Resolution We focus on short-term forecasting for next 48 hours, primarily because such outlooks have a more immediate impact on soci- etal functions and daily routines. Furthermore, this is the period that NWP models have optimal performances. The deterministic forecasting results of ScaleMixer and baselines are summarized in Table 1, and Figure 2, evaluating forecast skill acrossΔ푡=1∼48 hours lead time. Hindcast evaluation. For the ERA5-driven hindcast period (2024/10 – 2024/12), ScaleMixer achieves significant improvements across all seven surface variables (T2M, U10, V10, Q, P, TCC, and SSRD) compared to both standalone global/regional baselines and IFS- HRES (Table 1), verifying the effectiveness of coupling global and regional scales. ScaleMixer achieves 40.86% lower latitude-weighted RMSE and 9.96% higher ACC compared to IFS-HRES, indicating enhanced resolution capability for mesoscale convective systems and boundary layer dynamics. As shown in Figure 2a, performance advantages persist consistently across forecast horizons. Operational forecast evaluation. Under dynamically evolving op- erational initial conditions (2025/01–2025/04), ScaleMixer main- tains superior skill despite real-time analysis field uncertainties (Figure 2b). Compared to IFS-HRES at 0.1° resolution, statistically significant RMSE improvements are sustained through 48-hour lead times under operational constraints, with pronounced im- provements in 1–24-hours ahead predictions where regional-scale processes dominate. Evaluation against station observations. Although we have used downscaling models to map IFS-HRES to CLDAS, there may still ex- ist some fairness concerns since ScaleMixer is end-to-end trained on the ground truth. We make extensive comparisons against station observations, which is more fair. The real-time station observation 1 https://sites.research.google/gr/weatherbench/scorecards-2020/ Figure 4: Station distrbution map. The station observation dataset contains 2216 weather stations across China, which record hourly observations of meteorological variables such as temperature, air pressure, and wind speed. dataset is a product provided by China Meterological Administra- tions. It contains more than 2000 stations distributed across China, with a portion located in complex terrain areas such as plateaus and mountainous regions, as shown in Figure 4. These stations have undergone quality control and provide hourly observations of several meteorological variables. We focus on some key variables with high importance for downstream usages and provided in our model: 2m temperature (T2M), 2m dewpoint temperature (D2M) and 10m wind speed (WS10). In implementation, we interpolate the grid predicitions of ScaleMixer and IFS-HRES to stations according to their latitudes and longitudes. Table 2 shows the RMSE of IFS- HRES, Baguan and ScaleMixer. Compared to IFS-HRES and Baguan, ScaleMixer achieves an average improvement of 27.63% across the three meteorological variables. Table 2: Average RMSE acrossΔ푡=1∼48 hours lead time regional weather operational forecasting against weather station observations. Variable IFS-HRES Baguan ScaleMixerΔRMSE T2M3.1583.1081.822 ↓41.4% D2M2.8122.6551.789 ↓36.4% WS101.4011.4021.329 ↓5.10% 4.4 Case Studies Orographic-induced wind and temperature. As exemplified in Fig- ure 3 (left) for wind prediction of the complex terrain regions in the Hengduan Mountains (25.0–35.0°N, 95.0–105.0°E) China, ScaleMixer (0.05°) resolves wind characteristics across topographic gradients: maximum wind speed at mountain crests (exceeding 10 m/s) and deceleration within valleys (<2 m/s). This contrasts with IFS-HRES (0.1°) which exhibits systematic underestimation of orographic wind characteristics due to insufficient subgrid-scale orographic parametrization. Conference’17, July 2017, Washington, DC, USAWeiqi Chen, Wenwei Wang, Qilong Yuan, Lefei Shen, Bingqing Peng, Jiawei Chen, Bo Wu, and Liang Sun Table 3: Ablation study of ScaleMixer components on 48-hour forecast performance. Model Variant Configuration DetailsT2M RMSEU10 RMSE ScaleMixer Adaptive Sampling + Bidirectional1.3821.644 Variant ARandom Sampling + Bidirectional1.605 (+16.1%) 1.882 (+14.5%) Variant BFixed Uniform Grid + Bidirectional1.512 (+9.4%)1.795 (+9.2%) Variant CAdaptive Sampling + Unidirectional1.468 (+6.2%)1.721 (+4.7%) Variant DNo Interaction (Standalone)1.991 (+44.1%) 1.967 (+19.6%) Moreover, the same orographic forcing that generates wind heterogeneity also drives temperature variations. ScaleMixer re- solves pronounced temperature contrasts across elevation gradients (Fig. 3, right), with leeward slopes exhibiting 4–8°C warming rela- tive to windward sides, a canonical Foehn effect signature 2 , arising from adiabatic compression of descending air masses. In contrast, IFS-HRES underestimates these temperature gradients, failing to capture dependencies between terrain steepness and temperature variation. The enhanced resolution with data-driven method in ScaleMixer enables superior representation fine-grained weather features in complex terrain. Additional visualizations of forecasts are provided in Section E, which demonstrate the framework’s capability to capture high- resolution meteorological details. 4.5 Ablation Studies To rigorously validate the architectural design of ScaleMixer, we conducted fine-grained ablation studies focusing on two core di- mensions: the sampling strategy for key position identification and the directionality of cross-scale coupling. Furthermore, we analyzed the sensitivity of the model to critical hyperparameters. All ablation experiments were conducted on the validation set with a forecast lead time ofΔ푡= 24 hours. Effectiveness of ScaleMixer Components. We compared our proposed framework against four variants: (A) Random Sampling, replacing adaptive identification with random selection; (B) Fixed Uniform Grid, utilizing a static grid for interaction; (C) Unidirectional Coupling, allowing only global-to-regional information flow; and (D) No Interaction, equivalent to the standalone regional model. The results are summarized in Table 3. The results demonstrate: (1) Effectiveness of Adaptive Sam- pling: The proposed adaptive key position identification signifi- cantly outperforms the Fixed Uniform Grid (Variant B), reducing T2M RMSE by 9.4%. This demonstrates that dynamically focusing computation on meteorologically active regions (e.g., high-gradient boundaries) is far more efficient than uniform processing, which may waste capacity on static areas; and (2) Effectiveness of Bidi- rectional Coupling: Compared to unidirectional coupling (Variant C), our bidirectional mechanism achieves a 6.2% improvement in T2M RMSE. This confirms that allowing high-resolution regional features to explicitly refine global tokens creates a necessary closed- loop feedback, enhancing the consistency of the synoptic-scale context. 2 https://en.wikipedia.org/wiki/Foehn_wind Hyperparameter Sensitivity. We further investigated the sen- sitivity of the regional encoder depth (푘). As shown in Table 5, increasing the number of layers from푘=2 to푘=4 yields signif- icant gains, while푘=8 offers diminishing returns with doubled computational cost. Based on these results, we adopted푘=4 as the default configuration to balance forecasting accuracy and inference efficiency. Table 4: Sensitivity analysis of Regional Encoder Layers (푘). Metric Regional Encoder Layers (푘 ) 푘= 2푘= 4푘= 8 T2M RMSE1.4851.3821.379 U10 RMSE1.7521.6441.641 Inference Time22ms28ms51ms 5 Conclusion In this paper, we present a multiscale deep learning framework for high-resolution regional weather forecasting that bridges synoptic- scale dynamics with localized mesoscale processes. By integrating a pretrained global foundation model and a novel bidirectional global-regional coupling module, ScaleMixer achieves state-of-the- art performance in resolving complex weather phenomena at 0.05 ◦ (∼5km) resolution. Experimental results establish ScaleMixer as a robust data-driven approach for regional weather forecasting. In the future, we will extend the framework to probabilistic forecasting and assimilate multi-modal observations (e.g., radar, satellite) for real-time forecasting. Limitations and Ethical Considerations Limitations. While ScaleMixer demonstrates significant advance- ments in regional weather forecasting, it is a purely data-driven model and lacks explicit equation-based constraints (e.g., Navier–Stokes or hydrostatic balance), which can cause unphysical artifacts in long roll-outs, especially in extreme regimes. Future work will add physics-informed regularization to better enforce conservation laws. Ethical Considerations Both global reanalysis and global op- erational analysis data are publicly available. The CLDAS dataset and weather station observations were obtained from the China Meteorological Administration, and we have been granted permis- sion to use them for academic research, which poses no potential ethical risks. Skillful Kilometer-Scale Regional Weather Forecasting via Global and Regional CouplingConference’17, July 2017, Washington, DC, USA References [1]Simon Adamov, Joel Oskarsson, Leif Denby, Tomas Landelius, Kasper Hintz, Simon Christiansen, Irene Schicker, Carlos Osuna, Fredrik Lindsten, Oliver Fuhrer, et al.2025. Building Machine Learning Limited Area Models: Kilometer-Scale Weather Forecasting in Realistic Settings. arXiv preprint arXiv:2504.09340 (2025). [2]Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian. 2023. Accurate medium-range global weather forecasting with 3D neural networks. Nature 619, 7970 (2023), 533–538. [3]Zied Ben Bouallègue, Mariana C. A. Clare, Linus Magnusson, Estibaliz Gascón, Michael Maier-Gerber, Martin Janoušek, Mark Rodwell, Florian Pinault, Jesper S. Dramsch, Simon T. K. Lang, Baudouin Raoult, Florence Rabier, Matthieu Cheval- lier, Irina Sandu, Peter Dueben, Matthew Chantry, and Florian Pappenberger. 2024. The Rise of Data-Driven Weather Forecasting: A First Statistical Assess- ment of Machine Learning–Based Weather Forecasts in an Operational-Like Context. Bulletin of the American Meteorological Society 105, 6 (2024), E864 – E883. doi:10.1175/BAMS-D-23-0162.1 [4] Kang Chen, Tao Han, Junchao Gong, Lei Bai, Fenghua Ling, Jing-Jia Luo, Xi Chen, Leiming Ma, Tianning Zhang, Rui Su, et al.2023. FengWu: Pushing the Skillful Global Medium-range Weather Forecast beyond 10 Days Lead. arXiv preprint arXiv:2304.02948 (2023). [5]Lei Chen, Xiaohui Zhong, Feng Zhang, Yuan Cheng, Yinghui Xu, Yuan Qi, and Hao Li. 2023. FuXi: A cascade machine learning forecasting system for 15-day global weather forecast. npj Climate and Atmospheric Science 6, 1 (2023), 190. [6]Jean Coiffier. 2011. Fundamentals of numerical weather prediction. Cambridge University Press, Cambridge; New York. [7]Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al.2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020). [8] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations. [9] ECMWF. 2023. IFS Documentation CY48R1. ECMWF. [10] European Centre for Medium-Range Weather Forecasts. 2023. Description of the Integrated Forecasting System (IFS) (cycle 48r1 ed.). https://w.ecmwf.int/en/ publications/manuals/deterministic-model [11] Yuan Gao, Hao Wu, Ruiqi Shu, Huanshuo Dong, Fan Xu, Rui Chen, Yibo Yan, Qingsong Wen, Xuming Hu, Kun Wang, et al.2025. OneForecast: A Univer- sal Framework for Global and Regional Weather Forecasting. arXiv preprint arXiv:2502.00338 (2025). [12] Hans Hersbach, Bill Bell, Paul Berrisford, Shoji Hirahara, András Horányi, Joaquín Muñoz-Sabater, Julien Nicolas, Carole Peubey, Raluca Radu, Dinand Schepers, et al.2020. The ERA5 global reanalysis. Quarterly Journal of the Royal Meteoro- logical Society 146, 730 (2020), 1999–2049. [13]James Hurrell, Marika Holland, Peter Gent, Steven Ghan, Jennifer Kay, Paul Kushner, J-F Lamarque, William Large, D Lawrence, Keith Lindsay, et al.2013. The community earth system model: a framework for collaborative research. Bulletin of the American Meteorological Society 94, 9 (2013), 1339–1360. [14]Ryan Keisler. 2022. Forecasting global weather with graph neural networks. arXiv preprint arXiv:2202.07575 (2022). [15]Dmitrii Kochkov, Janni Yuval, Ian Langmore, Peter Norgaard, Jamie Smith, Griffin Mooers, Milan Klöwer, James Lottes, Stephan Rasp, Peter Düben, et al.2024. Neural general circulation models for weather and climate. Nature 632, 8027 (2024), 1060–1066. [16]Remi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Wirnsberger, Meire Fortunato, Ferran Alet, Suman Ravuri, Timo Ewalds, Zach Eaton-Rosen, Weihua Hu, et al.2023. Learning skillful medium-range global weather forecasting. Science 382, 6677 (2023), 1416–1421. [17]Remi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Wirnsberger, Meire Fortunato, Ferran Alet, Suman Ravuri, Timo Ewalds, Zach Eaton-Rosen, Weihua Hu, et al.2023. Learning skillful medium-range global weather forecasting. Science 382, 6677 (2023), 1416–1421. [18]Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer us- ing shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision. 10012–10022. [19]Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101 (2017). [20]Morteza Mardani, Noah Brenowitz, Yair Cohen, Jaideep Pathak, Chieh-Yu Chen, Cheng-Chin Liu, Arash Vahdat, Mohammad Amin Nabian, Tao Ge, Akshay Subramaniam, et al.2025. Residual corrective diffusion modeling for km-scale atmospheric downscaling. Communications Earth & Environment 6, 1 (2025), 124. [21]Thomas Nils Nipen, Håvard Homleid Haugen, Magnus Sikora Ingstad, Even Mar- ius Nordhagen, Aram Farhad Shafiq Salihi, Paulina Tedesco, Ivar Ambjørn Seier- stad, Jørn Kristiansen, Simon Lang, Mihai Alexe, et al.2024. Regional data-driven weather modeling with a global stretched-grid. arXiv preprint arXiv:2409.02891 (2024). [22]Peisong Niu, Ziqing Ma, Tian Zhou, Weiqi Chen, Lefei Shen, Rong Jin, and Liang Sun. 2025. Utilizing strategic pre-training to reduce overfitting: Baguan-a pre-trained weather forecasting model. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 2186–2197. [23] Joel Oskarsson, Tomas Landelius, and Fredrik Lindsten. 2023. Graph-based neural weather prediction for limited area modeling. arXiv preprint arXiv:2309.17370 (2023). [24]William Peebles and Saining Xie. 2023. Scalable diffusion models with transform- ers. In Proceedings of the IEEE/CVF international conference on computer vision. 4195–4205. [25] Ilan Price, Alvaro Sanchez-Gonzalez, Ferran Alet, Timo Ewalds, Andrew El-Kadi, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson. 2023. GenCast: Diffusion-based ensemble forecasting for medium-range weather. arXiv preprint arXiv:2312.15796 (2023). [26]Haoyu Qin, Yungang Chen, Qianchuan Jiang, Pengchao Sun, Xiancai Ye, and Chao Lin. 2024. Metmamba: Regional weather forecasting with spatial-temporal mamba model. arXiv preprint arXiv:2408.06400 (2024). [27]Florence Rabier and Zhiquan Liu. 2003. Variational data assimilation: theory and overview. In Proc. ECMWF Seminar on Recent Developments in Data Assimilation for Atmosphere and Ocean, Reading, UK. 29–43. [28]Stephan Rasp, Peter D Dueben, Sebastian Scher, Jonathan A Weyn, Soukayna Mouatadid, and Nils Thuerey. 2020. WeatherBench: a benchmark data set for data-driven weather forecasting. Journal of Advances in Modeling Earth Systems 12, 11 (2020), e2020MS002203. [29]Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017). [30] Ruolan Xiang, Christian R Steger, Shuping Li, Loïc Pellissier, Silje Lund Sør- land, Sean D Willett, and Christoph Schär. 2024. Assessing the regional climate response to different Hengduan Mountains geometries with a high-resolution regional climate model. Journal of Geophysical Research: Atmospheres 129, 6 (2024), e2023JD040208. [31] Pengbo Xu, Xiaogu Zheng, Tianyan Gao, Yu Wang, Junping Yin, Juan Zhang, Xuanze Zhang, San Luo, Zhonglei Wang, Zhimin Zhang, et al.[n. d.]. YingLong- weather: AI-Based Limited Area Models for Forecasting. ([n. d.]). Conference’17, July 2017, Washington, DC, USAWeiqi Chen, Wenwei Wang, Qilong Yuan, Lefei Shen, Bingqing Peng, Jiawei Chen, Bo Wu, and Liang Sun A Implementation Details In our multiscale regional weather forecasting framework, the back- bones of the global model (M global ) and regional model (M regional ) are based on ViT [8]. The framework contains 1.07 billion parame- ters, with the global modelM global accounting for 736 million. The hyperparameter configurations of the model are summarized in Table 5. B Dataset Details and Experimental Settings In our experiments, we use the preprocessed ERA5 data from WeatherBench [28]. EAR5 is a well-acknowledged weather fore- casting benchmark dataset and it is widely used in data-driven weather forecasting methods. WeatherBench processed the raw ERA5 dataset 3 , which includes 8 atmospheric variables across 13 pressure levels, 6 surface variables, and 3 static variables. We nor- malize all the inputs via z-score normalization for each variable at each pressure level. Also, we apply the inverse normalization for the predictions of future states for performance evaluation. We collected and processed operational analysis data, which are used for operational forecast, from initial conditions of ECMWF’s High-Resolution Deterministic Prediction (HRES) system, assim- ilating observations with 4D-variational data assimilation. The 0.1° analysis fields (interpolated to ERA5 resolution, 0.25°) provide dynamically real-time initial conditions. We set and process the atmosphere variables consistent with the ERA5 Dataset. We also collected and processed the China Meteorological Ad- ministration’s Land Data System data (CLDAS), which offers 0.01° resolution meteorological fields over East Asia (0–65°N, 60–160°E). The regional analysis dataset is used to train and evaluate regional weather forecasting model. The dataset includes 7 critical surface variables: wind components (U, V), temperature (T), specific hu- midity (Q), pressure (P), radiation fluxes (SSRD), and total cloud cover (TCC). We normalize all the inputs via z-score normalization for each variable at each pressure level. Also, we apply the inverse normalization for the predictions of future states for performance evaluation. B.1 ERA5 and operational anlysis with 0.25° resolution we selected 6 atmospheric variables at all 13 pressure levels, 3 surface variables, and 3 static variables for the ERA5 dataset with 0.25° resolution, as detailed in Table 6. In our model training, we choose all variables as input variables, and all variables except three static variables as output variables that are used for loss calculation to pretrain global modelM global . B.2 CLDAS with 0.05° resolution We selected 7 surface variables for the CLDAS dataset with 0.25° resolution, as detailed in Table 7. In our model training, we choose all variables as input variables, and all variables as output variables that are used for loss calculation to train global-regional model M global−regional . 3 More details of ERA5 data can be found in https://confluence.ecmwf.int/display/CKB/ ERA5%3A+data+documentation. C Evaluation Metrics for Regional Weather Forecasting This section provides detailed explanations of all the evaluation met- rics for regional weather forecasting used in the main experiments. For each metric,풖and ˆ 풖represent the predicted and ground truth values, respectively, both shaped asℎ×푤 ×푉 reg , where푉 reg is the number of total weather factors, andℎ×푤is the spatial resolution of latitude (ℎ) and longitude (푤). To account for the non-uniform grid cell areas, the latitude weighting term 훼(·) is introduced. Latitude-weighted Root Mean Square Error (RMSE). assesses model accuracy while considering the Earth’s curvature. The lati- tude weighting adjusts for the varying grid cell areas at different latitudes, ensuring that errors are appropriately measured. Lower RMSE values indicate better model performance. RMSE= 1 푉 reg 푉 reg ∑︁ 푘=1 v u t 1 ℎ푤 ℎ ∑︁ 푖=1 푤 ∑︁ 푗=1 훼(푖) ˆ 풖 푖,푗,푘 −풖 푖,푗,푘 2 , 훼(푖)= cos(lat(푖)) 1 ℎ Í ℎ 푖 ′ =1 cos ( lat ( 푖 ′ )) . Anomaly Correlation Coefficient (ACC). measures a model’s ability to predict deviations from the mean. Higher ACC values indicate better accuracy in capturing anomalies, which is crucial in meteorology and climate science. ACC= Í 푖,푗,푘 ˆ 풖 ′ 푖,푗,푘 풖 ′ 푖,푗,푘 √︃ Í 푖,푗,푘 훼(ℎ)( ˆ 풖 ′ 푖,푗,푘 ) 2 Í 푖,푗,푘 훼(ℎ)(풖 ′ 푖,푗,푘 ) 2 , where풖 ′ =풖−퐶and ˆ 풖 ′ = ˆ 풖−퐶, with climatology퐶representing the temporal mean of the same period over the training set. D Broader Impacts This research focuses on high-resolution regional weather fore- casting, which has an essential influence on relevant fields such as energy, transportation, and agriculture. As an AI application for so- cial good, our model boosts predictions for various weather factors such as temperature, wind speed, and radiation flux. It is essential to note that our work focuses solely on scientific issues, and we also ensure that ethical considerations are carefully taken into account. Thus, we believe that there is no ethical risk associated with our research. E Visualization of Forecasts To intuitively demonstrate the forecasting capacity of our model, we present the showcases of weather forecasting results in China and zoomed-in regions in Figures 5, 6, 7 8, 9, 10, 11, 12, 13, 14, 15, and 16. F Comparison with Weatherbench2’s Baselines The global branch of our frameworkM 푔푙표푏푎푙 demonstrates promis- ing performance that aligns with top-tier AI weather models. Fig- ure 17 shows RMSE and ACC for some key variables of our global Skillful Kilometer-Scale Regional Weather Forecasting via Global and Regional CouplingConference’17, July 2017, Washington, DC, USA Table 5: Default hyperparameters of the framework. ModuleHyperparameterDescriptionValue Global Model M global 푃Patch size of global tokens6 푑hidden dimension1536 푀Number of Transformer encoder layers of in the global model24 HeadsNumber of attention heads8 MLP ratioExpansion factor for MLP4. Depth of prediction headNumber of deconvolution layers of the final prediction head2 Drop pathStochastic depth rate0.1 DropoutDropout rate0.1 Regional Model M regional 푝Patch size of regional tokens30 푑hidden dimension1536 푘Number of Transformer encoder layers of in the regional model4 HeadsNumber of attention heads8 MLP ratioExpansion factor for MLP4. Depth of prediction headNumber of deconvolution layers of the final prediction head2 Drop pathStochastic depth rate0.1 DropoutDropout rate0.1 ScaleMixer Depthtotal number of ScaleMix modules4 Depth of position identification blocknumber of convolution layers in position identification block1 Kernel sizekernel size of convolution layers in position identification block3 푚number of key positions64 model and Pangu-Weather [2], Fuxi [5], Graphcast [17] and IFS- HRES [9], reported on Weatherbench2 4 . It is evaluated on 2022 ERA5 dataset, with lead time ranging from 6 to 240 hours. For both surface variables and pressure level variables, our model out- performs Pangu-Weather, Granphcast and EC-IFS and is compa- rable to Fuxi in most cases. Notably, our global model achieves an average improvement of 12.59%, 0.90%, 3.56%, 16.30% on RMSE over pangu-weather, fuxi, graphcast, EC-IFS, respectively, and an average improvement of 8.07%, 0.87%, 4.32%, 9.51% on ACC over pangu-weather, fuxi, graphcast, EC-IFS, respectively. 4 https://sites.research.google/gr/weatherbench/deterministic-scores/ Conference’17, July 2017, Washington, DC, USAWeiqi Chen, Wenwei Wang, Qilong Yuan, Lefei Shen, Bingqing Peng, Jiawei Chen, Bo Wu, and Liang Sun Table 6: Summary of ECMWF variables utilized in the ERA5 and operational analysis dataset with 0.25° resolution. The variables 푙푠푚 and 표푟표 are constant and invariant with time. TypeVariable NameAbbrev.DescriptionPressure Levels Static Variable Land-sea mask푙푠푚Binary mask distinguishing land (1) from sea (0)N/A Orography표푟표Height of Earth’s surfaceN/A Latitude푙푎푡Latitude of each grid pointN/A Surface Variable 2 metre temperature푡2푚Temperature measured 2 meters above the surfaceSingle level 10 metre U wind component 푢10East-west wind speed at 10 meters above the surfaceSingle level 10 metre V wind component 푣10North-south wind speed at 10 meters above the surfaceSingle level Mean sea level presure푚푠푙 Pressure of the atmosphere adjusted to the height of mean sea level Single level Surface pressure푠푝 Pressure of the atmosphere on the surface of land, sea and in- land water Single level 2 metre dewpoint temperature 푑2푚 Temperature to which the air, at 2 metres above the surface of the Earth Single level Upper-air Variable Geopotential푧Height relative to a pressure level 50, 100,150, 200, 250,300, 400, 500, 600, 700, 850, 925,1000 hPa U wind component푢Wind speed in the east-west direction50, 100,150, 200, 250,300, 400, 500, 600, 700, 850, 925,1000 hPa V wind component푣Wind speed in the north-south direction 50, 100,150, 200, 250,300, 400, 500, 600, 700, 850, 925,1000 hPa Temperature푡Atmospheric temperature50, 100,150, 200, 250,300, 400, 500, 600, 700, 850, 925,1000 hPa Specific humidity푞Mixing ratio of water vapor to total air mass50, 100,150, 200, 250,300, 400, 500, 600, 700, 850, 925,1000 hPa Table 7: Summary of variables utilized in CLDAS with 0.05° resolution. TypeVariable NameAbbrev.Description Surface Variable 2 metre temperature푇Temperature measured 2 meters above the surface 10 metre U wind component푈East-west wind speed at 10 meters above the surface 10 metre V wind component푉North-south wind speed at 10 meters above the surface Surface specific humidity푄Mixing ratio of water vapor to total air mass at 2 meters above the surface Surface pressure푃Pressure of the atmosphere on the surface of land, sea and in-land water Total cloud cover푇퐶Cloud occurring at different model levels through the atmosphere Radiation flux (surface solar radiation flux downwards) 푆푅퐷Flux of solar radiation that reaches a horizontal plane at the surface of the Earth Skillful Kilometer-Scale Regional Weather Forecasting via Global and Regional CouplingConference’17, July 2017, Washington, DC, USA 2024/11/1923UTC(Forecastleadtime=11) 2024/11/1911UTC(Forecastleadtime=23) Figure 5: 2 metre temperature forecasts over China Conference’17, July 2017, Washington, DC, USAWeiqi Chen, Wenwei Wang, Qilong Yuan, Lefei Shen, Bingqing Peng, Jiawei Chen, Bo Wu, and Liang Sun 2024/11/1923UTC(Forecastleadtime=11) 2024/11/1911UTC(Forecastleadtime=23) Figure 6: 2 metre temperature forecasts over a subregion of latitudes in [30, 40] and longitudes in [95, 115] Skillful Kilometer-Scale Regional Weather Forecasting via Global and Regional CouplingConference’17, July 2017, Washington, DC, USA 2024/11/1923UTC(Forecastleadtime=11) 2024/11/1911UTC(Forecastleadtime=23) Figure 7: surface pressure forecasts over China Conference’17, July 2017, Washington, DC, USAWeiqi Chen, Wenwei Wang, Qilong Yuan, Lefei Shen, Bingqing Peng, Jiawei Chen, Bo Wu, and Liang Sun 2024/11/1923UTC(Forecastleadtime=11) 2024/11/1911UTC(Forecastleadtime=23) Figure 8: surface pressure forecasts over a subregion of latitudes in [30, 40] and longitudes in [95, 115] Skillful Kilometer-Scale Regional Weather Forecasting via Global and Regional CouplingConference’17, July 2017, Washington, DC, USA 2024/11/1923UTC(Forecastleadtime=11) 2024/11/1911UTC(Forecastleadtime=23) Figure 9: 10 metre Wind speed U component forecasts over China Conference’17, July 2017, Washington, DC, USAWeiqi Chen, Wenwei Wang, Qilong Yuan, Lefei Shen, Bingqing Peng, Jiawei Chen, Bo Wu, and Liang Sun 2024/11/1923UTC(Forecastleadtime=11) 2024/11/1911UTC(Forecastleadtime=23) Figure 10: 10 metre wind speed U component forecasts over a subregion of latitudes in [16.3, 26] and longitudes in [115, 125] Skillful Kilometer-Scale Regional Weather Forecasting via Global and Regional CouplingConference’17, July 2017, Washington, DC, USA 2024/11/1923UTC(Forecastleadtime=11) 2024/11/1911UTC(Forecastleadtime=23) Figure 11: 10 metre Wind speed V component forecasts over China Conference’17, July 2017, Washington, DC, USAWeiqi Chen, Wenwei Wang, Qilong Yuan, Lefei Shen, Bingqing Peng, Jiawei Chen, Bo Wu, and Liang Sun 2024/11/1923UTC(Forecastleadtime=11) 2024/11/1911UTC(Forecastleadtime=23) Figure 12: 10 metre wind speed V component forecasts over a subregion of latitudes in [16.3, 26] and longitudes in [115, 125] Skillful Kilometer-Scale Regional Weather Forecasting via Global and Regional CouplingConference’17, July 2017, Washington, DC, USA 2025/01/192UTC(Forecastleadtime=14) 2025/01/194UTC(Forecastleadtime=16) Figure 13: surface solar radiation downwards forecasts over China Conference’17, July 2017, Washington, DC, USAWeiqi Chen, Wenwei Wang, Qilong Yuan, Lefei Shen, Bingqing Peng, Jiawei Chen, Bo Wu, and Liang Sun 2025/01/192UTC(Forecastleadtime=14) 2025/01/194UTC(Forecastleadtime=16) Figure 14: surface solar radiation downwards forecasts over a subregion of latitudes in [35, 50] and longitudes in [115, 135] Skillful Kilometer-Scale Regional Weather Forecasting via Global and Regional CouplingConference’17, July 2017, Washington, DC, USA 2025/01/1813UTC(Forecastleadtime=1) 2025/01/1820UTC(Forecastleadtime=8) Figure 15: total cloud cover forecasts over China Conference’17, July 2017, Washington, DC, USAWeiqi Chen, Wenwei Wang, Qilong Yuan, Lefei Shen, Bingqing Peng, Jiawei Chen, Bo Wu, and Liang Sun 2025/01/1813UTC(Forecastleadtime=1) 2025/01/1820UTC(Forecastleadtime=8) Figure 16: total cloud cover forecasts over a subregion of latitudes in [35, 50] and longitudes in [115, 135] Figure 17: RMSE and ACC for some key variables of various AI-based weather models on Weatherbench2