Paper deep dive
Enhancing AI-Based Tropical Cyclone Track and Intensity Forecasting via Systematic Bias Correction
Peisong Niu, Haifan Zhang, Yang Zhao, Tian Zhou, Ziqing Ma, Wenqiang Shen, Junping Zhao, Huiling Yuan, Liang Sun
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/26/2026, 1:29:00 AM
Summary
BaguanCyclone is a unified AI framework designed to improve tropical cyclone (TC) track and intensity forecasting by addressing discretization errors and smoothing effects inherent in coarse-resolution reanalysis data. It introduces a probabilistic center refinement module for sub-grid track precision and a region-aware intensity forecasting module that utilizes high-resolution internal representations. The system outperforms operational NWP models and existing AI baselines, demonstrating success in complex scenarios like re-intensification and multi-TC interactions, and has been operationally deployed at the Zhejiang Meteorological Observatory.
Entities (5)
Relation Signals (3)
BaguanCyclone → deployedat → Zhejiang Meteorological Observatory
confidence 100% · The system has been fully operational at the Zhejiang Meteorological Observatory
BaguanCyclone → evaluatedon → IBTrACS
confidence 100% · Evaluated on the global IBTrACS dataset across six major TC basins
BaguanCyclone → improves → Tropical Cyclone Forecasting
confidence 95% · Enhancing AI-Based Tropical Cyclone Track and Intensity Forecasting via Systematic Bias Correction
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Tropical cyclones (TCs) pose severe threats to life, infrastructure, and economies in tropical and subtropical regions, underscoring the critical need for accurate and timely forecasts of both track and intensity. Recent advances in AI-based weather forecasting have shown promise in improving TC track forecasts. However, these systems are typically trained on coarse-resolution reanalysis data (e.g., ERA5 at 0.25 degree), which constrains predicted TC positions to a fixed grid and introduces significant discretization errors. Moreover, intensity forecasting remains limited especially for strong TCs by the smoothing effect of coarse meteorological fields and the use of regression losses that bias predictions toward conditional means. To address these limitations, we propose BaguanCyclone, a novel, unified framework that integrates two key innovations: (1) a probabilistic center refinement module that models the continuous spatial distribution of TC centers, enabling finer track precision; and (2) a region-aware intensity forecasting module that leverages high-resolution internal representations within dynamically defined sub-grid zones around the TC core to better capture localized extremes. Evaluated on the global IBTrACS dataset across six major TC basins, our system consistently outperforms both operational numerical weather prediction (NWP) models and most AI-based baselines, delivering a substantial enhancement in forecast accuracy. Remarkably, BaguanCyclone excels in navigating meteorological complexities, consistently delivering accurate forecasts for re-intensification, sweeping arcs, twin cyclones, and meandering events. Our code is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2603.22314v1
- Canonical: https://arxiv.org/abs/2603.22314v1
Trouble viewing inline? Open PDF directly →
Full Text
53,000 characters extracted from source content.
Expand or collapse full text
Enhancing AI-Based Tropical Cyclone Track and Intensity Forecasting via Systematic Bias Correction Peisong Niu ∗ niupeisong.nps@alibaba-inc.com DAMO Academy, Alibaba Group Hangzhou, China Haifan Zhang ∗ zhanghaifan.zhf@alibaba-inc.com DAMO Academy, Alibaba Group Hangzhou, China Yang Zhao ∗ zhaoy2024@smail.nju.edu.cn State Key Laboratory of Severe Weather Meteorological Science and Technology, Nanjing University Nanjing, China Tian Zhou ∗ tian.zt@alibaba-inc.com DAMO Academy, Alibaba Group Hangzhou, China Ziqing Ma maziqing.mzq@alibaba-inc.com DAMO Academy, Alibaba Group Hangzhou, China Wenqiang Shen wqshen91@163.com Zhejiang Meteorological Observatory, China Meteorological Administration Hangzhou, China Junping Zhao jpzhaozj@q.com Zhejiang Meteorological Observatory, China Meteorological Administration Hangzhou, China Huiling Yuan yuanhl@nju.edu.cn State Key Laboratory of Severe Weather Meteorological Science and Technology, Nanjing University Nanjing, China Liang Sun † liang.sun@alibaba-inc.com DAMO Academy, Alibaba Group Bellevue, USA Abstract Tropical cyclones (TCs) pose severe threats to life, infrastructure, and economies in tropical and subtropical regions, underscoring the critical need for accurate and timely forecasts of both track and intensity. Recent advances in AI-based weather forecasting have shown promise in improving TC track forecasts. However, these systems are typically trained on coarse-resolution reanalysis data (e.g., ERA5 at 0.25 ◦ ), which constrains predicted TC positions to a fixed grid and introduces significant discretization errors. More- over, intensity forecasting remains limited especially for strong TCs by the smoothing effect of coarse meteorological fields and the use of regression losses that bias predictions toward conditional means. To address these limitations, we propose BaguanCyclone, a novel, unified framework that integrates two key innovations: (1) a probabilistic center refinement module that models the continuous spatial distribution of TC centers, enabling finer track precision; and (2) a region-aware intensity forecasting module that lever- ages high-resolution internal representations within dynamically defined sub-grid zones around the TC core to better capture lo- calized extremes. Evaluated on the global IBTrACS dataset across ∗ Authors contributed equally to this research. † Corresponding authors. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. Conference acronym ’X, Woodstock, NY © 2018 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-X-X/2018/06 https://doi.org/X.X six major TC basins, our system consistently outperforms both operational numerical weather prediction (NWP) models and most AI-based baselines, delivering a substantial enhancement in forecast accuracy. Remarkably, BaguanCyclone excels in navigating meteo- rological complexities, consistently delivering accurate forecasts for re-intensification, sweeping arcs, twin cyclones, and meander- ing events. The system has been fully operational at the Zhejiang Meteorological Observatory, China Meteorological Administration (CMA) during the 2025 Western Pacific typhoon season. It pro- vided real-time forecasts that critically underpinned emergency response, public warnings, and maritime safety operations. This work represents a significant step toward the operational integra- tion of physically informed high-fidelity deep learning models in high-stakes meteorological applications. Our code is available at https://github.com/DAMO-DI-ML/Baguan-cyclone. CCS Concepts • Applied computing→Earth and atmospheric sciences;• Computing methodologies→ Neural networks. Keywords Tropical Cyclone, Track Forecasting, Intensity Forecasting, Deep Learning, Bias Correction ACM Reference Format: Peisong Niu, Haifan Zhang, Yang Zhao, Tian Zhou, Ziqing Ma, Wenqiang Shen, Junping Zhao, Huiling Yuan, and Liang Sun. 2018. Enhancing AI- Based Tropical Cyclone Track and Intensity Forecasting via Systematic Bias Correction. In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (Conference acronym ’X). ACM, New York, NY, USA, 10 pages. https://doi.org/X.X arXiv:2603.22314v1 [cs.LG] 20 Mar 2026 Conference acronym ’X, June 03–05, 2018, Woodstock, NYTrovato et al. 1 Introduction Tropical cyclones (TCs) [12] represent one of the most hazardous weather phenomena on Earth, routinely causing catastrophic dam- age to life, infrastructure, and economies across tropical and ad- jacent subtropical regions. Accurate and timely prediction of TC tracks and intensity forecasts remains a cornerstone of operational meteorology, directly influencing evacuation planning, maritime safety, emergency response coordination, and insurance risk assess- ment. Numerical weather prediction (NWP) [4] models have been long served as the backbone of operational forecasts and research, predicting weather through explicit physical equations. However, the practical deployment of NWP suffers from incomplete compre- hension for certain atmospheric physical processes, alongside high computational costs and model spin-up [30]. In recent years, AI-based weather forecasting models [2,3,6,7, 21,24] have emerged as promising alternatives or complements to NWP. A growing body of literature demonstrates that deep learning architectures can outperform traditional NWP systems in predicting TC trajectories. Although these AI models demonstrate substantial reductions in track forecast errors at extended lead times, highlight- ing their potential for operational use, their performance within the 72-hour window still lags behind that of NWP [31]. Furthermore, a fundamental, yet often overlooked constraint remains inherent to nearly all TC forecasting systems that rely on AI-based global weather models: they are trained exclusively on coarse-resolution reanalysis data, most commonly ERA5 at 0.25 ◦ . As a result, the pre- dicted TC positions are inherently constrained to this fixed grid, limiting the spatial fidelity of forecasts and introducing discretization errors that can be significant relative to the scale of TC motion. Despite the emergence of specialized models [14,31], specifically designed for extreme events, they still underperform compared to HRES within the critical 72-hour forecast window. To address this resolution bottleneck, we propose a novel probabilis- tic correction framework that explicitly models the continuous spatial distribution of likely TC center locations. By computing the expectation of this distribution, we recover TC position estimates at arbitrary spatial precision effectively. Beyond track prediction, accurate forecasting of TC intensity, typically quantified by maximum sustained wind speed, remains a persistent challenge for both numerical and data-driven models. While recent AI-based systems have shown promise in capturing large-scale steering flows, they consistently underperform in pre- dicting rapid intensification or decay events [14]. This limitation stems from two interrelated factors. First, the widely used ERA5 reanalysis at 0.25 ◦ resolution provides only grid-averaged meteorological fields, which inherently smooth out the sharp gradients and localized extremes characteristic of TC cores. Second, most deep learning models for TC forecasting are trained with regression losses such as mean squared error (MSE) or mean absolute error (MAE), which inherently bias predic- tions toward the conditional mean of the target distribution. To overcome these limitations, we introduce a region-aware inten- sity correction algorithm that operates on dynamically defined sub-grid zones surrounding the TC center. Within each region, high-resolution proxies derived from the model’s internal represen- tations are aggregated and calibrated against historical intensity statistics. To enable holistic TC forecasting, we further couple the prob- abilistic center refinement module with the region-aware intensity correction module into a unified system. This inte- grated architecture ensures coherence between predicted locations and local structures while streamlining the handling of complex real-world TCs. Consequently, BaguanCyclone effectively models both standard sweeping-arc tracks and intricate scenarios, such as re-intensification, multi-TC interactions, and meandering, that typ- ically defeat conventional numerical methods. We rigorously eval- uate our system on the IBTrACS (International Best Track Archive for Climate Stewardship) dataset, the global standard for histor- ical TC records. Experiments across six major TC-prone basins, including the Western North Pacific (WP), Eastern North Pacific (EP), North Atlantic (NA), North Indian Ocean (NI), South Indian Ocean (SI) and South Pacific (SP) (Fig. 1), demonstrate consistent and significant improvements over both operational NWP models and AI-based baselines. Note that we skip the South Atlantic (SA) basin from our analysis due to its insufficient number of tropical cyclone events. 180°120°W60°W0°60°E120°E 60°S 30°S 0° 30°N 60°N Western North Pacific Eastern North Pacific Southern Pacific North Indian South Indian North Atlantic South Atlantic Tropical depression (W < 34 knots) Tropical storm (34 <= W < 64 knots) Category 1 (64 <= W < 83 knots) Category 2 (83 <= W < 96 knots) Category 3 (96 <= W < 113 knots) Category 4 (113 <= W < 137 knots) Category 5 (W >= 137 knots) Figure 1: Tropical cyclone intensities and tracks. IBTrACS data (2000–2024) colored by Saffir-Simpson category. More importantly, the system has been operationally deployed at Zhejiang Meteorological Observatory, China Meteorological Ad- ministration (CMA), where it has been continuously running for over a year in a real-time forecasting environment. Throughout this period, it has consistently outperformed all other AI-based forecasting systems in use, delivering more accurate and reliable TC predictions that directly support emergency decision-making, public warnings, and maritime safety operations. This successful deployment represents a significant milestone in the operational integration of deep learning into high-stakes meteorological appli- cations, demonstrating that AI models, when carefully designed with physical awareness and sub-grid-scale refinement, can meet the rigorous demands of real-world weather services. The main contributions of this work are as follows: (1)Our analysis reveals fundamental data-centric limitations that lead to the degraded performance of existing AI-based global weather models during intensity forecasting. Enhancing AI-Based Tropical Cyclone Track and Intensity Forecasting via Systematic Bias CorrectionConference acronym ’X, June 03–05, 2018, Woodstock, NY (2)We introduce BaguanCyclone, a framework designed to re- fine both track and intensity forecasting. It significantly outperforms existing AI-based models, achieving a 16% reduction in tracking error and a 34% improvement in intensity precision, respectively across all 6 regions in comparison with the vanilla model based on global weather foundation model. (3)By resolving the non-linear feedbacks between internal TC processes and the environment, BaguanCyclone maintains high predictive skill across complex synoptic regimes, such as re-intensification, multi-TC, and meandering tracks, that often elude traditional numerical TC prediction. (4)BaguanCyclone have been operationally deployed at the Zhejiang Meteorological Observatory, CMA during the 2025 WP typhoon season, achieving average track and intensity errors of 85.61 km and 6.20 m/s. During Typhoon Co-May in July, 2025, forecasts have supported the evacuation of 97,000 people from vulnerable coastal and low-lying regions in Zhejiang. 2 Related Work 2.1 Challenges in Typhoon Prediction Tropical cyclone is a rapidly rotating system characterized by a low- pressure center, closed low-level atmospheric circulation, strong winds, and a spiral arrangement of thunderstorms that generate torrential rain and squalls. While operational TC forecasting has seen steady progress in recent years, significant challenges remain. First, the prediction of rapid intensification (RI) event—defined as an increase in TC intensity of at least 30 knots (approximately 15.4 m/s) within a 24-hour period [16]—has proven especially diffi- cult [5,10,17]. This difficulty primarily stems from a limited un- derstanding of the physical mechanisms underlying such extreme events [16]. Second, as a TC dissipates into a tropical depression (maximum sustained wind speed < 17.1 m/s or 34 knots, according to diffirent standards [19]), its structural integrity tends to become loose and highly asymmetric, leading to significantly inflated errors in both intensity estimation and center localization. Notably, the rare instances where a system undergoes re-intensification back into a stronger system—an "intensity rebound"—frequently result in major forecast discrepancies [8,11]. Additionally, the complexity of multi-TCs environments poses a significant predictive hurdle; even in the absence of direct binary interactions (Fujiwhara effect [15]), the competitive depletion of ambient moisture and energy between coexisting systems can severely degrade model fidelity. Finally, the synergistic evolution between mid-latitude westly troughs and large-scale circulation patterns often induces erratic trajectories, further amplifying the prediction model inherent non-linearity and uncertainty [1,8,23]. In summary, achieving the precise prediction of typhoon intensity evolution and movement trajectories remains a persistent and cutting-edge challenge in atmospheric science. 2.2 Deep Learning for Tropical Cyclones Global weather foundation models (GFMs), such as FourCastNet [20], Pangu-Weather [2], FuXi [7], GraphCast [21], and Baguan [24], have revolutionized TC track forecasting, with many now con- sistently surpassing the operational HRES benchmark. However, intensity prediction remains a persistent “Achilles’ heel” for these systems. Recent benchmarking [9,14] reveals that GFMs system- atically underpredict peak wind speeds, often performing worse than rudimentary statistical baselines. This performance gap stems from the difficulty of capturing high-resolution structural nuances and physical consistency within global-scale architectures. Concur- rently, specialized deep learning models—such as ConvLSTM [26], VQLTI [29], and Deep-Hurricane-Tracker [18], have been developed to address TC structural evolution by incorporating physical con- straints or multi-source data. Yet, a critical limitation persists: these specialized methods typically function as standalone systems, fail- ing to leverage the robust, large-scale atmospheric representations inherent in GFMs. To bridge this divide, We propose a Baguan-based refinement framework that synergizes global forecasting power with targeted optimizations for superior track and intensity fidelity. 3 Methodology The architecture of BaguanCyclone, as shown in Fig. 2, integrates capabilities for TC track prediction and intensity forecasting. The Probabilistic center refinement model is designed to predict TC latitude and longitude coordinates with arbitrary precision by modeling location as a probability distribution map, using the initial position and a continuous sequence of atmospheric states of 0.25 ◦ generated by an AI model as input. The region-aware intensity forecasting model partitions the atmospheric states into multiple regions and then predicts the maximum wind speed for each region. During inference, these two modules are coupled, with the predicted latitude and longitude localized within the intensity forecasting regions to retrieve the corresponding intensity information. 3.1 Probabilistic Center Refinement Model Our methodology transitions from a deterministic tracking paradigm to a probabilistic density learning framework, effec- tively bypassing the spatial quantization (grid-locking) inherent in the 0.25 ◦ reanalysis data. This process is performed in three stages: kinematic initialization, density mapping, and neural correction. By optimizing a softened spatial distribution via KL-divergence and decoding it through an expectation operator, the model achieves sub-grid localization. This allows for trajectory resolution with arbi- trary precision, transcending the native resolution of the underlying atmospheric grid. 3.1.1 Kinematic Initialization. To establish a robust physical an- chor, we first generate a deterministic trajectory prior by integrat- ing a multi-step procedural tracker [3] with steering flow vectors from ECMWF [28]. An initial position estimate is obtained via the kinematic extrapolation of the historical trajectory, weighted by the local steering flow. This candidate position is then iteratively refined within a bounding box, which progressively narrows to con- verge upon the nearest local pressure minimum. While this provides a physically consistent estimate( ˆ 푥 푐 , ˆ 푦 푐 ), it remains constrained by the discrete grid resolution of the input fields. 3.1.2 Probabilistic Density Mapping. To represent spatial uncer- tainty and enable sub-grid refinement, the discrete coordinates are transformed into a continuous probability density field. Given a region of size퐻 × 푊, we apply a truncated Gaussian kernel to Conference acronym ’X, June 03–05, 2018, Woodstock, NYTrovato et al. (b) Probabilistic Center Refinement Model(a) Input Forecast[t1] Initial Position Forecast[t0] (c) Region-aware Intensity Forecasting Model Window Partition of Forecast[t1] Int. 1 Int. 2Int. 3 Int. 4 Int. 5Int. 6 Int. 7 Int. 8Int. 9 Kinematic Initialization Extrapolation steering flow argmin MSLP Shift Window MSA Patch Merging Multi - head Self - Attention × N BaguanModel Layer Norm Feed - Forward Network Forecast[t1] Probabilistic Mapping Position Embedding Module Feature Embedding Module 푳풐풏=% 풋 풚 풋 % 풊 푷(풙 # ,풚 풋 ) Region Selection × N Corrected Position Intensity Layer NormMulti - head Self - Attention Feed - Forward Network 푳풂풕 = % 풊 풙 # % 풋 푷 ( 풙 # , 풚 풋 ) True Position Tracked Position error error improvement Figure 2: Overview of BaguanCyclone’s architecture. (a) The input consists of an initial position and 2 subsequent forecast frames. (b) The probabilistic center refinement model first employs a nonparametric algorithm to obtain an initial tracking distribution, which is then refined using a tracking correction model. (c) Region-aware intensity forecasting model leverages a Swin-transformer with patch merging to compute the per-region intensity. map the cyclone center into a softened spatial representation. The local weights ̃ 푔 푖푗 are computed based on the Euclidean distance 푑 푖푗 between the grid point(푖, 푗)and the center, subject to a fixed truncation radius 푟 : ̃ 푔 푖푗 = ( exp(− 푑 2 푖푗 2휎 2 ),if 푑 푖푗 ≤ 푟, 0,otherwise. (1) To ensure that the density field represents a valid probability mass, the weights are normalized: 푤 푖푗 = ̃ 푔 푖푗 Í 푘 Í 푙 ̃ 푔 푘푙 .(2) This probabilistic encoding transforms a single point into a field of spatial influence, allowing the deep learning model to perceive the center’s location relative to its surrounding atmospheric environ- ment. 3.1.3 Neural Correction and Reconstruction. The tracking correc- tion model ingests this probability field as an auxiliary spatial chan- nel, concatenated with the multi-level atmospheric state features. This allows the model to learn the non-linear relationship between local meteorological dynamics and systematic tracking biases. The model outputs a predicted density field ˆ 푊 , which is normalized via a softmax operation to maintain the property Í 푖푗 ˆ 푊 푖푗 =1. The optimization is guided by the KL divergence between the predicted distribution ˆ 푊 and the ground-truth distribution푊 푔푡 : L 퐾퐿 = 퐷 퐾퐿 (푊 푔푡 || ˆ 푊).(3) Finally, to recover high-precision coordinates, we apply an expect- ation-based reconstruction. The final latitude and longitude are calculated by computing the expected values along the spatial axes of the density field. By leveraging the continuous nature of the predicted probability mass, this mechanism produces coordinates that are no longer confined to the 0.25 ◦ grid intervals, thereby achieving sub-grid localization and superior tracking fidelity. 3.2 Why AI Models Failed in Intensity Forecasting: A Data-centric Perspective The majority of AI models in the literature report tracking accu- racy only [2,24], omitting an assessment of intensity forecasting performance, and there is growing evidence [14] that such models exhibit significant limitations in predicting storm intensity. Table 1: Comparison of ERA5 (0.25 ◦ ) and EC Analysis (0.1 ◦ ) over the Western North Pacific (WP) and North Atlantic (NA). Results are presented only for 6-hour. dist. (km)wind (m/s) ERA5 (0.25 ◦ )EC Analysis (0.1 ◦ )ERA5 (0.25 ◦ )EC Analysis (0.1 ◦ ) WP45.356.97.745.69 NA32.748.111.538.87 Before introducing our intensity forecasting model, we first ex- amine, from a data-centric perspective, the reasons underlying the poor performance of existing AI models in tropical cyclone track forecasting. We utilize the widely used ground truth data in the training of many AI models, the 0.25 ◦ ERA5 reanalysis and the 0.1 ◦ EC analysis. As shown in Tab. 1, using the nonparametric track- ing algorithm, we find that differences in spatial resolution lead to intensity forecasts that diverge by approximately 20%, whereas their impact on track forecasts is considerably smaller. Owing to its grid-averaged nature, ERA5 data is inherently limited in its ability to resolve extreme values—such as the peak intensity of tropical cyclones. As a result, AI models trained on ERA5 consistently under- perform in intensity forecasting relative to NWP systems initialized with higher-resolution 0.1° analyses. Enhancing AI-Based Tropical Cyclone Track and Intensity Forecasting via Systematic Bias CorrectionConference acronym ’X, June 03–05, 2018, Woodstock, NY 3.3 Region-aware Intensity Forecasting Model Additionally, AI-based models are typically trained with MSE loss, which inherently biases predictions toward the conditional mean, further leading to suboptimal performance on extreme-value tasks. Table 2: Minimum inter-TC distance ( ◦ ) within each region from 2000 to 2023. A value of –1 indicates no twin TCs oc- curred. BasinWPEPNANISISP Dist.3.37 ◦ 1.20 ◦ -1 16.37 ◦ 4.70 ◦ 5.92 ◦ To this end, as shown in Fig. 2 (c), we propose the region-aware algorithm based on patch merging, a mechanism that naturally partitions the input into multiple regions. To better determine the region partitioning size, we conducted a statistical analysis of the minimum inter-TC distance across all regions, as summarized in Tab. 2. Each region is therefore designed to have a size smaller than 1/3 of this minimum distance. Specifically, the input size of 푉×퐻×푊is transformed into size of 퐻 2 푁 ·푝 × 푊 2 푁 ·푝 , where푝denotes the patch size used in the patch embedding layer,푁is the number of patch merging blocks and the first dimension. 3.4 Joint Deployment of Tropical Cyclone Forecasting Models at Inference In the inference phase, our framework employs a dual-stream architecture that jointly deploys the probabilistic center refine- ment model alongside the region-aware intensity forecasting model. Upon receiving the input atmospheric states, the system processes the data through both channels concurrently: the center refinement module generates the predicted coordinates of the cyclone’s eye, while the intensity forecasting module produces a comprehensive map of predicted maximum intensities across various spatial parti- tions. To synthesize these outputs, the predicted center coordinates are utilized as a spatial query to index into the partitioned inten- sity map. By identifying the specific region that encompasses the refined center, the corresponding intensity value is extracted and designated as the final, authoritative intensity prediction for the cyclone. 4 Experiments 4.1 Settings We utilize IBTrACS [19] as the ground-truth best-track reposi- tory. Training features are derived from ERA5 reanalysis, while Baguan forecasts [24] are employed during evaluation to simu- late operational conditions. Input features comprise atmospheric variables (Z, T, U, V, W, Q) at three pressure levels (500, 700, 850 hPa), supplemented by surface-level parameters including 2m tem- perature/dewpoint, 10m wind components, mean sea level (MSL), surface pressure, and total column water/vapor. Due to limited 2024 samples in the NI and SP basins, the training and evaluation periods are set to 2000–2022 and 2023–2024, respectively. For all other regions, the model is trained on 2000–2023 and evaluated on 2024. The tracking performance is measured by the great-circle distance (km) between forecast and IBTrACS coordinates using the Haversine formula. Intensity is evaluated via the mean absolute error (m/s) of the maximum wind speed. Detailed descriptions of the datasets and evaluation metrics are provided in App. A.1. 4.2 Main Results 4.2.1 Baselines. We employ the most recent AI-based model, Ba- guan, to generate atmospheric states for evaluation. To enable a meaningful and fair comparison, we select both NWP and AI-based methods. For the NWP method, we select the leading ECMWF HRES. In terms of AI-based methods, we select those methods based on global weather foundation models due to their outstanding performance, including Vanilla Baguan (without the refinement module), Pangu-Weather and GraphCast. We evaluate perfor- mance across six major global tropical cyclone basins (WP, EP, NA, NI, SI and SP). Furthermore, to contextualize the performance of BaguanCyclone relative to reanalysis quality, we also visualize the intrinsic error of the ERA5 dataset at a 6-hour lead time. The SA basin was not included in this study due to the scarcity of historical samples required for robust model training. Since TC forecasts within a 3-day lead time are of greater practi- cal significance, we focus our comparison on this period. Further- more, the HRES only provides forecasts when the initial intensity exceeds a specific threshold [25], leading to insufficiant data samples in certain regions. Thus, for the WP and NA regions, the evaluation is based on data samples where HRES forecasts are successfully matched with IBTrACS observations. In other regions, the compar- ison is limited to AI models and excludes HRES, as these models provide comprehensive coverage for all cases documented in the IBTrACS. 4.2.2 Tracking. As illustrated in Fig. 3 (a), BaguanCyclone deliv- ers superior tracking performance across all six regions. Notably, BaguanCyclone consistently surpasses the HRES benchmark across the majority of forecast lead times. This achievement is particularly significant as it maintains a competitive edge despite the inherent constraints of a relatively coarse 0.25 ◦ input resolution, underscor- ing the model’s exceptional architectural efficiency. By integrating our proposed methodology, BaguanCyclone yields an average im- provement of 16% over all AI-based baselines. Specifically, within the highly active WP basin, the model demonstrates significant gains over Vanilla Baguan, achieving a 16.14% error reduction in 24-hour short-range forecasts and an 11.3% improvement at the 72- hour long-range horizon. Performance gains are even more promi- nent in the NI basin, where BaguanCyclone achieves a substantial reduction in tracking error, exceeding 30%, relative to AI-based baselines. 4.2.3Intensity Prediction. As illustrated in Fig. 3 (b), conventional AI-based baselines (including Vanilla Baguan) exhibit a persistent deficiency in intensity forecasting. This limitation, well-documented in existing literature [14,31], is further exacerbated by the inherent intensity biases within the ERA5 reanalysis dataset, as visualized in the Fig. 3 (b). While such data-quality bottlenecks typically con- strain the performance of data-driven frameworks, BaguanCyclone overcomes these limitations, delivering substantial gains that sur- pass the HRES benchmark across most lead times. This achievement Conference acronym ’X, June 03–05, 2018, Woodstock, NYTrovato et al. 50 100 150 200 Sample count 50 100 150 200 Sample count 50 100 150 200 Track Error (km) (a) WPNANISPEPSI 61830425466 Lead Time (hour) 2 4 6 8 10 Wind Error (m/s) (b) 61830425466 Lead Time (hour) 61830425466 Lead Time (hour) 61830425466 Lead Time (hour) 61830425466 Lead Time (hour) 61830425466 Lead Time (hour) OursHRESVanilla BaguanPanguGraphcastERA5 Figure 3: Bias in track and intensity predictions. Solid lines correspond to forecast models:BaguanCyclone (ours), HRES (ECMWF), Vanilla Baguan, Pangu-Weather, and Graphcast. As ERA5 (the horizontal dashed lines) represents reanalysis data rather than predictive forecasts, its 6-hour intrinsic error is depicted as a constant baseline. The gray histograms indicate the sample count of the validation set. effectively bridges the gap between reanalysis-trained research models and stringent operational standards. Notably, in the WP basin, our model consistently outperforms HRES across all lead times by an average margin of 12.23%. Compared with AI-based baselines, BaguanCyclone demonstrates a substantial performance leap, yielding a 34% average reduction in forecast error. Notably, this improvement reaches a remarkable 50% in the NI basin, demonstrat- ing BaguanCyclone’s profound utility for real-world meteorological applications. 4.3 Generalization Capabilities across Various Tropical Cyclone Varieties Tropical cyclones (TCs) exhibit diverse track and intensity evolu- tions controlled by different mechanisms, including steering flow, land–sea energy-budget transitions, and nonlinear interactions be- tween TC and background fields. Rather than treating these as isolated “special cases”, we use representative events to test the following science-oriented hypothesis: Hypothesis (AI4Science). BaguanCyclone generalizes across TC varieties by correcting two key failure modes of coarse-resolution AI forecasts: (i) discretization-induced center errors and (i) core-intensity underestimation, thereby remaining consistent with the governing drivers (steering flow, surface energy supply, and synoptic competition) even when the dominant regime shifts. Across the cases below, BaguanCyclone’s gains are not only numerical; they reflect mechanism-aware robustness under three recurring drivers: (1) steering-flow guidance, (2) land–sea regime transitions, and (3) nonlinear competition under weak or interacting systems. 4.3.1Steering-flow guidance under sweeping-arc tracks (Beryl). Many TCs follow sweeping-arc tracks (including recurvature), primar- ily controlled by large-scale steering flow from subtropical highs and midlatitude troughs. In this regime, small center errors on coarse grids can amplify rapidly, especially near turning points. We examine TC Beryl (Fig. 4), rapid intensification (RI) event, weak- ening, land interaction, and post-landfall evolution. Incorporat- ing physical guidance from meteorological fields into the learning of nonlinear rules, BaguanCyclone provides more accurate track and intensity forecasts over most of the 72-hour window initial- ized at 18:00 UTC on 29 June, indicating improved robustness for intensification-related evolution. Beryl’s track was strongly influenced by the Azores (Bermuda) High [13]. BaguanCyclone captures this control by maintaining multi-level dynamical consistency across 500/700/850 hPa and sur- face winds, stabilizing the inferred steering signal through recur- vature and landfall. It also responds to land transition via storm- relative humidity and low-level wind gradients consistent with increased surface friction and associated kinetic energy dissipation. Figure 4: Forecast evaluation for tropical cyclone Beryl: (a) track error, (b) max wind error, and (c) observation vs. pre- diction. Black line and gray shading indicate landfall and re-emergence; dashed line denotes mean error. 4.3.2Regime shifts in intensity: landfall weakening and re-intensification (Pulasan). Intensity forecasting is particularly difficult for AI models Enhancing AI-Based Tropical Cyclone Track and Intensity Forecasting via Systematic Bias CorrectionConference acronym ’X, June 03–05, 2018, Woodstock, NY trained on coarse reanalysis due to smoothed core extremes at 0.25 ◦ and regression-to-mean effects. We evaluate TC Pulasan (Fig. 5), which weakened after landfall and re-intensified after returning to the ocean. Landfall largely cuts off oceanic heat/moisture fluxes and in- creases surface friction, often yielding a negative energy budget and rapid decay. Pulasan, however, remained over land briefly and retained structural coherence, enabling re-intensification. Baguan- Cyclone anticipates the rebound over the Yellow Sea under sup- portive thermodynamic conditions (warm underlying and abun- dant moisture), suggesting improved generalization across land–sea transitions by anchoring intensity estimation to the storm-centered local environment rather than coarse-grid smoothing. Figure 5: Forecast evaluation for tropical cyclone Pulasan, following the same format as Fig. 4. 4.3.3 Nonlinear interactions: multiple cyclones and rapid steering reversals (Trami and Kong-rey). Multiple or closely spaced cyclones require maintaining cyclone identity while disentangling storm- relative dynamics under shared environmental flow. We analyze the concurrent Trami and Kong-rey event (Fig. 6). BaguanCyclone decouples their steering environments from multi-level fields, yield- ing a more accurate northwestward track for Kong-rey under the subtropical high. It also better captures peak intensity over the open ocean, with lower errors near maximum winds. This case further includes Trami’s rare 180 ◦ post-landfall rever- sal over the Indochina Peninsula. Such U-turns are typically tied to rapid adjustments of the midlatitude trough and subtropical high that flip the steering flow. BaguanCyclone reproduces the rever- sal, indicating sensitivity to fast synoptic transitions rather than regression-to-mean behavior. 4.3.4 Weak-steering environments: meandering and stalling trajec- tories (Shanshan). When steering flow is weak, TC motion becomes highly nonlinear and sensitive to subtle circulation gradients, often leading to meandering and stalling tracks. We evaluate TC Shan- shan (Fig. 7), which exhibited weak-steering meandering and later abrupt deflections. BaguanCyclone captures the slow, erratic motion by responding to circulation gradients that reflect the competition between the subtropical high and midlatitude westerly troughs. After landfall, it remains robust during dissipation, reproducing the abrupt, nearly 90 ◦ deflection in the terminal stage, consistent with interactions be- tween the decaying vortex and an approaching midlatitude trough. Figure 6: Forecast evaluation for tropical cyclone Trami and Kong-rey, following the same format as Fig. 4. Summary (AI4Science). Across sweeping arcs, re-intensification, multiple cyclones, and meandering events, BaguanCyclone remains skillful as the dominant driver shifts: (i) steering flow control, (i) surface energy-budget transitions, and (i) nonlinear competition under weak or interacting systems. These cases support systematic bias correction as a mechanism-aware refinement layer that yields more physically coherent TC track and intensity evolutions from coarse-resolution AI forecasts. Figure 7: Forecast evaluation for tropical cyclone Shanshan, following the same format as Fig. 4. 4.4 Ablations 4.4.1Probabilistic Forecasting vs. Residual Forecasting. To validate the design choice of the probabilistic model for TC center locations, we compare our probabilistic formulation against a deterministic variant that directly forecasts the residual of latitude and longitude Conference acronym ’X, June 03–05, 2018, Woodstock, NYTrovato et al. coordinates using the same backbone architecture. Both models are trained under identical conditions and evaluated on the same test set. As shown in Tab. 3, the probabilistic approach mostly outper- forms the direct regression baseline in all evaluation metrics, par- ticularly in terms of localization accuracy and robustness to trajec- tory outliers. More importantly, the probabilistic formulation pro- vides calibrated uncertainty estimates, enabling risk-aware decision- making. For example, high predictive entropy correlates strongly with challenging forecasting scenarios such as rapid direction shifts or landfall events. In contrast, the deterministic model offers no measure of confidence, and tends to produce overconfident yet inaccurate predictions under distributional shift. These results un- derscore the value of explicit uncertainty modeling in geospatial forecasting tasks where prediction reliability is as critical as point accuracy. Table 3: A comparative analysis of probabilistic and residual forecasting in the tracking refinement module in WP and NA. The error is measured in km. Lower value is better. LeadWPNA TimeVanillaProb.Res.VanillaProb.Res. 24h60.9 53.060.753.8 48.649.8 48h98.8 91.5101.192.888.5 87.6 72h157.1 147.9158.4112.8 109.6111.4 4.4.2Coupling Tracking and Intensity Module. To evaluate the ef- fectiveness of the coupled architecture, we conducted a comparative analysis by decoupling the track and intensity modules. Specifi- cally, we used both the vanilla Baguan predicted tracks and the corrected tracks as inputs to the intensity module for post-hoc in- tensity estimation. These results serve as baselines for quantifying the performance gains achieved by our coupled modeling approach. As shown in Tab. 4, our experiments in the SI region demonstrate that the coupled architecture yields significant improvements in the vast majority of cases. This improvement can be attributed to the fact that the refined tracks provide more precise locations of the TCs within the spatial window. Consequently, this reduces intensity forecast biases that typically stem from track displacement errors, ensuring the intensity module extracts features from the correct atmospheric environment Table 4: Performance comparison of coupled vs. separate SI models. Values indicate % improvement; positive values favor the coupled model. Lead12h24h36h48h60h72h wind (m/s)+0.89+1.50+2.40+2.88+3.74+0.15 pres (hpa)+1.96+0.82+1.87+0.81+1.81-0.55 4.5 Real-world Deployment In July 2025, BaguanCyclone is officially operationalized at the Zhejiang Meteorological Observatory, CMA. By delivering high- precision trajectory predictions for all WP typhoons, the system provided critical decision support for regional disaster mitigation. In July 2025, the operational value of BaguanCyclone is demonstrated through its accurate predictions of Typhoon Co-May. By providing reliable landfall and intensity data, the model supported the organized evacuation of 97,000 people from vulnerable coastal and low-lying regions in Zhejiang. 122436486072 40 60 80 100 120 140 Track Error (km) 122436486072 6 8 10 12 14 Wind Error (m/s) BaguanCyclonePanguFengwuFuxiGraphCastAIFS Figure 8: Average real-time forecasting performance during the 2025 WP typhoon season. Note that most baseline models (e.g., FuXi and Pangu-Weather) are not the latest versions. Fig. 8 presents the operational performance of BaguanCyclone alongside several AI-based models, including FuXi [7], Pangu-Wea- ther [2], GraphCast [21], FengWu [6], and AIFS [22]. While these models are operational at the CMA, it should be noted that their performance may reflect specific deployed versions rather than the latest iterations. Among these, BaguanCyclone demonstrates clear superiority across both key metrics. It leads the majority of models in track forecasting with a mean error of 85.61 km. Most no- tably, its performance in intensity prediction represents a landmark achievement—its average error of 6.20 m/s constitutes an unprece- dented 50% reduction compared to the second-best performer. This paradigm-shifting precision effectively overcomes long-standing bottlenecks in intensity forecasting, offering a level of reliability previously considered unattainable in operational settings. 5 Conclusion In this study, we proposed a novel AI-based modeling framework specifically designed for TC’s track and intensity forecasting. Our models have achieved substantial performance gains over existing AI-based weather forecasting models. Beyond theoretical valida- tion, the proposed system has been successfully deployed at the Zhejiang Meteorological Observatory, CMA, for operational use. Over a continuous one-year evaluation period, the system demon- strated exceptional stability and predictive accuracy for typhoon events in the WP basin, significantly outperforming baseline AI models in real-world scenarios. This work bridges the gap between experimental AI research and practical meteorological operations. 6 Limitations and Ethical Considerations This study uses only gridded meteorological/geophysical datasets, with no personal data or human subjects; all data licenses are fol- lowed. There are two limitations: the need for higher spatial reso- lution (0.1 ◦ ) to resolve finer atmospheric features, and the inherent ‘black-box’ nature of the AI architecture. Future efforts will fo- cus on bridging these gaps by integrating higher-granularity data and exploring the physical drivers behind the model’s decisions to improve both precision and transparency. Enhancing AI-Based Tropical Cyclone Track and Intensity Forecasting via Systematic Bias CorrectionConference acronym ’X, June 03–05, 2018, Woodstock, NY GenAI Disclosure During the preparation of this work the authors used GenAI in order to improve language. After using this tool, the authors re- viewed and edited the content as needed and take full responsibility for the content. References [1] George R Alvey I and Andrew Hazelton. 2022. How do weak, misaligned tropi- cal cyclones evolve toward alignment? A multi-case study using the hurricane analysis and forecast system. Journal of Geophysical Research: Atmospheres 127, 20 (2022), e2022JD037268. [2]Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi Tian. 2023. Accurate medium-range global weather forecasting with 3D neural networks. Nature 619, 7970 (2023), 533–538. [3] Cristian Bodnar, Wessel P. Bruinsma, Ana Lucic, Megan Stanley, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A. Weyn, Haiyu Dong, Anna Vaughan, Jayesh K. Gupta, Kit Thambiratnam, Alex Archibald, Eliza- beth Heider, Max Welling, Richard E. Turner, and Paris Perdikaris. 2024. Au- rora: A Foundation Model of the Atmosphere. ArXiv abs/2405.13063 (2024). https://api.semanticscholar.org/CorpusID:269983273 [4]Roberto Buizza, M Alonso Balmaseda, Andrew Brown, S English, Richard Forbes, Alan Geer, T Haiden, Martin Leutbecher, L Magnusson, Mark Rodwell, et al. 2018. The development and evaluation process followed at ECMWF to upgrade the Integrated Forecasting System (IFS). European Centre for Medium Range Weather Forecasts. [5] John P. Cangialosi, Eric S. Blake, Mark DeMaria, Andrew Penny, Andrew Latto, Edward N. Rappaport, and Vijay Tallapragada. 2020. Recent Progress in Trop- ical Cyclone Intensity Forecasting at the National Hurricane Center. Weather and Forecasting 35, 5 (2020), 1747–1766. doi:10.1175/WAF-D-20-0059.1 Online publication: 17 August 2020; Print publication: 1 October 2020. [6]Kan Chen, Tao Han, Junchao Gong, Lei Bai, Fenghua Ling, Jingyao Luo, Xi Chen, Lei Ma, Tianning Zhang, Rui Su, Yuanzheng Ci, Bin Li, Xiaokang Yang, and Wanli Ouyang. 2023. FengWu: Pushing the Skillful Global Medium-range Weather Forecast beyond 10 Days Lead. ArXiv abs/2304.02948 (2023). https: //api.semanticscholar.org/CorpusID:257985330 [7] Lei Chen, Xiaohui Zhong, Feng Zhang, Yuan Cheng, Yinghui Xu, Yuan Qi, and Hao Li. 2023. FuXi: a cascade machine learning forecasting system for 15-day global weather forecast. npj Climate and Atmospheric Science 6, 1 (2023), 190. doi:10.1038/s41612-023-00512-1 [8]Xiaomin Chen, Jian-Feng Gu, Jun A Zhang, Frank D Marks, Robert F Rogers, and Joseph J Cione. 2021. Boundary layer recovery and precipitation symmetrization preceding rapid intensification of tropical cyclones under shear. Journal of the Atmospheric Sciences 78, 5 (2021), 1523–1544. [9]Mark DeMaria, James L Franklin, Galina Chirokova, Jacob Radford, Robert De- Maria, Kate D Musgrave, and Imme Ebert-Uphoff. 2024. Evaluation of tropical cyclone track and intensity forecasts from Artificial Intelligence Weather Predic- tion (AIWP) models. arXiv preprint arXiv:2409.06735 (2024). [10] Mark DeMaria, James L. Franklin, Matthew J. Onderlinde, and John Kaplan. 2021. Operational Forecasting of Tropical Cyclone Rapid Intensification at the National Hurricane Center. Atmosphere 12 (2021), 683. doi:10.3390/atmos12060683 [11] Difei Deng, Elizabeth A Ritchie-Tyo, Noel E Davidson, and Christopher A Davis. 2025. A new pathway of tropical cyclone–trough interaction as illustrated by the over-land re-intensification of Oswald (2013). Geophysical Research Letters 52, 14 (2025), e2025GL115502. [12]Kerry Emanuel. 2003. Tropical cyclones. Annual review of earth and planetary sciences 31, 1 (2003), 75–104. [13]Robert M. Hordon and Mark Binkley. 2005. Azores (Bermuda) High. Springer Netherlands, Dordrecht, 154–155. doi:10.1007/1-4020-3266-8_27 [14]Cheng Huang, Pan Mu, Jinglin Zhang, Sixian Chan, Shiqi Zhang, Hanting Yan, Shengyong Chen, and Cong Bai. 2025. Benchmark dataset and deep learning method for global tropical cyclone forecasting. Nature Communications 16, 1 (2025), 5923. [15]Kosuke Ito, Soichiro Hirano, Jae-Deok Lee, and Johnny C. L. Chan. 2023. Three- Dimensional Fujiwhara Effect for Binary Tropical Cyclones in the Western North Pacific. Monthly Weather Review 151, 7 (2023), 1779–1795. doi:10.1175/MWR-D- 22-0239.1 [16] John Kaplan and Mark DeMaria. 2003. Large-Scale Characteristics of Rapidly In- tensifying Tropical Cyclones in the North Atlantic Basin. Weather and Forecasting 18, 6 (2003), 1093–1108. doi:10.1175/1520-0434(2003)018<1093:LCORIT>2.0.CO;2 [17]John Kaplan, Mark DeMaria, and John A. Knaff. 2010. A Revised Tropical Cyclone Rapid Intensification Index for the Atlantic and Eastern North Pacific Basins. Weather and Forecasting 25, 1 (2010), 220–241. doi:10.1175/2009WAF2222280.1 [18]Sookyung Kim, Hyojin Kim, Joonseok Lee, Sangwoong Yoon, Samira Ebrahimi Kahou, Karthik Kashinath, and Mr Prabhat. 2019. Deep-hurricane-tracker: Track- ing and forecasting extreme climate events. In 2019 IEEE winter conference on applications of computer vision (WACV). IEEE, 1761–1769. [19]Kenneth R Knapp, Michael C Kruk, David H Levinson, Howard J Diamond, and Charles J Neumann. 2010. The international best track archive for climate stewardship (IBTrACS) unifying tropical cyclone data. Bulletin of the American Meteorological Society 91, 3 (2010), 363–376. [20]Thorsten Kurth, Shashank Subramanian, Peter Harrington, Jaideep Pathak, Morteza Mardani, David Hall, Andrea Miele, Karthik Kashinath, and Anima Anandkumar. 2023. FourCastNet: Accelerating Global High-Resolution Weather Forecasting Using Adaptive Fourier Neural Operators. In Proceedings of the Plat- form for Advanced Scientific Computing Conference, PASC 2023, Davos, Switzerland, June 26-28, 2023, Axel Huebl, Cristina Silvano, and Timothy Robinson (Eds.). ACM, 13:1–13:11. doi:10.1145/3592979.3593412 [21] Remi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Wirnsberger, Meire Fortunato, Ferran Alet, Suman Ravuri, Timo Ewalds, Zach Eaton-Rosen, Weihua Hu, Alexander Merose, Stephan Hoyer, George Holland, Oriol Vinyals, Jacklynn Stott, Alexander Pritzel, Shakir Mohamed, and Peter Battaglia. 2023. Learning skillful medium-range global weather forecasting. Science 382, 6677 (2023), 1416– 1421. arXiv:https://w.science.org/doi/pdf/10.1126/science.adi2336 doi:10.1126/ science.adi2336 [22]Simon Lang, Mihai Alexe, Matthew Chantry, Jesper Dramsch, Florian Pin- ault, Baudouin Raoult, Mariana C. A. Clare, Christian Lessig, Michael Maier- Gerber, Linus Magnusson, Zied Ben Bouallègue, Ana Prieto Nemesio, Peter D. Dueben, Andrew Brown, Florian Pappenberger, and Florence Rabier. 2024. AIFS – ECMWF’s data-driven forecasting system. arXiv:2406.01465 [physics.ao-ph] https://arxiv.org/abs/2406.01465 [23] Leon T Nguyen and John Molinari. 2015. Simulation of the downshear refor- mation of a tropical cyclone. Journal of the Atmospheric Sciences 72, 12 (2015), 4529–4551. [24]Peisong Niu, Ziqing Ma, Tian Zhou, Weiqi Chen, Lefei Shen, Rong Jin, and Liang Sun. 2025. Utilizing strategic pre-training to reduce overfitting: Baguan-a pre-trained weather forecasting model. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 2186–2197. [25]F. Richardson, J. Bidlot, L. Ferranti, A. Ghelli, and M. Miller. 2012. New tropical cyclone products on the web. ECMWF Newsletter 130 (2012), 17–23. [26] Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo. 2015. Convolutional LSTM network: A machine learning approach for precipitation nowcasting. Advances in neural information processing systems 28 (2015). [27] R. W. Sinnott. 1984. Virtues of the Haversine. skytel 68, 2 (Dec. 1984), 158. [28] G. van der Grijn. 2002. Tropical cyclone forecasting at ECMWF: new products and validation. 13 pages. doi:10.21957/c8525o38f [29]Xinyu Wang, Lei Liu, Kang Chen, Tao Han, Bin Li, and Lei Bai. 2025. VQLTI: Long-Term Tropical Cyclone Intensity Forecasting with Physical Constraints. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 28476–28484. [30]James L Warner, Jon Petch, Chris J Short, and Caroline Bain. 2023. Assessing the impact of a NWP warm-start system on model spin-up over tropical Africa. Quarterly Journal of the Royal Meteorological Society 149, 751 (2023), 621–636. [31]Xiaohui Zhong, Lei Chen, Jun Liu, Chensen Lin, Yuan Qi, and Hao Li. 2024. FuXi- Extreme: Improving extreme rainfall and wind forecasts with diffusion model. Science China Earth Sciences 67, 12 (2024), 3696–3708. A Appendix A.1 Settings A.1.1 Datasets. IBTrACS [19] serves as the ground truth, being the most comprehensive global repository of TC best-track data. It integrates recent and historical TC observations from multiple meteorological agencies into a unified, publicly available dataset. The atmospheric state used during training is derived from ERA5 reanalysis data. For evaluation, to better reflect real-world opera- tional conditions, we instead employ atmospheric forecasts from Baguan [24]. Our input features include 5 standard vertical pressure levels (500, 700, and 850 hPa), with key atmospheric variables at each level: geopotential, temperature, u-component of wind, v-component of wind, w-component of wind, and specific humidity. Additionally, we incorporate the following surface-level variables: 2-meter tem- perature (T2M), 2-meter dewpoint temperature (D2M), 10-meter wind components (U10, V10, W10), mean sea-level pressure (MSL), surface pressure (SP), total column water (TCW), and total column Conference acronym ’X, June 03–05, 2018, Woodstock, NYTrovato et al. water vapor (TCVW). For NI and SP, the model is trained on data from 2000–2022 and evaluated on 2023–2024. For other regions, the training period spans 2000–2023, with evaluation conducted on 2024 data. A.1.2 Metrics. For track evaluation, we compute the great-circle distance (in kilometers) between forecasts and observations using the Haversine formula [27]. Given two geographic coordinates- (휙 1 ,휆 1 )for the forecast location and(휙 2 ,휆 2 )for the best-track ob- servation, where휙denotes latitude and휆denotes longitude (both expressed in radians)—the central angle푐between the points is calculated as: 푎= sin 2 ( Δ휙 2 )+ cos휙 1 cos휙 2 sin 2 ( Δ휆 2 )(4) 푐= 2arctan2( √ 푎, √ 1−푎).(5) whereΔ휙= 휙 2 −휙 1 andΔ휆= 휆 2 − 휆 1 . The corresponding surface distance 푑 is then obtained as: 푑= 푅·푐,(6) with푅=6371kmrepresenting the mean radius of the Earth. All input coordinates are first converted from degrees to radians prior to computation. For intensity evaluation, we use absolute error directly in m/s for maximum wind speed. Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009