Paper deep dive
ReasonCast: Agentic Demand Forecasting with Selective Semantic Reasoning
Ziyue Yang, Chaolin Xu, Yijing Wang, Tiankai Gu, Hui Yang, Yanhong Lin, Kaiyuan Liu, Fei Xiao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/18/2026, 6:06:09 AM
Summary
ReasonCast is an agentic demand forecasting framework that combines time-series foundation models (TSFMs) with large language models (LLMs) via selective semantic intervention. It uses an agent to route instances to Skip, Basic, or Tool reasoning modes based on context and uncertainty. Structured semantic fields (relevance, direction, shape, amplitude, peak) are generated and fused with the TSFM through additive and multiplicative paths, allowing for exact recovery of the no-text forecast when intervention is unnecessary. The model is trained using a curriculum of Schema SFT, semantic-field RL, and forecast-utility RL to align reasoning with marginal forecast improvement.
Entities (13)
Relation Signals (11)
ReasonCast ā supportsrouting ā Basic
confidence 95% Ā· The agent selects Skip, low-cost Basic semantics, or tool-augmented Tool semantics
ReasonCast ā supportsrouting ā Skip
confidence 95% Ā· The agent selects Skip, low-cost Basic semantics, or tool-augmented Tool semantics
ReasonCast ā supportsrouting ā Tool
confidence 95% Ā· The agent selects Skip, low-cost Basic semantics, or tool-augmented Tool semantics
ReasonCast ā uses ā Chronos-2
confidence 95% Ā· A pretrained Chronos-2 model fĪø provides the no-text forecast
ReasonCast ā uses ā Qwen3-32b
confidence 95% Ā· For either non-skip action, Qwen3-32B produces structured semantic tokens
ReasonCast ā achievesimprovementon ā M5
confidence 90% Ā· ReasonCast lowers WMAPE by ... 0.47 percentage points on ... M5 event windows
ReasonCast ā implements ā Additive Path
confidence 90% Ā· An additive path corrects local trends and temporal shapes
ReasonCast ā implements ā Multiplicative Path
confidence 90% Ā· a multiplicative path captures event-driven level shifts
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Demand forecasting increasingly requires combining two complementary sources of information: historical sales reveal recurring numerical dynamics, while future promotions, holidays, price changes, and platform interventions provide forward-looking knowledge. Existing text-enhanced forecasting methods often encode such context into generic representations and fuse it uniformly with time-series features, without explicitly distinguishing which semantic effects are forecast-relevant or how they should modify future dynamics. We introduce ReasonCast, a structured semantic intervention framework that translates event knowledge into forecast-specific operations. An agent examines the event context, the no-text forecast, and its uncertainty to determine whether textual reasoning is needed. Rather than injecting free-form text, ReasonCast represents event knowledge through structured fields describing event relevance, demand direction, temporal shape, amplitude, and peak intensity. These fields interact selectively with temporal components of a time-series foundation model. An additive path corrects local trends and temporal shapes, while a multiplicative path captures event-driven level shifts. ReasonCast introduces a forecast-grounded post-training curriculum. Schema SFT establishes semantic fields; semantic-field RL calibrates direction, shape, amplitude, and peak judgments; and forecast-utility RL evaluates semantic interventions through a frozen forecaster, aligning reasoning outputs with marginal forecast improvement. ReasonCast lowers WMAPE by 3.29, 1.25, and 0.47 percentage points on holiday-sensitive categories, mega-sale-sensitive categories, and M5 event windows, respectively. On stable-sales periods, indiscriminate semantic intervention increases WMAPE by 1.68 percentage points, whereas suppressing unnecessary intervention preserves the numerical backbone.
Tags
Links
- Source: https://arxiv.org/abs/2608.15291v1
- Canonical: https://arxiv.org/abs/2608.15291v1
Trouble viewing inline? Open PDF directly ā
Full Text
75,188 characters extracted from source content.
Expand or collapse full text
ReasonCast: Agentic Demand Forecasting with Selective Semantic ReasoningConference: Preprint; August 2026; Ziyue Yang Note: Ziyue Yang and Chaolin Xu contributed equally to this work. Affiliation: Taobao and Tmall Group, Alibaba Group , China , Chaolin Xu Affiliation: Taobao and Tmall Group, Alibaba Group , China , Yijing Wang Affiliation: Taobao and Tmall Group, Alibaba Group , China , Tiankai Gu Affiliation: Taobao and Tmall Group, Alibaba Group , China , Hui Yang Affiliation: Taobao and Tmall Group, Alibaba Group , China , Yanhong Lin Affiliation: Taobao and Tmall Group, Alibaba Group , China , Kaiyuan Liu Affiliation: Taobao and Tmall Group, Alibaba Group , China and Fei Xiao Note: Corresponding author. Affiliation: Taobao and Tmall Group, Alibaba Group , China 2026Ā© none; Abstract. Demand forecasting increasingly requires combining two complementary sources of information: historical sales reveal recurring numerical dynamics, while future promotions, holidays, price changes, and platform interventions provide forward-looking knowledge that may not be recoverable from past observations alone. Time-series foundation models are strong numerical forecasters, and large language models can interpret heterogeneous event context. However, existing text-enhanced forecasting methods often encode such context into generic representations and fuse it uniformly with time-series features, without explicitly distinguishing which semantic effects are forecast-relevant, where they should interact with temporal representations, or how they should modify future dynamics. We introduce ReasonCast, a structured semantic intervention framework that translates event knowledge into forecast-specific operations. Before constructing an intervention, a lightweight agent examines the event context, the no-text forecast, and its predictive uncertainty to determine whether textual reasoning is needed and whether additional evidence or temporal-statistics tools should be invoked. For known regime-changing events, including holidays and mega-sales, semantic reasoning is activated by default. Rather than injecting free-form text, ReasonCast represents event knowledge through structured fields describing event relevance, demand direction, temporal shape, amplitude, and peak intensity. These fields interact selectively with event-related temporal components of a pretrained time-series foundation model. An additive path corrects local trends and temporal shapes, a multiplicative path captures event-driven level shifts, and an instance-wise mechanism calibrates intervention strength. When semantic intervention is unnecessary, the original numerical forecast can be preserved exactly. To make the resulting semantics not only well-formed but forecast-effective, ReasonCast introduces a forecast-grounded post-training curriculum. Schema SFT establishes valid and internally consistent semantic fields; semantic-field RL calibrates direction, shape, amplitude, and peak judgments using forecast-derived supervision; and forecast-utility RL evaluates candidate semantic interventions through a frozen forecaster, aligning reasoning outputs with marginal forecast improvement while penalizing negative transfer and unnecessary tool cost. ReasonCast lowers WMAPE by 3.29, 1.25, and 0.47 percentage points on holiday-sensitive categories, mega-sale-sensitive categories, and M5 event windows, respectively. On stable-sales periods, indiscriminate semantic intervention increases WMAPE by 1.68 percentage points, whereas suppressing unnecessary intervention preserves the numerical backbone. These results demonstrate the value of moving beyond generic text fusion toward structured, forecast-aligned semantic intervention. Keywords: time series forecasting, demand forecasting, large language models, agentic reasoning, multimodal learning, textātime-series fusion, event-aware forecasting 1. Introduction Commerce demand reflects two fundamentally different sources of variation. Recurring regularitiesāsuch as trend, seasonality, autocorrelation, and repeated campaign patternsācan be learned from historical observations. Future promotions, holidays, price changes, platform interventions, and product-specific events, however, may cause shifts that are not identifiable from past sales alone, a distinction also reflected in demand benchmarks and models of known future covariates (23; 13; 33). Time-series foundation models (TSFMs) provide strong numerical backbones by transferring recurring structure across large collections of series (4; 2; 35; 22; 5), with recent work further scaling model capacity, context length, and pattern specialization (17; 21; 20). Yet a numerical backbone cannot infer a future event that is absent from its inputs. Large language models (LLMs) offer a complementary capability: they can interpret heterogeneous event records and convert contextual knowledge into forecasting-relevant judgments. The value of such context is nevertheless conditional. Text helps when it supplies future events, constraints, or background information unavailable from numerical history (34); broad empirical evidence further shows that multimodal gains depend strongly on whether text contributes signals not already captured by the temporal input and backbone (40). Event-aware studies likewise find that unfiltered news can degrade forecasting relative to carefully selected evidence (32). Thus, the central question is not simply how to fuse more text, but whether, where, and to what extent semantic reasoning should alter a strong numerical forecast. Representative forecasting designs, summarized in Table 1, leave three challenges unresolved. First, semantic value is heterogeneous across products, events, and forecast windows, requiring instance-adaptive control. Second, useful semantics must enter the computation at an appropriate location and in an appropriate form: some events change local trends and shapes, whereas major campaigns produce multiplicative level shifts. Third, language supervision rewards plausible reasoning but does not ensure downstream predictive value. These challenges call for a forecaster that separates semantic reasoning from numerical prediction and treats text as a controlled intervention rather than an always-on modality. Table 1. Comparison of representative forecasting designs. ā , ā³ , and Ć denote explicit, partial or method-dependent, and no explicit support. Reasoning over exogenous events goes beyond encoding external text to infer forecast-relevant effects. Selective semantic intervention combines instance-level Skip/Basic/Tool routing with continuous control over intervention strength; designs without an explicit abstention route provide only partial support. Reasoner adaptation covers both parametric post-training and non-parametric use of forecasting feedback. Forecasting design Numerical forecast owner Reasoning over exogenous events Selective semantic intervention Exact no-text preservation Reasoner adaptation TS-only / TSFM (2; 4) TSFM Ć Ć ā N/A LLM numerical forecasting (6; 10) LLM Ć Ć Ć Frozen / prompting Joint multimodal forecasting (8; 12; 29; 36; 41) Joint model ā³ Ć Ć Task-level training / SFT From News to Forecast (32) LLM ā Ć Ć Reflection + evidence reselection VoT (31) Hybrid branches ā ā³ Ć Retrieval + historical ICL ReasonCast TSFM ā ā ā SFT + semantic-field RL + forecast-utility RL We introduce ReasonCast, an agentic demand forecasting framework following a routeāreasonāinterveneāforecast paradigm. Given event context, the no-text forecast, and its entropy, the agent selects Skip, low-cost Basic semantics, or tool-augmented Tool semantics; holidays and mega-sales rule out Skip. A pretrained TSFM remains responsible for numerical prediction. For non-skip actions, event-related temporal components alone query semantic tokens. An instance-wise gate controls intervention strength, while additive and multiplicative paths model local and level changes. Skip disables both paths and exactly recovers the no-text computation. Alignment pretraining grounds event semantics, and forecast-aligned policy optimization rewards routing and outputs by their marginal improvement over the backbone. We evaluate ReasonCast on large-scale commerce data at multiple granularities and public item-level demand benchmarks, separating normal and event-driven periods and stress-testing corrupted context. Our contributions are: ⢠We formulate event-enhanced demand forecasting as selective semantic intervention, shifting the focus from extracting more text to controlling its marginal influence on a strong numerical forecaster. ⢠We propose a three-action agent policy that routes each instance to no intervention, low-cost structured reasoning, or tool-augmented reasoning, with mandatory semantic activation for holidays and mega-sales. ⢠We combine selective event interaction, per-instance gating, additive and multiplicative corrections, and exact recovery of the no-text backbone whenever the agent skips intervention. 2. Related Work 2.1. Numerical and Language-Based Forecasting Numerical forecasting has progressed from probabilistic and covariate-aware models (25; 13) to specialized linear, patch-based, inverted, and multiscale architectures (39; 24; 18; 30; 33). Time-series foundation models further transfer recurring structure across datasets through value tokenization, decoder-style generation, or heterogeneous-series pretraining (4; 2; 35; 22; 5; 17). These models are strong numerical specialists, but they cannot infer a future event absent from their inputs. Other work repurposes pretrained language models as numerical forecasters or representation learners (6; 42; 27; 10; 3; 19; 16). Because language components do not consistently explain the resulting gains (28), ReasonCast retains a TSFM as forecast owner and assigns the LLM a distinct role: reasoning over event knowledge rather than generating sales values. 2.2. Multimodal Forecasting and Cross-Modal Alignment Text-enhanced forecasting uses either endogenous descriptions derived from the series (10; 27; 15; 41; 9) or exogenous evidence such as news, events, and domain conditions (8; 14; 12; 36; 29). The former can improve representation transfer but may duplicate numerical information; the latter can reveal shifts unavailable from history. Empirical studies accordingly find that multimodal gains depend on backbone capacity, alignment, data scale, and whether text adds predictive information (34; 40). ReasonCast turns this conditional utility into an instance-level control objective. 2.3. Event Reasoning and Forecast-Aware Agents Language agents can interleave reasoning, actions, and tool use (38; 26). In forecasting, From News to Forecast refines evidence selection using validation feedback (32), while VoT combines event and numerical branches with retrieval and adaptive fusion (31). ReasonCast instead routes each instance among no intervention, low-cost reasoning, and tool-augmented reasoning; structured semantics modify a dedicated TSFM only when selected. Forecast-aligned optimization trains this policy for marginal utility, and the no-text forecast is retained exactly when the agent skips. 3. Problem Formulation For each item or leaf category i, let (1) i=[xi,1,ā¦,xi,L]x_i=[x_i,1,ā¦,x_i,L] denote an observed demand history of length L, and let ic_i denote the event records and contextual information available before the forecast origin. The goal is to predict the next H=7H=7 observations, (2) ^i=[y^i,1,ā¦,y^i,H]. y_i=[ y_i,1,ā¦, y_i,H]. A pretrained Chronos-2 model fĪøf_Īø (1) provides the no-text forecast ^its=fĪøā(i) y^ts_i=f_Īø(x_i) and entropy āiH_i. With i=(i,^its,āi)o_i=(c_i, y^ts_i,H_i), the agent chooses (3) aiā¼ĻĻrouteā(i),aiāSkip,Basic,Tool,a_i _Ļ^route(o_i), a_iā\ Skip, Basic, Tool\, where holidays and mega-sales enforce aiā Skipa_iā Skip. We set i=ā T_i= for Skip and otherwise generate semantic tokens with ĻĻsemā(i,ai) _Ļ^sem(o_i,a_i). The forecast FĪø,Ļā(i,i)F_Īø,Ļ(x_i,T_i) satisfies (4) ai=Skipā¹FĪø,Ļā(i,ā )=fĪøā(i).a_i= Skip F_Īø,Ļ(x_i, )=f_Īø(x_i). 4. ReasonCast 4.1. Framework Overview ReasonCast is trained with a three-phase curriculum. First, schema SFT establishes valid structured outputs, after which semantic-field RL calibrates forecast-derived direction, temporal shape, amplitude, and peak fields. Second, the semantic-field policy regenerates the fusion-training corpus, on which we learn orthogonal event-subspace grounding, multimodal alignment, and selective dual-path fusion. Third, the trained forecaster is frozen and reused as a stable downstream evaluator for forecast-utility RL, which jointly optimizes routing and semantic outputs without updating the numerical model. Figure 1 summarizes the framework. Figure 1. Overview of ReasonCast. (a) ReasonCast follows a reasonāalignāinterveneāforecast pipeline. An evidence-augmented LLM agent combines the historical series with temporal statistics, promotion intelligence, platform signals, and external signals to produce structured event semantics. A text encoder and an MLP aligner map these semantics to aligned tokens T. In parallel, a pretrained time-series foundation model (TSFM) produces a future token Fh_F, which is orthogonally decomposed at the fusion interface into event-related and complementary components Eh_E and Nh_N. Only Eh_E queries T through event-selective cross-attention. An adaptive semantic gate modulates the additive latent correction for local trend and shape changes, while a complementary multiplicative path handles large event-driven level shifts. Without semantic context, both correction paths are disabled and the fusion interface exactly recovers the original TSFM computation. (b) The training curriculum comprises schema SFT, semantic-field RL, orthogonal event-subspace grounding and selective fusion training, and forecast-utility RL. During forecast-utility RL, candidate semantics are compared with the cached semantic-field-policy fusion result, while degradation beyond the no-text TSFM is explicitly penalized.A two-panel overview. The first panel traces event evidence through an LLM reasoner, text alignment, event-selective fusion, and a time-series forecast. The second panel shows the SFT, semantic-field RL, fusion training, and forecast-utility RL curriculum. 4.2. Agentic Event Routing and Semantic Reasoning Skip terminates semantic processing, Basic uses available context, and Tool retrieves evidence or invokes temporal-statistics tools for ambiguous or high-entropy cases. Hard-triggered holidays and mega-sales use at least Basic. For either non-skip action, Qwen3-32B (37) produces (5) i=(ri,ei,di,qi,ui,pi,zi),s_i=(r_i,\,e_i,\,d_i,\,q_i,\,u_i,\,p_i,\,z_i), where rir_i denotes event relevance, eie_i the dominant event, did_i demand direction, qiq_i temporal shape, uiu_i the amplitude multiplier, pip_i the peak multiplier, and ziz_i a concise forecast-trend narrative. A text encoder and MLP aligner produce D-dimensional tokens iT_i. SFT teaches the schema; forecast-aligned optimization calibrates routing and semantic utility. The agent decides whether and how much to reason, while the TSFM retains numerical forecasting. Reasoner supervision. We stratify 40,000 itemāwindow candidates by trajectory regime and obtain 11,627 teacher generations. Deterministic field-consistency filters retain 7,312 responses, which are compressed, deduplicated, and rebalanced into 4,590 SFT examples; Appendix B provides the full construction protocol. 4.3. Event-Selective Temporal Interaction Let iāāMĆDH_i ^MĆ D be temporal tokens at a fusion layer. From a learned basis āāDĆdEB ^DĆ d_E, QR factorization yields an orthonormal event basis Q with ā¤ā=Q Q=I. We define the orthogonal projector E=ā¤P_E=QQ and decompose (6) iE=iāE,iN=iā(āE).H^E_i=H_iP_E, ^N_i=H_i(I-P_E). Because Eā¤=EP_E =P_E and E2=EP_E^2=P_E, the decomposition guarantees iE+iN=iH^E_i+H^N_i=H_i and āØiE,iNā©F=0 ^E_i,H^N_i _F=0. Textual interaction is computed with the event component as query: (7) Īāi=CrossAttnā”(Q=iE,K=i,V=i). _i=CrossAttn(Q=H^E_i,K=T_i,V=T_i). The fused representation is (8) ~i=iN+iE+miāαāgiāĪāi, H_i=H^N_i+H^E_i+m_iα g_i _i, where mim_i is a text-availability mask, α is a residual scale, and gig_i is an instance-wise gate. The interface is applied to the Chronos-2 future-token representation before the prediction head. Thus, only the event component queries text; because the attention output is not reprojected, we make no stronger claim that the semantic update itself remains in the event component. 4.4. Multimodal Alignment Pretraining Before end-to-end forecasting training, we align pooled event semantics iTz^T_i with event-related temporal representations iEz^E_i. Motivated by contrastive cross-modal alignment (7) and its recent use in languageātime-series forecasting (15; 41), a contrastive objective encourages matched eventāseries pairs to be close and mismatched pairs to remain separated: (9) āalign=ā1Bāi=1Blogexpā”(simā”(iE,iT)/Ļ)āj=1Bexpā”(simā”(iE,jT)/Ļ).L_align=- 1B _i=1^B (sim(z^E_i,z^T_i)/Ļ) _j=1^B (sim(z^E_i,z^T_j)/Ļ). The denominator uses the other text embeddings in the same minibatch as negatives. 4.5. Instance-Adaptive Dual-Path Correction The value of semantic context varies across items and forecast windows. ReasonCast uses an instance-wise gate (10) gi=gmaxāĻā(MLPā”(i)),g_i=g_ Ļ\! (MLP(u_i) ), where iu_i concatenates pooled semantic features, recent-demand volatility, entropy of the no-text forecast, and encoded event descriptors, all available at inference time. The gate modulates the additive latent residual in Eq. (8). Latent residuals are well suited to fine-grained temporal pattern changes but may have limited authority over large level shifts. We therefore introduce a complementary multiplicative correction: (11) ^i=^ilatāexpā”(miāβāgiāi), y_i= y^lat_i \! (m_iβ g_i Ļ_i ), where ^ilat y^lat_i is decoded from the additively fused representation, i Ļ_i is a bounded semantic level-shift signal, and β controls the correction scale. 4.6. Exact No-Text Identity When the agent selects Skip, or when textual context is unavailable, mi=0m_i=0. Equations (8) and (11) then reduce to (12) ~i=i,^i=^its. H_i=H_i, y_i= y^ts_i. The property is enforced by the routing action and an explicit mask rather than learned approximately. Consequently, a skipped stable-sales instance recovers the Chronos-2 backbone prediction exactly. The guarantee applies to the fusion interface. Although its event subspace is geometrically orthogonal to the complement, this constraint does not make the subspace causally or globally identifiable. 4.7. Forecast-Aligned Reasoner Post-Training Language supervision alone can yield plausible but forecast-irrelevant text, whereas direct loss optimization may reward malformed semantics that happen to improve a prediction. Following schema SFT, we optimize the reasoner with two RL stages: semantic-field RL calibrates verifiable semantic fields, and forecast-utility RL optimizes their downstream value against a frozen forecaster. Semantic-field RL: structured semantic calibration. For prompt i, GRPO samples G structured responses. Forecast-derived targets score direction, temporal shape, amplitude, and peak: (13) Aiāj=ākād,q,a,pwkāaiājkāCiājstruct,A_ij= _kā\d,q,a,p\w_ka^k_ij-C_ij^struct, where CiājstructC_ij^struct penalizes invalid formats and cross-field contradictions. Group-normalized rewards and a KL reference to the SFT policy yield the field-calibrated reasoner Ļ1 _1. Forecast-utility RL: frozen-forecaster optimization. We use Ļ1 _1 to regenerate the fusion corpus, train the alignment and dual-path forecaster F1F_1, and then freeze it. Each rollout selects Skip, Basic, or Tool; holidays and mega-sales disallow Skip. Non-skip candidates must remain parseable, cross-field consistent, and close to the cached semantic-field score before being evaluated by F1F_1. Let MiājM_ij be candidate MAE, MirefM_i^ref the cached semantic-field-policy fusion MAE, MitsM_i^ts the no-text MAE, and SiS_i a sample-specific scale. We measure downstream improvement and one-sided harm as (14) Diāj=MirefāMiājSi,Hiāj=[Miājā(1+Ī·)āMits]+Si.D_ij= M_i^ref-M_ijS_i, H_ij= [M_ij-(1+Ī·)M_i^ts]_+S_i. With semantic-validity indicator viājv_ij, semantic margin CiājC_ij, and route cost cā”(aiāj)c(a_ij), the reward is (15) Riāj=b+Ī»uāclipā”(Diāj),aiāj=Skip,Aiājāγ,viāj=0,b+Ī»cāCiāj+Ī»uāclipā”(DiājāĪ»hāHiāj)āĪ»costācā(aiāj),viāj=1.R_ij= casesb+ _uclip(D_ij),&a_ij= Skip,\\ A_ij-γ,&v_ij=0,\\ b+ _cC_ij+ _uclip(D_ij- _hH_ij)- _costc(a_ij),&v_ij=1. cases Thus, abstention competes directly with semantic intervention, while tool use must justify its additional cost. GRPO is initialized from Ļ1 _1 and uses the semantic-field policy globally, plus SFT conditionally, as KL references to preserve semantic fidelity. Appendix B provides the validity threshold, reference conditions, and optimization settings. 5. Experiments 5.1. Research Questions We evaluate ReasonCast through four questions: ⢠RQ1: Does ReasonCast improve event-period forecasting and transfer to M5 while preserving stable-period performance? ⢠RQ2: Do discrete routing and entropy-conditioned gating allocate semantic intervention more effectively than always-on, ungated, or globally weighted fusion? ⢠RQ3: What do orthogonal decomposition, alignment, gating, and additiveāmultiplicative correction each contribute? ⢠RQ4: How do SFT, semantic-field RL, and forecast-utility RL affect semantic fidelity and downstream forecast utility? 5.2. Datasets and Evaluation Protocol The proprietary commerce dataset contains 15,696 daily sales series at the leaf-category level, spanning August 10, 2024 to May 30, 2026. We adopt a chronological split: August 10, 2024 to August 31, 2025 for training, September 1 to October 31, 2025 for model selection, and November 1, 2025 to May 30, 2026 for evaluation. The test set comprises 15 seven-day forecasting windows, including three mega-sale, ten holiday, and two stable-sales windows. To prevent temporal leakage, all textual descriptions and event information are restricted to records available before the corresponding forecast origin. For M5 (23), we evaluate 3,049 item-level series aggregated across all ten stores over 14 seven-day event-centered windows, using only calendar, event, and price covariates available at each forecast origin. 5.3. Baselines We compare Seasonal Naive with LightGBM (11), PatchTST (24), and iTransformer (18); pretrained TimesFM (4), Chronos (2), and Chronos-2 small (1); and the text-enhanced Time-LLM (10) and VoT (31). Chronos-2 small is evaluated zero-shot and with the Uni-FT and Multi-FT variants shown in the tables. Numerical models receive historical demand, while covariate-enabled and text-enhanced variants additionally receive only calendar, event, price, or text fields available before the forecast origin. All methods use the same split, rolling origins, seven-day horizon, and metrics; trainable baselines are selected by validation WMAPE without test-window tuning. 5.4. Metrics We evaluate point forecasting accuracy using weighted mean absolute percentage error (WMAPE) and mean absolute error (MAE). Let sD_s denote the set of leaf-categoryāday observations in an evaluation slice s. We compute (16) WMAPEs=ā(i,t)ās|yi,tāy^i,t|ā(i,t)ās|yi,t|,WMAPE_s= _(i,t) _s|y_i,t- y_i,t| _(i,t) _s|y_i,t|, and (17) MAEs=1|s|āā(i,t)ās|yi,tāy^i,t|.MAE_s= 1|D_s| _(i,t) _s|y_i,t- y_i,t|. WMAPE is the primary metric because it measures aggregate forecasting error relative to realized demand and places greater emphasis on high-volume observations. MAE complements WMAPE by retaining the absolute magnitude of forecasting errors. Both metrics are computed separately for each evaluation slice, including stable-sales periods, holiday and mega-sale periods, event-sensitive and event-insensitive leaf categories, and the M5 dataset. Lower values indicate better performance. 5.5. Main Results Table 2 compares event-centered performance on the proprietary dataset and M5. Holiday and mega-sale categories are split into sensitive and insensitive groups using training data only; M5 uses windows around major calendar events. Table 3 separately tests the boundary case in which regular temporal dynamics already explain most demand. Table 2. Main forecasting results on event-centered evaluations. Cells report WMAPE/MAE; lower values are better. Sens. and Insens. denote event-sensitive and event-insensitive leaf categories, respectively. The best results are highlighted in bold, and the second-best results are underlined. Method Holiday Mega-sale M5 event windows Sens. Insens. All Sens. Insens. All WMAPE MAE Seasonal Naive 119.96% 2738.47 56.69% 609.09 47.39% 571.90 27.94% 766.85 12.49% 283.57 16.55% 434.69 44.55% 5.47 LightGBM 74.61% 1671.02 43.20% 489.11 34.23% 447.08 26.94% 732.37 12.83% 290.49 16.23% 425.58 34.06% 4.17 PatchTST 84.90% 1771.26 47.80% 516.75 39.78% 485.47 22.94% 652.38 9.91% 214.34 13.41% 355.74 48.31% 5.91 iTransformer 79.29% 1849.54 42.03% 486.83 33.64% 450.65 23.07% 653.62 11.30% 248.18 14.49% 382.73 34.39% 4.21 TimesFM 67.23% 1586.46 48.82% 557.09 39.37% 517.45 24.20% 701.77 14.57% 316.56 17.63% 465.66 34.63% 4.23 Chronos1 44.54% 1163.21 37.25% 450.04 29.70% 417.03 26.99% 730.93 10.97% 248.37 14.90% 391.61 35.38% 4.33 Chronos2-small 55.42% 1412.40 38.50% 464.16 31.28% 437.37 24.59% 681.33 10.83% 245.81 14.64% 389.11 34.40% 4.20 Chronos2-small (Uni-FT) 42.39% 1185.03 34.94% 415.31 27.89% 388.79 25.20% 694.97 10.75% 242.43 14.55% 382.88 33.21% 4.08 Chronos2-small (Multi-FT) 45.30% 1268.46 31.64% 390.67 25.59% 370.86 22.14% 622.10 10.02% 222.37 13.35% 352.24 33.10% 4.06 Time-LLM 87.76% 1972.79 57.31% 602.99 47.03% 556.71 25.89% 721.11 13.57% 306.44 16.97% 445.65 38.99% 4.78 VoT 61.51% 1514.34 32.38% 417.44 27.33% 400.96 24.74% 711.91 13.42% 295.53 16.83% 444.76 38.05% 4.69 ReasonCast (agent-routed) 39.10% 1080.71 25.08% 331.09 21.06% 322.24 20.89% 589.94 9.71% 215.61 12.87% 339.68 32.63% 4.00 ReasonCast obtains the lowest WMAPE on all six proprietary event slices and on M5 event windows, with the largest margins during holidays. It also achieves the best WMAPE and second-best MAE on mega-sale-insensitive categories, indicating that the gain transfers beyond the most responsive groups and beyond the proprietary dataset. Table 3. Forecasting performance during stable-sales periods on the proprietary commerce dataset. Lower values are better. The best results are highlighted in bold, and the second-best results are underlined. Method WMAPE MAE Seasonal Naive 10.89% 254.82 LightGBM 11.10% 259.88 PatchTST 9.72% 227.22 iTransformer 10.26% 239.83 TimesFM 13.19% 308.73 Chronos1 8.33% 194.61 Chronos2-small 7.91% 185.15 Chronos2-small (Uni-FT) 9.57% 223.80 Chronos2-small (Multi-FT) 8.85% 206.99 Time-LLM 13.34% 312.68 VoT 11.31% 264.57 ReasonCast (always-on semantics) 9.59% 224.43 ReasonCast (agent-routed) 7.91% 185.15 Table 3 isolates routing. Always-on semantics degrade the Chronos2-small backbone from 7.91%/185.15 to 9.59%/224.43. Agent-routed ReasonCast selects Skip and recovers the backbone exactly. Event windows, by contrast, require at least Basic reasoning. 5.6. Dissecting Selective Semantic Intervention Table 4 tests where, how strongly, and in what form semantics enter the forecast. Its stable-sales column forces Basic to isolate fusion from routing; positive changes denote degradation. Table 4. Selective-fusion ablations (WMAPE/MAE). The stable-sales column forces Basic to isolate fusion from routing. Gray values report changes from the full fusion mechanism as WMAPE percentage points / relative MAE (%); positive values indicate degradation. Lower is better. Design axis Variant Holiday Sens. Mega-sale Sens. Stable-sales (forced Basic) M5 Full fusion mechanism 39.10% / 1080.71 [-1pt] reference 20.89% / 589.94 [-1pt] reference 9.59% / 224.43 [-1pt] reference 32.63% / 4.00 [-1pt] reference Interaction target w/o orthogonal decompositionāalignment 40.04% / 1155.23 [-1pt] Ī+0.94p/+6.90% \;+0.94\,p\,/\,+6.90\% 21.14% / 596.19 [-1pt] Ī+0.25p/+1.06% \;+0.25\,p\,/\,+1.06\% 11.31% / 264.57 [-1pt] Ī+1.72p/+17.89% \;+1.72\,p\,/\,+17.89\% 32.64% / 4.00 [-1pt] Ī+0.01p/+0.00% \;+0.01\,p\,/\,+0.00\% Intervention strength w/o adaptive semantic gate 39.76% / 1112.21 [-1pt] Ī+0.66p/+2.91% \;+0.66\,p\,/\,+2.91\% 21.54% / 608.01 [-1pt] Ī+0.65p/+3.06% \;+0.65\,p\,/\,+3.06\% 9.64% / 225.53 [-1pt] Ī+0.05p/+0.49% \;+0.05\,p\,/\,+0.49\% 32.64% / 4.01 [-1pt] Ī+0.01p/+0.25% \;+0.01\,p\,/\,+0.25\% Correction form w/o additive correction 40.12% / 1189.98 [-1pt] Ī+1.02p/+10.11% \;+1.02\,p\,/\,+10.11\% 22.14% / 622.09 [-1pt] Ī+1.25p/+5.45% \;+1.25\,p\,/\,+5.45\% 8.85% / 206.99 [-1pt] Īā0.74p/ā7.77% \;-0.74\,p\,/\,-7.77\% 32.84% / 4.03 [-1pt] Ī+0.21p/+0.75% \;+0.21\,p\,/\,+0.75\% w/o multiplicative correction 44.29% / 1159.19 [-1pt] Ī+5.19p/+7.26% \;+5.19\,p\,/\,+7.26\% 20.89% / 589.95 [-1pt] Ī+0.00p/+0.00% \;+0.00\,p\,/\,+0.00\% 9.59% / 224.43 [-1pt] Ī+0.00p/+0.00% \;+0.00\,p\,/\,+0.00\% 32.65% / 4.01 [-1pt] Ī+0.02p/+0.25% \;+0.02\,p\,/\,+0.25\% Reference Base fusion only 44.30% / 1229.05 [-1pt] Ī+5.20p/+13.73% \;+5.20\,p\,/\,+13.73\% 21.50% / 606.18 [-1pt] Ī+0.61p/+2.75% \;+0.61\,p\,/\,+2.75\% 8.74% / 204.43 [-1pt] Īā0.85p/ā8.91% \;-0.85\,p\,/\,-8.91\% 33.03% / 4.04 [-1pt] Ī+0.40p/+1.00% \;+0.40\,p\,/\,+1.00\% Removing orthogonal decompositionāalignment degrades every internal slice, although its M5 effect is negligible. Removing the adaptive gate also increases WMAPE in all four slices, with much larger changes during event-sensitive periods; this supports adaptive aggregate control but does not establish sample-level harm rates. The correction paths are complementary: the additive path improves both event regimes, whereas the multiplicative path is especially important for the large holiday level shift. Under forced Basic, simplified fusion can perform better on stable sales, confirming that unnecessary semantics still add noise; discrete routing removes this failure mode through the exact no-text path. 5.7. Reasoner Post-Training: From Semantic Fidelity to Forecast Utility We evaluate cumulative reasoner stages with the same frozen fusion forecaster; only the reasoning policy changes. Event WMAPE pools the absolute-error numerator and demand denominator over holiday- and mega-sale-sensitive slices. For seven-day instance i, let āim=āh|yi,hāy^i,hm|/(āh|yi,h|+ϵ) _i^m= _h|y_i,h- y_i,h^m|/( _h|y_i,h|+ε). We retain the sample-level negative transfer rate NTR=Nā1āi[āitext>āinoā-ātext]NTR=N^-1 _iI[ _i^text> _i^no -text], using the exact no-text path of the same frozen forecaster. Table 5. Semantic fidelity and downstream forecast utility across cumulative reasoner post-training stages. Event WMAPE pools holiday-sensitive and mega-sale-sensitive observations. Negative Transfer Rate is the proportion of instances for which text-conditioned seven-day normalized absolute error exceeds the exact no-text forecast from the same frozen forecaster. Arrows indicate whether higher or lower values are better. Training stage Parse Rate ā Direction MAE ā Shape Weighted-F1 ā Amplitude log-MAE ā Peak log-MAE ā Event WMAPE ā Negative Transfer Rate ā Original LLM 91.36% 1.14 0.25 0.44 0.54 20.52% 23.59% + Schema SFT 100.00% 0.91 0.34 0.44 0.53 19.48% 16.20% + Semantic-Field RL (GRPO) 100.00% 0.85 0.33 0.39 0.52 19.16% 16.03% + Forecast-Utility RL (GRPO) 100.00% 0.83 0.35 0.38 0.53 18.96% 15.08% SFT establishes parseability and improves direction and shape; semantic-field RL further calibrates direction, amplitude, and peak. Forecast-utility RL targets Event WMAPE and NTR while preserving semantic fidelity, separating well-formed text from interventions that improve the final forecast. 5.8. Entropy-Adaptive Semantic Intervention We isolate continuous gate control after a non-skip route. Within each forecast window, itemāwindow pairs are grouped into entropy tertiles and compared with a global-α control that replaces the instance-wise gate with one shared additive coefficient. Figure 2. Entropy-adaptive semantic intervention across within-window forecast-entropy tertiles. (a) WMAPE of the no-text backbone; (b) mean additive-injection magnitude |gi||g_i|, with the dashed line denoting the global-α control; and (c) WMAPE gain over that control in percentage points.Three panels group forecasts into low-, medium-, and high-entropy tertiles, showing increasing no-text backbone WMAPE, increasing additive-injection magnitude, and positive WMAPE gains over a global-$α$ control. Figure 2 links higher no-text entropy to stronger semantic correction. Adaptive gating outperforms global α in every tier, with larger gains at medium and high entropy, showing that intervention strength tracks the reliability of the numerical forecast. Semantic inversion stress test. We invert all six structured semantic fields and the trend narrative while holding the time series, checkpoint, and inference configuration fixed. Incorrect semantics increase WMAPE in 13 of 15 windows, with mean and median increases of 2.67 and 0.76 percentage points and a worst-case increase of 25.33 points. Thus, the forecast responds to semantic content rather than merely to the presence of text; gating attenuates many corruptions but does not guarantee per-window immunity. Full window-level results appear in Table 10. 6. Discussion Our results support interpreting event-enhanced forecasting as conditional intervention rather than uniform multimodal fusion. ReasonCast yields its clearest gains when future context supplies information that is difficult to recover from historical demand alone: relative to the strongest non-ReasonCast baselines, it reduces WMAPE by 7.8% on holiday-sensitive categories, 5.6% on mega-sale-sensitive categories, and 1.4% on M5 event windows. In contrast, forcing semantic intervention during stable-sales periods increases WMAPE from 7.91% to 9.59%. This asymmetry is central to our formulation: semantic context has regime-dependent marginal utility and should modify a strong numerical forecaster only when it contributes information beyond the observed history. The analyses further explain how ReasonCast realizes this principle. Discrete routing and continuous gating determine when semantic reasoning is invoked and how strongly it influences the forecast; the orthogonal decomposition determines where interaction occurs by restricting the semantic query interface to event-related temporal components; and the additive and multiplicative paths express different forms of correction. Removing orthogonal decomposition and alignment degrades all evaluated slices, while the multiplicative path is particularly important for holiday-driven level shifts. The additive path improves event-sensitive forecasts but can be harmful when intervention is unnecessarily forced on stable demand. Adaptive gating also outperforms a global-strength control across all forecast-entropy tiers, assigning greater influence to semantics when the numerical forecast is less reliable. Finally, semantic inversion increases WMAPE in 13 of 15 windows, confirming that the forecast responds to semantic content rather than merely to the presence of text. These results support forecast-conditioned semantic interaction, although they do not imply that the learned event subspace is a causally identifiable factor. The stable-sales analysis clarifies the distinct roles of routing and fusion. In Table 4, the stable-sales column deliberately overrides the routing policy and forces every instance to take the Basic route. It therefore diagnoses how the fusion mechanism behaves under unnecessary intervention rather than reporting end-to-end ReasonCast performance. In deployment, the agent may instead choose Skip, disable both correction paths, and exactly recover the numerical backbone. The evidence thus supports two narrower claims: the fusion components improve forecasting when semantic intervention is warranted, while the exact-identity skip path prevents the model from systematically modifying the backbone when no intervention is selected. It does not imply that the fusion gate alone guarantees the absence of sample-level negative transfer. The post-training results highlight a related distinction between semantic fidelity and forecast utility. Schema SFT establishes parseable and internally consistent outputs, and semantic-field RL improves forecast-critical judgments such as direction, amplitude, and peak intensity. Forecast-utility RL then evaluates valid candidates through a frozen forecaster, incorporating marginal predictive improvement, potential harm, and tool cost into the reasoning objective. This separation is important because well-formed event descriptions are not necessarily useful interventions. Several broader challenges remain for deploying event-aware forecasting systems. Event effects are intrinsically non-stationary across products, regions, and seasons, requiring semantic interpretation and forecast intervention to be recalibrated under distribution shift. External signals also vary in timeliness, completeness, and reliability, making evidence validation and uncertainty-aware reasoning essential. At production scale, tool-augmented reasoning introduces a fundamental trade-off among predictive accuracy, latency, and computational cost, while sparse counterfactual observations make it difficult to separate forecast-relevant associations from causal event effects. A specific limitation of the current study is that it does not yet provide a complete accounting of tool-call frequency, end-to-end latency, or realized inference cost. Future work may combine online adaptation, calibrated abstention, source-reliability modeling, and resource-aware routing to make semantic intervention more robust and operationally efficient. 7. Conclusion We presented ReasonCast, an agentic framework that treats event semantics as a routed intervention rather than an always-on forecasting input. The agent selects among no intervention, low-cost structured reasoning, and tool-augmented reasoning, while continuous gating controls intervention strength after a non-skip decision. Schema SFT, semantic-field RL, and forecast-utility RL progressively align structured reasoning with downstream predictive value. The results show that forecast-aligned routing can exploit contextual signals during event-driven shifts while exactly preserving the numerical backbone when semantics are unnecessary, thereby avoiding systematic negative transfer in stable-sales periods. References Ansari et al. (2025) A. F. Ansari, O. Shchur, J. Küken, A. Auer, B. Han, P. Mercado, S. S. Rangapuram, H. Shen, L. Stella, X. Zhang, et al. Chronos-2: from univariate to universal forecasting. External Links: 2510.15821, Document, Link Cited by: §3, §5.3. Ansari et al. (2024) A. F. Ansari, L. Stella, A. C. Turkmen, X. Zhang, P. Mercado, H. Shen, O. Shchur, S. S. Rangapuram, S. Pineda-Arango, S. Kapoor, J. Zschiegner, D. C. Maddix, H. Wang, M. W. Mahoney, K. Torkkola, A. G. Wilson, M. Bohlke-Schneider, and B. Wang Chronos: learning the language of time series. Note: Transactions on Machine Learning Research External Links: Link Cited by: Table 1, §1, §2.1, §5.3. Cao et al. (2024) D. Cao, F. Jia, S. Ć. Arik, T. Pfister, Y. Zheng, W. Ye, and Y. Liu TEMPO: prompt-based generative pre-trained transformer for time series forecasting. Note: International Conference on Learning Representations External Links: Link Cited by: §2.1. Das et al. (2024) A. Das, W. Kong, R. Sen, and Y. Zhou A decoder-only foundation model for time-series forecasting. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, Vienna, Austria, p. 10148ā10167. External Links: Link Cited by: Table 1, §1, §2.1, §5.3. Goswami et al. (2024) M. Goswami, K. Szafer, A. Choudhry, Y. Cai, S. Li, and A. Dubrawski MOMENT: a family of open time-series foundation models. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, Vienna, Austria, p. 16115ā16152. External Links: Link Cited by: §1, §2.1. Gruver et al. (2023) N. Gruver, M. Finzi, S. Qiu, and A. G. Wilson Large language models are zero-shot time series forecasters. Note: Advances in Neural Information Processing Systems External Links: Link Cited by: Table 1, §2.1. Jia et al. (2021) C. Jia, Y. Yang, Y. Xia, Y. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig Scaling up visual and vision-language representation learning with noisy text supervision. In Proceedings of the 38th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 139, Virtual, p. 4904ā4916. External Links: Link Cited by: §4.4. Jia et al. (2024) F. Jia, K. Wang, Y. Zheng, D. Cao, and Y. Liu GPT4MTS: prompt-based large language model for multimodal time-series forecasting. Proceedings of the AAAI Conference on Artificial Intelligence 38 (21), p. 23343ā23351. External Links: Document Cited by: Table 1, §2.2. Jia et al. (2026) S. Jia, B. Song, C. Ye, and C. Yuan M3Time: LLM-enhanced multi-modal, multi-scale, and multi-frequency multivariate time series forecasting. Proceedings of the AAAI Conference on Artificial Intelligence 40 (27), p. 22265ā22273. External Links: Document Cited by: §2.2. Jin et al. (2024) M. Jin, S. Wang, L. Ma, Z. Chu, J. Y. Zhang, X. Shi, P. Chen, Y. Liang, Y. Li, S. Pan, and Q. Wen Time-LLM: time series forecasting by reprogramming large language models. Note: International Conference on Learning Representations External Links: Link Cited by: Table 1, §2.1, §2.2, §5.3. Ke et al. (2017) G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T. Liu LightGBM: a highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems, Vol. 30, Long Beach, CA, USA, p. 3146ā3154. External Links: Link Cited by: §5.3. Li et al. (2026) Z. Li, X. Lin, Z. Liu, J. Zou, Z. Wu, L. Zheng, D. Fu, Y. Zhu, H. Hamann, H. Tong, and J. He Language in the flow of time: time-series-paired texts weaved into a unified temporal narrative. Note: International Conference on Learning Representations External Links: Link Cited by: Table 1, §2.2. Lim et al. (2021) B. Lim, S. Ć. Arik, N. Loeff, and T. Pfister Temporal fusion transformers for interpretable multi-horizon time series forecasting. International Journal of Forecasting 37 (4), p. 1748ā1764. External Links: Document Cited by: §1, §2.1. Liu et al. (2024a) H. Liu, S. Xu, Z. Zhao, L. Kong, H. Kamarthi, A. B. Sasanur, M. Sharma, J. Cui, Q. Wen, C. Zhang, and B. A. Prakash Time-MMD: multi-domain multimodal dataset for time series analysis. Note: Advances in Neural Information Processing Systems External Links: Document, Link Cited by: §2.2. Liu et al. (2025a) P. Liu, H. Guo, T. Dai, N. Li, J. Bao, X. Ren, Y. Jiang, and S. Xia CALF: aligning LLMs for time series forecasting via cross-modal fine-tuning. Proceedings of the AAAI Conference on Artificial Intelligence 39 (18), p. 18915ā18923. External Links: Document Cited by: §2.2, §4.4. Liu et al. (2024b) X. Liu, J. Hu, Y. Li, S. Diao, Y. Liang, B. Hooi, and R. Zimmermann UniTime: a language-empowered unified model for cross-domain time series forecasting. In Proceedings of the ACM Web Conference 2024, Singapore, p. 4095ā4106. External Links: Document Cited by: §2.1. Liu et al. (2025b) X. Liu, J. Liu, G. Woo, T. Aksu, Y. Liang, R. Zimmermann, C. Liu, J. Li, S. Savarese, C. Xiong, and D. Sahoo Moirai-MoE: empowering time series foundation models with sparse mixture of experts. In Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 267, Vancouver, BC, Canada, p. 38940ā38962. External Links: Link Cited by: §1, §2.1. Liu et al. (2024c) Y. Liu, T. Hu, H. Zhang, H. Wu, S. Wang, L. Ma, and M. Long ITransformer: inverted transformers are effective for time series forecasting. Note: International Conference on Learning Representations External Links: Link Cited by: §2.1, §5.3. Liu et al. (2024d) Y. Liu, G. Qin, X. Huang, J. Wang, and M. Long AutoTimes: autoregressive time series forecasters via large language models. Note: Advances in Neural Information Processing Systems External Links: Document, Link Cited by: §2.1. Liu et al. (2025c) Y. Liu, G. Qin, X. Huang, J. Wang, and M. Long Timer-XL: long-context transformers for unified time series forecasting. Note: International Conference on Learning Representations External Links: Link Cited by: §1. Liu et al. (2025d) Y. Liu, G. Qin, Z. Shi, Z. Chen, C. Yang, X. Huang, J. Wang, and M. Long Sundial: a family of highly capable time series foundation models. In Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 267, Vancouver, BC, Canada, p. 39295ā39317. External Links: Link Cited by: §1. Liu et al. (2024e) Y. Liu, H. Zhang, C. Li, X. Huang, J. Wang, and M. Long Timer: generative pre-trained transformers are large time series models. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, Vienna, Austria, p. 32369ā32399. External Links: Link Cited by: §1, §2.1. Makridakis et al. (2022) S. Makridakis, E. Spiliotis, and V. Assimakopoulos M5 accuracy competition: results, findings, and conclusions. International Journal of Forecasting 38 (4), p. 1346ā1364. External Links: Document Cited by: §1, §5.2. Nie et al. (2023) Y. Nie, N. H. Nguyen, P. Sinthong, and J. Kalagnanam A time series is worth 64 words: long-term forecasting with transformers. Note: International Conference on Learning Representations External Links: Link Cited by: §2.1, §5.3. Salinas et al. (2020) D. Salinas, V. Flunkert, J. Gasthaus, and T. Januschowski DeepAR: probabilistic forecasting with autoregressive recurrent networks. International Journal of Forecasting 36 (3), p. 1181ā1191. External Links: Document Cited by: §2.1. Schick et al. (2023) T. Schick, J. Dwivedi-Yu, R. DessƬ, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom Toolformer: language models can teach themselves to use tools. Note: Advances in Neural Information Processing Systems External Links: Link Cited by: §2.3. Sun et al. (2024) C. Sun, H. Li, Y. Li, and S. Hong TEST: text prototype aligned embedding to activate LLMās ability for time series. Note: International Conference on Learning Representations External Links: Link Cited by: §2.1, §2.2. Tan et al. (2024) M. Tan, M. A. Merrill, V. Gupta, T. Althoff, and T. Hartvigsen Are language models actually useful for time series forecasting?. Note: Advances in Neural Information Processing Systems External Links: Document, Link Cited by: §2.1. Wang et al. (2025) C. Wang, Q. Qi, J. Wang, H. Sun, Z. Zhuang, J. Wu, L. Zhang, and J. Liao ChatTime: a unified multimodal time series foundation model bridging numerical and textual data. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, Philadelphia, PA, USA, p. 12694ā12702. External Links: Document, Link Cited by: Table 1, §2.2. Wang et al. (2024a) S. Wang, H. Wu, X. Shi, T. Hu, H. Luo, L. Ma, J. Y. Zhang, and J. Zhou TimeMixer: decomposable multiscale mixing for time series forecasting. Note: International Conference on Learning Representations External Links: Link Cited by: §2.1. Wang et al. (2026) S. Wang, P. Chen, Y. Wang, W. Qiu, C. Guo, B. Yang, and Y. Shu Unlocking the value of text: event-driven reasoning and multi-level alignment for time series forecasting. Note: International Conference on Learning Representations External Links: Link Cited by: Table 1, §2.3, §5.3. Wang et al. (2024b) X. Wang, M. Feng, J. Qiu, J. Gu, and J. Zhao From news to forecast: integrating event analysis in LLM-based time series forecasting with reflection. In Advances in Neural Information Processing Systems, Vol. 37, Vancouver, BC, Canada, p. 58118ā58153. External Links: Document, Link Cited by: Table 1, §1, §2.3. Wang et al. (2024c) Y. Wang, H. Wu, J. Dong, G. Qin, H. Zhang, Y. Liu, Y. Qiu, J. Wang, and M. Long TimeXer: empowering transformers for time series forecasting with exogenous variables. Note: Advances in Neural Information Processing Systems External Links: Document, Link Cited by: §1, §2.1. Williams et al. (2025) A. R. Williams, A. Ashok, Ć. Marcotte, V. Zantedeschi, J. Subramanian, R. Riachi, J. Requeima, A. Lacoste, I. Rish, N. Chapados, and A. Drouin Context is key: a benchmark for forecasting with essential textual information. In Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 267, Vancouver, BC, Canada, p. 66887ā66944. External Links: Link Cited by: §1, §2.2. Woo et al. (2024) G. Woo, C. Liu, A. Kumar, C. Xiong, S. Savarese, and D. Sahoo Unified training of universal time series forecasting transformers. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, Vienna, Austria, p. 53140ā53164. External Links: Link Cited by: §1, §2.1. Wu et al. (2026) X. Wu, J. Jin, W. Qiu, P. Chen, Y. Shu, B. Yang, and C. Guo Aurora: towards universal generative multimodal time series forecasting. Note: International Conference on Learning Representations External Links: Link Cited by: Table 1, §2.2. Yang et al. (2025) A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al. Qwen3 technical report. External Links: 2505.09388, Document, Link Cited by: Table 13, §4.2. Yao et al. (2023) S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao ReAct: synergizing reasoning and acting in language models. Note: International Conference on Learning Representations External Links: Link Cited by: §2.3. Zeng et al. (2023) A. Zeng, M. Chen, L. Zhang, and Q. Xu Are transformers effective for time series forecasting?. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, Washington, DC, USA, p. 11121ā11128. External Links: Document Cited by: §2.1. Zhang et al. (2026) X. Zhang, B. Han, H. Fang, A. F. Ansari, S. Zhang, D. C. Maddix, C. Hu, A. G. Wilson, M. W. Mahoney, H. Wang, Y. Liu, M. Bohlke-Schneider, H. Rangwala, G. Karypis, and B. Wang When does multimodality lead to better time series forecasting?. Note: Transactions on Machine Learning Research External Links: Link Cited by: §1, §2.2. Zhou et al. (2025) S. Zhou, H. Schƶner, H. Lyu, E. FouchĆ©, and S. Wang BALM-TSF: balanced multimodal alignment for LLM-based time series forecasting. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management, Seoul, Republic of Korea, p. 4498ā4508. External Links: Document Cited by: Table 1, §2.2, §4.4. Zhou et al. (2023) T. Zhou, P. Niu, X. Wang, L. Sun, and R. Jin One fits all: power general time series analysis by pretrained LM. In Advances in Neural Information Processing Systems, Vol. 36, New Orleans, LA, USA, p. 43322ā43355. External Links: Link Cited by: §2.1. Appendix A Detailed Window-Level Results This appendix reports the unaggregated seven-day forecast windows underlying the main-text averages. Each cell contains WMAPE and MAE for exactly the same items, forecast origin, and horizon; no window-specific model selection is performed. For compactness, Chronos2-small is abbreviated as C2-small and the time-series backbone as TS backb. A.1. Commerce Dataset Tables 6 and 7 contain the 15 commerce windows in chronological blocks: three mega-sale windows, ten holiday windows, and two stable-sales windows. The split across two tables is purely presentational; the rows and evaluation protocol are identical. The hardest cases are the two Spring FestivalāValentine windows beginning on February 12 and February 15. On February 12, ReasonCast reduces WMAPE from 108.75% for the strongest non-ReasonCast comparator in its table to 83.85%; on February 15 it further reduces the best comparator from 28.77% to 24.53%. These windows combine a sharp holiday regime change with overlapping event semantics, precisely the setting in which history-only extrapolation is least reliable. The stable-sales rows provide an important counterpoint. ReasonCast is not the best model on January 13 or March 23: several purely numerical models attain lower error. The aggregate benefit therefore does not come from uniformly overriding the backbone. Instead, the detailed rows support the intended division of labor: semantic intervention is most useful when future events add information absent from the observed history, whereas the numerical backbone remains highly competitive when demand is stable. A.2. M5 Dataset Tables 8 and 9 report all 14 M5 event windows. Each holiday is evaluated at two forecast origins, which prevents a single favorable alignment between the event day and horizon from determining the aggregate result. The public benchmark has a different error profile from the commerce data. Absolute differences are smaller, but ReasonCast remains consistently competitive across Halloween, New Year, Super Bowl, Easter, and Motherās Day origins. It does not dominate every origin: the time-series backbone is slightly better on the first Christmas origin, and the covariate baseline is slightly better on the first Thanksgiving origin. Reporting both origins therefore exposes where semantic context helps and where a strong numerical model already captures most of the predictable event pattern. Table 6. Commerce window-level results, baseline group A. Each method occupies a WMAPE (%) and MAE pair. Rows 1ā3 are mega-sale windows, rows 4ā13 are holiday windows, and rows 14ā15 are stable-sales windows. SNaive LightGBM PatchTST iTransformer TimesFM C2-small Window WMAPE MAE WMAPE MAE WMAPE MAE WMAPE MAE WMAPE MAE WMAPE MAE 2025-11-11 19.61% 547.89 18.44% 515.27 19.68% 550.02 19.63% 548.61 23.87% 667.14 18.03% 509.35 2025-12-08 15.55% 392.93 17.21% 434.92 10.72% 270.99 12.40% 313.53 15.12% 382.17 15.94% 406.74 2025-12-11 14.51% 363.26 13.04% 326.56 9.84% 246.23 11.43% 286.05 13.89% 347.66 9.94% 251.24 2026-01-01 18.68% 446.77 17.61% 421.21 18.82% 450.14 18.36% 439.06 19.72% 471.68 29.83% 719.76 2026-02-12 210.75% 1614.50 169.72% 1300.13 175.83% 1347.00 159.53% 1222.10 186.35% 1427.55 138.84% 1070.39 2026-02-15 131.99% 1077.18 54.45% 444.35 107.73% 879.20 53.02% 432.70 67.55% 551.23 48.13% 395.08 2026-03-02 16.67% 397.63 15.35% 366.01 12.78% 304.67 14.74% 351.40 15.92% 379.54 15.11% 362.31 2026-03-05 14.25% 336.66 12.75% 301.27 12.37% 292.26 14.06% 332.20 17.03% 402.42 12.23% 290.56 2026-04-28 20.43% 388.59 19.54% 371.61 20.23% 384.60 21.66% 411.86 29.08% 553.03 20.86% 396.73 2026-05-04 23.19% 534.31 16.59% 382.22 16.47% 379.51 17.25% 397.43 19.30% 444.67 19.84% 457.16 2026-05-10 16.17% 391.96 14.31% 346.96 13.93% 337.68 14.39% 348.90 14.68% 355.83 11.73% 284.37 2026-05-16 11.51% 280.86 10.82% 263.98 9.52% 232.32 12.55% 306.24 10.80% 263.47 8.03% 195.89 2026-05-19 10.23% 250.58 11.14% 273.05 10.10% 247.35 10.80% 264.61 13.27% 325.10 8.22% 201.46 2026-01-13 10.02% 236.73 10.88% 257.11 8.46% 199.75 9.10% 214.96 13.12% 309.89 7.49% 176.95 2026-03-23 11.77% 272.90 11.33% 262.64 10.98% 254.70 11.41% 264.70 13.26% 307.57 8.34% 193.35 Table 7. Commerce window-level results, baseline group B and ReasonCast. Each method occupies a WMAPE (%) and MAE pair. TS backb. + covariates Chronos Time-LLM VoT ReasonCast Window WMAPE MAE WMAPE MAE WMAPE MAE WMAPE MAE WMAPE MAE WMAPE MAE 2025-11-11 18.04% 504.26 17.74% 495.69 17.67% 493.82 20.11% 561.93 22.86% 638.72 17.18% 480.21 2025-12-08 13.31% 336.44 11.61% 293.52 16.84% 425.55 16.42% 415.09 15.13% 382.42 11.25% 284.27 2025-12-11 12.30% 307.93 10.69% 267.51 10.20% 255.44 14.38% 359.93 12.51% 313.13 10.17% 254.56 2026-01-01 24.92% 596.10 27.36% 654.40 27.72% 663.06 15.69% 375.23 19.89% 475.70 22.29% 533.21 2026-02-12 131.66% 1008.59 116.26% 890.61 141.52% 1084.17 210.34% 1611.33 108.75% 833.09 83.85% 642.36 2026-02-15 35.95% 293.38 28.77% 234.77 33.26% 271.42 136.70% 1115.56 42.28% 345.02 24.53% 200.21 2026-03-02 11.09% 264.51 10.82% 258.05 14.24% 339.45 16.67% 397.43 16.19% 386.05 10.65% 253.93 2026-03-05 9.03% 213.45 8.88% 209.76 10.79% 254.92 15.62% 369.00 12.63% 298.44 8.78% 207.43 2026-04-28 14.44% 274.63 13.13% 249.78 20.35% 386.90 20.34% 386.70 17.95% 341.31 12.52% 238.09 2026-05-04 17.88% 411.98 16.83% 387.78 20.32% 468.08 19.92% 458.93 18.32% 422.06 16.34% 376.43 2026-05-10 15.42% 373.72 14.70% 356.27 13.05% 316.35 13.34% 323.52 16.22% 393.13 13.76% 333.46 2026-05-16 9.78% 238.51 10.01% 244.12 7.56% 184.29 10.45% 254.90 11.48% 280.11 9.07% 221.16 2026-05-19 8.69% 213.02 9.10% 223.05 8.23% 201.68 11.20% 274.46 9.58% 234.69 8.82% 216.14 2026-01-13 8.43% 199.07 8.43% 199.24 6.91% 163.30 15.35% 362.56 10.59% 250.25 9.60% 226.74 2026-03-23 10.72% 248.52 9.26% 214.74 9.74% 225.93 11.33% 262.80 12.03% 278.88 9.58% 222.12 Table 8. M5 window-level results, baseline group A. Each method occupies a WMAPE (%) and MAE pair. SNaive LightGBM PatchTST iTransformer TimesFM C2-small Event and origin WMAPE MAE WMAPE MAE WMAPE MAE WMAPE MAE WMAPE MAE WMAPE MAE Halloween 10-25 37.20% 4.55 29.25% 3.58 45.65% 5.59 29.45% 3.61 29.84% 3.65 29.29% 3.59 Halloween 10-31 38.36% 4.86 30.46% 3.86 44.79% 5.68 30.71% 3.89 30.17% 3.82 29.38% 3.72 Thanksgiving 11-21 46.84% 5.45 35.41% 4.12 53.78% 6.26 36.71% 4.27 37.91% 4.41 36.64% 4.27 Thanksgiving 11-24 50.78% 5.46 37.95% 4.08 56.76% 6.11 41.32% 4.44 41.86% 4.50 41.30% 4.44 Christmas 12-21 59.59% 6.17 49.74% 5.15 68.18% 7.06 49.61% 5.13 50.30% 5.21 49.69% 5.14 Christmas 12-24 64.00% 6.18 51.91% 5.01 71.71% 6.92 52.24% 5.04 53.95% 5.21 54.07% 5.22 NewYear 12-27 52.48% 6.19 35.05% 4.14 49.38% 5.83 34.11% 4.02 33.76% 3.98 33.91% 4.00 NewYear 12-30 49.21% 6.25 33.68% 4.28 46.54% 5.91 32.62% 4.14 32.30% 4.10 32.79% 4.16 SuperBowl 02-01 37.92% 5.22 29.63% 4.07 42.02% 5.78 30.21% 4.15 29.82% 4.10 30.04% 4.13 SuperBowl 02-07 40.13% 5.72 30.98% 4.41 41.19% 5.87 31.03% 4.42 30.65% 4.37 31.17% 4.44 Easter 03-23 37.29% 4.97 29.70% 3.96 41.53% 5.54 29.23% 3.90 29.57% 3.94 29.89% 3.98 Easter 03-26 37.90% 5.02 28.55% 3.78 41.63% 5.51 29.50% 3.91 30.03% 3.98 30.20% 4.00 Mothers day 05-04 36.55% 5.31 27.77% 4.03 36.62% 5.32 27.95% 4.06 28.01% 4.07 27.26% 3.96 Mothers day 05-07 35.45% 5.19 26.71% 3.91 36.58% 5.35 26.82% 3.93 26.65% 3.90 25.95% 3.80 Table 9. M5 window-level results, baseline group B and ReasonCast. Each method occupies a WMAPE (%) and MAE pair. TS backb. + covariates Chronos Time-LLM VoT ReasonCast Event and origin WMAPE MAE WMAPE MAE WMAPE MAE WMAPE MAE WMAPE MAE WMAPE MAE Halloween 10-25 29.59% 3.62 29.24% 3.58 30.50% 3.73 35.20% 4.31 33.53% 4.11 28.83% 3.53 Halloween 10-31 30.52% 3.87 29.76% 3.77 30.55% 3.87 35.21% 4.46 35.73% 4.53 29.52% 3.74 Thanksgiving 11-21 33.77% 3.93 33.46% 3.89 37.69% 4.39 42.81% 4.98 38.81% 4.52 33.81% 3.94 Thanksgiving 11-24 35.85% 3.86 36.40% 3.92 41.53% 4.47 45.59% 4.90 39.83% 4.28 35.71% 3.84 Christmas 12-21 47.10% 4.87 47.72% 4.94 50.25% 5.20 54.21% 5.61 53.72% 5.56 48.34% 5.00 Christmas 12-24 48.43% 4.67 48.65% 4.69 54.20% 5.23 57.83% 5.58 52.18% 5.04 48.46% 4.68 NewYear 12-27 33.74% 3.98 33.14% 3.91 35.64% 4.21 37.86% 4.47 37.64% 4.44 32.49% 3.83 NewYear 12-30 32.65% 4.14 32.59% 4.14 34.97% 4.44 38.05% 4.83 39.46% 5.01 31.53% 4.00 SuperBowl 02-01 29.33% 4.03 29.35% 4.04 31.00% 4.26 34.26% 4.71 35.39% 4.87 29.18% 4.01 SuperBowl 02-07 31.36% 4.47 30.78% 4.38 31.36% 4.47 33.89% 4.83 36.16% 5.15 30.69% 4.37 Easter 03-23 29.87% 3.98 30.21% 4.03 30.40% 4.05 34.76% 4.63 33.63% 4.48 29.28% 3.90 Easter 03-26 29.18% 3.87 29.07% 3.85 30.63% 4.06 34.19% 4.53 34.63% 4.59 28.62% 3.79 Mothers day 05-04 27.36% 3.97 27.08% 3.93 28.64% 4.16 31.35% 4.55 31.25% 4.54 26.87% 3.90 Mothers day 05-07 26.19% 3.83 25.98% 3.80 27.96% 4.09 30.69% 4.49 30.72% 4.50 25.54% 3.74 A.3. Semantic Corruption Stress Test Table 10 gives the per-window results for the controlled semantic inversion test. We invert the six control fields and the final trend narrative while holding the time-series input, checkpoint, and inference configuration fixed. Positive Ī denotes degradation. Incorrect semantics increase WMAPE in 13 of 15 windows, by 2.67 percentage points on average. The distribution is strongly right-skewed. The median increase is 0.76 percentage points, while the February 12 Spring FestivalāValentine window increases by 25.33 points. Thus, a single aggregate mean would understate the typical harm while obscuring the severe failure mode on a highly event-sensitive window. The two small negative deltas are also informative: gating attenuates many incorrect interventions, but it cannot guarantee that every corrupted description worsens every finite test slice. Table 10. Semantic inversion by window. Ī is inverted minus correct WMAPE in percentage points. Event Window Correct WMAPE Inverted WMAPE Ī Correct MAE Inverted MAE 520 2026-05-16 7.38% 7.77% 0.39 252.36 265.70 520 2026-05-19 5.78% 6.47% 0.69 218.79 244.81 Double 11 2025-11-05 13.27% 14.03% 0.76 154.35 163.25 Double 11 2025-11-08 17.54% 19.15% 1.61 959.36 1047.42 Double 11 2025-11-11 13.76% 13.62% -0.15 618.88 612.33 Double 12 2025-12-08 9.89% 11.09% 1.19 163.66 183.39 Double 12 2025-12-11 9.77% 12.15% 2.39 184.16 229.16 Labor Day 2026-04-28 7.61% 8.74% 1.14 175.76 202.01 Labor Day 2026-05-04 13.98% 14.71% 0.74 457.55 481.63 Lantern 2026-03-02 8.42% 9.61% 1.19 225.67 257.55 Motherās Day 2026-05-10 12.35% 12.89% 0.53 403.99 421.43 New Year 2026-01-01 15.13% 15.09% -0.04 403.91 402.86 SpringāValentine 2026-02-12 79.37% 104.69% 25.33 943.54 1244.66 SpringāValentine 2026-02-15 21.21% 24.95% 3.74 286.61 337.10 Womenās Day 2026-03-05 7.66% 8.16% 0.50 206.44 219.89 Appendix B Agent Policy, Semantic Interpretation, and Optimization Details B.1. Agent Policy and Basic Semantic Interpreter Prompt We explicitly separate the agent policy from the semantic interpreter. Given the event context, the no-text forecast, and its predictive entropy, the policy selects Skip, Basic, or Tool. Skip terminates semantic processing and preserves the numerical backbone. Basic invokes the semantic interpreter with the available context. Tool first augments that context with retrieved event evidence or temporal statistics and then invokes the same interpreter. Holidays and mega-sales disallow Skip; other instances may abstain when the backbone is sufficiently reliable. The fixed procedure below is therefore not the agent policy. We retain it as the Basic semantic interpreter prompt, invoked only after a non-skip policy decision. The Basic route supplies the original context, whereas the Tool route supplies tool-augmented context. All scale judgments are expressed relative to the most recent seven-day mean. Input: item hierarchy; forecast dates; event calendar and event positions; recent seven-day demand, mean, and slope; no-text forecast and entropy; current level C; aligned historical mean A and peak ApeakA_peak when available; platform activities; and any policy-approved tool observations. 1: Validate the reference scale. If aligned history exists, compute A/CA/C for the seven-day mean effect and Apeak/CA_peak/C for the single-day peak; otherwise activate the no-YoY branch. 2: Assess itemāevent relevance from historical evidence, item knowledge, recent slope, and event position. Do not infer relevance solely from the presence of an event. 3: Select the dominant event when multiple events or promotions overlap, and record why weaker candidates are excluded. 4: Assign demand direction and temporal shape, explicitly separating the seven-day mean from a concentrated single-day peak. 5: Calibrate amplitude and peak multipliers to the recent seven-day reference, using conservative priors when aligned history is absent. 6: Check cross-field consistency: relevance, direction, shape, amplitude, peak, and peak timing must describe the same trajectory. Return: six structured control fields plus one concise trend narrative. Never directly generate the numerical demand forecast. Figure 3. Basic semantic interpreter prompt used after a non-skip policy decision. The tool route augments its input evidence but does not change the output schema.A fixed semantic interpreter validates the reference scale, assesses event relevance, selects a dominant event, predicts direction and temporal shape, calibrates magnitude and peak effects, checks consistency, and returns structured semantic fields. This separation also defines the roles of the three post-training stages. Schema SFT and semantic-field RL train the semantic interpreter to produce valid, consistent, and calibrated forecast-specific fields. Forecast-utility RL trains the agent policyāincluding route selection, tool use, and the final intervention decisionāagainst realized downstream utility from the frozen forecaster, with explicit penalties for negative transfer and tool cost. B.2. Structured Fields and Representative Rationale Every final policy response begins with one machine-readable routing action. A Skip response terminates there and activates the no-text path. Basic and Tool responses additionally contain six machine-readable control fields and a short natural-language trend narrative. Tool calls and observations belong to the intermediate policy trajectory rather than the final output schema. The narrative is encoded together with the fields, but it is not counted as an additional control label because it has no closed label set. Table 11. ReasonCast reasoner output schema. Field Value space and forecasting role Intervention action Skip, basic, or tool. Controls whether semantic generation is bypassed, uses available context, or invokes additional evidence/statistics. Relevance Related / unrelated. Determines whether the context is a plausible intervention rather than incidental calendar text. Dominant event Event name / none. Resolves overlapping holidays and promotions to a primary driver. Direction Large decrease, small decrease, flat, small increase, or large increase; describes the seven-day mean relative to the recent mean. Temporal shape Increasing, decreasing, riseāfall, fallārise, or flat; specifies the within-horizon trajectory and turning behavior. Amplitude Positive scalar; ratio of the predicted seven-day mean to the recent seven-day mean. Peak Positive scalar; ratio of the largest predicted day to the recent seven-day mean. Trend narrative At most 80 Chinese characters; states event timing, turning point, and coarse magnitude for semantic encoding. Table 12 shows a representative teacher example. The stored compressed rationale is Chinese; an English translation is shown for readability. It is a concise forecasting rationale, not a raw token-by-token trace of the teacherās internal generation. Table 12. Representative compressed rationale and structured output for a laptop forecast from March 7 to March 13, with Womenās Day on horizon day 2. Compressed rationale (English translation) The window covers the March 8 platform campaign on day 2. Laptops are not a core Womenās Day gift category, but they respond to platform-wide promotions. Recent demand is already elevated, so the additional effect should be moderate rather than comparable with Double 11 or 618. We therefore expect a short peak of about 1.4Ć1.4Ć on March 8 followed by a return toward the recent level, producing a seven-day mean of about 1.2Ć1.2Ć. Structured target Relevance: related Dominant event: Womenās Day Direction: small increase Shape: riseāfall Amplitude: 1.2 Peak: 1.4 Trend: promotion-driven peak on March 8, followed by a return toward the recent level. B.3. Teacher Distillation and Rationale Compression We construct the initial SFT pool by stratifying 40,000 itemāwindow examples across upward (40%), downward (22%), flat (33%), and noisy (5%) demand regimes. Claude Opus 4.6 serves as the teacher and receives only the Basic semantic interpreter prompt, without agent routing or tool interaction. This stage deliberately supervises structured semantic construction rather than the agent policy. Of 11,627 successful teacher generations, deterministic filters on direction, amplitude, and peak consistency retain 7,312 candidates. The teacher rationales are substantially longer than needed by the student model. We therefore translate the English rationales into Chinese and compress them to a single 200ā300-character paragraph while preserving the original decision order: item and window, event position, recent baseline and C, historical A and ApeakA_peak or the no-YoY branch, itemāevent relevance, dominant-event selection, direction and shape, magnitude and peak, and any necessary recalibration. Numerical evidence may only be copied from the prompt or teacher response. The six structured labels are held fixed during compression. After compression validation, deduplication, and rebalancing, the final SFT set contains 4,590 examples. This data transformation shortens the student rationale without discarding the forecasting decision pattern. As shown in Table 5, mean rationale length decreases from 568.99 to 206.61 tokens after SFT and to 202.77 tokens after semantic-field RL, a 64.36% reduction relative to the original LLM. Direction, amplitude, and peak errors simultaneously decline. Because Table 5 reports cumulative post-training stages rather than an isolated compression ablation, we interpret this as an end-to-end distillation result rather than attributing the quality gain to compression alone. B.4. SFT, Semantic-Field RL, and Forecast-Utility RL Configuration Tables 13 and 14 record the main optimization settings. Both RL stages use group-relative policy optimization (GRPO). Semantic-field RL improves the forecast-critical fields; forecast-utility RL first applies a semantic validity gate and then evaluates only gate-passing candidates with the frozen fusion model. Table 13. SFT and semantic-field RL configuration. Component Configuration Student model Qwen3-32B (37) for SFT and both GRPO stages. SFT data 4,590 compressed teacher examples; 2% validation split. SFT adapter LoRA on all linear layers, rank 16, scale 32, dropout 0.05. SFT optimization Three epochs; learning rate 1Ć10ā41Ć 10^-4; cosine schedule; warmup ratio 0.03; BF16; effective batch size 32 (eight workers, batch size 1, four accumulation steps). Semantic-field adapter New LoRA initialized on the SFT-merged model; rank 8, scale 16, dropout 0; all linear layers. Semantic-field sampling Group size 8; temperature 0.8; top-p 0.95; top-k 50; maximum prompt/completion lengths 2816/512 tokens. Semantic-field optimization 1,500 updates (one pass over the consumed subset); learning rate 5Ć10ā65Ć 10^-6; cosine schedule; warmup ratio 0.03; KL coefficient 0.02 against the SFT policy. Semantic-field reward Direction/shape/amplitude/peak weights 0.20/0.20/0.30/0.30. Unparseable outputs receive ā1.5-1.5; additional penalties enforce peakāamplitude, directionāamplitude, and narrative consistency. Table 14. Forecast-utility RL (GRPO) configuration. Component Configuration Forecast-utility data and sampling 5,500 prompts, one epoch (1,834 updates); group size 8; the same maximum prompt/completion lengths and rollout sampling as semantic-field RL. Forecast-utility optimization Learning rate 2Ć10ā62Ć 10^-6; cosine schedule; warmup ratio 0.03; global KL coefficient 0.02 against the frozen semantic-field policy, plus a prompt-conditional SFT KL coefficient in [0,0.004][0,0.004]. Forecast-utility routing and gate Sample skip/basic/tool with each rollout; disallow skip for holidays and mega-sales. Skip bypasses field parsing and uses the cached no-text forecast. For non-skip routes, require field reward Aiā„maxā”(Airefā0.08,0.10)A_iā„ (A_i^ref-0.08,0.10) and reject hard cross-field contradictions before invoking the frozen fusion model. Forecast utility Compare candidate MAE with the cached semantic-field-policy fusion MAE and the no-text TSFM MAE; penalize harm beyond 1.051.05 times the no-text MAE. Utility is clipped on a 0.25 relative-error scale. Its weight is 0.5 in the reference run and 1.5 after the fixed step-500 diagnostic; a normalized route-cost penalty makes tool use more expensive than basic reasoning and assigns zero cost to skip.