Paper deep dive
RiskTraf: Risk-Extrapolated Residual Learning for Multi-Variate Traffic Flow Prediction
Guangyu Wang, Zhidan Liu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/24/2026, 4:59:04 AM
Summary
The paper introduces RiskTraf, a model-agnostic risk-extrapolated residual plug-in for multi-variate traffic flow prediction, and PEMSB-3V, a public benchmark suite preserving raw flow, speed, and occupancy measurements from PeMS detectors. RiskTraf addresses regime-dependent shortcut correlations by freezing a trained spatio-temporal backbone and learning a lightweight residual head from historical speed and occupancy, optimized with a risk extrapolation objective to mitigate distribution shifts between free-flow and congested states.
Entities (13)
Relation Signals (9)
PEMSB-3V → contains → flow
confidence 95% · PEMSB-3V, a public benchmark suite that preserves raw flow, speed, and occupancy measurements
PEMSB-3V → contains → speed
confidence 95% · PEMSB-3V, a public benchmark suite that preserves raw flow, speed, and occupancy measurements
PEMSB-3V → contains → occupancy
confidence 95% · PEMSB-3V, a public benchmark suite that preserves raw flow, speed, and occupancy measurements
PEMSB-3V → derivedfrom → PeMS
confidence 95% · preserves raw flow, speed, and occupancy measurements from PeMS detectors
RiskTraf → uses → Residual Head
confidence 95% · RiskTraf freezes the selected checkpoint and learns a lightweight zero-start residual head from historical speed and occupancy.
RiskTraf → improves → Forecasting Backbones
confidence 90% · RiskTraf consistently improves diverse forecasting backbones
RiskTraf → mitigates → regime-specific shortcut correlations
confidence 90% · thereby mitigating regime-specific shortcut correlations without modifying the backbone.
RiskTraf → outperforms → debiasing methods
confidence 85% · outperforms debiasing and distribution-shift adaptation methods.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Traffic sensors commonly record flow, speed, and occupancy, but standard traffic flow forecasting benchmarks and models rarely exploit all three raw measurements reliably. Although speed and occupancy provide sensor-native traffic-state information beyond flow alone, existing releases often omit these variables, replace them with proxies, or contain logically inconsistent records. Moreover, direct empirical risk minimization over three-variable inputs may exploit regime-dependent shortcuts, as the relationships among flow, speed, and occupancy vary substantially between free-flow and congested states. We introduce \textbf{PEMSB-3V}, a public benchmark suite that preserves raw flow, speed, and occupancy measurements from PeMS detectors for flow prediction. We also propose \textbf{RiskTraf}, a model-agnostic risk-extrapolated residual plug-in. For each trained spatio-temporal backbone, RiskTraf freezes the selected checkpoint and learns a lightweight zero-start residual head from historical speed and occupancy. The residual head constructs ordered traffic-risk environments and optimizes horizon-wise flow corrections with a risk extrapolation objective, thereby mitigating regime-specific shortcut correlations without modifying the backbone. Extensive experiments demonstrate that RiskTraf consistently improves diverse forecasting backbones and outperforms debiasing and distribution-shift adaptation methods. Our code and benchmark are available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.20656v1
- Canonical: https://arxiv.org/abs/2608.20656v1
Trouble viewing inline? Open PDF directly →
Full Text
67,865 characters extracted from source content.
Expand or collapse full text
RiskTraf: Risk-Extrapolated Residual Learning for Multi-Variate Traffic Flow PredictionConference: Proceedings of the 35th ACM International Conference on Information and Knowledge Management; November 7–11, 2026; Rome, Italy.Proceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM ’26), November 7–11, 2026, Rome, ItalyISBN: 979-8-4007-2539-5/2026/11DOI: 10.1145/3799682.3840707CCS: Computing methodologies Spatial and temporal reasoningCCS: Computing methodologies Neural networksCCS: Information systems Data mining Guangyu Wang email: hulnegy@gmail.com Affiliation: Dongbei University of Finance & Economics , Dalian , Liaoning , China and Zhidan Liu Note: Corresponding author. email: zhidanliu@hkust-gz.edu.cn Affiliation: Hong Kong University of Science and Technology (Guangzhou) , Guangzhou , Guangdong , China 2026; © c Abstract. Traffic sensors commonly record flow, speed, and occupancy, but standard traffic flow forecasting benchmarks and models rarely exploit all three raw measurements reliably. Although speed and occupancy provide sensor-native traffic-state information beyond flow alone, existing releases often omit these variables, replace them with proxies, or contain logically inconsistent records. Moreover, direct empirical risk minimization over three-variable inputs may exploit regime-dependent shortcuts, as the relationships among flow, speed, and occupancy vary substantially between free-flow and congested states. We introduce PEMSB-3V, a public benchmark suite that preserves raw flow, speed, and occupancy measurements from PeMS detectors for flow prediction. We also propose RiskTraf, a model-agnostic risk-extrapolated residual plug-in. For each trained spatio-temporal backbone, RiskTraf freezes the selected checkpoint and learns a lightweight zero-start residual head from historical speed and occupancy. The residual head constructs ordered traffic-risk environments and optimizes horizon-wise flow corrections with a risk extrapolation objective, thereby mitigating regime-specific shortcut correlations without modifying the backbone. Extensive experiments demonstrate that RiskTraf consistently improves diverse forecasting backbones and outperforms debiasing and distribution-shift adaptation methods. Our code and benchmark are available at https://github.com/Guangyu4/RiskTraf. Keywords: Traffic Flow Prediction, Multi-Variate Time Series Forecasting, Risk Extrapolation, Traffic Benchmark †c-license: by 1. Introduction Traffic flow prediction aims to forecast the number of vehicles passing through road segment stations over future time intervals based on historical observations. In widely used benchmarks such as PeMS (Freeway Performance Measurement System) (8), traffic flow is typically defined as the vehicle count recorded by loop detectors within fixed time windows, e.g., every 5 minutes. As a fundamental indicator of road utilization, accurate flow prediction plays a critical role in intelligent transportation systems (43), enabling traffic authorities to anticipate congestion, optimize signal control, and allocate transportation resources more effectively (35). Recent studies have made substantial progress in traffic forecasting by designing increasingly expressive spatio-temporal models, including graph neural networks (31; 45) and attention-based architectures (32). However, as the field matures, performance gains from purely architectural innovations have become increasingly marginal (40), while model complexity and computational costs continue to grow. This trend has motivated researchers to exploit richer contextual information for more accurate prediction. Existing efforts have incorporated external sources such as weather conditions (15; 30) and textual information (49). Nevertheless, incorporating such auxiliary sources often requires additional collection, storage, and maintenance, thereby increasing the data-side overhead of prediction systems (10). More importantly, they are difficult to align with traffic measurements: weather typically affects broad regions while sensors observe localized road segments (50), and textual signals such as news are updated at much coarser temporal granularity than traffic sensors (22). Such spatial and temporal mismatches may introduce additional noise during data fusion. Compared with external auxiliary sources, a more natural way to enrich traffic representations is to leverage multiple variables reported by the same traffic sensor. In addition to flow, loop detectors commonly record speed and occupancy, which describe complementary aspects of traffic states. Speed captures vehicle movement dynamics, while occupancy measures the fraction of time during which detectors are occupied by vehicles. Since these variables are collected from the same sensors and at the same temporal resolution as flow, they provide sensor-native auxiliary information without the spatial and temporal misalignment of external data. From an information-theoretic perspective (3), jointly modeling heterogeneous sensor variables can provide additional state information beyond any single variable. A few studies have therefore begun to exploit flow, speed, and occupancy for traffic forecasting (24; 39; 36). However, incorporating speed and occupancy into deep forecasting models is not a straightforward extension of flow-only prediction. Their effective use requires satisfying two coupled requirements: obtaining reliable three-variable measurements from the same set of sensors, and learning from these auxiliary variables without relying on regime-specific shortcuts. Challenge 1: Lack of reliable multi-variate benchmarks. Existing traffic sensor releases do not guarantee that flow, speed, and occupancy are simultaneously valid and usable. Raw loop-detector records often require strict screening based on validity thresholds and traffic-flow consistency (42), and field detectors may suffer from faults such as stuck-off, stuck-on, or hanging-on behaviors (7). Moreover, speed and occupancy are sensitive to detector configurations, lane definitions, loop lengths, and road types. As a result, a sensor that provides usable flow measurements may still be unreliable for raw three-variable modeling. This makes it difficult to systematically evaluate methods that exploit flow, speed, and occupancy under a standardized benchmark setting. Challenge 2: Regime-dependent auxiliary correlations. Although speed and occupancy contain useful traffic-state information, their relationships with flow are not invariant across traffic regimes. For example, under congested conditions, slower vehicles may occupy loop detectors for longer durations, increasing occupancy even when the observed flow is similar. Consequently, naively treating speed and occupancy as ordinary covariates may cause models to exploit regime-specific shortcuts rather than robust predictive patterns. Models trained with empirical risk minimization (ERM) can overfit such spurious correlations, leading to degraded generalization across traffic states. Simply discarding speed and occupancy avoids this risk but also wastes valuable sensor information. Existing distribution-shift adaptation (11; 16) and debiasing methods (53) partially address related issues, but they often rely on specialized shift or bias detectors, whose effectiveness is limited by the accuracy and availability of such detectors. To address these challenges, we introduce both a reliable benchmark suite and a risk-aware learning framework. First, we construct the PEMSB-3V Benchmark consisting of four datasets: PEMS03-B, PEMS04-B, PEMS07-B, and PEMS08-B. The benchmark is district-aligned with widely used PeMS datasets while preserving raw flow, speed, and occupancy measurements reported by PeMS detectors. By conducting sensor-level validation, PEMSB-3V provides a standardized testbed for studying historical three-variable inputs with flow-only forecasting targets. Second, we propose RiskTraf, a risk extrapolation plug-in for traffic prediction. Rather than using speed and occupancy as unrestricted covariates or replacing existing forecasting architectures, RiskTraf treats them as historical auxiliary signals for residual correction. Specifically, given a trained three-variable backbone, we freeze its parameters and attach a lightweight zero-initialized residual head driven by historical speed and occupancy. The residual head learns horizon-wise flow corrections across ordered low-to-high traffic-risk environments, encouraging robust improvements while reducing reliance on regime-specific auxiliary correlations. RiskTraf performs rollback when validation MAE does not improve. The main contributions of this work are summarized as follows: • We release PEMSB-3V, a district-aligned benchmark suite that enables standardized evaluation of flow prediction with historical flow, speed, and occupancy measurements. • We formulate regime-dependent auxiliary correlations as a key obstacle in multi-variate traffic forecasting and propose RiskTraf, a REx-based residual plug-in that improves trained backbones without replacing their architectures. • We conduct extensive experiments on multiple PEMSB-3V datasets and diverse spatio-temporal backbones, showing that RiskTraf achieves consistent improvements over strong baselines, debiasing methods, and distribution-shift adaptation approaches. The rest of this paper is organized as follows. Section 2 introduces the PEMSB-3V benchmark. Section 3 presents the RiskTraf framework. Section 4 reports experimental results and analyses. Section 5 discusses related work, and Section 6 concludes the paper. 2. PEMSB-3V Benchmark We construct PEMSB-3V, a benchmark suite consisting of four district-level datasets: PEMS03-B, PEMS04-B, PEMS07-B, and PEMS08-B. The suite is built from the California Department of Transportation Performance Measurement System (PeMS)11 1 https://pems.dot.ca.gov/ and its official source page22 2 https://dot.ca.gov/programs/traffic-operations/mpr/pems-source. PeMS is described by Caltrans as a statewide freeway monitoring system with nearly 40,000 detectors. Unlike existing releases that replace missing variables with proxies or temporal codes, PEMSB-3V retains only detectors whose native PeMS records contain all three raw measurements: flow, speed, and occupancy. The overall construction pipeline is summarized in Algorithm 1. Algorithm 1 PEMSB-3V Benchmark Construction Pipeline 0: District ID, time range [tstart,tend][t_start,t_end], completeness threshold ρ 0: Multi-variate traffic benchmark D and adjacency matrix A 1: Query detector metadata from PeMS 2: Keep only detectors with raw flow, speed, and occupancy channels 3: Filter detectors whose temporal completeness is below ρ 4: for each day in [tstart,tend][t_start,t_end] do 5: Download 5-minute flow, speed, and occupancy records 6: Remove malformed timestamps and interpolate only short gaps 7: end for 8: Aggregate the cleaned records into ∈ℝT×N×3X ^T× N× 3 9: Construct the road topology and compute pairwise distances 10: Build A with a distance kernel and sparsification threshold 11: Split the benchmark into train/val/test by ratio 6:2:2 12: return D, A Figure 1. High-quality PEMS03-B sensors after metadata-based screening. The map shows 1,013 merged mainline/HOV detectors with complete metadata.Map of high-quality PEMS03-B freeway mainline and HOV detector locations after metadata-based screening. To ensure that the benchmark is built from sensors with physically interpretable three-variable measurements, we apply a metadata-based progressive screening procedure before constructing the final sensor set. The screening follows the earliest-failed-check rule. We first perform an ID coverage check. A target sensor without a corresponding PeMS metadata record is marked as Missing Metadata, since its static information cannot be verified. We then perform a spatial localization check. If latitude or longitude is missing, invalid, or outside the legal geographic range, the detector is marked as Missing Geolocation and excluded from map-based topology construction. Next, we check static structural completeness using detector length and lane count. Records with empty, non-numeric, or non-positive values are marked as Invalid Static Attributes, because flow, speed, and occupancy depend on the detector segment and lane definition. Finally, we check road-type consistency. Detectors with Type∈ML,HVType∈\ML,HV\ are retained as freeway mainline or HOV sensors, whereas Type∈OR,FR,FType∈\OR,FR,F\ are treated as ramp or connector sensors and excluded from the mainline subset. Figure 1 illustrates the resulting PEMS03-B sensor footprint after metadata-based screening. Among 1,708 target sensor IDs, 41 are removed due to missing metadata and 2 due to invalid or missing geolocation. Among the remaining 1,665 geolocated records, 1,039 pass all metadata checks as complete mainline/HOV sensors, while 626 are identified as ramp or connector sensors. After merging duplicate records with the same location, road direction, and detector type, the final map contains 1,013 high-quality points, including 762 mainline and 251 HOV points. This procedure makes each discarded sensor traceable to an interpretable failure category, rather than removing sensors through opaque filtering. The subset design follows the administrative organization of PeMS rather than an arbitrary partition. Specifically, PEMS03-B, PEMS04-B, PEMS07-B, and PEMS08-B correspond to Caltrans districts 03, 04, 07, and 08, respectively33 3 https://cwwp2.dot.ca.gov/documentation/district-map-county-chart.htm. This preserves geographic provenance while maintaining comparability with widely used PeMS benchmarks. Table 1. Availability and usability of raw flow, speed, and occupancy measurements in representative traffic benchmarks. ✓: usable raw measurement; ✗: not provided; ○ : provided but logically inconsistent. Source Benchmark Nodes Interval Flow Speed Occupancy DCRNN (31) METR-LA 207 5 mins ✗ ✓ ✗ PEMS-BAY 325 5 mins ✗ ✓ ✗ LSTNet(29) Traffic 862 1 hour ✓ ○ ○ STGCN(48) PEMSD7(M) 228 5 mins ✓ ○ ○ PEMSD7(L) 1,026 5 mins ✓ ○ ○ ASTGCN(21) PEMSD4-I 228 5 mins ✓ ○ ○ PEMSD8-I 1,979 5 mins ✓ ○ ○ STSGCN(41) PEMS03 358 5 mins ✓ ○ ○ PEMS04 307 5 mins ✓ ○ ○ PEMS07 883 5 mins ✓ ○ ○ PEMS08 170 5 mins ✓ ○ ○ LargeST(33) CA 8,600 5 mins ✓ ✗ ✗ GLA 3,834 5 mins ✓ ✗ ✗ GBA 2,352 5 mins ✓ ✗ ✗ SD 716 5 mins ✓ ✗ ✗ XXLTraffic(47) Full_PEMS03 1,809 5 mins ✓ ○ ○ Full_PEMS04 4,089 5 mins ✓ ✗ ✗ Full_PEMS05 573 5 mins ✓ ✗ ✗ Full_PEMS06 705 5 mins ✓ ✗ ✗ Full_PEMS07 4,888 5 mins ✓ ✗ ✗ Full_PEMS08 2,059 5 mins ✓ ✗ ✗ Full_PEMS10 1,378 5 mins ✓ ✗ ✗ Full_PEMS11 1,440 5 mins ✓ ✗ ✗ Full_PEMS12 2,587 5 mins ✓ ✗ ✗ tfNSW 27 60 mins ✓ ✗ ✗ TraffiDent (19) TraffiDent∗ 16,972 5 mins ✓ ✓ ✓ PEMSB-3V PEMS03-B 1,013 5 mins ✓ ✓ ✓ PEMS04-B 2,474 5 mins ✓ ✓ ✓ PEMS07-B 2,788 5 mins ✓ ✓ ✓ PEMS08-B 1,515 5 mins ✓ ✓ ✓ ∗TraffiDent provides all three variables but does not apply sensor-type filtering or provide sufficient metadata to distinguish detector types. We further audit representative public traffic benchmarks to examine whether they support the raw flow-speed-occupancy forecasting setting considered in this paper. As shown in Table 1, most existing benchmarks either omit speed and occupancy, replace them with proxy variables, or contain logically inconsistent records. For example, PEMS03 and PEMS07 use daily and weekly temporal encodings in place of raw speed and occupancy, while PEMS04 and PEMS08 include physically implausible speed records, such as positive speed when flow is zero and truncated low-speed ranges. These issues prevent existing benchmarks from serving as reliable testbeds for studying how raw speed and occupancy contribute to flow prediction. 3. Methodology In this section, we present RiskTraf, a model-agnostic risk-extrapolation plug-in for multi-variate traffic forecasting. As illustrated in Figure 2, RiskTraf adopts a paired two-stage design without modifying the forecasting backbone: • Stage I trains a given standard spatio-temporal backbone with 12 historical steps of traffic flow, speed, and occupancy to predict future flow. • Stage I freezes the trained backbone and attaches a lightweight residual head driven by historical speed and occupancy. The residual head is optimized with a risk extrapolation objective to learn robust flow corrections across traffic-risk environments. Figure 2. Overview of RiskTraf. Stage I trains a backbone to predict future flow from historical flow, speed, and occupancy. Stage I freezes the backbone, constructs traffic-risk environments from historical speed and occupancy, and learns a lightweight residual head with risk extrapolation to correct the baseline prediction. 3.1. Problem Formulation and Paired Protocol Let F, S, and O denote traffic flow, speed, and occupancy, respectively. For a sample ending at time t−1t-1, the historical input is (1) Xt=Ft−T:t−1,St−T:t−1,Ot−T:t−1,X_t=\F_t-T:t-1,S_t-T:t-1,O_t-T:t-1\, where T is the input length and each variable is observed over all sensor nodes. The prediction target is future flow only: (2) Yt=Ft:t+H−1,Y_t=F_t:t+H-1, where H is the prediction horizon. Thus, speed and occupancy are available only as historical auxiliary measurements. RiskTraf is applied under a paired protocol for each dataset-backbone pair. Let BϕB_φ be an arbitrary spatio-temporal backbone. The Stage-I baseline prediction is given by (3) Y^base=Bϕ(Xt). Y_base=B_φ(X_t). After selecting the baseline checkpoint by validation MAE, RiskTraf starts from that checkpoint, freezes the backbone parameters ϕφ, and learns a residual correction head gψg_ψ from historical speed and occupancy: (4) Y^risk=Y^base+sr⋅gψ(St−T:t−1,Ot−T:t−1), Y_risk= Y_base+s_r· g_ψ(S_t-T:t-1,O_t-T:t-1), where srs_r is a small residual scaling factor; during Stage I only the residual parameters ψ are updated. 3.2. Stage I: Baseline Backbone Training The first stage follows the standard empirical risk minimization (ERM) setting adopted by existing traffic forecasting models. Given a historical three-variable sequence X, the backbone BϕB_φ predicts the future flow sequence YtY_t. The optimization objective is (5) minϕ(X,Y)[ℓ(Bϕ(X),Y)], _φ\;E_(X,Y) [ (B_φ(X),Y) ], where ℓ is the masked mean absolute error (MAE) computed on inverse-scaled flow values. The checkpoint with the best validation MAE is selected as the reference checkpoint for both the vanilla backbone and the subsequent RiskTraf stage. 3.3. Stage I: Risk-Aware Residual Plug-in The second stage attaches RiskTraf to the trained backbone obtained in Stage I. Although the Stage-I backbone has access to all three variables, its ERM objective may exploit regime-specific correlations between speed, occupancy, and flow. RiskTraf therefore freezes the backbone prediction Y^base Y_base and learns only a constrained horizon-wise residual correction, which keeps the plug-in model-agnostic and preserves the spatio-temporal representations learned by the backbone. RiskTraf consists of three components: a zero-start residual head, risk environment construction from historical covariates, and a REx-based objective for robust residual learning. 3.3.1. Zero-Start Residual Head RiskTraf uses a lightweight residual head to correct the fixed backbone prediction. Since traffic sensors are not exchangeable, the same speed–occupancy pattern may imply different flow corrections at different road segments due to location, lane configuration, detector type, and upstream–downstream context. Motivated by recent studies on spatio-temporal and heterogeneity-aware embeddings (32; 14), RiskTraf introduces a minimal node-identity embedding to provide node-specific context without modifying the frozen backbone. Specifically, RiskTraf maintains a trainable node-identity embedding table E=[e1,…,eN]⊤∈ℝN×deE=[e_1,…,e_N] ^N× d_e, where ene_n denotes the embedding of sensor node n. The table is initialized at the beginning of Stage I and optimized together with ψ; it is not copied from the backbone and requires no external sensor metadata. In addition, the head uses node-level historical summaries Q, including mean speed, mean occupancy, and occupancy-minus-speed over the input window. The residual correction is computed as (6) ΔY^=gψ(St−T:t−1,Ot−T:t−1,E,Q), Y=g_ψ(S_t-T:t-1,O_t-T:t-1,E,Q), where E and Q are broadcast to the batch and horizon dimensions as needed. The last linear layer of gψg_ψ is initialized to zero, so RiskTraf initially reproduces the baseline prediction and any nonzero correction must be learned during the debiasing stage. 3.3.2. Risk Environment Construction Risk environments are constructed only from historical speed and occupancy. For each sample in a training batch, we summarize the covariates over the historical window and all nodes: (7) v¯ v =1TN∑t,nvt,n,o¯=1TN∑t,not,n, = 1TN _t,nv_t,n, o= 1TN _t,no_t,n, where T is the input length and N is the number of sensor nodes. Inspired by mean alignment (18), we define a normalized traffic-risk score as (8) z(x)=x−μxσx,ρ=z(o¯)−z(v¯),z(x)= x- _x _x, ρ=z( o)-z( v), where μx _x and σx _x are estimated from the training data. A larger ρ indicates higher occupancy and lower speed, corresponding to more congested or high-risk traffic states. This ordering is grounded in the fundamental diagram of traffic flow, where congestion is characterized jointly by low speed and high occupancy, so ρ ranks samples along the physical free-flow–congestion axis. Samples are sorted by ρ and partitioned into K ordered environments as follows: (9) ei=(Xj,Yj)∣ρj∈quantilei,i=1,…,K.e^i=\(X_j,Y_j) _j _i\, i=1,…,K. This produces low-to-high traffic-risk environments without requiring incident annotations or manually defined congestion labels. 3.3.3. REx-Based Residual Learning Given the ordered risk environments, RiskTraf trains the residual head to improve flow prediction while keeping residual errors stable across traffic regimes. For each environment eie^i, we compute (10) ℒei=1|ei|∑(X,Y)∈eiℓ(Y^base+sr⋅gψ(S,O,E,Q),Y).L^e^i= 1|e^i| _(X,Y)∈ e^i ( Y_base+s_r· g_ψ(S,O,E,Q),Y). Let ℒ¯=1K∑iℒei L= 1K _iL^e^i. The REx penalty measures the variance of environment losses: (11) ℛREx=1K∑i=1K(ℒei−ℒ¯)2.R_REx= 1K _i=1^K(L^e^i- L)^2. The full Stage-I objective is (12) ℒtotal=ℒ¯+λτ(ℛREx+β⋅pair+η⋅extra),L_total= L+ _τ (R_REx+β·P_pair+η·P_extra ), where β and η control the relative weights of the two order-aware penalties, defined on the ordered environment losses as (13) pair _pair =1K−1∑i=1K−1[max(0,ℒei+1−ℒei)]2, = 1K-1 _i=1^K-1 [ (0,\,L^e^i+1-L^e^i ) ]^2, extra _extra =[max(0,ℒeK−sg(ℒe1))]2, = [ (0,\,L^e^K-sg(L^e^1) ) ]^2, where sg(⋅)sg(·) denotes the stop-gradient operator. pairP_pair penalizes any adjacent pair whose higher-risk environment incurs the larger residual loss, so that errors do not grow with the risk level, and extraP_extra penalizes the highest-risk loss when it exceeds the detached lowest-risk reference, pushing down only the former. The penalty weight is warmed up as (14) λτ=λ⋅min(1,max(0,τ−τwarmupτtotal−τwarmup)), _τ=λ· (1, (0, τ- _warmup _total- _warmup ) ), where τ denotes the current Stage-I epoch. This warmup allows the residual head to first learn useful corrections before enforcing stronger cross-environment consistency. 3.3.4. Validation Safeguard Since RiskTraf is designed as a plug-in correction, it should not degrade a strong backbone. We therefore evaluate the Stage-I checkpoint on the validation set and keep the corrected model only when it improves validation MAE over the frozen baseline. Otherwise, RiskTraf rolls back to the Stage-I prediction. This keeps the plug-in conservative: it improves the backbone when useful residual patterns exist and avoids harmful corrections otherwise. 3.4. Inference At inference, the selected Stage-I backbone first produces the baseline forecast Y^base Y_base. If the Stage-I residual head passes the validation safeguard, RiskTraf computes ΔY Y from historical speed and occupancy, node embeddings, and node-level summaries. The final prediction is (15) Y^risk=Y^base+sr⋅ΔY^. Y_risk= Y_base+s_r· Y. If the residual head is rolled back, the output remains the Stage-I baseline prediction. 4. Experiments We conduct experiments on the PEMSB-3V benchmark suite introduced in Section 2. The experiments are designed to evaluate whether the proposed benchmark provides physically meaningful three-variable traffic data, and whether RiskTraf can improve traffic flow prediction as a model-agnostic residual plug-in. We organize the evaluation around the following research questions: • RQ1: Data validity. Does PEMSB-3V exhibit physically meaningful flow–speed–occupancy relationships? • RQ2: Overall effectiveness. Can RiskTraf consistently improve diverse spatio-temporal forecasting backbones? • RQ3: Method comparison. How does RiskTraf compare with existing debiasing and distribution-shift adaptation methods? • RQ4: Component analysis. How do auxiliary variables, risk penalties, and environment granularity affect RiskTraf? • RQ5: Practicality and interpretability. What overhead does RiskTraf introduce, and what corrections does it learn under ambiguous traffic states? 4.1. Experimental Setup Datasets and forecasting setting. We evaluate RiskTraf on four PEMSB-3V subsets: PEMS03-B, PEMS04-B, PEMS07-B, and PEMS08-B. All experiments follow the same three-variable-to-flow forecasting setting. The input contains 12 historical time steps of flow, speed, and occupancy, corresponding to one hour of observations, and the target contains the next 12 time steps of flow. All metrics are computed on inverse-scaled flow values. We report MAE, RMSE, and MAPE, where lower values indicate better performance. Backbones. To evaluate model-agnostic applicability, we apply RiskTraf to representative spatio-temporal forecasting backbones from different architectural generations, including STGCN (48), DCRNN (31), AGCRN (4), Graph WaveNet (46), GMAN (51), GTS (38), STEMGNN (5), STNorm (13), STWA (17), MegaCRN (26), HimNet (14), and STDN (6). These backbones cover recurrent, graph-based, attention-based, normalization-based, and recent heterogeneity-aware designs. Paired plug-in protocol. For each dataset–backbone pair, we first train a vanilla three-variable backbone using standard empirical risk minimization. RiskTraf then starts from the same validation-selected checkpoint, freezes the backbone, and trains only the lightweight residual head driven by historical speed and occupancy. This paired protocol ensures that any improvement comes from the RiskTraf plug-in rather than from a different backbone initialization or training recipe. Table 2. Main plug-in results on PEMSB-3V, averaged over horizons 3, 6, and 12. Lower is better; MAPE is in percentage form. Green/red arrows mark improvement/degradation over the backbone row. Method PEMS03-B PEMS04-B PEMS07-B PEMS08-B MAE↓ RMSE↓ MAPE(%)↓ MAE↓ RMSE↓ MAPE(%)↓ MAE↓ RMSE↓ MAPE(%)↓ MAE↓ RMSE↓ MAPE(%)↓ STGCN 16.30 27.04 31.20 24.73 37.94 14.70 19.66 32.95 10.81 19.43 32.87 19.06 +RiskTraf 15.86↑ 26.27↑ 28.27↑ 24.23↑ 37.32↑ 13.77↑ 18.66↑ 32.29↑ 9.57↑ 18.75↑ 32.13↑ 17.86↑ DCRNN 47.38 71.99 203.44 56.39 80.63 41.98 110.81 141.90 141.00 53.82 77.69 78.09 +RiskTraf 25.72↑ 39.23↑ 140.30↑ 30.93↑ 45.61↑ 19.03↑ 25.59↑ 41.89↑ 13.37↑ 25.80↑ 39.82↑ 29.89↑ AGCRN 17.24 29.39 25.17 25.71 39.59 14.46 20.03 33.78 9.10 19.96 33.76 17.81 +RiskTraf 16.61↑ 27.65↑ 24.45↑ 25.27↑ 38.84↑ 14.17↑ 19.05↑ 32.77↑ 8.62↑ 19.31↑ 32.92↑ 17.14↑ GWNet 19.65 33.05 25.05 32.13 48.89 17.34 24.61 41.25 10.97 23.92 39.89 19.47 +RiskTraf 17.43↑ 29.07↑ 23.92↑ 27.75↑ 42.20↑ 15.35↑ 21.69↑ 36.84↑ 9.91↑ 21.22↑ 35.56↑ 18.22↑ GMAN 14.07 23.68 23.38 23.26 36.00 13.06 16.98 30.65 9.28 17.06 30.55 15.95 +RiskTraf 13.74↑ 23.46↑ 23.66↓ 22.12↑ 34.91↑ 12.43↑ 15.94↑ 29.79↑ 7.55↑ 16.43↑ 29.81↑ 15.43↑ STEMGNN 16.09 26.58 26.97 26.18 40.31 14.67 45.99 79.86 74.35 20.85 34.59 21.96 +RiskTraf 15.68↑ 25.80↑ 26.54↑ 25.29↑ 38.99↑ 14.18↑ 40.64↑ 64.59↑ 59.04↑ 19.73↑ 33.02↑ 19.40↑ STNorm 17.31 28.78 25.07 27.16 41.73 14.89 21.63 37.30 9.78 21.86 36.89 19.93 +RiskTraf 15.78↑ 26.10↑ 23.23↑ 25.20↑ 38.71↑ 14.18↑ 19.87↑ 34.68↑ 8.98↑ 20.51↑ 34.37↑ 19.20↑ GTS 28.45 50.08 49.64 41.99 62.02 24.72 44.76 70.24 24.12 37.53 57.01 63.22 +RiskTraf 21.45↑ 34.87↑ 46.49↑ 31.53↑ 47.18↑ 17.66↑ 28.32↑ 48.27↑ 16.00↑ 23.51↑ 39.50↑ 36.94↑ STWA 15.38 25.63 21.69 24.74 38.12 14.29 19.29 32.63 10.26 19.85 33.59 17.99 +RiskTraf 14.86↑ 24.79↑ 21.06↑ 24.15↑ 37.34↑ 13.66↑ 18.42↑ 31.96↑ 8.36↑ 18.94↑ 32.54↑ 17.47↑ MegaCRN 15.07 25.14 20.33 24.50 37.70 13.16 18.97 33.36 7.97 18.67 32.06 15.24 +RiskTraf 14.89↑ 24.82↑ 20.82↓ 23.98↑ 37.14↑ 12.99↑ 18.12↑ 32.44↑ 7.61↑ 18.22↑ 31.45↑ 15.43↓ HimNet 15.09 25.42 20.92 23.92 37.28 13.28 19.70 33.82 8.78 18.41 31.82 15.63 +RiskTraf 14.72↑ 24.79↑ 21.26↓ 23.46↑ 36.74↑ 13.10↑ 18.55↑ 32.36↑ 8.44↑ 17.73↑ 30.96↑ 15.57↑ STDN 13.78 22.67 33.80 22.56 35.17 12.54 15.86 29.58 7.36 16.59 30.34 15.61 +RiskTraf 13.24↑ 22.35↑ 20.36↑ 21.85↑ 34.64↑ 12.44↑ 14.95↑ 29.04↑ 6.71↑ 16.04↑ 29.59↑ 14.95↑ Implementation details. For RiskTraf, each training batch is sorted by the speed–occupancy risk score and partitioned into four ordered traffic-risk environments unless otherwise specified. The residual head is trained with masked MAE and the warmed-up risk extrapolation penalty described in Section 3.3. Baseline backbones and compared methods are implemented using official code when available or reproduced following the original papers, with hyper-parameters selected by validation MAE under the same protocol. We use Adam for optimization, validation MAE for model selection, and rollback to the Stage-I checkpoint when the plug-in does not improve validation MAE. The batch size is 32 unless otherwise specified. All experiments are conducted on NVIDIA L20 GPUs. Line plots showing that higher occupancy states correspond to lower speeds under similar traffic flow levels on PEMS03-B and PEMS04-B. Figure 3. Flow–speed relationships under different occupancy states on the four PEMSB-3V subsets. Samples are grouped into low-, mid-, and high-occupancy states; each curve shows the average speed at different flow levels.Line plots showing that higher occupancy states correspond to lower speeds under similar traffic flow levels on PEMS03-B and PEMS04-B. Stacked bar charts for PEMS03-B, PEMS04-B, PEMS07-B, and PEMS08-B showing baseline MAE, RiskTraf MAE, and reduced MAE across prediction horizons 1 to 12 for representative backbones. Figure 4. Per-horizon MAE decomposition for GTS, GWNet, and STNorm on the four PEMSB-3V subsets. For each horizon, the full bar denotes the baseline MAE, the solid segment denotes the MAE after adding RiskTraf, and the hatched cap denotes the absolute error reduced by RiskTraf.Stacked bar charts for PEMS03-B, PEMS04-B, PEMS07-B, and PEMS08-B showing baseline MAE, RiskTraf MAE, and reduced MAE across prediction horizons 1 to 12 for representative backbones. 4.2. Dataset Characteristics (RQ1) We first examine whether PEMSB-3V preserves meaningful relationships among flow, speed, and occupancy. Figure 3 plots flow–speed curves under low-, mid-, and high-occupancy states. Across the four displayed subsets, higher occupancy generally corresponds to lower speed, which is consistent with the loop-detector mechanism that vehicles occupy detectors for longer durations under congested conditions. For the same flow level, speed can vary substantially across occupancy states, indicating that flow alone is insufficient to identify the underlying traffic regime. The figure also shows that occupancy is not a trivial proxy for flow. High-occupancy states span a wide range of flow levels and exhibit distinct speed–flow patterns, rather than simply corresponding to the highest flow values. Moreover, the slope and shape of the flow–speed curves differ across occupancy groups and benchmark subsets, suggesting regime-dependent auxiliary correlations. These observations support the use of speed and occupancy as informative traffic-state indicators, while motivating RiskTraf to exploit them through risk-aware residual learning rather than direct unconstrained fusion. 4.3. Overall Performance (RQ2) Table 2 reports MAE, RMSE, and MAPE averaged over horizons 3, 6, and 12. For each backbone, the vanilla row denotes the Stage-I model trained with historical flow, speed, and occupancy, while the +RiskTraf row denotes the paired residual plug-in result. The results lead to the following observations. • RiskTraf consistently improves heterogeneous backbones. Across the displayed backbones and PEMSB-3V subsets, adding RiskTraf reduces MAE and RMSE in nearly all cases, supporting its model-agnostic plug-in design. • Larger gains appear on backbones more sensitive to shifted auxiliary correlations. DCRNN (31) and GTS (38) obtain particularly large improvements, suggesting that risk-aware residual learning is especially useful when direct three-variable ERM is unstable. • Strong recent backbones still benefit from the plug-in. Competitive models such as STWA (17), MegaCRN (26), HimNet (14), and STDN (6) also improve with RiskTraf, showing that the plug-in complements rather than replaces advanced backbone architectures. • MAPE degradation reflects denominator sensitivity. A few worse MAPE results do not contradict the MAE/RMSE gains, since MAPE is unstable for zero or near-zero actual values (23) and behaves like MAE weighted by inverse target magnitude (12). Thus, better abnormal-regime corrections can still be penalized by small-flow denominators. Figure 4 further examines whether the improvements persist at individual prediction horizons. Each bar decomposes the baseline MAE into the remaining error after applying RiskTraf and the absolute error reduced by the plug-in, where the hatched cap directly represents the removed MAE. Across all four subsets, RiskTraf reduces errors on most horizons, showing that the gains are not caused by averaging a few favorable steps. The improvement is especially visible for GTS (38), while GWNet (46) and STNorm (13) obtain smaller but still consistent reductions. This suggests that RiskTraf provides larger corrections for backbones with larger horizon-wise errors, while still refining more stable backbones. The reduction pattern is also horizon-dependent. On PEMS04-B, PEMS07-B, and PEMS08-B, the hatched caps often become more pronounced from short to medium and long horizons, although some subsets also show clear short-horizon corrections when the baseline is unstable. This pattern is consistent with the increasing uncertainty of long-range forecasting: near-term flow is constrained by recent observations, whereas later horizons require recognizing the underlying traffic regime from speed and occupancy. Therefore, the per-horizon decomposition supports that RiskTraf does not simply apply a uniform offset, but learns residual corrections that are more useful when regime-dependent uncertainty accumulates. Table 3. Comparison with robust traffic forecasting methods on PEMSB-3V. Results are averaged over horizons 3, 6, and 12. The first four methods are standalone robustness methods, while the last column reports RiskTraf plugged into the strongest backbone from Table 2. Best results are in bold. Dataset Metric CauSTG ST-SSDL Dish-TS STEVE RiskTraf (STDN) PEMS03-B MAE↓ 32.82 14.94 14.51 16.95 13.24 RMSE↓ 53.95 24.91 24.29 27.24 22.35 PEMS04-B MAE↓ 53.43 23.55 23.03 24.86 21.85 RMSE↓ 75.93 37.06 35.90 36.10 34.64 PEMS07-B MAE↓ 62.79 17.32 17.72 20.62 14.95 RMSE↓ 86.79 31.52 33.37 31.91 29.04 PEMS08-B MAE↓ 54.41 17.08 18.47 18.77 16.04 RMSE↓ 82.58 30.71 32.25 29.90 29.59 4.4. Comparison with Robust Forecasting Methods (RQ3) Table 4. Ablation study on auxiliary measurements used for traffic-risk environment construction. Results are averaged over horizons 3, 6, and 12; MAPE is reported in percentage form. The best result within each backbone–dataset block is in bold. Backbone Variant PEMS03-B PEMS04-B PEMS07-B PEMS08-B MAE↓ RMSE↓ MAPE(%)↓ MAE↓ RMSE↓ MAPE(%)↓ MAE↓ RMSE↓ MAPE(%)↓ MAE↓ RMSE↓ MAPE(%)↓ DCRNN w/o speed 35.70 53.36 149.01 43.50 62.29 34.10 45.85 70.52 33.63 36.86 54.11 56.48 w/o speed&occ 39.45 59.49 154.12 49.01 71.24 34.47 61.43 83.21 66.09 46.54 68.41 61.30 Full (speed+occ) 25.72 39.23 140.30 30.93 45.61 19.03 25.59 41.89 13.37 25.80 39.82 29.89 GWNet w/o speed 19.50 32.85 24.90 31.30 47.75 17.11 23.85 40.21 10.78 23.39 39.07 19.58 w/o speed&occ 19.52 32.88 25.16 31.80 48.46 17.14 24.48 41.09 10.96 23.71 39.60 19.63 Full (speed+occ) 17.43 29.07 23.92 27.75 42.20 15.35 21.69 36.84 9.91 21.22 35.56 18.22 STWA w/o speed 15.24 25.49 21.53 24.56 37.92 13.69 18.95 32.44 8.67 19.58 33.27 18.27 w/o speed&occ 15.28 25.56 21.68 24.68 38.07 13.75 19.11 32.58 8.68 19.73 33.52 18.14 Full (speed+occ) 14.86 24.79 21.06 24.15 37.34 13.66 18.42 31.96 8.36 18.94 32.54 17.47 To answer RQ3, we compare RiskTraf with representative robust traffic forecasting methods, including debiasing-based methods CauSTG (52) and ST-SSDL (18), and distribution-shift adaptation methods Dish-TS (16) and STEVE (25). Since these methods are standalone forecasting frameworks rather than backbone models, Table 3 reports their all-horizon MAE/RMSE separately from the paired plug-in comparison in Table 2. We include STDN (6)+RiskTraf as the representative RiskTraf configuration, as STDN is the strongest backbone in the main comparison. RiskTraf obtains the best MAE and RMSE on all four PEMSB-3V subsets. Compared with the strongest standalone competitor on each subset, it reduces MAE by 8.75%, 5.12%, 13.68%, and 6.09%, and reduces RMSE by 7.99%, 3.51%, 7.87%, and 1.04% on PEMS03-B, PEMS04-B, PEMS07-B, and PEMS08-B, respectively. The consistent improvements on both MAE and RMSE indicate that RiskTraf reduces not only average prediction bias but also larger forecasting deviations. Moreover, existing robust methods exhibit dataset-dependent rankings, suggesting that their robustness mechanisms may be sensitive to district-level traffic dynamics. In contrast, RiskTraf remains consistently strong across subsets while preserving the original backbone architecture. This supports the effectiveness of risk-aware residual correction for handling regime-dependent auxiliary correlations without replacing the forecasting model. 4.5. Component and Sensitivity Analysis (RQ4) Auxiliary-variable ablation. Table 4 ablates the auxiliary measurements used for traffic-risk environment construction while keeping the backbone, residual head, and training protocol unchanged. The w/o speed variant constructs environments using occupancy only, the w/o speed&occ variant removes both speed and occupancy from environment partitioning, and the full variant uses both measurements. The results show that occupancy alone already provides useful traffic-state information: w/o speed consistently outperforms w/o speed&occ on most backbone–dataset pairs. This is expected because occupancy directly reflects detector occupation intensity and is closely related to congestion states. However, the full speed+occ variant achieves the best overall performance across all backbone blocks, showing that speed provides complementary information about vehicle movement dynamics. The improvement is particularly clear for DCRNN, while stronger backbones such as GWNet and STWA obtain smaller but still consistent gains. These results confirm that RiskTraf benefits from constructing traffic-risk environments with both auxiliary measurements. Figure 5. Sensitivity analysis of RiskTraf on PEMS03-B with respect to the risk penalty weight λ and the number of risk environments K. Stars mark the lowest MAE in each sweep. Sensitivity to λ and K. Figure 5 studies the sensitivity of RiskTraf on PEMS03-B by varying the risk penalty weight λ and the number of environments K, while keeping β and η fixed. The weight λ controls the strength of the REx, pairwise, and extrapolation regularizers. A small λ weakens cross-environment constraints, whereas an overly large λ may sacrifice pointwise accuracy for environment consistency. The best performance is achieved around λ=0.05λ=0.05. The environment number K controls the granularity of traffic-risk partitioning. The non-monotonic curve shows that environment construction is not simply improved by using more partitions. Although K=3K=3 obtains the lowest MAE in this PEMS03-B sweep, we use K=4K=4 in the main experiments as a fixed quartile split for consistent evaluation across all dataset–backbone pairs. This avoids tuning K separately for each setting. Figure 6. Targeted STAEformer forecasting cases on PEMS04-B and PEMS08-B. The black curve denotes observed flow, the red dashed curve denotes the baseline prediction, and the blue curve denotes the prediction after adding RiskTraf. The vertical dashed line marks the prediction start, and the shaded region highlights the residual correction. 4.6. Efficiency Study (RQ5) Table 5. Plug-in efficiency comparison on PEMS03-B with batch size 32. Each pair compares a backbone with the same backbone equipped with RiskTraf. Training time denotes the observed wall-clock time of the recorded run, where RiskTraf trains only the residual head with the backbone frozen. Backbone Variant Params (M) GFLOPs/Batch Infer ms/Sample Infer Mem (MB) Train Time (s) Train Mem (MB) Test MAE GTS Base 8.307 21.229 0.941 231.7 18.08 962.2 28.45 +RiskTraf 8.312 21.262 1.036 239.9 6.55 274.2 21.45 GWNet Base 0.031 11.484 0.162 193.9 4.95 330.7 19.65 +RiskTraf 0.035 11.518 0.159 193.9 1.72 196.8 17.43 STNorm Base 0.042 1.685 0.069 143.1 5.27 287.8 17.31 +RiskTraf 0.046 1.719 0.078 143.2 2.42 145.9 15.78 Table 5 evaluates RiskTraf from a plug-in efficiency perspective. Since RiskTraf is not designed as a standalone forecasting architecture, the relevant comparison is between each backbone and the same backbone equipped with RiskTraf. Across GTS (38), GWNet (46), and STNorm (13) on PEMS03-B, RiskTraf introduces only about 0.004M additional trainable parameters and marginal GFLOP increases. The inference latency and inference memory remain in the same practical range, indicating that the residual head adds limited deployment overhead. The small overhead is accompanied by clear accuracy gains. Test MAE decreases from 28.45 to 21.45 for GTS, from 19.65 to 17.43 for GWNet, and from 17.31 to 15.78 for STNorm. These improvements show that the auxiliary risk environments and residual correction provide a favorable accuracy–efficiency trade-off. Training wall-clock time is reported as an auxiliary measurement, since it is affected by Stage-I early stopping and by the fact that only the residual head is optimized while the backbone is frozen. Figure 7. RiskTraf correction over STAEformer on PEMS03-B. Samples are grouped by speed and occupancy bins; each cell reports mean(Y^RiskTraf−Y^Base)mean( Y_RiskTraf- Y_Base), where larger positive values indicate stronger upward residual corrections. 4.7. Case Study (RQ5) Trajectory-level correction. Figure 6 visualizes representative STAEformer forecasting cases on PEMS04-B and PEMS08-B. The baseline sometimes over-reacts or under-reacts after the prediction start, leading to trajectories that deviate from the observed flow. RiskTraf adjusts the baseline prediction toward the observed trajectory in these cases, especially around turning points or short-term trend changes after the prediction boundary. This shows that the residual head does not simply apply a constant shift, but can produce horizon-wise corrections that reshape the forecast trajectory. State-dependent correction. Figure 7 further groups samples by speed and occupancy states and reports the average correction mean(Y^RiskTraf−Y^Base)mean( Y_RiskTraf- Y_Base). The correction is positive in all bins, indicating that RiskTraf generally raises the STAEformer baseline in the selected PEMS03-B setting. More importantly, the correction magnitude is strongly associated with occupancy: low-occupancy states receive relatively small corrections, whereas high-occupancy states receive much larger corrections. In contrast, the variation along the speed axis is less monotonic. This pattern suggests that occupancy provides a stronger congestion-state signal for the residual head, while speed helps refine the correction within each occupancy regime. Figure 8. t-SNE visualization of PEMS03-B samples colored by RiskTraf risk environments. The three colors denote ordered environments constructed from historical speed and occupancy, showing separated regimes with overlapping transition regions. Risk-environment structure. Figure 8 visualizes the risk-environment partition on PEMS03-B using t-SNE. The three environments form a coarse low-to-high risk organization with visible separation, while still retaining overlap in transition regions. The separated regions suggest that the speed–occupancy risk score captures distinct traffic regimes, whereas the overlapping regions correspond to samples whose raw traffic patterns are less clearly separable. This structure is consistent with the motivation of REx-based residual learning: RiskTraf does not require perfectly separated environments, but encourages the residual head to remain stable across related risk regimes while still adapting to high-risk states. Overall, these case studies provide qualitative evidence for the mechanism behind RiskTraf. The plug-in changes individual trajectories, applies larger corrections under high-occupancy traffic states, and relies on risk environments that reflect meaningful structure in the raw traffic data. 5. Related Work 5.1. Traffic Flow Prediction Early deep approaches relied on recurrent architectures such as LSTM, which are limited in capturing complex spatial dependencies of road networks; later work turned to graph-based designs: DCRNN (31) combines diffusion convolutions with sequence-to-sequence learning, and STGCN (48) adopts a fully convolutional graph architecture, both relying on predefined adjacency matrices from road topology. Graph WaveNet (46) instead learns adaptive graphs for latent spatial dependencies, and MegaCRN (26) adds memory and meta-learning for heterogeneous patterns. STAEformer (32) showed that Transformers with spatio-temporal adaptive embeddings perform strongly, STDN (6) and HimNet (14) explored seasonality and heterogeneity, and ST-SSDL (18) aligns inputs with historical means to reduce stochastic deviations. As architectural gains become marginal (40), some works turn to richer sources such as event information (EastNet (44)) or global-scale multimodal data (Terra (9)), which require additional collection and alignment. In contrast, speed and occupancy are sensor-native measurements collected with flow, yet no existing benchmark systematically studies these raw auxiliary measurements for robust flow prediction; PEMSB-3V fills this gap. 5.2. Invariant Learning and Causal Inference Invariant learning assumes that stable mechanisms remain predictive across environments while spurious correlations vary. IRM (2) seeks representations that admit a shared optimal classifier across environments, IRM Games (1) recasts this as Nash equilibrium finding, REx (28) minimizes the variance of risks across environments to improve extrapolation, and Group DRO (37) minimizes worst-group risk for subgroup robustness. Their effectiveness, however, hinges on how environments are defined: DomainBed (20) shows limited gains over empirical risk minimization when partitions do not align with the relevant spurious correlations. This is central in traffic prediction, where environments are not directly given. RiskTraf therefore builds ordered risk environments from historical speed and occupancy—physically interpretable indicators of traffic regimes—and applies REx-style regularization to a residual correction head. 5.3. Distribution Shift in Time Series Non-stationarity between training and test periods is commonly handled by statistical normalization: RevIN (27) removes instance-level statistics before prediction and restores them afterward, Dish-TS (16) models shifts through learnable transformations, and Non-stationary Transformers (34) introduce de-stationary attention against over-stationarization, while the multi-order wavelet derivative transform (54) captures evolving non-stationary dynamics in the wavelet domain. These methods target shifts in the statistical or spectral properties of the series itself. The challenge studied here is different: the relationships among flow, speed, and occupancy themselves vary across traffic regimes, so treating speed and occupancy as ordinary covariates may introduce shortcut learning. RiskTraf instead uses them to define traffic-risk environments and learns a constrained residual correction that remains stable across these environments. 6. Conclusion and Future Works In this paper, we study how to effectively use speed and occupancy for traffic flow prediction. While these sensor-native variables provide useful traffic-state information, direct three-variable training can exploit regime-dependent correlations and hurt generalization. To address this, we propose RiskTraf, a paired plug-in that freezes a trained three-variable backbone and learns a lightweight residual correction from historical speed and occupancy under a REx objective. By constructing traffic-risk environments, RiskTraf performs state-dependent flow correction on the frozen backbone. We also introduce PEMSB-3V, a public multi-variate traffic benchmark with raw flow, speed, and occupancy measurements. Experiments show that RiskTraf consistently improves diverse spatio-temporal backbones and outperforms robust forecasting methods. Future work may explore batch-independent REx objectives and extend RiskTraf to other spatio-temporal prediction tasks with auxiliary measurements. Acknowledgements. This work was supported in part by National Natural Science Foundations of China under Grant No. 62572416 and the Guangdong Provincial Key Lab of Integrated Communication, Sensing and Computation for Ubiquitous Internet of Things under Grant No. 2023B1212010007. GenAI Usage Disclosure Generative AI tools were used only to assist with language polishing, wording refinement, and formatting during manuscript preparation. The research ideas, method design, experiments, analyses, figures, tables, and conclusions were produced and verified by the authors, who take full responsibility for the content of this paper. References Ahuja et al. (2020) K. Ahuja, K. Shanmugam, K. R. Varshney, and A. Dhurandhar Invariant risk minimization games. In Proceedings of the 37th International Conference on Machine Learning, ICML’20. Cited by: §5.2. Arjovsky et al. (2020) M. Arjovsky, L. Bottou, I. Gulrajani, and D. Lopez-Paz Invariant risk minimization. External Links: 1907.02893, Link Cited by: §5.2. Ash (2012) R. B. Ash Information theory. Courier Corporation. Cited by: §1. Bai et al. (2020) L. Bai, L. Yao, C. Li, X. Wang, and C. Wang Adaptive graph convolutional recurrent network for traffic forecasting. Advances in Neural Information Processing Systems 33, p. 17804–17815. Cited by: §4.1. Cao et al. (2020) D. Cao, Y. Wang, J. Duan, C. Zhang, X. Zhu, C. Huang, Y. Tong, B. Xu, J. Bai, J. Tong, et al. Spectral temporal graph neural network for multivariate time-series forecasting. Advances in Neural Information Processing Systems 33, p. 17766–17778. Cited by: §4.1. Cao et al. (2025) L. Cao, B. Wang, G. Jiang, Y. Yu, and J. Dong Spatiotemporal-aware trend-seasonality decomposition network for traffic flow forecasting. Proceedings of the AAAI Conference on Artificial Intelligence 39 (11), p. 11463–11471. External Links: Link, Document Cited by: 3rd item, §4.1, §4.4, §5.1. Chen et al. (2003) C. Chen, J. Kwon, J. Rice, A. Skabardonis, and P. Varaiya Detecting errors and imputing missing data for single-loop surveillance systems. Transportation Research Record 1855 (1), p. 160–167. External Links: Document Cited by: §1. Chen et al. (2001) C. Chen, K. Petty, A. Skabardonis, P. Varaiya, and Z. Jia Freeway performance measurement system: mining loop detector data. Transportation Research Record 1748 (1), p. 96–102. External Links: Document Cited by: §1. Chen et al. (2024) W. Chen, X. Hao, Y. Wu, and Y. Liang Terra: a multimodal spatio-temporal dataset spanning the earth. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, External Links: Link Cited by: §5.1. Chen and Liang (2025) W. Chen and Y. Liang Expand and compress: exploring tuning principles for continual spatio-temporal graph forecasting. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §1. Chen et al. (2021) X. Chen, J. Wang, and K. Xie TrafficStream: a streaming traffic flow forecasting framework based on graph neural networks and continual learning. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, p. 3620–3626. Cited by: §1. de Myttenaere et al. (2016) A. de Myttenaere, B. Golden, B. Le Grand, and F. Rossi Mean absolute percentage error for regression models. Neurocomputing 192, p. 38–48. Note: Advances in artificial neural networks, machine learning and computational intelligence External Links: ISSN 0925-2312, Document, Link Cited by: 4th item. Deng et al. (2021) J. Deng, X. Chen, R. Jiang, X. Song, and I. W. Tsang ST-norm: spatial and temporal normalization for multi-variate time series forecasting. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, p. 269–278. Cited by: §4.1, §4.3, §4.6. Dong et al. (2024) Z. Dong, R. Jiang, H. Gao, H. Liu, J. Deng, Q. Wen, and X. Song Heterogeneity-informed meta-parameter learning for spatiotemporal time series forecasting. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, p. 631–641. Cited by: §3.3.1, 3rd item, §4.1, §5.1. Dunne and Ghosh (2013) S. Dunne and B. Ghosh Weather adaptive traffic prediction using neurowavelet models. IEEE Transactions on Intelligent Transportation Systems 14 (1), p. 370–379. Cited by: §1. Fan et al. (2023) W. Fan, P. Wang, D. Wang, D. Wang, Y. Zhou, and Y. Fu Dish-ts: a general paradigm for alleviating distribution shift in time series forecasting. In Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI’23. External Links: Link, Document Cited by: §1, §4.4, §5.3. Fang et al. (2024) Y. Fang, Y. Liang, B. Hui, Z. Shao, L. Deng, X. Liu, X. Jiang, and K. Zheng Efficient large-scale traffic forecasting with transformers: a spatial data management perspective. arXiv preprint arXiv:2412.09972. Cited by: 3rd item, §4.1. Gao et al. (2025) H. Gao, Z. Dong, J. Yong, S. Fukushima, K. Taura, and R. Jiang How different from the past? spatio-temporal time series forecasting with self-supervised deviation learning. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §3.3.2, §4.4, §5.1. Gou et al. (2026) X. Gou, Z. Li, T. Lan, J. Lin, Z. Li, B. Zhao, C. Zhang, D. Wang, and X. Zhang TraffiDent: a dataset for understanding the interplay between traffic dynamics and incidents. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, External Links: Link Cited by: Table 1. Gulrajani and Lopez-Paz (2021) I. Gulrajani and D. Lopez-Paz In search of lost domain generalization. In International Conference on Learning Representations, External Links: Link Cited by: §5.2. Guo et al. (2019) S. Guo, Y. Lin, N. Feng, C. Song, and H. Wan Attention based spatial-temporal graph convolutional networks for traffic flow forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33, p. 922–929. Cited by: Table 1. He et al. (2013) J. He, W. Shen, P. Divakaruni, L. Wynter, and R. Lawrence Improving traffic prediction with tweet semantics. In Proceedings of the Twenty-Third International Joint Conference on Artificial Intelligence, IJCAI ’13, p. 1387–1393. Cited by: §1. Hyndman and Koehler (2006) R. J. Hyndman and A. B. Koehler Another look at measures of forecast accuracy. International Journal of Forecasting 22 (4), p. 679–688. External Links: ISSN 0169-2070, Document, Link Cited by: 4th item. Ji et al. (2022) J. Ji, J. Wang, Z. Jiang, J. Jiang, and H. Zhang STDEN: towards physics-guided neural networks for traffic flow prediction. In Proceedings of the AAAI conference on artificial intelligence, Vol. 36, p. 4048–4056. Cited by: §1. Ji et al. (2025) J. Ji, W. Zhang, J. Wang, and C. Huang Seeing the unseen: learning basis confounder representations for robust traffic prediction. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, KDD ’25, p. 577–588. External Links: Link, Document Cited by: §4.4. Jiang et al. (2023) R. Jiang, Z. Wang, J. Yong, P. Jeph, Q. Chen, Y. Kobayashi, X. Song, S. Fukushima, and T. Suzumura Spatio-temporal meta-graph learning for traffic forecasting. Proceedings of the AAAI Conference on Artificial Intelligence 37, p. 8078–8086. External Links: Link, Document Cited by: 3rd item, §4.1, §5.1. Kim et al. (2021) T. Kim, J. Kim, Y. Tae, C. Park, J. Choi, and J. Choo Reversible instance normalization for accurate time-series forecasting against distribution shift. In International Conference on Learning Representations, External Links: Link Cited by: §5.3. Krueger et al. (2021) D. Krueger, E. Caballero, J. Jacobsen, A. Zhang, J. Binas, D. Zhang, R. L. Priol, and A. Courville Out-of-distribution generalization via risk extrapolation (rex). In Proceedings of the 38th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 139, p. 5815–5826. Cited by: §5.2. Lai et al. (2018) G. Lai, W. Chang, Y. Yang, and H. Liu Modeling long- and short-term temporal patterns with deep neural networks. In Proceedings of the 41st International ACM SIGIR Conference on Research & Development in Information Retrieval (SIGIR ’18), p. 105–114. External Links: Document Cited by: Table 1. Lee et al. (2015) J. Lee, B. Hong, K. Lee, and Y. Jang A prediction model of traffic congestion using weather data. In 2015 IEEE International Conference on Data Science and Data Intensive Systems, p. 81–88. Cited by: §1. Li et al. (2018) Y. Li, R. Yu, C. Shahabi, and Y. Liu Diffusion convolutional recurrent neural network: data-driven traffic forecasting. In International Conference on Learning Representations (ICLR ’18), Cited by: §1, Table 1, 2nd item, §4.1, §5.1. Liu et al. (2023a) H. Liu, Z. Dong, R. Jiang, J. Deng, J. Deng, Q. Chen, and X. Song Spatio-temporal adaptive embedding makes vanilla transformer sota for traffic forecasting. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management, p. 4125–4129. Cited by: §1, §3.3.1, §5.1. Liu et al. (2023b) X. Liu, Y. Xia, Y. Liang, J. Hu, Y. Wang, L. Bai, C. Huang, Z. Liu, B. Hooi, and R. Zimmermann LargeST: a benchmark dataset for large-scale traffic forecasting. In Advances in Neural Information Processing Systems, Cited by: Table 1. Liu et al. (2022a) Y. Liu, H. Wu, J. Wang, and M. Long Non-stationary transformers: exploring the stationarity in time series forecasting. In Advances in Neural Information Processing Systems, Cited by: §5.3. Liu et al. (2022b) Z. Liu, J. Li, and K. Wu Context-aware taxi dispatching at city-scale using deep reinforcement learning. IEEE Transactions on Intelligent Transportation Systems 23 (3), p. 1996–2009. External Links: Document Cited by: §1. Lu et al. (2023) J. Lu, C. Li, X. B. Wu, and X. S. Zhou Physics-informed neural networks for integrated traffic state and queue profile estimation: a differentiable programming approach on layered computational graphs. Transportation Research Part C: Emerging Technologies 153, p. 104224. Cited by: §1. Sagawa et al. (2020) S. Sagawa, P. W. Koh, T. B. Hashimoto, and P. Liang Distributionally robust neural networks. In International Conference on Learning Representations, External Links: Link Cited by: §5.2. Shang et al. (2021) C. Shang, J. Chen, and J. Bi Discrete graph structure learning for forecasting multiple time series. arXiv preprint arXiv:2101.06861. Cited by: 2nd item, §4.1, §4.3, §4.6. Shao et al. (2025) F. Shao, H. Shao, X. Wu, Q. Cheng, and W. H. Lam A physics-informed machine learning framework for speed-flow prediction: integrating an s-shaped traffic stream model with deep learning models. Transportation Research Part C: Emerging Technologies 180, p. 105362. Cited by: §1. Shao et al. (2024) Z. Shao, F. Wang, Y. Xu, W. Wei, C. Yu, Z. Zhang, D. Yao, T. Sun, G. Jin, X. Cao, et al. Exploring progress in multivariate time series forecasting: comprehensive benchmarking and heterogeneity analysis. IEEE Transactions on Knowledge and Data Engineering 37 (1), p. 291–305. Cited by: §1, §5.1. Song et al. (2020) C. Song, Y. Lin, S. Guo, and H. Wan Spatial-temporal synchronous graph convolutional networks: a new framework for spatial-temporal network data forecasting. Proceedings of the AAAI Conference on Artificial Intelligence 34 (01), p. 914–921. External Links: Link, Document Cited by: Table 1. Turochy and Smith (2000) R. E. Turochy and B. L. Smith New procedure for detector data screening in traffic management systems. Transportation Research Record 1727 (1), p. 127–131. External Links: Document Cited by: §1. Wang and Tong (2026) G. Wang and J. Tong Frequency as identity: a fourier hypernetwork for spatiotemporal forecasting. Applied Soft Computing 187, p. 114297. External Links: Document Cited by: §1. Wang et al. (2022) Z. Wang, R. Jiang, H. Xue, F. D. Salim, X. Song, and R. Shibasaki Event-aware multimodal mobility nowcasting. Proceedings of the AAAI Conference on Artificial Intelligence 36 (4), p. 4228–4236. External Links: Link, Document Cited by: §5.1. Wu et al. (2020) Z. Wu, S. Pan, G. Long, J. Jiang, X. Chang, and C. Zhang Connecting the dots: multivariate time series forecasting with graph neural networks. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Cited by: §1. Wu et al. (2019) Z. Wu, S. Pan, G. Long, J. Jiang, and C. Zhang Graph wavenet for deep spatial-temporal graph modeling. In Proceedings of the 28th International Joint Conference on Artificial Intelligence, IJCAI’19, p. 1907–1913. Cited by: §4.1, §4.3, §4.6, §5.1. Yin et al. (2025) D. Yin, H. Xue, A. Prabowo, S. Ao, and F. Salim XXLTraffic: Expanding and Extremely Long Traffic forecasting beyond test adaptation. arXiv. External Links: 2406.12693, Document Cited by: Table 1. Yu et al. (2018) B. Yu, H. Yin, and Z. Zhu Spatio-temporal graph convolutional networks: a deep learning framework for traffic forecasting. In Proceedings of the 27th International Joint Conference on Artificial Intelligence (IJCAI 2018), p. 3634–3640. External Links: Link Cited by: Table 1, §4.1, §5.1. Zhang et al. (2024) C. Zhang, Y. Zhang, Q. Shao, B. Li, Y. Lv, X. Piao, and B. Yin ChatTraffic: text-to-traffic generation via diffusion model. IEEE Transactions on Intelligent Transportation Systems, p. 1–13. External Links: Document Cited by: §1. Zhang et al. (2017) J. Zhang, Y. Zheng, and D. Qi Deep spatio-temporal residual networks for citywide crowd flows prediction. Proceedings of the AAAI Conference on Artificial Intelligence 31. External Links: Link, Document Cited by: §1. Zheng et al. (2020) C. Zheng, X. Fan, C. Wang, and J. Qi Gman: a graph multi-attention network for traffic prediction. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34, p. 1234–1241. Cited by: §4.1. Zhou et al. (2023a) Z. Zhou, Q. Huang, K. Yang, K. Wang, X. Wang, Y. Zhang, Y. Liang, and Y. Wang Maintaining the status quo: capturing invariant relations for ood spatiotemporal learning. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’23, New York, NY, USA, p. 3603–3614. External Links: ISBN 9798400701030, Link, Document Cited by: §4.4. Zhou et al. (2023b) Z. Zhou, Q. Huang, K. Yang, K. Wang, X. Wang, Y. Zhang, Y. Liang, and Y. Wang Maintaining the status quo: capturing invariant relations for ood spatiotemporal learning. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’23, p. 3603–3614. Cited by: §1. Zhou et al. (2026) Z. Zhou, J. Hu, Q. Wen, J. T. Kwok, and Y. Liang Multi-order wavelet derivative transform for deep time series forecasting. External Links: 2505.11781 Cited by: §5.3.