Paper deep dive
STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting
Sicong Lai, Yuehong Hu, Siru Zhong, Si Qiao, Yuxuan Liang, Guangyin Jin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/18/2026, 6:38:13 AM
Summary
The paper introduces STKAN, a spatio-temporal forecasting architecture that integrates Taylor-polynomial Kolmogorov-Arnold Network (KAN) modules into spatial and temporal token mixing. It employs a learnable soft node-group assignment mechanism and self-attention layers to capture heterogeneous spatial correlations and nonlinear temporal dynamics in traffic data. Evaluated on five traffic forecasting benchmarks, STKAN demonstrates competitive performance, outperforming MLP-based variants and highlighting the importance of nonlinear function approximator design.
Entities (31)
Relation Signals (34)
Guangyin Jin → affiliatedwith → Chang’an University
confidence 100% · Yuxuan Liang The Hong Kong University of Science and Technology (Guangzhou)China and Guangyin Jin Chang’an UniversityChina
Sicong Lai → affiliatedwith → The Hong Kong University of Science and Technology (Guangzhou)
confidence 100% · Sicong Lai The Hong Kong University of Science and Technology (Guangzhou)China
Yuehong Hu → affiliatedwith → The Hong Kong University of Science and Technology (Guangzhou)
confidence 100% · Sicong Lai The Hong Kong University of Science and Technology (Guangzhou)China , Yuehong Hu The Hong Kong University of Science and Technology (Guangzhou)China
Siru Zhong → affiliatedwith → The Hong Kong University of Science and Technology (Guangzhou)
confidence 100% · Yuehong Hu The Hong Kong University of Science and Technology (Guangzhou)China , Siru Zhong The Hong Kong University of Science and Technology (Guangzhou)China
Si Qiao → affiliatedwith → The Hong Kong University of Science and Technology (Guangzhou)
confidence 100% · Siru Zhong The Hong Kong University of Science and Technology (Guangzhou)China , Si Qiao The Hong Kong University of Science and Technology (Guangzhou)China
Yuxuan Liang → affiliatedwith → The Hong Kong University of Science and Technology (Guangzhou)
confidence 100% · Si Qiao The Hong Kong University of Science and Technology (Guangzhou)China , Yuxuan Liang The Hong Kong University of Science and Technology (Guangzhou)China
Guangyin Jin → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Real-world traffic data exhibit heterogeneous spatial correlations and nonlinear temporal dynamics, posing substantial challenges for accurate spatio-temporal forecasting. Existing approaches have developed increasingly sophisticated graph, attention, and decomposition architectures, while the influence of the underlying nonlinear function approximator has received comparatively less attention. In this work, we propose STKAN, a spatio-temporal forecasting architecture that introduces Taylor-polynomial Kolmogorov--Arnold Network modules into spatial and temporal token mixing. STKAN first constructs high-level spatial representations through a learnable soft node-group assignment mechanism, applies group-wise spatial mixing, and subsequently models temporal dependencies over the compressed sequence. Spatial and temporal self-attention layers are further employed to capture long-range interactions. Experiments on five traffic forecasting benchmarks show that STKAN achieves competitive performance and performs better than the evaluated MLP-based variant in the tested settings. These results suggest that the design of nonlinear function approximators can serve as a useful complement to architectural design in spatio-temporal forecasting.
Tags
Links
- Source: https://arxiv.org/abs/2607.13108v1
- Canonical: https://arxiv.org/abs/2607.13108v1
Trouble viewing inline? Open PDF directly →
Full Text
41,046 characters extracted from source content.
Expand or collapse full text
STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting Sicong Lai The Hong Kong University of Science and Technology (Guangzhou)China , Yuehong Hu The Hong Kong University of Science and Technology (Guangzhou)China , Siru Zhong The Hong Kong University of Science and Technology (Guangzhou)China , Si Qiao The Hong Kong University of Science and Technology (Guangzhou)China , Yuxuan Liang The Hong Kong University of Science and Technology (Guangzhou)China and Guangyin Jin Chang’an UniversityChina China University of GeosciencesChina Abstract. Real-world traffic data exhibit heterogeneous spatial correlations and nonlinear temporal dynamics, posing substantial challenges for accurate spatio-temporal forecasting. Existing approaches have developed increasingly sophisticated graph, attention, and decomposition architectures, while the influence of the underlying nonlinear function approximator has received comparatively less attention. In this work, we propose STKAN, a spatio-temporal forecasting architecture that introduces Taylor-polynomial Kolmogorov–Arnold Network modules into spatial and temporal token mixing. STKAN first constructs high-level spatial representations through a learnable soft node-group assignment mechanism, applies group-wise spatial mixing, and subsequently models temporal dependencies over the compressed sequence. Spatial and temporal self-attention layers are further employed to capture long-range interactions. Experiments on five traffic forecasting benchmarks show that STKAN achieves competitive performance and performs better than the evaluated MLP-based variant in the tested settings. These results suggest that the design of nonlinear function approximators can serve as a useful complement to architectural design in spatio-temporal forecasting. †copyright: none 1. Introduction Spatio-temporal forecasting serves as a foundational pillar within the AI4Transport ecosystem. It drives intelligent transportation technologies (Chen et al., 2018; Wang et al., 2022; Lin et al., 2022). By predicting future road traffic conditions from historical data, AI agents enable services such as autonomous traffic management, route optimization, and road safety support (Wang et al., 2020; Deng et al., 2024; Han et al., 2022). Traffic signals exhibit spatial heterogeneity and complex temporal dynamics. Spatially, interactions between regions are intricate and non-Euclidean. Temporally, signals exhibit strong periodicity entangled with abrupt changes and non-stationary patterns. These properties motivate forecasting models that can capture both spatial dependencies and temporal variations. Figure 1. Illustration of different spatio-temporal models. To address these challenges, researchers have developed sophisticated architectures to explicitly model these dependencies. As illustrated in Figure 1, Graph-based methods perform forecasting by assuming predefined or learned graph structures. These structures capture spatial correlations among nodes (Jin et al., 2023). Representative works such as STGCN (Yu et al., 2018) and Graph WaveNet (Wu et al., 2019) leverage graph convolutions and adaptive adjacency matrices. Meanwhile, Transformer-based methods utilize attention mechanisms. They model long-range dependencies across spatio-temporal tokens (Liu et al., 2023). Additionally, spectral decomposition techniques like StemGNN (Cao et al., 2020) and STWave (Fang et al., 2023) focus on frequency-domain transforms. Frameworks like STID (Shao et al., 2022) utilize identity-based embeddings. These decoupled methods aim to obtain latent spatio-temporal representations that support more accurate predictions. Despite diverse structural designs, many spatio-temporal architectures employ MLPs with fixed activation functions as their default nonlinear mapping modules. Although MLPs are expressive general-purpose approximators, their nonlinearities are usually selected in advance and shared across units. This observation motivates a complementary research question: beyond architectural design, to what extent does the choice of nonlinear function approximator affect spatio-temporal forecasting performance? In this work, we investigate whether the choice of nonlinear function approximator constitutes an underexplored factor in spatio-temporal forecasting performance. We propose STKAN, a Kolmogorov–Arnold Network (KAN)-based architecture that decomposes spatial and temporal dependencies while introducing TaylorKAN layers into the spatial and temporal token-mixing mappings. STKAN provides a controlled architectural setting for examining whether Taylor-polynomial KAN token mixers can complement conventional spatio-temporal modeling components. The channel mixer and Transformer feed-forward network remain MLP-based, so the model does not replace all MLP components. Our main contributions are summarized as follows: • We investigate the role of nonlinear function-approximator design in spatio-temporal forecasting and introduce STKAN, which incorporates Taylor-polynomial KAN mappings into spatial and temporal token-mixing modules. • We develop a learnable soft node-group assignment mechanism together with dedicated spatial and temporal mixing blocks. The grouping mechanism constructs compact spatial representations, while the two token mixers model interactions along the spatial-group and temporal dimensions, respectively. • We evaluate STKAN on five traffic forecasting benchmarks. The results demonstrate competitive forecasting accuracy, while the existing ablation results suggest that the TaylorKAN token mixers, adaptive spatial grouping, and attention components each contribute to the overall model in the evaluated configurations. 2. Related Work 2.1. Spatio-Temporal Forecasting Spatio-temporal forecasting extends traditional time-series forecasting by incorporating both temporal dynamics and spatial dependencies, such as in traffic management, where multiple traffic sensors’ data is used to predict future conditions. Early deep learning approaches combined Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) to capture spatial and temporal dependencies (Shi et al., 2015; Yao et al., 2018; Lai et al., 2018). However, grid-based CNNs may not effectively handle non-Euclidean spatial relationships, leading to the development of Graph Convolutional Networks (GCNs) (Defferrard et al., 2016; Kipf and Welling, 2016) and spatio-temporal Graph Neural Networks (STGNNs) (Li et al., 2017; Yu et al., 2018; Li et al., 2023). These models, such as DCRNN (Li et al., 2017), ST-MetaNet (Pan et al., 2019), and DGCRN (Li et al., 2023), integrate GCNs with RNNs (Cho et al., 2014), while others like Graph WaveNet (Wu et al., 2019) and STGCN (Yu et al., 2018) combine GCNs with gated Temporal Convolutional Networks (TCNs). Attention mechanisms have also been widely adopted in STGNNs (Zheng et al., 2020). Some studies criticize the reliance on pre-defined graphs, suggesting alternatives like AGCRN (Bai et al., 2020) and MTGNN (Wu et al., 2020), which learn latent graph structures. Recent non-GCN solutions, such as STNorm (Deng et al., 2021) and STID (Shao et al., 2022), further highlight the need for a better understanding of spatial dependencies in forecasting tasks. 2.2. Kolmogorov-Arnold Network KANs are motivated by the Kolmogorov–Arnold representation theorem (Liu et al., 2024), which represents multivariate functions through compositions of univariate functions. Building upon this paradigm, several variants have been proposed, including wavelet-based WavKAN (Bozorgasl and Chen, 2024), Taylor polynomial-based TaylorKAN (Yu et al., 2025), fKAN with trainable Jacobi basis functions (Aghaei, 2025), and FastKAN (Li, 2024), which employs radial basis function approximations of B-spline bases. Although KANs are motivated by the Kolmogorov–Arnold representation theorem, the theorem itself does not imply that a finite KAN model will universally outperform an MLP in empirical forecasting tasks. In this work, KANs are therefore treated as an alternative nonlinear parameterization with a different inductive bias, rather than as a theoretically guaranteed replacement for MLPs. KANs have recently attracted attention in time series forecasting. Prior work has explored KAN-based models for time series analysis (Vaca-Rubio et al., 2024), mixture-of-experts KAN designs (Han et al., 2024), symbolic-regression-oriented T-KAN and MT-KAN models (Xu et al., 2024), collaborative time-frequency learning with iTFKAN (Liang et al., 2025), and frequency-decomposition learning with TimeKAN (Huang et al., 2025). BiLSTM-KAN integrates KAN layers with a bidirectional LSTM backbone for traffic flow forecasting (Chen et al., 2024). STKAN differs from these studies by replacing the token-mixing mappings in its spatial and temporal mixer blocks with Taylor-polynomial KAN layers, while retaining conventional MLPs in the channel-mixing and Transformer feed-forward components. 3. Preliminary Spatio-temporal forecasting is a specialized multivariate time-series forecasting problem. Given a historical input sequence t−Lin+1:t=[t−Lin+1,…,t]∈ℝLin×N×CinX_t-L_in+1:t=[X_t-L_in+1,…,X_t] ^L_in× N× C_in, the goal is to predict ^t+1:t+H=[^t+1,…,^t+H]∈ℝH×N×Cout Y_t+1:t+H=[ Y_t+1,…, Y_t+H] ^H× N× C_out. Here, LinL_in denotes the historical input length, H denotes the prediction horizon, N is the number of spatial nodes, and CinC_in and CoutC_out are the input and output channel dimensions. 4. Methodology Figure 2. Overview of our proposed STKAN framework. In this paper, we propose STKAN to model spatial interactions and temporal dynamics by introducing TaylorKAN token mixers into a spatio-temporal forecasting architecture. The overall architecture of STKAN is shown in Figure 2. 4.1. Embedding Layers To capture spatio-temporal dependencies in traffic sequences, we embed the raw input ∈ℝLin×N×CinX ^L_in× N× C_in with a fully connected layer as f=FC()∈ℝLin×N×dfE_f=FC(X) ^L_in× N× d_f. We use a learnable spatial embedding s∈ℝN×dsE_s ^N× d_s. The spatial embedding is shared across time steps and is broadcast along the temporal dimension before concatenation. To incorporate temporal periodicity, we use two embedding dictionaries: a time-of-day embedding with Ttod=288T_tod=288 slots for five-minute data and a day-of-week embedding with Tdow=7T_dow=7 categories. The resulting embeddings are denoted as todE_tod and dowE_dow after lookup and broadcasting to the node dimension. 4.2. Temporal Convolution Blocks All embeddings are concatenated along the feature dimension to form the initial representation: (1) (0)=f‖s‖tod∥dow∈ℝLin×N×dh,X^(0)=E_f\|E_s\|E_tod\|E_dow ^L_in× N× d_h, where dh=df+ds+dtod+ddowd_h=d_f+d_s+d_tod+d_dow. To capture local temporal context while reducing the temporal resolution, we define a patch-extraction operator that applies a 1×w1× w convolution along the time axis with stride s: (2) p=PatchConv((0);w,s)∈ℝLp×N×dh,X_p=PatchConv (X^(0);w,s ) ^L_p× N× d_h, where Lp=⌊Lin−ws⌋+1L_p= L_in-ws +1. Thus, pX_p serves as a patch-level spatio-temporal feature map for subsequent mixer and attention blocks. 4.3. Spatio-Temporal KAN Blocks The STKAN blocks combine learnable spatial grouping, spatial TaylorKAN token mixing, and temporal TaylorKAN token mixing. STKAN first uses a soft assignment matrix to aggregate raw nodes into macro-level spatial tokens. The aggregated representations are then processed by the Spatial KAN (SKAN) block. The spatial output is mapped back to node-level features and passed to the Temporal KAN (TKAN) block, which models dependencies over the compressed temporal dimension. Learnable Spatial Grouping. To transform the original spatial nodes into macro-level spatial tokens, we introduce an unnormalized learnable parameter matrix ∈ℝN×GA ^N× G. The assignment matrix is obtained by applying softmax along the group dimension: (3) n,g=exp(n,g)∑g′=1Gexp(n,g′).G_n,g= (A_n,g) _g =1^G (A_n,g ). Thus, ∑g=1Gn,g=1 _g=1^GG_n,g=1 for each node n. Given p∈ℝLp×N×dhX_p ^L_p× N× d_h, the group-level representation ~∈ℝLp×G×dh X ^L_p× G× d_h is computed as (4) ~τ,g,c=∑n=1Nn,gp,τ,n,c. X_τ,g,c= _n=1^NG_n,gX_p,τ,n,c. Spatial KAN Block. The SKAN block follows a mixer-style design. Although MLPs are expressive general-purpose approximators, their activation functions are usually fixed before training. KAN-based mappings provide a different parameterization in which learnable univariate functions are associated with network connections. The transmission from the j-th neuron in layer ℓ+1 +1 to all neurons in layer ℓ is formulated as: (5) zℓ+1,j=∑i=1nℓϕℓ,j,i(zℓ,i),z_ +1,j= _i=1^n_ _ ,j,i(z_ ,i), where zℓ,iz_ ,i is the i-th neuron in layer ℓ , nℓn_ is the number of neurons in that layer, and ϕℓ,j,i(⋅) _ ,j,i(·) is a learnable univariate mapping. In our model, ϕφ is instantiated by a TaylorKAN layer. A Taylor expansion motivates the polynomial basis: (6) f(x)≈∑r=0Kf(r)(0)r!xr.f(x)≈ _r=0^K f^(r)(0)r!x^r. The implemented layer uses learnable coefficients: (7) ϕq()=∑p=1C∑r=0Kθq,p,rxpr+bq. _q(x)= _p=1^C _r=0^K _q,p,rx_p^r+b_q. The coefficients θq,p,r _q,p,r are optimized directly and are not constrained to equal the analytical derivatives of an underlying function. Accordingly, the adopted layer is more precisely interpreted as a Taylor-inspired polynomial-basis KAN layer. The polynomial basis allows the mapping to combine nonlinear terms of different orders. Its effectiveness for spatial and temporal mixing is evaluated empirically in the subsequent experiments. To avoid ambiguity when mixing tensor dimensions, we define a spatial permutation operator s:ℝLp×G×dh→ℝLp×dh×GP_s:R^L_p× G× d_h ^L_p× d_h× G. The spatial token-mixing stage is then (8) s=s(~)+TaylorKANs(LN(s(~))),U_s=P_s( X)+TaylorKAN_s (LN (P_s( X) ) ), where s∈ℝLp×dh×GU_s ^L_p× d_h× G. We restore the original group-feature order by (9) ¯s=s−1(s)∈ℝLp×G×dh. U_s=P_s^-1(U_s) ^L_p× G× d_h. The channel mixer remains an MLP applied to the feature dimension: (10) s=¯s+2σ(1LN(¯s)+1)+2,V_s= U_s+W_2σ (W_1LN( U_s)+b_1 )+b_2, where s∈ℝLp×G×dhV_s ^L_p× G× d_h and σ denotes the GELU activation. The node-level spatial output is (11) τ,n,c=p,τ,n,c+∑g=1Gn,gs,τ,g,c.S_τ,n,c=X_p,τ,n,c+ _g=1^GG_n,gV_s,τ,g,c. Temporal KAN Blocks. The TKAN block applies TaylorKAN token mixing along the compressed temporal dimension. We define a temporal permutation operator t:ℝLp×N×dh→ℝN×dh×LpP_t:R^L_p× N× d_h ^N× d_h× L_p. The temporal token mixer is (12) t=t()+TaylorKANt(LN(t())),U_t=P_t(S)+TaylorKAN_t (LN (P_t(S) ) ), where t∈ℝN×dh×LpU_t ^N× d_h× L_p. The channel mixer is written as (13) t=t−1(t)+MLPt(LN(t−1(t))),V_t=P_t^-1(U_t)+MLP_t (LN (P_t^-1(U_t) ) ), and the final temporal representation is (14) =t∈ℝLp×N×dh.T=V_t ^L_p× N× d_h. 4.4. Spatio-Temporal Fusion Blocks To complement the TaylorKAN token mixers, we introduce Transformer layers along both spatial and temporal axes. Given a hidden representation ∈ℝLp×N×dhT ^L_p× N× d_h, the attention block projects it into query, key, and value matrices: (15) =Q,=K,=V,Q=TW_Q, =TW_K, =TW_V, where Q,K,V∈ℝdh×dhW_Q,W_K,W_V ^d_h× d_h. The scaled dot-product attention is (16) Attention(,,)=Softmax(⊤dk).Attention(Q,K,V)=Softmax ( QK d_k )V. Spatial attention is computed independently at each time step along the node dimension, so the softmax operates over the N spatial nodes. Temporal attention is computed independently at each node along the compressed temporal dimension, so the softmax operates over the LpL_p temporal tokens. Figure 2 indicates that the attention block contains LayerNorm, FFN, and residual connections in addition to the attention operation. 4.5. Prediction Head After feature extraction and fusion via STKAN blocks, the prediction head aggregates temporal information across compressed time steps for each node and applies a linear projection to generate multi-step forecasts. Let Z denote the output representation after the fusion block. The prediction head is expressed as: (17) ^t+1:t+H=fout(), Y_t+1:t+H=f_out(Z), where ^t+1:t+H∈ℝH×N×Cout Y_t+1:t+H ^H× N× C_out. The output function foutf_out denotes the reshape and linear projection operations used to map the fused representation to the future horizon. 5. Experiment 5.1. Experimental Setup Datasets. We evaluate our model on five traffic forecasting datasets, including PEMS04, PEMS07, PEMS08, PEMS-BAY and METR-LA. Following previous work, we divide the PEMS04, PEMS07 and PEMS08 dataset into training, validation, and test sets in a ratio of 6:2:2. For the remaining datasets, we adopt a split ratio of 7:1:2. Detailed statistics of these datasets are shown in Table 1. Dataset PEMS04 PEMS07 PEMS08 PEMS-BAY METR-LA Sensors 307 883 170 325 207 Time Steps 16992 28224 17856 52116 34272 Time Interval 5min 5min 5min 5min 5min Table 1. Summary of Five Spatio-temporal Benchmarks Baselines. We compare 11 representative baselines with our proposed STKAN. (i) Non-spatial modeling-based: STID (Shao et al., 2022), which adopts identity spatio-temporal embeddings and avoids explicit spatial dependency modeling. (i) Static spatial-based methods: STGCN (Yu et al., 2018), GWNet (Wu et al., 2019), AGCRN (Bai et al., 2020), GMAN (Zheng et al., 2020), MTGNN (Wu et al., 2020) and STDN (Cao et al., 2025) combine pre-defined or learned static graph structures with temporal modeling modules. (i) Dynamic spatial-based methods: STAEformer (Liu et al., 2023) and STWave (Fang et al., 2023) capture time-varying spatial dependencies through adaptive or attention-based mechanisms. (iv) Spatio-temporal decomposition-based: StemGNN (Cao et al., 2020) and STNorm (Deng et al., 2021) decompose spatio-temporal series into separate components for modeling, focusing on disentangling spatial and temporal patterns. Settings. In the experiments, we use the traffic flow of the last 12 time steps to predict the traffic flow of the next 12 time steps, and record the prediction performance of the 3rd, 6th, 12th steps and the average. We set the Adam optimizer with an initial learning rate of 0.002, where the learning rate follows a step-wise decay strategy, and the batch size is set as 64. During the training phase, we employ the early stopping strategy with tolerance 30 for 200 epochs. For performance evaluation, we adopt three widely used metrics to quantify the accuracy of traffic forecasting results: Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and Mean Absolute Percentage Error (MAPE). All baseline methods are implemented within the BasicTS (Shao et al., 2024) framework to ensure consistent training pipelines and fair performance comparison. Dataset PEMS04 PEMS07 PEMS08 PEMS-BAY METR-LA Method Metric @3 @6 @12 Avg. @3 @6 @12 Avg. @3 @6 @12 Avg. @3 @6 @12 Avg. @3 @6 @12 Avg. GWNet (2019) MAE 17.89 18.80 20.35 18.81 18.71 20.14 22.35 20.10 13.67 14.59 15.99 14.58 1.31 1.65 1.99 1.59 2.69 3.08 3.52 3.04 RMSE 28.81 30.40 32.66 30.38 30.70 33.20 36.58 33.11 21.65 23.54 25.82 23.46 2.76 3.74 4.54 3.66 5.17 6.20 7.28 6.15 MAPE 12.23% 12.99% 14.24% 12.97% 8.04% 8.50% 9.73% 8.59% 9.20% 9.69% 10.41% 9.69% 2.77% 3.80% 4.84% 3.66% 6.93% 8.33% 9.84% 8.15% STGCN (2018) MAE 19.09 19.98 21.74 20.03 20.67 22.23 25.04 22.28 15.97 16.86 18.64 16.96 1.41 1.75 2.08 1.70 2.76 3.15 3.63 3.12 RMSE 30.1 31.57 34.07 31.63 32.76 35.71 40.4 35.83 24.56 26.29 29.00 26.38 2.93 3.89 4.69 3.81 5.30 6.32 7.47 6.29 MAPE 12.95% 13.44% 14.73% 13.73% 8.97% 9.53% 10.71% 9.59% 10.91% 11.56% 12.58% 11.50% 3.06% 3.98% 4.85% 3.84% 7.11% 8.61% 10.40% 8.49% AGCRN (2020) MAE 18.55 19.50 20.77 19.45 19.29 20.82 22.81 20.74 14.78 15.96 17.63 15.91 1.35 1.67 1.96 1.61 2.87 3.24 3.63 3.19 RMSE 29.86 31.60 33.51 31.46 31.64 34.64 38.05 34.50 22.98 25.01 27.79 25.03 2.85 3.80 4.54 3.69 5.61 6.66 7.58 6.52 MAPE 12.88% 13.44% 14.18% 13.40% 8.15% 8.70% 9.70% 8.83% 9.53% 11.72% 12.17% 10.86% 2.93% 3.84% 4.68% 3.69% 7.74% 9.03% 10.30% 8.85% StemGNN (2020) MAE 19.14 20.82 24.05 21.00 20.78 23.25 27.91 23.41 14.63 16.05 18.76 16.20 1.39 1.78 2.20 1.73 2.97 3.50 4.24 3.49 RMSE 30.38 32.78 37.09 33.06 32.78 36.56 43.05 36.89 22.98 25.42 29.45 25.62 2.92 3.95 4.94 3.90 5.82 7.04 8.59 7.06 MAPE 13.68% 14.82% 17.44% 15.05% 9.29% 10.21% 12.45% 10.39% 9.28% 10.50% 12.26% 10.55% 2.94% 4.09% 5.32% 3.97% 7.97% 10.06% 13.01% 10.04% GMAN (2020) MAE 18.23 18.78 20.12 18.81 19.31 20.41 22.20 20.48 13.76 14.59 15.83 14.81 1.35 1.66 1.93 1.58 2.81 3.15 3.49 3.07 RMSE 29.38 30.91 31.25 30.99 31.25 33.32 36.51 33.40 22.78 24.15 26.49 24.23 2.92 3.84 4.51 3.69 5.56 6.50 7.36 6.43 MAPE 12.71% 13.27% 13.41% 13.22% 8.22% 8.71% 9.44% 8.65% 9.40% 9.53% 10.56% 9.71% 2.88% 3.75% 4.54% 3.69% 7.42% 8.75% 10.11% 8.65% MTGNN (2020) MAE 18.29 19.12 20.57 19.12 19.52 21.11 23.87 21.16 14.23 15.30 16.97 15.31 1.32 1.65 1.95 1.59 2.70 3.07 3.53 3.04 RMSE 29.82 31.34 33.57 31.28 31.37 34.19 38.46 34.26 22.38 24.33 26.78 24.25 2.78 3.73 4.50 3.65 5.21 6.17 7.24 6.14 MAPE 12.62% 13.09% 14.31% 13.14% 8.77% 9.10% 10.34% 9.27% 9.42% 10.57% 12.17% 10.60% 2.75% 3.68% 4.55% 3.53% 6.85% 8.17% 9.81% 8.08% STNorm (2021) MAE 18.30 19.12 20.27 19.05 19.21 20.57 22.66 20.51 14.48 15.45 17.03 15.45 1.33 1.66 1.97 1.58 2.80 3.18 3.56 3.12 RMSE 29.82 31.52 33.22 31.28 31.65 34.66 38.30 34.48 23.05 25.38 27.93 25.22 2.85 3.81 4.56 3.67 5.49 6.52 7.47 6.41 MAPE 12.32% 12.83% 13.69% 12.81% 8.29% 8.69% 9.61% 8.70% 9.27% 9.79% 10.90% 9.88% 2.85% 3.77% 4.63% 3.59% 7.44% 8.89% 10.26% 8.65% STID (2022) MAE 17.62 18.40 19.72 18.41 18.40 19.66 21.54 19.62 13.29 14.22 15.55 14.20 1.31 1.64 1.91 1.56 2.79 3.17 3.54 3.11 RMSE 28.61 29.95 31.93 29.93 30.45 32.82 36.04 32.75 21.53 23.40 25.72 23.34 2.77 3.73 4.40 3.60 5.52 6.57 7.53 6.47 MAPE 11.95% 12.42% 13.50% 12.51% 7.77% 8.28% 9.22% 8.31% 8.65% 9.29% 10.32% 9.31% 2.77% 3.73% 4.52% 3.55% 7.66% 9.27% 10.77% 9.01% STAEformer (2023) MAE 17.48 18.24 19.30 18.19 18.00 19.40 21.42 19.33 12.71 13.55 14.84 13.55 1.30 1.61 1.87 1.54 2.65 2.96 3.33 2.93 RMSE 28.89 30.31 31.99 30.18 30.42 33.30 37.02 33.21 21.63 23.48 25.80 23.44 2.77 3.68 4.34 3.57 5.11 6.01 7.02 5.98 MAPE 11.78% 12.21% 13.00% 12.25% 7.61% 8.19% 9.03% 8.14% 8.33% 8.92% 9.85% 8.90% 2.74% 3.63% 4.41% 3.46% 6.90% 8.20% 9.77% 8.10% STWave (2023) MAE 17.57 18.17 19.42 18.25 18.57 19.91 21.75 19.93 12.78 13.76 14.86 13.69 1.32 1.63 1.89 1.56 2.83 3.22 3.58 3.15 RMSE 28.88 29.95 31.78 29.99 31.59 34.36 37.35 34.09 21.59 23.79 25.77 23.57 2.80 3.71 4.35 3.59 5.63 6.71 7.60 6.56 MAPE 11.65% 12.02% 13.13% 12.16% 7.63% 8.17% 9.07% 8.20% 8.63% 9.16% 10.03% 9.10% 2.76% 3.66% 4.44% 3.50% 7.72% 9.49% 11.03% 9.20% STDN (2025) MAE 18.15 18.89 20.14 18.92 19.92 21.16 23.51 21.29 13.85 14.43 15.71 14.53 1.38 1.66 1.93 1.61 2.79 3.15 3.53 3.10 RMSE 33.14 34.64 35.85 34.33 33.56 35.88 39.61 36.02 22.31 23.90 26.21 23.96 2.95 3.83 4.47 3.66 5.59 6.61 7.56 6.51 MAPE 19.34% 19.24% 19.80% 19.22% 12.73% 12.12% 14.78% 12.85% 12.45% 11.64% 10.92% 11.43% 3.03% 3.81% 4.47% 3.66% 7.61% 9.08% 10.73% 8.93% STKAN(Ours) MAE 17.40 18.13 19.15 18.09 17.94 19.27 20.97 19.16 12.62 13.48 14.80 13.45 1.29 1.61 1.87 1.54 2.69 2.99 3.39 2.97 RMSE 28.63 30.00 31.59 29.87 30.13 32.96 36.12 32.74 21.25 23.24 25.56 23.17 2.73 3.68 4.33 3.56 5.13 6.08 7.14 6.06 MAPE 11.78% 12.24% 12.95% 12.23% 7.60% 8.05% 8.93% 8.07% 8.28% 8.90% 9.84% 8.88% 2.70% 3.62% 4.35% 3.44% 6.98% 8.25% 9.71% 8.11% Table 2. Forecasting performance on the five benchmark datasets. We bold the best results and underline the second-best results. 5.2. Performance Comparisons Table 2 reports the forecasting results on the five benchmark datasets. STKAN achieves the best average MAE and RMSE on PEMS04, although its average MAPE is slightly higher than that of STWave. On PEMS07 and PEMS08, STKAN obtains the best average results across the three reported metrics. On PEMS-BAY, STKAN ties with STAEformer in average MAE and obtains the lowest average RMSE and MAPE. On METR-LA, STKAN remains competitive but does not outperform the strongest baseline in average MAE or RMSE. Overall, the results indicate that STKAN is particularly effective on the evaluated traffic-flow datasets, whereas its advantage is less evident on the METR-LA speed dataset. The gains over the best-performing baselines are generally modest. Since the current evaluation reports single-run results, small numerical differences should be interpreted cautiously. These results demonstrate the competitiveness of the overall STKAN architecture. The comparison with the MLP variant further suggests that the adopted TaylorKAN token mixers are useful within the evaluated configuration, although the current experiments do not isolate function-approximator capacity as the sole source of the observed gains. 5.3. Ablation Study Effectiveness of KAN Modules. To evaluate the role of KAN components within STKAN, we design three model variants: • MLPs: replacing the TaylorKAN token mixers with MLP layers in the evaluated implementation. • w/o SKAN: removing the spatial block while retaining TKAN to test the importance of inter-node mixing. • w/o TKAN: removing the temporal block while retaining the spatial modeling component. As shown in Figure 3, replacing the TaylorKAN token mixers with the evaluated MLP implementation increases the forecasting errors on PEMS04 and PEMS08. This result suggests that the Taylor-polynomial mappings contribute positively within the current architecture and hyperparameter setting. Removing either SKAN or TKAN also degrades at least part of the reported performance, indicating that spatial-group mixing and temporal mixing provide complementary information. These observations are specific to the evaluated model configurations and should not be interpreted as establishing the general advantage of KANs over all parameter-matched MLP alternatives. Figure 3. KAN modules ablation on PEMS04 and PEMS08. Effectiveness of Attention Mechanisms. To evaluate the role of attention in STKAN, we conduct ablation studies on PEMS04 and PEMS08 with the following variants: • w/o S-Attention: disabling spatial attention while preserving temporal modeling. • w/o T-Attention: removing temporal attention while keeping spatial modeling. • w/o ST-Attention: removing both attention modules, leaving only KAN-based token and channel mixing. Table 3 shows that the complete model achieves the best result for every reported metric on PEMS04 and PEMS08. However, the relative influence of spatial and temporal attention varies across metrics. For example, removing temporal attention produces the largest MAE degradation on PEMS08, whereas removing spatial attention produces the largest RMSE degradation. Removing both branches is therefore not uniformly the worst variant for every metric. The results indicate that the two attention branches contribute differently across datasets and evaluation criteria. The ablation results suggest that attention provides an additional refinement over the KAN-based mixer blocks, but the current experiments do not establish a strict ranking between the contributions of attention and KAN components. Dataset PEMS04 PEMS08 Metric MAE RMSE MAPE MAE RMSE MAPE w/o S-Attention 18.24 30.13 12.43% 13.70 23.68 9.06% w/o T-Attention 18.28 30.27 12.40% 13.84 23.39 9.20% w/o ST-Attention 18.31 30.05 12.52% 13.57 23.30 8.99% STKAN 18.09 29.87 12.23% 13.45 23.17 8.88% Table 3. Ablation study of the attention block. 5.4. Hyper-parameter Study Table 4 reports the sensitivity of STKAN to the number of spatial groups G. The results exhibit a non-monotonic pattern: using either fewer or more groups than the selected value leads to slightly higher forecasting errors. Among the tested settings, G=16G=16 performs best on PEMS04, while G=20G=20 performs best on PEMS07. This observation suggests that the group number should be selected according to the dataset rather than increased monotonically with the number of nodes. PEMS04 PEMS07 G MAE RMSE MAPE G MAE RMSE MAPE 8 18.28 30.76 12.55% 12 19.41 33.18 8.19% 12 18.20 30.16 12.40% 16 19.18 32.77 8.10% 16 18.09 29.87 12.23% 20 19.16 32.74 8.08% 20 18.20 30.00 12.44% 24 19.26 32.86 8.17% 24 18.25 30.25 12.61% 28 19.24 32.93 8.13% Table 4. Hyper-parameter study on varying G values. 5.5. Case Study Figure 4 provides a qualitative visualization of the learned soft node-to-group assignment matrix on PEMS-BAY. Several groups exhibit relatively concentrated assignment patterns, and some corresponding hard assignments appear to cover spatially contiguous road segments. This observation suggests that the learned grouping mechanism may capture corridor-level regularities in the sensor network. However, the visualization should be interpreted as qualitative evidence rather than a formal validation of geographical consistency or model interpretability. Figure 4. Adaptive grouping matrices visualization on PEMS-BAY. 6. Conclusion In this work, we introduced STKAN, a spatio-temporal forecasting architecture that incorporates Taylor-polynomial KAN mappings into spatial and temporal token-mixing modules. The model combines learnable soft node grouping, group-wise spatial mixing, temporal mixing, and spatial–temporal attention to model traffic dynamics. Experiments on five benchmark datasets show that STKAN achieves competitive forecasting performance, with particularly strong results on the evaluated traffic-flow datasets. The comparison with the tested MLP variant suggests that TaylorKAN token mixers can provide a useful alternative nonlinear parameterization within the proposed architecture. These findings indicate that function-approximator design is a relevant component of spatio-temporal modeling, alongside spatial and temporal architectural design. Future work may examine parameter-matched comparisons, computational efficiency, statistical variability, and broader spatio-temporal applications. References A. A. Aghaei (2025) Fkan: fractional kolmogorov–arnold networks with trainable jacobi basis functions. Neurocomputing 623, p. 129414. Cited by: §2.2. L. Bai, L. Yao, C. Li, X. Wang, and C. Wang (2020) Adaptive graph convolutional recurrent network for traffic forecasting. Advances in neural information processing systems 33, p. 17804–17815. Cited by: §2.1, §5.1. Z. Bozorgasl and H. Chen (2024) Wav-kan: wavelet kolmogorov-arnold networks. arXiv preprint arXiv:2405.12832. Cited by: §2.2. D. Cao, Y. Wang, J. Duan, C. Zhang, X. Zhu, C. Huang, Y. Tong, B. Xu, J. Bai, J. Tong, et al. (2020) Spectral temporal graph neural network for multivariate time-series forecasting. Advances in neural information processing systems 33, p. 17766–17778. Cited by: §1, §5.1. L. Cao, B. Wang, G. Jiang, Y. Yu, and J. Dong (2025) Spatiotemporal-aware trend-seasonality decomposition network for traffic flow forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 11463–11471. Cited by: §5.1. L. Chen, Q. Zhong, X. Xiao, Y. Gao, P. Jin, and C. S. Jensen (2018) Price-and-time-aware dynamic ridesharing. In 2018 IEEE 34th international conference on data engineering (ICDE), p. 1061–1072. Cited by: §1. Y. Chen, S. Li, N. Zhao, R. Zheng, and Y. Li (2024) BiLSTM-kan: a time series-based traffic flow forecasting model. In Proceedings of the 2024 13th International Conference on Computing and Pattern Recognition, p. 314–319. Cited by: §2.2. K. Cho, B. Van Merriënboer, D. Bahdanau, and Y. Bengio (2014) On the properties of neural machine translation: encoder-decoder approaches. arXiv preprint arXiv:1409.1259. Cited by: §2.1. M. Defferrard, X. Bresson, and P. Vandergheynst (2016) Convolutional neural networks on graphs with fast localized spectral filtering. Advances in neural information processing systems 29. Cited by: §2.1. J. Deng, X. Chen, R. Jiang, X. Song, and I. W. Tsang (2021) St-norm: spatial and temporal normalization for multi-variate time series forecasting. In Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining, p. 269–278. Cited by: §2.1, §5.1. J. Deng, F. Ye, D. Yin, X. Song, I. Tsang, and H. Xiong (2024) Parsimony or capability? decomposition delivers both in long-term time series forecasting. Advances in Neural Information Processing Systems 37, p. 66687–66712. Cited by: §1. Y. Fang, Y. Qin, H. Luo, F. Zhao, B. Xu, L. Zeng, and C. Wang (2023) When spatio-temporal meet wavelets: disentangled traffic forecasting via efficient spectral graph attention networks. In 2023 IEEE 39th International Conference on Data Engineering (ICDE), p. 517–529. Cited by: §1, §5.1. J. Han, H. Liu, H. Xiong, and J. Yang (2022) Semi-supervised air quality forecasting via self-supervised hierarchical graph neural network. IEEE Transactions on Knowledge and Data Engineering 35 (5), p. 5230–5243. Cited by: §1. X. Han, X. Zhang, Y. Wu, Z. Zhang, and Z. Wu (2024) Are kans effective for multivariate time series forecasting?. arXiv preprint arXiv:2408.11306. Cited by: §2.2. S. Huang, Z. Zhao, C. Li, and L. Bai (2025) Timekan: kan-based frequency decomposition learning architecture for long-term time series forecasting. arXiv preprint arXiv:2502.06910. Cited by: §2.2. G. Jin, Y. Liang, Y. Fang, Z. Shao, J. Huang, J. Zhang, and Y. Zheng (2023) Spatio-temporal graph neural networks for predictive learning in urban computing: a survey. IEEE transactions on knowledge and data engineering 36 (10), p. 5388–5408. Cited by: §1. T. N. Kipf and M. Welling (2016) Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907. Cited by: §2.1. G. Lai, W. Chang, Y. Yang, and H. Liu (2018) Modeling long-and short-term temporal patterns with deep neural networks. In The 41st international ACM SIGIR conference on research & development in information retrieval, p. 95–104. Cited by: §2.1. F. Li, J. Feng, H. Yan, G. Jin, F. Yang, F. Sun, D. Jin, and Y. Li (2023) Dynamic graph convolutional recurrent network for traffic prediction: benchmark and solution. ACM Transactions on Knowledge Discovery from Data 17 (1), p. 1–21. Cited by: §2.1. Y. Li, R. Yu, C. Shahabi, and Y. Liu (2017) Diffusion convolutional recurrent neural network: data-driven traffic forecasting. arXiv preprint arXiv:1707.01926. Cited by: §2.1. Z. Li (2024) Kolmogorov-arnold networks are radial basis function networks. arXiv preprint arXiv:2405.06721. Cited by: §2.2. Z. Liang, R. An, W. Fan, Y. Rao, and Y. Liang (2025) ITFKAN: interpretable time series forecasting with kolmogorov-arnold network. arXiv preprint arXiv:2504.16432. Cited by: §2.2. H. Lin, Z. Gao, Y. Xu, L. Wu, L. Li, and S. Z. Li (2022) Conditional local convolution for spatio-temporal meteorological forecasting. In Proceedings of the AAAI conference on artificial intelligence, Vol. 36, p. 7470–7478. Cited by: §1. H. Liu, Z. Dong, R. Jiang, J. Deng, J. Deng, Q. Chen, and X. Song (2023) Spatio-temporal adaptive embedding makes vanilla transformer sota for traffic forecasting. In Proceedings of the 32nd ACM international conference on information and knowledge management, p. 4125–4129. Cited by: §1, §5.1. Z. Liu, Y. Wang, S. Vaidya, F. Ruehle, J. Halverson, M. Soljačić, T. Y. Hou, and M. Tegmark (2024) Kan: kolmogorov-arnold networks. arXiv preprint arXiv:2404.19756. Cited by: §2.2. Z. Pan, Y. Liang, W. Wang, Y. Yu, Y. Zheng, and J. Zhang (2019) Urban traffic prediction from spatio-temporal data using deep meta learning. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, p. 1720–1730. Cited by: §2.1. Z. Shao, F. Wang, Y. Xu, W. Wei, C. Yu, Z. Zhang, D. Yao, T. Sun, G. Jin, X. Cao, et al. (2024) Exploring progress in multivariate time series forecasting: comprehensive benchmarking and heterogeneity analysis. IEEE Transactions on Knowledge and Data Engineering 37 (1), p. 291–305. Cited by: §5.1. Z. Shao, Z. Zhang, F. Wang, W. Wei, and Y. Xu (2022) Spatial-temporal identity: a simple yet effective baseline for multivariate time series forecasting. In Proceedings of the 31st ACM international conference on information & knowledge management, p. 4454–4458. Cited by: §1, §2.1, §5.1. X. Shi, Z. Chen, H. Wang, D. Yeung, W. Wong, and W. Woo (2015) Convolutional lstm network: a machine learning approach for precipitation nowcasting. Advances in neural information processing systems 28. Cited by: §2.1. C. J. Vaca-Rubio, L. Blanco, R. Pereira, and M. Caus (2024) Kolmogorov-arnold networks (kans) for time series analysis. arXiv preprint arXiv:2405.08790. Cited by: §2.2. J. Wang, J. Ji, Z. Jiang, and L. Sun (2022) Traffic flow prediction based on spatiotemporal potential energy fields. IEEE Transactions on Knowledge and Data Engineering 35 (9), p. 9073–9087. Cited by: §1. S. Wang, J. Cao, and S. Y. Philip (2020) Deep learning for spatio-temporal data mining: a survey. IEEE transactions on knowledge and data engineering 34 (8), p. 3681–3700. Cited by: §1. Z. Wu, S. Pan, G. Long, J. Jiang, X. Chang, and C. Zhang (2020) Connecting the dots: multivariate time series forecasting with graph neural networks. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining, p. 753–763. Cited by: §2.1, §5.1. Z. Wu, S. Pan, G. Long, J. Jiang, and C. Zhang (2019) Graph wavenet for deep spatio-temporal graph modeling. arXiv preprint arXiv:1906.00121. Cited by: §1, §2.1, §5.1. K. Xu, L. Chen, and S. Wang (2024) Kolmogorov-arnold networks for time series: bridging predictive power and interpretability. arXiv preprint arXiv:2406.02496. Cited by: §2.2. H. Yao, F. Wu, J. Ke, X. Tang, Y. Jia, S. Lu, P. Gong, J. Ye, and Z. Li (2018) Deep multi-view spatio-temporal network for taxi demand prediction. In Proceedings of the AAAI conference on artificial intelligence, Vol. 32. Cited by: §2.1. B. Yu, H. Yin, and Z. Zhu (2018) Spatio-temporal graph convolutional networks: a deep learning framework for traffic forecasting. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, p. 3634–3640. Cited by: §1, §2.1, §5.1. S. Yu, Z. Chen, Z. Yang, J. Gu, B. Feng, and Q. Sun (2025) Exploring kolmogorov-arnold networks for realistic image sharpness assessment. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1–5. Cited by: §2.2. C. Zheng, X. Fan, C. Wang, and J. Qi (2020) Gman: a graph multi-attention network for traffic prediction. In Proceedings of the AAAI conference on artificial intelligence, Vol. 34, p. 1234–1241. Cited by: §2.1, §5.1.