Paper deep dive
VideoSEMA: a scalable and efficient Mamba-like attention for video understanding
Nhat Thanh Tran, Fanghui Xue andShuai Zhang, Jiancheng Lyu, Yunling Zheng, Yingyong Qi, Jack Xin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 7/18/2026, 3:42:36 AM
Summary
The paper introduces VideoSEMA, a scalable and efficient video understanding model that utilizes a split space-time attention framework. It combines a Mamba-like SEMA block for spatial processing (integrating local window attention and global averaging) with softmax temporal attention. Theoretical analysis demonstrates that split space-time attention is mathematically equivalent to full space-time attention under specific rank conditions. Empirically, VideoSEMA outperforms heavier Vision Transformer and Mamba-based models on the K400 and SSv2 benchmarks, showing superior accuracy and graceful resolution scaling without fine-tuning.
Entities (10)
Relation Signals (9)
Jack Xin → authored → VideoSEMA
confidence 99% · Jack Xin Department of Mathematics University of California, Irvine
VideoSEMA → employs → Split Space-Time Attention
confidence 96% · We present for video understanding (classification) a split space-time attention model, VideoSEMA
VideoSEMA → evaluatedon → SSv2
confidence 95% · On benchmark SSv2 data, VideoSEMA leads in top-1 accuracy among models of similar parameter sizes.
VideoSEMA → uses → SEMA
confidence 95% · VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space
SEMA → inspiredby → Mamba
confidence 94% · SEMA is designed to mimic the recurrent relation of Mamba in the large token number limit
VideoSEMA → uses → Softmax Temporal Attention
confidence 93% · and a softmax temporal attention in time
VideoSEMA → outperforms → VideoMamba
confidence 92% · VideoSEMA degrades much more gracefully than VideoMamba in accuracy
VideoSEMA → outperforms → TimeSformer
confidence 91% · VideoSema made considerable computational savings from TimeSformer... while achieving better accuracies
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time. In each frame, SEMA attention applies a local window attention in parallel with a global averaging in a Mamba macro-architecture, which is called Mamba-like. Under certain rank conditions, we prove that the computationally cheaper split space-time attention is equivalent to full space-time attention. On benchmark K400 data sets, VideoSEMA out-performs heavier vision transformer and Mamba models. On benchmark SSv2 data, VideoSEMA leads in top-1 accuracy among models of similar parameter sizes. As image resolution scales up from standard $224^2$ to $1024^2$ on K400 and without fine-tuning, VideoSEMA degrades much more gracefully than VideoMamba in accuracy. It is promising to extend VideoSEMA to longer videos with a dilated/sparse temporal attention.
Tags
Links
- Source: https://arxiv.org/abs/2607.14711v1
- Canonical: https://arxiv.org/abs/2607.14711v1
Trouble viewing inline? Open PDF directly →
Full Text
49,207 characters extracted from source content.
Expand or collapse full text
VideoSEMA: a scalable and efficient Mamba-like attention for video understanding Nhat Thanh Tran Department of Mathematics University of California, Irvine Irvine, USA nhattt@uci.edu &Fanghui Xue Qualcomm AI Research San Diego, USA fangxue@qti.qualcomm.com &Shuai Zhang Qualcomm AI Research San Diego, USA shuazhan@qti.qualcomm.com &Jiancheng Lyu Qualcomm AI Research San Diego, USA jianlyu@qti.qualcomm.com &Yunling Zheng Qualcomm AI Research San Diego, USA yunlzhen@qti.qualcomm.com &Yingyong Qi Qualcomm AI Research San Diego, USA yingyong@qti.qualcomm.com &Jack Xin Department of Mathematics University of California, Irvine Irvine, USA jack.xin@uci.edu Corresponding authorQualcomm AI Research is an initiative of Qualcomm Technologies, Inc. Abstract We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time. In each frame, SEMA attention applies a local window attention in parallel with a global averaging in a Mamba macro-architecture, which is called Mamba-like. Under certain rank conditions, we prove that the computationally cheaper split space-time attention is equivalent to full space-time attention. On benchmark K400 data sets, VideoSEMA out-performs heavier vision transformer and Mamba models. On benchmark SSv2 data, VideoSEMA leads in top-1 accuracy among models of similar parameter sizes. As image resolution scales up from standard 2242224^2 to 102421024^2 on K400 and without fine-tuning, VideoSEMA degrades much more gracefully than VideoMamba in accuracy. It is promising to extend VideoSEMA to longer videos with a dilated/sparse temporal attention. 1 Introduction Understanding complex spatiotemporal patterns in videos remains a fundamental challenge in computer vision, with application ranging from classification and segmentation to world modeling for autonomous system such as robotics and self-driving vehicles. With the widespread of usage of multimodal models in combination with language models, recent advances have leveraged transformer based architecture and large foundation models to capture long range dependencies, achieving state of the art performance Bertasius et al. (2021); Fan et al. (2021); Li et al. (2022b); Wang et al. (2025b, a); Assran et al. (2025). Despite their success, these models predominately rely on Vision Transformer (ViT) backbones, which incur high computational costs for large video inputs. To alleviate this problem, some works explores Mamba Gu and Dao (2024) as backbone, which reduces the computation cost for long sequences, such as in VideoMamba Li et al. (2024). Linear attention is another approach to reduce the cost of classical attention mechanism and WLiT demonstrates that it is an effective way to process video Sun et al. (2023). Lately, Mamba-like attention models have been developed as an efficient replacement of classical softmax attention in image application Han et al. (2024). Mamba-like means to keep the macro-architecture of Mamba yet embed in it light weight non-Mamba attention blocks. Along this line, SEMA Tran et al. (2026) is designed to mimic the exponential forgetting property of Mamba by an asymptotically guided global approximation of softmax attention. For a duality view and treatment of softmax and efficient attentions, see Nguyen et al. (2023); Zheng et al. (2024). An operator splitting and integro-differential equation perspective of transformer is in Tai et al. (2025). In this work, we explore the efficacy of SEMA backbone in video applications. Our main contributions are: • We develop a novel video model in the split space-time attention framework based on a Mamba-like scalable and efficient image backbone (SEMA,Tran et al. (2026)). • To our best knowledge, this is the first work to explore Mamba-like backbones for videos. • We show theoretically that the split space-time attention is equivalent to the full space-time attention under specific rank conditions. • We demonstrate on K400 (SSv2) classification tasks that VideoSEMA is an efficient model, outperforming much larger (comparable) parameter and flops size transformer and Mamba video models to date. Post-training and without fine-tuning on K400, it is much more robust than VideoMamba Li et al. (2024) as video frames scale up in size. 2 Related Works Video representation learning has evolved from CNN to attention based models to hierarchical and efficient sequence architectures such as Mamba Gu and Dao (2024). TimeSformer Bertasius et al. (2021) demonstrated that factorized space time attention can replace 3D convolutions for video recognition, while MViT and MViTv2 Fan et al. (2021); Li et al. (2022b) introduced multiscale hierarchical Transformer that improves efficiency and scalability across image and video tasks. More recently, state space models have been explored for long range temporal modeling, with VideoMamba achieving competitive performance on video understanding tasks Li et al. (2024). At scale, foundation models such as InternVideo2 unify self-supervised, constrastive, and generative objectives for multimodal video understanding Wang et al. (2025b). Other efforts focus on efficiency and prediction such as adaptive token strategies improving training and inference flexibility Wang et al. (2025a), while V-JEPA 2 advanced self-supervised world modeling for motion prediction and long horizon reasoning in video Assran et al. (2025). However, many of these rely on ViT as a vision backbone, whereas our work explores a light weight Mamba-like model (SEMA Tran et al. (2026)) as an alternative backbone to capture spatial and temporal features efficiently. 3 Methodology 3.1 Preliminary 3.1.1 Attention and SEMA Attention is a core component of Transformer which is the main driving force of deep learning in recent years. Formally, given an input x∈ℝn×dx ^n× d, then softmax full attention is defined as: A(Q,K,V)=softmax(QKT)V,A(Q,K,V)=softmax(QK^T)V, (1) where Q=xWQ+bQ,K=xWK+bK,V=xWV+bVQ=xW_Q+b_Q,K=xW_K+b_K,V=xW_V+b_V, for WQ,WK,WV∈ℝd×dW_Q,W_K,W_V ^d× d and bQ,bK,bV∈ℝn×db_Q,b_K,b_V ^n× d. We observe that compute the attention matrix (softmax(QKT)softmax(QK^T)) requires (n2)O(n^2) operations. Also as n→∞n→∞, the attention matrix disperses, i.e. tending to zero uniformly Tran et al. (2026). Thus, it is unable to distinguish the variations across the keys. On the other hand, SEMA Tran et al. (2026) is designed to mimic the recurrent relation of Mamba Gu and Dao (2024) in the large token number limit with an averaging operation to approximate the global aspect of the attention matrix. This alleviates the burden in calculating full attention, with an access to global information in an efficient manner. Concretely, SEMA(Q,K,V):=Aw(Q,K,V)+[1n∑j=1nvj],SEMA(Q,K,V):=A_w(Q,K,V)+ [ 1n _j=1^nv_j ], (2) where [⋅][·] broadcasts the row n times to permit matrix addition and AwA_w is window attention Liu et al. (2021) defined as: Aw(Q,K,V):=[∑j∈J(1)exp(q1kjT)vj∑i∈J(1)exp(q1kiT)⋮∑j∈J(n)exp(qnkjT)vj∑i∈J(n)exp(qnkiT)],A_w(Q,K,V):= bmatrix _j∈ J(1) (q_1k_j^T)v_j _i∈ J(1) (q_1k_i^T)\\ \\ _j∈ J(n) (q_nk_j^T)v_j _i∈ J(n) (q_nk_i^T) bmatrix, (3) for some index set J(m)J(m). An example is J(m)=Mw+1,…,(M+1)wJ(m)=\Mw+1,…,(M+1)w\, where M=⌊m−1w⌋M= m-1w . The latter term of Eq. (2) is proven to be a good approximation for softmax(QKT)softmax(QK^T) as n→∞n→∞ under realistic assumption of the input feature x, and also with high probability that such an approximation holds Tran et al. (2026). Thus SEMA is an effective mechanism to process large input sequence with (n)O(n) computational complexity. 3.1.2 Mamba as a Recursive Attention Mamba Gu and Dao (2024) is a state space models to handle long input sequence with computational complexity of (n)O(n). For complete derivation of state space models, we refer to Gu and Dao (2024); Gu et al. (2020). Concretely, Mamba is a map from x to y through the dynamical system: ht h_t =At⊙ht−1+Bt(Δt⊙xt), =A_t h_t-1+B_t( _t x_t), (4) yt y_t =Ctht+D⊙xt, =C_th_t+D x_t, (5) where xt,Δt∈ℝ1×d,At,ht∈ℝd×d,Bt∈ℝd×1x_t, _t ^1× d,A_t,h_t ^d× d,B_t ^d× 1 and yt∈ℝ1×d,Ct∈ℝ1×d,D∈ℝ1×dy_t ^1× d,C_t ^1× d,D ^1× d, and ⊙ denotes the Hadamard product. It follows in discrete form that ym=∑i=1mqmk~iTv~i+D⊙xm,y_m= _i=1^mq_m k_i^T v_i+D x_m, (6) where k~iT=(∏j=1m−iAm−(j−1))⊙Bi k_i^T= ( _j=1^m-iA_m-(j-1) ) B_i, v~i=Δi⊙xi v_i= _i x_i, ∏Π is short for multiple matrix products in the elementwise (Hadamard) sense, with the convention that ∏Π acts as identity if the upper index is zero. The first term in (6) reveals the (Q,K,V)(Q,K,V) structure implicit in Mamba, while the second term can be understood as a skip connection (a casual masked attention). 3.2 Proposed Method - VideoSEMA Given an input video x∈ℝC×T×h×wx ^C× T× h× w. VideoSEMA first tokenizes using a 3D convolution (e.g. kernel=(1,4,4)) to project the input video into X∈ℝC×T×H×WX ^C× T× H× W of non-overlapping spatiotemporal patches where H=h/4,W=w/4H=h/4,W=w/4. Second, we append a learnable classifier token at the end of the sequence. The last step of pre-processing is to add learned positional embedding and then temporal embedding. Concretely, X X =3DConv(x) =3DConv(x) (7a) X X =[X,Xcls]+Embpos+Embtemp, =[X,X_cls]+Emb_pos+Emb_temp, (7b) where Xcls∈ℝH×W×CX_cls ^H× W× C is classifier token, Embpos∈ℝH×W×CEmb_pos ^H× W× C and Embtemp∈ℝT×CEmb_temp ^T× C are learned positional and temporal embedding respectively. Here the addition is broadcast along the appropriate dimension to enable matrix addition. Next, we will process the spatial and temporal information of the input via the VideoSEMA block. The VideoSEMA block is constructed as y y =Spatial_SEMA(X), =Spatial\_SEMA(X), (8a) y′ y =LayerNorm(y), =LayerNorm(y), (8b) z′ z =Temp_Attn(y′), =Temp\_Attn(y ), (8c) z z =X+z′, =X+z , (8d) where spatial function operates over the H×WH× W dimension of the input, while the temporal function operates over the T dimension. To be precise, for Q,K,V∈ℝT×H×W×CQ,K,V ^T× H× W× C, we have Spatial_SEMA(qt,h,w) Spatial\_SEMA(q_t,h,w) :=1HW∑i=1H∑j=1Wvt,i,j := 1HW _i=1^H _j=1^Wv_t,i,j +∑l,p∈I(h,w)exp(qt,h,wkt,l,pT)∑i,j∈I(h,w)exp(qt,h,wkt,i,jT)vt,l,p, + _l,p∈ I(h,w) (q_t,h,wk_t,l,p^T) _i,j∈ I(h,w) (q_t,h,wk_t,i,j^T)v_t,l,p, (9) where I(h,w)I(h,w) is the index set for the window, e.g. window size of 77. And temporal attention Temp_Attn(qt,h,w):=∑i=1Texp(qt,h,wki,h,wT)∑j=1Texp(qt,h,wkj,h,wT)vi,h,w.Temp\_Attn(q_t,h,w):= _i=1^T (q_t,h,wk_i,h,w^T) _j=1^T (q_t,h,wk_j,h,w^T)v_i,h,w. (10) The alternating spatial and temporal attention treatment in (8a) and (8c) can be viewed as an efficient operator splitting approximation of the full SEMA obtained by directly extending SEMA formula (2) to space and time. For video frames of moderate lengths considered here, a softmax attention in (8c) is affordable and effective as supported by our ablation study, making linear complexity approximations unnecessary in time. As seen later, VideoSEMA outperforms much heavier networks (Tab.1). As a future direction for handling longer videos, a sparse or dilated attention Ding et al. (2023) in the temporal attention step is a promising alternative to leverage continuity of information in time at reduced computational costs. VideoSEMA’s macro-structure is shown in Fig.1, and is repeated in the network. Between two adjacent repeats, we use a convolutional layer to down-sample the spatial dimension of the video signal to produce a hierarchical network (similar to Han et al. (2024); Tran et al. (2026); Liu et al. (2021)). Lastly in Algorithm 1, the representation of the [CLS][CLS] token is normalized before passing through a linear classification head. Figure 1: Overview of VideoSEMA macro-structure. frame t−δt- ttframe t+δt+ -TimeAttention (ST)SplitSpace-TimeAttention (T+S)SEMA SplitSpace-Time Attention (T+window+average) Figure 2: Visualization of space time attention types. For illustration, the dark dot denotes the query patch and colored patches show its self-attention space-time neighborhood under each scheme. Patches without color are not used for the self-attention computation of the query patch. Multiple colors within a scheme denote attentions separately applied along different dimensions, e.g., space and time for (T+S)(T+S). Note that self-attention is computed for every single patch in the video clip, i.e., every patch serves as a query. Although the attention pattern is shown for only two adjacent frames, it extends in the same fashion to all frames of the clip. Algorithm 1 Proposed VideoSEMA algorithm. Here cls, temp, pos denote classifier tokens, temporal embedding, and positional embedding respectively. When there are dimensions mismatch in binary operations, there are implicit reshaping/broadcasting to match the dimensions. On the right hand side, we denote the current running shape of x. Input: Input x (C×T×H×WC× T× H× W) Learnable Params: cls ∈ℝH×W×C ^H× W× C, temp ∈ℝT×C ^T× C, pos ∈ℝHW×C ^HW× C 1: x=x= reshape(x) (T×H×W×CT× H× W× C) 2: x=x= concat((x, cls), dim=0) ((T+1)×H×W×C(T+1)× H× W× C) 3: x = reshape(x) ((T+1)×HW×C(T+1)× HW× C) 4: x = x+x+pos ((T+1)×HW×C(T+1)× HW× C) 5: x=x+x=x+temp ((T+1)×HW×C(T+1)× HW× C) 6: x=x= reshape(x) ((T+1)×HW×C(T+1)× HW× C) 7: x = VideoSEMA(x)((T+1)×HW×C(T+1)× HW× C) 8: Process spatial information via SEMA 9: Process temporal information via classical attention 10: x = AvgPooling(x)((T+1)×C(T+1)× C) 11: Return x[−1,:]x[-1,:] (1×C1× C) 3.3 Complexity Given an input video x∈ℝT×H×W×Cx ^T× H× W× C, the joint space-time attention treats the video as a sequence of length N=THWN=THW. Therefore, the attention matrix has size N×N× N, yielding computational complexity ((THW)2C)O((THW)^2C). In practice, this formulation is prohibitively expensive for videos due to the quadratic memory and compute cost of the attention matrix. Consequently, fully joint space time attention is rarely used in large scale video models Bertasius et al. (2021). The split space time attention approach works with a factorized (or separate spatial and temporal) attention, thereby lower the computational costs similar to operator splitting Strang (1968); Glowinski et al. (2017). In particular, TimeSformer Bertasius et al. (2021) applies a split 2-step procedure: (1) spatial attention independently in each frame, and (2) temporal attention independently across frames at each spatial location. The spatial attention complexity in (1) is (T(HW)2C)O(T(HW)^2C) as each of the T frames performs full attention over the HWHW spatial tokens. The temporal attention complexity in (2) is (T2HWC)O(T^2HWC) as each spatial location performs full attention over T tokens. So the total complexity of the split attention design becomes ((T2HW+T(HW)2)C)O((T^2HW+T(HW)^2)C). In contrast, VideoSEMA replaces full spatial attention with SEMA, whose complexity scales linearly with spatial dimension (THWwC)O(THWwC) where w denotes the window size. The temporal attention is still performed globally, thus the total computational complexity of VideoSEMA is ((T2HW+THWw)C)O((T^2HW+THWw)C), linear in spatial resolution HWHW. In the datasets of this paper, the number of video frames is at most 32 (i.e. T≤32T≤ 32 in Tab. 1, and T≤16T≤ 16 in Tab. 2 and Tab. 3). The main contribution to the complexity comes from spatial resolution HW=(224)2HW=(224)^2. With a linear complexity in HWHW, VideoSema made considerable computational savings from TimeSformer Bertasius et al. (2021), as seen in Table 1 on 16×224216× 224^2 resolution of K400 dataset where VideoSema’s parameter size is a fraction 31/121 of TimeSformer’s while achieving better accuracies. We present visualization of each space-time attention in Fig. 2. 4 Theoretical Explanation 4.1 A Simplified Setting In this section, we show that in an ideal setting, the split space-time attention can be equivalent to full attention. Given an input x∈ℝN×T×dx ^N× T× d, where N,T,dN,T,d represent spatial, temporal and channel dimensions respectively. For simplicity, let V=xV=x, then the full attention for a fixed query q becomes: AST=∑i=1N∑j=1Tηijxij,ηij=ϕ(qkijT)∑ijϕ(qkijT),A_ST= _i=1^N _j=1^T _ijx_ij,\;\; _ij= φ(q\,k_ij^T) _ijφ(q\,k_ij^T), (11) where kijk_ij and xijx_ij are vectors in ℝdR^d at space-time location (i,j)(i,j), 1≤i≤N1≤ i≤ N, 1≤j≤T1≤ j≤ T. Here we define ϕ:ℝ→ℝ+φ:R ^+ to be any continuous function. Now we compute split space-time attention AT+SA_T+S for the fixed query q. Without loss of generality, we have the split space-time function as the composite function of space and then time. The order does not matter, as it is up to a swap in the first and second dimensions of the input x. We have AT+S=∑j=1Tαj(∑i=1NβN(j−1)+ixij),A_T+S= _j=1^T _j ( _i=1^N _N(j-1)+i\,x_ij ), (12) where αj _j’s represent attention score in the time dimension, and β⋅ _·’s represent the frame-wise attention score. The equivalence AST=AT+SA_ST=A_T+S is realized if the following system of equations hold: ηij=αjβN(j−1)+i, _ij= _j _N(j-1)+i, (13) ∑i,j=1N,Tηij=1, _i,j=1^N,T _ij=1, (14) ∑i=1NβN(j−1)+i=1,∀j∈[1,…,T], _i=1^N _N(j-1)+i=1,\;\;∀ j∈[1,…,T], (15) ∑j=1Tαj=1. _j=1^T _j=1. (16) An ideal solution for this system exists as follows. Suppose the standard full attention weights ηij _ij are given as in (11) and so (14) is satisfied. To match these attention scores by the split space-time attention (12), we let βN(j−1)+i=ηij∑l=1Nηlj, _N(j-1)+i= _ij _l=1^N _lj, (17) and αj=∑l=1Nηlj, _j= _l=1^N _lj, (18) then (15) and (16) hold automatically. In practice however, one computes the weights (α⋅,β⋅)( _·, _·) of the split attention (12) without knowledge of full attention weights. The corresponding factorization (13) may not satisfy normalization condition (14), and may not be equal to the (ηij)( _ij) in (11). The ideal solution (17)-(18) would only be an indirect theoretical target for the training of AT+SA_T+S. 4.2 Attention Factorization and Rank Conditions Now we turn to standard attention. The previous section only addresses a single query, however in practice one works with multiple queries. Assume we have M queries, and let m∈1,…,Mm∈\1,…,M\ index the query row, and let ηmij _mij be the full softmax attention coefficient from query m to spatial location i in frame j. The single-query construction above applies row-wise, so define αmj=∑l=1Nηmlj,βmi|j=ηmijαmj. _mj= _l=1^N _mlj, _mi|j= _mij _mj. (19) Since softmax coefficients are strictly positive, αmj>0 _mj>0 and the definition is well-defined. Moreover, ∑j=1Tαmj=1,∑i=1Nβmi|j=1,αmjβmi|j=ηmij. _j=1^T _mj=1, _i=1^N _mi|j=1, _mj _mi|j= _mij. (20) Thus the split output equals the full output for every query: AST(m)=∑j=1T∑i=1Nηmijxij=∑j=1Tαmj(∑i=1Nβmi|jxij)=AT+S(m).A_ST^(m)= _j=1^T _i=1^N _mijx_ij= _j=1^T _mj ( _i=1^N _mi|jx_ij )=A_T+S^(m). (21) Therefore split space time attention is an exact marginal-conditional decomposition of each row of the full attention matrix. The only remaining question is whether a chosen dot-product parameterization can realize the required temporal and spatial logits. For positive normalized attention, αmj>0 _mj>0, and these coefficients again satisfy αmjβmi|j=ηmij _mj _mi|j= _mij. A particular split attention module realizes this decomposition exactly when its temporal branch can produce the distribution αm,: _m,: and its spatial branch can produce the conditional distributions βm:|j _m:|j under the same normalized ϕφ operation. If ϕφ is invertible on its positive range, one possible choice of temporal and spatial scores is any rmjr_mj and smijs_mij satisfying ϕ(rmj)=cmαmj,ϕ(smij)=cmjβmi|j,φ(r_mj)=c_m _mj, φ(s_mij)=c_mj _mi|j, (22) for arbitrary positive constants cmc_m and cmjc_mj. Theorem 4.1 (Exact rank condition for normalized ϕφ). Assume ϕ:ℝ→ℝ+φ:R ^+ is invertible on its positive range. Choose positive constants cmc_m and cmjc_mj so that cmαmjc_m _mj and cmjβmi|jc_mj _mi|j lie in the range of ϕφ, and define the target temporal score matrix Rϕ∈ℝM×TR^φ ^M× T and target spatial score matrix Sϕ∈ℝM×NTS^φ ^M× NT by Rmjϕ=ϕ−1(cmαmj),Sm,(j,i)ϕ=ϕ−1(cmjβmi|j).R^φ_mj=φ^-1(c_m _mj), S^φ_m,(j,i)=φ^-1(c_mj _mi|j). (23) Let dT,dS∈ℕd_T,d_S be the temporal and spatial head dimensions. If dT≥rank(Rϕ),dS≥rank(Sϕ),d_T (R^φ), d_S (S^φ), (24) then there exist learned query and key matrices Qτ∈ℝM×dT,Kτ∈ℝT×dTQ^τ ^M× d_T,K^τ ^T× d_T and QS∈ℝM×dS,KS∈ℝNT×dSQ^S ^M× d_S,K^S ^NT× d_S such that Qτ(Kτ)T=Rϕ,QS(KS)T=Sϕ.Q^τ(K^τ)^T=R^φ, Q^S(K^S)^T=S^φ. (25) Consequently, normalized ϕφ attention over the temporal scores produces αm,: _m,:, normalized ϕφ attention over the spatial scores inside each frame produces βm:|j _m:|j, and split space time attention exactly equals full attention for every query row m. Proof. The rank assumptions imply matrix factorizations Rϕ=Qτ(Kτ)TR^φ=Q^τ(K^τ)^T and Sϕ=QS(KS)TS^φ=Q^S(K^S)^T, for example by the compact singular value decomposition. For each query row m, normalizing ϕ(Rm,:ϕ)φ(R^φ_m,:) over the temporal index gives ϕ(Rmjϕ)∑n=1Tϕ(Rmnϕ)=cmαmj∑n=1Tcmαmn=αmj. φ(R^φ_mj) _n=1^Tφ(R^φ_mn)= c_m _mj _n=1^Tc_m _mn= _mj. (26) Likewise, for each fixed (m,j)(m,j), normalizing ϕ(Sm,(j,:)ϕ)φ(S^φ_m,(j,:)) over the spatial index gives ϕ(Sm,(j,i)ϕ)∑l=1Nϕ(Sm,(j,l)ϕ)=cmjβmi|j∑l=1Ncmjβml|j=βmi|j. φ(S^φ_m,(j,i)) _l=1^Nφ(S^φ_m,(j,l))= c_mj _mi|j _l=1^Nc_mj _ml|j= _mi|j. (27) Therefore the split coefficient is αmjβmi|j=ηmij _mj _mi|j= _mij for all m,i,jm,i,j. ∎ Remark 4.2 (Softmax as a special case). If ϕ(x)=exp(x)φ(x)= (x), then the above construction recovers the usual softmax factorization. In this case βmi|j=exp(emij)∑l=1Nexp(emlj),αmj=exp(τmj)∑n=1Texp(τmn), _mi|j= (e_mij) _l=1^N (e_mlj), _mj= ( _mj) _n=1^T ( _mn), (28) where the temporal score can be chosen as the frame-level log-sum-exp τmj=log(∑l=1Nexp(emlj)),emij=qmkijT. _mj= ( _l=1^N (e_mlj) ),e_mij=q_mk_ij^T. (29) More generally, if the desired normalized coefficients ηmij _mij are already known, softmax scores can be chosen as rmj=log(αmj)+cm,smij=log(βmi|j)+cmj,r_mj= ( _mj)+c_m, s_mij= ( _mi|j)+c_mj, (30) since adding a constant to all scores in a normalized softmax does not change the resulting distribution. This is the special case of Theorem 4.1 with ϕ−1(x)=log(x)φ^-1(x)= (x). Now suppose that we are given a full rank condition as in Theorem 4.1, then we can realize the matrices Qτ,Kτ,QS,KSQ^τ,K^τ,Q^S,K^S under a typical condition, for example, given y∈ℝM×dTy ^M× d_T, we want Qτ=yWQQ^τ=yW_Q, for some WQ∈ℝdT×dTW_Q ^d_T× d_T. Thus we solve the optimization problem minW‖Qτ−yW‖2, _W\|Q^τ-yW\|_2, (31) and this has an exact solution if col(Qτ)⊂col(y)col(Q^τ)⊂ col(y). Theorem 4.3 (Low-rank approximation for normalized ϕφ). Assume ϕφ is invertible on its positive range, and let RϕR^φ and SϕS^φ be the target score matrices from Theorem 4.1. Let R R and S S be rank-constrained approximations with rank(R^)≤dT,rank(S^)≤dS.rank( R)≤ d_T, ( S)≤ d_S. (32) Define the approximate normalized coefficients α^mj=ϕ(R^mj)∑n=1Tϕ(R^mn),β^mi|j=ϕ(S^m,(j,i))∑l=1Nϕ(S^m,(j,l)), α_mj= φ( R_mj) _n=1^Tφ( R_mn), β_mi|j= φ( S_m,(j,i)) _l=1^Nφ( S_m,(j,l)), (33) and let η^mij=α^mjβ^mi|j η_mij= α_mj β_mi|j. If ϵT=maxm‖α^m,:−αm,:‖1,ϵS=maxm,j‖β^m:|j−βm:|j‖1, _T= _m\| α_m,:- _m,:\|_1, _S= _m,j\| β_m:|j- _m:|j\|_1, (34) then for every query row m, ∑j=1T∑i=1N|η^mij−ηmij|≤ϵT+ϵS. _j=1^T _i=1^N| η_mij- _mij|≤ _T+ _S. (35) If additionally ‖xij‖2≤B\|x_ij\|_2≤ B for all i,ji,j, then ‖A^T+S(m)−AST(m)‖2≤B(ϵT+ϵS).\| A_T+S^(m)-A_ST^(m)\|_2≤ B( _T+ _S). (36) Proof. Add and subtract α^mjβmi|j α_mj _mi|j: |η^mij−ηmij| | η_mij- _mij| =|α^mjβ^mi|j−αmjβmi|j| =| α_mj β_mi|j- _mj _mi|j| (37) ≤α^mj|β^mi|j−βmi|j|+|α^mj−αmj|βmi|j. ≤ α_mj| β_mi|j- _mi|j|+| α_mj- _mj| _mi|j. (38) Summing over i,ji,j, and using that α α and β are probability distributions, gives ∑j=1T∑i=1N|η^mij−ηmij|≤‖α^m,:−αm,:‖1+∑j=1Tα^mj‖β^m:|j−βm:|j‖1≤ϵT+ϵS. _j=1^T _i=1^N| η_mij- _mij|≤\| α_m,:- _m,:\|_1+ _j=1^T α_mj\| β_m:|j- _m:|j\|_1≤ _T+ _S. (39) The output bound follows from ‖A^T+S(m)−AST(m)‖2≤∑j=1T∑i=1N|η^mij−ηmij|‖xij‖2.\| A_T+S^(m)-A_ST^(m)\|_2≤ _j=1^T _i=1^N| η_mij- _mij|\|x_ij\|_2. (40) ∎ Remark 4.4. In particular, one may choose R R and S S as best rank-dTd_T and rank-dSd_S approximations of RϕR^φ and SϕS^φ in Frobenius norm. By the Eckart-Young theorem, ‖Rϕ−R^‖F2=∑r>dTσr(Rϕ)2,‖Sϕ−S^‖F2=∑r>dSσr(Sϕ)2,\|R^φ- R\|_F^2= _r>d_T _r(R^φ)^2, \|S^φ- S\|_F^2= _r>d_S _r(S^φ)^2, (41) where σr _r is the r singular value. Thus the approximation quality is controlled by the singular value tails of the target temporal and spatial score matrices, together with how sensitively the normalized ϕφ map converts score errors into coefficient errors. Remark 4.5. We are mostly dealing with long sequence, thus the number of query M or spatial N is large. Therefore the rank condition is quite strict, that is we are most likely to be in the scenario of Theorem 4.3. However when we work with short video clips as in SSv2 dataset, i.e. T is small, QτQ^τ and KτK^τ may be found exactly during training. 5 Experiments Table 1: Performance reported on standard K400 dataset. Here F×m×nF× m× n denotes by F the total FLOPs, m the number of testing segments, n the number of testing crops. Res: resolution, T: transformer. P(M): parameter size in millions. T1: top-1, T5: top-5. C: CNN, CT: CNN +T, M: Mamba, Ml: Mamba-like. Method Type P(M) FLOPs (G) Res T1(%) T5(%) X3D-XL Feichtenhofer (2020) C 20 194×3×10194× 3× 10 16×224216× 224^2 80.4 94.6 Swin-T Liu et al. (2022) T 28 88×3×488× 3× 4 32×224232× 224^2 78.8 93.6 MViTv1-B Fan et al. (2021) CT 37 70×1×570× 1× 5 32×224232× 224^2 80.2 94.4 MViTv2-S Li et al. (2022b) CT 35 64×1×564× 1× 5 16×224216× 224^2 81.0 94.6 Uniformer-S Li et al. (2022a) CT 21 42×1×442× 1× 4 16×224216× 224^2 80.8 94.7 WLiT Sun et al. (2023) T 22 21×1×421× 1× 4 8×22428× 224^2 74.6 92.0 TimeSformer-L Bertasius et al. (2021) T 121 2380×3×12380× 3× 1 16×224216× 224^2 80.7 94.7 Uniformer-S Li et al. (2022a) CT 311 3992×3×43992× 3× 4 16×224216× 224^2 81.3 94.7 Mformer-HR Patrick et al. (2021) T 311 959×3×10959× 3× 10 16×336216× 336^2 81.1 95.2 VideoMamba-M Li et al. (2024) M 74 202×3×4202× 3× 4 16×224216× 224^2 81.9 95.4 VideoMamba-S Li et al. (2024) M 26 34×3×434× 3× 4 8×22428× 224^2 79.3 94.2 VideoMamba-S Li et al. (2024) M 26 68×3×468× 3× 4 16×224216× 224^2 80.8 94.8 VideoMamba-S Li et al. (2024) M 26 135×3×4135× 3× 4 32×224232× 224^2 81.5 95.2 VideoSEMA (Ours) Ml 31 46×3×446× 3× 4 8×22428× 224^2 81.3 94.9 VideoSEMA (Ours) Ml 31 87×3×487× 3× 4 16×224216× 224^2 82.4 95.4 VideoSEMA (Ours) Ml 31 172×3×4172× 3× 4 32×224232× 224^2 82.6 95.2 Table 2: Performance reported on standard SSv2 dataset. Here F×m×nF× m× n denotes by F the total FLOPs, m the number of testing segments, and n the number of testing crops. Res: resolution, T: transformer. P(M): parameter size in millions. T1: top-1, T5: top-5. C: CNN, CT: CNN +T, M: Mamba, Ml: Mamba-like. Method Type P(M) FLOPs (G) Res T1(%) T5(%) CT-NetR50 Li et al. (2021) C 21 75×1×175× 1× 1 16×224216× 224^2 64.5 89.3 TDNR50 Wang et al. (2021) C 26 75×1×175× 1× 1 16×224216× 224^2 65.3 91.6 WLiT Sun et al. (2023) T 22 50×3×150× 3× 1 16×224216× 224^2 66.3 91.5 VideoMAE Tong et al. (2022) T 22 57×2×357× 2× 3 16×224216× 224^2 66.8 90.3 MViTv1-B Fan et al. (2021) CT 37 71×3×171× 3× 1 16×224216× 224^2 64.7 89.2 VideoMamba-S Li et al. (2024) M 26 34×3×234× 3× 2 8×22428× 224^2 65.2 89.6 VideoMamba-S Li et al. (2024) M 26 68×3×268× 3× 2 16×224216× 224^2 66.0 90.2 VideoSEMA (Ours) Ml 31 46×3×246× 3× 2 8×22428× 224^2 67.2 91.0 VideoSEMA (Ours) Ml 31 87×3×287× 3× 2 16×224216× 224^2 67.3 90.4 5.1 Dataset We evaluate our approach on two widely used large-scale video action recoginition benchmarks: Kinetic-400 (K400) Kay et al. (2017) and Something-Something v2 (SSv2) Goyal et al. (2017). K400 contains around 240000 training, 20000 validation, and 40000 testing videos spanning 400 human action categories, with clips source from YouTube and trimmed to focus on a single action that lasted around 10 seconds. The dataset is focused on human action from diverse scenes, actors, and viewpoints, making it a standard benchmark for learning high-level semantic. In contrast, SSv2 consists of around 220000 videos, with approximately 170000 training, 25000 validation and 27000 testing videos that last around 2-6 seconds. SSv2 places a strong temporal reasoning and fine-grained actions. Together, these datasets provide complementary evaluation settings. 5.2 Setting A common strategy to train a video model is to pretrain a model with an image-based architecture Bertasius et al. (2021); Li et al. (2024); Liu et al. (2022); Tong et al. (2022) on ImageNet-1K or ImageNet-21K, and then inflate the image model into a video model. For fair comparison with VideoMamba, we pretrain the SEMA model on ImageNet-1K, then incorporate the temporal component and train the model on K400 and SSv2 to obtain the VideoSEMA results. For SEMA, we use the same training setup as described in Tran et al. (2026). We use the setting of the T model variant, i.e. 4 stages with hidden dimension of 64, 128, 256, and 512 respectively. To train VideoSEMA, we use a similar setup to VideoMamba. In particular, for K400 we use a set of 5 warmup epochs, 50 total epochs, a 0.35 stochastic depth rate, and 0.05 weight decay, with an initial learning rate of 2×10−42× 10^-4 using the AdamW optimizer. For SSv2, we use a set of 5 warmup epochs, 30 total epochs, a 0.35 stochastic depth rate, and 0.05 weight decay, with initial learning rate of 4×10−44× 10^-4. We train and test VideoSEMA and VideoMamba on 8 NVIDIA RTX A6000 GPUs, each with 46G of memory. Code will be available upon publication. 5.3 Experimental Results (a) example depicts a person peeling potatoes, where VideoSEMA correctly predicts the action with high (P=0.9P=0.9), while VideoMamba incorrectly classifies it as peeling apples. (b) example shows a person drinking shots, which VideoSEMA correctly recognizes with moderate confidence P=0.5P=0.5. (c) example contains a video of a person ripping paper played in reverse time; VideoSEMA correctly identifies the action with low confidence, whereas VideoMamba fails to classify it correctly despite assigning moderate confidence to its prediction. Figure 3: Qualitative comparison of VideoSEMA and VideoMamba on the K400 test set. Shown are representative test samples along with the predicted labels and corresponding highest confidence scores. 5.3.1 K400 We present the overall performance of VideoSEMA on K400 dataset in Table 1. Compared to VideoMamba under a similar computational budget, VideoSEMA’s top-1 accuracies are 1-2% better across different input resolutions. Compared to other state of the art methods under similar training methodologies and computational budgets shown in the first row group of Table 1, VideoSEMA performs significantly better in both top-1 and top-5 metrics. For example, at an input resolution of 16×224216× 224^2, it is 1.4%1.4\% better than MViTv2-S in top-1 and 0.8%0.8\% in top-5. Notably, the second row group of Table 1 includes models with substantially larger parameter counts and FLOPs. Despite operating under much smaller computational budget, VideoSEMA consistently achieves superior performance. This demonstrates that VideoSEMA is a In addition, we present a few visual example comparisons between VideoSEMA and VideoMamba on somewhat challenging predictions in Fig. 3. Example (A) depicts a video of a person peeling potatoes. VideoSEMA correctly predicts the action with a high probability of 90%90\%, while VideoMamba incorrectly classifies it as the similar label peeling apples. The global action of peeling is correct for both models; however, this particular video is challenging due to the local distinction between apples and potatoes, which occupy only a small region of the image. This could be explained by the window attention of SEMA, which keeps the fine details. Example (B) shows a person drinking shots. VideoSEMA correctly identifies the action with a medium probability of 50%50\%, while VideoMamba misidentifies it as tasting beer. Drinking shots involves rapidly consuming a small amount of liquid, while tasting beer is a slower action. This shows that the attention in time between frames allows VideoSEMA to understand quick changes in motion between frames, while VideoMamba processes these tokens in sequential order, thus reducing its temporal capability. Lastly, example (C) shows a video of a person ripping paper recorded in reverse time. VideoSEMA classifies it correctly with low confidence of 20%20\%, while VideoMamba predicts folding napkins. Because the video plays in reverse, the model needs strong temporal understanding. VideoSEMA with temporal attention allows the model to access frames in both directions simultaneously, which helps the model prediction. In contrast, VideoMamba’s forward-direction processing observes the person putting the paper together and thus predicts folding napkins. Table 3: Performance (ablation study) of VideoSEMA using various temporal processing units on K400 dataset. Here F×m×nF× m× n denotes by F the total FLOPs, m the number of testing segments, n the number of testing crops. Time Component # Params (M) FLOPs (G) Res Top-1 (%) 3DConv (ker=7) 26 40×3×440× 3× 4 8×22428× 224^2 80.0 3DConv (ker=7) 26 76×3×476× 3× 4 16×224216× 224^2 79.3 3DConv (ker=13) 26 76×3×476× 3× 4 16×224216× 224^2 80.8 Mamba 36 47×3×447× 3× 4 8×22428× 224^2 76.3 Attention 31 46×3×446× 3× 4 8×22428× 224^2 81.3 5.3.2 SSv2 We present the overall performance of VideoSEMA on the SSv2 dataset in Table 2. We observe that with 8 frames inputs, VideoSEMA achieve a 2%2\% improvement over VideoMamba-S, and with 16 frames inputs, it outperforms VideoMAE by 0.5%0.5\%. Notably, using only 8 frames, VideoSEMA is able to outperform other models that utilize all 16 frames. This demonstrates that VideoSEMA achieves strong efficacy and temporal efficiency. We also note that the performance improvement with 16 frames is smaller compared to that with 8 frames. One possible explanation is the short temporal duration of videos in SSv2, thus limiting the benefit of longer input sequences. Table 4: Top-1 accuracy (%) at higher input resolutions to models pretrained on 16×224216× 224^2 videos of K400, without fine-tuning. Method 16×224216× 224^2 (baseline) 16×512216× 512^2 16×1024216× 1024^2 VideoMamba 80.8 74.4 40.8 VideoSEMA 82.4 75.3 54.6 5.4 Ablation Study We extend the SEMA model from image domain to processing videos by adding a temporal processing unit. There are many standard choices for this component such as convolution, attention, and Mamba. In this subsection, we examine the performance of each choice. For convolution, to maintain relatively low parameter count, we employed a depthwise convolution. For Mamba, we use a single directional Mamba from VideoMamba Li et al. (2024). And lastly, for attention we use softmax full attention. From Table 3, we observe that on K400, convolution performs relatively well, and as the number of frames increases, we need larger temporal convolutional kernel. Mamba on the other hand, performs poorly as a temporal processing unit. One possible explanation is that Mamba is applied only in the temporal dimension, resulting in independent operations on spatial positions. Lastly, since attention provides the best trade-off between accuracy and efficiency, we use attention as a temporal processing unit for VideoSEMA on the datasets here. 5.5 Scalability to Higher Resolution Videos SEMA is designed to handle high-resolution frames. In this subsection, we examine the adaptability of the model when applied to larger input resolutions. We use the model pretrained on 2242224^2 and evaluate it on larger input sizes of 5122512^2 and 102421024^2 with 16 frames on K400. From Tab. 4, we observe that VideoMamba and VideoSEMA perform similarly on 2242224^2; however, at 102421024^2, the performance of VideoMamba degrades much more quickly compared to VideoSEMA. In particular, the performance gap between VideoMamba and VideoSEMA is about 13.8 percentage points. This demonstrates the robustness and scalability of VideoSEMA for large input resolutions. 6 Conclusion We introduced a space-time attention model (VideoSEMA) consisting of a scalable and efficient Mamba-like attention (SEMA) in space and a regular attention in time. For moderate number of video frames, the model out-performs recent space-time transformers and video-mamba of larger (comparable) sizes on benchmark K400 (SSv2) datasets. In future work, we plan to scale the model to longer videos by using dilated/sparse attention Ding et al. (2023) in time and evaluate the model’s robustness on higher spatial resolutions across more datasets. Acknowledgments The work was partly supported by NSF grants DMS-2219904, DMS-2309520, and a Qualcomm Gift Award. NTT was also funded by a Faculty Endowed Fellowship and the Graduate Scholar Success Fund from the University of California, Irvine. References M. Assran, A. Bardes, D. Fan, Q. Garrido, R. Howes, Mojtaba, Komeili, M. Muckley, A. Rizvi, C. Roberts, K. Sinha, A. Zholus, S. Arnaud, A. Gejji, A. Martin, F. R. Hogan, D. Dugas, P. Bojanowski, V. Khalidov, P. Labatut, F. Massa, M. Szafraniec, K. Krishnakumar, Y. Li, X. Ma, S. Chandar, F. Meier, Y. LeCun, M. Rabbat, and N. Ballas (2025) V-jepa 2: self-supervised video models enable understanding, prediction and planning. arXiv:2506.09985. Cited by: §1, §2. G. Bertasius, H. Wang, and L. Torresani (2021) Is space-time attention all you need for video understanding?. In Proceedings of the International Conference on Machine Learning (ICML), Cited by: §1, §2, §3.3, §3.3, §3.3, §5.2, Table 1. J. Ding, S. Ma, L. Dong, X. Zhang, S. Huang, W. Wang, N. Zheng, and F. Wei (2023) LongNet: scaling transformers to 1,000,000,000 tokens. arXiv preprint arXiv:2307.02486. Cited by: §3.2, §6. H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer (2021) Multiscale vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 6824–6835. Cited by: §1, §2, Table 1, Table 2. C. Feichtenhofer (2020) X3D: expanding architectures for efficient video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 1. R. Glowinski, S. J. Osher, and W. Yin (2017) Splitting methods in communication, imaging, science, and engineering. Springer. Cited by: §3.3. R. Goyal, S. E. Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, F. Hoppe, C. Thurau, I. Bax, and R. Memisevic (2017) The “something something” video database for learning and evaluating visual common sense. In 2017 IEEE International Conference on Computer Vision (ICCV), Vol. , p. 5843–5851. External Links: Document Cited by: §5.1. A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré (2020) HiPPO: recurrent memory with optimal polynomial projections. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33, p. 1474–1487. External Links: Link Cited by: §3.1.2. A. Gu and T. Dao (2024) Mamba: linear-time sequence modeling with selective state spaces. In First Conference on Language Modeling, External Links: Link Cited by: §1, §2, §3.1.1, §3.1.2. D. Han, Z. Wang, Z. Xia, Y. Han, Y. Pu, C. Ge, J. Song, S. Song, B. Zheng, and G. Huang (2024) Demystify Mamba in Vision: A Linear Attention Perspective. NeurIPS. Cited by: §1, §3.2. W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman (2017) The kinetics human action video dataset. arXiv:1705.06950. Cited by: §5.1. K. Li, X. Li, Y. Wang, J. Wang, and Y. Qiao (2021) CT-Net: channel tensorization network for video classification. In International Conference on Learning Representations, External Links: Link Cited by: Table 2. K. Li, X. Li, Y. Wang, Y. He, Y. Wang, L. Wang, and Y. Qiao (2024) VideoMamba: state space model for efficient video understanding. arXiv:2403.06977, in ECCV. Cited by: 4th item, §1, §2, §5.2, §5.4, Table 1, Table 1, Table 1, Table 1, Table 2, Table 2. K. Li, Y. Wang, G. Peng, G. Song, Y. Liu, H. Li, and Y. Qiao (2022a) UniFormer: unified transformer for efficient spatial-temporal representation learning. In International Conference on Learning Representations, External Links: Link Cited by: Table 1, Table 1. Y. Li, C. Wu, H. Fan, K. Mangalam, B. Xiong, J. Malik, and C. Feichtenhofer (2022b) MViTv2: improved multiscale vision transformers for classification and detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 4804–4814. Cited by: §1, §2, Table 1. Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo (2021) Swin transformer: hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 10012–10022. Cited by: §3.1.1, §3.2. Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu (2022) Video swin transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 3202–3211. Cited by: §5.2, Table 1. T. Nguyen, T. Nguyen, N. Ho, A. Bertozzi, R. Baraniuk, and S. Osher (2023) A primal-dual framework for transformers and neural networks. in Proc. of ICLR. Cited by: §1. M. Patrick, D. Campbell, Y. Asano, I. Misra, F. Metze, C. Feichtenhofer, A. Vedaldi, and J. F. Henriques (2021) Keeping your eye on the ball: trajectory attention in video transformers. In Advances in Neural Information Processing Systems, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (Eds.), External Links: Link Cited by: Table 1. G. Strang (1968) On the construction and comparison of difference schemes. SIAM Journal on Numerical Analysis 5(3), p. 506–517. Cited by: §3.3. R. Sun, T. Zhang, Y. Wan, F. Zhang, and J. Wei (2023) WLiT: windows and linear transformer for video action recognition. Sensors (Basel, Switzerland) 23 (3), p. 1616. External Links: Document, Link Cited by: §1, Table 1, Table 2. X. Tai, H. Liu, L. Li, and R. H. Chan (2025) A mathematical explanation of transformers. arXiv:2510.03989. Cited by: §1. Z. Tong, Y. Song, J. Wang, and L. Wang (2022) VideoMAE: masked autoencoders are data-efficient learners for self-supervised video pre-training. In Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35, p. 10078–10093. External Links: Link Cited by: §5.2, Table 2. N. T. Tran, F. Xue, S. Zhang, J. Lyu, Y. Zheng, Y. Qi, and J. Xin (2026) SEMA: a scalable and efficient mamba like attention via token localization and averaging. arXiv:2506.08297; in Proc. of ICML. External Links: 2506.08297, Link Cited by: 1st item, §1, §2, §3.1.1, §3.1.1, §3.2, §5.2. C. Wang, K. Li, T. Jiang, X. Zeng, Y. Wang, and L. Wang (2025a) Make your training flexible: towards deployment-efficient video models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 23880–23891. Cited by: §1, §2. L. Wang, Z. Tong, B. Ji, and G. Wu (2021) TDN: temporal difference networks for efficient action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 1895–1904. Cited by: Table 2. Y. Wang, K. Li, X. Li, J. Yu, Y. He, G. Chen, B. Pei, R. Zheng, Z. Wang, Y. Shi, T. Jiang, S. Li, J. Xu, H. Zhang, Y. Huang, Y. Qiao, Y. Wang, and L. Wang (2025b) InternVideo2: scaling foundation models for multimodal video understanding. In Computer Vision – ECCV 2024, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol (Eds.), Cham, p. 396–416. External Links: ISBN 978-3-031-73013-9 Cited by: §1, §2. Y. Zheng, Z. Xu, F. Xue, B. Yang, J. Lyu, S. Zhang, Y. Qi, and J. Xin (2024) AFIDAF: Alternating Fourier and Image Domain Adaptive Filters as an Efficient Alternative to Attention in ViTs. International Symposium of Visual Computing, Reno, NV 15046, p. 17–30. Cited by: §1.