Paper deep dive
LITEWAY: LIghtweight HAR via Temporal Efficient highWAY
Dominique Nshimyimana, Vitor Fortes Rey, Mengxi Liu, Bo Zhou, Paul Lukowicz
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Wearable human activity recognition (HAR) remains challenging due to the computational and energy constraints of deep learning models on resource-limited devices. Existing lightweight approaches often rely on recurrent architectures (e.g., GRU and LSTM), limiting parallelism and increasing inference latency. We propose LITEWAY, a modality-agnostic, fully convolutional framework for multichannel sensor time series that replaces recurrent temporal modeling with structured convolutional decomposition. LITEWAY combines lightweight convolutional blocks, strided temporal processing, and convolution-attention pooling to efficiently capture temporal dependencies while reducing computational complexity. We evaluate LITEWAY on 16 HAR datasets against TinyHAR, TinierHAR, and MLP-HAR. LITEWAY achieves competitive macro F1 while reducing model size by 4.06x-9.52x (Light) and 3.87x-9.07x (Full) compared with TinyHAR and TinierHAR. Deployment experiments further show energy reductions of 2.29x-3.14x (Light) and 1.46x-2.01x (Full) compared with TinierHAR and MLP-HAR, highlighting efficient fully convolutional temporal modeling for wearable HAR. The source code is publicly available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.09421v1
- Canonical: https://arxiv.org/abs/2608.09421v1
Trouble viewing inline? Open PDF directly →
Full Text
43,119 characters extracted from source content.
Expand or collapse full text
LITEWAY: LIghtweight HAR via Temporal Efficient highWAY Dominique Nshimyimana ✉ RPTU, DFKI Kaiserslautern, Germany dominique.nshimyimana@dfki.de Vitor Fortes Rey RPTU, DFKI Kaiserslautern, Germany vitor.fortes_rey@dfki.de Mengxi Liu RPTU, DFKI Kaiserslautern, Germany mengxi.liu@dfki.de Bo Zhou RPTU, DFKI Kaiserslautern, Germany bo.zhou@dfki.de Paul Lukowicz RPTU, DFKI Kaiserslautern, Germany Paul.Lukowicz@dfki.de Abstract Wearable human activity recognition (HAR) remains challenging due to the computational and energy constraints of deep learn- ing models on resource-limited devices. Existing lightweight ap- proaches often rely on recurrent architectures (e.g., GRU and LSTM), limiting parallelism and increasing inference latency. We propose LITEWAY, a modality-agnostic, fully convolutional framework for multichannel sensor time series that replaces recurrent temporal modeling with structured convolutional decomposition. LITEWAY combines lightweight convolutional blocks, strided temporal pro- cessing, and convolution-attention pooling to efficiently capture temporal dependencies while reducing computational complex- ity. We evaluate LITEWAY on 16 HAR datasets against TinyHAR, TinierHAR, and MLP-HAR. LITEWAY achieves competitive macro F1 while reducing model size by 4.06×–9.52×(Light) and 3.87×– 9.07×(Full) compared with TinyHAR and TinierHAR. Deploy- ment experiments further show energy reductions of 2.29×–3.14× (Light) and 1.46×–2.01×(Full) compared with TinierHAR and MLP- HAR, highlighting efficient fully convolutional temporal model- ing for wearable HAR. The source code is publicly available at https://github.com/dominique-nshimyimana/liteway CCS Concepts • Human-centered computing→ Ubiquitous and mobile de- vices;• Computer systems organization→Embedded systems; • Computing methodologies→ Neural networks. Keywords Time series, computing methodologies, human activity recognition, edge AI 1 Introduction Wearable sensor-based human activity recognition has attracted sig- nificant research interest due to applications in healthcare [37,40], sports analytics [29,41], smart homes [11], and industrial safety monitoring [6,31,34]. These systems use inertial sensors such as accelerometers and gyroscopes to recognize human activities from multivariate time-series signals, enabling applications in personal- ized health monitoring and assisted living. Deploying HAR models on wearable devices remains challenging due to limited memory, computational capability, and battery ca- pacity [21]. Many high-performing HAR models require substantial computational resources, limiting their practicality for real-time on- device inference [26]. As a result, improving hardware efficiency while maintaining strong recognition performance has become increasingly important. Recent lightweight HAR models such as TinierHAR [9], Tiny- HAR [43], and MLP-HAR [42] have reduced model complexity while maintaining competitive accuracy. However, many approaches still rely on recurrent temporal modeling (RNN) or computationally intensive feature extraction. Although recurrent networks such as GRU and LSTM model temporal dependencies, their sequential computation limits parallelization and can introduce latency and energy overhead on resource-constrained hardware [19, 22]. To address these limitations, we propose LITEWAY, a fully convo- lutional HAR framework for efficient on-device inference. LITEWAY uses only convolutional layers and a single linear layer, achieving low computational cost and strong hardware efficiency while main- taining competitive recognition performance. This paper makes the following contributions. (i) We propose LITEWAY, a fully convolutional HAR framework that replaces RNN with structured convolutional decomposition, enabling efficient temporal modeling and being resource-aware. (i) We introduce a lightweight architecture optimized for wearable HAR, balancing memory, compute, and representation capacity via modular convo- lutional blocks. (i) Evaluated on 16 datasets, LITEWAY achieves competitive macro F1, with the Light and Full variants reducing model size by 4.06×–9.52×and 3.87×–9.07×, respectively, com- pared to TinierHAR and TinyHAR. (iv) Ablation and deployment show LITEWAY Light maximizes efficiency with 2.29×–3.14×lower energy, while LITEWAY Full improves macro F1 with 1.46×–2.01× lower energy, both versus TinierHAR and MLP-HAR, highlighting the accuracy–efficiency trade-off. 2 Related Work Deep HAR for Wearable Sensing. Deep learning has become the dominant approach for wearable HAR due to its ability to learn discriminative representations from multimodal sensor streams. Early architectures such as DeepConvLSTM (DCL) [22] combined convolutional layers for local feature extraction with recurrent layers for sequence modeling, establishing a widely adopted CNN- RNN paradigm for inertial sensing applications. Despite their effectiveness, recurrent architectures require se- quential computation, limiting parallelism and increasing inference latency on resource-constrained wearable devices. To address these limitations, several works explored convolution-based temporal arXiv:2608.09421v1 [cs.LG] 10 Aug 2026 Nshimyimana et al. modeling approaches for sequence processing [4,35], demonstrat- ing strong performance while enabling low latency. However, capturing long-range temporal dependencies through deeper or dilated convolutions may still increase computational cost and memory usage, particularly on resource-constrained wear- able platforms [43]. Consequently, efficient long-range temporal modeling remains a key challenge for real-time HAR. Lightweight HAR Architectures. Several studies have explored efficient HAR models for on-device inference. TinyHAR [43] pro- posed a lightweight architecture optimized for edge deployment by reducing computational complexity through efficient convolutional operations and compact feature extraction modules. Beyond reducing parameter count, lightweight HAR research has focused on balancing temporal modeling capability with de- ployment efficiency. For example, TinierHAR [9], SPECTRA [15] and MLPHAR [42] further reduced model complexity while pre- serving competitive performance. However, existing approaches often rely on CNN–RNN architectures. Although MLPHAR does not use recurrent modules, it is not fully end-to-end learnable. Although prior lightweight architectures improve efficiency, ex- isting methods still face challenges in jointly optimizing tempo- ral receptive field, inference latency, and model compactness for streaming wearable HAR. These limitations motivate the design of convolution-only architectures that provide efficient temporal mod- eling while remaining suitable for low-power HAR deployment. 3 Methodology We propose LITEWAY (Figure 1), a fully convolutional architec- ture for time-series classification on resource-limited devices. The model has three components: (i) a feature extraction backbone that downsamples and refines features using residual and depth- wise-separable convolutions with attention-based channel recali- bration; (i) a Structured Convolutional Temporal Modeling (SCTM) module capturing long-range dependencies via depthwise convolu- tions, shared projections, and gated pathways without recurrence; (i) a lightweight classification head that aggregates features via attention-based pooling followed by a linear layer. Each component minimizes redundant parameters while preserving capacity. 3.1Convolutional Feature Extraction Backbone The goal of the backbone is to reduce temporal resolution early while avoiding expensive feature transformations and refinement. The backbone consists of six convolutional blocks in two stages. Step 1: Temporal Downsampling. The first stage uses two residual blocks with batch norm and leaky-relu for temporal downsam- pling. The first block applies residual depthwise convolution, while the second block performs depthwise separable convolution. Each block is followed by pointwise mixing. This design enables efficient temporal downsampling. Step 2: Feature refinement. The subsequent four layers operate on reduced temporal resolution. These blocks use depthwise separable convolutions to decouple temporal filtering from channel mixing while preserving representational capacity. Each block also includes a squeeze-and-excitation (SE) submodule [18,27]. We implement SE using 1×1 convolutions for channel attention to maintain a fully convolutional design. For the efficient variant, we replace conventional convolution and pooling with strided depthwise convolutions (StrideConv), which combine feature extraction and downsampling to reduce MACs. Consequently, we remove the SE modules at the cost of a tolerable performance loss. Overall, this backbone establishes an early and sustained re- duction in feature dimensionality, forming the foundation for a compact model design. 3.2 SCTM SCTM is the core component of LITEWAY and integrates efficient feature transformation principles from gated architectures and multi-branch convolutional designs. Rather than introducing a new gating operation, SCTM combines these concepts into a wearable- oriented temporal modeling module optimized for HAR efficiency. 3.2.1 SCTM-Full. This block draws on and unifies several estab- lished design principles into a single parameter-efficient module. Given an input푋 ∈R 퐵×퐶×푇 , the block first applies a depthwise temporal convolution followed by a pointwise activation,퐻= 휙(DWConv(푋)), decoupling temporal filtering from channel mix- ing in the spirit of depthwise separable convolutions [17], which have been shown to approximate full convolutions at a fraction of the parameter cost. A shared pointwise projection푍=푊 푝 (퐻) then produces a single latent representation reused across both pathways, avoiding the parameter duplication inherent in standard two-branch designs such as the Gated Linear Unit [14], where two independent projections푊 1 ,푊 2 are learned. The block then con- structs two complementary signals. The first,푌 푓 = 휎(푍)⊙ tanh(푍), is a collapsed Gated Tanh Unit [35,36] in which the filter and gate weights are tied to the same projection; the sigmoid acts as a soft content gate, selecting which features to pass, while the tanh provides a bounded, zero-centered nonlinear transformation, a combination empirically shown to outperform rectified activa- tions for sequential and audio modeling [14,35]. The second signal, 푌 푏 =(1−휎(푍))⊙푊 푝 (푋), applies the complement of the same gate to a projection of the raw input, directly instantiating the carry gate of Highway networks [30], where퐶=1−푇was introduced to allow unimpeded information flow through deep networks. Critically, the gate is derived from푍, which encodes temporal structure via the preceding depthwise convolution, rather than from the raw input as in the original highway formulation; the gating decision is there- fore informed by processed temporal features rather than channel statistics alone. Similarly, the carry stream applies푊 푝 to푋rather than bypassing it as an identity, ensuring that even the preserved pathway undergoes channel mixing, preserving the complemen- tary relationship between the two streams while maintaining a shared projection. The two streams are fused by channel-wise con- catenation,푌= Concat(푌 푓 ,푌 푏 ), rather than by the addition used in highway networks [30] and residual connections [16]. Additive fusion combines transformed and carried information into a single representation, whereas concatenation preserves both streams sep- arately and allows subsequent layers to learn their interaction. This design is motivated by the split-transform-merge strategy employed in Inception architectures [26,32], where projected representations are processed independently before concatenation. Similarly, SCTM applies complementary transformations within parallel pathways LITEWAY: LIghtweight HAR via Temporal Efficient highWAY DWConv(k=5) BN ... 1-x ... C T C T/2 F 2F T/4 C C 2F T/4 C T/4 Cx2F Residual Downsample Conv(k=(5,1)) + BN + ReLU+ Pool Residual Separable Downsample Conv(k=(5,1)) + BN + ReLU + Pool Feature Refinement: 4x ConvAttention Conv(k=(5,1)) + BN + ReLU + SE Flatten SCTM-Full: Conv1D(k=5) & Conv1D(k=1) Global Temporal Pooling & Classify Conv1D attention T/4 F * X T Conv1D (k=1) Softmax (0~1) (Bx1xT/4) Linear ... Activity Sum T/4 F * T/4 Cx2F F * /2 F * /2 F * /2 T/4 T/4 T/4 F * /2 T/4 F * /2 T/4 * W p Sigmoid * W p Tanh * W p FULL LIGHT T/4 F * T/4 T/4 Cx2F DWConv(k=5) BN + GELU ... C T C T/2 F 2F T/4 C C 2F T/4 C T/4 Cx2F Residual Downsample Strided Conv(k=(5,1)) + BN + Leaky ReLU Residual Separable Downsample Strided Conv(k=(5,1)) + BN + Leaky ReLU Feature Refinement: 4x Separable Convolution Conv(k=(5,1)) + BN + ReLU Flatten SCTM-Light: Conv1D(k=5) & Conv1D(k=1) Global Temporal Pooling & Classify: Conv1D attention T/4 F * X T Conv1D (k=1) Softmax (0~1) (Bx1xT/4) Linear ... Activity Sum ... F * /2 F * /2 T/4 * W p * W p + ELU * W p 1x1 convolution (1D); weight-shared F=4; F * =32 For feature size multiplication concatenation Figure 1: LITEWAY architectures (Full and Light), based on a fully convolutional design with a single linear classification layer and merges them through concatenation, increasing local repre- sentational capacity without requiring multiple full-dimensional transformations. Together, these decisions yield a block that per- forms temporal modeling through a single depthwise convolution and a single shared projection, without recurrent state, without sep- arate branch weights, and without additive fusion losses, offering a principled reduction in parameter count relative to both recurrent models [13] and standard gated convolutional baselines [14]. 3.2.2SCTM-Light. The lightweight variant of SCTM that reduces computational cost. Like the full block, it applies a depthwise tem- poral convolution followed by a pointwise projection with GELU activation. To further reduce MACs, a residual pathway is pro- jected using ELU and concatenated with the main stream, yielding a compact yet expressive output. This design retains a compressed residual shortcut for the input while omitting separate gate multipli- cation, simplifying computation while preserving complementary feature flow. Formally,푌= Concat 푍, ELU(푊 p (푋)) , where푍is the GELU-activated projection of the depthwise convolution. 3.3 Global Temporal Pooling and Classification We aim to aggregate temporal features without introducing addi- tional heavy sequence modeling. Step 1: We use attention-based temporal pooling with a single learnable projection that computes importance weights over time steps i.e.훼= softmax(푊 푎 푋). Step 2: Aggregation. The final rep- resentation is a weighted sum of temporal features, producing a compact global embedding. Step 3: Classification. A single linear layer maps this embedding to output classes, ensuring minimal parameter overhead in the decision stage. 3.4 Efficiency design choices Beyond architectural design, efficiency is enforced through system- atic reduction of redundant computation, especially when param- eter optimization is sensitive. (i) Strided Convolutions: Replace pooling operations to eliminate redundant layers. (i) 1D Convo- lution over Recurrent Models: Avoid sequential hidden-state computations. (i) Lightweight Activations: Prefer lightweight activation where appropriate. (iv) Selective Residual Connec- tions: Apply only where optimization stability requires it. Collec- tively, these design choices ensure parameter and computational efficiency throughout the entire architecture. 4 Experimental Results 4.1 Experiment Setup Datasets and Preprocessing. We evaluate the proposed method on 16 widely used HAR datasets covering diverse sensing modalities, sampling frequencies, and activity types (Table 1). All sensor signals are segmented using dataset-specific sliding windows with 50% overlap. Each sensor channel is independently standardized using the mean and standard deviation computed from the training set only. Oppo and oppoloc share data but use different labels. Nshimyimana et al. Table 1: Summary of the evaluated HAR datasets. #푆푢푏푗de- notes the number of subjects, #퐶푙푠the number of activity classes,퐶ℎthe number of sensor channels,퐹(Hz) the sam- pling frequency, and 푆푊 the sliding-window in seconds. DatasetSensor#Subj #Cls Ch 퐹(Hz) SW Dg [3]Acc1099641 Uschad [39]Acc/Gyro71261001 Skodar [38]Acc11030334 Pamap2 [23]Acc/Gyro/Mag91218334 Dsads [1]Acc/Gyro/Mag81945254 Hapt [24]Acc/Gyro10126502.56 Rw [33]Acc15821504 Oppo [25]IMU/Mag/Quat41877304 Oppoloc [25]IMU/Mag/Quat4677304 Recgym [10]Acc/Gyro/Cap10712204 MotionSense [20] Acc/Gyro24126504 Mhealth [5]IMU/ECG101223504 Sho [28]IMU/LAcc10760504 Uci [2]Acc/Gyro/LAcc3069502.56 Realdisp [25]Acc/Gyro/Mag/Quat173381504 Wear [12]Acc221912504 Evaluation Protocol. We follow a subject-independent evaluation protocol for robust results. For most datasets, Leave-One-Subject- Out (LOSO) cross-validation is used to assess generalization to unseen users. For large-scale datasets (MotionSense and uci), group- based subject hold-out is adopted to reduce training time. An ex- ception is skodar, which contains a single subject; thus, Leave-One- Session-Out is used. Training and Metrics. All experiments are conducted on NVIDIA RTX 3090 GPU. To ensure reproducibility and reduce variance due to random initialization, each experiment is repeated with five ran- dom seeds (1–5), and the average performance is reported. Models are trained for up to 150 epochs using the AdamW optimizer with cross-entropy loss. The initial learning rate of 1×10 −3 is reduced by a factor of 0.1 if no improvement is observed for 7 epochs. Early stopping is applied with a patience of 15 epochs. We report Macro- F1 (퐹1 푀 ) as the primary performance metric due to class imbalance across datasets. In addition, model efficiency is evaluated using the number of parameters (푛푃) and MACs (multiply–accumulate oper- ations), enabling a direct accuracy–efficiency trade-off assessment. Baselines. We compare the proposed method against representa- tive HAR models, including TinierHAR, TinyHAR, and MLP-HAR as state-of-the-art efficient models, as well as DeepConvLSTM, the most commonly reported efficient baseline. All baselines are evalu- ated under the same training and evaluation protocol to ensure a fair comparison. 4.2 Experimental Results 4.2.1Per-Dataset Performance. We first evaluate all methods across 16 datasets to assess generalization ability. Figure 2 reports detailed results for each dataset in terms of (1) macro F1 score, (2) model complexity (MACs), and (3) number of parameters. The LITEWAY Light variant achieves top-2 macro F1 scores on 9 out of 16 datasets. It ranks first on mhealth, motionsense, and recgym, and achieves second place on dsads, hapt, pamap2, realdisp, sho, and skodar. Similarly, the LITEWAY Full variant reaches top-2 performance on 10 datasets, securing first place on dg, motionsense, and pamap2, and second place on dsads, hapt, mhealth, realdisp, rw, sho, and skodar. These results demonstrate that both LITEWAY variants consistently rank near the top despite using substantially fewer parameters and MACs than larger models such as TinyHAR. Even on more challenging datasets such as oppo and oppoloc, where macro F1 scores are lower across all methods, LITEWAY Light and Full maintain competitive rankings, typically second or third. This highlights their robustness and strong generalization ability under diverse and difficult conditions. Overall, these findings indicate that the proposed methods provide a favorable balance of high accuracy and efficiency across all evaluated datasets. Finding: Both LITEWAYs achieve comparable performance across divers datasets, demonstrating that our proposed methods deliver strong generalization while remaining lightweight and efficient. 4.2.2 Overall Performance. Figure 3 summarizes the average per- formance across all 16 datasets, reporting macro F1-score alongside relative computational cost (MACs) and model size (parameters). Among all methods, LITEWAY Full achieves the highest macro F1-score (0.813), followed by LITEWAY Light (0.808). Both models outperform existing baselines, including TinierHAR (0.801), Tiny- HAR (0.805), MLPHAR (0.807) and DeepConvLSTM (0.801). In terms of efficiency, LITEWAY Light requires the lowest cost (988.8K MACs) and model size (6.5K parameters), while LITEWAY Full maintains a similarly compact footprint with only a small increase in complexity. Compared to existing baselines, this corre- sponds to reductions of approximately 2.51×–146×in MACs and 4.06×–176×in parameters, highlighting the substantially lower resource requirements of the proposed designs. Finding: LITEWAY variants achieve competitive macro F1, while LITEWAY Light sets the lowest model size and MAC. 4.2.3 Trade-off between Efficiency and Performance. We analyze the relationship between efficiency and recognition performance, as shown in Figure 3. Two key trends emerge: (i) increasing model complexity does not consistently improve accuracy, with gains often saturating or remaining limited despite higher MACs and parameter counts, and (i) parameter count is a poor proxy for computational cost, as models with similar sizes can exhibit substantially different MACs. In contrast, the proposed models operate in a more efficient regime. LITEWAY Light represents the extreme low-cost setting, while LITEWAY Full achieves higher accuracy with only marginal additional cost, indicating better utilization of model capacity. Finding: The proposed networks achieve substantially lower energy consumption without compromising recognition perfor- mance. 4.3 Ablation Study 4.3.1Ablation setup. The impact of key architectural components was evaluated using the following ablations. Three residual config- urations: Res-Down (default, residuals in early layers), Res-All (residuals in all layers), and NoRes (no residuals). Channel recalibration was studied via LightSE, an efficient vari- ant of squeeze-and-excitation that reduces computation while pro- viding a middle ground between SE and no-SE (NoSE). LITEWAY: LIghtweight HAR via Temporal Efficient highWAY 60% 65% Macro F1 66.1±2.162.6±4.063.7±1.961.5±2.4 dg 500K 1M 1.5M MACs 226.6K122.6K318.7K1.65M 10K 20K Params 2.35K2.16K9.46K29.7K 85% 90% 90.3±1.288.5±0.985.7±1.285.1±1.6 dsads 5M 10M 1.77M951.8K2.29M14.6M 20K 40K 9.54K9.35K37.7K59.2K 75% 80% 76.9±1.378.8±1.281.7±0.781.1±0.3 hapt 1M 2M 302.7K164.5K446.4K2.69M 10K 20K 2.13K1.94K7.48K27.7K 87% 90% 92% 91.1±1.493.0±1.792.5±0.890.6±1.1 mhealth 5M 10M 15M 1.81M974.2K2.39M15.8M 15K 30K 5.26K5.07K20.5K41.3K 87% 90% 92% 92.7±1.291.7±0.391.7±0.591.8±1.0 motionsense 2.5M 5M 7.5M 942.8K509.8K1.29M8.31M 15K 30K 3.04K2.85K11.9K32.2K 35% 37% 38.6±2.337.7±1.637.9±1.638.8±1.5 oppo 10M 20M 30M 3.62M1.95M4.66M33.0M 30K 60K 15.4K15.2K62.2K84.7K 65% 70% 75% 76.0±4.073.0±2.467.5±2.472.8±3.0 oppoloc 10M 20M 30M 3.62M1.95M4.66M33.0M 30K 60K 15.0K14.8K61.8K84.2K 70% 75% 79.3±1.576.2±1.473.3±2.072.9±3.4 pamap2 2.5M 5M 7.5M 933.1K503.9K1.25M7.76M 15K 30K 4.34K4.15K16.7K37.3K 85% 87% Macro F1 90.0±0.889.5±0.788.1±0.689.3±0.3 realdisp 25M 50M MACs 6.35M3.42M8.17M62.3M 30K 60K Params 16.6K16.4K65.8K88.5K 87% 90% 92% 91.1±0.490.9±0.292.7±0.392.0±0.2 recgym 500K 1M 1.5M 220.8K119.9K319.0K1.75M 10K 20K 2.31K2.12K8.25K28.5K 65% 70% 72.2±1.068.9±2.168.1±1.769.5±1.5 rw 5M 10M 1.65M889.7K2.19M14.4M 15K 30K 4.76K4.57K18.9K39.5K 90% 95% 98.7±1.098.1±1.194.6±1.498.9±0.8 sho 20M 40M 4.71M2.54M6.08M44.0M 25K 50K 11.9K11.7K48.8K70.7K 92% 95% 97% 97.5±0.898.0±0.398.3±0.498.9±0.2 skodar 5M 10M 1.55M838.0K2.04M13.1M 15K 30K 45K 6.48K6.29K25.8K46.8K 90% 95% 92.5±2.896.6±0.396.0±0.596.5±0.5 uci 1M 2M 3M 453.0K245.3K637.4K3.86M 10K 20K 2.48K2.29K9.59K29.8K 65% 70% 66.0±0.868.0±0.570.7±0.969.3±1.1 uschad 1M 2M 236.6K128.6K348.8K2.01M 10K 20K 2.13K1.94K7.48K27.7K 75% 80% 81.4±0.580.7±1.179.4±0.778.8±0.6 wear 2.5M 5M 7.5M 943.3K510.3K1.30M8.31M 15K 30K 3.47K3.27K12.3K32.8K LITEWAY-FLITEWAY-LTinierHARTinyHAR Figure 2: Performance of LITEWAYs in terms of (1) macro F1 score, (2) MACs, and (3) number of parameters, compared with TinierHAR, and TinyHAR. The average across 16 datasets is shown in Figure 3. 10 0 10 1 10 2 MACs (log) LITEWAY-L LITEWAY-F TinierHAR MLPHAR TinyHAR DeepConvLSTM 988.8K 1.85x 2.51x 9.88x 16.02x 146x 10 0 10 1 10 2 #Params (log) 6.5K 1.05x 4.06x 102x 9.52x 176x 0.7500.7750.800 Macro F1 0.808 0.813 0.801 0.807 0.805 0.801 Figure 3: Comparison of macro F1, MACs, and parameters. Aggregation strategies were compared, including ConvAtt (pro- posed convolutional attention), LinAtt (linear attention, similar to TinyHAR and TinierHAR), and MMX (max-mean pooling), to assess their influence on accuracy and efficiency. Finally, activation strategies were compared, including homoge- neous GELU, homogeneous Leaky ReLU, GELU→Leaky (uniform replacement of GELU by Leaky ReLU), and a heterogeneous design in LITEWAY Light, where blocks use ReLU, Leaky ReLU, or GELU. This setup allows us to quantify the contributions of residual connections, attention, and activation functions to both accuracy and efficiency. 4.3.2Design Ablation. Table 2 evaluates the impact of LITEWAY ar- chitectural blocks on performance and efficiency. Parameter counts are similar across variants (6.7–6.9K), but MACs vary significantly. LITEWAY Full achieves the highest퐹1 푀 (81.3), serving as the perfor- mance reference. Removing or simplifying components generally Table 2: Ablation of LITEWAY variants; each row shows a single architectural modification. Empty cells indicate the default modules used in LITEWAY Full. DownsampleRefineTemporalAggregate퐹1 푀 nPMAC LITEWAY-FRes-DownSESCTM-FConvAtt81.36.7K1.8M NoResNoRes80.66.6K1.6M StrideConvStrideConv80.76.7K1.3M LightSE LightSE80.66.6K1.8M NoSE NoSE80.96.5K1.8M SCTM-L SCTM-L80.66.7K1.6M LinAttLinAtt80.86.7K1.8M MMXMMX79.76.7K1.8M LITEWAY-LStrideConvNoSESCTM-LConvAtt80.86.5K989K reduces퐹1 푀 : NoRes and SCTM-L decrease퐹1 푀 by 0.7 points, while MMX achieves the lowest퐹1 푀 (79.7), highlighting the importance of ConvAtt aggregation. LITEWAY Light achieves a favorable trade- off, reducing MACs by over 45% relative to the full model, with only a minor 퐹1 푀 drop (from 81.3 to 80.8). Finding: Residual connection, SCTM, and ConvAtt are criti- cal for high퐹1 푀 , while efficient blocks allow LITEWAY Light to maintain strong performance with minimal resource demand. 4.3.3 Activation function. We evaluate activation strategies (Ta- ble 4) considering homogeneous GELU, homogeneous Leaky ReLU, GELU→Leaky replacement, and the heterogeneous design in LITE- WAY. Homogeneous ReLU was excluded as it is less accurate. Nshimyimana et al. Model퐹1 푀 nPMACs NoRes80.16.4K904K Res-All80.06.8K1.3M LITEWAY-L 80.86.5K989K Table 3: Effect of residual con- nections (16 datasets). ModelMACnPF1 GELU1.4M6.5K80.6 Leaky ReLU1.3M6.5K80.5 GELU→ Leaky978K6.5K80.4 LITEWAY-L989K6.5K80.8 Table 4: Effect of activation functions (16 datasets). Table 5: Comparison of SOTA and LITEWAY on STM32L4S5 Model Inf. Time (ms) Weight (KiB) Activation (KiB) Cycles /MAC CPU (% load) Energy (mJ/Inf ) TinyHAR249.01± 0.07107.4862.7013.5124 19.14± 1.03 MLPHAR114.81± 0.02342.4140.5411.99118.73± 0.39 TinierHAR81.42± 0.0139.5414.4325.5886.36± 0.14 LITEWAY-F56.71± 0.0010.6316.3533.8654.35± 0.21 LITEWAY-L37.44± 0.0010.0716.0330.5432.90± 0.16 Activation functions have limited impact on accuracy (≤ 0.2 F1 variation) but affect computation. Leaky ReLU lowers MACs com- pared to GELU, and GELU→Leaky reduces computation from 1.4M to 978K MACs without changing parameter count. The proposed LITEWAY Light achieves the best trade-off, obtaining the highest macro F1 score (80.8) with low cost (989K MACs). This indicates that block-wise assignment of activation functions is more effective than using a single activation throughout the network. Finding: Mixed activations yield better accuracy–efficiency bal- ance than uniform designs. 4.3.4Effect of Residual Connection. We evaluate residual connec- tions using three configurations: NoRes, Res-All, and LITEWAY (Res-Down). Table 3 summarizes results across 16 datasets. NoRes achieves the lowest execution cost (904K MACs) with slightly fewer parameters while maintaining competitive perfor- mance (80.1 F1). Res-All increases computation (1.3M MACs, 6.8K params) without improving accuracy (80.0 F1), indicating limited benefit from applying residuals uniformly across all layers in light- weight HAR models. In contrast, LITEWAY achieves the best trade-off (80.8 F1) with near-NoRes cost (989K MACs, 6.5K params), showing that residual connections are most effective when applied only in early layers. Finding: Selective residual connections in early layers are suffi- cient for stable optimization in lightweight HAR models. 4.3.5 Deployment on Hardware. To evaluate our models on edge devices, we deploy them on the low-power STM32L4S5 microcon- troller running at 120 MHz). Power is measured with the ST X- NUCLEO-LPM01A shield. Average inference time is recorded over 16 cycles via the system clock and UART, and energy is measured after disabling all GPIOs. Table 5 summarizes the hardware effi- ciency metrics, including inference latency, memory footprint, CPU utilization, and energy consumption. Among all evaluated methods, LITEWAY Light achieves the best overall efficiency on the STM32L4S5, requiring only 37.44 ms per inference, 3% CPU load, and 2.90 mJ per inference. Compared with TinierHAR, it reduces both latency and energy by more than 2×, while using only 10.07 KiB of weights, corresponding to an approximately 4×smaller parameter footprint. Although LITEWAY has a higher cycles/MACC than some baselines, its substantially lower MAC count (Figure 3) leads to lower total execution cost. Finding: LITEWAY achieves superior embedded deployment efficiency across latency, memory, CPU, and energy. 4.4 Discussion Summary and Future Work Generalization. LITEWAY models show consistent performance across diverse HAR datasets with varying modalities and complexi- ties. Even on more challenging datasets (e.g., dg, oppo), they remain competitive, demonstrating robust feature learning despite their lightweight design. Following the Bayesian analysis of classifier comparisons advocated by Benavoli et al. [7], we ran the Bayesian signed-rank test [8] on the per-dataset macro-F1 scores across the 16 datasets, using a region of practical equivalence (ROPE) of one F1 point. For LITEWAY-F, the posterior probability of outperforming each baseline ranges from 0.59 to 0.86 and never favours a baseline (all푃(baseline better) ≤0.13); since no comparison crosses the conventional 0.95 decision threshold, we conclude that our archi- tecture is at least on par with, and most likely superior to, the SOTA baselines while remaining substantially smaller. Efficiency–Accuracy Trade-off. Increasing model size does not reliably improve performance. Larger models raise cost with limited gains, while the proposed models maintain strong accuracy with far lower MACs and parameters. The low energy demand and latency of LITEWAY further confirm its superiority, placing the models on the Pareto frontier and emphasizing efficient design over scale. Design Insights. Ablation results show that (i) residual connec- tions are most effective in early layers, (i) attention-based aggre- gation improves performance over simpler methods, (i) temporal modeling remains essential, and (iv) heterogeneous activations pro- vide better efficiency–accuracy balance than uniform choices. Limitations and Future Work. LITEWAY was evaluated on a single microcontroller, which may limit generalization across hardware. Real-world deployment could require hardware- and data-aware adaptations. We did not explore further optimizations such as quan- tization, pruning, hardware-specific acceleration, or architectural choices including SCTM depth, kernel size, and multi-sensor fusion. Improving cycle/MAC efficiency is also left for future work. 5 Conclusion This paper presents LITEWAY, an efficient HAR framework that substantially reduces computation and model size. It replaces recur- rent architectures with structured convolutions, enabling efficient, expressive feature learning. Ablation shows MACs can be reduced despite limited parameter compression. Low energy use and fast inference confirm deployability on resource-constrained devices. Acknowledgments This research was supported by the Carl Zeiss Stiftung, Germany, through the Sustainable Embedded AI project (P2021-02-009) and by BMFTR in the project Cross-Act (01IW25001). LITEWAY: LIghtweight HAR via Temporal Efficient highWAY References [1]Kerem Altun, Billur Barshan, and Orkun Tunçel. 2010. Comparative study on classifying human activities with miniature inertial and magnetic sensors. Pattern Recognition 43, 10 (2010), 3605–3620. [2]D. Anguita, Alessandro Ghio, L. Oneto, Xavier Parra, and Jorge Luis Reyes- Ortiz. 2013. A Public Domain Dataset for Human Activity Recognition using Smartphones. In The European Symposium on Artificial Neural Networks. https: //api.semanticscholar.org/CorpusID:6975432 [3]Marc Bachlin, Daniel Roggen, Gerhard Troster, Meir Plotnik, Noit Inbar, Inbal Meidan, Talia Herman, Marina Brozgol, Eliya Shaviv, Nir Giladi, et al.2009. Potentials of enhanced context awareness in wearable assistants for Parkin- son’s disease patients with the freezing of gait syndrome. In 2009 International Symposium on Wearable Computers. IEEE, 123–130. [4]Shaojie Bai, J Zico Kolter, and Vladlen Koltun. 2018. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint arXiv:1803.01271 (2018). [5] Oresti Banos, Rafael Garcia, Juan A Holgado-Terriza, Miguel Damas, Hector Pomares, Ignacio Rojas, Alejandro Saez, and Claudia Villalonga. 2014. mHealth- Droid: a novel framework for agile development of mobile health applications. In International workshop on ambient assisted living. Springer, 91–98. [6]Hymalai Bello, Daniel Geißler, Sungho Suh, Bo Zhou, and Paul Lukowicz. 2024. TSAK: Two-Stage Semantic-Aware Knowledge Distillation for Efficient Wear- able Modality and Model Optimization in Manufacturing Lines. arXiv preprint arXiv:2408.14146 (2024). [7] Alessio Benavoli, Giorgio Corani, Janez Demšar, and Marco Zaffalon. 2017. Time for a Change: a Tutorial for Comparing Multiple Classifiers Through Bayesian Analysis. Journal of Machine Learning Research 18, 77 (2017), 1–36. [8]Alessio Benavoli, Giorgio Corani, Francesca Mangili, Marco Zaffalon, and Fab- rizio Ruggeri. 2014. A Bayesian Wilcoxon Signed-Rank Test Based on the Dirich- let Process. In Proceedings of the 31st International Conference on Machine Learning (ICML). 1026–1034. [9] Sizhen Bian, Mengxi Liu, Vitor Fortes Rey, Daniel Geissler, and Paul Lukowicz. 2025. TinierHAR: Towards Ultra-Lightweight Deep Learning Models for Efficient Human Activity Recognition on Edge Devices. In Proceedings of the 2025 ACM International Symposium on Wearable Computers. 163–169. [10] Sizhen Bian, Vitor Fortes Rey, Siyu Yuan, and Paul Lukowicz. 2022. The con- tribution of human body capacitance/body-area electric field to individual and collaborative activity recognition. arXiv preprint arXiv:2210.14794 (2022). [11] Valentina Bianchi, Marco Bassoli, Gianfranco Lombardo, Paolo Fornacciari, Mon- ica Mordonini, and Ilaria De Munari. 2019. IoT wearable sensor and deep learning: An integrated approach for personalized human activity recognition in a smart home environment. IEEE Internet of Things Journal 6, 5 (2019), 8553–8562. [12] Marius Bock, Hilde Kuehne, Kristof Van Laerhoven, and Michael Moeller. 2023. Wear: An outdoor sports dataset for wearable and egocentric activity recognition. arXiv preprint arXiv:2304.05088 (2023). [13]Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio. 2014. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. arXiv preprint arXiv:1412.3555 (2014). [14]Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. 2017. Lan- guage Modeling with Gated Convolutional Networks. In Proceedings of the 34th International Conference on Machine Learning. PMLR, 933–941. [15] Deepika Gurung, Lala Shakti Swarup Ray, Mengxi Liu, Bo Zhou, and Paul Lukow- icz. 2026. SPECTRA: An Efficient Spectral-Informed Neural Network for Sensor- Based Activity Recognition. arXiv preprint arXiv:2603.26482 (2026). [16]Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition. 770–778. [17]Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017). [18]Jie Hu, Li Shen, and Gang Sun. 2018. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition. 7132–7141. [19]Ji Lin, Wei-Ming Chen, Yujun Lin, Chuang Gan, Song Han, et al.2020. Mcunet: Tiny deep learning on iot devices. Advances in neural information processing systems 33 (2020), 11711–11722. [20] Mohammad Malekzadeh, Richard G Clegg, Andrea Cavallaro, and Hamed Had- dadi. 2019. Mobile sensor data anonymization. In Proceedings of the international conference on internet of things design and implementation. 49–58. [21]Aimé Cedric Muhoza, Emmanuel Bergeret, Corinne Brdys, and Francis Gary. 2023. Power consumption reduction for IoT devices thanks to Edge-AI: Application to human activity recognition. Internet of Things 24 (2023), 100930. [22]Francisco Javier Ordóñez and Daniel Roggen. 2016. Deep convolutional and lstm recurrent neural networks for multimodal wearable activity recognition. Sensors 16, 1 (2016), 115. [23]Attila Reiss and Didier Stricker. 2012. Introducing a new benchmarked dataset for activity monitoring. In 2012 16th international symposium on wearable computers. IEEE, 108–109. [24] Jorge-L Reyes-Ortiz, Luca Oneto, Albert Samà, Xavier Parra, and Davide An- guita. 2016. Transition-aware human activity recognition using smartphones. Neurocomputing 171 (2016), 754–767. [25] Daniel Roggen, Alberto Calatroni, Mirco Rossi, Thomas Holleczek, Kilian Förster, Gerhard Tröster, Paul Lukowicz, David Bannach, Gerald Pirkl, Alois Ferscha, et al.2010. Collecting complex activity datasets in highly rich networked sensor environments. In 2010 Seventh international conference on networked sensing systems (INSS). IEEE, 233–240. [26]Mutegeki Ronald, Alwin Poulose, and Dong Seog Han. 2021. iSPLInception: an inception-ResNet deep learning architecture for human activity recognition. IEEE Access 9 (2021), 68985–69001. [27]Abhijit Guha Roy, Nassir Navab, and Christian Wachinger. 2018. Concurrent spatial and channel ‘squeeze & excitation’in fully convolutional networks. In International conference on medical image computing and computer-assisted inter- vention. Springer, 421–429. [28]Muhammad Shoaib, Stephan Bosch, Ozlem Durmaz Incel, Hans Scholten, and Paul JM Havinga. 2014. Fusion of smartphone motion sensors for physical activity recognition. Sensors 14, 6 (2014), 10146–10176. https://doi.org/10.3390/ s140610146 [29]Davinder Pal Singh, Lala Shakti Swarup Ray, Bo Zhou, Sungho Suh, and Paul Lukowicz. 2024. A Novel Local-Global Feature Fusion Framework for Body- Weight Exercise Recognition with Pressure Mapping Sensors. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 6375–6379. [30] Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber. 2015. Training Very Deep Networks. In Advances in Neural Information Processing Systems, Vol. 28. [31]Sungho Suh, Vitor Fortes Rey, Sizhen Bian, Yu-Chi Huang, Jože M Rožanec, Hooman Tavakoli Ghinani, Bo Zhou, and Paul Lukowicz. 2023. Worker Activity Recognition in Manufacturing Line Using Near-body Electric Field. IEEE Internet of Things Journal (2023). [32] Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition. 1–9. [33] Timo Sztyler, Heiner Stuckenschmidt, and Wolfgang Petrich. 2017. Position- aware activity recognition with wearable devices. Pervasive and mobile computing 38 (2017), 281–295. [34]Wenjin Tao, Ze-Hao Lai, Ming C Leu, and Zhaozheng Yin. 2018. Worker ac- tivity recognition in smart manufacturing using IMU and sEMG signals with convolutional neural networks. Procedia Manufacturing 26 (2018), 1159–1166. [35]Aaron Van Den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, Koray Kavukcuoglu, et al. 2016. Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499 12, 1 (2016). [36] Aäron van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu. 2016. Conditional Image Generation with Pix- elCNN Decoders. In Advances in Neural Information Processing Systems, Vol. 29. [37]Cheng Xu, Jie He, Xiaotong Zhang, Cui Yao, and Po-Hsuan Tseng. 2018. Geo- metrical kinematic modeling on human motion using method of multi-sensor fusion. Information Fusion 41 (2018), 243–254. [38]Piero Zappi, Clemens Lombriser, Thomas Stiefmeier, Elisabetta Farella, Daniel Roggen, Luca Benini, and Gerhard Tröster. 2008. Activity recognition from on- body sensors: accuracy-power trade-off by dynamic sensor selection. In European Conference on Wireless Sensor Networks. Springer, 17–33. [39]Mi Zhang and Alexander A Sawchuk. 2012. USC-HAD: A daily activity dataset for ubiquitous activity recognition using wearable sensors. In Proceedings of the 2012 ACM conference on ubiquitous computing. 1036–1043. [40]Xiaolong Zheng, Jiliang Wang, Longfei Shangguan, Zimu Zhou, and Yunhao Liu. 2017. Design and implementation of a CSI-based ubiquitous smoking detection system. IEEE/ACM transactions on networking 25, 6 (2017), 3781–3793. [41]Bo Zhou, Sungho Suh, Vitor Fortes Rey, Carlos Andres Velez Altamirano, and Paul Lukowicz. 2022. Quali-mat: Evaluating the quality of execution in body- weight exercises with a pressure sensitive sports mat. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 6, 2 (2022), 1–45. [42] Yexu Zhou, Tobias King, Haibin Zhao, Yiran Huang, Till Riedel, and Michael Beigl. 2024. Mlp-har: Boosting performance and efficiency of har models on edge devices with purely fully connected layers. In Proceedings of the 2024 ACM International Symposium on Wearable Computers. 133–139. [43] Yexu Zhou, Haibin Zhao, Yiran Huang, Till Riedel, Michael Hefenbrock, and Michael Beigl. 2022. Tinyhar: A lightweight deep learning model designed for human activity recognition. In Proceedings of the 2022 ACM International Symposium on Wearable Computers. 89–93.