Paper deep dive
Rotation-Invariant Multi-IMU Activity Recognition under Independent Per-Location Orientation Shifts
Seungyeol Baek, Yoonbyung Chai, Yonghyeon Lee, Sungjoon Choi, Sungho Suh
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/22/2026, 2:39:30 AM
Summary
The paper introduces TRI-HAR, a rotation-invariant framework for Human Activity Recognition (HAR) using multi-IMU wearable data. It addresses the challenge of independent per-location orientation shifts caused by reattaching sensors in self-administered settings. Unlike existing methods relying on rotation augmentation or calibration, TRI-HAR structurally embeds robustness by using an SO(3)-equivariant backbone (FER-VN-DGCNN) to process triaxial vector streams per IMU location, followed by invariant projection and fusion. This approach preserves macro-F1 performance under fixed independent SO(3) rotations across four benchmarks (PAMAP2, DSADS, Opportunity, RealDISP) without requiring rotational augmentation.
Entities (10)
Relation Signals (9)
TRI-HAR → evaluatedon → PAMAP2
confidence 99% · We evaluate TRI-HAR on four public multi-IMU HAR benchmarks: PAMAP2, DSADS, Opportunity, and RealDISP.
TRI-HAR → evaluatedon → DSADS
confidence 99% · We evaluate TRI-HAR on four public multi-IMU HAR benchmarks: PAMAP2, DSADS, Opportunity, and RealDISP.
TRI-HAR → evaluatedon → Opportunity
confidence 99% · We evaluate TRI-HAR on four public multi-IMU HAR benchmarks: PAMAP2, DSADS, Opportunity, and RealDISP.
TRI-HAR → evaluatedon → REALDISP
confidence 99% · We evaluate TRI-HAR on four public multi-IMU HAR benchmarks: PAMAP2, DSADS, Opportunity, and RealDISP.
TRI-HAR → providesrobustnessto → independent per-location orientation shifts
confidence 95% · TRI-HAR makes robustness to independent per-location IMU orientation offsets a structural model property.
TRI-HAR → solves → Human Activity Recognition
confidence 95% · We present Truly Rotation-Invariant HAR (TRI-HAR), a rotation-invariant framework... for multi-IMU HAR under independent per-location orientation shifts.
TRI-HAR → uses → FER-VN-DGCNN
confidence 92% · We instantiate B with FER-VN-DGCNN (25), which combines Vector Neuron operations (9) with Frequency-based Equivariant Feature Representation (FER).
FER-VN-DGCNN → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Human Activity Recognition (HAR) with self-administered wearables, such as at-home rehabilitation and exercise monitoring, often requires reattaching inertial measurement units (IMUs) across sessions. In multi-IMU settings, this can induce independent orientation offsets across body locations, a deployment shift that conventional scalar HAR models do not structurally handle. Existing remedies rely on rotation augmentation, whose robustness depends on sampled transformations, or calibration and orientationnormalization pipelines requiring additional reference-frame assumptions or explicit procedures. We present Truly Rotation-Invariant HAR (TRI-HAR), a rotation-invariant framework that makes robustness to independent per-location IMU orientation offsets a structural model property. TRI-HAR reshapes accelerometer and gyroscope streams into triaxial vectors, applies a shared SO(3)-equivariant backbone and invariant projection to each IMU location, and fuses the resulting invariant features for activity classification. Across four multi-IMU benchmarks, TRI-HAR preserves macro-F1 under fixed independent per-location SO(3) rotations and outperforms rotation-augmented baselines under this target shift without requiring rotational augmentation.
Tags
Links
- Source: https://arxiv.org/abs/2608.15621v1
- Canonical: https://arxiv.org/abs/2608.15621v1
Trouble viewing inline? Open PDF directly →
Full Text
45,960 characters extracted from source content.
Expand or collapse full text
Rotation-Invariant Multi-IMU Activity Recognition under Independent Per-Location Orientation ShiftsConference: Proceedings of the 2026 ACM International Symposium on Wearable Computers; October 11–15, 2026; Shanghai, ChinaProceedings of the 2026 ACM International Symposium on Wearable Computers (ISWC ’26), October 11–15, 2026, Shanghai, ChinaDOI: 10.1145/3830727.3834824ISBN: 979-8-4007-2872-3/2026/10CCS: Human-centered computing Ubiquitous computingCCS: Human-centered computing Ubiquitous and mobile computing theory, concepts and paradigms Seungyeol Baek Affiliation: Korea University , Seoul , Republic of Korea email: mbaek01@korea.ac.kr , Yoonbyung Chai Affiliation: Korea University , Seoul , Republic of Korea email: yoonbyung-chai@korea.ac.kr , Yonghyeon Lee Affiliation: Massachusetts Institute of Technology , Cambridge , Massachusetts , USA email: yhlee.gabe@gmail.com , Sungjoon Choi Affiliation: Dept. of Artificial Intelligence , Korea University , Seoul , Republic of Korea email: sungjoon-choi@korea.ac.kr and Sungho Suh Affiliation: Department of Artificial Intelligence , Korea University , Seoul , Republic of Korea email: sungho_suh@korea.ac.kr 2026; © c Abstract. Human Activity Recognition (HAR) with self-administered wearables, such as at-home rehabilitation and exercise monitoring, often requires reattaching inertial measurement units (IMUs) across sessions. In multi-IMU settings, this can induce independent orientation offsets across body locations, a deployment shift that conventional scalar HAR models do not structurally handle. Existing remedies rely on rotation augmentation, whose robustness depends on sampled transformations, or calibration and orientation-normalization pipelines requiring additional reference-frame assumptions or explicit procedures. We present Truly Rotation-Invariant HAR (TRI-HAR), a rotation-invariant framework that makes robustness to independent per-location IMU orientation offsets a structural model property. TRI-HAR reshapes accelerometer and gyroscope streams into triaxial vectors, applies a shared SO(3)-equivariant backbone and invariant projection to each IMU location, and fuses the resulting invariant features for activity classification. Across four multi-IMU benchmarks, TRI-HAR preserves macro-F1 under fixed independent per-location SO(3) rotations and outperforms rotation-augmented baselines under this target shift without requiring rotational augmentation. Keywords: Human Activity Recognition, Rotation Invariance, SO(3)-Equivariance †c-license: by 1. Introduction Figure 1. Body-worn IMUs may be reattached with independent orientation offsets across uses. TRI-HAR performs per-location equivariant encoding and invariant projection before fusion for stable activity predictions.Teaser diagram for TRI-HAR. The left side shows a person wearing IMUs at the chest, wrist, thigh, and ankle during an initial placement. The middle shows the same activity during inference after sensors have been reattached, with each IMU coordinate frame rotated differently and labeled with independent rotations R1 through R4. The right side shows a simplified TRI-HAR block with equivariant encoding, invariant projection, and location-wise fusion, producing the same activity label before and after reattachment. Human Activity Recognition (HAR) classifies human activities from wearable inertial sensor data and supports applications such as healthcare monitoring, sports and exercise analytics, and gesture-based interaction (8; 21; 7). A central challenge in wearable HAR is the distribution shift caused by variations in sensor orientation, or rotational misalignment. When an IMU is worn or reattached in varying orientations, the same body motion can produce different accelerometer and gyroscope measurements. These orientation shifts change the representation of sensor signals and can substantially degrade deep HAR models that lack built-in robustness to rotation-induced distribution shift (3; 15). This problem becomes especially relevant in self-administered multi-IMU settings, such as at-home rehabilitation or exercise monitoring, where sensors may be reattached across sessions, and each body-worn IMU can acquire its own orientation offset (21). Existing approaches to orientation variability typically rely on either data augmentation or calibration and orientation-normalization pipelines (13; 12). Rotation augmentation exposes a model to synthetic variants of the training data, but its robustness is tied to the sampled transformations and does not guarantee invariance to arbitrary unseen rotations (17; 5; 14; 35). Calibration-based pipelines can reduce orientation mismatch, but they add preprocessing overhead and introduce additional assumptions about available reference frames, static intervals, magnetic reliability, or explicit user or system procedures (33; 13; 12). These assumptions may be difficult to satisfy in uncontrolled self-administered test settings (3; 4). A more structural approach is to encode rotational symmetry directly into the model. SO(3)-equivariant architectures produce intermediate features that co-rotate predictably with the input, enabling rotation-invariant representations to be constructed by design (9; 25). However, wearable HAR requires more than shared-global rotation invariance. In self-administered multi-IMU settings, sensors at different locations can be reattached with their own orientation offsets. We present Truly Rotation-Invariant HAR (TRI-HAR), a rotation-invariant framework for multi-IMU HAR under independent per-location orientation shifts. As shown in Figure 1, TRI-HAR treats accelerometer and gyroscope measurements as triaxial vector streams, groups them by physical IMU location, and maps each group through a shared equivariant-to-invariant location encoder. This encoder lifts each group into a multi-frequency SO(3)-equivariant representation and projects it to an invariant location feature. The resulting invariant location features are then concatenated in a fixed body-location order for classification. This design preserves separate rotation actions until invariant projection, enabling invariance to independent per-location SO(3) offsets without rotational augmentation or an additional test-time orientation-calibration step. We evaluate TRI-HAR on four public multi-IMU HAR benchmarks: PAMAP2, DSADS, Opportunity, and RealDISP. Against supervised HAR baselines with and without matched per-location rotation augmentation, TRI-HAR preserves macro-F1 under fixed independent SO(3) test rotations and outperforms rotation-augmented baselines under this target shift. We also measure live host-side latency on 3-IMU and 5-IMU streams to characterize the runtime cost of the location-wise encoder. Our contributions are summarized as follows: • We propose TRI-HAR, a multi-IMU HAR framework that learns equivariant-to-invariant representations per physical IMU location, enabling structural robustness to independent orientation offsets without rotation augmentation or additional test-time orientation calibration. • We formulate an independent per-location SO(3) shift protocol, where each IMU location can receive a distinct fixed orientation offset, matching self-administered multi-IMU reattachment settings. • Across four public benchmarks, we show that TRI-HAR preserves macro-F1 under fixed independent per-location rotations and outperforms rotation-augmented baselines under this target orientation-shift condition. 2. Related Work Prior HAR work mainly handles orientation variability through data augmentation, ranging from simple signal-space rotations to synthetic IMU generation methods (33; 13). More recent cross-dataset approaches extend this idea with large-scale representation learning; for example, CrossHAR (16) combines hierarchical self-supervised pretraining with physics-informed augmentation, while oneHAR (30) uses LLM-assisted virtual IMU generation to improve generalization across sensor positions and orientations. However, these methods remain sampling-dependent. Most HAR augmentation schemes apply only restricted families of rotations, often around gravity or body axes, so robustness depends on the synthesized transformations seen during training rather than being guaranteed for arbitrary SO(3) rotations (14; 5; 35). Calibration, canonicalization, and orientation-normalization pipelines provide an alternative by mapping measurements to a consistent reference frame (27; 34; 12). While often effective within distribution, such preprocessing introduces additional engineering assumptions, such as static intervals or reliable magnetic references, that may not hold in practical wearable deployments (27; 34; 4). A complementary line of work encodes SO(3)-equivariance directly in the network architecture. Earlier examples include Tensor Field Networks and 3D Steerable CNNs (28; 31). More directly relevant to our setting, Vector Neurons (VN) lift scalar features to 3D vectors and redefine standard neural-network operations so that they remain SO(3)-equivariant, enabling equivariant variants of backbones such as PointNet and DGCNN (9; 23; 29). FER extends VN with a higher-dimensional multi-frequency equivariant embedding that captures richer angular structure while preserving strict rotation equivariance (25). TRI-HAR builds on this FER-VN line and adapts it from 3D geometric data to windowed IMU sequences for rotation-invariant HAR. Figure 2. Architecture of TRI-HAR. Windowed input signals are vectorized into triaxial streams and grouped by physical IMU location. A shared location encoder EθE_θ is applied to each group X(l)X^(l) to produce an invariant feature hlh_l; the features h1,…,hLh_1,…,h_L are then concatenated in a fixed location order and passed to an MLP classifier. Dimensions inside EθE_θ are shown for one location group, and the encoder weights are shared across locations.Diagram of the TRI-HAR architecture. A sensor input window is first vectorized from flattened accelerometer and gyroscope channels into triaxial streams, then grouped by IMU location. Each location group is passed through the same shared encoder, labeled $E_θ$, which contains VN-EdgeConv layers, a VN invariant layer, and pooling, producing one invariant feature $h_l$ per location. The invariant features from all locations are concatenated and passed through an MLP to produce activity-class predictions. 3. Method Given a window of accelerometer and gyroscope signals, TRI-HAR reshapes scalar channels into triaxial vector streams, groups them by physical IMU location, and maps each group through a shared SO(3)SO(3)-equivariant backbone followed by invariant projection. The backbone lifts the streams into a multi-frequency angular Vector Neuron representation, and the invariant projection converts the resulting equivariant features to a rotation-invariant location feature (9; 25). These invariant location features are then concatenated in a fixed order for activity classification. Figure 2 summarizes the TRI-HAR architecture. By applying invariant projection before cross-location fusion, TRI-HAR matches self-administered wearable settings in which each body-worn IMU may be reattached with its own orientation offset. 3.1. Location-Grouped IMU Vector Representation A windowed IMU sample is initially represented as flattened scalar channels Xraw∈ℝT×3SX_raw ^T× 3S, where T is the window length and S is the number of triaxial streams. TRI-HAR reshapes these flat channels into triaxial vector streams, X∈ℝT×S×3X ^T× S× 3. In this representation, rotations are applied to the final 3D vector dimension of each stream. The triaxial streams are then grouped by physical IMU location: (1) X=(X(1),X(2),…,X(L)),X(l)∈ℝT×Sloc×3,l=1,…,L,X= (X^(1),X^(2),…,X^(L) ), X^(l) ^T× S_loc× 3, l=1,…,L, where L is the number of physical IMU locations and SlocS_loc is the number of triaxial streams assigned to each location. These location groups define the units over which independent orientation offsets are modeled. Accelerometer and gyroscope streams from the same body-worn IMU are grouped together and share the same rotation action. The ordered tuple in Eq. 1 defines the fixed location order used for fusion. Thus, S=LSlocS=LS_loc in the grouped representation. TRI-HAR preserves body-location identity and is not designed to be invariant to arbitrary permutations of sensor locations. We use Rl⋅X(l)R_l· X^(l) to denote applying Rl∈SO(3)R_l (3) to the 3D vector dimension of every triaxial stream at every time step within the window for location l. This notation models independent mounting offsets across physical IMU locations. 3.2. Shared SO(3)SO(3)-Equivariant Backbone TRI-HAR uses a single shared SO(3)SO(3)-equivariant backbone, denoted B, for all physical IMU locations. We instantiate B with FER-VN-DGCNN (25), which combines Vector Neuron operations (9) with Frequency-based Equivariant Feature Representation (FER). Vector Neurons replace scalar hidden features with vector-valued channels that transform predictably under 3D rotations. In the multi-frequency FER extension, a time-indexed latent feature is represented as Vt∈ℝC×NV_t ^C× N, where C is the number of vector-valued channels and N is the lifted representation dimension. Given the location-grouped input from Eq. 1, the backbone maps each location group to a sequence of equivariant latent features: (2) V1:T(l)=B(X(l)),Vt(l)∈ℝC×N,l=1,…,L.V_1:T^(l)=B (X^(l) ), V_t^(l) ^C× N, l=1,…,L. Below, we describe the computation for one location group and omit the superscript (l)(l) when no ambiguity arises. Let D:SO(3)→SO(N)D:SO(3) (N) denote the lifted rotation action used by FER. Under the right-action convention used here, a rotation R∈SO(3)R (3) acts on the latent feature as VtD(R)V_tD(R). A learnable mapping f is equivariant if (3) f(VtD(R))=f(Vt)D(R).f (V_tD(R) )=f (V_t )D(R). Thus, rotating the input causes intermediate features to co-rotate in the lifted feature space, rather than changing the underlying activity information. FER performs the initial lift from 3D IMU vectors to this multi-frequency equivariant space. For an input vector u∈ℝ3u ^3, FER defines (4) ψ(u):=φ(‖u‖)D(Rz^(u^))e^,ψ(u):= (\|u\|)D\! (R z( u) ) e, where u^=u/‖u‖ u=u/\|u\|, Rz^(u^)∈SO(3)R z( u) (3) is the rotation that aligns a reference axis z z to u u, e e is a basis vector in the target embedding space, and φ(⋅) (·) is a learnable radial function applied to the input magnitude. The direction of u determines the angular component of the lifted feature, while φ(‖u‖) (\|u\|) modulates its magnitude-dependent frequency coefficients. This gives TRI-HAR a multi-frequency angular representation while preserving SO(3)-equivariance. Within B, each location group is processed as a dynamic graph over time-indexed IMU states. The nodes are the T time steps within the window, and the node feature at time t is the colocated multi-stream sensor state Xt(l)∈ℝSloc×3X_t^(l) ^S_loc× 3. After the FER lift, each VN-EdgeConv block recomputes a k-nearest-neighbor graph among these time-step nodes in the current equivariant latent feature space. Thus, the backbone relates feature-similar motion states from the same physical IMU location, rather than constructing nodes over individual scalar channels or sensor-time pairs. Here, VN-EdgeConv denotes the Vector-Neuron implementation of the EdgeConv update used for graph feature aggregation. For a central node feature ViV_i and neighbor feature VjV_j, the edge message is computed as mij=VN-MLP([Vi,Vj−Vi])m_ij=VN -MLP([V_i,V_j-V_i]), where the concatenation is along the vector-channel dimension. This construction combines the current node feature with its relative difference to a neighbor while preserving equivariance through VN linear layers, pooling, and nonlinearities. We use the norm-based VN nonlinearity from FER-VN-DGCNN, which rescales vector features using functions of their rotation-invariant norms, rather than the original VN-ReLU (9; 25). Stacking these VN-EdgeConv blocks yields the equivariant feature sequence V1:T(l)V_1:T^(l) used by the invariant projection described next. 3.3. Rotation-Invariant Location Features The equivariant backbone produces features that co-rotate with the input. TRI-HAR converts these features into invariant scalar representations using the VN invariant layer. Let Vt∈ℝC×NV_t ^C× N be an equivariant feature and let Ft∈ℝC′×NF_t ^C × N be a learned equivariant frame. If both transform under the same lifted rotation D(R)D(R), then (5) (VtD(R))(FtD(R))⊤=VtD(R)D(R)⊤Ft⊤=VtFt⊤. (V_tD(R) ) (F_tD(R) ) =V_tD(R)D(R) F_t =V_tF_t . This allows invariant scalars to be constructed from inner products between the latent feature and a co-rotating frame. We first compute a window-level context V¯ V. Then, for each time step t, we estimate an equivariant frame FtF_t and obtain the invariant time-step feature ztz_t: (6) V¯ V :=1T∑τ=1TVτ, := 1T _τ=1^TV_τ, Ft F_t :=VN-MLP([Vt,V¯]), :=VN -MLP([V_t, V]), zt z_t :=VN-In(Vt)=VtFt⊤. :=VN -In(V_t)=V_tF_t . Thus, ztz_t is invariant to the rotation of the input that produced VtV_t. For a physical IMU location l, the invariant time-step features zt(l)t=1T\z_t^(l)\_t=1^T are aggregated over time with global max and average pooling, denoted by PoolPool, to obtain a single location-level invariant representation hlh_l. We use EθE_θ to denote the full shared location encoder: (7) hl:=Eθ(X(l)),Eθ:=Pool∘VN-In∘B.h_l:=E_θ(X^(l)), E_θ:=Pool -In B. Here, θ collects the learnable parameters of the shared location encoder, including those in the equivariant backbone B and the VN invariant projection. The same EθE_θ is used for every location group, but it is evaluated separately on each X(l)X^(l), so TRI-HAR shares weights across locations without learning location-specific equivariant encoders. 3.4. Per-Location Invariant Fusion TRI-HAR fuses information across IMU locations only after each location group has been mapped to an invariant feature. The resulting location features h1,…,hLh_1,…,h_L are concatenated in a fixed dataset-specific order and passed to the final classifier: (8) ^=MLPω([h1,h2,…,hL]). y=MLP_ω([h_1,h_2,…,h_L]). Although location features are computed separately, the final classifier receives their fixed-order concatenation, so activity prediction can still depend on multi-location patterns in the invariant feature space. This fusion order gives TRI-HAR invariance to independent per-location orientation offsets. In Eq. 5, the invariant projection cancels the lifted rotation because both the latent feature VtV_t and the learned frame FtF_t transform by the same D(R)D(R). Combining vector features from multiple IMU locations before invariant projection would therefore impose a shared-global rotation assumption on the joint representation. TRI-HAR instead applies invariant projection within each location group, canceling the local action D(Rl)D(R_l) before fusing locations. Suppose that location l receives an independent rotation Rl∈SO(3)R_l (3). Equivariance of B gives B(Rl⋅X(l))=B(X(l))D(Rl)B(R_l· X^(l))=B(X^(l))D(R_l). The invariant projection removes this location-specific rotation action, and pooling aggregates the invariant time-step features. Thus, the full location encoder satisfies (9) Eθ(Rl⋅X(l))=Eθ(X(l)),l=1,…,L.E_θ(R_l· X^(l))=E_θ(X^(l)), l=1,…,L. Consequently, replacing each location input X(l)X^(l) in Eq. 8 with its independently rotated version Rl⋅X(l)R_l· X^(l) leaves every hlh_l unchanged, and therefore leaves y unchanged. This establishes invariance to independent per-location rotations, with shared global rotation as a special case. It is also the central distinction from a joint-fusion equivariant control, which does not preserve separate location groups through invariant projection and therefore matches a shared-global rotation assumption rather than the independent reattachment setting targeted here. 4. Experiments 4.1. Datasets and Preprocessing We evaluate TRI-HAR on PAMAP2 (24), DSADS (1), Opportunity (6), and RealDISP (2). Table 1 summarizes the sampling rate, window length, number of subjects, number of classes, and number of IMU locations used in our experiments. We use the listed sampling rates and window lengths following the selected baselines (37; 20). For consistency, we retain accelerometer and gyroscope channels only. For PAMAP2, we remove subject 109 due to missing data and use the ±16g± 16g accelerometer stream. For Opportunity, we exclude the Null label and retain seven on-body IMU locations; for the shoe IMUs, we use the sensor/body-frame accelerometer and gyroscope streams and omit navigation-frame streams. For RealDISP, we include the ideal, self-displacement, and mutual-displacement recordings. Signals are segmented with 50%-overlap sliding windows using the listed window lengths and standardized channel-wise using training-split statistics. Table 1. Characteristics of the datasets used in the evaluation Dataset Freq. WL #Subj. #Cls. #Locs. PAMAP2 (24) 33 5.12 8 12 3 DSADS (1) 25 5.00 8 19 5 Opportunity (6) 30 1.00 4 17 7 RealDISP (2) 50 2.40 17 33 9 Note: Freq. denotes sampling rate (Hz); WL denotes window length (s); Subj. and Cls. indicate the numbers of subjects and activity classes, respectively; and Locs. denotes the number of body-worn sensor locations. Table 2. Benchmark comparison under fixed independent per-location test rotations. Values are macro-F1 (%, mean ± standard deviation) over cross-subject folds. Columns report performance on the original test data I and under SO(3)loc-fixSO(3)_loc-fix; entries under I are omitted for loc-sample rows for compactness. Bold denotes the best value within each dataset/test column. Model Train Aug. PAMAP2 (24) DSADS (1) Opportunity (6) RealDISP (2) I SO()loc-fix SO(3)_loc-fix I SO()loc-fix SO(3)_loc-fix I SO()loc-fix SO(3)_loc-fix I SO()loc-fix SO(3)_loc-fix MC-CNN (32) none 80.45± 8.18 31.52± 6.21 82.80± 8.09 30.63± 5.80 41.18± 6.62 11.82± 3.90 87.48± 10.39 48.95± 9.81 loc-sample — 79.14± 6.75 — 81.75± 6.43 — 33.56± 8.10 — 87.19± 8.35 DeepConvLSTM (22) none 71.13± 10.39 21.84± 10.05 76.05± 11.27 25.85± 6.74 39.44± 7.99 6.16± 1.85 86.99± 10.21 53.37± 16.25 loc-sample — 76.65± 8.80 — 77.97± 7.94 — 30.01± 9.06 — 87.65± 7.65 MLP-HAR (36) none 80.08± 6.73 28.52± 9.27 83.23± 7.80 22.55± 6.46 42.01± 3.93 5.45± 1.32 88.49± 6.82 39.70± 4.80 loc-sample — 79.67± 8.74 — 86.11± 6.83 — 32.67± 4.61 — 88.38± 5.64 TinyHAR (37) none 82.10± 6.52 31.64± 4.77 82.82± 7.25 26.60± 10.33 38.13± 4.54 7.37± 1.12 88.13± 8.78 52.43± 8.74 loc-sample — 79.99± 7.46 — 82.50± 8.01 — 29.98± 5.44 — 88.36± 8.71 SA-HAR (20) none 74.74± 10.27 12.65± 5.55 78.11± 7.87 16.14± 3.32 28.78± 9.56 4.84± 1.15 81.78± 16.14 50.54± 14.29 loc-sample — 70.59± 7.20 — 72.95± 8.01 — 21.23± 6.13 — 82.51± 15.26 TRI-HAR none 83.64± 5.56 83.64± 5.56 89.07± 4.49 89.07± 4.49 38.06± 3.77 38.06± 3.77 91.94± 4.99 91.94± 4.99 4.2. Baselines We compare TRI-HAR with five supervised HAR baselines: MC-CNN (32), DeepConvLSTM (22), MLP-HAR (36), TinyHAR (37), and SA-HAR (20). Unless noted otherwise, we use the original architectures from the cited papers, adapting only the input and output dimensions to each dataset. All baselines use the same preprocessing, cross-validation splits, and test-rotation protocol as TRI-HAR. Our comparison focuses on supervised HAR backbones under a matched independent-rotation augmentation protocol, isolating the architectural effect of replacing sampled augmentation with built-in per-location invariance. For MLP-HAR, τ was set to evenly divide each dataset-specific window, yielding three temporal patches per window except for DSADS, which uses five one-second patches. 4.3. Implementation Details 4.3.1. Batched Location-Encoder Implementation For efficient computation, our implementation11 1 https://github.com/mbaek01/TRI-HAR folds the location axis into the batch axis, so location groups are evaluated by the shared encoder EθE_θ in one batched forward pass. The location features are then reshaped to restore the batch axis and fixed location order before fusion. 4.3.2. Evaluation Protocol and Rotation Setup We adopted a unified cross-subject evaluation protocol across all benchmarks, with folds defined by subject or subject group. For PAMAP2 (8 folds), DSADS (8 folds), and Opportunity (4 folds), we used leave-one-subject-out cross-validation. For RealDISP, we used a 5-fold grouped split designed to evaluate robustness to sensor displacement. Four folds contain subject groups from the Ideal and Self-displacement scenarios, while the fifth fold contains all Mutual-displacement samples. In this scenario, the instructor repositions and rotates the sensors to purposely introduce substantial variation in rotational orientation and asymmetric placement errors. This design ensures that the displacement and orientation variations in the Mutual-displacement setting are not observed during training, providing a stringent test of robustness under realistic placement variability. To evaluate robustness to orientation variability in multi-IMU settings, we consider two test conditions: the original test data, I, and fixed independent per-location SO(3) offsets, SO(3)loc-fixSO(3)_loc-fix. In SO(3)loc-fixSO(3)_loc-fix, one SO(3) rotation is sampled independently for each physical IMU location and held fixed across the held-out fold. This protocol is motivated by the target deployment setting, in which different body-worn IMUs may be reattached with different mounting orientations. In contrast to prior equivariant-learning work, which commonly evaluates shared-global SO(3) augmentation (11; 9), our setting focuses on independent per-location offsets. For the augmented baselines, we use a matched independent per-location SO(3) augmentation, denoted by loc-sample: for each training window, one random SO(3) rotation is sampled independently for each location group and applied jointly to all channels from that IMU location group (accelerometer and gyroscope). These per-location rotations are resampled on the fly for every training window. We use this augmentation rather than shared-global SO(3) augmentation because it better matches the target deployment condition and empirically yields stronger baseline robustness under SO(3)loc-fixSO(3)_loc-fix. TRI-HAR is trained without rotational augmentation and instead achieves invariance to the modeled per-location SO(3) offsets by construction. 4.3.3. Hyperparameters All models were trained with cross-entropy loss and Adam (18) (β1=0.9,β2=0.999 _1=0.9, _2=0.999, learning rate 1×10−31× 10^-3) for up to 150 epochs on an NVIDIA RTX 4090 (24GB) GPU, using batch size 128 and early stopping on validation macro-F1. For TRI-HAR, we set Cin=14C_in=14, Cout=224C_out=224. We use k=20k=20 for Opportunity and k=5k=5 for the remaining datasets, selected by validation macro-F1. 4.3.4. Latency Evaluation Latency was measured on a Ryzen 7 9800X3D host with an RTX 5090 using native-rate ESP32-S3 streams matched to the PAMAP2-derived 3-IMU setting at 33 Hz and the DSADS-derived 5-IMU setting at 25 Hz. The ESP32-S3 streamed IMU samples only; windowing, preprocessing, and inference ran on the host. We report p99 live model and window-to-label latency over 1000 windows, where window-to-label latency is measured from final-sample arrival to label production. A stream is considered feasible for steady-state real-time operation if p99 window-to-label latency is below the window-update period H/fsH/f_s. These measurements characterize live host-side streaming inference in the tested setup. 5. Results 5.1. Benchmark Comparison Table 2 compares TRI-HAR with supervised HAR baselines on the original test data I and under fixed independent per-location offsets SO(3)loc-fixSO(3)_loc-fix. Without rotation augmentation, all non-equivariant baselines degrade substantially under SO(3)loc-fixSO(3)_loc-fix, whereas TRI-HAR preserves its macro-F1 under SO(3)loc-fixSO(3)_loc-fix across all four datasets, consistent with its invariance by construction and relative invariance error below 10−1010^-10. Loc-sample augmentation improves baseline robustness, but TRI-HAR still achieves the highest macro-F1 under SO(3)loc-fixSO(3)_loc-fix on all four datasets without rotation augmentation. Under I, TRI-HAR is best on PAMAP2, DSADS, and RealDISP, while trailing the strongest scalar baseline on Opportunity. 5.2. RealDISP Mutual-Displacement Fold Beyond the synthetic SO(3)loc-fixSO(3)_loc-fix evaluation, we examine the held-out RealDISP Mutual-displacement fold, where sensors are physically repositioned and reoriented. This fold is consistently among the most challenging of the five RealDISP folds. As shown in Table 3, TRI-HAR achieves the highest macro-F1, exceeding the strongest non-augmented and rotation-augmented baselines, MLP-HAR and TinyHAR, by 7.74 and 6.80 points, respectively. Relative to the corresponding five-fold RealDISP means under I, the non-augmented baselines are 10.94–31.89 macro-F1 points lower on the Mutual-displacement fold, compared with 6.7 points for TRI-HAR. Because Mutual-displacement includes both orientation and positional changes, it is not an isolated test of rotation robustness. Nevertheless, TRI-HAR’s advantage suggests that making predictions invariant to the modeled per-location rotations by construction remains beneficial under this combined physical placement shift. Moreover, rotation augmentation is not consistently beneficial: it improves four of the five baselines by 1.4–6.9 points but decreases MLP-HAR by 2.9 points. TRI-HAR instead attains the highest macro-F1 with structural invariance and without the augmentation-induced degradation observed for MLP-HAR. Table 3. Macro-F1 (%) on RealDISP Mutual-displacement fold. Δaug _aug is augmentation-induced change in baseline score. Bold marks the highest score. Model none loc-sample _aug MC-CNN 67.95 74.84 +6.89+6.89 DeepConvLSTM 67.68 72.90 +5.22+5.22 MLP-HAR 77.55 74.67 −2.88-2.88 TinyHAR 72.55 78.49 +5.94+5.94 SA-HAR 49.89 51.28 +1.39+1.39 TRI-HAR 85.29 — 5.3. Location-wise Fusion and Host-side Latency To isolate the effect of location-wise invariant fusion, we compare TRI-HAR with TRI-HAR-joint, an architectural control that retains the same equivariant components but replaces location-wise invariant fusion with one joint all-location invariant projection. For PAMAP2/DSADS/Opportunity/RealDISP, TRI-HAR-joint loses 37.71/50.95/16.80/50.12 macro-F1 points under SO(3)loc-fixSO(3)_loc-fix relative to I, whereas TRI-HAR loses none. A joint invariant projection can therefore cancel a shared global rotation but not independent per-location rotations. TRI-HAR avoids this failure by canceling each location-specific rotation before fusion. Under I, TRI-HAR also exceeds TRI-HAR-joint by 3.76/2.25/3.61/5.20 macro-F1 points; shared-global SO(3) tests are equivalent to I for both equivariant models and are omitted. This location-wise formulation creates TRI-HAR’s main runtime tradeoff. Although the encoder is shared, the equivariant backbone and invariant projection are executed once per physical IMU location. We therefore report live latency on 3-IMU and 5-IMU streams to quantify the host-side runtime cost introduced by the location-wise encoder. Table 4 reports live p99 model and window-to-label latency. Our feasibility check uses window-to-label latency. TRI-HAR is slower than scalar baselines, reflecting both the heavier equivariant-to-invariant encoder and its repeated per-location application. Nevertheless, its CPU window-to-label latency remains far below the window-update period. With the implemented hops, H/fs=84/33≈2.55H/f_s=84/33≈ 2.55 s for 3-IMU and H/fs=62/25=2.48H/f_s=62/25=2.48 s for 5-IMU, whereas TRI-HAR’s corresponding latencies are 67.61 ms and 85.78 ms. The update periods are therefore 37.6 and 28.9 times larger than the corresponding p99 latencies, supporting host-side real-time streaming inference despite the per-location computation. Table 4. Host-side live p99 latency for the 3-IMU and 5-IMU settings. CPU and GPU entries report model/window-to-label in ms. FP32 denotes parameter memory. Setting Model FP32 (MiB) CPU (ms) GPU (ms) 3-IMU MLP-HAR 0.67 0.86/1.08 0.69/1.09 MC-CNN 23.14 1.79/2.02 0.53/0.94 TRI-HAR 13.40 67.38/67.61 2.58/2.97 5-IMU MLP-HAR 1.84 0.97/1.21 1.14/1.48 MC-CNN 28.46 2.61/2.84 0.88/1.25 TRI-HAR 21.41 85.56/85.78 3.01/3.36 5.4. Discussion and Limitations TRI-HAR targets calibration-light, self-administered multi-IMU HAR, such as at-home rehabilitation or exercise monitoring, where sensors may be reattached across sessions and orientation should be treated as a nuisance factor. This is reflected in the multi-sensor benchmarks, where TRI-HAR remains stable under unseen rotations while the non-equivariant baselines degrade sharply. The location-wise formulation is particularly motivated by settings in which different body-worn IMUs have independent mounting orientations. Comparisons with explicit calibration, reference-frame normalization, and handcrafted orientation-invariant pipelines are complementary and left for future work. The latency results show that the current host-side implementation can keep up with live 3-IMU and 5-IMU streams, supporting off-sensor inference from streaming wearable IMUs. TRI-HAR nevertheless remains a robustness-first architecture rather than an edge-optimized wearable model. Its computational cost comes from two sources: the shared encoder is evaluated once per IMU location, and the encoder itself is costly, with the invariant projection as the dominant stage. Because the first source is the architectural cost of the targeted per-location invariance, future efficiency work should focus on reducing the encoder-internal cost while preserving the location-wise fusion. Promising directions include reducing the feature width before invariant projection, lower-rate frame estimation, graph sparsification, and distillation or simplification for equivariant GNNs (10; 19; 26). 6. Conclusion We presented TRI-HAR, a rotation-invariant wearable HAR framework that applies a shared SO(3)-equivariant backbone and invariant projection separately to each physical IMU location before fixed-order fusion. Across four multi-IMU benchmarks, TRI-HAR preserved macro-F1 under fixed independent per-location SO(3) rotations and outperformed rotation-augmented baselines under this target shift. These results support architecture-level rotation invariance for calibration-light, self-administered wearable HAR, with efficient on-device deployment left for future work. Acknowledgments This work was supported by the Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) (No. RS-2019-I190079, AI Graduate School Program (Korea University), 20%), the IITP-ITRC (Information Technology Research Center) grant (IITP-2026-RS-2024-00436857, 30%), and the IITP grant (No. RS-2026-25519380, 50%). References Altun et al. (2010) K. Altun, B. Barshan, and O. Tunçel Comparative study on classifying human activities with miniature inertial and magnetic sensors. Pattern Recognition 43 (10), p. 3605–3620. Cited by: §4.1, Table 1, Table 2. Baños et al. (2012) O. Baños, M. Damas, H. Pomares, I. Rojas, M. A. Tóth, and O. Amft A benchmark dataset to evaluate sensor displacement in activity recognition. In Proceedings of the 2012 ACM Conference on Ubiquitous Computing, p. 1026–1035. Cited by: §4.1, Table 1, Table 2. Barcelo-Ordinas et al. (2019) J. M. Barcelo-Ordinas, M. Doudou, J. Garcia-Vidal, and N. Badache Self-calibration methods for uncontrolled environments in sensor networks: a reference survey. Ad Hoc Networks 88, p. 142–159. Cited by: §1, §1. Batista et al. (2010) P. Batista, C. Silvestre, P. Oliveira, and B. Cardeira Accelerometer calibration and dynamic bias and gravity estimation: analysis, design, and experimental evaluation. IEEE transactions on control systems technology 19 (5), p. 1128–1137. Cited by: §1, §2. Caramaschi et al. (2023) S. Caramaschi, G. B. Papini, and E. G. Caiani Device orientation independent human activity recognition model for patient monitoring based on triaxial acceleration. Applied Sciences 13 (7), p. 4175. Cited by: §1, §2. Chavarriaga et al. (2013) R. Chavarriaga, H. Sagha, A. Calatroni, S. T. Digumarti, G. Tröster, J. d. R. Millán, and D. Roggen The opportunity challenge: a benchmark database for on-body sensor-based activity recognition. Pattern Recognition Letters 34 (15), p. 2033–2042. Cited by: §4.1, Table 1, Table 2. Chen et al. (2021) K. Chen, D. Zhang, L. Yao, B. Guo, Z. Yu, and Y. Liu Deep learning for sensor-based human activity recognition: overview, challenges, and opportunities. ACM Computing Surveys (CSUR) 54 (4), p. 1–40. Cited by: §1. De et al. (2015) D. De, P. Bharti, S. K. Das, and S. Chellappan Multimodal wearable sensing for fine-grained activity recognition in healthcare. IEEE Internet Computing 19 (5), p. 26–35. Cited by: §1. Deng et al. (2021) C. Deng, O. Litany, Y. Duan, A. Poulenard, A. Tagliasacchi, and L. J. Guibas Vector neurons: a general framework for so (3)-equivariant networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 12200–12209. Cited by: §1, §2, §3.2, §3.2, §3, §4.3.2. Ekström Kelvinius et al. (2023) F. Ekström Kelvinius, D. Georgiev, A. Toshev, and J. Gasteiger Accelerating molecular graph neural networks via knowledge distillation. Advances in Neural Information Processing Systems 36, p. 25761–25792. Cited by: §5.4. Esteves et al. (2018) C. Esteves, C. Allen-Blanchette, A. Makadia, and K. Daniilidis Learning so (3) equivariant representations with spherical cnns. In Proceedings of the european conference on computer vision (ECCV), p. 52–68. Cited by: §4.3.2. Gil-Martín et al. (2023) M. Gil-Martín, J. López-Iniesta, F. Fernández-Martínez, and R. San-Segundo Reducing the impact of sensor orientation variability in human activity recognition using a consistent reference system. Sensors 23 (13), p. 5845. Cited by: §1, §2. Halmich et al. (2025) C. Halmich, L. Höschler, C. Schranz, and C. Borgelt Data augmentation of time-series data in human movement biomechanics: a scoping review. PloS one 20 (7), p. e0327038. Cited by: §1, §2. Han et al. (2021) D. Han, C. Lee, and H. Kang Gravity control-based data augmentation technique for improving vr user activity recognition. Symmetry 13 (5), p. 845. Cited by: §1, §2. Haresamudram et al. (2025) H. Haresamudram, C. I. Tang, S. Suh, P. Lukowicz, and T. Ploetz Past, present, and future of sensor-based human activity recognition using wearables: a surveying tutorial on a still challenging task. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 9 (2), p. 1–44. Cited by: §1. Hong et al. (2024) Z. Hong, Z. Li, S. Zhong, W. Lyu, H. Wang, Y. Ding, T. He, and D. Zhang Crosshar: generalizing cross-dataset human activity recognition via hierarchical self-supervised pretraining. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8 (2), p. 1–26. Cited by: §2. Huang et al. (2023) X. Huang, Y. Xue, S. Ren, and F. Wang Sensor-based wearable systems for monitoring human motion and posture: a review. Sensors 23 (22), p. 9047. Cited by: §1. Kingma and Ba (2014) D. P. Kingma and J. Ba Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980. Cited by: §4.3.3. Li et al. (2021) Y. Li, H. Chen, Z. Cui, R. Timofte, M. Pollefeys, G. S. Chirikjian, and L. Van Gool Towards efficient graph convolutional networks for point cloud handling. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 3752–3762. Cited by: §5.4. Mahmud et al. (2020) S. Mahmud, M. Tanjid Hasan Tonmoy, K. Kumar Bhaumik, A. Mahbubur Rahman, M. Ashraful Amin, M. Shoyaib, M. Asif Hossain Khan, and A. Ahsan Ali Human activity recognition from wearable sensor data using self-attention. In ECAI 2020, p. 1332–1339. Cited by: §4.1, §4.2, Table 2. Müller et al. (2024) P. N. Müller, A. J. Müller, P. Achenbach, and S. Göbel Imu-based fitness activity recognition using cnns for time series classification. Sensors 24 (3), p. 742. Cited by: §1. Ordóñez and Roggen (2016) F. J. Ordóñez and D. Roggen Deep convolutional and lstm recurrent neural networks for multimodal wearable activity recognition. Sensors 16 (1), p. 115. Cited by: §4.2, Table 2. Qi et al. (2017) C. R. Qi, H. Su, K. Mo, and L. J. Guibas Pointnet: deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 652–660. Cited by: §2. Reiss and Stricker (2012) A. Reiss and D. Stricker Introducing a new benchmarked dataset for activity monitoring. In 2012 16th international symposium on wearable computers, p. 108–109. Cited by: §4.1, Table 1, Table 2. Son et al. (2024) D. Son, J. Kim, S. Son, and B. Kim An intuitive multi-frequency feature representation for so (3)-equivariant networks. arXiv preprint arXiv:2405.04537. Cited by: §1, §2, §3.2, §3.2, §3. Tailor et al. (2021) S. A. Tailor, R. De Jong, T. Azevedo, M. Mattina, and P. Maji Towards efficient point cloud graph neural networks through architectural simplification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 2095–2104. Cited by: §5.4. Tedaldi et al. (2014) D. Tedaldi, A. Pretto, and E. Menegatti A robust and easy to implement method for imu calibration without external equipments. In 2014 IEEE international conference on robotics and automation (ICRA), p. 3042–3049. Cited by: §2. Thomas et al. (2018) N. Thomas, T. Smidt, S. Kearnes, L. Yang, L. Li, K. Kohlhoff, and P. Riley Tensor field networks: rotation-and translation-equivariant neural networks for 3d point clouds. arXiv preprint arXiv:1802.08219. Cited by: §2. Wang et al. (2019) Y. Wang, Y. Sun, Z. Liu, S. E. Sarma, M. M. Bronstein, and J. M. Solomon Dynamic graph cnn for learning on point clouds. ACM Transactions on Graphics (tog) 38 (5), p. 1–12. Cited by: §2. Wei et al. (2025) Q. Wei, J. Huang, Y. Gao, and W. Dong One model to fit them all: universal imu-based human activity recognition with llm-assisted cross-dataset representation. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 9 (3), p. 1–22. Cited by: §2. Weiler et al. (2018) M. Weiler, M. Geiger, M. Welling, W. Boomsma, and T. S. Cohen 3d steerable cnns: learning rotationally equivariant features in volumetric data. Advances in Neural information processing systems 31. Cited by: §2. Yang et al. (2015) J. B. Yang, M. N. Nguyen, P. P. San, X. L. Li, and S. Krishnaswamy Deep convolutional neural networks on multichannel time series for human activity recognition. In Proceedings of the Twenty-Fourth International Joint Conference on Artificial Intelligence, p. 3995–4001. Cited by: §4.2, Table 2. Yu et al. (2022) X. Yu, T. Ma, J. Jang, and S. Xiong Data augmentation to address various rotation errors of wearable sensors for robust pre-impact fall detection. IEEE journal of biomedical and health informatics 27 (5), p. 2197–2207. Cited by: §1, §2. Yurtman et al. (2018) A. Yurtman, B. Barshan, and B. Fidan Activity recognition invariant to wearable sensor unit orientation using differential rotational transformations represented by quaternions. Sensors 18 (8), p. 2725. Cited by: §2. Yurtman and Barshan (2017) A. Yurtman and B. Barshan Activity recognition invariant to sensor orientation with wearable motion sensors. Sensors 17 (8), p. 1838. Cited by: §1, §2. Zhou et al. (2024) Y. Zhou, T. King, H. Zhao, Y. Huang, T. Riedel, and M. Beigl Mlp-har: boosting performance and efficiency of har models on edge devices with purely fully connected layers. In Proceedings of the 2024 ACM International Symposium on Wearable Computers, p. 133–139. Cited by: §4.2, Table 2. Zhou et al. (2022) Y. Zhou, H. Zhao, Y. Huang, T. Riedel, M. Hefenbrock, and M. Beigl Tinyhar: a lightweight deep learning model designed for human activity recognition. In Proceedings of the 2022 ACM International Symposium on Wearable Computers, p. 89–93. Cited by: §4.1, §4.2, Table 2.