Paper deep dive
Zero-Fi: Zero-Shot Wi-Fi-Based Human Activity Recognition via Contrastive Signal-Language Alignment
Yitong Shen, Cheng Guo, Peiliang Wang, Jingzhe Zhang, Yi Sheng, Haopeng Zhang, Hongfei Xue, Yili Ren
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/4/2026, 10:32:17 AM
Summary
The paper introduces Zero-Fi, a framework for zero-shot Wi-Fi-based human activity recognition (HAR). It aligns Wi-Fi signal features (Doppler, phase, amplitude) with natural language descriptions of activities using contrastive learning. The method addresses hardware noise via conjugate multiplication and environmental variance via a domain discriminator, enabling the recognition of unseen activity classes without labeled training data.
Entities (8)
Relation Signals (6)
Zero-Fi → targets → Wi-Fi-based Human Activity Recognition
confidence 98% · We present Zero-Fi, a contrastive signal-language alignment framework for zero-shot Wi-Fi-based human activity recognition.
Zero-Fi → uses → Contrastive Learning
confidence 95% · Zero-Fi learns unified representations from complementary Wi-Fi signal features and aligns them with the semantic representations of natural-language activity descriptions in a shared embedding space.
Zero-Fi → utilizes → Channel State Information
confidence 92% · To capture Wi-Fi signal changes caused by a target, we utilize the Channel State Information (CSI).
Zero-Fi → extracts → Doppler Frequency Shift
confidence 90% · We then extract three complementary representations from Wi-Fi signals: Doppler frequency shift (DFS), phase, and amplitude.
Zero-Fi → employs → Domain Discriminator
confidence 88% · we introduce a domain discriminator that encourages the activity feature encoder... to learn domain-invariant activity representations.
Zero-Fi → mitigates → Conjugate Multiplication
confidence 85% · To suppress hardware-induced noise, we leverage the fact that distortions are shared across antennas... and cancel them out by computing the conjugate multiplication
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Wi-Fi-based human activity recognition has advanced substantially, but most existing methods assume a closed set of activities and require labeled Wi-Fi samples for every target class, limiting their ability to recognize unseen activities. We present Zero-Fi, a contrastive signal-language alignment framework for zero-shot Wi-Fi-based human activity recognition. Zero-Fi learns unified representations from complementary Wi-Fi signal features and aligns them with the semantic representations of natural-language activity descriptions in a shared embedding space. This cross-modal alignment enables Zero-Fi to recognize new activity classes without requiring labeled Wi-Fi samples or model adaptation for those classes. Experiments on large-scale public benchmark datasets demonstrate effective zero-shot recognition of held-out activity classes, highlighting the potential of signal-language alignment to extend Wi-Fi sensing beyond predefined activity classes.
Tags
Links
- Source: https://arxiv.org/abs/2607.26381v1
- Canonical: https://arxiv.org/abs/2607.26381v1
Trouble viewing inline? Open PDF directly →
Full Text
52,052 characters extracted from source content.
Expand or collapse full text
Zero-Fi: Zero-Shot Wi-Fi-Based Human Activity Recognition via Contrastive Signal-Language Alignment Yitong Shen1 , Cheng Guo2 , Peiliang Wang1, Jingzhe Zhang1, Yi Sheng1, Haopeng Zhang2, Hongfei Xue2 , Yili Ren1 Abstract Wi-Fi-based human activity recognition has advanced substantially, but most existing methods assume a closed set of activities and require labeled Wi-Fi samples for every target class, limiting their ability to recognize unseen activities. We present Zero-Fi, a contrastive signal-language alignment framework for zero-shot Wi-Fi-based human activity recognition. Zero-Fi learns unified representations from complementary Wi-Fi signal features and aligns them with the semantic representations of natural-language activity descriptions in a shared embedding space. This cross-modal alignment enables Zero-Fi to recognize new activity classes without requiring labeled Wi-Fi samples or model adaptation for those classes. Experiments on large-scale public benchmark datasets demonstrate effective zero-shot recognition of held-out activity classes, highlighting the potential of signal-language alignment to extend Wi-Fi sensing beyond predefined activity classes. Introduction Human activity recognition (HAR) is an important technology that automatically detects and identifies human physical movements (Kaur et al. 2024). Over the decades, HAR has been extensively studied (Zhang et al. 2020) and has demonstrated broad applicability in various areas such as healthcare (Ge et al. 2022), security (Jiang et al. 2013), human-computer interaction (Kellogg et al. 2014), and smart homes (Du et al. 2019). Owing to its substantial practical value, HAR has been realized through various conventional sensing modalities, including cameras (Tran et al. 2018), inertial measurement units (IMUs) (Zhang et al. 2022), wearable sensors (Ordóñez and Roggen 2016), millimeter-wave (mmWave) radar (Singh et al. 2019), and LiDAR (Luo et al. 2020). However, each modality has inherent limitations. Specifically, IMU- and wearable-based systems require users to continuously carry or wear dedicated devices/sensors, which can be intrusive and inconvenient. Radar- and LiDAR-based systems (Cui and Dahnoun 2021; Bouazizi et al. 2023) often incur non-negligible hardware costs and deployment complexity. The performance of camera-based HAR systems (Zhang et al. 2019; Liu et al. 2019) is highly sensitive to illumination conditions and may deteriorate substantially in low-light or dark environments. In recent years, Wi-Fi has emerged as a promising sensing modality for HAR. Human movements alter the propagation paths of Wi-Fi signals through reflection, diffraction, and scattering, producing measurable variations in signals (Tan et al. 2022). These variations encode motion-related information that can be analyzed to distinguish different human activities. Compared with conventional sensing modalities, Wi-Fi enables contactless sensing and remains robust under varying ambient illumination conditions. Moreover, existing Wi-Fi infrastructure and widely available Wi-Fi devices can be repurposed for HAR, thereby reducing deployment costs. The pervasive coverage of Wi-Fi signals in indoor environments further provides a practical foundation for scalable and widely deployable HAR systems. Owing to these advantages, Wi-Fi-based HAR has attracted growing research interest, leading to the development of numerous systems (Adib and Katabi 2013; Wang et al. 2015; Zheng et al. 2019; Xiao et al. 2021; Li et al. 2021; Zhang et al. 2026). However, most existing systems have limited ability to generalize to previously unseen activity categories. They typically rely on annotated training data and strong supervision under closed-set assumptions, and thus cannot recognize activities that are absent from the training set. This limitation reduces their practical applicability because deployed HAR systems may encounter activity categories that are absent from the training data. Given the breadth and diversity of human behavior, collecting and annotating representative Wi-Fi samples for every possible activity category is prohibitively expensive and ultimately infeasible. The ability to recognize unseen activities is therefore essential for Wi-Fi-based HAR systems. In this paper, we present Zero-Fi, a novel framework for zero-shot Wi-Fi-based human activity recognition that can effectively recognize unseen activities. Realizing this capability, however, requires addressing three key challenges. First, zero-shot Wi-Fi-based HAR requires learning representations from Wi-Fi signals that generalize beyond the activity categories observed during training. The key challenge is to learn Wi-Fi representations that capture shared, transferable motion characteristics, rather than features specific to a fixed set of predefined activities. To address this challenge, we align Wi-Fi representations with the semantic space of natural language using contrastive learning, which offers a scalable interface for describing human activities (Guadarrama et al. 2013), including categories absent from the Wi-Fi training data. Unlike discrete labels that treat each activity as an independent category, language descriptions explicitly encode motion attributes, such as body parts, movement directions, and speeds, that are shared across activities. Although Wi-Fi signals and language differ substantially in form, both characterize the same underlying human motion: Wi-Fi signals capture the physical effects of body movements on wireless signal propagation, while language expresses those movements at a semantic level. Aligning the two modalities therefore enables the model to associate Wi-Fi signals with these shared motion attributes. Since unseen activities can often be characterized as new combinations of attributes already observed in seen activities, the resulting signal-language correspondence naturally extends beyond the training classes, enabling zero-shot recognition. Second, Wi-Fi signals are inherently sensitive to environmental variations. Since the signal component reflected off the human body is often superimposed with signals reflected from surrounding objects, even minor environmental changes (e.g., furniture relocation) can cause the received signals to differ substantially for the same activity. In addition, Wi-Fi devices introduce hardware-induced phase noise, such as random phase offset and carrier frequency offset (Ma et al. 2019), which further distorts the signal. The challenge is therefore to reliably extract activity-relevant features from Wi-Fi signals that remain robust to such environmental and hardware noise. To address these, we introduce a domain discriminator that encourages the feature encoder to suppress environment-specific information while preserving motion patterns. To suppress hardware-induced noise, we leverage the fact that distortions are shared across antennas on the same Wi-Fi device, and cancel them out by computing the conjugate multiplication of signals received across antennas, thereby removing the common noise component while preserving activity-relevant signal variations. Third, existing Wi-Fi HAR datasets typically provide only simple activity labels. Directly aligning Wi-Fi representations with such brief text labels is unlikely to yield satisfactory generalization, as simple labels fail to capture detailed motion attributes and the subtle relationships among activities. For example, “answering the phone” and “combing hair” are semantically distant as class names, yet both activities involve moving the arm toward the head. The challenge is therefore to obtain semantically rich language descriptions that explicitly characterize the underlying human motion. To address this, we leverage the broad semantic knowledge encoded in large language models (LLMs) to decompose each activity label into motion attributes associated with the torso, head, arms, hands, legs, and overall trajectory. This allows us to explicitly model the semantic relationships between activities that share common motion components: although “answering the phone” and “combing hair” are distinct as class names, their attribute-level descriptions both involve moving the arm toward the head, resulting in high similarity between their attribute embeddings. We evaluate Zero-Fi on public datasets under a strict zero-shot setting, in which the Wi-Fi signals and language descriptions of all test activity categories are unseen during training. Results demonstrate that Zero-Fi effectively recognizes unseen activities, achieving an average accuracy of 69.58%, and consistently outperforms existing baselines. Our main contributions are summarized as follows: • We present Zero-Fi, a contrastive framework for zero-shot Wi-Fi-based human activity recognition that aligns Wi-Fi signals with semantically rich language descriptions. • We design mechanisms to extract robust Wi-Fi features and to construct rich, attribute-level language descriptions for each activity, enabling cross-modal alignment. • We conduct extensive experiments on multiple public datasets, demonstrating the effectiveness and generalization of our work under strict zero-shot settings. Related Work Human Activity Recognition HAR has been implemented using diverse sensing modalities (Kaur et al. 2024). For example, IF-ConvTransformer (Zhang et al. 2022) fuses measurements from multiple IMU sensors and employs a convolutional Transformer for activity recognition. TSAM (Li et al. 2025a) augments a pretrained visual backbone with a sequential perceiver adapter to capture spatial and temporal features. mmCLIP (Cao et al. 2024) aligns mmWave radar heatmaps with textual embeddings to recognize unseen activities, while Luo et al. (Luo et al. 2020) combine LSTM and temporal convolutional networks to classify activities from LiDAR point clusters. However, these modalities face inherent limitations in deployment cost, hardware requirements, and robustness to illumination conditions. In contrast, Wi-Fi-based HAR offers a practical and scalable alternative by leveraging the ubiquitous Wi-Fi devices and signals. Wi-Fi Sensing Wi-Fi sensing has been widely studied for object sensing (Ren et al. 2020; Wang et al. 2024c), localization (Kotaru et al. 2015; Fan et al. 2024), pose estimation (Ren et al. 2022; Yan et al. 2024), security (Jiang et al. 2013; Lin et al. 2023), healthcare (Ge et al. 2022; Wang et al. 2024b), smart homes (Du et al. 2019; Li et al. 2025c), and HAR (Li et al. 2021; Xiao et al. 2021; Li et al. 2025b; Zhang et al. 2026). For HAR, THAT (Li et al. 2021) models time-channel dependencies using a two-stream convolution-augmented Transformer, while OneFi (Xiao et al. 2021) applies meta-learning (Hospedales et al. 2021) for one-shot gesture adaptation. However, they cannot support strict zero-shot recognition. Wi-Chat (Ren et al. 2025) uses an LLM to infer activities without task-specific training, but its reliance on manually designed signal-description prompts limits its demonstrated zero-shot capability to four activity categories. Preliminary Wi-Fi Sensing Basics Wi-Fi has evolved beyond its traditional role as a communication technology to become a promising sensing modality (Tan et al. 2022). As illustrated in Figure 1, a Wi-Fi transmitter emits signals that propagate through the environment, interact with the human body through reflection, diffraction, and scattering, and are subsequently captured by a Wi-Fi receiver. Human movements alter these propagation paths, producing measurable variations in the received signals. Because different activities generate distinct signal patterns, these variations can be analyzed to infer human motion and recognize activities. Figure 1: Wi-Fi sensing illustration. Wi-Fi Channel State Information To capture Wi-Fi signal changes caused by a target, we utilize the Channel State Information (CSI). It describes how a transmitted signal propagates through the wireless medium and captures the effects of human movements and the environment. Wi-Fi CSI could be expressed as H(f,t)=∑iNAie−j2πdi(t)λ.H(f,t)= _i^NA_ie^-j2π d_i(t)λ. Here, AiA_i is the complex attenuation, di(t)d_i(t) is the length of the ithi^th path, N is the total number of paths, and λ is the wavelength. As shown in Figure 1, the received Wi-Fi signal can be divided into static and dynamic components. Specifically, static components are the signal reflections from static objects (e.g., furniture) and dynamic components are the signals reflected off the dynamic human body. Thus, CSI can be further expressed as H(f,t)=Hs(f,t)+Hdyn(f,t)=Hs(f,t)+a(f,t)e−j2πd(t)λH(f,t)=H_s(f,t)+H_dyn(f,t)=H_s(f,t)+a(f,t)e^-j2π d(t)λ, where Hs(f,t)H_s(f,t) and Hdyn(f,t)H_dyn(f,t) denote the static and dynamic components, respectively. a(f,t)a(f,t) represents the amplitude attenuation of the dynamic component, e−j2πd(t)λe^-j2π d(t)λ denotes the corresponding phase, d(t)d(t) is the propagation path length of the dynamic component, and λ is the signal wavelength. Methodology Figure 2: Overview of Zero-Fi. As shown in Figure 2, Zero-Fi comprises signal, text, and signal-text contrastive learning modules. Signal Module Activity Feature Extractor. Commodity Wi-Fi devices introduce hardware noise including random phase offset (RPO) and carrier frequency offset (CFO) (Kotaru et al. 2015; Li et al. 2016). In particular, RPO and CFO induce time-varying phase offsets that distort the temporal phase variations of CSI, thereby degrading the fidelity of the extracted signal representations. Accordingly, Wi-Fi CSI can be modeled as H(f,t)=e−jθ(Hs(f,t)+a(f,t)e−j2πd(t)λ),H(f,t)=e^-jθ(H_s(f,t)+a(f,t)e^-j2π d(t)λ), where e−jθe^-jθ represents the aggregation of these phase offsets. Because these phase offsets are the same across the antennas of the same Wi-Fi device, conjugate multiplication of CSIs between antennas can be applied to suppress them: Hcm(f,t) H_cm(f,t) =H1(f,t)H¯2(f,t) =H_1(f,t) H_2(f,t) =(e−jθ(H1,s(f,t)+a1(f,t)e−j2πd1(t)λ)) =(e^-jθ(H_1,s(f,t)+a_1(f,t)e^-j2π d_1(t)λ)) (ejθ(H¯2,s(f,t)+a2(f,t)ej2πd2(t)λ)) (e^jθ( H_2,s(f,t)+a_2(f,t)e^j2π d_2(t)λ)) =H1,s(f,t)H¯2,s(f,t)⋯① =H_1,s(f,t) H_2,s(f,t) ·s ① +a1(f,t)a2(f,t)e−j2πd1(t)−d2(t)λ⋯② +a_1(f,t)a_2(f,t)e^-j2π d_1(t)-d_2(t)λ ·s ② +H¯2,s(f,t)a1(f,t)e−j2πd1(t)λ))⋯③ + H_2,s(f,t)a_1(f,t)e^-j2π d_1(t)λ)) ·s ③ +H1,s(f,t)a2(f,t)ej2πd2(t)λ.⋯④ +H_1,s(f,t)a_2(f,t)e^j2π d_2(t)λ. ·s ④ (1) Here, Hcm(f,t)H_cm(f,t) denotes the result of conjugate multiplication, while H1(f,t)H_1(f,t) and H¯2(f,t) H_2(f,t) denote the CSI obtained from one antenna and the complex-conjugated CSI obtained from another antenna, respectively. Term ① is motion-invariant and removed by high-pass filtering. Term ② has a small magnitude and is therefore negligible, and terms ③ and ④ retain the dominant motion-induced variations. The denoised CSIs are subsequently used for feature extraction. We then extract three complementary representations from Wi-Fi signals: Doppler frequency shift (DFS), phase, and amplitude. DFS captures the time-frequency dynamics of human motion, phase preserves fine-grained activity variations, and amplitude reflects coarse-grained motion patterns (Qian et al. 2017; Ren et al. 2020). For DFS extraction, we first compute the maximum mean-to-variance ratio of the amplitude for each transmit-receive antenna pair’s CSI. The CSIs with higher amplitude values and lower amplitude variance tend to contain less dynamic information (Qian et al. 2017). Therefore, removing these CSIs helps retain signals with richer motion-related dynamics. The Short-Time Fourier Transform (STFT) is subsequently applied to obtain the time-frequency representation, from which the frequency bins within the target DFS range are retained (Zheng et al. 2019). An additional ℓ1 _1 normalization step is applied to the extracted DFS representations to reduce scale differences across samples. The signal phase can be extracted using the angle(⋅)angle(·) function as follows HP=unwrap(angle(Hcm(f,t))).H_P=unwrap(angle(H_cm(f,t))). Because phase measurements are periodic with a period of 2π2π, the unwrap(⋅)unwrap(·) function is applied to eliminate discontinuities caused by phase wrapping. Then, a sliding-window smoothing filter is applied to HPH_P to reduce residual fluctuations and improve the stability of the extracted phase. The signal amplitude HAH_A can be calculated as HA=abs(H(f,t)).H_A=abs(H(f,t)). We also apply a sliding-window smoothing filter to HAH_A to improve the stability of the amplitude. The extracted Wi-Fi representations for human activities are shown in Figure 3. Figure 3: DFS, phase, and amplitude features. Domain Discriminator. CSIs are affected by both dataset-specific acquisition conditions and environment-specific characteristics. Because these factors are often coupled in public Wi-Fi sensing datasets, we treat each unique dataset-environment configuration as an independent domain. To mitigate the influence of configuration-specific characteristics, we introduce a domain discriminator that encourages the activity feature encoder in Figure 2 to learn domain-invariant activity representations. Each training sample xix_i is therefore associated with a domain label di∈1,…,Ndd_i∈\1,…,N_d\, where NdN_d denotes the number of domains represented in the training data. Let i=FAFE(xi)z_i=F_AFE(x_i) denote the signal embedding produced by the activity feature encoder FAFE(⋅)F_AFE(·). We first calculate its cosine similarity sim(⋅)sim(·) to the text embedding ct_c of each seen activity class: i=softmax(exp(τ)sim(¯i,¯c)),q_i=softmax( (τ)sim ( z_i, t_c )), where ¯i z_i and ¯c t_c denote ℓ2 _2-normalized signal and text embeddings, respectively, τ is the learnable log-temperature parameter, and iq_i is the predicted activity distribution over the seen classes. Only descriptions of the seen activities are used to calculate iq_i during training. The domain discriminator is conditioned on both the signal representation and its predicted activity distribution (Zhao et al. 2017): ^id=Dϕ(GRLγd(i)⊕sg(i)), p_i^\,d=D_φ(GRL_ _d(z_i) (q_i)), where GRLγd(⋅)GRL_ _d(·) denotes a gradient-reversal layer with coefficient γd _d, sg(⋅)sg(·) denotes the stop-gradient operation, ⊕ denotes concatenation, and DϕD_φ denotes the domain discriminator. Conditioning the discriminator on the activity prediction helps account for activity-dependent differences between domains, while stopping the gradient through iq_i prevents the domain objective from directly altering the activity predictions. The discriminator is trained using the domain-classification loss ℒdomain=−1B∑i=1B∑j=1Ndℐ[di=j]logp^i,jd,L_domain=- 1B _i=1^B _j=1^N_dI[d_i=j] p_i,j^\,d, (2) where B denotes the batch size, ℐ[⋅]I[·] is the indicator function, and p^i,jd p_i,j^\,d is the predicted probability that sample i belongs to dataset-environment configuration j. We use domain-balanced mini-batches to reduce the influence of domain imbalance. The overall training objective is ℒtotal=ℒcon+ℒdomain.L_total=L_con+L_domain. (3) During backpropagation, the gradient-reversal layer multiplies the gradient from ℒdomainL_domain to the activity feature encoder by −γd- _d. Thus, the domain discriminator is optimized to correctly identify the dataset-environment configuration, whereas the activity feature encoder is optimized to make this prediction difficult while preserving signal-text alignment. This adversarial objective encourages the learned signal representations to become less dependent on configuration-specific acquisition and propagation characteristics. Activity Feature Encoder. We design three distinct, representation-specific Transformer-based signal encoders, one for each extracted CSI representation: amplitude, phase, and DFS. For the amplitude and phase encoders, the two-dimensional convolutional layer commonly adopted in standard Vision Transformer (ViT) architectures (Dosovitskiy et al. 2020) is replaced with a one-dimensional convolutional layer. This design captures the temporal structure of the amplitude and phase representations while preserving dependencies among their constituent measurements, which contain discriminative activity-related information. Specifically, given the amplitude and phase inputs HA,HP∈ℛt×NscH_A,H_P ^t× N_sc, where t denotes the temporal length and NscN_sc denotes the number of subcarriers, the corresponding encoders apply one-dimensional convolutional patch embedding, temporal positional encoding, and self-attention to each representation. The resulting embeddings for the encoded amplitude and phase representations are denoted by EmbA,EmbP∈ℛ⌊t/s⌋×dSEEmb_A,Emb_P t/s × d_SE, respectively, where s denotes the temporal downsampling factor and dSEd_SE denotes the output feature dimension of the signal encoders. For the DFS encoder, we retain the standard ViT architecture because the STFT-derived DFS jointly captures both temporal and frequency-domain features. For the DFS input HD∈ℛt×Nf×1H_D ^t× N_f× 1, where NfN_f denotes the number of frequency bins, a two-dimensional convolutional patch embedding with a window size of w×hw× h is applied. The result is projected to an embedding EmbD∈ℛ⌊t/w⌋×dSEEmb_D t/w × d_SE. The encoded amplitude, phase, and DFS embeddings have identical dimensions, ℛT×dSER^T× d_SE, where T=⌊t/w⌋T= t/w denotes the temporal length. All three representations preserve temporal information, and the amplitude and phase are inherently sequential; their encoded embeddings are concatenated along the feature dimension while maintaining temporal correspondence across representations. A learnable positional encoding PE(⋅)PE(·) is subsequently applied along the temporal dimension to capture temporal dependencies: Emb=PE(EmbD⊕EmbP⊕EmbA)Emb=PE(Emb_D Emb_P Emb_A). Here, EmbEmb denotes the concatenated embedding, and ⊕ denotes concatenation along the feature dimension. Consequently, the resulting embedding EmbEmb has dimensions (T,3dSE)(T,3d_SE). Next, a signal embedding extractor (SEE) is used to learn shared representations from the concatenated modality embeddings. The SEE employs multi-head self-attention to model cross-modal interactions before forwarding the resulting features to the subsequent aggregation stage. The attention output is further processed using layer normalization LN(⋅)LN(·) and a multilayer perceptron MLP(⋅)MLP(·), together with a residual connection: z=Attention(Q,K,V),z=Attention(Q,K,V), EmbSEE=MLP(LN(z))+z.Emb_SEE=MLP(LN(z))+z. Here, the query Q, key K, and value V are obtained through separate linear projections of the concatenated embedding EmbEmb. Finally, we employ a single transformer layer, termed the signal embedding aggregator (SEA), to further aggregate and refine the representations produced by the SEE: Embsignal=Linear(SEA(EmbSEE)).Emb_signal=Linear(SEA(Emb_SEE)). Here, Linear(⋅)Linear(·) denotes a linear projection layer. The resulting representation EmbsignalEmb_signal serves as the final CSI-derived embedding. Text Module LLM-Based Activity Describer. Effective signal-language alignment requires richer semantic supervision than that provided by a single activity label. We therefore employ an LLM, pretrained on large-scale text corpora, to construct the activity describer. The architecture of this component is illustrated in Figure 4. To constrain and standardize its output, we design a structured prompt that instructs the LLM to expand each activity label into a detailed description of body-part movements, including those of the head, arms, hands, torso, and legs, as well as the overall movement direction and trajectory. Specifically, the activity label and the structured prompt are jointly provided to the LLM to generate the corresponding activity description. To further improve descriptive accuracy, we exploit the multimodal reasoning capabilities of contemporary LLMs by supplying relevant visual references, such as images and videos, when available. These descriptions provide semantically informative representations of both seen and unseen daily activities, thereby facilitating subsequent alignment with the corresponding signal features. Figure 4: LLM-based activity describer. CLIP-Based Description Encoder. Because language descriptions and Wi-Fi CSI measurements are not inherently aligned, their representations must be projected into a shared latent space. CLIP (Radford et al. 2021) was originally trained to align textual and visual representations through contrastive learning. Its text encoder therefore captures visually grounded semantic information related to objects, actions, and human motion, which can provide a useful supervisory representation for activity-related signal features. Accordingly, we employ the frozen CLIP text encoder to encode the generated activity descriptions for subsequent contrastive alignment with the CSI representations. The resulting text embeddings are denoted by Embt−encEmb_t-enc. In addition, we introduce a single transformer layer, termed the Text Embedding Aggregator (TEA), to aggregate and refine the embeddings of multiple descriptions into a unified representation, denoted by EmbtextEmb_text, for downstream tasks. Signal-Text Alignment The objective of this stage is to train the overall framework to produce semantically meaningful signal embeddings and align them with the corresponding textual representations. This alignment is achieved through contrastive learning, which increases the similarity between positive signal-text pairs associated with the same activity while decreasing the similarity between negative pairs associated with different activities. Specifically, we employ the InfoNCE objective (Radford et al. 2021) to encourage bidirectional alignment between the signal and textual embeddings. The signal embeddingEmbsignalEmb_signal and text embedding EmbtextEmb_text are first ℓ2 _2-normalized. Their pairwise similarities are computed as ℓi,j=exp(τ)sim(Embsignali‖Embsignali‖2,Embtextj‖Embtextj‖2), _i,j= (τ)\,sim ( Emb_signal^i\|Emb_signal^i\|_2, Emb_text^j\|Emb_text^j\|_2 ), where τ is a learnable log-temperature parameter, sim(⋅,⋅)sim(·,·) denotes cosine similarity, and ℓi,j _i,j denotes the similarity between the iith signal sample and the jjth text sample. The resulting similarity matrix has dimensions B×B× B. To reduce the occurrence of false-negative pairs caused by multiple samples from the same activity class, we employ class-aware sampling to maximize activity-class diversity within each batch. Consequently, most batches contain at most one sample from each activity class, although duplicate classes may occasionally occur because of sampling constraints. Cross-entropy losses are then computed in both alignment directions: text-to-signal and signal-to-text. The overall contrastive loss is defined as Lcon=12(aLCE,t2s+bLCE,s2t),L_con= 12(aL_CE,t2s+bL_CE,s2t), (4) where LCE,t2sL_CE,t2s and LCE,s2tL_CE,s2t denote the text-to-signal and signal-to-text contrastive losses, respectively. The weighting coefficients a and b are set to 11 by default. Each directional loss is implemented using cross-entropy: LCE=−1B∑i=1B∑j=1Byi,jlogy^i,j,L_CE=- 1B _i=1^B _j=1^By_i,j y_i,j, where B denotes the batch size, yi,j=1y_i,j=1 when the iith signal and jjth text form the paired sample, i.e., i=ji=j, and yi,j=0y_i,j=0 otherwise. The predicted probability y^i,j y_i,j is obtained by applying softmax along the corresponding signal-to-text or text-to-signal direction of the similarity matrix. The contrastive loss is applied not only to the final aggregated embeddings but also to the intermediate signal and text embeddings, EmbSEEEmb_SEE and Embt−encEmb_t-enc, respectively. This intermediate-level supervision enables the signal encoders to learn more directly from the textual embeddings. Through bidirectional contrastive supervision at both the intermediate and final embeddings, the model receives semantic guidance from both the less-processed features and the refined embeddings. Experiment Datasets We evaluate our method on two of the largest Wi-Fi datasets, WiDAR 3.0 (Zheng et al. 2019) and XRF55 (Wang et al. 2024a), which are combined for training and zero-shot testing. WiDAR 3.0 is one of the largest publicly available datasets for human activity recognition. It contains approximately 270,000 samples covering 22 activities, 17 participants, and 3 environments. XRF55 is a large-scale and complex dataset for Wi-Fi-based human activity recognition. It contains 128,700 samples covering 55 human daily activities, 31 participants, and 4 environments. In our experiments, we retain all XRF55 activities and remove two overlapping activity classes from WiDAR 3.0, resulting in a total of 75 activity classes. Data Processing We follow the procedure described in the activity feature extractor to preprocess the CSI measurements and extract the corresponding signal representations using MATLAB. To enable batch processing and ensure consistent model inputs, all samples are resampled to a fixed temporal length of 1,000 Wi-Fi packets before being fed into the activity feature encoder. After preprocessing and flattening, the amplitude and phase representations have dimensions of 1000×901000× 90, whereas the DFS representation has dimensions of 1000×121×11000× 121× 1. Model and Environment Settings As described in Section Methodology, each signal encoder comprises two ViT-style Transformer layers, whereas the signal embedding extractor contains four self-attention blocks. The signal embedding aggregator and text embedding aggregator each consist of a single Transformer layer. The signal embedding aggregator includes an additional linear projection layer. Both the signal and text modules produce 512-dimensional output embeddings. For the LLM-based activity describer, we employ ChatGPT 5.5 as the backbone model. The visual aids are provided with the original datasets and are used only as class-level references for generating activity descriptions. During pretraining, the model is optimized for 60,000 iterations using Adam (Kingma and Ba 2014) with a learning rate of 10−410^-4 and a batch size of 32. Training requires approximately 7 hours. The framework is implemented in PyTorch 2.0.2 with Python 3.9.18 and CUDA 11.8. All experiments are conducted on a Linux 24.04 platform equipped with an Intel(R) Core(TM) i9-13900 CPU, an NVIDIA GeForce RTX 4090 GPU with 24 GB of video memory. Zero-Shot Evaluation Protocol We evaluate all zero-shot-compatible methods under a class-disjoint protocol. Specifically, in each trial, the activity categories are partitioned into a seen-class set and an unseen-class set, with no overlap between them. All model parameters are trained exclusively using Wi-Fi samples from the seen classes. Wi-Fi samples from the unseen classes are used only during final evaluation and are not involved in model training, hyperparameter selection, receiver selection, or checkpoint selection. Additionally, no classifier is trained or adapted using unseen-class Wi-Fi samples. During inference, the descriptions of the unseen activities are introduced as fixed candidate semantic representations, and only to classify the held-out Wi-Fi samples. Participants and environments are not treated as additional holdout factors in the current evaluation; therefore, the reported results specifically measure generalization across activity categories rather than cross-subject or cross-environment generalization. We designate seven activity categories out of all 75 classes as unseen classes. The unseen categories are randomly selected in six independent trials. We use classification accuracy as the evaluation metric for the unseen-activity classification task, measuring the proportion of unseen activity samples that are correctly classified. To reduce the influence of an exceptionally easy or difficult class partition, we report a one-sample-per-tail trimmed mean: for each method, the highest and lowest split accuracies are removed, and the remaining accuracies are averaged. Baselines We compare Zero-Fi with five representative baselines on the WiFi-based HAR task. The baselines include THAT (Li et al. 2021), a Transformer-based HAR system under supervised learning that achieves strong recognition performance. OneFi (Xiao et al. 2021), a Transformer-based meta-learning framework designed to recognize previously unseen activities in a one-shot setting. CLAR (Xiao et al. 2024), a diffusion-based contrastive learning system for HAR. FM-ZSL-IoT (Xue et al. 2024) and Wi-CLIP (Zhang et al. 2025), both employ Transformer-based contrastive learning for zero-shot WiFi-based HAR. For all baselines, we implement the architectures and training procedures based on their original papers and publicly released code, when available. Necessary modifications are made to accommodate our datasets and zero-shot evaluation protocol. For baselines that natively support signal-language zero-shot inference, i.e., FM-ZSL-IoT and Wi-CLIP, we retain their original semantic-inference procedures while adopting our activity splits and input data settings. THAT, OneFi, and CLAR do not natively provide a mechanism for predicting unseen activity labels. They demonstrate that potentially transferable signal representations are insufficient for semantic zero-shot recognition when the system lacks a mechanism for associating unseen signal patterns with unseen activity labels. Nevertheless, we pretrain their original architectures on the seen activities, use the resulting models as frozen signal encoders, and attach untrained, randomly initialized classifiers whose output dimensions correspond to the held-out activity classes. No Wi-Fi sample from unseen activities is used to optimize these classifiers. Results Method Accuracy (%) THAT (Li et al. 2021) 16.60 OneFi (Xiao et al. 2021) 11.33 CLAR (Xiao et al. 2024) 15.79 FM-ZSL-IoT (Xue et al. 2024) 23.86 Wi-CLIP (Zhang et al. 2025) 26.79 Zero-Fi 69.58 Table 1: Performance comparisons on unseen activities. Comparisons with Baselines. Following the zero-shot evaluation protocol described above, all methods are evaluated on identical seen-unseen activity-class splits The trimmed mean accuracy across these splits is reported in Table 1. Among the methods that directly support zero-shot inference, Zero-Fi achieves the best performance, with an accuracy of 69.58%, substantially outperforming Wi-CLIP and FM-ZSL-IoT. The performance of THAT, OneFi, and CLAR under this setting reflects the absence of a learned correspondence between their extracted signal representations and the held-out activity labels. Zero-Fi achieves this performance by explicitly aligning Wi-Fi signal representations with natural-language activity descriptions. Although FM-ZSL-IoT achieves an accuracy of 23.86%, its fine-tuning stage relies on GAN-generated synthetic data, which may introduce distributional discrepancies and artifacts that limit further performance improvements (Gong et al. 2025). Wi-CLIP (Zhang et al. 2025) achieves an accuracy of 26.79%. However, its reliance on relatively uninformative single-sentence descriptions and a BERT-based text encoder that is not explicitly optimized for signal-language alignment (Koroteev 2021) limits its zero-shot inference capability. Figure 5: Ablation study: impact of signal representations. Ablation Studies. We first evaluate the contribution of individual CSI representations and their combinations through an ablation study. As shown in Figure 5, amplitude achieves an accuracy of 61.09%, phase achieves the highest accuracy among the individual representations at 65.19%, and DFS achieves an accuracy of 48.99%. Combining all three representations yields the best accuracy of 69.58%, improving upon the strongest individual representation by 4.39%. Amplitude and phase jointly capture the overall characteristics of human motion, while DFS contributes complementary frequency-domain information that further improves zero-shot recognition accuracy for unseen activities. Together, amplitude, phase, and DFS provide complementary motion information, yielding a more comprehensive representation that improves zero-shot recognition of unseen activities. Text format Accuracy (%) Simple Activity Label 48.49 Language Description 69.58 Table 2: Ablation study: LLM-based activity describer. We further evaluate the effectiveness of the LLM-based activity describer. Specifically, we directly feed simple activity labels, such as “walking” and “running,” into the CLIP-based text encoder while keeping all other model components unchanged. As shown in Table 2, replacing the generated activity descriptions with simple labels leads to a substantial decrease in unseen-activity classification accuracy. This result indicates that the enriched descriptions provide more informative semantic supervision by explicitly characterizing activity-related motion attributes, thereby improving the model’s ability to transfer to held-out activity classes. Domain Discriminator Accuracy (%) Off 63.68 On 69.58 Table 3: Ablation study: domain discriminator. Moreover, we evaluate the contribution of the domain discriminator by removing it while keeping all other experimental settings and hyperparameters unchanged. As shown in Table 3, disabling the domain discriminator reduces the recognition accuracy from 69.58% to 63.68%. This result demonstrates that the domain discriminator improves unseen-activity recognition, suggesting that reducing dataset and environment-specific variations benefits the learned activity embeddings. Conclusion In this work, we propose Zero-Fi, a contrastive signal-language alignment framework for zero-shot Wi-Fi-based human activity recognition. The framework leverages semantically rich textual representations and aligns them with CSI-derived signal embeddings in a shared latent space, thereby enabling the recognition of previously unseen activity categories without requiring labeled samples from those classes. Extensive experiments demonstrate that Zero-Fi achieves strong zero-shot recognition performance and exhibits robust generalization to unseen activities. References F. Adib and D. Katabi (2013) See through walls with wifi!. In Proceedings of the ACM SIGCOMM 2013 conference on SIGCOMM, p. 75–86. Cited by: Introduction. M. Bouazizi, A. Lorite Mora, and T. Ohtsuki (2023) A 2d-lidar-equipped unmanned robot-based approach for indoor human activity detection. Sensors 23 (5), p. 2534. Cited by: Introduction. Q. Cao, H. Xue, T. Liu, X. Wang, H. Wang, X. Zhang, and L. Su (2024) Mmclip: boosting mmwave-based zero-shot har via signal-text alignment. In Proceedings of the 22nd ACM conference on embedded networked sensor systems, p. 184–197. Cited by: Human Activity Recognition. H. Cui and N. Dahnoun (2021) High precision human detection and tracking using millimeter-wave radars. IEEE Aerospace and Electronic Systems Magazine 36 (1), p. 22–32. Cited by: Introduction. A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint arXiv:2010.11929. Cited by: Signal Module. Y. Du, Y. Lim, and Y. Tan (2019) A novel human activity recognition and prediction in smart home based on interaction. Sensors 19 (20), p. 4474. Cited by: Introduction, Wi-Fi Sensing. N. Fan, Z. Tian, A. Dubey, S. Deshmukh, R. Murch, and Q. Chen (2024) Multitarget device-free localization via cross-domain wi-fi rss training data and attentional prior fusion. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, p. 91–99. Cited by: Wi-Fi Sensing. Y. Ge, A. Taha, S. A. Shah, K. Dashtipour, S. Zhu, J. Cooper, Q. H. Abbasi, and M. A. Imran (2022) Contactless wifi sensing and monitoring for future healthcare-emerging trends, challenges, and opportunities. IEEE Reviews in Biomedical Engineering 16, p. 171–191. Cited by: Introduction, Wi-Fi Sensing. C. Gong, B. Liang, W. Gao, and C. Xu (2025) Data can speak for itself: quality-guided utilization of wireless synthetic data. In Proceedings of the 23rd Annual International Conference on Mobile Systems, Applications and Services, p. 209–222. Cited by: Results. S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko (2013) Youtube2text: recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition. In Proceedings of the IEEE international conference on computer vision, p. 2712–2719. Cited by: Introduction. T. Hospedales, A. Antoniou, P. Micaelli, and A. Storkey (2021) Meta-learning in neural networks: a survey. IEEE transactions on pattern analysis and machine intelligence 44 (9), p. 5149–5169. Cited by: Wi-Fi Sensing. Z. Jiang, J. Zhao, X. Li, J. Han, and W. Xi (2013) Rejecting the attack: source authentication for wi-fi management frames using csi information. In 2013 Proceedings IEEE INFOCOM, p. 2544–2552. Cited by: Introduction, Wi-Fi Sensing. H. Kaur, V. Rani, and M. Kumar (2024) Human activity recognition: a comprehensive review. Expert Systems 41 (11), p. e13680. Cited by: Introduction, Human Activity Recognition. B. Kellogg, V. Talla, and S. Gollakota (2014) Bringing gesture recognition to all devices. In 11th USENIX Symposium on Networked Systems Design and Implementation (NSDI 14), p. 303–316. Cited by: Introduction. D. P. Kingma and J. Ba (2014) Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980. Cited by: Model and Environment Settings. M. V. Koroteev (2021) BERT: a review of applications in natural language processing and understanding. arXiv preprint arXiv:2103.11943. Cited by: Results. M. Kotaru, K. Joshi, D. Bharadia, and S. Katti (2015) Spotfi: decimeter level localization using wifi. In Proceedings of the 2015 ACM conference on special interest group on data communication, p. 269–282. Cited by: Wi-Fi Sensing, Signal Module. B. Li, W. Cui, W. Wang, L. Zhang, Z. Chen, and M. Wu (2021) Two-stream convolution augmented transformer for human activity recognition. In Proceedings of the AAAI conference on artificial intelligence, Vol. 35, p. 286–293. Cited by: Introduction, Wi-Fi Sensing, Baselines, Table 1. B. Li, M. Liu, G. Wang, and Y. Yu (2025a) Frame order matters: a temporal sequence-aware model for few-shot action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 18218–18226. Cited by: Human Activity Recognition. R. Li, T. Deng, S. Feng, M. Sun, and J. Jia (2025b) Consense: continually sensing human activity with wifi via growing and picking. In Proceedings of the AAAI Conference on Artificial intelligence, Vol. 39, p. 14292–14300. Cited by: Wi-Fi Sensing. S. Li, Z. Liu, Q. Lv, Y. Zou, Y. Zhang, and D. Zhang (2025c) WiLife: long-term daily status monitoring and habit mining of the elderly leveraging ubiquitous wi-fi signals. ACM Transactions on Computing for Healthcare 6 (1), p. 1–29. Cited by: Wi-Fi Sensing. X. Li, S. Li, D. Zhang, J. Xiong, Y. Wang, and H. Mei (2016) Dynamic-music: accurate device-free indoor localization. In Proceedings of the 2016 ACM international joint conference on pervasive and ubiquitous computing, p. 196–207. Cited by: Signal Module. C. Lin, P. Wang, C. Ji, M. S. Obaidat, L. Wang, G. Wu, and Q. Zhang (2023) A contactless authentication system based on wifi csi. ACM Transactions on Sensor Networks 19 (2), p. 1–20. Cited by: Wi-Fi Sensing. Y. Liu, Z. Lu, J. Li, T. Yang, and C. Yao (2019) Deep image-to-video adaptation and fusion networks for action recognition. IEEE Transactions on Image Processing 29, p. 3168–3182. Cited by: Introduction. F. Luo, S. Poslad, and E. Bodanese (2020) Temporal convolutional networks for multiperson activity recognition using a 2-d lidar. IEEE Internet of Things Journal 7 (8), p. 7432–7442. Cited by: Introduction, Human Activity Recognition. Y. Ma, G. Zhou, and S. Wang (2019) WiFi sensing with channel state information: a survey. ACM Computing Surveys (CSUR) 52 (3), p. 1–36. Cited by: Introduction. F. J. Ordóñez and D. Roggen (2016) Deep convolutional and lstm recurrent neural networks for multimodal wearable activity recognition. Sensors 16 (1), p. 115. Cited by: Introduction. K. Qian, C. Wu, Z. Zhou, Y. Zheng, Z. Yang, and Y. Liu (2017) Inferring motion direction using commodity wi-fi for interactive exergames. In Proceedings of the 2017 CHI conference on human factors in computing systems, p. 1961–1972. Cited by: Signal Module. A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. (2021) Learning transferable visual models from natural language supervision. In International conference on machine learning, p. 8748–8763. Cited by: Text Module, Signal-Text Alignment. Y. Ren, S. Tan, L. Zhang, Z. Wang, Z. Wang, and J. Yang (2020) Liquid level sensing using commodity wifi in a smart home environment. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 4 (1), p. 1–30. Cited by: Wi-Fi Sensing, Signal Module. Y. Ren, Z. Wang, Y. Wang, S. Tan, Y. Chen, and J. Yang (2022) GoPose: 3d human pose estimation using wifi. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 6 (2), p. 1–25. Cited by: Wi-Fi Sensing. Y. Ren, H. Zhang, H. Yuan, J. Zhang, and Y. Shen (2025) Wi-chat: large language model-powered wi-fi-based human activity recognition. In Proceedings of the International Workshop on Environmental Sensing Systems for Smart Cities, p. 8–13. Cited by: Wi-Fi Sensing. A. D. Singh, S. S. Sandha, L. Garcia, and M. Srivastava (2019) Radhar: human activity recognition from point clouds generated through a millimeter-wave radar. In Proceedings of the 3rd ACM Workshop on Millimeter-wave Networks and Sensing Systems, p. 51–56. Cited by: Introduction. S. Tan, Y. Ren, J. Yang, and Y. Chen (2022) Commodity wifi sensing in ten years: status, challenges, and opportunities. IEEE Internet of Things Journal 9 (18), p. 17832–17843. Cited by: Introduction, Wi-Fi Sensing Basics. D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri (2018) A closer look at spatiotemporal convolutions for action recognition. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, p. 6450–6459. Cited by: Introduction. F. Wang, Y. Lv, M. Zhu, H. Ding, and J. Han (2024a) Xrf55: a radio frequency dataset for human indoor action analysis. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8 (1), p. 1–34. Cited by: Datasets. M. Wang, J. Huang, X. Zhang, Z. Liu, M. Li, P. Zhao, H. Yan, X. Sun, and M. Dong (2024b) Target-oriented wifi sensing for respiratory healthcare: from indiscriminate perception to in-area sensing. IEEE Network 39 (5), p. 201–208. Cited by: Wi-Fi Sensing. W. Wang, A. X. Liu, M. Shahzad, K. Ling, and S. Lu (2015) Understanding and modeling of wifi signal based human activity recognition. In Proceedings of the 21st annual international conference on mobile computing and networking, p. 65–76. Cited by: Introduction. X. Wang, J. Wang, K. Niu, J. Xiong, F. Zhang, E. Yi, A. Yu, Z. Yao, and D. Zhang (2024c) Wi2DMeasure: wifi-based 2d object size measurement. In Proceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems, p. 253–266. Cited by: Wi-Fi Sensing. C. Xiao, Y. Han, W. Yang, Y. Hou, F. Shi, and K. Chetty (2024) Diffusion-model-based contrastive learning for human activity recognition. IEEE Internet of Things Journal 11 (20), p. 33525–33536. Cited by: Baselines, Table 1. R. Xiao, J. Liu, J. Han, and K. Ren (2021) Onefi: one-shot recognition for unseen gesture via cots wifi. In Proceedings of the 19th ACM Conference on Embedded Networked Sensor Systems, p. 206–219. Cited by: Introduction, Wi-Fi Sensing, Baselines, Table 1. D. Xue, X. Fan, T. Chen, G. Lan, and Q. Song (2024) Leveraging foundation models for zero-shot iot sensing. arXiv preprint arXiv:2407.19893. Cited by: Baselines, Table 1. K. Yan, F. Wang, B. Qian, H. Ding, J. Han, and X. Wei (2024) Person-in-wifi 3d: end-to-end multi-person 3d pose estimation with wi-fi. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 969–978. Cited by: Wi-Fi Sensing. H. Zhang, Y. Guo, Z. Wang, Z. Sun, B. Guo, and Z. Yu (2025) Wi-clip: toward zero-shot air gesture recognition based on rf-text foundation model. In International Conference on Artificial Intelligence of Things and Systems, p. 174–189. Cited by: Baselines, Results, Table 1. H. Zhang, Y. Zhang, B. Zhong, Q. Lei, L. Yang, J. Du, and D. Chen (2019) A comprehensive survey of vision-based human action recognition methods. Sensors 19 (5), p. 1005. Cited by: Introduction. R. Zhang, S. Tang, H. Yan, X. Zhang, and J. Guo (2026) Wi-cbr: salient-aware adaptive wifi sensing for cross-domain behavior recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 1552–1560. Cited by: Introduction, Wi-Fi Sensing. Y. Zhang, L. Wang, H. Chen, A. Tian, S. Zhou, and Y. Guo (2022) IF-convtransformer: a framework for human activity recognition using imu fusion and convtransformer. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 6 (2), p. 1–26. Cited by: Introduction, Human Activity Recognition. Y. Zhang, X. Wang, Y. Wang, and H. Chen (2020) Human activity recognition across scenes and categories based on csi. IEEE Transactions on Mobile Computing 21 (7), p. 2411–2420. Cited by: Introduction. M. Zhao, S. Yue, D. Katabi, T. S. Jaakkola, and M. T. Bianchi (2017) Learning sleep stages from radio signals: a conditional adversarial architecture. In International conference on machine learning, p. 4100–4109. Cited by: Signal Module. Y. Zheng, Y. Zhang, K. Qian, G. Zhang, Y. Liu, C. Wu, and Z. Yang (2019) Zero-effort cross-domain gesture recognition with wi-fi. In Proceedings of the 17th annual international conference on mobile systems, applications, and services, p. 313–325. Cited by: Introduction, Signal Module, Datasets.