Paper deep dive
Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding
Jeongwan Shin, Jaehyeon Kim, Donguk Ko, Jaeho Choi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/17/2026, 5:17:03 AM
Summary
This paper introduces mmWave-QA, the first benchmark for language-conditioned mmWave human perception, addressing the gap in integrating Large Language Models (LLMs) with millimeter-wave radar data. The authors propose a minimal textualization interface that serializes mmWave point clouds into natural language, enabling off-the-shelf LLMs to perform zero-shot reasoning on human actions without dataset-specific training. The benchmark aggregates heterogeneous public datasets (mmBody, MM-Fi, mRI), harmonizes them via calibration-aware preprocessing, and provides five QA tasks (Action Recognition, Spatial Movement Verification, Temporal Action Sequencing, Action Cardinality Estimation, Limb Motion Detection) across six environmental scenarios. Experiments demonstrate that LLMs exhibit robust zero-shot reasoning capabilities and that mmWave sensing is more robust than RGB under visual degradation.
Entities (14)
Relation Signals (11)
mmWave-QA â evaluates â Large Language Models
confidence 95% ¡ We further evaluate and analyze LLMs on our mmWave-QA
mmWave-QA â includestask â Temporal Action Sequencing
confidence 95% ¡ (III) Temporal Action Sequencing (ActOrder)
mmWave-QA â includestask â Action Recognition
confidence 95% ¡ mmWave-QA comprises five core categories... (I) Action Recognition (ActRec)
mmWave-QA â includestask â Spatial Movement Verification
confidence 95% ¡ (II) Spatial Movement Verification (TrajCheck)
mmWave-QA â includestask â Action Cardinality Estimation
confidence 95% ¡ (IV) Action Cardinality Estimation (ActNum)
mmWave-QA â includestask â Limb Motion Detection
confidence 95% ¡ (V) Limb Motion Detection (LimbFocus)
mmWave-QA â uses â mRI
confidence 95% ¡ we integrate three mmWave sensing datasets for human perception: mmBody [7], MM-Fi [61], and mRI [5].
mmWave-QA â uses â MM-Fi
confidence 95% ¡ we integrate three mmWave sensing datasets for human perception: mmBody [7], MM-Fi [61], and mRI [5].
â â
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) have shown remarkable reasoning and generative capabilities, motivating their use as universal reasoning engines for perception. While modern approaches such as vision-language models (VLMs) have attempted to incorporate reasoning capabilities into visual sensing, the integration of LLMs with the millimeter-wave (mmWave) modality-despite its unique advantages under low light and occlusion-remains largely unexplored. The principal bottlenecks stem from the scarcity of radar language pairs, severe cross-dataset heterogeneity, and the absence of a foundational mmWave encoder. We address this gap through a minimal textualization interface that serializes each mmWave point cloud into concise natural language, allowing off-the-shelf LLMs to operate in a question answering (QA) setting. Building on this, we present mmWave-QA, the first benchmark for language-conditioned mmWave human perception. mmWave-QA aggregates heterogeneous public mmWave datasets and harmonizes them via calibration-aware preprocessing and global taxonomy alignment, while providing natural language QA. Spanning six scenarios and five QA tasks, the benchmark enables standardized evaluation across diverse mmWave hardware and experimental conditions, establishing a foundation for scalable research on mmWave-LLM integration. We further evaluate and analyze LLMs on our mmWave-QA, highlighting their zero-shot reasoning potential for radar perception, as well as their robustness under visual degradation.
Tags
Links
- Source: https://arxiv.org/abs/2608.14179v1
- Canonical: https://arxiv.org/abs/2608.14179v1
Trouble viewing inline? Open PDF directly â
Full Text
63,360 characters extracted from source content.
Expand or collapse full text
Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding Jeongwan Shin 1,2 Jaehyeon Kim 1 Donguk Ko 1 Jaeho Choi â ,1,2 1 DGIST 2 KAIST InnoCORE LLM Model trained on Data A Model trained on Data B LLM Model trained on Data C Activity Instruction: Given a sequence of mmWave radar point cloud frames, identify the human action being performed. Each frame contains points with coordinates [x, y, z, doppler, intensity]. mmWave Point Cloud: [3.169, -0.2, 0.4, 12.3, 0.604], [3.12, 0.0, 0.494, 12.6, - 0.604], ... Question: Analyzing the point cloud and movement patterns, what motion is being captured? Rationale: The point cloud data shows a consistent pattern of points moving in a vertical direction with minimal horizontal displacement, which is characteristic of a squat motion. The doppler values indicate a repetitive up and down movement, typical of squatting, as opposed to lateral or twisting movements. Answer: Squat Requires dataset-specific training due to mmWave heterogeneity Transform (c) Dataset Composition Heterogeneity Scenario Normal Rain Smoke Furnished Occlusion Dark RGB Video (for reference) 1) ActRec: Action Recognition Q: Which action is represented by this point cloud? A: Basketball Dribble 2) TrajCheck: Spatial Movement Verification Q: Has the personâs position centroid changed location? A: No 4) ActNum: Action Cardinality Estimation Q: How many distinct actions appear in this segment? A: 1 3) ActOrder: Temporal Action Sequencing Q: What is the first action in this segment? A: Basketball Dribble 5) LimbFocus: Primary Limb Motion Detection Q: Which single body part is primarily moving in this action? A: left arm Data A Data B Data C (a) Current mmWave Human Perception (b) LLM-integrated mmWave Human Perception (Training-free) Works with any point cloud input mmWave Point Cloud Activity Activity mmWave QA Trainable Frozen Figure 1. Top: Conventional mmWave perception frameworks relying on per-dataset retraining vs. our LLM-based perception that interprets structured radar point cloud inputs in a zero-shot manner. Bottom: Composition of mmWave-QA, encompassing heterogeneous mmWave devices, six real-world scenarios, and five QA types for comprehensive multimodal reasoning assessment. Abstract Large language models (LLMs) have shown remarkable reasoning and generative capabilities, motivating their use as universal reasoning engines for perception. While mod- ern approaches such as vision-language models (VLMs) have attempted to incorporate reasoning capabilities into visual sensing, the integration of LLMs with the millimeter- wave (mmWave) modalityâdespite its unique advantages under low light and occlusionâremains largely unexplored. The principal bottlenecks stem from the scarcity of radar- language pairs, severe cross-dataset heterogeneity, and the absence of a foundational mmWave encoder. We address this gap through a minimal textualization interface that se- rializes each mmWave point cloud into concise natural lan- guage, allowing off-the-shelf LLMs to operate in a question- answering (QA) setting.Building on this, we present mmWave-QA, the first benchmark for language-conditioned mmWave human perception. mmWave-QA aggregates het- erogeneous public mmWave datasets and harmonizes them via calibration-aware preprocessing and global taxonomy alignment, while providing natural language QA. Spanning six scenarios and five QA tasks, the benchmark enables standardized evaluation across diverse mmWave hardware and experimental conditions, establishing a foundation for scalable research on mmWave-LLM integration. We further evaluate and analyze LLMs on our mmWave-QA, highlight- ing their zero-shot reasoning potential for radar perception, as well as their robustness under visual degradation. 1. Introduction Millimeter-wave (mmWave) sensing has emerged as a strong alternative for human perception due to its ability to measure human targets via electromagnetic reflections rather than appearance cues [5, 25, 28, 68]. By captur- ing human kinematics across range, Doppler, and angle â Corresponding author. arXiv:2608.14179v1 [cs.AI] 14 Aug 2026 domains, modern mmWave radar systems remain effective even under challenging conditions that routinely degrade camerasâlow/mixed lighting, partial/strong occlusions of the human body, and long standoff distances [11, 57]. Fur- thermore, the mmWave signals do not reveal facial identity or fine appearance, naturally offering a privacy advantage that is essential for continuous and ubiquitous monitoring in everyday home or office environments [2, 62, 71]. Meanwhile, recent LLMs [6, 22, 36, 37, 49] trained on hyper-scale corpora [24, 39, 48] have demonstrated power- ful contextual understanding [42, 69, 72] and broad gener- alization across diverse tasks [12, 30, 31, 38, 43, 60, 64, 65]. When further integrated with visual encoders [26, 40, 63], these models evolve into VLMs [1, 9, 10], where perception is conditioned on both linguistic and visual cues, allowing language-guided scene understanding with minimal super- vision [27, 54]. This context-conditioned paradigm offers two notable benefits for human perception: 1) a natural- language interface that enables humans to query and steer models through QA-based interactions and 2) the ability to leverage extensive knowledge of LLMs to interpret unseen behaviors in a zero-/few-shot manner [6, 51, 52]. However, integrating LLMs with the mmWave modal- ity presents several non-trivial challenges.First, an mmWaveâlanguage ecosystem remains underdeveloped: unlike common modalities (e.g., RGB or audio) with web-scale modality-language pairings, existing mmWave datasets are small, lab-specific, and hard to annotate, lead- ing to the absence of public mmWave datasets with radar- caption pairings. Second, this data scarcity is further com- pounded by heterogeneity: mmWave observations vary not only due to environmental clutter/multipath but also by hardware-specific factors (e.g., carrier frequency, waveform design, # antennas), which in turn induce severe distribu- tional shifts across datasets [4, 8, 29, 59]. Third, the field lacks standardized tokenizers or foundational encoders to bridge mmWave tensors and language backbones. Conse- quently, current mmWave systems for human understanding [45, 58, 62, 66, 67] remain scenario-tuned: models are opti- mized for each dataset and require retraining whenever the scenario (e.g., hardware, surroundings, or subject distribu- tion) changes, as examplified in Fig. 1 (a). Accordingly, our work begins with a concrete question: can we still leverage the reasoning power of LLMs for mmWave-based human understanding despite the aforementioned challenges? To investigate its feasibility, we render each 5D mmWave point cloud sample into a coordinate-formatted text rep- resentation (i.e., [x, y, z, Doppler, intensity]) and prompt LLMs to answer the corresponding human motion question (Fig. 1 (b)). Surprisingly, without any mmWave-specific tuning, LLMs exhibit reasonable zero-shot reasoning, re- flecting their potential to map textualized radar data to high- level behaviors. Building on this observation, we intro- duce mmWave-QA, the first benchmark designed to explore mmWave-LLM integration. The benchmark is constructed through a curation pipeline that integrates heterogeneous datasets, reformulates them into mmWave-based QA tasks aligned with the benchmark taxonomy, and generates QA pairs through an LLMâhuman collaboration process. Fi- nally, quality review and refinement are conducted to ensure the quality and balance of the resulting benchmark. After constructing the benchmark, we evaluate various prompt- ing strategies to comprehensively assess and analyze how LLMs interpret human motion from mmWave data. As shown in Fig. 1 (c), our mmWave-QA provides a unified testbed spanning six diverse scenarios with five complementary QA tasks that go beyond simple spatial action queries to probe spatial, temporal, and joint-level granular aspects of human motion.Furthermore, our benchmark encompasses different mmWave hardware de- vices and experimental conditions, thereby enabling stan- dardized evaluation in analyzing mmWave-LLM integra- tion under diverse co-factors: subject placements, envi- ronment/scenario shifts, and even hardware/waveform de- pendencies. Through extensive experiments on mmWave- QA, we investigate the behavior of off-the-shelf LLMs [19, 36] under different prompting strategies, revealing their ability to interpret mmWave-based human motion without task-specific fine-tuning. We further evaluate the poten- tial of mmWaveâLLM integration for zero-shot human un- derstanding, demonstrating that the mmWave modality re- mains more robust than RGB under visual degradation. In summary, our contributions are threefold: ⢠We present mmWave-QA, a unified benchmark that har- monizes heterogeneous mmWave human-motion datasets and recasts them as natural-language QA tasks spanning multiple aspects of action understanding. ⢠We show that integration with simple off-the-shelf LLMs can even lead to strong zero-shot reasoning on textualized mmWave radar inputs. ⢠We provide empirical evidence that the mmWave modal- ity is more robust than RGB under visual degradation. 2. Related Work mmWave Human Sensing Datasets. As shown in Ta- ble 1, several mmWave human sensing datasets have been introduced in recent years.Within action recognition, datasets have evolved from RadHAR [47] and MVDoppler [21], which capture basic body motions to XRF55 [50], which covers composite actions and diverse environments. Extending beyond action recognition, pose estimation [3, 5, 61] focuses on joint-level understanding, spanning from re- habilitation to natural daily motions, and further enriches datasets by expanding the number of subjects and envi- ronments. With recent advances in radar hardware, m- Body [7], HIBER [53], and MMVR [41] employ high- Benchmarks#Act #Subj #Env mmWaveSensorDataset ScenarioModalityAnnotations HWRes.SizeNorm. Furn. Rain Smoke Dark Occl.RVTAct Pose QA RadHAR [47]521ALowMâ-----â--â-- MVDoppler [21]4131BLow-â-----â--â-- XRF55 [50]55394CLowXLâ-----â-â-- MARS [3]1041ALowSâ-----â---â- mmBody [7]100206DHighMâ--â- HIBER [53]41010EHighMâ----â--â- MMVR [41]-256EHighMâ----â--â- mRI [5]12201ALowMâ-----â-â- M-Fi [61]27404CLowMâ-----â-â- Ours139(86) 8011 A, C, D Low, HighLâ Table 1. Comparison of mmWave-based human sensing benchmarks. Each column denotes the following â number of distinct human actions (#Act), number of subjects (#Subj), number of environmental settings (#Env), and radar hardware type (mmWave HW): A) TI IWR1443, B) TI AWR1843, C) TI IWR6843, D) Phoenix-type custom board, E) TI AWR2243. Scenario cover diverse sensing conditions including normal, furnished, rain, smoke, dark and occlusion environments. Modality specifies the available input types: radar, vision, and text. Annotations indicate the provided supervision signals: human action labels (Act), 3D pose annotations (Pose), and natural-language QA pairs (QA). Numbers in parentheses (e.g., (86)) represent newly categorized action classes. Datasets are categorized by the number of frames: S (<100k), M (100â400k), L (400â700k), and XL (>700k), as detailed in the supplementary material. resolution radar sensors, providing denser and more de- tailed observations. Despite these advances, the growing heterogeneity in data sources and sensing setups continues to hinder the development of universally adaptable models. mmWave-based Human Understanding. Considering the diversity of human sensing data, mmWave-based hu- man understanding has evolved into several distinct model- ing paradigms. MARS [3] and mmPose [44] employ CNN- based architectures that learn spatial representations from radar reflections, while RadHAR [47] and DP-CBL [20] further incorporate LSTM to capture temporal dependen- cies across consecutive frames. Inspired by the success of transformers in achieving powerful context understanding across NLP [14] and vision [16, 32, 35] domains, HMR- SCL [15], MVDoppler-Pose [11], and mmPoint [46] adopt attention mechanisms, underscoring how mmWave encod- ing paradigms largely track broader architectural trends in other domains.Despite growing applications of LLMs for contextual reasoning, their integration with mmWave sensing remains limited; mmWave-QA addresses this gap through a benchmark for systematic exploration. Sensing modalities with LLMs. By incorporating addi- tional modalities, LLMs can be extended into multimodal, allowing them to interpret multimodal inputs such as im- ages [13, 26, 30, 31], videos [17, 64, 65], audio [18, 33], and spatial measurements [55, 56]. Building on this progress, only a limited number of studies have explored mmWave sensing modalities. Among recent attempts, RadarLLM [23] is the framework that leverages LLMs for human mo- tion understanding using millimeter-wave radar, training on synthetically generated IF signals to bridge radar sensing and language understanding. Similarly, HoloLLM [70] in- troduces a multisensory foundation model that integrates sensing modalities with LLMs for language-grounded hu- man perception and reasoning. However, existing mod- els face inherent limitations: those relying on synthetic radar data struggle to generalize to real-world clutter, while those processing raw mmWave signals depend on modality- specific encoders.To overcome these limitations, we demonstrate that existing LLMs can interpret textualized mmWave point clouds in a zero-shot manner and establish a benchmark that quantifies their reasoning capability, laying the groundwork for future radarâlanguage integration. 3. mmWave-QA Benchmark 3.1. Data Construction As a step toward building the first mmWaveâlanguage benchmark, we integrate three mmWave sensing datasets for human perception: mmBody [7], M-Fi [61], and mRI [5]. This integration provides a comprehensive and diverse foundation covering multiple radar devices, environments, and subjects. Merging the datasets yielded 139 action la- bels, which were refined into 86 distinct categories after removing redundancies and ambiguities.Table 1 compares mmWave-QA with prior human sensing benchmarks, high- lighting its diversity and QA-driven design, while Fig. 2 il- lustrates the overall dataset construction pipeline. mmWave Data Collection and Textualization. We be- gin by addressing the insufficiencies of prior mmWave datasets, which contained dense and cluttered point clouds with limited visual information. RGB recordings are con- verted into visual references for annotators, presenting clean scenes as videos and degraded or occluded ones as skeleton videos to aid human understanding. To standard- ize data representation across heterogeneous point-cloud formats, we apply coordinate-based quantization to cluster I. mmWave Data Collection & TextualizationII. Action Taxonomy & AnnotationIII. QA Generation V. LLM Inference IV. Quality Review Data source Annotators LLM Images Transform Video 3D SkeletonVideo "frame_01": "vertices": [[0.076, 3.060, 0.2772], ...], "joints": [[0.076, 3.060, 0.2772], ...], "frame_02": "vertices": [[0.082, 3.058, 0.2801], ...], "joints": [[0.082, 3.058, 0.2801], ...],... Local Quantization GroundTruthMesh Human-interpretable Transformation Point-wise Compression for LLM Input Step1) Frame Select Step2) Action Annotation Walk Annotation Process Action Taxonomy Construction Paraphrased Question 1.What movement is shown in this radar data? 2.Analyzing the point cloud and movement patterns, what motion is being captured? 3.Which action is represented by this point cloud? 4.Based on the spatiotemporal distribution of radar points, what type of body movement or gesture can be inferred from this sequence? Writing Question âQuestionâ: âLooking at this radar point cloud data, what action is the person performing?â, âOptionâ: [âA. High Knee Runâ, âB. Mark Timeâ, âC. T- poseâ, âD. Walkâ], âAnswerâ: âD. Walkâ ActRec LimbFocus ActNum ActOrder TrajCheck [SideKick, Bowing, InsideKick, Walk] â [Side Kick, Bowing, Hug, Walk] Data source Data1 Step1) Annotator Review Refinement Process Step2) Option Refinement Step3) Data Rebalancing Data2 LLM Zero-shot Instruction Given the radar point cloud data and the question, Select exactly one option Chain-of-Thought Instruction Given the radar point cloud data and the question, provide a brief reasoning for your answer. OR Zero-shot Answer: D. Walk Chain-of-Thought Rationale: The point cloud data shows a consistent forward and backward movement pattern along the y-axis with minimal lateral movement, which is characteristic of walking. The doppler values also indicate a continuous motion, supporting the walking action. Answer: D. Walk OR âPoint Cloudâ: â[3.17, -0.23, 0.26, 12.33, 0.61], [3.12, 0.01, 0.49, 12.61, -0.24]â âQuestionâ: âLooking at this radar point cloud data, what action is the person performing?â âOptionâ: [âA. High Knee Runâ, âB. Mark Timeâ, âC. T-poseâ, âD. Walkâ] Data source Update ActRec LimbFocus ActNum ActOrder TrajCheck Figure 2. Pipeline of mmWave-QA construction and evaluation: I) Original mmWave Collection & Textualizationâradar data are preprocessed into human-interpretable and LLM-ready quantized formats; I) Action Taxonomy & Annotationâactions are categorized into hierarchical domains; I) QA Generationâdiverse QA pairs are generated through human-in-the-loop LLM; IV) Quality Reviewâ annotators refine and balance QA; and V) LLM Inferenceâthe benchmark enables evaluation of language models across diverse prompting. points into a form compatible with LLM token limits. Action Taxonomy and Annotation. As some of the orig- inal datasets do not provide action-wise segmentation, an- notators segment the videos into continuous action inter- vals, assigning the corresponding action label. Based on the annotated action segments, annotators construct a hier- archical taxonomy of question types, as illustrated on the left side of Fig. 3. Specifically, mmWave-QA comprises five core categories that capture different aspects of hu- man motion understanding, enabling multifaceted evalua- tion on how well an mmWave-conditioned LLM can per- ceive and align physical information with linguistic repre- sentations. (I) Action Recognition (ActRec) for recogniz- ing the performed action, (I) Spatial Movement Verifica- tion (TrajCheck) for determining whether the subjectâs cen- troid has changed position, (I) Temporal Action Sequenc- ing (ActOrder) for reasoning about the temporal relation- ships between consecutive actions, (IV) Action Cardinality Estimation (ActNum) for estimating the number of distinct actions, and (V) Limb Motion Detection (LimbFocus) for identifying the primary moving body part. Each category is divided into subtypes capturing motion attributes. QuestionâAnswer Generation. Following the annota- tion stage, annotators generate natural language ques- tionâanswer pairs based on the labeled actions.Each question is designed to evaluate different aspects of point cloud understanding in accordance with the taxonomy es- tablished in the previous step. Subsequently, annotators construct multiple-choice QA pairs by selecting answer op- tions from the predefined action taxonomy. To enhance lin- guistic diversity and ensure natural phrasing, we further em- ploy LLMs leveraging its paraphrastic generation capability [34, 37] to perform paraphrasing of the original questions while preserving their semantic intent. Quality Review. Existing human motion datasets often suffer from biased action distributions and inconsistent la- beling of visually similar actions. To ensure annotation reli- ability and dataset balance, we conduct a three-step quality review process. Three annotators independently review all generated QA pairs, and any pair that appears ambiguous or contains unclear action semantics is removed from the dataset if even one annotator marks it as invalid. During this step, we refine the answer options by replacing overly similar or redundant candidates with more distinctive ones to prevent confusion during QA. Finally, we adjust the dis- tribution of QA pairs to ensure a balanced coverage across different action domains and question types defined in the taxonomy. After completing refinement procedures, we fi- nalize the mmWave-QA benchmark and perform evaluation using LLM inference under diverse prompting strategies. 3.2. Dataset Statistics The top-right chart in Fig. 3 presents the distribution of question types across domains in mmWave-QA. Among all categories, ActRec and TrajCheck questions occur most fre- quently, whereas ActOrder, ActNum, and LimbFocus ap- pear less often yet enrich the benchmark by addressing se- 330 170 140 220 55 5050 55 112 80 60 40 125 160 125125125 0 150 300 450 Number of ďźą uestions Question Typ e Distrib uti onAcross Categories ActRecLimbFocusActNum ActOrderTrajCheck 0 50 100 150 200 NormalFurishedRainSmokeDarkOcclusion Distribu tion of Scene Con ditions Across Question Types ActRecActNum ActOrderLimbFocus Number of Questions Figure 3. Statistics of mmWave-QA dataset. Left: Hierarchical taxonomy of question domains: ActRec, TrajCheck, ActOrder, ActNum, and LimbFocus. Top-right: Distribution of subcategories within each top-level action domain, illustrating diverse coverage across human- motion questions. Bottom-right: Distribution of taxonomy actions across six scene conditions: normal, furnished, rain, smoke, dark, and occlusion. Note that the TrajCheck category is omitted in the bottom-right plot due to limited samples under specific conditions. quential reasoning, quantitative estimation, and localized motion analysis. Together, these categories provide a com- prehensive, balanced, and complementary set of tasks for evaluating multimodal reasoning capabilities of language models. The bottom-right chart presents the distribution of question types across six scene conditions: normal, fur- nished, rain, smoke, dark, and occlusion. Note that the Tra- jCheck category is excluded due to an insufficient number of samples in certain conditions. Each condition contains roughly 100â200 questions, with Normal scenes having the largest share due to their higher data availability and role as the baseline for comparison. This diversity enables con- sistent evaluation of model robustness and generalization across heterogeneous environmental settings. 4. LLM-driven mmWave-QA 4.1. Preliminary We first formalize the inference process of an LLM on the proposed mmWave-QA benchmark. Given a natural- language question Q and a mmWave point cloud sequence P =p i l i=1 consisting of l frames, where each frame p i exhibits varying point density depending on the sensing de- vice and environmental conditions, we formally define the inference process of an LLMM as: A =M(I, Q, P),(1) where I denotes the instruction prompt defining the infer- ence setting, and A represents the sequence of output to- kens generated byM. This formulation serves as the foun- dation for subsequent prompting strategies, including zero- shot, few-shot, and chain-of-thought inference. 4.2. Inference Strategies We investigate three prompting inference strategies to eval- uate the reasoning performance of LLMs on the proposed mmWave-QA benchmark: zero-shot, few-shot, and chain- of-thought (CoT) inference. Each strategy conditions the model with different inputs, leading to distinct reasoning processes and answer generation behaviors. Zero-shot. In the zero-shot setting, the model generates an answer directly from the given questionâpoint cloud pair (Q, P) without additional examples or reasoning context: A zero =M zero (I, Q, P),(2) where A zero denotes the generated answer, andM zero rep- resents the LLM conditioned on zero-shot prompting. Few-shot. Given k few-shot in-context examples consist- ing of questionâpoint cloudâanswer triplets (Q i , P i , A i ) k i=1 as contextual references for inference: A few =M few (I,(Q i , P i , A i ) k i=1 , Q, P),(3) enabling the LLM to utilize these examples to infer answers for the target input without updating model parameters. Chain-of-Thought (CoT). Under the CoT prompting set- ting, the model follows a rationale-generation instruction to produce a continuous text sequence that first presents inter- mediate reasoning steps before providing the final answer. Formally, the CoT inference can be expressed as: R CoT =M CoT (I rationale , Q, P),(4) ModelPrompt Type Action CategorymmWave HW Total Upper-body Lower-bodyTorsoFull-body Phoenix IWR6843 IWR1443 Gemini2.5-flash Zero-shot32.1223.5325.0029.0928.0629.1430.7728.49 Few-shot45.7624.1225.7129.5534.0334.2933.8534.07 CoT37.2725.2928.5730.9132.1032.0029.2331.86 Few-shot CoT46.9730.5927.1429.5536.1336.0035.3836.05 Gemini2.5-Pro Zero-shot32.7324.7125.7128.1827.9029.7135.3828.84 Few-shot46.3626.4727.1429.5535.1633.7136.9235.00 CoT41.5224.7127.8629.0932.9031.4335.3832.79 Few-shot CoT47.8828.2428.5730.9136.4536.5736.9236.51 GPT-4o Zero-shot29.0920.5924.2928.1826.1327.4326.1526.40 Few-shot45.4526.4727.8630.9136.7731.4329.2335.12 CoT34.8522.3525.0029.0929.1930.2927.6929.30 Few-shot CoT46.9728.8229.2933.1838.3933.7132.3136.98 GPT-5 Zero-shot32.1221.7623.5728.1826.4527.4326.1526.63 Few-shot45.7628.8230.0031.3637.4233.7130.7736.16 CoT35.4523.5327.1429.5529.8432.0029.2330.23 Few-shot CoT48.4831.7631.4333.6438.7137.1438.4638.37 Table 2. Evaluation of various LLMs across different prompt settings and action categories on multiple mmWave HW. All values denote accuracy (%), with the best results in bold and the second-best in italic where I rationale represents the rationale-generation instruc- tion, and R CoT denotes the LLM-generated sequence [rationale; A CoT ]. The CoT process enables the model to generate a rationale that explains inter-frame relationships within point clouds, thereby facilitating the derivation of the final answer. Lastly, we additionally introduce a hybrid in- ference strategy that integrates few-shot examples with CoT reasoning. All prompt examples and detailed cases are pre- sented in the supplementary material. 5. Experiments The proposed mmWave-QA benchmark evaluates commer- cial LLMs [19, 36] reasoning on radar-based human activity data collected under heterogeneous devices, environments, and scene conditions. Quantitative results under different prompting strategies are presented first, along with analy- ses comparing mmWave- and RGB-based QA and examin- ing the effect of frame count on performance. Finally, qual- itative examples further illustrate how LLMs reason over mmWave point cloud inputs. Detailed evaluation settings are provided in the supplementary material. 5.1. Quantitative Results Q: Can LLMs understand radar point clouds without task-specific training? To explore this question, we eval- uate various prompting strategies, where the results in Ta- ble 2 show consistent performance gains-from zero-shot to few-shot, CoT, and finally few-shot CoT prompting- across all models. This behavior highlights the comple- mentary benefits of in-context learning and rationale-based reasoning, demonstrating their potential to enhance under- standing of radar-derived motion representations. GPT-5 achieved the best overall performance in the few-shot CoT setting, improving overall accuracy by 1.39 and 1.86 per- centage points over GPT-4o and Gemini 2.5-Pro, respec- tively. Among action categories, arm motions yield the best results, followed by full-body movements, while leg and torso actions remain challenging due to subtler motion cues and frequent occlusions in radar reflections. Across mmWave hardware, Phoenix shows the highest accuracy owing to its high-quality point clouds enabled by an ad- vanced radar antenna configuration, whereas IWR6843 and IWR1443 show performance degradation caused by sensor noise and environmental variability. These results provide empirical evidence that LLMs can reason robustly over het- erogeneous radar datasets, suggesting their strong general- ization capability even under diverse sensing conditions. Q: How effectively do LLMs perform multi-aspect rea- soning across human-motion QA tasks? We extend our evaluation to multi-aspect reasoning QA tasks, as shown in Table 3. For TrajCheck in the zero-shot setting, both GPT-4o and Gemini 2.5-flash predominantly classify radar samples as âmoveâ, achieving 92.0% and 87.2% accuracy for moving cases but only 4.0% and 3.2% for static ones. When few-shot examples are provided, the static accuracy increases to 32.0% and 27.2%, respectively, suggesting that contextual information helps both models better distinguish between moving and static cases. For ActNum, even un- der the few-shot CoT setting, models achieve stable accu- racy for 2â3 actions at 43â58.75%, but sharply decline to 5-22.5% as the action count increases, indicating difficulty in sustaining consistent reasoning over extended temporal sequences in point clouds. In ActOrder, predictions for the present are most accurate, whereas reasoning about previ- ous and next actions remains challenging, likely because the excessive number of input tokens from dense point clouds hinders the modelâs ability to capture clear temporal rela- tionships. Finally, LimbFocus demonstrates strong recog- ModelPrompt Type ActNumActOrderLimbFocusTrajCheck 1234PrevPresent NextSingle Arm/Leg Multiple Static Move Gemini2.5-flash Zero-shot9.82 47.50 5.007.5024.80 26.88 27.2021.6724.8028.003.20 87.20 Few-shot38.39 52.50 36.67 17.508.0030.00 32.0030.0031.2038.0027.20 43.20 CoT11.60 50.00 30.00 15.00 26.40 28.75 28.8023.3325.6032.007.20 88.00 Few-shot-CoT41.07 55.00 45.00 22.50 30.40 30.63 35.2035.0033.6042.0062.40 70.40 GPT-4o Zero-shot11.61 52.50 5.005.0024.80 26.88 27.2021.6724.8028.004.00 92.00 Few-shot40.18 56.25 35.00 17.50 28.80 28.75 31.2026.6728.8038.0032.00 44.00 CoT30.36 36.25 25.00 12.50 26.40 27.50 29.6025.0026.4034.0010.40 88.00 Few-shot-CoT42.86 58.75 43.33 22.50 31.20 30.63 36.0038.3335.2042.0068.00 72.80 Table 3. Evaluation of various LLMs across different prompt settings, including ActNum, ActOrder, LimbFocus, and TrajCheck QA. All values denote accuracy (%), with the best results in bold and the second-best in italic. 10 20 30 40 50 60 70 Nomal Furnished Rain Smoke Dark Occulusion Accuracy ActRec RADARRGB 22 26 30 34 38 612163264 Accuracy Number of frames ActRec Zero-shotCoTFew-shot Few-shot CoT 32 42 52 62 72 612163264 Accuracy Number of frames TrajCheck 22 25 28 31 34 612163264 Accuracy Number of frames ActOrder 18 22 26 30 34 38 612163264 Accuracy Number of frames LimbFocus 17 23 29 35 41 612163264 Accuracy Number of frames ActNum 10 20 30 40 50 60 70 Nomal Furnished Rain Smoke Dark Occulusion Accuracy ActNum RADARRGB 10 20 30 40 50 60 70 Nomal Furnished Rain Smoke Dark Occulusion Accuracy ActOrder RADARRGB 10 20 30 40 50 60 70 Nomal Furnished Rain Smoke Dark Occulusion Accuracy LimbFocus RADARRGB Figure 4. Top: Comparison of mmWave- and RGB-based QA performance across six scenarios: normal, furnished, rain, smoke, dark, and occlusion. Bottom: Performance variation with respect to the number of frames used in each QA task. nition for single or limb-specific motions but lower accu- racy in multi-limb scenarios, likely due to multipath inter- ference and ghost reflections that distort body-part separa- bility in complex radar scenes. These findings promote fu- ture studies on understanding the structural composition of radar point clouds and selecting temporally salient frames to facilitate more effective temporal reasoning. 5.2. Ablation Studies Comparison with RGB Modality. We compare the per- formance of mmWave-based QA with its RGB counterpart across six scenarios, as shown in the upper part of Fig. 4. For the RGB modality, we evaluate using GPT-4o [36], which accepts RGB inputs. In normal and furnished en- vironments, RGB outperforms radar, as vision models have been trained on far larger and more diverse datasets, where richer visual cues facilitate easier reasoning for VLMs, whereas multipath ghost reflections from furniture degrade the quality of radar point clouds. However, in rain and smoke conditions, both modalities achieve comparable per- formance, demonstrating radarâs robustness against visibil- ity degradation. Notably, in dark and occlusion scenarios, the radar modality surpasses RGB, highlighting its unique advantage in perception tasks where visual information is heavily corrupted or unavailable. In contrast to other tasks, ActNum shows consistently better performance with radar, as motion counting relies on temporal Doppler cues that are hard for RGB models to capture with limited frames. Frame-wise Performance Analysis. As shown in the lower part of Fig. 4, the performance varies with the num- ber of frames used in each QA task, where we gradually in- crease the frame count from 6 to 64 during evaluation. The input frame length plays a critical role in interpreting hu- man activity from point cloud sequences, as it determines both temporal coverage and information density. Across most tasks, performance peaks around 16 frames, as too few frames limit temporal context, while excessive frames in- troduce redundant or noisy information that hinders reason- ing. This finding suggests that radar-based motion reason- ing requires an optimal temporal balanceâsegments should be long enough to capture key motion patterns but short enough to avoid redundant or noisy points. 5.3. Qualitative Analysis To qualitatively examine reasoning behavior, Fig. 5 illus- trates the rationales and responses produced by the VLM and LLM given video and point cloud inputs. When RGB provides clear visual evidence, its interpretations are largely consistent with radar, producing similar answers and rea- Question: What movement is shown in this radar data? Option: A. Front Lunge, B. Unilateral Upper-Lower Limb Extension, C. Walk, D. Both Upper-Limb Lateral Extension âRationaleâ: âAcross consecutive frames, the subjectâs entire body moves steadily forward along a horizontal direction, with consistent leg alternation and synchronized arm swing typical of natural gait. There is no stationary stance, lunge posture, or isolated limb extension; instead, the motion indicates full-body displacement across space.â âAnswerâ: âC. Walkâ âRationaleâ: âThe point cloud data shows a consistent pattern of movement along the z-axis with varying doppler values, indicating forward and backward motion typical of walking. The points are distributed in a manner that suggests a repetitive, rhythmic motion consistent with walking, rather than isolated limb movements or stationary exercises.â âAnswerâ: âC. Walkâ ActRec Question: What is the single moving body part in this pose? Options: A. arm,B. leg,C. none âRationaleâ:The video frames show a dark scene where the subjectâs body outline is visible but lacks clear detail.Across consecutive frames, no distinct or repetitive motion of the arms or legs can be confidently identified.Because no specific body part exhibits a noticeable movement pattern, itâs difficult to determine any primary motion. âAnswerâ: âC. noneâ âRationaleâ: "The point cloud shows concentrated high-Doppler activity in the upper-body region while lower-body points remain stable.This localized motion pattern indicates that the arm is the primary moving part.â âAnswerâ: âA. armâ LimbFocus Question: Which single body part is primarily moving in this action? Options: A. arm,B. leg,C. none âRationaleâ:The subject remains mostly upright with minimal lower-body movement, while one arm lifts gradually from the side to an overhead position.The consistent elevation of the arm and stationary posture indicate deliberate upper-limb motion.Overall, the motion pattern clearly identifies the arm as the primary moving body part. âAnswerâ: âA. armâ âRationaleâ: "High Doppler responses are mainly observed around the upper-body area, whereas the lower body exhibits minimal movement. This distribution suggests that the arm is responsible for the primary motion in the scene.â âAnswerâ: âA. armâ Question: Analyzing the point cloud and movement patterns, what motion is being captured? Option: A. Scoop Tap Motion, B. Walk, C. One Hand Raise, D. High Knee Run âRationaleâ: âThe frames mainly display alternating white and purple patterns, with almost no visible human silhouette.Because of the low visibility, the specific limb motion is hard to discern, yet the subtle periodic changes suggest a gentle, rhythmic movement consistent with walking.â âAnswerâ: âWalkâ âRationaleâ: âThe radar point cloud data shows a consistent pattern of points moving vertically with minimal horizontal displacement, indicating a vertical motion. The points are concentrated around a central axis with slight variations in height, suggesting an upward motion of a single hand. This pattern is consistent with a 'One Hand Raise' action, where the hand moves upwards without significant lateral movement.â âAnswerâ: âOne Hand Raiseâ ActRec VLM Time LLM RGB vs Radar RGB vs Radar Time LimbFocus Time RGB vs Radar VLM LLM Time RGB vs Radar LLM VLM VLM LLM Figure 5. Qualitative analysis of mmWave-QA. Each example from the ActRec and LimbFocus tasks visualizes temporal changes in video and radar point clouds along with corresponding rationales generated by the VLM and the LLM. When both modalities provide clear visual evidence, their reasoning remains consistent; however, under degraded conditions (e.g., occlusion or darkness), the VLM often hallucinates or misinterpret corrupted frames, whereas radar-based reasoning remains stable and physically grounded. soning patterns. However, under visually degraded condi- tions such as darkness or occlusion, the RGB-based model often exhibits hallucinated reasoning or over-explains cor- rupted regions, leading to incorrect predictions (e.g., âal- most no visible humanâ). In contrast, the radar modality facilitates Doppler-based motion reasoning, proving effec- tive when visual information is degraded or unavailable. In ActRec, the LLM correctly infers walking motion by track- ing coordinate displacement and Doppler variation across frames (e.g., âhigh-Doppler activity in the upper-body re- gionâ). That is, LLM reasoning relies on coordinate and Doppler cues from radar inputs, while VLM depends on vi- sual features, as detailed in the supplementary material. 6. Conclusion In this paper, we present the first mmWave-QA benchmark designed to assess a diverse range of human actions through natural language QA. Our benchmark integrates multiple mmWave sensing datasets to capture heterogeneity across devices, environments, and scene conditions through hardware-aware preprocessing and global taxonomy. Through extensive experiments, we demonstrate that LLMs can effectively interpret mmWave point cloud inputs and perform effective QA tasks. Notably, with the same underlying model, LLMs conditioned on mmWave point cloud inputs demonstrate greater robustness than those conditioned on RGB inputs, especially under visually degraded conditions such as low light or occlusion. This finding underscores the potential of leveraging LLMs for mmWave understanding to achieve reliable reasoning in vision-limited environments.While our study demon- strates the promising capabilities of LLMs for radar-based understanding using off-the-shelf models such as GPT and Gemini, exploration with diverse open-source LLMs remains limited. Future work will focus on improving open- source models through fine-tuning and exploring broader integration with reasoning-oriented frameworks such as retrieval-augmented generation systems and intelligent agents. We hope our work paves the way toward a broader understanding of multimodal reasoning that bridges non-visual sensing with language-based intelligence. References [1] J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millicah, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Saman- gooei, M. Monteiro, J. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan. Flamingo: a Visual Language Model for Few-Shot Learning. In Proceed- ings of the 36th International Conference on Neural Infor- mation Processing Systems, NIPS â22, Red Hook, NY, USA, 2022. Curran Associates Inc. ISBN 9781713871088. 2 [2] M. Alloulah and M. Arnold. Look, Radiate, and Learn: Self-Supervised Localisation via Radio-Visual Correspon- dence. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 17430â 17440, June 2023. 2 [3] S. An and U. Y. Ogras. MARS: mmWave-based Assistive Rehabilitation System For Smart Healthcare. ACM Trans- actions on Embedded Computing Systems (TECS), 20(5s): 1â22, 2021. 2, 3 [4] S. An and U. Y. Ogras. Fast and Scalable Human Pose Es- timation using mmWave Point Cloud. In Proceedings of the 59th ACM/IEEE Design Automation Conference, pages 889â 894, 2022. 2 [5] S. An, Y. Li, and U. Ogras. mRI: Multi-modal 3D Human Pose Estimation Dataset using mmWave, RGB-D, and In- ertial Sensors. In Thirty-sixth Conference on Neural Infor- mation Processing Systems Datasets and Benchmarks Track, 2022. URL https://openreview.net/forum?id= Oa2-cdfBxun. 1, 2, 3 [6] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei. Language Models are Few-Shot Learners. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors, Advances in Neural Information Process- ing Systems, volume 33, pages 1877â1901. Curran Asso- ciates, Inc., 2020.URL https://proceedings. neurips.c/paper_files/paper/2020/file/ 1457c0d6bfcb4967418bfb8ac142f64a - Paper . pdf. 2 [7] A. Chen, X. Wang, S. Zhu, Y. Li, J. Chen, and Q. Ye. mmBody Benchmark: 3D Body Reconstruction Dataset and Analysis for Millimeter Wave Radar. In Proceedings of the 30th ACM International Conference on Multimedia, M â22, page 3501â3510, New York, NY, USA, 2022. Associa- tion for Computing Machinery. ISBN 9781450392037. doi: 10.1145/3503161.3548262. URL https://doi.org/ 10.1145/3503161.3548262. 2, 3 [8] X. Chen and X. Zhang.RF Genesis: Zero-Shot Gen- eralization of mmWave Sensing through Simulation-Based Data Synthesis and Generative Diffusion Models. SenSys â23, page 28â42, New York, NY, USA, 2024. Association for Computing Machinery.ISBN 9798400704147.doi: 10.1145/3625687.3625798. URL https://doi.org/ 10.1145/3625687.3625798. 2 [9] X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer, A. Kolesnikov, J. Puigcerver, N. Ding, K. Rong, H. Akbari, G. Mishra, L. Xue, A. V. Thapliyal, J. Brad- bury, W. Kuo, M. Seyedhosseini, C. Jia, B. K. Ayan, C. R. Ruiz, A. P. Steiner, A. Angelova, X. Zhai, N. Houlsby, and R. Soricut. PaLI: A Jointly-Scaled Multilingual Language- Image Model. In The Eleventh International Conference on Learning Representations, 2023.URL https:// openreview.net/forum?id=mWVoBz4W0u. 2 [10] Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al. InternVL: Scal- ing up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24185â24198, 2024. 2 [11] J. Choi, S. Hor, S. Yang, and A. Arbabian. MVDoppler- pose: Multi-Modal Multi-View mmWave Sensing for Long- Distance Self-Occluded Human Walking Pose Estimation. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 27750â27759, 2025. 2, 3 [12] H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tai, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, A. Webson, S. S. Gu, Z. Dai, M. Suzgun, X. Chen, A. Chowdhery, A. Castro- Ros, M. Pellat, K. Robinson, D. Valter, S. Narang, G. Mishra, A. Yu, V. Zhao, Y. Huang, A. Dai, H. Yu, S. Petrov, E. H. Chi, J. Dean, J. Devlin, A. Roberts, D. Zhou, Q. V. Le, and J. Wei. Scaling Instruction-Finetuned Language Models. J. Mach. Learn. Res., 25(1), Jan. 2024. ISSN 1532-4435. 2 [13] W. Dai, J. Li, D. LI, A. Tiong, J. Zhao, W. Wang, B. Li, P. N. Fung, and S. Hoi. InstructBLIP: Towards General- purpose Vision-Language Models with Instruction Tuning. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Pro- cessing Systems, volume 36, pages 49250â49267. Curran Associates, Inc., 2023. URL https://proceedings. neurips.c/paper_files/paper/2023/file/ 9a6a435e75419a836fe47ab6793623e6 - Paper - Conference.pdf. 3 [14] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Lan- guage Understanding. In Proceedings of the 2019 conference of the North American chapter of the association for compu- tational linguistics: human language technologies, volume 1 (long and short papers), pages 4171â4186, 2019. 3 [15] C. Ding, L. Zhang, H. Chen, H. Hong, X. Zhu, and C. Li.Human Motion Recognition with Spatial- Temporal-ConvLSTM Network Using Dynamic Range- Doppler Frames Based on Portable FMCW Radar. IEEE Transactions on Microwave Theory and Techniques, 70(11): 5029â5038, 2022. 3 [16] A. Dosovitskiy. An Image is Worth 16x16 Words: Trans- formers for Image Recognition at Scale.arXiv preprint arXiv:2010.11929, 2020. 3 [17] Y. Fang, W. Menapace, A. Siarohin, T.-S. Chen, K.-C. Wang, I. Skorokhodov, G. Neubig, and S. Tulyakov. VIMI: Ground- ing Video Generation through Multi-modal Instruction. In Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, editors, Proceed- ings of the 2024 Conference on Empirical Methods in Natu- ral Language Processing, pages 4444â4456, Miami, Florida, USA, Nov. 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp- main.254. URL https: //aclanthology.org/2024.emnlp-main.254/. 3 [18] T. Geng, J. Zhang, Q. Wang, T. Wang, J. Duan, and F. Zheng.LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR), pages 18959â18969, June 2025. 3 [19] Google DeepMind. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities, 2025. URL https: //arxiv.org/abs/2507.06261. 2, 6 [20] Y. He, J. Wang, Y. Li, and Y. Luo. Dual-Path CNN-BiLSTM for mmWave-Based Human Skeletal Pose Estimation. IEEE Sensors Journal, 2025. 3 [21] S. Hor, S. Yang, J. Choi, and A. Arbabian. MVDoppler: Un- leashing the Power of Multi-View Doppler for Micromotion- Based Gait Classification. Advances in Neural Information Processing Systems, 36:58064â58074, 2023. 2, 3 [22] A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lam- ple, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed. Mis- tral 7B, 2023. URL https://arxiv.org/abs/2310. 06825. 2 [23] Z. Lai, J. Yang, S. Xia, L. Lin, L. Sun, R. Wang, J. Liu, Q. Wu, and L. Pei.RadarLLM: Empowering Large Language Models to Understand Human Motion from Millimeter-wave Point Cloud Sequence, 2025. URL https://arxiv.org/abs/2504.09862. 3 [24] H. Laurenc ̧on, L. Saulnier, T. Wang, C. Akiki, A. V. del Moral, T. Le Scao, L. von Werra, C. Mou, E. G. Ponferrada, H. Nguyen, J. Frohberg, M. Ë Sa Ë sko, Q. Lhoest, A. McMillan- Major, G. Dupont, S. Biderman, A. Rogers, L. Benallal, F. De Toni, G. Pistilli, O. Nguyen, S. Nikpoor, M. Masoud, P. Colombo, J. de la Rosa, P. Villegas, T. Thrush, S. Long- pre, S. Nagel, L. Weber, M. R. Mu Ě noz, J. Zhu, D. van Strien, Z. Alyafeai, K. Almubarak, V. M. Chien, I. Gonzalez-Dios, A. Soroa, K. Lo, M. Dey, P. O. Suarez, A. Gokaslan, S. Bose, D. I. Adelani, L. Phan, H. Tran, I. Yu, S. Pai, J. Chim, V. Lepercq, S. Ili Ě c, M. Mitchell, S. Luccioni, and Y. Jer- nite. The BigScience ROOTS Corpus: A 1.6TB Compos- ite Multilingual Dataset. In Proceedings of the 36th Inter- national Conference on Neural Information Processing Sys- tems, NIPS â22, Red Hook, NY, USA, 2022. Curran Asso- ciates Inc. ISBN 9781713871088. 2 [25] S.-P. Lee, N. P. Kini, W.-H. Peng, C.-W. Ma, and J.-N. Hwang.HuPR: A Benchmark for Human Pose Estima- tion Using Millimeter Wave Radar. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5715â5724, 2023. 1 [26] J. Li, D. Li, S. Savarese, and S. Hoi. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In Proceedings of the 40th International Conference on Machine Learning, ICMLâ23. JMLR.org, 2023. 2, 3 [27] F. Liang, B. Wu, X. Dai, K. Li, Y. Zhao, H. Zhang, P. Zhang, P. Vajda, and D. Marculescu. Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7061â7070, 2023. 2 [28] K. Liang, Y. Li, H. Hao, H. Zeng, A. Zhou, H. Ma, C. Wu, H. Jia, J. Zuo, and J. Fan. mmHolmes: Amodal Millimeter- wave Sensing by Understanding Human Kinetics. 9(3), Sept. 2025. doi: 10.1145/3749500. URL https://doi.org/ 10.1145/3749500. 1 [29] H. Liu, K. Cui, K. Hu, Y. Wang, A. Zhou, L. Liu, and H. Ma. mTransSee: Enabling Environment-Independent mmWave Sensing Based Gesture recognition via Transfer Learning. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 6(1):1â28, 2022. 2 [30] H. Liu, C. Li, Q. Wu, and Y. J. Lee. Visual Instruction Tun- ing. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS â23, Red Hook, NY, USA, 2023. Curran Associates Inc. 2, 3 [31] H. Liu, C. Li, Y. Li, and Y. J. Lee. Improved Baselines with Visual Instruction Tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26296â26306, June 2024. 2, 3 [32] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10012â 10022, 2021. 3 [33] J. Lu, C. Clark, S. Lee, Z. Zhang, S. Khosla, R. Marten, D. Hoiem, and A. Kembhavi. Unified-IO 2: Scaling Autore- gressive Multimodal Models with Vision Language Audio and Action. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26439â26455, June 2024. 3 [34] H. Luo, Y. Liu, P. Liu, and X. Liu. Vector-Quantized Prompt Learning for Paraphrase Generation. In The 2023 Confer- ence on Empirical Methods in Natural Language Process- ing, 2023. URL https://openreview.net/forum? id=9cALtYoAEy. 4 [35] S. Mehta and M. Rastegari.MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer. arXiv preprint arXiv:2110.02178, 2021. 3 [36] OpenAI. GPT-4o System Card, 2024. URL https:// arxiv.org/abs/2410.21276. 2, 6, 7 [37] OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, and . . . . GPT-4 Technical Report, 2023. URL https://arxiv.org/abs/2303.08774. 2, 4 [38] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wain- wright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe. Training language models to follow instructions with human feedback. In Proceedings of the 36th International Confer- ence on Neural Information Processing Systems, NIPS â22, Red Hook, NY, USA, 2022. Curran Associates Inc. ISBN 9781713871088. 2 [39] G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, H. Alobei- dli, A. Cappelli, B. Pannier, E. Almazrouei, and J. Lau- nay. The RefinedWeb Dataset for Falcon LLM: Outper- forming Curated Corpora with Web Data Only. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Process- ing Systems, volume 36, pages 79155â79172. Curran As- sociates, Inc., 2023. URL https://proceedings. neurips.c/paper_files/paper/2023/file/ fa3ed726c5073b9c31e3e49a807789c - Paper - Datasets_and_Benchmarks.pdf. 2 [40] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning Transferable Visual Models From Natural Language Supervision. In M. Meila and T. Zhang, editors, Proceedings of the 38th Interna- tional Conference on Machine Learning, volume 139 of Pro- ceedings of Machine Learning Research, pages 8748â8763. PMLR, 18â24 Jul 2021. URL https://proceedings. mlr.press/v139/radford21a.html. 2 [41] M. M. Rahman, R. Yataka, S. Kato, P. Wang, P. Li, A. Car- dace, and P. Boufounos. MMVR: Millimeter-wave Multi- View Radar Dataset and Benchmark for Indoor Perception. In European Conference on Computer Vision, pages 306â 322. Springer, 2024. 2, 3 [42] M. Ravaut, A. Sun, N. Chen, and S. Joty. On Context Uti- lization in Summarization with Large Language Models. In L.-W. Ku, A. Martins, and V. Srikumar, editors, Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 2764â 2781, Bangkok, Thailand, Aug. 2024. Association for Com- putational Linguistics. doi: 10.18653/v1/2024.acl-long.153. URL https://aclanthology.org/2024.acl- long.153/. 2 [43] V. Sanh, A. Webson, C. Raffel, S. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, A. Raja, M. Dey, M. S. Bari, C. Xu, U. Thakker, S. S. Sharma, E. Szczechla, T. Kim, G. Chhablani, N. Nayak, D. Datta, J. Chang, M. T.-J. Jiang, H. Wang, M. Manica, S. Shen, Z. X. Yong, H. Pandey, R. Bawden, T. Wang, T. Neeraj, J. Rozen, A. Sharma, A. Santilli, T. Fevry, J. A. Fries, R. Teehan, T. L. Scao, S. Biderman, L. Gao, T. Wolf, and A. M. Rush. Multitask Prompted Training Enables Zero-Shot Task Generalization. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id= 9Vrb9D0WI4. 2 [44] A. Sengupta, F. Jin, R. Zhang, and S. Cao. m-Pose: Real- Time Human Skeletal Posture Estimation Using mmWave Radars and CNNs. IEEE sensors journal, 20(17):10032â 10044, 2020. 3 [45] Z. Sheng, H. Xu, Q. Zhang, and D. Wang. Facilitating Radar- Based Gesture Recognition With Self-Supervised Learning. In 2022 19th Annual IEEE International Conference on Sensing, Communication, and Networking (SECON), pages 154â162. IEEE, 2022. 2 [46] X. Shi, M. Bouazizi, and T. Ohtsuki. mmPoint-Attention: A Unified Attention Framework for Human Pose and Activ- ity Recognition From mmWave Radar Point Clouds. IEEE Sensors Journal, 2025. 3 [47] A. D. Singh, S. S. Sandha, L. Garcia, and M. Srivastava. RadHAR: Human Activity Recognition from Point Clouds Generated through a Millimeter-wave Radar. In Proceedings of the 3rd ACM Workshop on Millimeter-wave Networks and Sensing Systems, pages 51â56, 2019. 2, 3 [48] L. Soldaini, R. Kinney, A. Bhagia, D. Schwenk, D. Atkin- son, R. Authur, B. Bogin, K. Chandu, J. Dumas, Y. Elazar, V. Hofmann, A. Jha, S. Kumar, L. Lucy, X. Lyu, N. Lam- bert, I. Magnusson, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. Peters, A. Ravichander, K. Richardson, Z. Shen, E. Strubell, N. Subramani, O. Tafjord, E. Walsh, L. Zettle- moyer, N. Smith, H. Hajishirzi, I. Beltagy, D. Groeneveld, J. Dodge, and K. Lo. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research. In L.-W. Ku, A. Martins, and V. Srikumar, editors, Pro- ceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15725â15788, Bangkok, Thailand, Aug. 2024. Association for Computational Linguistics.doi: 10 . 18653 / v1 / 2024 . acl- long.840. URL https://aclanthology.org/ 2024.acl-long.840/. 2 [49] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi ` ere, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample. LLaMA: Open and Efficient Foundation Language Models, 2023. URL https://arxiv.org/abs/2302.13971. 2 [50] F. Wang, Y. Lv, M. Zhu, H. Ding, and J. Han. XRF55: A Ra- dio Frequency Dataset for Human Indoor Action Analysis. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 8(1):1â34, 2024. 2, 3 [51] X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou. Self-Consistency Improves Chain of Thought Reasoning in Language Models. In The Eleventh International Conference on Learning Rep- resentations, 2023. URL https://openreview.net/ forum?id=1PL1NIMMrw. 2 [52] J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS â22, Red Hook, NY, USA, 2022. Curran Associates Inc. ISBN 9781713871088. 2 [53] Z. Wu, D. Zhang, C. Xie, C. Yu, J. Chen, Y. Hu, and Y. Chen. RFMask: A Simple Baseline for Human Silhouette Segmen- tation with Radio Signals. IEEE Transactions on Multime- dia, 25:4730â4741, 2022. 2, 3 [54] J. Xu et al. GroupViT: Semantic Segmentation Emerges from Text Supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. URL https://openaccess.thecvf. com/content/CVPR2022/papers/Xu_GroupViT_ Semantic_Segmentation_Emerges_From_Text_ Supervision_CVPR_2022_paper.pdf. 2 [55] R. Xu, X. Wang, T. Wang, Y. Chen, J. Pang, and D. Lin. PointLLM: Empowering Large Language Models to Under- stand Point Clouds. In Computer Vision â ECCV 2024: 18th European Conference, Milan, Italy, September 29âOctober 4, 2024, Proceedings, Part XXV, page 131â147, Berlin, Hei- delberg, 2024. Springer-Verlag.ISBN 978-3-031-72697- 2. doi: 10.1007/978- 3- 031- 72698- 9 8. URL https: //doi.org/10.1007/978-3-031-72698-9_8. 3 [56] R. Xu, S. Yang, X. Wang, T. Wang, Y. Chen, J. Pang, and D. Lin.PointLLM-V2: Empowering Large Lan- guage Models to Better Understand Point Clouds . IEEE Transactions on Pattern Analysis & Machine Intelligence, (01):1â15, July 5555.ISSN 1939-3539.doi:10 . 1109 / TPAMI . 2025 . 3590784.URL https://doi. ieeecomputersociety.org/10.1109/TPAMI. 2025.3590784. 3 [57] H. Xue, Y. Ju, C. Miao, Y. Wang, S. Wang, A. Zhang, and L. Su. mmMesh: Towards 3D Real-Time Dynamic Hu- man Mesh Construction Using Millimeter-Wave. MobiSys â21, page 269â282, New York, NY, USA, 2021. Associa- tion for Computing Machinery. ISBN 9781450384438. doi: 10.1145/3458864.3467679. URL https://doi.org/ 10.1145/3458864.3467679. 2 [58] H. Xue, Q. Cao, Y. Ju, H. Hu, H. Wang, A. Zhang, and L. Su. M4esh: mmWave-based 3D Human Mesh Construction for Multiple Subjects. In Proceedings of the 20th ACM Confer- ence on Embedded Networked Sensor Systems, pages 391â 406, 2022. 2 [59] H. Xue, Q. Cao, C. Miao, Y. Ju, H. Hu, A. Zhang, and L. Su. Towards Generalized mmWave-based Human Pose Estima- tion through Signal Augmentation. In Proceedings of the 29th Annual International Conference on Mobile Computing and Networking, pages 1â15, 2023. 2 [60] A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid. Zero- Shot Video Question Answering via Frozen Bidirectional Language Models. NIPS â22, Red Hook, NY, USA, 2022. Curran Associates Inc. ISBN 9781713871088. 2 [61] J. Yang, H. Huang, Y. Zhou, X. Chen, Y. Xu, S. Yuan, H. Zou, C. X. Lu, and L. Xie. M-Fi: Multi-Modal Non- Intrusive 4D Human Dataset for Versatile Wireless Sens- ing.In Thirty-seventh Conference on Neural Informa- tion Processing Systems Datasets and Benchmarks Track, 2023. URL https://openreview.net/forum?id= 1uAsASS1th. 2, 3 [62] X. Yang, J. Liu, Y. Chen, X. Guo, and Y. Xie. MU-ID: Multi- user Identification Through Gaits Using Millimeter Wave Radios. In IEEE INFOCOM 2020 - IEEE Conference on Computer Communications, pages 2589â2598, 2020. doi: 10.1109/INFOCOM41043.2020.9155471. 2 [63] X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer. Sigmoid Loss for Language Image Pre-Training. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 11941â11952, 2023. doi: 10.1109/ICCV51070.2023.01100. 2 [64] H. Zhang, X. Li, and L. Bing.Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding. In Y. Feng and E. Lefever, editors, Proceed- ings of the 2023 Conference on Empirical Methods in Nat- ural Language Processing: System Demonstrations, pages 543â553, Singapore, Dec. 2023. Association for Computa- tional Linguistics. doi: 10.18653/v1/2023.emnlp-demo.49. URL https://aclanthology.org/2023.emnlp- demo.49/. 2, 3 [65] Y. Zhang, J. Wu, W. Li, B. Li, Z. Ma, Z. Liu, and C. Li. LLaVA-Video: Video Instruction Tuning With Synthetic Data, 2025. URL https://arxiv.org/abs/2410. 02713. 2, 3 [66] M. Zhao, T. Li, M. Abu Alsheikh, Y. Tian, H. Zhao, A. Tor- ralba, and D. Katabi. Through-Wall Human Pose Estima- tion Using Radio Signals. In Proceedings of the IEEE con- ference on computer vision and pattern recognition, pages 7356â7365, 2018. 2 [67] M. Zhao, Y. Tian, H. Zhao, M. A. Alsheikh, T. Li, R. Hris- tov, Z. Kabelac, D. Katabi, and A. Torralba. RF-Based 3D Skeletons. In Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Communication, pages 267â281, 2018. 2 [68] P. Zhao, C. X. Lu, J. Wang, C. Chen, W. Wang, N. Trigoni, and A. Markham. mID: Tracking and Identifying People with Millimeter Wave Radar. In 2019 15th International Conference on Distributed Computing in Sensor Systems (DCOSS), pages 33â40, 2019. doi: 10.1109/DCOSS.2019. 00028. 1 [69] Z. Zhao, E. Monti, J. Lehmann, and H. Assem. Enhanc- ing Contextual Understanding in Large Language Models through Contrastive Decoding. In K. Duh, H. Gomez, and S. Bethard, editors, Proceedings of the 2024 Conference of the North American Chapter of the Association for Compu- tational Linguistics: Human Language Technologies (Vol- ume 1: Long Papers), pages 4225â4237, Mexico City, Mex- ico, June 2024. Association for Computational Linguistics. doi: 10 . 18653 / v1 / 2024 . naacl - long . 237. URL https: //aclanthology.org/2024.naacl-long.237/. 2 [70] C. Zhou and J. Yang. HoloLLM: Multisensory Foundation Model for Language-Grounded Human Sensing and Reason- ing, 2025. URL https://arxiv.org/abs/2505. 17645. 3 [71] B. Zhu, Z. He, W. Xiong, G. Ding, T. Huang, and W. Xi- ang. ProbRadarM3F: mmWave Radar-based Human Skeletal Pose Estimation with Probability Map Guided Multi-Format Feature Fusion. IEEE Transactions on Aerospace and Elec- tronic Systems, pages 1â11, 2025. doi: 10.1109/TAES.2025. 3594328. 2 [72] Y. Zhu, J. R. A. Moniz, S. Bhargava, J. Lu, D. Piraviperu- mal, S. Li, Y. Zhang, H. Yu, and B.-H. Tseng. Can Large Language Models Understand Context? In Y. Graham and M. Purver, editors, Findings of the Association for Compu- tational Linguistics: EACL 2024, pages 2004â2018, St. Ju- lianâs, Malta, Mar. 2024. Association for Computational Lin- guistics. URL https://aclanthology.org/2024. findings-eacl.135/. 2