Paper deep dive
The Language of Touch: Translating Vibrations into Text with Dual-Branch Learning
Jin Chen, Yifeng Lin, Chao Zeng, Si Wu, Tiesong Zhao
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 98%
Last extracted: 3/31/2026, 1:32:19 AM
Summary
The paper introduces Vibrotactile Periodic-Aperiodic Captioning (ViPAC), a novel framework for generating natural language descriptions from vibrotactile signals. It addresses the lack of paired data by constructing LMT108-CAP, a dataset using GPT-4o to annotate LMT-108 surface images. The model utilizes a dual-branch encoder to disentangle periodic and aperiodic signal components, combined with a dynamic fusion mechanism and a Transformer-based decoder to achieve superior semantic alignment.
Entities (5)
Relation Signals (3)
LMT108-CAP → derivedfrom → LMT-108
confidence 100% · construct LMT108-CAP... from the popular LMT-108 dataset
GPT-4o → generatedannotationsfor → LMT108-CAP
confidence 100% · using GPT-4o to generate five constrained captions per surface image from the popular LMT-108 dataset
ViPAC → usesdataset → LMT108-CAP
confidence 100% · Experiments show that ViPAC significantly outperforms the baseline methods... on the dataset [LMT108-CAP]
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The standardization of vibrotactile data by IEEE P1918.1 workgroup has greatly advanced its applications in virtual reality, human-computer interaction and embodied artificial intelligence. Despite these efforts, the semantic interpretation and understanding of vibrotactile signals remain an unresolved challenge. In this paper, we make the first attempt to address vibrotactile captioning, {\it i.e.}, generating natural language descriptions from vibrotactile signals. We propose Vibrotactile Periodic-Aperiodic Captioning (ViPAC), a method designed to handle the intrinsic properties of vibrotactile data, including hybrid periodic-aperiodic structures and the lack of spatial semantics. Specifically, ViPAC employs a dual-branch strategy to disentangle periodic and aperiodic components, combined with a dynamic fusion mechanism that adaptively integrates signal features. It also introduces an orthogonality constraint and weighting regularization to ensure feature complementarity and fusion consistency. Additionally, we construct LMT108-CAP, the first vibrotactile-text paired dataset, using GPT-4o to generate five constrained captions per surface image from the popular LMT-108 dataset. Experiments show that ViPAC significantly outperforms the baseline methods adapted from audio and image captioning, achieving superior lexical fidelity and semantic alignment.
Tags
Links
- Source: https://arxiv.org/abs/2603.26804v1
- Canonical: https://arxiv.org/abs/2603.26804v1
Trouble viewing inline? Open PDF directly →
Full Text
46,989 characters extracted from source content.
Expand or collapse full text
JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20211 The Language of Touch: Translating Vibrations into Text with Dual-Branch Learning Jin Chen, Yifeng Lin, Chao Zeng, Si Wu, and Tiesong Zhao, Senior Member, IEEE Abstract—The standardization of vibrotactile data by IEEE P1918.1 workgroup has greatly advanced its applications in vir- tual reality, human-computer interaction and embodied artificial intelligence. Despite these efforts, the semantic interpretation and understanding of vibrotactile signals remain an unresolved challenge. In this paper, we make the first attempt to ad- dress vibrotactile captioning, i.e., generating natural language descriptions from vibrotactile signals. We propose Vibrotactile Periodic-Aperiodic Captioning (ViPAC), a method designed to handle the intrinsic properties of vibrotactile data, including hybrid periodic-aperiodic structures and the lack of spatial semantics. Specifically, ViPAC employs a dual-branch strategy to disentangle periodic and aperiodic components, combined with a dynamic fusion mechanism that adaptively integrates signal features. It also introduces an orthogonality constraint and weighting regularization to ensure feature complementarity and fusion consistency. Additionally, we construct LMT108-CAP, the first vibrotactile-text paired dataset, using GPT-4o to generate five constrained captions per surface image from the popular LMT-108 dataset. Experiments show that ViPAC significantly outperforms the baseline methods adapted from audio and image captioning, achieving superior lexical fidelity and semantic alignment. Index Terms—Vibrotactile Captioning, Haptic Perception, Multimodal Learning, Haptic Signal Processing. I. INTRODUCTION Haptics plays a critical role in enhancing multimodal in- teraction by complementing visual and auditory modalities. It has been widely applied in industrial automation, telemedicine, hazardous environment exploration, and e-commerce. Typi- cally, haptic signals can be classified into kinesthetic and vibrotactile components [1]: the former involves motion, force, and torque, while the latter conveys surface-related properties such as friction, hardness, temperature, and roughness. Until now, vibrotactile signals can be captured by different types of sensors, including visuotactile sensors [2] and triaxial This work is supported by the National Science Foundation of China (Grant No. 62571131) (Corresponding author: Tiesong Zhao.) J. Chen and Y. Lin are with the Fujian Key Lab for Intelligent Processing and Wireless Transmission of Media Information, Fuzhou University, Fuzhou 350108, China(e-mails:231110036, 241120007@fzu.edu.cn). T. Zhao is with the Fujian Key Lab for Intelligent Processing and Wireless Transmission of Media Information, Fuzhou University, Fuzhou 350108, China and also with the Fujian Science and Technology Innovation Laboratory for Optoelectronic Information of China, Fuzhou 350108, China (e-mails: t.zhao@fzu.edu.cn). C. Zeng is with the School of Artificial Intelligence, Hubei University, Wuhan, China, and the Key Laboratory of Intelligent Sensing System and Security, Hubei University and Ministry of Education, China. (email: chao.zeng@hubu.edu.cn) S. Wu is with the School of Computer Science and Engineering, South China University of Technology. (email: cswusi@scut.edu.cn) Semantic Retrieval Material Inspection & Automated Reports VR Semantic Augmentation in Tactile Rendering Vibrotactile Captioning Vibration Signal Natural Language Caption This surface feels coarse and irregular... Fig. 1. Application scenarios of vibrotactile captioning. accelerometers [3]. Although visuotactile sensors such as Gel- Sight have attracted considerable attention from the computer vision community [4], the IEEE P1918.1 Working Group stan- dardized the representation of vibrotactile signals as multiple 1D time series in 2019 [5]. This step has significantly en- hanced the perception, transmission, reproduction, and cross- platform application of vibrotactile signals [6]. It has also promoted interoperability across tactile devices and enabled system-level integration in virtual reality, human–computer interaction, and embodied artificial intelligence applications [5]. Despite their broad applicability, vibrotactile signals are inherently complex and noisy, which makes their interpretation challenging [7]. To facilitate a deeper understanding of these signals, we adapt the concept of audio-visual captioning and introduce the task of vibrotactile captioning, which translates these 1D signals into structured natural language. Rather than relying on low-level features or fixed taxonomies, natural language aligns more closely with human perception and facilitates intuitive understanding. As illustrated in Fig. 1, vibrotactile captioning enables three practical applications: (i) Semantic indexing, where captions support language-based search, retrieval, and reasoning; (i) Material inspection, where standardized textual summaries assist in quality control and verification within industrial workflows; (i) Virtual per- ception, where captions offer semantic guidance to augment texture understanding in virtual reality environments with limited haptic resolution. Although Large Language Models (LLMs) have shown strengths in describing signals, their massive parameter and computational power requirements also limit their applications in the above scenarios. To the best of our knowledge, vibrotactile captioning or similar works have not been explored. This might be attributed arXiv:2603.26804v1 [cs.CV] 26 Mar 2026 JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20212 to data scarcity and technical challenges. Data Scarcity. First, there are few publicly available vibrotactile datasets, and most lack natural language annotations. For example, the LMT haptic texture database [8] offers triaxial acceleration signals but does not provide textual descriptions, reducing its value for cross-modal learning. Second, vibrotactile signals often contain relevant and irrelevant components, making it difficult to extract clean and meaningful features [9]. Third, human descriptions of tactile sensations vary widely depending on individual perception, expertise, and attention [10], which hin- ders consistent large-scale annotation. Technical Challenges. First, captioning models developed for image, video, or audio domains rely on structural priors such as spatial layouts, motion continuity, or acoustic regularity. The absence of these priors in vibrotactile signals leads to difficulties in cross- modal transfer. Second, vibrotactile data encodes material surface properties as complex temporal dynamics that combine periodic patterns (e.g., from regular textures) and aperiodic patterns (e.g., from irregular surfaces or noise). This hybrid structure poses significant modeling challenges for single- stream encoders. To address the above issues, we propose ViPAC, a caption- ing framework tailored to vibrotactile signals. To mitigate data scarcity, we employ GPT-4o to generate textual descriptions of material surface images and pair them with their corresponding vibration signals to construct a cross-modal dataset. To over- come modeling challenges, we design a dual-branch encoder that separately processes periodic and aperiodic components for stable and transient vibrotactile cues, respectively. These features are fused via adaptive weighting and decoded via a Transformer-based decoder. This design allows ViPAC to effectively model complex vibrotactile structures and generate perception-aligned descriptions. Our main contributions are as follows: • We introduce the task of vibrotactile captioning, which aims to convert 1D triaxial acceleration signals into structured natural language descriptions that reflect char- acteristics of material surface and haptic signals. This task establishes a new direction for semantic modeling and interpretation of tactile data. • We propose ViPAC, the first captioning framework specif- ically designed for vibrotactile signals. The model em- ploys a dual-branch encoder to separately extract periodic and aperiodic features, and integrates them through an adaptive fusion mechanism that captures the hybrid tem- poral structure of tactile inputs. The fused representation is then decoded into natural language using a standard Transformer decoder. • We construct LMT108-CAP, a vibrotactile-text paired dataset generated using GPT-4o under controlled lin- guistic constraints. Comprehensive experiments on the dataset validate the effectiveness of ViPAC in capturing the temporal structure of tactile signals and producing descriptive outputs that align with human interpretations. The remainder of this paper is organized as follows. Section I reviews related work on vibrotactile datasets and text description tasks. Section I introduces the construction of the proposed LMT108-CAP vibrotactile–text paired dataset. Section IV presents the ViPAC framework, including the dual- branch encoder, dynamic fusion mechanism, and decoding strategy. Section V reports experimental results, ablation stud- ies, and a retrieval demo to validate the effectiveness of the proposed method. Finally, Section VI concludes the paper and outlines directions for future work. I. RELATED WORK A. Datasets Existing vibrotactile datasets fall into two categories: visuo- tactile patterns and multiple 1D signals. Significant progress has been made in visuotactile datasets such as TVL [11] and Touch100k [12], which use vision-based sensors to produce representations compatible with computer vision techniques, thereby facilitating multimodal alignment and captioning. In contrast, multiple 1D vibrotactile datasets aligns with the IEEE P1918.1 standard that advocates multiple 1D vibration signals as the canonical format. For example, the LMT hap- tic texture database [8] provides triaxial acceleration signals without paired natural language descriptions, which restricts its applicability in cross-modal learning. To date, no public dataset aligns triaxial signals with textual descriptions, posing significant obstacles to model training and evaluation. Moti- vated by the success of LLMs in visuotactile research [11], we extend this paradigm to multiple 1D vibrotactile signals. We use GPT-4o to generate natural language descriptions for surface images and pair them with corresponding triaxial acceleration signals, resulting in a new dataset that supports tactile semantic modeling and cross-modal learning. B. Text Description Tasks 1) Image-to-Text: Image captioning aims to generate de- tailed textual descriptions of visual content by leveraging the structured spatial semantics present in images, such as object presence, spatial layout, and visual attributes [13]–[15]. These methods perform well in visual domains but cannot be directly applied to vibrotactile data. Unlike images, vibrotactile signals do not contain spatial structure and instead encode material properties through temporal and spectral patterns. This fun- damental difference makes conventional vision-based feature extractors, such as Convolutional Neural Networks (CNNs), ineffective when applied to multiple 1D tactile signals. 2) Video-to-Text: Video captioning extends image caption- ing into the temporal domain, modeling dynamic content such as motion trajectories, temporal coherence, and scene transitions [16]–[18]. Although vibrotactile signals also evolve over time, they capture localized, fine-grained surface inter- actions rather than global scene-level dynamics. As a result, video captioning models, which emphasize macroscopic visual motion, fail to represent the subtle frequency-based variations characteristic of tactile vibrations. 3) Audio-to-Text: Audio captioning generates textual de- scriptions from acoustic signals, benefiting from well-defined semantic categories such as speech, music, and environmental sounds, along with consistent recording conditions [19]–[21]. In contrast, vibrotactile signals exhibit high variability due JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20213 This material surface is rough-textured with irregular jagged edges and varying sizes of protruding fragments. This material surface has a rough uneven texture with slightly jagged edges. This material surface feels rough with small uneven closely-packed pebbles. This material surface feels coarse with uneven jagged small pebbles closely packed together. This material surface is rough with small irregularly shaped protrusions. Natural language description: Three-axis acceleration signal: Image: Prompt: Sentence pattern Exclude color information Length restriction Deterministic descriptions Source Data Generate Textual Descriptions Paired with Corresponding Signals GPT-4o VisionTouch Fig. 2. Illustration of the vibrotactile-text dataset generation process. Surface images from the LMT-108 dataset are provided as input to GPT-4o, which generates five textual descriptions per image under predefined linguistic constraints. These descriptions are then paired with the corresponding triaxial acceleration signals collected from the same material surfaces, resulting in the final vibrotactile-text dataset. Fig. 3. Examples of material surface images and their corresponding three- axis vibrotactile signals. Top: materials with regular textures exhibit strong periodicity. Bottom: irregular surfaces yield noisy, aperiodic signals. This motivates the use of distinct modeling pathways. to scanning speed and contact force, and lack standardized semantic labels. This variability introduces ambiguity and noise, making direct application of audio captioning models ineffective for vibrotactile interpretation. Above all, these modality-specific captioning models rely on structural priors that do not hold in the vibrotactile domain. Image models assume spatial semantics, video models empha- size motion continuity, and audio models exploit categorical regularity—none of which apply to tactile vibrations. To ad- dress this gap, we propose ViPAC, a captioning framework that models vibrotactile-specific features through separate periodic and aperiodic branches. The model further employs adaptive fusion to accommodate temporal variability and mitigate se- mantic ambiguity. This design provides a dedicated solution to the challenges of cross-modal captioning in tactile contexts. I. LMT108-CAP DATASET CONSTRUCTION To overcome the lack of publicly available paired vibrotactile-text data, we introduce LMT108-CAP, a novel captioned vibrotactile dataset. This dataset, is based on the LMT-108 Surface-Materials database [22] promoted by the IEEE P1918.1 workgroup. The LMT-108 Surface-Materials database contains 108 distinct material surfaces grouped into 10 categories, each recorded with triaxial acceleration, friction, sound, and surface image data. Each surface is measured 20 times, resulting in 2,160 vibrotactile samples represented as 1D triaxial acceleration signals. The dataset is widely recognized and applied for its representativeness. As one of the field’s most widely used tactile benchmarks, LMT-108 offers representative coverage and thus provides a suitable basis for constructing paired data. We construct the vibrotactile-text pairs with the following approach, which is inspired by image captioning dataset construction such as Flickr8K [23]. Specifically, we employ GPT-4o to generate five textual descriptions for each material surface image in the LMT-108 dataset. The generation process adheres to four constraints designed to ensure relevance and consistency with the characteristics of vibrotactile signals: (i) Sentence Pattern: Descriptions start with “This material surface...” to encourage vibrotactile-focused textual output. (i) Exclude Color Information: To align with vibrotactile signal characteristics, which do not capture color. (i) Length Restriction: Each description contains no more than 15 words. (iv) Deterministic Descriptions: To prevent uncertain or imaginative associations, ensuring consistency and relevance to vibrotactile properties. Images are used only once to bootstrap the textual side; no visual inputs are used during training or inference, so the model learns signal–semantic correspondences rather than reproducing visual semantics. LMT108-CAP contains surface images, their corresponding triaxial acceleration signals, and GPT-generated constrained captions, as illustrated in Fig. 2. In total, the dataset includes 2,160 samples, each with five captions, yielding 10,800 paired instances following the five-caption protocol of Flickr8K [23]. We split the data into training and testing subsets using a JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20214 Triaxial Acceleration Signals Aperiodic Branch Periodic Branch Input One-Dimensional Signals DFT321 Encoder periodicity L aperiodicity L orthogonality L Dynamic Feature Fusion L i n e a r + S o f t m a x P r e v i o u s T o k e n s Output "This material surface..." CE L Cross-Attention Decoder Word Embedding PosEnc Periodic Branch Aperiodic Branch Mel-Spectrogram Conv+Pool Cosine Sine Activation Function Mel-Spectrogram Learnable Weights LSTMLSTMLSTM ...... ...... Transformer ... ... ... T r a n s f o r m e r D e c o d e r Word probability distribution per time step. Fig. 4. ViPAC takes triaxial acceleration signals as input and applies DFT321 to obtain 1D vibration data. These signals are processed by a dual-branch encoder that separately models periodic and aperiodic components using FAN-based frequency analysis and Transformer+LSTM-based temporal modeling, respectively. The extracted features are dynamically fused based on estimated periodicity scores, and the fused representation is decoded into natural language using a Transformer decoder. 7:3 ratio, with 1,512 training samples (7,560 captions) and 648 test samples (3,240 captions), and perform corpus-level vocabulary control by filtering singleton words so that each remaining token appears at least once in both splits [24]. The selected materials span a broad range of textures, providing representative coverage for training and evaluating vibrotactile captioning models. To ensure the reliability of the generated text, we impose strict post-processing constraints. The LMT- 108 dataset has also been widely used in prior work on tactile sensing and surface analysis [25]–[27], further supporting its suitability as a foundation for paired caption construction. IV. PROPOSED VIPAC MODEL The primary innovation of this work lies in defining and systematically tackling the novel task of vibrotactile caption- ing. Vibrotactile signals, unlike visual or auditory signals, often combine periodic and aperiodic components that arise from structured (periodic) and unstructured (aperiodic) surface interactions [28], as illustrated in Fig. 3. This hybrid structure challenges single-stream models and motivates an approach that treats these components explicitly. We propose ViPAC, a dedicated framework that translates vibration signals into natural language through three stages: (i) dual-branch encoding to extract periodic and aperiodic fea- tures, (i) adaptive fusion guided by signal characteristics, and (i) sequence generation with a Transformer-based decoder. This design targets both repetitive and transient patterns in vibrotactile signals; the overall architecture is shown in Fig. 4. In the encoder, periodic and aperiodic components are mod- eled separately to match their distinct statistics—an approach consistent with practices in surface engineering and tactile sensing. Concretely, the periodic branch employs a Fourier Analysis Network (FAN) for frequency-domain processing of stable, repeating patterns, while the aperiodic branch uses an LSTM+Transformer stack to capture irregular, long-range tem- poral variation. The two streams are then fused adaptively via a periodicity embedding p i , enabling the model to emphasize the appropriate branch for each input and to form a unified representation for decoding. A. Encoder Design To simplify the input while preserving perceptually relevant cues, we apply DFT321 to compress triaxial acceleration into a single 1-D signal. DFT321 is a widely used, standard front end in vibrotactile processing, including work that processes the LMT-108 dataset in the same way [25], and it is consistent with the IEEE P1918.1 recommendations. Landin et al. [29] introduced DFT321 and showed that collapsing triaxial high- frequency vibration into a single magnitude channel causes no noticeable perceptual degradation. Our objective is to generate semantic descriptions of surface properties; the relevant cues are texture statistics and spectral structure rather than full 3- D force or direction vectors. Direction-invariant magnitude fusion therefore removes orientation nuisance while preserving the information needed for captioning. The resulting signal t i typically consists of both peri- odic patterns from regular textures and aperiodic fluctuations caused by irregular surfaces or noise: t i = t PER + t APER .(1) Conventional single-stream encoders fail to adequately cap- ture this hybrid structure: f = f enc (t i ).(2) To address this, we propose a dual-branch encoder that models periodic and aperiodic components independently: f PER = f PER (t i ), f APER = f APER (t i ).(3) JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20215 These features are then fused into a unified representation: f = Fuse(f PER , f APER ).(4) a) Periodic Branch: To capture stable, repetitive patterns that emerge from regular textures, we employ the Fourier Analysis Network (FAN) [30]. FAN is chosen for its effec- tiveness in extracting dominant frequency components from time-series signals. The transformed output is converted into a Mel-Spectrogram and further processed by a convolution- pooling module to obtain spectral representations: f PER,i = ConvPool(Mel(FAN(t i ))).(5) To enhance the learning of periodic structure, we introduce a periodicity loss computed as the variance of peak intervals in the autocorrelation function: L periodicity = Var(∆FAN(t i )),(6) where ∆t denotes the lag between autocorrelation peaks. b) Aperiodic Branch: To model irregular, non-repetitive variations commonly observed in natural surfaces, we use a Transformer encoder combined with a Long-Short Term Mem- ory (LSTM) layer. The layer captures short-term dynamics, while the Transformer provides long-range temporal modeling. This architecture is selected to address both local variability and global dependencies: f APER,i = Transformer(LSTM(Mel(t i ))).(7) To regularize the aperiodic features and avoid overly large activations, we apply a Mean Squared Error (MSE) style penalty to their magnitude: L aperiodicity = 1 D ∥f APER,i ∥ 2 2 .(8) c) Feature Decoupling and Fusion: To ensure that the two branches capture complementary information, we apply an orthogonality loss between periodic and aperiodic features: L orthogonality =∥⟨f PER,i , f APER,i ⟩∥ 2 .(9) We compute a scalar fusion weight w i ∈ [0, 1] based on the periodicity embedding p i via a sigmoid activation: w i = σ(α· (p i − τ )),(10) where τ is a learnable threshold and α controls the sharpness of the transition. The final fused representation is: f i = w i · f PER,i + (1− w i )· f APER,i .(11) B. Decoder Design The decoder transforms the fused tactile feature f i into a natural language description c i = c 1 ,c 2 ,...,c T in an auto-regressive manner. We adopt a standard Transformer- based decoder [31] consisting of a word embedding layer, a Transformer block with masked self-attention and cross- attention, and a linear projection layer with softmax activation. At each time step t, the decoder computes the probability of generating the next token conditioned on the previously generated tokens and the tactile feature: p(c t | c 1:t−1 , f i ,θ) = TransDec(c 1:t−1 , f i ;θ),(12) where θ denotes the decoder parameters. The sequence is trained to minimize the cross-entropy loss between the pre- dicted and reference captions: L CE (θ) =− 1 T T X t=1 logp(c t | c 1:t−1 , f i ,θ).(13) During training, we apply the teacher forcing strategy, where the ground-truth tokens c 1:t−1 are used as input to predict c t . While the decoder architecture itself is not novel, it plays a vital role in translating temporal tactile features into coherent, perception-aligned textual descriptions. C. Loss Functions The model is trained using a composite loss function that supervises both the tactile feature encoding and the text generation process. The overall objective is defined as: L total =L CE + λ 1 L periodicity + λ 2 L aperiodicity + λ 3 L orthogonality ,(14) where λ 1 ,λ 2 ,λ 3 are hyperparameters controlling the relative importance of each term. During test, ViPAC processes a raw triaxial acceleration sig- nal by computing periodic and aperiodic features, fusing them adaptively, and decoding a textual description conditioned on the fused representation and previously generated tokens. V. EXPERIMENTS A. Experimental Setup Compared Methods. To validate the effectiveness of our ViPAC framework, we compared it with existing methods from audio captioning tasks, including ACT [31], Kim et al. [32], and Recap [33], as well as image captioning methods adapted for vibrotactile signals via Mel-spectrogram conver- sion, such as ClipCap [34], ViECap [35] and RCMF [36]. This cross-domain comparison rigorously evaluates our frame- work’s ability to translate vibrotactile vibration signals into natural language under diverse methodological paradigms. All methods were trained under the same experimental conditions, including the same training, validation, and test data splits. Although ClipCap [34] was not published through a peer- reviewed venue, it has gained popularity due to its novel approach of injecting CLIP-based visual embeddings as prefix tokens into pretrained language models for image captioning. Evaluation Metrics. We adopted five widely used metrics to comprehensively evaluate caption quality. BLEU [37] mea- sures n-gram overlap using the geometric mean of modified precision with a brevity penalty. ROUGE-L [38] computes an F-score based on the longest common subsequence, reflecting structural alignment. METEOR [39] improves recall-oriented evaluation by incorporating synonym and stem matching. CIDEr [40] captures semantic relevance via TF-IDF-weighted cosine similarity of n-grams. SPICE [41] evaluates semantic content by comparing scene graph tuples (objects, attributes, relations). SPIDEr, the arithmetic mean of SPICE and CIDEr, balances semantic accuracy and lexical fidelity. JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20216 TABLE I PERFORMANCE COMPARISON OF DIFFERENT CAPTIONING MODELS. ModelBLEU1BLEU2BLEU3BLEU4ROUGE-LMETEORCIDErSPICESPIDEr ACT (DCASE 2021)0.72040.56190.44980.35100.59310.28820.31070.24540.2780 Kim et al. (ICASSP 2023)0.73650.57320.42780.30860.51630.25970.74280.19420.4685 Recap (ICASSP 2024)0.75240.58910.44150.32130.52780.27140.76520.21260.4889 ClipCap (2021) 0.73430.57920.46670.37480.58190.29850.47370.24800.4959 ViECap (ICCV 2023)0.74320.57640.45730.36820.58670.29460.69810.24130.4697 RCMF (TCSVT 2024)0.74800.59000.45200.33500.58000.27500.72000.22600.4730 ViPAC0.76350.60050.47820.38610.60470.30940.77950.25920.5194 ViPAC: This material surface has evenly spaced slightly raised rounded perforations providing a textured feel. ViPAC: This material surface has a smooth texture with subtle evenly spaced linear grooves. ViPAC: This material surface has a slightly rough texture with fine evenly spaced bumps. ViPAC: This material surface is rough and uneven with small irregular gaps and interwoven texture. GT1: This material surface has evenly spaced rounded perforations with a slightly raised texture at each hole. GT2: This material surface features evenly spaced slightly raised rounded perforations. GT3: This material surface feels smooth with evenly spaced slightly raised rounded perforations. GT4: This material surface has evenly spaced slightly raised rounded perforations. GT5: This material surface is perforated with evenly spaced slightly raised rounded holes. GT1: This material surface feels smooth with a faint uniform grid texture underneath. GT2: This material surface is smooth with a fine material pattern of tiny evenly spaced dots. GT3: This material surface has a fine evenly spaced grid-like texture. GT4: This material surface has a fine grid-like texture subtly raised and evenly spaced. GT5: This material surface has a fine regular pattern with a smooth slightly reflective texture. GT1: This material surface has a rough texture with uniformly distributed small glitter particles. GT2: This material surface is slightly rough with finely distributed raised reflective speckles. GT3: This material surface has fine rough microglitter particles evenly dispersed. GT4: This material surface has fine uneven glitter particles creating a slightly rough texture. GT5: This material surface is rough with small scattered slightly raised irregular speckles. GT1: This material surface has a textured uneven pattern with noticeable indentations and raised areas. GT2: This material surface is uneven textured and slightly rough with minor creases. GT3: This material surface features an irregular bumpy texture with slight depressions throughout. GT4: This material surface has a rough uneven texture with small bumps and indentations. GT5: This material surface features an uneven texture with multiple small indistinct bumps and slight indentation. Fig. 5. Qualitative comparisons between ViPAC generated captions and five GPT-4o reference descriptions for four representative materials. Matched phrases are highlighted to emphasize semantic consistency. The selected samples—covering regular perforations, fine grids, rough glitter, and irregular bumps—demonstrate ViPAC’s ability to produce accurate and diverse textual descriptions directly from vibrotactile signals. TABLE I ABLATION STUDY ON MODEL COMPONENTS. WE EVALUATE THE EFFECTS OF REMOVING THE PERIODIC BRANCH, APERIODIC BRANCH, AND ADAPTIVE FUSION MODULE. THE FULL MODEL CONSISTENTLY ACHIEVES THE HIGHEST PERFORMANCE, CONFIRMING THAT BOTH BRANCHES AND THE DYNAMIC WEIGHTING MECHANISM CONTRIBUTE TO THE EFFECTIVENESS OF VIPAC. VariantBLEU1BLEU2BLEU3BLEU4ROUGE-LMETEORCIDErSPICESPIDEr Periodic Only0.61240.48930.35820.26450.50120.22870.42150.16830.2949 Aperiodic Only0.66570.53280.41270.30610.52370.26140.53620.19470.3655 No Fusion 0.68290.55350.43120.31980.53790.27450.56830.20540.3869 ViPAC (Full)0.76350.60050.47820.38610.60470.30940.77950.25920.5194 Experimental Configuration. All experiments were con- ducted on a Windows 10 64-bit workstation equipped with an Intel Core i7-14700KF CPU (3.40 GHz), 32 GB RAM, and an NVIDIA GeForce RTX 4090 D GPU (24 GB). Model development and training were implemented using Python 3.7.12, PyTorch 2.1.0, and CUDA Toolkit 11.1. B. Experimental Results Qualitative and Quantitative Analysis. We evaluate ViPAC using standard captioning metrics and compare it against baseline methods adapted from image and audio do- mains. As shown in Table I, ViPAC achieves the highest scores across all metrics, with particularly notable gains in CIDEr and SPICE, indicating improved lexical precision and semantic consistency. Fig. 5 illustrates qualitative comparisons between generated captions and GPT-4o-generated references for four representative material types: regular grid, fine mesh, glittery roughness, and bumpy irregularities. ViPAC reliably captures salient tactile features such as smoothness, periodic spacing, and surface irregularity. However, challenging textures occa- JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20217 TABLE I COMPLETE EVALUATION METRICS WHEN A SPECIFIC MATERIAL CATEGORY (G1 TO G9) IS EXCLUDED FROM TRAINING. EACH ROW CORRESPONDS TO A MODEL TRAINED WITHOUT THE RESPECTIVE CATEGORY. Excluded CategoryBLEU1BLEU2BLEU3BLEU4ROUGE-LMETEORCIDErSPICESPIDEr G10.63250.46740.35920.28410.50350.23940.48920.15530.3223 G20.71520.54830.43010.34560.56580.27650.65180.21690.4344 G30.69500.52980.41640.33400.55360.26610.61420.20070.4075 G40.61590.45630.35080.27560.49270.23280.46710.14750.3073 G50.75240.58910.46670.37480.59310.29850.71000.24800.4380 G60.64370.47790.36600.29010.51240.24500.50940.16250.3360 G70.70530.53850.42350.34040.55970.27130.63320.20740.4203 G80.59980.44270.33840.26670.47930.22600.44250.13940.2909 G90.54120.39060.29750.23280.43510.19250.37640.11520.2458 Full Data0.76350.60050.47820.38610.60470.30940.77950.25920.5194 TABLE IV ABLATION ON INPUT REPRESENTATION WITH THE FULL METRIC SUITE. TRAINING THREE INDEPENDENT MODELS WITH ONLY ONE AXIS (X/Y/Z) YIELDS CONSISTENTLY WORSE CAPTIONS THAN USING DFT321 TO FUSE TRIAXIAL SIGNALS INTO ONE CHANNEL. InputB1B2B3B4R-LMETCIDErSPICESPIDEr X-only0.68200.52200.40100.30870.54500.26680.59540.21300.4042 Y-only0.70900.54500.42500.33700.56600.27900.65120.20680.4290 Z-only 0.66400.50600.38900.29240.53800.25830.56270.20450.3836 Mean(X,Y,Z)0.68500.52400.40500.31270.55000.26800.60310.20810.4056 DFT3210.76350.60050.47820.38610.60470.30940.77950.25920.5194 Fig. 6. Demo interface for caption-based material retrieval. The left panel displays matched material images; the right panel shows vibrotactile files. Images are used only for visualization. sionally result in lower semantic alignment, highlighting the need for improved sensitivity to subtle signal variations. These results collectively demonstrate the effectiveness of ViPAC in producing accurate, fluent, and perception-consistent textual descriptions of vibrotactile signals. Demo of ViPAC Application in Retrieval. To demonstrate a practical application of vibrotactile captioning, we designed a web-based demo interface that supports text-based retrieval over material samples. Users can input a keyword or sentence to retrieve all generated captions containing the query. The interface displays the corresponding material name and ref- erence image in the left panel and the matched vibrotactile files in the right panel. This proof-of-concept system operates solely on vibrotactile input; no visual data are used during training or inference. Reference images are included only for visualization, helping users interpret the results. As shown in Fig. 6, the interface supports both keyword and sentence queries and exemplifies the semantic indexing and retrieval scenario introduced earlier. C. Ablation Study Component Ablation. We conduct ablation experiments to assess the impact of each model component. As shown in Ta- ble I, the full ViPAC model achieves the highest performance across all metrics. Removing the dynamic fusion module and using fixed weights leads to noticeable degradation, confirming the effectiveness of adaptive weighting. The aperiodic-only variant performs better than the periodic-only one, indicating that irregular features contain more standalone information. These results validate that periodic and aperiodic features JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20218 are complementary and that dynamic fusion is essential for modeling hybrid tactile structures. Generalization under Missing Categories. To assess gen- eralization, we retrain nine models, each with one material category (G1–G9) held out, and evaluate on that category. Table I reports the full metric suite. Across all held-out cate- gories, ViPAC remains stable. The strongest zero-shot transfer appears on G5, followed by G2 and G7; the weakest is G9, suggesting its textures are less represented by the remaining categories. Trends are consistent across lexical and semantic metrics, indicating that periodic/aperiodic cues learned from other groups transfer broadly. The “Full Data” row provides the upper bound when no category is removed. Justifying DFT321 Axis Fusion. To verify that compress- ing triaxial acceleration into a single magnitude channel via DFT321 is reasonable for captioning, we replace DFT321 with single-axis inputs and train three independent models using only x, only y, or only z signals. All settings are identical to the main model. Table IV reports the full metric suite. DFT321 clearly outperforms any single-axis alternative across lexical (BLEU/ROUGE/METEOR) and semantic met- rics (CIDEr/SPICE/SPIDEr). The best single-axis model (Y- only) still trails DFT321 by 4.9 BLEU4 points and 0.090 SPIDEr; the mean of single-axis runs remains notably lower. VI. CONCLUSION This work introduces vibrotactile captioning, which aims to translate multiple 1D triaxial acceleration signals into natural- language descriptions that convey material surface characteris- tics. To enable this task, we construct a paired vibrotactile-text dataset using an LLM and propose a dual-branch architecture that separately models the periodic and aperiodic components of vibrotactile signals. The proposed method, ViPAC, gener- ates accurate and semantically rich captions, as demonstrated by both quantitative metrics and qualitative comparisons. We further develop a retrieval demo to showcase the usefulness of vibrotactile captioning for caption-based material search. Limitations of this study include the modest dataset size, the lack of large-scale human-authored captions, and the reliance on triaxial acceleration signals only. Future work will expand the dataset with more diverse materials and human-verified descriptions, improve robustness under real-world conditions, and explore multimodal fusion with additional sensory signals. REFERENCES [1] R. Hassen and E. Steinbach, “Vibrotactile signal compression based on sparse linear prediction and human tactile sensitivity function,” in Proc. IEEE World Haptics Conf. (WHC), 2019, p. 301–306. [2] S. Li, Z. Wang, C. Wu, X. Li, S. Luo, B. Fang, F. Sun, X.-P. Zhang, and W. Ding, “When vision meets touch: A contemporary review for visuotactile sensors from the signal processing perspective,” IEEE Journal of Selected Topics in Signal Processing, vol. 18, no. 3, p. 267– 287, 2024. [3] T. Zhao, Y. Fang, K. Wang, Q. Liu, and Y. Niu, “High efficiency vibrotactile codec based on gate recurrent network,” IEEE Transactions on Multimedia, vol. 25, p. 5043–5052, 2023. [4] Y. Li, J.-Y. Zhu, R. Tedrake, and A. Torralba, “Connecting touch and vision via cross-modal prediction,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, p. 10 601– 10 610. [5] O. Holland, E. Steinbach, R. V. Prasad, Q. Liu, Z. Dawy, A. Aijaz, N. Pappas, K. Chandra, V. S. Rao, S. Oteafy, M. Eid, M. Luden, A. Bhardwaj, X. Liu, J. Sachs, and J. Ara ́ ujo, “The ieee 1918.1 “tactile internet” standards working group and its standards,” Proceedings of the IEEE, vol. 107, no. 2, p. 256–279, 2019. [6] B. Wu and Q. Liu, “Integrating point spread function into taxel-based tactile pattern super resolution,” IEEE Transactions on Haptics, vol. 17, no. 4, p. 637–649, 2024. [7] C. Bernard, E. Thoret, N. Huloux, and S. Ystad, “The high/low frequency balance drives tactile perception of noisy vibrations,” IEEE Transactions on Haptics, vol. 17, no. 4, p. 614–624, 2024. [8] M. Strese, L. Bruderm ̈ uller, J. Kirsch, and E. Steinbach, “Haptic Material Analysis and Display Inspired by Human Exploratory Patterns,” 2019. [9] L. Zou, C. Ge, Z. J. Wang, E. Cretu, and X. Li, “Novel tactile sensor technology and smart tactile sensing systems: A review,” Sensors, vol. 17, no. 11, 2017. [10] G. Pati ̃ no-Lakatos, H. Genevois, and B. Navarret, “From vibrotactile sensation to semiotics. mediations for the experience of music,” Hybrid, 2019. [Online]. Available: https://api.semanticscholar.org/CorpusID: 246393140 [11] L. Fu, G. Datta, H. Huang, W. C.-H. Panitch, J. Drake, J. Ortiz, M. Mukadam, M. Lambeta, R. Calandra, and K. Goldberg, “A touch, vi- sion, and language dataset for multimodal alignment,” in Proceedings of the 41st International Conference on Machine Learning, ser. ICML’24. JMLR.org, 2025. [12] N. Cheng, J. Xu, C. Guan, J. Gao, W. Wang, Y. Li, F. Meng, J. Zhou, B. Fang, and W. Han, “Touch100k: A large-scale touch-language-vision dataset for touch-centric multimodal representation,” p. 103305, 2025. [13] Q. Huang, Y. Liang, J. Wei, Y. Cai, H. Liang, H.-f. Leung, and Q. Li, “Image difference captioning with instance-level fine-grained feature representation,” IEEE Transactions on Multimedia, vol. 24, p. 2004– 2017, 2022. [14] M. Al-Qatf, X. Wang, A. Hawbani, A. Abdussalam, and S. H. Alsamhi, “Image captioning with novel topics guidance and retrieval-based topics re-weighting,” IEEE Transactions on Multimedia, vol. 25, p. 5984– 5999, 2023. [15] S. Kornblith, L. Li, Z. Wang, and T. Nguyen, “Guiding image captioning models toward more specific captions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2023, p. 15 259–15 269. [16] Y. Zhong, L. Wang, J. Chen, D. Yu, and Y. Li, “Comprehensive image captioning via scene graph decomposition,” in Computer Vision – ECCV 2020, A. Vedaldi, H. Bischof, T. Brox, and J.-M. Frahm, Eds. Cham: Springer International Publishing, 2020, p. 211–229. [17] Z. Zhang, Z. Qi, C. Yuan, Y. Shan, B. Li, Y. Deng, and W. Hu, “Open- book video captioning with retrieve-copy-generate network,” in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, p. 9832–9841. [18] G. Xu, S. Niu, M. Tan, Y. Luo, Q. Du, and Q. Wu, “Towards accurate text-based image captioning with content diversity exploration,” in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, p. 12 632–12 641. [19] X. Mei, X. Liu, J. Sun, M. D. Plumbley, and W. Wang, “Towards generating diverse audio captions via adversarial training,” IEEE/ACM Trans. Audio, Speech and Lang. Proc., vol. 32, p. 3311–3323, Jun. 2024. [20] S. Deshmukh, B. Elizalde, D. Emmanouilidou, B. Raj, R. Singh, and H. Wang, “Training audio captioning models without audio,” in 2024 International Conference on Acoustics, Speech, and Signal Processing, April 2024. [21] X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang, “Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, p. 3339–3354, 2024. [22] M. Strese, C. Schuwerk, A. Iepure, and E. Steinbach, “Multimodal Feature-Based Surface Material Classification,” IEEE Transactions on Haptics, vol. 10, no. 2, p. 226–239, 2017. [23] M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: data, models and evaluation metrics,” J. Artif. Int. Res., vol. 47, no. 1, p. 853–899, May 2013. [24] K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP). Barcelona, Spain: IEEE, May 2020, p. 736–740. [25] S. Cai, K. Zhu, Y. Ban, and T. Narumi, “Visual-Tactile Cross-Modal Data Generation Using Residue-Fusion GAN With Feature-Matching and Perceptual Losses,” IEEE Rob. Autom, p. 7525–7532, Jul. 2021. JOURNAL OF L A T E X CLASS FILES, VOL. 14, NO. 8, AUGUST 20219 [26] H. Liu, D. Guo, X. Zhang, W. Zhu, B. Fang, and F. Sun, “Toward image-to-tactile cross-modal perception for visually impaired people,” IEEE Trans. Autom. Sci. Eng., p. 521–529, Feb. 2021. [27] Y. Ujitoko and Y. Ban, “Vibrotactile Signal Generation from Texture Im- ages or Attributes Using Generative Adversarial Network,” in Proc. Int. Conf. Hum. Haptic Sens. Touch Enabled Comput. Appl. (EuroHaptics), 2018, p. 25–36. [28] P. Pawlus, R. Reizer, and M. Wieczorowski, “A review of methods of random surface topography modeling,” Tribology International, vol. 152, p. 106530, 2020. [29] N. Landin, J. M. Romano, W. McMahan, and K. J. Kuchenbecker, “Dimensional reduction of high-frequency accelerations for haptic ren- dering,” in Haptics: Generating and Perceiving Tangible Sensations, 2010, p. 79–86. [30] Y. Dong, G. Li, Y. Tao, X. Jiang, K. Zhang, J. Li, J. Su, J. Zhang, and J. Xu, “Fan: Fourier analysis networks,” 2024. [Online]. Available: https://arxiv.org/abs/2410.02675 [31] X. Mei, X. Liu, Q. Huang, M. D. Plumbley, and W. Wang, “Audio captioning transformer,” in Proceedings of the 6th Detection and Clas- sification of Acoustic Scenes and Events 2021 Workshop (DCASE2021), Barcelona, Spain, November 2021, p. 211–215. [32] M. Kim, K. Sung-Bin, and T.-H. Oh, “Prefix tuning for automated audio captioning,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).IEEE, 2023, p. 1–5. [33] S. Ghosh, S. Kumar, C. K. Reddy Evuru, R. Duraiswami, and D. Manocha, “Recap: Retrieval-augmented audio captioning,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, p. 1161–1165. [34] R. Mokady, A. Hertz, and A. H. Bermano, “Clipcap: Clip prefix for image captioning,” arXiv preprint arXiv:2111.09734, 2021. [35] J. Fei, T. Wang, J. Zhang, Z. He, C. Wang, and F. Zheng, “Transfer- able decoding with visual entities for zero-shot image captioning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2023, p. 3136–3146. [36] L. Wang, H. Chen, Y. Liu, and Y. Lyu, “Regular constrained multimodal fusion for image captioning,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 11, p. 11 900–11 913, 2024. [37] K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, P. Isabelle, E. Charniak, and D. Lin, Eds., Jul. 2002, p. 311–318. [38] C.-Y. Lin and F. J. Och, “Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics,” in Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics (ACL-04), Barcelona, Spain, Jul. 2004, p. 605–612. [Online]. Available: https://aclanthology.org/P04-1077/ [39] M. Denkowski and A. Lavie, “Meteor universal: Language specific translation evaluation for any target language,” in Proceedings of the Ninth Workshop on Statistical Machine Translation, O. Bojar, C. Buck, C. Federmann, B. Haddow, P. Koehn, C. Monz, M. Post, and L. Specia, Eds., Jun. 2014, p. 376–380. [40] R. Vedantam, C. L. Zitnick, and D. Parikh, “Cider: Consensus-based image description evaluation,” in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, p. 4566–4575. [41] P. Anderson, B. Fernando, M. Johnson, and S. Gould, “Spice: Semantic propositional image caption evaluation,” in Computer Vision – ECCV 2016, B. Leibe, J. Matas, N. Sebe, and M. Welling, Eds. Cham: Springer International Publishing, 2016, p. 382–398.