Paper deep dive
Decision-Level Fusion for Robust Wearable Affect Recognition
Lokesh Singh, Athina Georgara, Jayati Deshmukh, Tan Viet Tuyen Nguyen, Sarvapali D. Ramchurn
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/8/2026, 3:29:16 PM
Summary
This paper introduces a non-stationary signal processing pipeline combining Fourier-Bessel Series Expansion (FBSE) and Empirical Wavelet Transform (EWT) for wearable affect recognition. It proposes an uncertainty-aware decision-level fusion strategy that weights multi-sensor predictions by predictive entropy and F1 reliability, demonstrating improved robustness over feature-level fusion on the WESAD dataset across baseline, stress, and amusement states.
Entities (14)
Relation Signals (11)
FBSE-EWT โ extractsfeaturesfrom โ EMG
confidence 95% ยท We use five sensors, electrocardiogram (ECG), electrodermal activity (EDA), blood volume pulse (BVP), electromyogram (EMG)... FBSE-EWT signal analysis method
FBSE-EWT โ extractsfeaturesfrom โ EDA
confidence 95% ยท We use five sensors, electrocardiogram (ECG), electrodermal activity (EDA)... FBSE-EWT signal analysis method
FBSE-EWT โ extractsfeaturesfrom โ BVP
confidence 95% ยท We use five sensors, electrocardiogram (ECG), electrodermal activity (EDA), blood volume pulse (BVP)... FBSE-EWT signal analysis method
FBSE-EWT โ extractsfeaturesfrom โ ACC
confidence 95% ยท We use five sensors, electrocardiogram (ECG), electrodermal activity (EDA), blood volume pulse (BVP), electromyogram (EMG) and three-axis acceleration... FBSE-EWT signal analysis method
FBSE-EWT โ extractsfeaturesfrom โ ECG
confidence 95% ยท We use five sensors, electrocardiogram (ECG)... FBSE-EWT signal analysis method
WESAD โ supportsclassificationof โ Amusement
confidence 95% ยท We study this problem on WESAD, using baseline, stress, and amusement conditions
WESAD โ supportsclassificationof โ Stress
confidence 95% ยท We study this problem on WESAD, using baseline, stress, and amusement conditions
Decision-level aggregation โ โ
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Automatic recognition of affective state from wearable physiology has clear societal impact for public health, preventive care, and stress-aware interventions, but real deployments require robustness to non-stationary dynamics, artefacts, and missing sensors. We study this problem on WESAD, using baseline, stress, and amusement conditions, where common fixed-basis spectral features such as FFT bandpower and Welch PSD can oversmooth short-lived discriminative patterns. We propose a non-stationary pipeline that combines Fourier-Bessel Series Expansion (FBSE) with EWT data-driven spectral segmentation to extract mode-wise transient descriptors. For multimodal integration, we adopt decision-level aggregation over per-modality predictors and weight each modality by predictive uncertainty and modality reliability. Results on WESAD, using 15 subjects and ECG, EDA, BVP, EMG, and ACC signals across three classes, indicate that decision-level aggregation is approximately 84 percent of the time at least as good as feature-level aggregation, and approximately 48 percent of the time strictly better, suggesting improved robustness under heterogeneous and partially reliable sensing.
Tags
Links
- Source: https://arxiv.org/abs/2605.14878v1
- Canonical: https://arxiv.org/abs/2605.14878v1
Trouble viewing inline? Open PDF directly โ
Full Text
25,733 characters extracted from source content.
Expand or collapse full text
Decision-Level Fusion for Robust Wearable Affect Recognition Lokesh Singh University of Southampton Southampton, United Kingdom lb5e23@soton.ac.uk Athina Georgara University of Southampton Southampton, United Kingdom ag1g24@soton.ac.uk Jayati Deshmukh University of Southampton Southampton, United Kingdom jd2r24@soton.ac.uk Tan Viet Tuyen Nguyen University of Southampton Southampton, United Kingdom tvtn1c23@soton.ac.uk Sarvapali D. Ramchurn University of Southampton Southampton, United Kingdom sdr1@soton.ac.uk ABSTRACT Automatic recognition of affective state from wearable physiology has clear societal impact for public health, preventive care, and stress-aware interventions, but real deployments require robust- ness to non-stationary dynamics, artefacts, and missing sensors. We study this problem on WESAD (baseline, stress, amusement), where common fixed-basis spectral features (e.g., FFT bandpower and Welch PSD) can oversmooth short-lived discriminative patterns. We propose a non-stationary pipeline that combines FourierโBessel Series Expansion (FBSE) with EWT data driven spectral segmen- tation to extract mode wise transient descriptors. For multimodal integration, we adopt decision-level aggregation over per-modality predictors and weight each modality by predictive uncertainty and modality reliability. On WESAD (15 subjects; ECG, EDA, BVP, EMG, ACC three classes) indicate that decision-level aggregation isโ84% times at least as good as feature level aggregation, andโ48% strictly better, suggesting improved robustness under heterogeneous and partially reliable sensing. KEYWORDS Wearable affect recognition, physiological signals, FourierโBessel series expansion, decision-level aggregation, uncertainty-aware fusion. ACM Reference Format: Lokesh Singh, Athina Georgara, Jayati Deshmukh, Tan Viet Tuyen Nguyen, and Sarvapali D. Ramchurn. 2026. Decision-Level Fusion for Robust Wear- able Affect Recognition. In In 7th International Workshop on Agents for Societal Impact (ASI 2026) @ AAMAS 2026, Paphos, Cyprus, May 25, 2026, IFAAMAS, 5 pages. 1 INTRODUCTION Automatic affective state detection (a personโs current subjective, psychological, and physiological state) from physiological signals is a key problem in AI for health, with applications such as continuous mental-health monitoring, stress-aware interventions, and adaptive humanโcomputer interaction [9,11]. Among affective states, stress and amusement capture complementary high-arousal conditions and enable continuous assessment of emotional well-being using wearable sensing in naturalistic settings [13]. Recent advances in In 7th International Workshop on Agents for Societal Impact (ASI 2026) @ AAMAS 2026, C. Amato, L. Dennis, V. Mascardi, J. Thangarajah (eds.), May 25, 2026, Paphos, Cyprus. ยฉ 2026 International Foundation for Autonomous Agents and Multiagent Systems (w.ifaamas.org). This work is licenced under the Creative Commons Attribution 4.0 International (C-BY 4.0) licence. wearable technology make it possible to collect multimodal signals such as electrodermal activity (EDA), electrocardiography (ECG), electromyography (EMG), blood volume pulse (BVP), and motion (ACC), which reflect autonomic and somatic correlates of affect. Benchmark datasets such as WESAD facilitate systematic evalu- ation of stress and affect recognition under standard protocols, accelerating progress toward deployable affect-aware health sys- tems [14]. Despite this progress, accurate detection remains chal- lenging because physiological time series are noisy, non-stationary, and subject-dependent, and affect-related dynamics often appear as transient, context-dependent changes rather than stable periodic patterns. A major limitation of many pipelines is feature extraction. Com- mon approaches rely on statistical summaries or Fourier-based rep- resentations computed over fixed windows (e.g., FFT bandpower or Welch PSD). These methods implicitly assume approximate local stationarity and can discard time-varying structure that is critical for distinguishing affective state [10]. Timeโfrequency methods such as STFT and wavelets partially mitigate this limitation, but they retain trade-offs between time resolution, frequency resolu- tion, and adaptability to signal-specific structure [16]. Motivated by these limitations, we focus on non-stationarity decomposition meth- ods and, in particular, we propose the FBSE-EWT signal analysis method, which combines Fourier Bessel-series expansion with Em- pirical Wavelet Transform like data-driven segmentation to better preserve transient spectral structure in physiological windows [7]. Beyond limitations in physiological signal analysis, wearable affect monitoring highly relies on multimodal sensing to lever- age complementary physiological mechanisms. Combining various physiological signals (such as ECG, EMG, EDA, etc.) using machine learning has proven to improve the affective state detection [3]. However, this introduces a second major challenge: the robust inte- gration under sensor noise, signal failure, and missing modalities, common issues in real-life deployments [8]. Early fusion [5] ap- proaches and models that assume complete sensor availability can therefore fail under realistic conditions. To address this challenge, we adopt a modular approach in which each modality computes an independent prediction, followed by decision-level aggregation that accounts for predictive uncertainty. This design is suitable for agent-based decision support because it outputs calibrated belief states that can be used to trigger interventions conservatively under uncertainty. In summary, we propose a three-stage pipeline for affective state detection from wearable physiology. First, we extract meaningful arXiv:2605.14878v1 [cs.MA] 14 May 2026 ํ 1 ํ 2 ํ 3 ํ 4 ํ 1 . . . ํ ํ ํ ํ ํ ํ ํ ํ Classifier Classifier Classifier Classifier Classifier Team Formatidon Aggregation Affective state Decision Final Class PredictionFusion Module Multimodal Affective State Identification Figure 1: Wearable affect recognition flow: window- ing, FBSEโEWT features, per-modality prediction, and uncertainty-weighted decision fusion. representations using feature extraction techniques tailored to non- stationary time series. Second, we train single-sensor predictors to classify sensor-specific features into affective state, enabling each modality to independently learn affect-discriminative patterns. Fi- nally, we aggregate the multiple sensor classifiers to reach one final prediction. We do so by weighting each sensorโs contribution ac- cording to its predictive uncertainty in order to enhance robustness to sensor variability and partial failure. We evaluate the proposed framework on WESAD for three-class affect recognition (baseline / stress/ amusement). Our results suggest that non-stationary fea- ture extraction and decision-level aggregation improve reliability in multimodal settings. The key contributions of this paper are as follows: (1)A non-stationary representation for wearable physiology based on FBSEโEWT that preserves transient spectral struc- ture via data-driven mode isolation. (2)A modular per-modality prediction design and an uncertainty- aware, entropy-weighted decision fusion rule that remains stable under noisy or missing modalities. (3)An empirical evaluation on WESAD demonstrating the ben- efit of decision-level aggregation across sensors. 2 RELATED WORK 2.1 Wearable affect detection Wearable affect recognition is commonly framed as supervised classification from multimodal physiological Signals, with WESAD dataset becoming a standard benchmark because it provides syn- chronized chest and wrist modalities and three affective conditions (baseline, stress, amusement) [14]. Recent work explores both clas- sical Machine learning and deep learning approaches. CNN/LSTM- attention style models are frequently reported as strong baselines for multimodal WESAD classification [15], reflecting the value of temporal modelling when affect clues over time. However, cross- subject generalization remains difficult due to subject-dependent physiology and non-stationary responses. Many methods on WE- SAD still depends on representation that are either overly aggre- gated (loosing transient nature) or not designed to be useful across modalities and subjects, leaving space for representation learning that better matches the signal properties of wearable physiology. 2.2 Frequency and time-frequency representation A major line of work relies on frequency domain coefficients, Fourier- domain descriptors and bandpower features are widely used due to simplicity and interpretability. Welchโs method is a standard PSD estimator that reduces variance relative to raw FFT periodograms through averaging of modified periodogram [17]. These estimators are effective when signals are approximately stationary within short windows, but physiological affect responses can be non-stationary and event-like, where averaging may smooth discriminative struc- ture. Time-frequnecy methods partially address this limitation. STFT provides a localised spectrum but requires a fixed window and thus faces inherent timeโfrequency resolution trade-offs [2]. Wavelet decompositions (e.g., DWT) provide multi-resolution anal- ysis and are widely used for physiological signals, but still rely on pre-defined basis structure rather than fully data-adaptive seg- mentation [12]. Fixed basic spectral and time frequency methods impose analysis choices (band, windows, wavelet families) that can underfit the transient and signal specific structure present in wear- able affect data, motivation more adaptive decomposition method that followed the observed spectrum. 2.3Adaptive decompositions for non-stationary physiological signals To better match non-stationary physiology, research has increas- ingly used adaptive decompositions such as empirical mode decom- position (EMD/HHT), Varitonal mode decomposition (VMD). EMD decomposes a signal into intrinsic mode functions that capture time varying oscillatory modes [10]. VMD formulates decomposition as a variational optimisation problem with band-limited modes, offering stronger control over mode bandwidth and noise sensi- tivity [6]. EWT learns wavelet bands from the signalโs spectrum via data driven segmentation, providing an adaptive alternative to fixed wavelet banks [7]. These methods motivate our focus to affect related cues in physiology often seen as transient spectral shifts, where adaptive segmentation can improve separability. In short noisy wearable windows, mode separation can still be unstable, and many decompositions remains sensitive to boundary selection and closely spaced components. moreover, their output are not always converted into fixed dimension coefficient in a way that consistently benefits machine learning model across modalities. 2.4 Multi model fusion Multimodal fusion can improve affect recognition by combining complementary physiological cues, as supported by benchmark settings such as WESAD. However, wearable data in practice is affected by artefacts and occasional modality failure, so fusion strategies that operate at the decision level are often preferred over approaches that assume all modalities are always available [4,8]. In this context, uncertainty aware weighting (e.g., using predic- tive entropy) provides a principled way to down weight unreliable modality predictions and achieve graceful degradation when signals are missing or noisy [1]. Motivated by these gaps, our approach targets two limitations in wearable affect detection. Representation quality under non- stationary, transient physiological dynamics, and robustness of mul- timodal integration under missing or unreliable modalities. We in- troduce an adaptive transient-preserving representation (FBSEโEWT) and couple it with uncertainty and reliability weighted decision- level fusion for multimodal inference in realistic sensing conditions. 3 AFFECTIVE STATE DETECTION: A THREE-STAGE PIPELINE 3.1 Feature Extraction 3.1.1 Problem formulation and windowing. Let subjects be indexed byํ โSand modalities byํ โ 1, . . .,ํ. Each modality provides a discrete time physiological signal ํฅ ํ ,ํ [ํ], ํ= 0, . . .,ํ ํ ,ํ โ 1.(1) Signals are segmented into overlapping windows of fixed duration ํฟwith overlap ratioํผ=0.75. With sampling rateํ ํ , the window length isํ= ํฟํ ํ samples and the hop size isํป=(1โ ํผ)ํ. The ํ -th window is defined as ํฅ (ํ) ํ ,ํ [ํ]= ํฅ ํ ,ํ [ํํป +ํ], ํ= 0, . . .,ํ โ 1.(2) The WESAD label stream is provided at 700 Hz. Each window receives a label by majority voting with purity threshold ํ : ํฆ (ํ) ํ = arg max ํโC โ๏ธ ํกโI ํ 1โ ํ [ํก]=ํ,(3) max ํโC ร ํกโI ํ 1โ ํ [ํก]=ํ |I ํ | โฅ ํ,(4) whereC=Baseline, Stress, AmusementandI ํ denotes the cor- responding label index range. 3.1.2 Fourier Bessel representation. Given a windowed signalํฆ[ํ] of lengthํ= ํ, the zero-order FBSE representsํฆusing Bessel bases: ํฆ[ํ]= ํ โ๏ธ ํ=1 ํถ ํ ํฝ 0 ํฝ ํ ํ ํ , ํ= 0, . . .,ํ โ 1,(5) with coefficients ํถ ํ = 2 ํ 2 ํฝ 1 (ํฝ ํ ) 2 ํโ1 โ๏ธ ํ=0 ํํฆ[ํ] ํฝ 0 ํฝ ํ ํ ํ ,(6) whereํฝ ํ is theํ-th positive root ofํฝ 0 (ํฝ)=0. Ordersํmap to physical frequencies via the approximation ํฝ ํ โ 2ํํ ํ ํ ํ ํ , โ ํ โ 2ํ ํ ํ ํ ํ ,(7) so the FBSE spectrum is|ํถ ํ | as a function of ํ ํ . This basis is advantageous for wide-band, non-stationary sig- nals because Bessel bases exhibit AM-like behaviour and can yield compact spectral representations for AFM like components. EWT constructs empirical wavelet filters whose supports de- pend on the signalโs spectral content. Boundaries are obtained by locating meaningful minima in a spectrum derived histogram using a scale space persistence criterion, and an automatic threshold (e.g., Otsu) selects which minima are retained. Let the resulting ordered boundaries be 0= ํ 0 < ํ 1 < ยท< ํ ํ = ํ.(8) Given a transition parameterํ, empirical scaling and wavelet func- tions in the Fourier domain are defined piecewise. These filters form a tight frame under suitableํ, and EWT coefficients are computed by inner products with the empirical wavelets/scaling function, reconstruction follows from the same frame. In FBSEโEWT, boundary detection and filter-bank construction operate on the FBSE spectrum rather than the conventional FFT spectrum, improving mode separation in challenging cases (e.g., closely spaced components, short-duration activity). 3.1.3 Feature construction from modes. Letํพ= ํbe the number of resulting modes for windowํฅ (ํ) ํ ,ํ , and letํ (ํ) ํ ,ํ,ํ [ํ]denote theํ-th reconstructed mode. A fixed-dimensional feature vectorํ ํ ํฅ (ํ) ํ ,ํ is formed by aggregating per-mode descriptors. In particular, the mode energy and log-energy are ํธ ํ = 1 ํ ํโ1 โ๏ธ ํ=0 ํ (ํ) ํ ,ํ,ํ [ํ] 2 ,(9) log ( ํธ ํ +ํ ) ,(10) and the mode entropy is computed from the normalised energy distribution ํ ํ [ํ]= ํ (ํ) ํ ,ํ,ํ [ํ] 2 ร ํโ1 ํข=0 ํ (ํ) ํ ,ํ,ํ [ํข] 2 ,(11) as ํป ํ =โ ํโ1 โ๏ธ ํ=0 ํ ํ [ํ] logํ ํ [ํ].(12) This representation preserves transient spectral structure while remaining compatible with standard supervised learning. 3.2 Single-Sensor Prediction The physiological features extracted from each modality are fed into a single-sensor predictor (SSP) to compute the probability of the predicted classesํ ํ along with the predictorโs confidence score ํน1 ํ . In this study, we adopt a Multi-Layer Perceptron (MLP) as the predictor, denoted asํํ ํํฟํ (S ํ ), to model non-linear relationships in the feature data collected from a single sensorS i as presented in Eq. 13. The choice of MLP is motivated by the nature of the input features, which provide fixed-dimensional representations that al- ready encode transient spectral dynamics. This enables effective learning without the need for explicit temporal models such as CNNs or LSTMs, while maintaining computational efficiency and robustness for relatively small datasets. In our designed classifier, the MLP consists of multiple fully con- nected layers with ReLU activation, followed by a softmax output layer that produces class probabilities. Dropout and L2 regularisa- tion are applied to mitigate overfitting. The model is trained using categorical cross-entropy loss and optimised with Adam, with early stopping based on validation performance. We also explore varia- tions in the number of hidden layers and select the configuration that achieves the highest prediction accuracy. The outputํ ํ is in- terpreted as a probabilistic estimate of the affective state, enabling uncertainty quantification via entropy, while the confidence score ํน1 ํ represents the overall reliability of the sensor-specific predic- tor. These two outputs are subsequently used in the multi-sensor aggregation phase discussed in the following section. ํ ํ ,ํน1 ํ โ ํํ ํํฟํ (S ํ ) โํ โ 1, 2, . . .,ํ(13) 3.3 Multi-Sensor Aggregation Wearable affect recognition benefits from combining complemen- tary physiological cues, but in realistic settings different modalities can vary in quality across time due to motion artefacts, skin con- tact variation, or temporary sensor failure. We therefore perform aggregation at the decision level using a weighted average of per- modality probability vectors. The information can be aggregated at the data, feature, or decision level [8], and each approach has its challenges and benefits. In this work, we propose to perform the aggregation on decision-level, and use a weighted average of the individual sensorsโ decisions. Letํ โ 1,2, . . .,ํbe a team of sensors, and each sensor feeds data to a single-sensor predictor. The decision ํ ํ yielded by the SSP operating on sensor ํ โ ํ , con- tributes to the final aggregated decision according to its entropy and to the confidence score of the SSP. Specifically, we obtain the aggregated decision as: ํ ํ (ํ)= 1 ํพ โ๏ธ ํโํ 1โ ํป(ํ ํ ) ํน1 ํ ยท ํ ํ (ํ) โ ํ โ classes(14) whereํis a team of sensors,ํ ํ is the probability distribution over the classes obtained by sensorํ โ ํ;ํป(ยท)denotes the information entropy function;ํน1 ํ is the F1 score of the single-sensor predictor of sensorํ โ ํ; andํพ= ร ํโํ ํป(ํ ํ ) ํน1 ํ is a normalisation factor. Intuitively, the more certain an SSP is (i.e., the lower the entropy ํป(ํ ํ )), the higher is its contribution to the final decision. Similarly, the more confident the SSP is (i.e., higher confidence scoreํน1 ํ ), the more the decision is valued. Notably, the entropy is predicted- instance specific, while the confidence score is descriptive of the SSP in general. As such, using the entropy as the base and the confidence score as the power, we regulate the influence of each individual decision, prioritising primarily according to the entropy, and secondly according to the confidence score. 4 IMPLEMENTATION AND RESULTS We performed some preliminary evaluation using the WESAD dataset [14] for three-class affect recognition (baseline, stress, amuse- ment). We use five sensors, electrocardiogram (ECG), electrodermal activity (EDA), blood volume pulse (BVP), electromyogram (EMG) and three-axis acceleration, for 15 subjects. In these experiments, we compare the decision-level aggregation against the feature-level one. Specifically, Figure 2 shows the overall accuracy in correctly predicting an class when the aggregation was performed at the feature (F) and decision (D) levels. We note that inโ84% cases, decision aggregation is as good as or better than feature aggrega- tion. Next, we explore the impact of the team size on the accuracy. That is, instead of considering all five sensors, we test the accuracy of a team consisting of varying team-size combinations of sensors D>F 48.39% D=F 35.48% D<F 16.13% Figure 2: Overall Accuracy all quartet triad pair single 0.6 0.8 1 accuracy Figure 3: Team level accuracy (including teams of a single sensor as well). Figure 3 shows the comparison of feature-level and decision-level aggregation of all the team sizes. Notably, regardless of the team size, decision-level aggregation is better. Additionally, we observe that when more sensors participate in the aggregation results in better accuracy. To conclude, our preliminary results indicate that the proposed decision-level aggregation achieves better accuracy than the typical feature-level one. Decision-level aggregation is more generic and can be used to aggregate a diverse set of sensors without expert knowledge regarding specific signal features. It is also robust, since if one of the sensors fail, the decisions of the other sensors can still be aggregated. 5 CONCLUSIONS AND FUTURE WORK This paper presented a three-stage pipeline for wearable affect recognition that combines a non-stationary representation (FBSE โ EWT) with uncertainty and reliability weighted decision-level fusion. The proposed approach is motivated by deployment realities in societally important health and care settings, where physiologi- cal sensing is noisy, non-stationary, and occasionally incomplete, and where downstream systems should act on calibrated beliefs rather than overconfident predictions. Preliminary experiments on WESAD (baseline, stress, amusement) indicate that decision- level fusion is competitive with, and often superior to, feature-level fusion across sensor-team sizes, suggesting improved robustness under heterogeneous sensing conditions. In future, we plan to work and extend this line of work in some of the following ways: (i) use multi-modal input, like physiologi- cal sensors, video, audio, images and natural language input, (i) augment the data by using data generators for different sensors, (i) detect a variety of relevant classes, (iv) deploy the model in different types of care homes. (v) process temporal information, (vi) communication and coordination among diverse robots from different manufacturers and having a different set of sensors. We also plan to validate on additional datasets and integrate the result- ing uncertainty-aware affect estimates into agent-based decision support pipelines for stress-aware interventions and continuous wellbeing monitoring. ACKNOWLEDGMENTS This work is supported by Responsible Ai UK (RAi UK) (EP/Y009800/1) coordinating keystone project on โEmbodied AI in Social Spaces: Responsible and Adaptive Robots in Complex Settingsโ. REFERENCES [1]Jose Ignacio Aizpurua, Victoria M Catterson, Brian G Stewart, Stephen DJ McArthur, Brandon Lambert, and James G Cross. 2018. Uncertainty-aware fusion of probabilistic classifiers for improved transformer diagnostics. IEEE Transactions on Systems, Man, and Cybernetics: Systems 51, 1 (2018), 621โ633. [2]Jont B Allen and Lawrence R Rabiner. 2005. A unified approach to short-time Fourier analysis and synthesis. Proc. IEEE 65, 11 (2005), 1558โ1564. [3]Deฤer Ayata, Yusuf Yaslan, and Mustafa E Kamasak. 2020. Emotion recognition from multimodal physiological signals for emotion aware healthcare systems. Journal of Medical and Biological Engineering 40, 2 (2020), 149โ157. [4]Tadas Baltruลกaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2018. Multi- modal machine learning: A survey and taxonomy. IEEE transactions on pattern analysis and machine intelligence 41, 2 (2018), 423โ443. [5]Jing Chen, Bin Hu, Lixin Xu, Philip Moore, and Yun Su. 2015. Feature-level fusion of multimodal physiological signals for emotion recognition. In 2015 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). IEEE, 395โ399. [6] Konstantin Dragomiretskiy and Dominique Zosso. 2013. Variational mode de- composition. IEEE transactions on signal processing 62, 3 (2013), 531โ544. [7]Jerome Gilles. 2013. Empirical wavelet transform. IEEE transactions on signal processing 61, 16 (2013), 3999โ4010. [8] Raffaele Gravina, Parastoo Alinia, Hassan Ghasemzadeh, and Giancarlo Fortino. 2017. Multi-sensor fusion in body sensor networks: State-of-the-art and research challenges. Information Fusion 35 (2017), 68โ80. [9]Jennifer A Healey and Rosalind W Picard. 2005. Detecting stress during real- world driving tasks using physiological sensors. IEEE Transactions on intelligent transportation systems 6, 2 (2005), 156โ166. [10]Norden E Huang, Zheng Shen, Steven R Long, Manli C Wu, Hsing H Shih, Quanan Zheng, Nai-Chyuan Yen, Chi Chao Tung, and Henry H Liu. 1998. The empirical mode decomposition and the Hilbert spectrum for nonlinear and non- stationary time series analysis. Proceedings of the Royal Society of London. Series A: mathematical, physical and engineering sciences 454, 1971 (1998), 903โ995. [11]Sylvia D Kreibig. 2010. Autonomic nervous system activity in emotion: A review. Biological psychology 84, 3 (2010), 394โ421. [12] Stephane G Mallat. 1989. Multifrequency channel decompositions of images and wavelet models. IEEE Transactions on Acoustics, speech, and signal processing 37, 12 (1989), 2091โ2110. [13] Rosalind W Picard. 2000. Affective computing. MIT press. [14] Philip Schmidt, Attila Reiss, Robert Duerichen, Claus Marberger, and Kristof Van Laerhoven. 2018. Introducing wesad, a multimodal dataset for wearable stress and affect detection. In Proceedings of the 20th ACM international conference on multimodal interaction. 400โ408. [15] Ritu Tanwar, Orchid Chetia Phukan, Ghanapriya Singh, Pankaj Kumar Pal, and Sanju Tiwari. 2024. Attention based hybrid deep learning model for wearable based stress recognition. Engineering Applications of Artificial Intelligence 127 (2024), 107391. [16] Zhiguang Wang and Tim Oates. 2015. Imaging time-series to improve classifica- tion and imputation. arXiv preprint arXiv:1506.00327 (2015). [17] Peter Welch. 2003. The use of fast Fourier transform for the estimation of power spectra: A method based on time averaging over short, modified periodograms. IEEE Transactions on audio and electroacoustics 15, 2 (2003), 70โ73.