Paper deep dive
FUSE: Frame-Unified Stress Estimation from Facial Video
Stefanos Gkikas, Thomas Kassiotis, Yang Guo, Guangliang Li, Giorgos Giannakakis
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 8/13/2026, 5:33:18 AM
Summary
The paper introduces FUSE (Frame-Unified Stress Estimation), a deep learning framework for automatic stress detection from facial video. Unlike traditional methods that segment videos into short temporal windows, FUSE processes complete recordings as a single input by folding the temporal dimension into the channel dimension. It utilizes an asymmetric-attention architecture to handle the resulting high-dimensional spatial representation. Experiments on a 58-subject dataset demonstrate that FUSE achieves competitive accuracy (up to 69.44%) while offering a trade-off between temporal density and computational efficiency.
Entities (6)
Relation Signals (5)
FUSE → achievesaccuracy → 69.44%
confidence 95% · FUSE achieves the highest test accuracy of 69.44% at t = 15
FUSE → uses → Asymmetric Attention
confidence 95% · the resulting high-dimensional input is processed using a unified asymmetric-attention architecture.
FUSE → employs → Axis Folding
confidence 92% · This unification is realized by folding the temporal dimension into the channel dimension of the spatial representation
FUSE → evaluatedon → Stress Dataset
confidence 90% · Experiments on a 58-subject stress dataset using a stratified subject-level protocol evaluate seven temporal-stride configurations
FUSE → adaptsprinciplefrom → Perceiver
confidence 85% · FUSE instead adapts the asymmetric-attention principle of the Perceiver
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Automatic stress detection from facial video offers a practical path to non-intrusive affect monitoring, yet existing video-based approaches commonly decompose full recordings into short temporal windows before classification. This design introduces additional choices regarding window length, overlap, and aggregation, while limiting direct analysis of temporal information across the entire recording. In this study, we present FUSE (Frame-Unified Stress Estimation), a facial-video stress detection framework that processes complete recordings as a single input without temporal windowing or external segmentation. The name reflects the defining operation of the method: rather than dividing a recording into short clips, all frames are fused into one unified two-dimensional representation from which the stress state is estimated. This unification is realized by folding the temporal dimension into the channel dimension of the spatial representation, and the resulting high-dimensional input is processed using a unified asymmetric-attention architecture. At a temporal stride of t = 1, FUSE retains the full 120-second recording as one input, corresponding to 3,600 frames at 30 fps. Experiments on a 58-subject stress dataset using a stratified subject-level protocol evaluate seven temporal-stride configurations, ranging from full-frame input to sparse subsampling. FUSE achieves the highest test accuracy of 69.44% at t = 15, while the full-frame configuration remains competitive at 69.03%. Across the stride range, computational cost varies from 12.48 to 348.78 GFLOPs, showing the trade-off between temporal density and efficiency. These results demonstrate that temporal windowing is not required for effective facial-video stress detection in this setting, and that complete-recording inference can be achieved within a single unified architecture.
Tags
Links
- Source: https://arxiv.org/abs/2608.10442v1
- Canonical: https://arxiv.org/abs/2608.10442v1
Trouble viewing inline? Open PDF directly →
Full Text
37,092 characters extracted from source content.
Expand or collapse full text
FUSE: Frame-Unified Stress Estimation from Facial Video Stefanos Gkikas Honda Research Institute Japan Wako City, Japan stefanos.gkikas@jp.honda-ri.com Thomas Kassiotis Department of Electronic Engineering Hellenic Mediterranean University Chania, Greece ddk305@edu.hmu.gr Yang Guo Faculty of Information Science and Engineering Ocean University of China Qingdao, China diw85827@gmail.com Guangliang Li Faculty of Information Science and Engineering Ocean University of China Qingdao, China guangliangli@ouc.edu.cn Giorgos Giannakakis Department of Electronic Engineering Hellenic Mediterranean University Chania, Greece ggian@hmu.gr Abstract—Automatic stress detection from facial video offers a practical path to non-intrusive affect monitoring, yet existing video-based approaches commonly decompose full recordings into short temporal windows before classification. This design introduces additional choices regarding window length, overlap, and aggregation, while limiting direct analysis of temporal information across the entire recording. In this study, we present FUSE (Frame-Unified Stress Estimation), a facial-video stress detection framework that processes complete recordings as a single input without temporal windowing or external segmentation. The name reflects the defining operation of the method: rather than dividing a recording into short clips, all frames are fused into one unified two-dimensional representation from which the stress state is estimated. This unification is realized by folding the temporal dimension into the channel dimension of the spatial representation, and the resulting high-dimensional input is processed using a unified asymmetric-attention architecture. At a temporal stride ofτ = 1, FUSE retains the full120-second recording as one input, corresponding to3,600frames at30fps. Experiments on a58-subject stress dataset using a stratified subject-level protocol evaluate seven temporal-stride configurations, ranging from full-frame input to sparse subsampling. FUSE achieves the highest test accuracy of69.44%atτ = 15, while the full- frame configuration remains competitive at69.03%. Across the stride range, computational cost varies from12.48to348.78 GFLOPs, showing the trade-off between temporal density and efficiency. These results demonstrate that temporal windowing is not required for effective facial-video stress detection in this setting, and that complete-recording inference can be achieved within a single unified architecture. Index Terms—Stress recognition, mental health, affective computing, transformer I. INTRODUCTION Stress arises as a coordinated physiological and psychological response to perceived demands, engaging the autonomic nervous system and triggering neuroendocrine cascades that vary in intensity and duration [1], [2]. Its manifestations range from brief, situationally bounded episodes to prolonged chronic states, each associated with distinct physiological signatures and long-term health consequences. Questionnaire- based instruments such as the Perceived Stress Scale are widely used in clinical and research settings to quantify subjective stress burden [3], but retrospective self-reports are susceptible to recall bias and poorly suited to capturing fine-grained temporal variation in stress levels [4]. Salivary cortisol serves as a validated neuroendocrine index of the stress response, yet the invasive nature of sample collection and the delayed temporal dynamics of cortisol secretion limit its applicability for continuous, real-time monitoring [2]. The scale of stress as a public health problem has grown substantially over recent decades. A large-scale analysis of nationally representative survey data spanning146countries found that reported stress levels roughly doubled over an18- year period, with disparities widening across demographic and socioeconomic groups [5]. In occupational contexts, psychoso- cial work-related exposures have been estimated to account for a measurable share of cardiovascular disease and depression cases across European countries [6]. At the biological level, chronic psychological stress has been linked to disruption of immune regulation, elevated cardiovascular risk, and depressive disorder, underscoring the broader clinical relevance of reliable stress monitoring [7]. Reliable stress monitoring is most valuable precisely in the contexts where existing assessment methods are least practical. In clinical, occupational, and operational settings, self- report is often delayed, incomplete, or systematically biased by social desirability and demand characteristics. Contact- based wearable devices provide a continuous physiological alternative, but face persistent challenges around user com- pliance, motion-induced signal degradation, and scalability to diverse deployment environments [8]. Field deployments of such systems have consistently identified data quality and signal integrity as limiting factors that constrain generalization beyond laboratory conditions [9]. These considerations motivate the development of passive, non-intrusive automated systems arXiv:2608.10442v1 [cs.CV] 11 Aug 2026 capable of objectively recognizing stress from signals naturally available in everyday environments [10]. Among non-intrusive modalities, facial video recorded by standard cameras offers a particularly practical channel for automated stress monitoring, requiring no physical contact, specialized hardware, or user compliance beyond proximity to a camera. Stress-related changes in facial behavior, including eye activity, mouth movements, and head dynamics, carry discriminative information for separating stressed from neutral and relaxed states [11], and Facial Action Unit representations provide a structured encoding of these signals for automatic recognition [12]. Deep learning has substantially advanced the state of the art in video-based stress recognition, with spatiotemporal architectures and AU-based pipelines achieving consistent improvements over handcrafted feature approaches [13]. A shared limitation of existing methods, however, is their reliance on fixed-length temporal windows or short clips as the unit of analysis: full recordings are partitioned into segments of a few seconds before classification, introducing additional design choices and discarding the temporal context that spans the window boundaries [14]–[16]. In this work, we present FUSE, a framework for automatic stress detection in facial video that processes entire recordings as a single input, without any temporal windowing or external segmentation. The full video sequence is encoded through axis folding and processed by a unified asymmetric attention model, whose internal spatial token segmentation handles the resulting high-dimensional representation without imposing any temporal partitioning on the input. Experiments on a 58-subject stress dataset, evaluated under a stratified subject- level protocol, assess performance across temporal stride configurations ranging from dense sampling to full-frame input, demonstrating that the architecture natively accommodates complete recordings without modification. Automatic human- state recognition has been investigated across a range of signals, including stress and pain estimation from electrodermal activity and other biosignals [17]–[20], emotion recognition from EEG [21], and cognitive workload assessment [22], complementing the facial-video stress detection addressed here. I. RELATED WORK Video-based approaches to stress recognition have been explored as a contact-free alternative to physiological mon- itoring, leveraging the behavioral information encoded in facial dynamics. Initial methods extracted handcrafted features from facial regions, including eye-related activity, mouth movements, head-motion statistics, and camera-based heart rate estimates, demonstrating their discriminative power in distinguishing stressed from neutral and anxious states [11]. Deep learning methods replaced handcrafted pipelines by jointly learning face- and action-level representations from video, achieving improved accuracy over feature-engineering baselines on a purpose-built video stress dataset [14]. Spatiotemporal architectures extended this direction by modeling both spatial regions and temporal changes in facial appearance, using fixed- length clips as the input unit for learning stress-related facial dynamics [15]. Beyond laboratory-induced stress protocols, recent work has investigated facial video-based stress detection under naturalistic conditions, collecting data without artificial contextual constraints to improve ecological validity [23]. Facial Action Units provide a physiologically grounded encoding of facial muscle activity and have been adopted as an interpretable intermediate representation for automatic stress recognition from video. AU-based classifiers have been applied to distinguish stressed from neutral states in a contact-free setting, with specific action unit patterns consistently emerging as discriminative across subjects and stressor types [12]. Deep learning pipelines for automatic AU recognition have been adapted to the stress domain, combining geometric deformation features with deep appearance descriptors extracted from facial video to classify affective states [24]. Explainable AI methods have been integrated into AU-based stress recognition to identify the facial muscle activations that are most predictive of stress, thereby providing interpretable evidence alongside model outputs [25]. An explainable graph attention network operating on differential facial action units has been proposed for stress recognition [26]. Comprehensive surveys of deep learning for stress detection identify facial video analysis as a prominent research direction, with video-based approaches showing consistent improvements alongside advances in spatiotemporal representation learning [13]. A common design choice in video-based stress recognition is to convert continuous recordings into predefined temporal units before classification. For example, Zhang et al. [14] partitioned each 2-min recording into 15-s samples, while Jeon et al. [15] modeled stress from short 2-s facial clips. More recent work has followed a similar temporal-reduction strategy, including non- overlapping windows for Transformer-based stress estimation [16], or frame-level facial feature extraction followed by temporal and frequency-domain aggregation [23]. Although these strategies make learning computationally tractable, they introduce additional design choices, such as clip duration, window length, overlap, and aggregation strategy. These choices are dataset- and context-dependent and may limit the model’s ability to exploit temporal information that spans predefined boundaries. Recent surveys confirm the growing role of deep learning in stress detection across facial, behavioral, physiological, and multimodal signals [13], but full-recording video analysis without external temporal windowing remains underexplored. Efficient long-video modeling has been widely studied through spatiotemporal token designs, such as the tubelet embedding of ViViT [27], which nonetheless retain an explicit temporal axis. FUSE instead adapts the asymmetric-attention principle of the Perceiver [28], folding time into the channel dimension so that the full recording is processed as a single spatial representation rather than a growing spatiotemporal to- ken set. Transformer-based and modality-agnostic architectures have similarly been employed for affective assessment from facial video and physiological signals [29]–[32]. I. METHODOLOGY A. Video Tokenization The proposed framework encodes facial video input into a token representation without temporal segmentation or modality-specific components, enabling a single model to process complete recordings of varying length. For a video ofTframes captured at30fps and subsampled at temporal strideτ, the retained sequence containsL = ⌊T/τ⌋frames, each a224 × 224RGB image. At a strideτ = 1applied to a120-second recording, this yieldsL = 3,600frames, corresponding to the complete recording without any external windowing or temporal segmentation. The temporal dimension is folded into the channel dimension—a step referred to as axis folding—creating a tensor of shapeH × W × 3L. This preserves the spatial structure of each frame while packing the entire temporal sequence into a single 2D representation: X∈R B×H×W×3L ,(1) whereBdenotes the batch size andH = W = 224. Geometric information is incorporated by encoding each spatial position p ∈ [−1, 1] 2 using Fourier features. The model usesK = 6 frequency bands and a maximum frequencyf max = 10. Since the input hasD = 2spatial axes, the Fourier encoding adds D(2K + 1) = 26 positional features. The encoding is: γ(p) = sin(πs 1 p), cos(πs 1 p), ..., sin(πs K p), cos(πs K p), p , (2) wheres k K k=1 spans[1,f max /2]. The spatial axes are flattened into a sequence ofN = H × W = 50176tokens, with data channels and positional features concatenated per token to form the token matrix: T∈R B×N×C ′ , C ′ = 3L + D(2K + 1),(3) whereD = 2denotes the number of spatial axes. SinceD = 2 and K = 6, the token dimension becomes: C ′ = 3L + 26.(4) The token sequence is partitioned intoS = 4contiguous spatial groups of lengthn s = N/S = 12544tokens. These spatial token groups, referred to as segments throughout, divide the folded 2D representation along the token axis and carry no correspondence to temporal windows of the input video. The tokens of spatial segment s are denoted ̃ T s ∈R B×n s ×C ′ . B. Asymmetric Attention The model processes the segmented token sequence through four layers, each comprising a cross-attention block followed byR ℓ self-attention blocks. A single latent state is associated with each spatial segment, instantiated at runtime by replicating a shared initialization vector derived from a set ofM 0 = 32 learnable global parameters ℓ m M 0 m=1 : ℓ init = 1 M 0 M 0 X m=1 ℓ m ∈R d 0 ,(5) whered 0 = 128. These segment states are not independently learnable; segment-specific representations emerge through the attention updates. Cross-attention. At each layerℓ, each segment state ag- gregates information exclusively from its corresponding token subset through cross-attention: e (ℓ) s = e (ℓ−1) s + Attn e (ℓ−1) s , ̃ T s ,(6) wheree (ℓ−1) s ∈R B×1×d ℓ provides the queries and ̃ T s ∈ R B×n s ×C ′ provides the keys and values. This operation is asymmetric: the query side consists of a single vector of dimensiond ℓ , while the key-value side spansn s ≫ 1token vectors of dimensionC ′ . The resulting attention matrix is 1× n s , not square, and the query and key-value spaces differ in both size and dimensionality. AllSsegments are processed in parallel by packing into the batch dimension, without altering the underlying computation. Cross-attention uses a single head at all layers, with per-layer head dimensions of64, 48, 32, 16. Self-attention. After cross-attention, all segment states are stacked to form the segment-state matrixE (ℓ) ∈R B×S×d ℓ . Self-attention is applied across allSsegment states, enabling global information exchange: E (ℓ) ← E (ℓ) + Attn E (ℓ) , E (ℓ) ,(7) repeatedR ℓ ∈8, 6, 4, 2times per layer forℓ = 0,..., 3. Self- attention uses multi-head attention with per-layer head counts of 8, 6, 4, 2 and per-layer head dimensions of 64, 48, 32, 16. Hierarchical segment-state compression. Across the four layers, the segment-state dimensionality decreases progressively asd ℓ ∈ 128, 112, 96, 80. At each layer transition, the segment representations are projected to the new dimensionality through a linear transformation when required. The number of spatial segment states remains fixed atS = 4throughout all layers, with one state per token group. After the final layer, the final segment statesE (4) ∈R B×S×80 are averaged across segments and passed through a linear classification head. Both attention operations use pre-layer normalization and residual connections, with attention and feedforward dropout of0.10applied uniformly. The complete set of architectural hyperparameters is reported in Table I and the layer structure is illustrated in Fig. 1. IV. EXPERIMENTAL EVALUATION & RESULTS This section presents the experimental evaluation of FUSE across multiple temporal stride configurations. All experiments are evaluated under a binary classification setting. Validation performance is reported using macro-averaged accuracy, preci- sion, and F1 score. Test performance is reported using macro- averaged accuracy. A. Dataset and Protocol A stress dataset comprising58adults (24men,34women) aged26.9 ± 4.8years was used in this study. The experi- mental protocol comprised four stress-induction phases: social exposure, emotional recall, mental workload, and stressful video stimuli. Each participant completed11tasks in total, TABLE I: Architectural hyperparameters of FUSE. HyperparameterValue Depth4 Latent pool size (M 0 )32 Latent dimension (d ℓ )128, 112, 96, 80 Cross-attention heads1, 1, 1, 1 Cross-attention head dimension64, 48, 32, 16 Self-attention heads8, 6, 4, 2 Self-attention head dimension64, 48, 32, 16 Self-attention blocks per cross (R ℓ )8, 6, 4, 2 Spatial segments (S)4 Attention dropout0.10 Feedforward dropout0.10 Fourier frequency bands (K)6 Maximum frequency (f max )10 Per-layer values are listed from layer 1 to layer 4. + + + + Input LayerNorm Cross-Attention LayerNorm LayerNorm LayerNorm Self-Attention FFN FFN ×4 Fig. 1: FUSE layer block, repeated four times. Each layer comprises a cross-attention sub-block, in which a single latent segment state attends over its corresponding spatial token group, followed byR ℓ self-attention sub-blocks that exchange information globally across all segment states. All sub-blocks use pre-layer normalization and residual connections. comprising4neutral,6stress-inducing, and1relaxation task, as detailed in Table I. The social exposure phase included a psychologist-led interview emphasizing negative personality traits. The emotional recall phase required participants to relive a past stressful event in real time. Mental workload was induced through a modified Stroop Color-Word Test [33] and the Paced Auditory Serial Addition Test [34]. The stressful stimuli phase presented videos depicting accidents and acrophobia; a relaxing video served as a physiological recovery baseline between induction phases to minimize carryover effects and was excluded from classification. Stress induction was verified through heart rate monitoring, which showed a statistically significant increase during stress tasks (p < 0.05). Subjec- tive validation was further confirmed using Self-Assessment Manikin scales, with participants reporting significantly higher TABLE I: Experimental tasks employed in this study. #TaskDuration (sec)State Social Exposure 1Neutral reference120N 2Baseline description120N 3Interview120S Emotional Recall 4Neutral reference120N 5Recall stressful event120S Mental Workload 6Reading reference120N 7Stroop Colour-Word Test120S 8PASAT task120S Stressful Stimuli 9Relaxing video ∗ 120R 10Adventure video120S 11Psychological pressure120S N = neutral S = stressR = relaxed. *: Used as a physiological recovery baseline between induction phases; excluded from binary classification. arousal and lower valence during stressful phases compared with neutral baselines. The facial video was recorded at60fps and subsampled to30fps with a resolution of608 × 800 pixels. ECG was recorded continuously on a single channel at a sampling rate of1kHz. This study uses facial video as the sole input modality. Binary classification is applied to distinguish neutral from stress conditions. The study received approval from the local Research Ethics Committee (approval no. 155/12- 09-2022). All participants provided informed consent. The dataset is available for non-commercial research upon request. 1 Subjects are partitioned into training, validation, and testing sets at the subject level, ensuring no participant appears in more than one set. To avoid performance inflation due to subject-difficulty imbalance, a stratified split protocol is used. Leave-one-subject-out cross-validation is conducted across all recorded modalities to estimate per-subject difficulty, providing a ranking that is not biased toward any single signal source. Subjects are ranked by the combined z-score and assigned to four quartiles. The final split comprises38training,8validation, and12testing subjects, with each set containing subjects from all four groups in proportion. The exact subject-level partition is reported in Table I to support reproducibility and direct comparison with future work. B. Video Table IV reports performance and computational cost across all seven temporal stride configurations; Fig. 2 visualizes the accuracy and efficiency trends jointly. At the densest setting (τ = 1), the full120-second recording is retained at30fps, yieldingL = 3,600frames and a token channel dimension of C ′ = 10,826. This configuration has the highest parameter count (9.16M), computational cost (348.78GFLOPs), and inference latency (133.02ms), with a corresponding throughput 1 https://github.com/ggian/stress dataset TABLE I: Subject-level split by difficulty group. Subjects are ranked by combined z-score and assigned to four quartile-based groups (Q1 = hardest, Q4 = easiest). Split Difficulty Group Q1 – HardQ2 – Med-HardQ3 – Med-EasyQ4 – Easy Training (38) P017, P018, P022, P026, P034, P035, P042, P045, P050, P056 P001, P002, P003, P007, P012, P021, P033, P040, P048 P004, P014, P016, P032, P036, P046, P047, P052, P053, P054 P005, P010, P020, P028, P029, P037, P039, P041, P057 Validation (8)P038, P055P009, P023P006, P013P019, P030 Testing (12)P008, P025, P044P011, P024, P043P015, P031, P058P027, P051, P059 Q1: z <−0.46; Q2: −0.46≤ z <−0.05; Q3: −0.05≤ z < +0.40; Q4: z ≥ +0.40. of7.52samples per second, yet achieves a test accuracy of 69.03%. Increasing the stride reducesLproportionally, which con- tracts the token channel dimensionC ′ = 3L + 26and the associated input projection weights. Atτ = 30, only120 frames are retained, reducing the parameter count to5.82M, GFLOPs to12.48, and latency to14.77ms, corresponding to an approximately28-fold reduction in compute relative to τ = 1. As shown in Fig. 2(b), computational cost decreases sharply as the stride increases, while latency drops rapidly up toτ = 10and then changes more gradually. Throughput similarly increases steeply at lower strides before saturating near 62–63 samples per second for τ ≥ 15. Test accuracy does not decrease monotonically with stride. The highest test accuracy is reached atτ = 15(69.44%), with τ = 1yielding the second-best result (69.03%). Validation accuracy peaks atτ = 20(70.42%). The configurationτ = 5 produces the lowest test accuracy (60.56%) despite its second- highest validation accuracy (68.80%), reflecting per-stride generalization variability over the12-subject test partition. The remaining configurations fall between63.75%and66.25%on the test set. These results indicate that temporal density beyond a moderate threshold does not yield consistent discriminative benefit, and that FUSE accommodates the full stride range without structural modification. C. Overall Analysis & Discussion The central capability demonstrated by this evaluation is that FUSE processes the complete facial recording as a single input, without temporal windowing or external segmentation. Atτ = 1, this corresponds to3,600frames from a120- second recording ingested in one model pass. This differs from common video-based stress-recognition pipelines that first decompose recordings into short clips or predefined temporal windows before classification. In FUSE, the full sequence is encoded via axis folding and processed with asymmetric attention, whereas the internal spatial token segmentation does not correspond to temporal windows. The reorganization of the temporal axis is consistent with prior evidence that restructuring facial spatiotemporal representations can benefit affective assessment [35]. The stride sweep shows that this capability is not limited to dense sampling. The same architecture accommodates inputs ranging from120frames atτ = 30to3,600frames atτ = 60 62 64 66 68 70 72 Test Accuracy (%) (a) size Params (M) Test Accuracy Params = 5.82 M Params = 9.16 M 12510152030 Temporal Stride 10 20 50 100 200 500 GFLOPs (b) Samples/s: s= 1: 7.5 s= 2: 15.6 s= 5: 35.9 s=10: 58.9 s=15: 61.0 s=20: 62.5 s=30: 63.1 10 20 50 100 200 Latency (ms) Fig. 2: Performance and computational cost of FUSE across temporal stride configurations. (a) Test accuracy and parameter count as a function of strideτ; circle area encodes the number of parameters, with the smallest and largest sizes corresponding to5.82M (τ = 30) and9.16M (τ = 1), respectively. (b) GFLOPs (solid red, left axis) and inference latency in milliseconds (dashed green, right axis) as a function of stride, shown on logarithmic axes; throughput in samples per second is annotated for each configuration. Inference cost is measured on an NVIDIA A100 GPU. 1 , with the token channel dimensionC ′ = 3L + 26scaling with the retained sequence length. No clip-level aggregation, temporal pooling, or external segmentation is introduced. This supports the main design goal of FUSE: any-length facial video processing without changing the model structure across temporal-stride settings. The best test accuracy is obtained atτ = 15(69.44%), using24.07GFLOPs, compared with348.78GFLOPs atτ = 1. The full-frame configuration still achieves a competitive test accuracy of69.03%while processing the entire3,600- frame sequence. The non-monotonic relationship between stride and test accuracy, visible in Fig. 2(a), indicates that denser temporal sampling does not necessarily improve generalization. Instead, moderate subsampling can preserve stress-relevant TABLE IV: Performance and computational cost using the video modality. InputComputational CostInference CostValidationTesting ModalityStrideParams (M)GFLOPsLatency (ms) GPU↓Samples/s GPU↑Accuracy PrecisionF1Accuracy Video19.16348.78133.027.5263.9066.7060.4869.03 Video27.43174.8364.0015.6366.4267.1366.5266.25 Video56.3970.4627.8435.9268.8069.3967.3960.56 Video106.0535.6716.9758.9266.3966.2066.1463.75 Video155.9324.0716.3961.0167.7768.3367.8969.44 Video205.8718.2715.0362.4770.4271.8668.4565.97 Video305.8212.4814.7763.1468.6069.9168.7264.86 Inference Cost measured on an NVIDIA A100 GPU. facial information while substantially reducing computational cost. The higher strides likely perform comparably because consecutive frames at30fps are largely redundant, so retaining every frame mainly enlarges the input projection (5.82M to 9.16M parameters) without adding useful information. A limitation of the current evaluation is that it uses a single dataset and a binary neutral-versus-stress classification setting. In addition, the study does not include a direct comparison against windowed baselines under the same subject-level split, which is left for future work. Windowing itself offers a practical benefit, as fixed-length clips yield a smaller and uniform input that avoids the growing channel dimension of the full-recording formulation. Consequently, the results support the feasibility of full-recording inference, but they should not be interpreted as a complete replacement for all window-based video stress- recognition strategies. V. CONCLUSION This paper presented FUSE, a facial-video stress detection framework designed to process complete recordings without temporal windowing or external segmentation. The proposed ap- proach folds the temporal dimension into the channel dimension and uses asymmetric attention to process the resulting high- dimensional representation through a compact set of segment states. This enables the same architecture to operate across a wide range of temporal stride settings, from sparse subsampling to full-frame input. Experiments on a58-subject stress dataset show that FUSE can process a full120-second recording at τ = 1, corresponding to3,600frames, while also supporting more efficient stride configurations without structural changes. The best test accuracy is achieved atτ = 15(69.44%), whereas the full-frame configuration remains competitive at69.03%. These results indicate that temporal windowing is not required for effective facial-video stress detection in this setting, and that complete-recording inference can be achieved within a single unified architecture. Overall, FUSE shifts the unit of analysis from short temporal clips to complete facial recordings. This provides a simpler evaluation pipeline, avoids choices about window length and aggregation, and preserves the ability to model stress-related facial information over the full duration of the recording. ACKNOWLEDGMENTS The authors used large language model (LLM)-based tools for language editing and improvement. All scientific content, results, and conclusions are solely the work of the authors. REFERENCES [1] D. S. Goldstein, “Stress and the autonomic nervous system,” Autonomic Neuroscience, vol. 247, p. 103096, 2023. [2] D. H. Hellhammer, S. W ̈ ust, and B. M. Kudielka, “Salivary cortisol as a biomarker in stress research,” Psychoneuroendocrinology, vol. 34, no. 2, p. 163–171, 2009. [3] S. Cohen, T. Kamarck, and R. Mermelstein, “A global measure of perceived stress,” Journal of Health and Social Behavior, vol. 24, no. 4, p. 385–396, 1983. [4]S. Shiffman, A. A. Stone, and M. R. Hufford, “Ecological momentary assessment,” Annual Review of Clinical Psychology, vol. 4, p. 1–32, 2008. [5]E. F. Canaletti, P. Lun, L. D. Stutzman, M. Chan, and F. Cheung, “Rising tide of stress: Global trends and structural predictors over 18 years,” Wellbeing, Space and Society, vol. 10, p. 100319, 2026. [6] H. Sultan-Ta ̈ ıeb, T. Villeneuve, J.-F. Chastang, and I. Niedhammer, “Bur- den of cardiovascular diseases and depression attributable to psychosocial work exposures in 28 European countries,” European Journal of Public Health, vol. 32, no. 4, p. 586–592, 2022. [7] S. Cohen, D. Janicki-Deverts, and G. E. Miller, “Psychological stress and disease,” JAMA, vol. 298, no. 14, p. 1685–1687, 2007. [8]M. Hosseini, R. Gottumukkala, R. Bhupatiraju, A. Maida, and H. Chu, “Wearable-based stress detection for real-world data: Perspective on challenges and recommendations,” 2026, JMIR Preprints, Preprint ID: 93741. [9]P. Neigel, A. Vargo, B. Tag, and K. Kise, “Unobtrusive stress detection using wearables: application and challenges in a university setting,” Frontiers in Computer Science, vol. 7, p. 1575404, 2025. [10]G. Giannakakis, D. Grigoriadis, K. Giannakaki, O. Simantiraki, A. Roni- otis, and M. Tsiknakis, “Review on psychological stress detection using biosignals,” IEEE Transactions on Affective Computing, vol. 13, no. 1, p. 440–460, 2022. [11]G. Giannakakis, M. Pediaditis, D. Manousos, E. Kazantzaki, F. Chiarugi, P. G. Simos, K. Marias, and M. Tsiknakis, “Stress and anxiety detection using facial cues from videos,” Biomedical Signal Processing and Control, vol. 31, p. 89–101, 2017. [12] G. Giannakakis, M. R. Koujan, A. Roussos, and K. Marias, “Automatic stress detection evaluating models of facial action units,” in 2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020), 2020, p. 728–733. [13]M. Kyrou, I. Kompatsiaris, and P. C. Petrantonakis, “Deep learning approaches for stress detection: a survey,” IEEE Transactions on Affective Computing, vol. 16, no. 2, p. 499–517, 2025. [14]H. Zhang, L. Feng, N. Li, Z. Jin, and L. Cao, “Video-based stress detection through deep learning,” Sensors, vol. 20, no. 19, p. 5552, 2020. [15]T. Jeon, H. B. Bae, Y. Lee, S. Jang, and S. Lee, “Deep-learning-based stress recognition with spatial-temporal facial information,” Sensors, vol. 21, no. 22, p. 7498, 2021. [16]P. Valergaki, V. C. Nicodemou, I. Oikonomidis, A. Argyros, and A. Roussos, “Combining facial videos and biosignals for stress estimation during driving,” 2026. [17] S. Gkikas, T. Kassiotis, Y. Guo, G. Li, E. Nichols, H. Asadi, N. Smyrnis, and G. Giannakakis, “Beyond the Raw Waveform: Fusing Visual Representations of EDA for Stress Detection,” in 2026 International Conference on Pattern Recognition and Artificial Intelligence (PRAI), 2026, accepted for publication. [18]S. Gkikas, I. Kyprakis, and M. Tsiknakis, “Multi-representation diagrams for pain recognition: Integrating various electrodermal activity signals into a single image,” in Companion Proceedings of the 27th International Conference on Multimodal Interaction, ser. ICMI Companion ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 162–171. [19]—, “Tiny-biomoe: a lightweight embedding model for biosignal analysis,” in Companion Proceedings of the 27th International Conference on Multimodal Interaction, ser. ICMI Companion ’25.New York, NY, USA: Association for Computing Machinery, 2025, p. 117–126. [20] —, “Efficient pain recognition via respiration signals: A single cross- attention transformer multi-window fusion pipeline,” in Companion Pro- ceedings of the 27th International Conference on Multimodal Interaction, ser. ICMI Companion ’25.New York, NY, USA: Association for Computing Machinery, 2025, p. 70–79. [21]S. Gkikas, Y. Guo, G. Li, R. F. Rojas, G. Giannakakis, and R. Gomez, “A Multi-Scale Temporal Framework with Dynamic Fusion for EEG- Based Emotion Recognition,” in 2026 International Conference on Pattern Recognition and Artificial Intelligence (PRAI), 2026, accepted for publication. [22] S. Gkikas, C. A. Cruz, C. Joseph, G. Giannakakis, and R. F. Rojas, “Towards a Unified Modality-Agnostic Multimodal Framework for Cognitive Workload Assessment,” in 2026 14th International Conference on Affective Computing and Intelligent Interaction (ACII).IEEE, 2026. [23]D. Ding, W. Xu, X. Liu, and T. Zhu, “Facial video based stress detection for enhancing ecological validity,” Acta Psychologica, vol. 255, p. 104877, 2025. [24] G. Giannakakis, M. R. Koujan, A. Roussos, and K. Marias, “Automatic stress analysis from facial videos based on deep facial action units recognition,” Pattern Analysis and Applications, vol. 25, no. 3, p. 521– 535, 2022. [25]G. Giannakakis, A. Roussos, C. Andreou, S. Borgwardt, and A. I. Korda, “Stress recognition identifying relevant facial action units through explainable artificial intelligence and machine learning,” Computer Methods and Programs in Biomedicine, vol. 259, p. 108507, 2025. [26]T. Kassiotis, S. Gkikas, N. Smyrnis, and G. Giannakakis, “Explainable Graph Attention Network for Stress Recognition (StressGAT) via Differential Action Units,” in 2026 14th International Conference on Affective Computing and Intelligent Interaction (ACII). IEEE, 2026. [27]A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lu ˇ ci ́ c, and C. Schmid, “ViViT: A video vision transformer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, p. 6836– 6846. [28]A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira, “Perceiver: General perception with iterative attention,” in Proceedings of the 38th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 139. PMLR, 2021, p. 4651–4664. [29]S. Gkikas and M. Tsiknakis, “A full transformer-based framework for automatic pain estimation using videos,” in 2023 45th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), 2023, p. 1–6. [30] S. Gkikas, N. S. Tachos, S. Andreadis, V. C. Pezoulas, D. Zaridis, G. Gkois, A. Matonaki, T. G. Stavropoulos, and D. I. Fotiadis, “Mul- timodal automatic assessment of acute pain through facial videos and heart rate signals utilizing transformer-based architectures,” Frontiers in Pain Research, vol. 5, 2024. [31] S. Gkikas and M. Tsiknakis, “Twins-painvit: Towards a modality-agnostic vision transformer framework for multimodal automatic pain assessment using facial videos and fnirs,” in 2024 12th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW), 2024, p. 13–21. [32]S. Gkikas, R. F. Rojas, and M. Tsiknakis, “Painformer: A vision foundation model for automatic pain assessment,” IEEE Transactions on Affective Computing, vol. 16, no. 4, p. 3369–3386, 2025. [33]J. R. Stroop, “Studies of interference in serial verbal reactions,” Journal of Experimental Psychology, vol. 18, no. 6, p. 643–662, 1935. [34] T. N. Tombaugh, “A comprehensive review of the paced auditory serial addition test (pasat),” Archives of Clinical Neuropsychology, vol. 21, no. 1, p. 53–76, 2006. [35]S. Gkikas, Y. Fang, C. A. Cruz, M. U. Khan, and R. F. Rojas, “ReFace: Reorganizing Facial Spatiotemporal Representations for Improved Pain Assessment,” in 2026 14th International Conference on Affective Com- puting and Intelligent Interaction (ACII). IEEE, 2026.