Paper deep dive
CardiacMamba: Fair and Robust RGB-RF Fusion for Remote Heart Rate Estimation via State Space Modeling
Bo Zhao, Zheng Wu, Yiping Xie, Zitong YU
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/22/2026, 3:07:19 AM
Summary
The paper introduces CardiacMamba, a state-space modeling framework for remote heart rate estimation that fuses RGB video and radio-frequency (RF) signals. It addresses limitations of RGB-only methods (illumination/skin-tone sensitivity) and RF-only methods (low resolution/motion artifacts) by using a Temporal Difference Mamba Module (TDMM), bidirectional SSM interaction, and Channel-wise Fast Fourier Transform (CFFT). CardiacMamba achieves state-of-the-art performance on the EquiPleth dataset, significantly reducing skin-tone bias and maintaining robustness under missing modalities.
Entities (8)
Relation Signals (6)
CardiacMamba → evaluatedon → EquiPleth
confidence 98% · On the EquiPleth dataset, CardiacMamba achieves state-of-the-art performance
CardiacMamba → usesmodule → Temporal Difference Mamba Module
confidence 95% · CardiacMamba introduces a Temporal Difference Mamba Module (TDMM) to enhance subtle RF temporal variations
CardiacMamba → usesmodule → Channel-wise Fast Fourier Transform
confidence 95% · a Channel-wise Fast Fourier Transform (CFFT) module for channel-domain spectral refinement.
CardiacMamba → fusesmodalities → Remote Photoplethysmography
confidence 90% · CardiacMamba, a fair and robust RGB-RF fusion framework that integrates optical facial cues and radio-frequency cardiac motion cues
CardiacMamba → outperforms → Vilesov et al.
confidence 90% · Compared with Vilesov et al. [7], it further reduces MAE and RMSE by 14.3% and 10.5%.
CardiacMamba → usesbackbone → Vision Mamba
confidence 90% · Second, Vision Mamba (Vim) is used to model long-range temporal dependencies in both modalities
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Remote photoplethysmography (rPPG) enables non-contact heart rate (HR) monitoring from facial videos, but RGB-only methods are vulnerable to illumination changes, motion artifacts, and skin-tone-dependent optical reflectance. We propose CardiacMamba, a fair and robust RGB-RF fusion framework that integrates optical facial cues and radio-frequency cardiac motion cues through state space modeling. CardiacMamba introduces a Temporal Difference Mamba Module (TDMM) to enhance subtle RF temporal variations, a bidirectional SSM-based interaction mechanism to align heterogeneous RGB-RF dynamics, and a Channel-wise Fast Fourier Transform (CFFT) module for channel-domain spectral refinement. On the EquiPleth dataset, CardiacMamba achieves state-of-the-art performance with 0.96 bpm MAE, 3.06 bpm RMSE, and 0.97 Pearson correlation, while reducing the observed light-dark skin-tone MAE gap to 0.26 bpm and maintaining robustness under RGB degradation and RF-missing conditions
Tags
Links
- Source: https://arxiv.org/abs/2608.15831v1
- Canonical: https://arxiv.org/abs/2608.15831v1
Trouble viewing inline? Open PDF directly →
Full Text
37,600 characters extracted from source content.
Expand or collapse full text
CardiacMamba: Fair and Robust RGB-RF Fusion for Remote Heart Rate Estimation via State Space Modeling Bo Zhao Zheng Wu Yiping Xie Affiliation: Great Bay University Zitong Yu [0.4em] These authors contributed equally to this work. Corresponding author Abstract Remote photoplethysmography (rPPG) enables non-contact heart rate (HR) monitoring from facial videos, but RGB-only methods are vulnerable to illumination changes, motion artifacts, and skin-tone-dependent optical reflectance. We propose CardiacMamba, a fair and robust RGB-RF fusion framework that integrates optical facial cues and radio-frequency cardiac motion cues through state space modeling. CardiacMamba introduces a Temporal Difference Mamba Module (TDMM) to enhance subtle RF temporal variations, a bidirectional SSM-based interaction mechanism to align heterogeneous RGB-RF dynamics, and a Channel-wise Fast Fourier Transform (CFFT) module for channel-domain spectral refinement. On the EquiPleth dataset, CardiacMamba achieves state-of-the-art performance with 0.96 bpm MAE, 3.06 bpm RMSE, and 0.97 Pearson correlation, while reducing the observed light-dark skin-tone MAE gap to 0.26 bpm and maintaining robustness under RGB degradation and RF-missing conditions. Keywords remote photoplethysmography, RGB-RF fusion, SSM 1 Introduction Remote photoplethysmography (rPPG) enables non-contact heart rate (HR) monitoring from facial videos, offering an unobtrusive alternative to contact-based ECG and PPG. Despite progress in physics-based [1] and deep learning-based [2] methods, RGB-based rPPG remains limited by its dependence on optical reflectance: illumination changes, motion artifacts, and reduced pulsatile contrast under darker skin tones degrade the signal-to-noise ratio. Radio Frequency (RF) sensing offers a complementary mechanism. By capturing minute chest-wall vibrations through electromagnetic reflections, RF signals are largely insensitive to ambient lighting and skin pigmentation, but suffer from lower spatial resolution and vulnerability to body motion and multipath interference [3]. These complementary properties motivate RGB-RF fusion: RGB provides rich facial appearance cues, while RF supplies illumination- and skin-tone-invariant mechanical cardiac information. Effective RGB-RF fusion, however, remains challenging. Existing strategies often rely on shallow feature concatenation or late-stage fusion, insufficiently modeling the heterogeneous dynamics between optical BVP signals and RF-induced chest motion. Three issues remain underexplored: aligning RGB and RF features with temporal characteristics in a shared representation space, enhancing physiological components via frequency-domain interaction and mitigating demographic disparities caused by the optical dependence of RGB-based rPPG. To address these challenges, we propose CardiacMamba, an SSM-based RGB-RF fusion framework for fair and robust remote HR estimation. It introduces a Temporal Difference Mamba Module (TDMM) to enhance subtle RF temporal variations, a bidirectional SSM-based interaction mechanism to align heterogeneous RGB-RF dynamics, and a Channel-wise Fast Fourier Transform (CFFT) module for channel-domain spectral refinement. On EquiPleth, CardiacMamba achieves state-of-the-art accuracy, reduces skin-tone-related performance disparities, improves robustness under RGB degradation and RF-missing conditions. Our main contributions are summarized as follows: • We propose CardiacMamba, a state-space RGB-RF fusion framework that jointly exploits optical facial cues and illumination-invariant RF cardiac cues for robust and fair remote HR estimation. • We develop a dynamic multimodal architecture where TDMM enhances RF temporal variations, a bidirectional SSM-based mechanism aligns heterogeneous RGB-RF dynamics, and CFFT performs channel-wise frequency-domain refinement. • Extensive experiments on EquiPleth show that CardiacMamba achieves 0.96 bpm MAE, 3.06 bpm RMSE, and 0.97 Pearson correlation, reducing the observed light-dark skin-tone MAE gap to 0.26 bpm while maintaining robustness under RGB degradation and RF-missing conditions. 2 Related Work 2.1 RGB Video-Based Methods RGB video-based rPPG estimates physiological signals from subtle facial appearance variations. Early methods relied on hand-crafted decomposition such as PCA and ICA [4], while deep learning later advanced rPPG with CNN [2] and Transformer [5, 23] architectures. Nevertheless, RGB-based rPPG remains constrained by optical sensing mechanism, motivating non-visual modalities. Beyond physiological measurement, deep learning has also advanced face-video understanding in related tasks such as face anti-spoofing [21, 22, 24, 25, 27, 28], where robustness to appearance variation and multimodal cues is likewise essential. 2.2 RF Radar-Based Methods RF radar measures minute chest displacements caused by cardiac motion; early frequency-domain pipelines [3] gave way to deep learning approaches [6]. RF is robust to illumination and skin pigmentation, but has lower spatial resolution and is affected by body motion, so it is best used as a complement to RGB. 2.3 Multi-modal Fusion Methods Prior work fused RGB with Infrared signals, and Vilesov et al. [7] studied RGB-RF fusion with camera and 77 GHz radar. However, RGB videos and RF signals capture different manifestations of cardiac activity (optical blood-volume variations vs. mechanical chest-wall motion), making simple concatenation or late fusion insufficient; existing methods often lack explicit cross-modal temporal alignment and frequency-domain interaction. 2.4 Mamba and State Space Models State Space Models (SSMs) [8] model long sequences through structured state transitions, and Vision Mamba (Vim) [9] introduces bidirectional state-space modeling with lower cost than Transformers, well suited to dense video or radar sequences with weak quasi-periodic dynamics. We build on these strengths for RGB-RF fusion. In parallel, recent works explore large language models [10], physics-grounded harmonic attention [11], multi-agent frameworks [12], optimal-transport feature warping [13], causal self-supervised learning [33], and dual-branch structured attention [34] for remote physiological measurement; LLM-based multimodal understanding is also advancing rapidly across affective computing, e.g., audio-visual reasoning [26], micro-expression action units [29, 30], and emotion world models [31, 32]. 3 Methodology 3.1 Preliminaries We adopt the discretized state space model (Mamba) [8] for sequence modeling; the formal definitions of the continuous SSM, its discretization, and the convolutional equivalence are given in Appendix C. 3.2 Overview Figure 1: The overall architecture of CardiacMamba. It consists of three stages: Dual-level Feature Extraction and Alignment, Bidirectional Feature Interaction, and Bidirectional Feature Fusion. As shown in Fig. 1, CardiacMamba estimates physiological signal by integrating complementary observations. Given an RGB video IC∈ℝ3×T1×H×WI_C ^3× T_1× H× W and an RF input If∈ℝC×T2I_f ^C× T_2, the framework contains three stages: dual-level feature extraction, SSM-based temporal interaction and frequency-domain fusion. First, modality-specific encoders extract physiological representations from the two streams. In the RGB branch, BDCF enhances BVP-related temporal color variations, while SCFM aggregates informative spatial-channel responses. In the RF branch, TDMM captures cardiac-induced temporal variations and two RFAMs refine RF features while aligning temporal resolution with RGB stream: Hcn2=SCFM(BDCF(IC)),Hfn2=RFAM(RFAM(TDMM(If))).H_c^n_2=SCFM(BDCF(I_C)), H_f^n_2=RFAM(RFAM(TDMM(I_f))). (1) Second, Vision Mamba (Vim) is used to model long-range temporal dependencies in both modalities under a Figure 2: Channel-wise Fast Fourier Transform (CFFT) for refining RGB and RF representations through channel-domain spectral interaction. shared SSM-based dynamic prior. Residual connections and linear projections stabilize the representation refinement: Hcn4=Linearc(Hcn2+Vimc(Hcn2)),Hfn4=Linearf(Hfn2+Vimf(Hfn2)).H_c^n_4=Linear_c(H_c^n_2+Vim_c(H_c^n_2)), H_f^n_4=Linear_f(H_f^n_2+Vim_f(H_f^n_2)). (2) This stage encourages RGB and RF features to share a coherent temporal evolution while preserving modality-specific cues. Finally, the Channel-wise Fast Fourier Transform (CFFT) module refines the channel spectra of both branches (Appendix D), and the refined representations are fused by the prediction head to reconstruct the BVP waveform: y^=Predictor(Fuse(Hcn5,Hfn5)), y=Predictor(Fuse(H_c^n_5,H_f^n_5)), (3) where Fuse(⋅)Fuse(·) denotes multimodal aggregation and y y is the predicted signal. 3.3 Dual-level Feature Extraction and Alignment 3.3.1 Low-level Feature Extraction Temporal Difference Mamba Module (TDMM). RF signals reflect cardiac activity through subtle chest-wall displacement, but such weak temporal variations are often buried by static reflections and low-frequency body motion. TDMM (Fig. 3) highlights local temporal changes by frame differencing (Appendix E) and captures long-range dependencies with an SSM-based Mamba block, Figure 3: Temporal Difference Mamba Module (TDMM) for extracting RF dynamic temporal features. whose output is computed as XTDMM X_TDMM =ReLU(Linearo(SSM(G⊙Conv7×1(Linearu(X0))))), =ReLU (Linear_o(SSM(G _7× 1(Linear_u(X_0)))) ), (4) G G =σ(Linearg(X0)). =σ(Linear_g(X_0)). Bifurcated Diff-Conv Fusion (BDCF). For the RGB branch, BDCF enhances weak BVP-related color variations by processing the raw sequence and its temporal difference representation in parallel branches. The temporal differences follow the formulation in Appendix E. The two branches are fused as Xfu X_fu =αStem2(Xori)+βStem2(αXori+βXdiff), = _2(X_ori)+ _2(α X_ori+β X_diff), (5) Xori X_ori =Stem1(X),Xdiff=Stem1(Concat(D−2,D−1,D1,D2)). =Stem_1(X), X_diff=Stem_1(Concat(D_-2,D_-1,D_1,D_2)). where Stem1Stem_1 consists of a 7×77× 7 convolution, batch normalization, ReLU activation, and max pooling; Stem2Stem_2 consists of a 7×77× 7 convolution, batch normalization, and ReLU activation. We set α=β=0.5α=β=0.5. 3.3.2 High-level Feature Extraction Spatial-Channel Fusion Module (SCFM). SCFM suppresses uninformative spatial responses and produces compact high-level RGB features: a lightweight 5×55× 5 convolutional stem generates a normalized spatial attention (Appendix E), and the attended representation is globally aggregated and refined as The attended representation is then globally aggregated and refined: Xstem=BN(Conv(GAP(Xattn))).X_stem=BN (Conv (GAP(X_attn) ) ). (6) RF Alignment Module (RFAM). RFAM refines RF features and aligns their temporal resolution with the RGB branch through local temporal convolution, channel attention, and strided downsampling (Appendix E). Its output is XRFAM=ReLU(Conv7×1s=2(X~r)).X_RFAM=ReLU (Conv_7× 1^s=2 ( X_r ) ). (7) Through temporal refinement, channel reweighting, and resolution alignment, RFAM produces RF features compatible with subsequent RGB-RF fusion. 3.4 Bidirectional Feature Fusion After modality-specific temporal modeling, the branches produce representations Hcn4H_c^n_4 and Hfn4H_f^n_4 that still contain modality-specific noise and redundant channels. We refine them with the Channel-wise Fast Fourier Transform (CFFT) module (Fig. 2), which performs spectral mixing along the feature-channel dimension: given H∈ℝB×C×TH ^B× C× T, a learnable complex-valued transformation refines the channel-frequency spectrum, and the real part of the inverse transform is retained as the output (Appendix D): CFFT(H)=Re(H′).CFFT(H)=Re(H ). (8) CFFT is applied to both modality branches: Hcn5=CFFT(Hcn4),Hfn5=CFFT(Hfn4).H_c^n_5=CFFT(H_c^n_4), H_f^n_5=CFFT(H_f^n_4). (9) For bidirectional multimodal fusion, each modality modulates the other through a lightweight cross-gating mechanism: H~c=Hcn5⊙σ(Linearf→c(Hfn5)),H~f=Hfn5⊙σ(Linearc→f(Hcn5)), H_c=H_c^n_5 σ(Linear_f→ c(H_f^n_5)), H_f=H_f^n_5 σ(Linear_c→ f(H_c^n_5)), (10) where σ(⋅)σ(·) is the Sigmoid function and ⊙ denotes element-wise multiplication. The refined features are aggregated and passed to the prediction head: y^=Predictor(Fuse(H~c,H~f)). y=Predictor(Fuse( H_c, H_f)). (11) 4 Experiment The specifics of the datasets, evaluation metrics, and experimental setup, covering data splits, metric formulations are given in Appendices A and B. 4.1 Comparison with State-of-the-Art Methods Table 1: Comparison of different methods on the EquiPleth dataset. The best results are marked in bold. Method Input MAE RMSE ρ DeepPhys [14] RGB 5.54 18.51 0.66 PhysNet [2] RGB 8.06 19.71 0.61 MTTS-CAN [15] RGB 3.69 13.8 0.82 PhysFormer [5] RGB 12.92 24.36 0.47 EfficientPhys [16] RGB 5.47 17.04 0.71 RhythmMamba [17] RGB 2.87 9.58 0.92 Ours (RGB-Only) RGB 1.2 4.23 0.95 Tu et al. RF 5.5 11.68 0.64 Mercuri et al. RF 4.73 9.6 0.7 FFT-based [3] RF 13.51 21.07 0.24 Ours (RF-Only) RF 5.2 7.4 0.8 Vilesov et al. [7] RGB+RF 1.12 3.42 0.95 Ours (Full model) RGB+RF 0.96 3.06 0.97 Table 1 compares CardiacMamba with representative RGB-only, RF-based, and RGB-RF multimodal baselines [7]. CardiacMamba achieves the best overall performance, with 0.96 bpm MAE, 3.06 bpm RMSE, and 0.97 Pearson correlation. Compared with the best RGB-only baseline, it reduces MAE and RMSE by 66.6% and 68.1%, respectively; compared with Vilesov et al. [7], it further reduces MAE and RMSE by 14.3% and 10.5%. These results validate the effectiveness of integrating RF cues with RGB features through TDMM, SSM-based temporal modeling, and CFFT-based channel-domain refinement. In the RF-only setting, our RF branch achieves the lowest RMSE and highest correlation among RF-based methods, although its MAE remains slightly higher than the strongest RF-only baseline, indicating that RF is most effective as a complementary modality rather than a replacement for RGB. 4.2 Measuring Skin Tone Bias and Fairness We assess skin-tone fairness by measuring light-dark group performance gaps in Table 5, where smaller gaps indicate more consistent HR estimation across demographic groups. CardiacMamba achieves the smallest MAE gap of 0.26 bpm, reducing the gap by 61.2% compared with Vilesov et al. [7] and substantially outperforming ICA [18] and PhysNet [2], whose MAE gaps are 4.42 bpm and 2.22 bpm, respectively. It also obtains a 1.28 bpm RMSE gap and a 0.05 Pearson correlation gap, lower in magnitude than those of most RGB-based baselines [2]. These results suggest that RGB-RF fusion mitigates skin-tone-related imbalance by reducing reliance on optical skin reflectance, while fairness is reported as an observed EquiPleth performance disparity. 4.3 Measurement in Missing Modality Scenarios Table 2: Testing under missing-modality conditions. Method Train Test MAE RMSE Base RGB&RF RGB 21.7 25.7 Base RGB&RF RF 21.4 24.3 Base RGB&RF RGB&RF 20.3 24.8 Vilesov et al. [7] RGB&RF RGB 6.82 13.32 Vilesov et al. [7] RGB&RF RF 7.25 9.62 Vilesov et al. [7] RGB&RF RGB&RF 1.12 3.42 CardiacMamba (Ours) RGB&RF RGB 1.2 3.41 CardiacMamba (Ours) RGB&RF RF 11.0 13.0 CardiacMamba (Ours) RGB&RF RGB&RF 0.96 3.06 Table 2 evaluates missing-modality robustness by training with RGB-RF inputs and testing with full or partial modalities. With both modalities available, CardiacMamba outperforms Vilesov et al. [7], reducing MAE and RMSE by 14.3% and 10.5%, respectively. When RF is absent, it maintains strong RGB-only inference with 1.20 bpm MAE, close to the full-modality result of 0.96 bpm, and reduces MAE by 82.4% over Vilesov et al. [7] under the same setting. However, when RGB is unavailable, performance degrades to 11.00 bpm MAE and 13.00 bpm RMSE, revealing asymmetric robustness. Thus, CardiacMamba is robust to RF missingness and RGB degradation, but RF-only 4.4 Ablation Study Table 3: Ablation on different modules, including ‘Vim’ (short for Vision Mamba), ‘CFFT’ (short for Channel-wise Fast Fourier Transform), ‘RFAM’ (short for RF Alignment Module), and ‘TDMM’ (short for Temporal Difference Mamba Module). Vim CFFT SSM RFAM TDMM MAE RMSE ρ × ✓ ✓ ✓ ✓ 1.7 4.9 0.91 ✓ × ✓ ✓ ✓ 4.92 6.33 0.77 ✓ ✓ × ✓ ✓ 1.86 5.51 0.89 ✓ ✓ ✓ × ✓ 1.85 5.31 0.9 ✓ ✓ ✓ ✓ × 3.82 8.06 0.81 ✓ ✓ ✓ ✓ ✓ 0.96 3.06 0.97 Table 4: Ablation study on model robustness under simulated RGB signal degradation. We added Gaussian noise to the RGB input to test performance. Best results under noisy conditions are in bold. Model Condition MAE RMSE ρ Ours (RGB-Only) Normal 1.20 4.23 0.95 + Gaussian Noise 8.54 15.12 0.41 Ours (Full model) Normal 0.96 3.06 0.97 + Gaussian Noise 2.15 5.80 0.88 Impact of Key Modules. Table 3 evaluates the contributions of Vim, CFFT, SSM, RFAM, and TDMM. Removing any component degrades performance, confirming that each module benefits RGB-RF representation learning. CFFT is the most critical: without it, MAE increases from 0.96 bpm to 4.92 bpm and ρ drops from 0.97 to 0.77, indicating that channel-domain spectral refinement is essential for stabilizing multimodal fusion. TDMM is also important for RF temporal modeling; removing it increases MAE and RMSE to 3.82 bpm and 8.06 bpm, respectively. Removing Vim raises MAE to 1.70 bpm, validating the benefit of SSM-based long-range temporal encoding, while removing SSM or RFAM leads to smaller but consistent degradation, with MAE increasing to 1.86 bpm and 1.85 bpm. These results show that CFFT and TDMM are influential, while SSM and RFAM further improve temporal stability and RF alignment. Robustness to RGB Degradation. Table 4 evaluates robustness under Gaussian noise added to the RGB input. The RGB-only model degrades severely, with MAE increasing from 1.20 bpm to 8.54 bpm and ρ dropping to 0.41. In contrast, the full RGB-RF model maintains substantially better performance under the same corruption, achieving 2.15 bpm MAE and 0.88 correlation. This demonstrates that RF provides a stable complementary signal when visual observations are degraded, supporting RGB-RF fusion for robust physiological measurement in challenging visual conditions. 4.5 Visualization and Analysis Figure 4: Visual representations of features (from left to right): human face, feature heat map, radar spectrum diagram, and its feature heat map. Fig. 4 visualizes the learned RGB and RF representations. Both modalities contribute distinct yet complementary (a) CardiacMamba (b) Vilesov et al. Figure 5: Bland-Altman plots comparing estimated and ground-truth heart rates. Figure 6: Comparison of ground-truth and predicted PPG signals. information: the optical branch encodes superficial blood-volume changes, whereas the RF branch reflects deeper mechanical cardiac activity, jointly supporting robust estimation under diverse conditions. The RGB heatmap highlights BVP-related facial regions, while the RF heatmap emphasizes radar responses associated with physiological motion, indicating that CardiacMamba captures optical and mechanical cardiac cues. Bland-Altman Analysis. Fig. 6 compares HR agreement with ground truth. CardiacMamba exhibits tighter clustering within the confidence intervals than Vilesov et al. [7], showing improved estimation consistency. Signal Reconstruction Quality. Fig. 6 shows that the predicted PPG signal follows the ground-truth periodic pattern, providing qualitative evidence of physiological waveform recovery. This phase-consistent alignment further suggests that CardiacMamba preserves beat-to-beat temporal dynamics rather than merely matching the dominant heart-rate frequency. 5 Conclusion We introduced CardiacMamba, an RGB-RF fusion framework integrating TDMM, bidirectional SSM, and CFFT for robust and fair remote HR estimation. It achieves state-of-the-art accuracy on EquiPleth, reduces observed skin-tone-related disparities, and remains robust under RGB degradation and RF-missing conditions. Future work will address the limited RF-only fallback and validate generalization across larger populations. References [1] Haan, G.D., Jeanne, V.: Robust pulse rate from chrominance-based rPPG. IEEE Trans. Biomed. Eng. 60(10), 2878–2886 (2013) [2] Yu, Z., Li, X., Zhao, G.: Remote photoplethysmograph signal measurement from facial videos using spatio-temporal networks. In: British Machine Vision Conference (BMVC) (2019) [3] Alizadeh, M., Shaker, G., Almeida, J.C.M.D., Morita, P.P., Safavi-Naeini, S.: Remote monitoring of human vital signs using m-wave FMCW radar. IEEE Access 7, 54958–54968 (2019) [4] Poh, M.Z., McDuff, D.J., Picard, R.W.: Noncontact, automated cardiac pulse measurements using video imaging and blind source separation. Opt. Express 18(10), 10762–10774 (2010) [5] Yu, Z., Shen, Y., Shi, J., Zhao, H., Torr, P.H., Zhao, G.: PhysFormer: Facial video-based physiological measurement with temporal difference transformer. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 4186–4196 (2022) [6] Zhang, H., Jian, P., Yao, Y., Liu, C., Wang, P., Chen, X., Du, L., Zhuang, C., Fang, Z.: Radar-Beat: Contactless beat-by-beat heart rate monitoring for life scenes. Biomed. Signal Process. Control 86, 105360 (2023) [7] Vilesov, A., Chari, P., Armouti, A., Harish, A.B., Kulkarni, K., Deoghare, A., Jalilian, L., Kadambi, A.: Blending camera and 77 GHz radar sensing for equitable, robust plethysmography. ACM Trans. Graph. 41(4), 1–14 (2022) [8] Gu, A., Goel, K., Ré, C.: Efficiently modeling long sequences with structured state spaces. In: International Conference on Learning Representations (ICLR) (2022) [9] Zhu, L., Liao, B., Zhang, Q., Wang, X., Liu, W., Wang, X.: Vision Mamba: Efficient visual representation learning with bidirectional state space model. In: International Conference on Machine Learning (ICML) (2024) [10] Xie, Y., Zhao, B., Dai, M., Zhou, J.P., Sun, Y., Tan, T., Xie, W., Shen, L., Yu, Z.: PhysLLM: Harnessing large language models for cross-modal remote physiological sensing. In: International Conference on Learning Representations (ICLR) (2026) [11] Zhao, B., Guo, D., Cao, J., Xu, Y., Zou, B., Tan, T., Sun, Y., Yu, Z.: PHASE-Net: Physics-grounded harmonic attention system for efficient remote photoplethysmography measurement. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 21198–21207 (2026) [12] Yang, Y., Zhao, B., Cao, J., Ma, H., Sun, Y., Wang, W., Yu, Z.: PhysAgent: A multi-agent framework for reliable remote heart rate estimation. arXiv preprint arXiv:2608.00066 (2026) [13] Zhao, B., Cao, J., Guo, D., Huang, D., Wang, W., Tan, T., Sun, Y., Yu, Z.: FLOW: Optimal transport-driven feature warping for generalized remote physiological measurement. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 28481–28491 (2026) [14] Chen, W., McDuff, D.: DeepPhys: Video-based physiological measurement using convolutional attention networks. In: European Conference on Computer Vision (ECCV), p. 349–365 (2018) [15] Liu, X., Fromm, J., Patel, S., McDuff, D.: Multi-task temporal shift attention networks for on-device contactless vitals measurement. In: Advances in Neural Information Processing Systems (NeurIPS), p. 1–23 (2020) [16] Liu, X., Hill, B., Jiang, Z., Patel, S., McDuff, D.: EfficientPhys: Enabling simple, fast and accurate camera-based cardiac measurement. In: IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), p. 5008–5017 (2023) [17] Zou, B., Guo, Z., Hu, X., Ma, H.: RhythmMamba: Fast remote physiological measurement with arbitrary length videos. arXiv preprint arXiv:2404.06483 (2024) [18] Poh, M.Z., McDuff, D.J., Picard, R.W.: Advancements in noncontact, multiparameter physiological measurements using a webcam. IEEE Trans. Biomed. Eng. 58(1), 7–11 (2010) [19] Balakrishnan, G., Durand, F., Guttag, J.: Detecting pulse from head motions in video. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), p. 3430–3437. IEEE, Portland, OR, USA (2013) [20] Zhang, K., Zhang, Z., Li, Z., Qiao, Y.: Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Process. Lett. 23(10), 1499–1503 (2016) [21] Yu, Z., Qin, Y., Li, X., Zhao, C., Lei, Z., Zhao, G.: Deep learning for face anti-spoofing: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 45(4), 5609–5631 (2023) [22] Yu, Z., Wan, J., Qin, Y., Li, X., Li, S.Z., Zhao, G.: NAS-FAS: Static-dynamic central difference network search for face anti-spoofing. IEEE Trans. Pattern Anal. Mach. Intell. 43(9), 3005–3023 (2021) [23] Yu, Z., Shen, Y., Shi, J., Zhao, H., Cui, Y., Zhang, J., Torr, P., Zhao, G.: PhysFormer++: Facial video-based physiological measurement with SlowFast temporal difference transformer. Int. J. Comput. Vis. 131, 1307–1330 (2023) [24] Yu, Z., Cai, R., Cui, Y., Liu, X., Hu, Y., Kot, A.C.: Rethinking vision transformer and masked autoencoder in multimodal face anti-spoofing. Int. J. Comput. Vis. 132, 5217–5238 (2024) [25] Lin, X., Liu, A., Yu, Z., Cai, R., Wang, S., Yu, Y., Wan, J., Lei, Z., Cao, X., Kot, A.: Reliable and balanced transfer learning for generalized multimodal face anti-spoofing. IEEE Trans. Pattern Anal. Mach. Intell. 47(8), 7608–7625 (2025) [26] Ye, Q., Yu, Z., Shao, R., Cui, Y., Kang, X., Liu, X., Torr, P., Cao, X.: CAT+: Investigating and enhancing audio-visual understanding in large language models. IEEE Trans. Pattern Anal. Mach. Intell. 47(10), 8674–8690 (2025) [27] Cai, R., Cui, Y., Yu, Z., Lin, X., Chen, C., Kot, A.: Rehearsal-free and efficient continual learning for cross-domain face anti-spoofing. IEEE Trans. Pattern Anal. Mach. Intell. 47(10), 11348–11365 (2025) [28] Lin, X., Wang, S., Yu, Y., Yu, Z., Zhou, J., Liu, Y., Cao, X., Kot, A., Zheng, Y.: Learning representation and synergy invariances: A provable framework for generalized multimodal face anti-spoofing. IEEE Trans. Pattern Anal. Mach. Intell. (2026) [29] Liu, Z., Yuan, K., Zhao, B., Xu, Y., Yu, Z.: AU-LLM: Micro-expression action unit detection via enhanced LLM-based feature fusion. In: Chinese Conference on Biometric Recognition (CCBR), p. 355–365 (2026) [30] Liu, Z., Yuan, K., Zhao, B., Ma, H., Yu, Z.: AULLM++: Structured-token-conditioned large language models for micro-expression action unit detection. arXiv preprint arXiv:2603.08387 (2026) [31] Wang, Z., Yu, Z., Zhu, Y., Zhao, B., Liang, H., Wang, T., Xia, W., Zhang, J., Liu, Z., et al.: AffectAgent: Collaborative multi-agent reasoning for retrieval-augmented multimodal emotion recognition. arXiv preprint arXiv:2604.12735 (2026) [32] Zhao, B., Ye, F., Ji, Y., Zhao, S., Peng, X., Yu, Z.: AffectVerse: Emotional world models for multimodal affective computing. arXiv preprint arXiv:2605.19950 (2026) [33] Niu, Z., Tu, X., Zhao, B., Cao, J., Guo, D., Yu, Z.: Intervention-based self-supervised learning: A causal probe paradigm for remote photoplethysmography. arXiv preprint arXiv:2605.00882 (2026) [34] Cao, J., Zhao, B., Niu, Z., Guo, D., Sun, Y., Liang, H., Xu, Y., Yu, Z.: PhysNeXt: Next-generation dual-branch structured attention fusion network for remote photoplethysmography measurement. arXiv preprint arXiv:2603.19752 (2026) Appendix A Datasets and Metrics Datasets. We evaluate CardiacMamba on EquiPleth [7], a synchronized RGB-RF benchmark for equitable remote physiological measurement. It contains recordings from 91 subjects, including 28 light-skin, Table 5: Comparison of the Fairness of various methods on the RGB and RF fusion task, with the differences in performance (Δ ) between light and dark skin tones in terms of MAE, RMSE, and correlation coefficient (ρ). Method Input Δ Δ Δ ρ ICA [18] RGB 4.42 3.15 -0.36 CHROM [1] RGB 4.97 4.17 -0.38 BCG [19] RGB 0.99 1.25 0.05 PhysNet [2] RGB 2.22 4.05 -0.25 FFT-based RF [3] RF 1.32 2.06 0.32 Vilesov et al. [7] RGB&RF 0.67 1.44 -0.10 CardiacMamba (Ours) RGB&RF 0.26 1.28 0.05 49 medium-skin, and 14 dark-skin subjects. Each subject participated in six 30-second sessions captured by an RGB camera at 30 fps and a 77 GHz FMCW radar. The paired streams provide complementary cardiac observations: RGB videos encode facial blood-volume-related appearance variations, while RF signals capture mechanical chest-wall motion. We follow a subject-independent protocol using the predefined training, validation, and testing folds, ensuring that no subject appears in more than one subset. This prevents identity leakage and enables reliable evaluation of cross-subject generalization. Metrics. Heart rate estimation is evaluated using Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and Pearson correlation coefficient (ρ). MAE and RMSE are reported in beats per minute (bpm), where lower values indicate better accuracy, while higher ρ indicates stronger agreement with ground-truth HR. Appendix B Experimental Setup B.1 Experimental Setup For RGB preprocessing, facial regions are detected by MTCNN [20], cropped from each frame, resized to 128×128128× 128, converted to floating-point tensors, and normalized by 255. For RF preprocessing, raw IQ samples are transformed into range profiles using Discrete Fourier Transform (DFT) and stacked over time to form range-time representations. We select the range bin with the highest average energy and retain a 25 cm neighboring window to focus on torso-related cardiac displacement while suppressing background reflections. The RF inputs are normalized by 1.255×1051.255× 10^5, processed with IQ rotation, and reshaped into channel-temporal representations. The model is trained for 30 epochs on an NVIDIA RTX 4090 GPU using Adam with a batch size of 32, an initial learning rate of 3×10−43× 10^-4, and a weight decay of 1×10−21× 10^-2. The best checkpoint is selected on the validation set and evaluated on the held-out test set. Unless otherwise specified, all learning-based baselines follow the same subject-independent protocol. Appendix C State Space Model Preliminaries State Space Models (SSMs) provide an efficient formulation for long-range sequence modeling through latent state evolution. Given a continuous input x(t)∈ℝx(t) , an SSM maps it to an output y(t)∈ℝy(t) via a hidden state h(t)∈ℝNh(t) ^N: dh(t)dt dh(t)dt =h(t)+x(t), =Ah(t)+Bx(t), (12) y(t) y(t) =h(t)+x(t), =Ch(t)+Dx(t), where A, B, C, and D denote the state transition, input projection, output projection, and skip connection parameters, respectively. For neural sequence modeling, the continuous system is discretized with step size Δ using Zero-Order Hold: ¯=exp(Δ),¯=(Δ)−1(exp(Δ)−)Δ. A= ( ), B=( )^-1 ( ( )-I ) . (13) The resulting recurrence is hk=¯hk−1+¯xk,yk=hk+xk,h_k= Ah_k-1+ Bx_k, y_k=Ch_k+Dx_k, (14) which can be equivalently computed as a structured convolution: =¯∗+,¯=(¯,¯¯,…,¯L−1¯).y= K*x+Dx, K= (C B,C A B,…,C A^L-1 B ). (15) This state-evolution structure enables efficient modeling of weak and quasi-periodic physiological dynamics over long RGB and RF sequences. Appendix D Channel-wise Fast Fourier Transform Details Unlike temporal Fourier analysis that directly estimates physiological frequencies, CFFT performs spectral mixing along the feature-channel dimension. This design models inter-channel dependencies in a transformed feature space, allowing informative responses to be enhanced while noisy or redundant channels are suppressed. Given H∈ℝB×C×TH ^B× C× T, CFFT applies a discrete Fourier transform along the channel dimension: H^b,k,t=∑c=0C−1Hb,c,texp(−j2πkcC),k=0,…,C−1, H_b,k,t= _c=0^C-1H_b,c,t (-j 2π kcC ), k=0,…,C-1, (16) where c and k denote the channel index and channel-frequency index, respectively. The transformed feature is decomposed as H^=H^re+jH^im. H= H^re+j H^im. (17) To enable learnable spectral interaction, CFFT applies a complex-valued linear transformation to the Fourier coefficients: H~re H^re =ϕ(H^rere−H^imim+re), =φ ( H^reW^re- H^imW^im+b^re ), (18) H~im H^im =ϕ(H^imre+H^reim+im), =φ ( H^imW^re+ H^reW^im+b^im ), where reW^re, imW^im, reb^re, and imb^im are learnable parameters, and ϕ(⋅)φ(·) denotes a nonlinear activation. This complex transformation jointly updates amplitude- and phase-related channel-frequency responses. The refined spectrum is written as H~=H~re+jH~im. H= H^re+j H^im. (19) The feature representation is reconstructed by the inverse transform: Hb,c,t′=1C∑k=0C−1H~b,k,texp(j2πkcC),c=0,…,C−1,H _b,c,t= 1C _k=0^C-1 H_b,k,t (j 2π kcC ), c=0,…,C-1, (20) Appendix E Additional Module Details E.1 Temporal Difference Computation in TDMM Given an RF feature sequence If∈ℝB×Ci×TiI_f ^B× C_i× T_i, TDMM constructs a five-frame temporal neighborhood Xt−2,Xt−1,Xt,Xt+1,Xt+2\X_t-2,X_t-1,X_t,X_t+1,X_t+2\ and computes adjacent differences: D−2=Xt−1−Xt−2,D−1=Xt−Xt−1, D_-2=X_t-1-X_t-2, D_-1=X_t-X_t-1, D1=Xt+1−Xt,D2=Xt+2−Xt+1. D_1=X_t+1-X_t, D_2=X_t+2-X_t+1. (21) The difference maps are concatenated and aggregated by a temporal convolution: X0=BN(Conv7×1(Concat(D−2,D−1,D1,D2))).X_0=BN (Conv_7× 1 (Concat(D_-2,D_-1,D_1,D_2) ) ). (22) E.2 Spatial Attention in SCFM Given XfuX_fu, a lightweight 5×55× 5 convolutional stem generates a spatial attention response A=σ(Stem(Xfu))A=σ(Stem(X_fu)), which is normalized and applied to the feature map: M=(H′W′)A2‖A‖1+ϵ,Xattn=M⊙Xfu.M= (H W )A2\|A\|_1+ε, X_attn=M X_fu. (23) E.3 Intermediate Refinement in RFAM RFAM first extracts local temporal patterns using a 7×17× 1 convolution: RFAM first extracts local temporal patterns using a 7×17× 1 convolution: Xr=BN(Conv7×1(X)).X_r=BN (Conv_7× 1(X) ). (24) A channel attention vector is computed from average- and max-pooled temporal descriptors: s=σ(MLP(AvgPool(Xr))+MLP(MaxPool(Xr))),X~r=s⊙Xr. gathereds=σ (MLP(AvgPool(X_r))+MLP(MaxPool(X_r)) ),\\[6.99997pt] X_r=s X_r. gathered (25) E.4 Channel-Spectral Refinement in Overview Finally, the Channel-wise Fast Fourier Transform (CFFT) module performs spectral refinement along the feature-channel dimension. Instead of estimating physiological frequency directly from raw temporal signals, CFFT enhances informative inter-channel dependencies and suppresses redundant responses in the learned feature space: Hcn5=CFFT(Hcn4),Hfn5=CFFT(Hfn4).H_c^n_5=CFFT(H_c^n_4), H_f^n_5=CFFT(H_f^n_4). (26) Appendix F Loss Function F.1 Loss Function We train with the Negative Pearson Loss [2]: ℒNP(y,y^)=1−ρ(y,y^),L_NP(y, y)=1-ρ(y, y), (27) where ρ(⋅,⋅)ρ(·,·) denotes the Pearson correlation coefficient between the predicted and ground-truth physiological waveforms.