Paper deep dive
GBU-Palm: A Multimodal Video Dataset and Benchmark for Palm Presentation Attack Detection
Yingjie Ma, Zitong Yu, Wei Jia, Ajay Kumar, Linlin Shen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/23/2026, 3:04:57 AM
Summary
The paper introduces GBU-Palm, a large-scale multimodal video dataset and benchmark for palm presentation attack detection (PAD). It contains 21,326 videos from 105 subjects across six acquisition environments, including bona fide, Print, and Replay attacks, with synchronized RGB-NIR samples. The study benchmarks four video architectures (R(2+1)D-18, ViViT, Video Swin-T, MViT-V2-S) under environment-matched and cross-environment settings, revealing significant architecture-dependent performance degradation under environmental shifts. Key findings include that RGB-NIR fusion does not consistently outperform RGB-only input and that temporal order sensitivity varies significantly across models.
Entities (12)
Relation Signals (12)
GBU-Palm → contains → 21,326 videos
confidence 99% · GBU-Palm, a large-scale multimodal video dataset and benchmark containing 21,326 videos from 105 subjects
GBU-Palm → provides → RGB
confidence 95% · synchronized RGB-NIR samples
GBU-Palm → provides → NIR
confidence 95% · synchronized RGB-NIR samples
GBU-Palm → supports → Presentation Attack Detection
confidence 95% · GBU-Palm: A Multimodal Video Dataset and Benchmark for Palm Presentation Attack Detection
R(2+1)D-18 → evaluatedon → GBU-Palm
confidence 92% · We benchmark four representative video backbones... R(2+1)D-18
ViViT → evaluatedon → GBU-Palm
confidence 92% · We benchmark four representative video backbones... ViViT
Video Swin-T → evaluatedon → GBU-Palm
confidence 92% · We benchmark four representative video backbones... Video Swin-T
MViT-V2-S → evaluatedon → GBU-Palm
confidence 92% · We benchmark four representative video backbones... MViT-V2-S
GBU-Palm → includes →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Existing palm presentation attack detection (PAD) datasets are often limited by static imagery, restricted acquisition conditions, or insufficient multimodal video data, hindering systematic evaluation across environments, modalities, and attack types. We present GBU-Palm, a large-scale multimodal video dataset and benchmark containing 21,326 videos from 105 subjects and 210 palms across six acquisition environments, including bona fide, Print, and Replay presentations, with 6,310 synchronized RGB-NIR samples. We construct leakage-controlled protocols that separate palm identity and attack lineage and benchmark four representative video architectures under environment-matched and held-out-environment settings. Results reveal substantial architecture-dependent degradation under environmental shift and show that RGB-NIR fusion does not consistently outperform RGB-only input. We further analyze model behavior through true accept (TA), true reject (TR), false accept (FA), and false reject (FR) decomposition, spectral masking, temporal-order intervention, and frozen-backbone NIR probing, revealing distinct failure patterns and evidence utilization across architectures. GBU-Palm provides a unified and challenging benchmark for developing and evaluating robust multimodal palm PAD methods under cross-environment conditions.
Tags
Links
- Source: https://arxiv.org/abs/2608.14389v1
- Canonical: https://arxiv.org/abs/2608.14389v1
Trouble viewing inline? Open PDF directly →
Full Text
27,506 characters extracted from source content.
Expand or collapse full text
GBU-Palm: A Multimodal Video Dataset and Benchmark for Palm Presentation Attack Detection Yingjie Ma 1,2 , Zitong Yu 2,3,4⋆ , Wei Jia 5 , Ajay Kumar 6 , and Linlin Shen 1,4⋆ 1 Shenzhen University 2 Great Bay University 3 Dongguan Key Laboratory for Intelligence and Information Technology 4 Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University 5 Hefei University of Technology 6 The Hong Kong Polytechnic University Abstract. Existing palm presentation attack detection (PAD) datasets are often limited by static imagery, restricted acquisition conditions, or insufficient multimodal video data, hindering systematic evaluation across environments, modalities, and attack types. We present GBU-Palm, a large-scale multimodal video dataset and benchmark containing 21,326 videos from 105 subjects and 210 palms across six acquisition environ- ments, including bona fide, Print, and Replay presentations, with 6,310 synchronized RGB-NIR samples. We construct leakage-controlled pro- tocols that separate palm identity and attack lineage and benchmark four representative video architectures under environment-matched and held-out-environment settings. Results reveal substantial architecture- dependent degradation under environmental shift and show that RGB- NIR fusion does not consistently outperform RGB-only input. We further analyze model behavior through true accept (TA), true reject (TR), false accept (FA), and false reject (FR) decomposition, spectral masking, temporal-order intervention, and frozen-backbone NIR probing, revealing distinct failure patterns and evidence utilization across architectures. GBU-Palm provides a unified and challenging benchmark for develop- ing and evaluating robust multimodal palm PAD methods under cross- environment conditions. The dataset will be released soon. Keywords: Palm presentation attack detection; Multimodal; RGB-NIR 1 Introduction Palm biometrics provide a convenient and contactless means of authentication, but printed or screen-replayed palm content can be presented directly to the sensor, creating practical presentation attacks. Presentation attack detection (PAD) is therefore essential for reliable deployment. Early studies demonstrated vulnerabilities of palmprint and palm-vein systems to Print and Display/Replay attacks and introduced representative resources such as PALMspoof and VERA ⋆ Corresponding authors arXiv:2608.14389v1 [cs.CV] 14 Aug 2026 2Y. Ma et al. Table 1. Representative palm PAD resources. Scale follows the evaluation unit reported by each source and is not directly comparable across rows. DatasetAccess UnitScale ModalityAttack VERA [9]Request Image2,000NIRPrint PALMspoof [2]Private Image–RGBPrint/display XJTU-PalmReplay [13] Request Image96,000RGBReplay PVASD [11]PublicImage 1,187,519NIR2D/3D GBU-PalmPublic Video21,326 RGB/NIR Print/replay Spoofing PalmVein [2,9]. More recent work has explored cross-domain adapta- tion, domain generalization, frequency cues, and larger NIR palm-vein PAD data [13,6,5,11]. Nevertheless, representative palm PAD resources remain largely based on static imagery, single-spectrum sensing, or limited acquisition conditions, making it difficult to systematically evaluate video PAD generalization across attack types, sensing modalities, and environments. Video and multimodal sensing introduce two important dimensions for palm PAD. Bona fide palms, printed media, and replay displays differ not only in appearance, but also in motion consistency, surface stability, reflections, and display dynamics. Video PAD in other biometric domains has demonstrated the value of spatiotemporal information [7,12]. Meanwhile, RGB and near-infrared (NIR) capture different responses from skin, printed materials, and electronic displays, while prior multimodal PAD studies show that additional sensing channels do not necessarily improve performance [14,3]. These factors may also change across acquisition environments, and palm PAD has already exhibited substantial cross-device and cross-domain degradation [13,6]. Reliable comparison of video, multimodal, and cross-environment generalization therefore requires these factors to be evaluated jointly within a controlled data design. However, as summarized in Table 1, existing palm PAD datasets do not jointly provide large-scale native video, synchronized RGB-NIR observations, multiple acquisition environments, and controlled presentation-attack provenance within a unified benchmark. Consequently, several practically important questions remain difficult to study reproducibly, including how strongly performance degrades under environmental shift, whether RGB and NIR provide complementary information for different architectures, whether video models actually depend on temporal ordering, and whether different attack types and error categories exhibit the same failure trends. Addressing these questions requires not only larger-scale data, but also a standardized benchmark that controls identity, attack origin, modality, temporal structure, and environment. To address this gap, we introduce GBU-Palm, a large-scale multimodal video dataset and benchmark for palm presentation attack detection. GBU-Palm contains 21,326 videos from 105 subjects and 210 palms across six acquisition environments, covering bona fide, Print, and Replay presentations, with 6,310 synchronized RGB-NIR samples. We construct evaluation protocols with dis- joint palm identities and attack lineages and benchmark representative video architectures under RGB, NIR, and RGB-NIR settings in both environment- GBU-Palm: Multimodal Video Palm PAD Benchmark3 Table 2. Presentation-attack factors retained in GBU-Palm metadata. Attack FactorValues PrintSpectrum RGB color, grayscale, NIR PrintMaterialA4, photo, matte, pearl, coated paper PrintCropFull, local, hand-shape, partial real-palm exposure ReplaySpectrum RGB, NIR ReplayInterfaceFull-screen, UI borders matched (In-Env) and held-out-environment (Cross-Env) conditions. Beyond the overall benchmark, we further study model failure modes and spectral-temporal information utilization through TA/TR/FR/FA decision-outcome decomposition, spatiotemporal evidence analysis, spectral masking, temporal-order intervention, and frozen-backbone NIR probing. Our contributions are threefold: – We introduce GBU-Palm, a large-scale multimodal video dataset for palm PAD containing 21,326 videos from 105 subjects and 210 palms across six acquisition environments, including bona fide, Print, and Replay presentations, 6,310 synchronized RGB-NIR samples, and attack-lineage annotations. –We establish standardized, leakage-controlled benchmark protocols with identity- and attack-lineage-disjoint splits, and systematically evaluate repre- sentative video architectures under RGB, NIR, and RGB-NIR inputs in both In-Env and Cross-Env conditions. –Beyond the overall benchmark, we further analyze model error patterns and spatiotemporal and spectral information utilization. The results reveal architecture-dependent cross-environment degradation, non-uniform benefits from RGB-NIR fusion, and substantially different dependence on temporal ordering and NIR information across video models. 2 GBU-Palm Dataset Acquisition and Attack Scenarios. The released GBU-Palm dataset contains 21,326 videos from 105 subjects and 210 palms acquired under six illumination en- vironments (E1–E6) using consumer phones, tablets, laptops, and a synchronized RGB-NIR acquisition system. E1 represents uniform indoor normal illumination, including artificial, mixed artificial-natural, and indirect natural light sampled at 1:1:1±10%; E2 represents indoor low illumination, with weak artificial and weak artificial-natural light sampled at 2:1±10%; E3 represents indoor directional illumination, covering natural or artificial side and back lighting at 1:1±10%; E4 represents uniform outdoor indirect natural light under cloudy, overcast, or shaded conditions; E5 represents uniform outdoor indirect natural light under explicit shelter; and E6 represents outdoor directional natural light with side and back lighting sampled at 1:1±10%. All environments follow a common data organization and ROI convention, enabling acquisition conditions to be compared under a consistent representation. The dataset contains three presentation classes. Bona fide videos record genuine palms. Print attacks recapture printed palm media generated from genuine source samples, whereas Replay attacks present 4Y. Ma et al. B o n a f i d e E1E2 E3E4E5E6 P r i n t R e p l a y R G B N I R N I R N I R R G B R G B Fig. 1. Representative GBU-Palm samples across six acquisition environments (E1–E6) and three presentation classes. RGB and NIR rows show synchronized observations of the same physical presentations, illustrating variation across environment, attack type, and sensing spectrum. palm content through electronic displays and recapture the resulting presentation. For each attack sample, the metadata retains its source bona fide session, attack generation, and material identity, forming an attack lineage. Beyond the attack labels, GBU-Palm retains interpretable generation factors including printing spec- trum, physical medium, crop strategy, replay spectrum, and playback interface, enabling analyses beyond the coarse Print/Replay taxonomy. Native Video and Multimodal Construction. Each GBU-Palm sample preserves a continuous 16-s physical observation window together with its native temporal ordering and frame timestamps, with palm observations stored as canonical 256×256 ROIs. Unlike video sequences constructed from static imagery, these native videos retain motion, reflection, and presentation dynamics produced during physical acquisition. For synchronized RGB-NIR samples, both modalities record the same physical presentation over the same observation window. RGB serves as the common spatial-geometry reference and is mapped to NIR through the fixed acquisition-system geometry, preserving consistent ROI definitions and spatial correspondence across modalities, while both streams retain their native frame timestamps. Dataset Statistics. The released set contains 11,108 bona fide videos, 5,584 Print attacks, and 4,634 Replay attacks. Among the 21,326 samples, 15,016 are RGB-only and 6,310 provide synchronized RGB-NIR observations, including 4,471 bona fide, 1,407 Print, and 432 Replay samples. In total, the canonical media GBU-Palm: Multimodal Video Palm PAD Benchmark5 Table 3. Split composition of the GBU-Palm P1 (In-Env) and P2 (Cross-Env) protocols. RGB+NIR denotes samples with synchronized paired observations. Protocol Split Total Bona fide Print Replay RGB+NIR P1Train 12,1527,4332,0492,6703,740 Val3,9201,6491,2949771,115 Test3,4221,6419028791,031 P2Train4,8273,4554459271,482 Val1,9751,040618317529 Test2,1741,471279424836 store contains 27,636 modality sequences and 5,293,639 frames; unobserved NIR data are not synthesized. All 105 subjects have recorded age and sex metadata, comprising 51 male and 54 female participants. Ages range from 18 to 70 years, with a mean of 44.1±16.6 years and a median of 45. The 18–30, 31–45, 46–60, and 61+ age groups contain 26, 27, 26, and 26 subjects, respectively, resulting in near-balanced distributions across sex and the predefined age groups. Benchmark Design. The GBU-Palm benchmark formulates palm PAD as binary discrimination between bona fide and attack presentations. For each video, models select 32 unique real-frame positions, and the canonical 256×256 ROI is transformed to a 112×112 model input. No temporal interpolation, frame duplication, or fabricated frames are introduced, so all frames presented to the model originate from physically captured video content. RGB, NIR, and RGB-NIR settings follow common training and evaluation rules. We define two complementary evaluation protocols. P1 evaluates environment-matched (In-Env ) performance: subjects are disjoint across training, validation, and test splits, while all six acquisition environments are represented in each partition. P2 evaluates cross-environment (Cross-Env ) generalization: E2, E4, E5, and E6 are used for training and validation, whereas E1 and E3 are completely withheld from training and used exclusively for testing. Both protocols enforce disjoint subjects, palm identities, and attack lineages between training and evaluation partitions. Table 3 reports the resulting split composition. 3 Benchmark Results and Evidence Analysis Experimental Setup. We benchmark four representative video backbones with different temporal modeling inductive biases: R(2+1)D-18 [10], a factorized ViViT baseline adapted from ViViT [1], Video Swin-T [8], and MViT-V2-S [4]. R(2+1)D represents factorized spatiotemporal convolution, ViViT represents transformer- based video modeling, while Video Swin-T and MViT-V2-S cover hierarchical and multi-scale spatiotemporal transformer architectures. This diversity enables analysis of how different video models exploit temporal and spectral evidence in palm PAD. Each model is trained independently under P1 and P2 protocols. RGB-only samples provide RGB streams, while synchronized RGB-NIR samples provide both modalities. The two streams share the same backbone and their available embeddings are averaged before the classification head. This design 6Y. Ma et al. Table 4. Overall video-based PAD results on the complete GBU-Palm test splits. AUC and HTER are reported in percentage; HTER uses the operating threshold fixed on validation. Method P1: In-EnvP2: Cross-Env AUC↑ (%) HTER↓ (%)AUC↑ (%) HTER↓ (%) R(2+1)D-18 [10]97.488.3994.6113.05 ViViT [1]93.4914.5582.2125.73 Video Swin-T [8]72.9532.8971.2036.36 MViT-V2-S [4]97.338.3993.3017.00 Table 5. Environment-specific PAD performance (AUC, %). P1 Range denotes the maximum–minimum AUC across the six P1 environments. P2 values compare the two held-out environments within the same protocol. MethodP1 Best P1 Worst Range P2 E1 P2 E3 R(2+1)D-18E1 / 98.7 E6 / 96.52.295.294.1 ViViTE2 / 95.8 E5 / 91.24.682.681.8 Video Swin-T E6 / 77.1 E4 / 67.110.071.770.6 MViT-V2-SE1 / 98.3 E6 / 96.71.794.692.0 avoids introducing an additional fusion module and keeps the comparison focused on evidence utilization rather than fusion architecture optimization. All models use identical optimization settings. Inputs consist of 32 real-frame positions with spatial resolution of 112×112. Training uses AdamW with learning rate 2×10 −4 , weight decay 10 −4 , a maximum of 100 epochs, and patience of 15. Checkpoints are selected according to validation AUC, and test data are never used for checkpoint selection or threshold determination. We report AUC to measure threshold-independent ranking performance and HTER, defined as the average of false acceptance rate (FAR) and false rejection rate (FRR), to evaluate PAD errors under a fixed decision threshold. Overall benchmark performance. Table 4 reports results on the complete test splits. Under P1 environment-matched evaluation, performance already varies substantially across architectures: R(2+1)D and MViT reach approximately 97.4% AUC, whereas Video Swin-T performs considerably worse. Under P2 Cross-Env evaluation, AUC decreases for all four architectures, but by markedly different amounts: ViViT loses 11.28 points, compared with 2.87 for R(2+1)D, 1.75 for Video Swin-T, and 4.03 for MViT. Cross-environment generalization is therefore strongly architecture-dependent rather than a uniform increase in task difficulty. Environment-specific generalization. As shown in Table 5, environment sensitivity is strongly architecture-dependent. Across the six P1 environments, Video Swin-T spans 10.0 AUC points and ViViT 4.6 points, whereas R(2+1)D and MViT vary by only 2.2 and 1.7 points, respectively. The best and worst environments also differ across architectures, indicating no universal environment- difficulty ordering. Within P2, E3 yields lower AUC than E1 for all four models, with the largest gap of 2.7 points observed for MViT. GBU-Palm: Multimodal Video Palm PAD Benchmark7 FA TA TR FR 퐀 1 퐀 2 퐀 3 퐀 4 Fig. 2. Decision-conditioned examples from GBU-Palm. TA, TR, FA, and FR cases show representative temporal observations (t 1 –t 4 ) and corresponding attack-probability trajectories. The visualization highlights different decision behaviors without making unsupported pixel-level attribution claims. Decision-conditioned error structure. To expose failure directions hidden by aggregate AUC, we decompose PAD decisions into true accept (TA), true reject (TR), false accept (FA), and false reject (FR). TA denotes correctly accepted bona fide samples and TR correctly rejected attacks; FA denotes attacks incorrectly accepted as bona fide and directly reflects security risk, whereas FR denotes bona fide samples incorrectly rejected and reflects usability cost. Figure 2 shows representative temporal observations and attack-probability trajectories for the four outcomes, illustrating that similar aggregate performance changes can correspond to fundamentally different decision failures. This asymmetry is pronounced under environmental shift. MViT FA rises from 7.52% on P1 to 26.32% on P2, while FR decreases from 9.26% to 7.68%. R(2+1)D FA increases from 9.77% to 18.21%, whereas FR remains relatively stable. ViViT, in contrast, increases in both FA and FR. Environmental shift can therefore primarily increase attack-acceptance risk for some architectures while simultaneously affecting both security and usability for others, behavior that is not captured by the magnitude of AUC degradation alone. Attack-family vulnerability. Further decomposing FA by attack family shows that Replay generally presents a higher attack-acceptance risk than Print across model–protocol combinations. For example, under MViT P2, Print FA is 11.47%, whereas Replay FA reaches 36.08%. Cross-environment vulnerability therefore depends not only on model architecture, but also on the physical presentation mechanism. 8Y. Ma et al. Table 6. Controlled spectral and temporal evidence. (a) Paired-test AUC/HTER (%) from the same multimodal-trained checkpoint and identical sample identities. (b) Full-test AUC under normal, shuffled, and reversed temporal order; shuffled results report mean±std over five fixed permutations. All temporal interventions preserve the same set of observed frames. (a) Spectral observation: AUC / HTER Method P1: In-EnvP2: Cross-Env RGBNIRRGB+NIRRGBNIRRGB+NIR R(2+1)D99.45/4.33 84.09/25.03 99.66/3.0296.94/7.97 91.54/16.19 98.65/4.08 ViViT95.61/9.13 67.71/44.54 98.82/4.8283.75/21.43 67.14/45.41 87.15/19.14 Swin-T 87.42/19.74 30.69/65.64 83.67/22.9878.84/25.33 77.97/20.52 82.39/21.92 MViTv2-S97.25/6.66 99.66/2.61 99.74/1.7797.20/6.77 99.07/5.47 99.38/5.58 (b) Temporal-order intervention: AUC Method P1: In-EnvP2: Cross-Env Normal Shuffled ReversedNormal Shuffled Reversed R(2+1)D97.48 93.34±0.3993.2694.61 89.58±0.2490.98 ViViT 93.49 93.49±0.0093.4982.21 82.21±0.0082.21 Swin-T72.95 73.31±0.3572.9071.20 72.21±0.1271.17 MViTv2-S97.33 95.95±0.1394.6993.30 90.90±0.1491.39 Spectral Evidence Analysis. Synchronized RGB-NIR samples enable spectral comparisons under the same physical presentation conditions. This analysis investigates whether additional spectral information can be effectively utilized by different architectures rather than assuming that multimodal input is always beneficial. Table 6 shows that RGB-NIR benefits are strongly architecture- dependent. R(2+1)D improves from 4.33% HTER with RGB input to 3.02% with RGB+NIR on P1, and from 7.97% to 4.08% on P2. ViViT also benefits from additional NIR information. However, Swin increases from 19.74% HTER with RGB to 22.98% with RGB+NIR on P1, indicating that additional modalities do not automatically translate into useful discriminative evidence. The Replay breakdown further shows how spectral changes affect False Accept errors. Since FA represents attacks incorrectly accepted as bona fide, Replay FA directly reflects security impact under different spectral inputs. On P1, NIR-only Replay FA reaches 67%, 93%, and 94% for R(2+1)D, ViViT, and Swin, respectively, while RGB+NIR reduces them to 6%, 0%, and 38%. MViT maintains 0% FA under both conditions. These results indicate that RGB and NIR are not simply additive information sources; their effectiveness depends on whether a model can learn and exploit complementary cross-modal cues. Temporal Evidence Analysis. Temporal-order interventions preserve the same 32 real frames and modify only their ordering. As shown in Table 6(b), R(2+1)D degrades under both shuffling and reversal, while MViT shows smaller but consis- tent drops. Video Swin-T exhibits no stable degradation under either intervention, revealing substantial architecture-dependent differences in temporal-order sen- sitivity. The evaluated ViViT is an adapted factorized baseline rather than a GBU-Palm: Multimodal Video Palm PAD Benchmark9 30507090100 NIR AUC (%) R(2+1)D-18 ViViT Video Swin-T MViT-V2-S 84.198.2 67.790.4 30.750.0 99.799.7 P1: In-Env 30507090100 NIR AUC (%) 91.594.8 67.285.0 78.085.8 99.199.0 P2: Cross-Env Shared headLinear probe Fig. 3. Frozen-backbone NIR diagnosis. The shared head uses the original multimodal classifier, while the linear probe is trained on paired-validation NIR embeddings and evaluated on paired test samples. complete reproduction of the original architecture. It uses no temporal positional encoding and applies temporal mean pooling, making it inherently insensitive to frame permutations; accordingly, neither shuffling nor reversal changes its AUC. Overall, video input does not guarantee temporal-order utilization, and temporal sensitivity alone does not predict Cross-Env robustness. NIR Representation Diagnosis. NIR-only performance differences may orig- inate from different mechanisms. To distinguish classifier-head misalignment from limited linear separability of single-stream representations, we compare the original shared classifier with a frozen-backbone linear probe. Figure 3 il- lustrates this diagnosis: the shared head directly uses the original multimodal classifier, whereas the linear probe freezes the NIR backbone and trains a new linear classifier on paired-validation NIR embeddings before evaluation on paired test samples. Because P1 and P2 use independently trained checkpoints and different test partitions, we interpret the probe as a within-protocol recoverabil- ity analysis rather than a direct comparison of absolute NIR difficulty across protocols. MViT remains near 99% AUC with both the shared head and linear probe, indicating that its original classifier already exploits the available NIR representation effectively. R(2+1)D improves from 84.09% to 98.17% on P1, while ViViT improves from 67.71/67.14% to 90.42/85.01% on P1/P2, indicating that their NIR representations retain substantial linearly accessible information that is not fully utilized by the original shared classifier. In contrast, Swin reaches only 50.00% AUC with the P1 linear probe, indicating weaker linear separability of its single-stream NIR representation under this diagnostic. Thus, NIR-only degradation has no single cause across architectures, but can arise from either classifier adaptation limitations or restricted linear separability of the learned representation. 4 Conclusion We introduced GBU-Palm, a large-scale multimodal video dataset and bench- mark for palm presentation attack detection. Systematic evaluation shows architecture- dependent sensitivity to environmental shift, non-uniform RGB-NIR complemen- tarity, and substantially different reliance on temporal order. Decision-conditioned analysis and NIR probing further reveal that similar aggregate performance can hide different security and usability failures, while single-modality degradation 10Y. Ma et al. may arise from either classifier- or representation-related limitations. Overall, benchmark performance, information utilization, and cross-environment trans- fer are related but distinct properties of palm PAD models. Future extensions will include three-dimensional and more challenging attacks, additional sensing modalities, and richer spatiotemporal annotations. References 1.Arnab, A., Dehghani, M., Heigold, G., Sun, C., Luˇci ́c, M., Schmid, C.: Vivit: A video vision transformer. In: IEEE/CVF International Conference on Computer Vision (ICCV). p. 6836–6846 (2021) 2.Bhilare, S., Kanhangad, V., Chaudhari, N.S.: A study on vulnerability and presen- tation attack detection in palmprint verification system. Pattern Analysis and Ap- plications 21(3), 769–782 (2018).https://doi.org/10.1007/s10044-017-0606-y 3.George, A., Marcel, S.: Cross modal focal loss for rgbd face anti-spoofing. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 7882–7891 (2021) 4.Li, Y., Wu, C.Y., Fan, H., Mangalam, K., Xiong, B., Malik, J., Feichtenhofer, C.: Mvitv2: Improved multiscale vision transformers for classification and detection. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 4804–4814 (2022) 5.Liu, C., Shao, H., Zhong, D.: Learning discriminative palmprint anti-spoofing features via high-frequency spoofing regions adaptation. IET Image Processing 19(1) (2025). https://doi.org/10.1049/ipr2.70029 6.Liu, C., Shao, H., Zhong, D.: Learning domain-adaptive palmprint anti-spoofing feature from multi-source domains. Displays 86, 102871 (2025).https://doi.org/ 10.1016/j.displa.2024.102871 7. Liu, Y., Jourabloo, A., Liu, X.: Learning deep models for face anti-spoofing: Binary or auxiliary supervision. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 389–398 (2018).https://doi.org/10.1109/CVPR.2018. 00048 8.Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., Hu, H.: Video swin transformer. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 3202–3211 (2022) 9.Tome, P., Marcel, S.: On the vulnerability of palm vein recognition to spoofing attacks. In: International Conference on Biometrics (ICB). p. 319–325 (2015). https://doi.org/10.1109/ICB.2015.7139056 10.Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., Paluri, M.: A closer look at spatiotemporal convolutions for action recognition. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 6450–6459 (2018). https://doi.org/10.1109/CVPR.2018.00675 11. Yan, C., Lan, Z., Li, H., Li, Y., Meng, Z.: A comprehensive framework for palm vein anti-spoofing with preprocessing pipeline, dataset, and benchmark. IEEE Transactions on Information Forensics and Security 21, 945–959 (2026).https: //doi.org/10.1109/TIFS.2025.3650391 12.Yang, X., Luo, W., Bao, L., Gao, Y., Gong, D., Zheng, S., Li, Z., Liu, W.: Face anti-spoofing: Model matters, so does data. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 3507–3516 (2019) GBU-Palm: Multimodal Video Palm PAD Benchmark11 13.Yao, D., Shao, H., Zhong, D.: Palmprint anti-spoofing based on domain-adversarial training and online triplet mining. In: IEEE International Conference on Image Processing (ICIP). p. 1235–1239 (2023).https://doi.org/10.1109/ICIP49359. 2023.10223182 14. Zhang, S., Wang, X., Liu, A., Zhao, C., Wan, J., Escalera, S., Shi, H., Wang, Z., Li, S.Z.: A dataset and benchmark for large-scale multi-modal face anti-spoofing. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 919–928 (2019)