Paper deep dive
Detecting Clear Contact Lenses for Iris Recognition: A Two-Stage Mask-Guided Attention Approach
Parisa Farmanifard, Arun Ross
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/12/2026, 3:07:15 AM
Summary
This paper proposes a two-stage framework for detecting clear contact lenses in iris recognition to mitigate verification errors. Stage 1 uses D-NetPAD to filter patterned lenses, while Stage 2 employs a ConvNeXt-Base model with a novel Mask-Guided Spatial Attention (MGSA) module to distinguish clear lenses from normal irises. The MGSA module integrates a Hough-derived anatomical ROI mask with learned spatial attention and Squeeze-and-Excitation channel recalibration. The method achieves 90.0%-98.8% accuracy across four datasets and reduces Equal Error Rate (EER) by 4.1%-28.3% via z-score calibration of VeriEye match scores.
Entities (10)
Relation Signals (8)
ConvNeXt-Base â equippedwith â Mask-Guided Spatial Attention (MGSA)
confidence 95% · ConvNeXt-Base model equipped with Mask-Guided Spatial Attention (MGSA)
Mask-Guided Spatial Attention (MGSA) â incorporates â Hough-derived anatomical ROI mask
confidence 95% · The proposed MGSA module incorporates a Hough-derived anatomical ROI mask
Two-Stage Mask-Guided Attention Approach â uses â D-NetPAD
confidence 95% · Stage 1 uses an existing PAD model known as D-NetPAD
Two-Stage Mask-Guided Attention Approach â uses â ConvNeXt-Base
confidence 95% · Stage 2 focuses on the more challenging clear-lens versus no-lens distinction using a ConvNeXt-Base model
Two-Stage Mask-Guided Attention Approach â improves â Iris Recognition
confidence 92% · reliable clear contact lens detection can directly improve iris verification performance
Clear Contact Lenses â degrades â Iris Recognition
confidence 90% · clear lenses marginally degrade genuine match scores and increase verification error
Mask-Guided Spatial Attention (MGSA) â incorporates â Squeeze-and-Excitation
confidence 90% · MGSA module incorporates ... Squeeze-and-Excitation channel recalibration
VeriEye â usedby â
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This work focuses on the impact and detection of clear contact lenses in the context of iris recognition. While the detection of cosmetic or patterned contact lenses has been extensively studied under the presentation attack detection (PAD) paradigm, clear prescription contact lenses, that are typically transparent, have received comparatively less attention despite their widespread use. Unlike patterned lenses, clear lenses introduce no salient texture artifact, making them difficult to detect and are often assumed to have no impact on iris recognition. We first examine this assumption using the commercial VeriEye matcher on four benchmark datasets and show that clear lenses marginally degrade genuine match scores and increase verification error. We then propose a two-stage contact-lens detection framework. Stage~1 uses an existing PAD model to identify patterned lenses, while Stage~2 focuses on the more challenging clear-lens versus no-lens distinction using a ConvNeXt-Base model equipped with Mask-Guided Spatial Attention (MGSA). The proposed MGSA module incorporates a Hough-derived anatomical ROI mask together with learned spatial attention and Squeeze-and-Excitation channel recalibration, allowing the network to focus on subtle limbal cues associated with clear lens wear. Across four datasets, the full pipeline consisting of both patterned and clear contact lens detection achieves between 90.0\%--98.8\% accuracy. Finally, we introduce a z-score calibration method that adjusts VeriEye match scores when a clear lens is detected in the input images. This calibration reduces EER by 4.1\%--28.3\% across datasets, demonstrating that reliable clear contact lens detection can directly improve iris verification performance.
Tags
Links
- Source: https://arxiv.org/abs/2608.08977v1
- Canonical: https://arxiv.org/abs/2608.08977v1
Trouble viewing inline? Open PDF directly â
Full Text
52,074 characters extracted from source content.
Expand or collapse full text
Paper accepted at IEEE/IAPR International Joint Conference on Biometrics (IJCB), September 2026. Detecting Clear Contact Lenses for Iris Recognition: A Two-Stage Mask-Guided Attention Approach Parisa Farmanifard and Arun Ross Michigan State University, East Lansing, MI 48824 farmanif,rossarun@msu.edu Abstract This work focuses on the impact and detection of clear contact lenses in the context of iris recognition. While the detection of cosmetic or patterned contact lenses has been extensively studied under the presentation attack detection (PAD) paradigm, clear prescription contact lenses, that are typically transparent, have received comparatively less attention despite their widespread use. Unlike patterned lenses, clear lenses introduce no salient texture artifact, making them difficult to detect and are often assumed to have no impact on iris recognition. We first examine this assumption using the commercial VeriEye matcher on four benchmark datasets and show that clear lenses marginally degrade genuine match scores and increase verification error. We then propose a two-stage contact-lens detection framework. Stage 1 uses an existing PAD model to identify patterned lenses, while Stage 2 focuses on the more challenging clear-lens versus no-lens distinction using a ConvNeXt-Base model equipped with Mask-Guided Spatial Attention (MGSA). The proposed MGSA module incorporates a Hough-derived anatomical ROI mask together with learned spatial attention and Squeeze-and-Excitation channel recalibration, allowing the network to focus on subtle limbal cues associated with clear lens wear. Across four datasets, the full pipeline consisting of both patterned and clear contact lens detection achieves between 90.0%â98.8% accuracy. Finally, we introduce a z-score calibration method that adjusts VeriEye match scores when a clear lens is detected in the input images. This calibration reduces EER by 4.1%â28.3% across datasets, demonstrating that reliable clear contact lens detection can directly improve iris verification performance. 1 Introduction Original ROI Mask Iris Crop Clear Patterned Normal Figure 1: Sample inputs to the proposed pipeline for each lens class. Each row shows the full periocular image, the anatomical ROI mask M (Eq. 3), and the localised iris crop I I used by Stage 1. In the mask column: red = pupil region, green = iris annulus, blue = outer arc beyond the limbus where clear-lens evidence appears. The iris is the annular region of the eye surrounding the pupil and is a powerful biometric cue whose intricate texture can even be used to distinguish identical twins [8, 9]. However, this efficacy depends on an implicit assumption: that the captured iris is an unobstructed, unaltered view of bona fide ocular tissue. Contact lensesâincluding patterned cosmetic lenses and the far more prevalent clear111In the iris presentation attack detection (PAD) literature, clear contact lenses are also commonly referred to as transparent or soft contact lenses [14]. lensesâchallenge this assumption in clinically distinct ways (samples can be seen in Fig.1). This is a widespread concern, as more than 140 million people worldwide were estimated to wear contact lenses in 2025 [31]. Patterned lenses are a well-studied threat. Cosmetic (patterned) contact lenses typically overlay a bold222Although most cosmetic contact lenses studied in the iris biometrics literature have a âstrongâ textured pattern, there are other such lenses with subtle texture patterns. synthetic texture over the iris, corrupting the Gabor-phase feature codes used for iris comparison and constituting a well-documented presentation attack [4, 17, 43, 3, 40, 41]. They are detectable precisely because the foreign texture is often visually conspicuous: a compact, well-aligned iris crop reveals the periodic pattern, and even lightweight classifiers can learn to separate it from a natural iris with high reliability [24, 17, 15, 38]. On the other hand, clear lenses are an underexplored challenge. Clear contact lensesâworn by people for vision correctionâhave attracted far less attention in the iris literature, on the assumption that an optically clear lens leaves the iris appearance unchanged. This assumption does not always hold in practice. Baker et al. [2] first showed, across more than 12,000 images, that clear lenses produce false non-matches through subtle deformation and edge artifacts at the iris limbus. Later work also confirmed that this degradation is not random noise but a systematic, condition-dependent bias: the refractive index mismatch at the lens boundary creates a predictable reflectance deviation in the periocular ring [12, 39]. Later, another paper [29], proposed ContlensNet, a method to detect clear contact lenses in ocular images. Despite over a decade of research, detecting clear contact lenses still remains an open problem [27]. The difficulty arises from a combination of factors that together make it a challenging fine-grained classification task in biometrics: âą Optical invisibility: Clear lenses carry no colour or texture signature, unlike cosmetic contact lenses. The only physical evidence is a faint reflectance deviation at the irisâsclera limbus, invisible to the naked eye and easily confused with natural periocular variation. âą Extreme spatial localization: Clear lens evidence is often confined to a narrow annular band at the lens edge. Standard classification backbones trained to aggregate global information tend to ignore this small region in favour of more salient but uninformative background features. âą High intra-class variability: Different lens brands, geometries, and moisture levels produce different limbal signatures, further modulated by eye colour, pupil dilation, and illumination angle. âą Two fundamentally different sub-problems under one label: Patterned and clear lenses share very less commonality. A detection model trained on the bold periodic textures of a cosmetic lens is suboptimal for clear lenses. âą Sensor and acquisition variability: Iris sensors differ in resolution, near-infrared wavelength, and depth of field, all of which alter the appearance of the lens edge and surrounding periocular tissue across devices. Our approach, motivated by the fundamental disparity between the two detection sub-problems, i.e., cosmetic patterned lens detection and clear lens detection, utilizes a two-stage method pipeline. Stage 1 uses an existing iris PAD model known as D-NetPAD [33] to screen for patterned lens artefacts from a compact iris crop and passing only non-patterned images to Stage 2. Stage 2 introduces a novel Mask-Guided Attention Network that injects an anatomical ROI mask derived from classical Hough-based iris localization [9], with no learned segmentation and no pixel-level annotation, directly into the backboneâs multi-level feature representation. The mask encodes the hard spatial prior that all clear lens evidence is often confined to the iris annulus and the periocular arc at the limbal boundary. Fusing this prior multiplicatively into the feature representation forces the network to commit to the anatomically correct region, eliminating the shortcut of latching onto uninformative background cues. A learned spatial gate within the same module further modulates which parts of the masked region are most diagnostic, while a Squeeze-and-Excitation block [18] recalibrates feature channels globallyâtogether asking where to look and what to look for. We implement Stage 2 with three backbone architectures to assess how discriminative capacity and receptive field size interact with mask guidance: EfficientNet-B4 [34], ResNet-101 [16], and ConvNeXt-Base [22]. Once a clear contact lens is detected, we apply a z-score calibration for adjusting verification scores. This work makes the following contributions: 1. Quantitative evidence that clear lenses degrade iris verification. We quantify, across four large-scale datasets using a commercial iris-matching SDK (VeriEye) [26], the systematic genuine-score drop introduced by clear lenses. 2. A two-stage method for contact-lens detection. Decomposes the three-class classification problem (ânormalâ, âclearâ, âpatternedâ) into a texture sub-problem (Stage 1, patterned vs. non-patterned) and a spatial localization sub-problem (Stage 2, clear vs. normal), substantially outperforming single-stage baselines. 3. A novel mask-guided attention mechanism. Injects an annotation-free anatomical ROI mask into the intermediate feature maps of a deep backbone. This combines a learned spatial attention gate with a fixed anatomy-based prior, allowing the contact-lens detector to focus on the most relevant iris regions during feature learning. Moreover, a backbone study comparing EfficientNet-B4, ResNet-101, and ConvNeXt-Base under identical mask-guidance conditions, showing that large-kernel depthwise convolutions best resolve the subtle limbal evidence of clear lens wear. 4. Downstream verification improvement. Upto 28.3% relative EER reduction via predicted-label score calibration. 2 Related Work Iris Presentation Attack Detection: Iris PAD has been an active area of research since the discovery that printed irides, replayed videos, and textured contact lenses can successfully fool commercial matchers [6, 30]. Early methods extracted hand-crafted descriptors, viz., BSIF, LBP, and co-occurrence matrices, from the iris annulus and fed them to shallow classifiers [1]. Deep learning substantially raised the bar: Menotti et al. [24] showed that convolutional networks fine-tuned end-to-end on iris crops outperform handcrafted pipelines across multiple attack types. D-NetPAD [33] built on DenseNet-121 [19] with interpretability constraints, demonstrated strong generalization to cosmetic lens artifacts across sensors [7]; we adopt it as our Stage 1 patterned-lens filter. More recent work has explored multi-task and attention-based formulations. Chen and Ross [5] proposed an attention-guided PAD framework that localizes discriminative iris regions using channel and spatial attention gates. The LivDet-Iris competition series [42, 7, 35] provides standardised cross-sensor benchmarks; the 2023 edition introduced GAN-generated irides as an attack category, highlighting the growing diversity of the threat landscape. The recent top-performing LivDet-Iris competition [25] achieved good performance, showing that PAD is a well-studied problem for patterned lenses. Clear Lens Detection: Unlike patterned contact lenses, clear lenses leave no strong texture signature, making detection fundamentally harder. Baker et al. [2] established the empirical baseline, demonstrating that clear lenses increase false non-match rates even for high-quality NIR sensors. Doyle and Bowyer [12] later showed that the lens boundary introduces a measurable 3D surface distortion detectable with photometric stereo, though this approach requires controlled multi-light acquisition. Raghavendra et al. [29] proposed ContlensNet, a 15-layer CNN that classifies normalized iris patches as no lens, soft lens, or textured lens and determines the image-level label through majority voting. Yadav et al. [39] conducted a large-scale analysis of the effect of textured and incidentally clear lenses on commercial matchers, quantifying the score drop across multiple cohorts. Unlike ContlensNet [29], which only reports lens classification performance, our work goes beyond that by reporting iris recognition and score calibration results. Further, we use the pre-normalized ocular image, which includes the scleral region, thereby capturing the outer boundary of lenses. Attention in Deep Recognition Models: Attention mechanisms have become a standard tool for directing network capacity to informative spatial regions [18, 37]. The Squeeze-and-Excitation (SE) block [18] introduced channel-wise recalibration via global average pooling, allowing the network to selectively amplify informative feature channels. CBAM [37] extended this with a sequential spatial gate, computing a 2D attention map from both average- and max-pooled features. Non-local networks [36] and Vision Transformers [11] model long-range dependencies via self-attention, but at a cost in data efficiency on small biometric datasets. In the medical imaging domain, mask-guided attention networks have been proposed to incorporate anatomical segmentations as spatial priors, improving sensitivity to lesion regions without requiring lesion-level supervision [20]. Our approach is inspired by this line of work: rather than learning attention maps entirely from the training signal, we inject a Hough-derived iris ROI mask as a hard prior directly into the backboneâs intermediate feature hierarchy, coupling it with a learned spatial gate and a SE recalibration block. This is, to our knowledge, the first application of anatomy-driven external mask guidance to contact-lens detection. Backbone Architectures: ResNet [16] established deep residual learning as the standard baseline for visual recognition, and it remains competitive on biometric tasks due to its simplicity and well-understood inductive bias. EfficientNet [34] introduced compound coefficient scaling of width, depth, and resolution, achieving strong accuracy-efficiency trade-offs on ImageNet while keeping parameter counts low attractive for resource-constrained biometric deployment. ConvNeXt [22] revisited the ResNet design space in light of Vision Transformer innovations, replacing 3Ă33\!Ă\!3 convolutions with 7Ă77\!Ă\!7 depthwise kernels, adopting LayerNorm, and applying the inverted major challenge structure; the resulting architecture matches or exceeds Swin Transformer on standard benchmarks while remaining fully convolutional. For our task, the enlarged receptive field of ConvNeXt is especially relevant: the limbal evidence of a clear lens spans a narrow but spatially extended annular band, and larger kernels are better positioned to capture the subtle curvature and reflectance gradients that distinguish a lens boundary from natural periocular tissue. We compare all three backbones under identical mask-guidance conditions to isolate the contribution of receptive field size to clear-lens discriminability. 3 Proposed Methodology 3.1 Pipeline Overview Figure 2: Two-Stage Contact Lens Detection Framework. Top: Stage 1 uses cropped iris images to detect patterned lenses, while Stage 2 uses full periocular images and ROI masks to classify clear vs. normal lenses. Bottom: At inference, images predicted as non-patterned by Stage 1 are passed to Stage 2 for final classification. I denotes the full iris image, M denotes the corresponding mask, and I I denotes the cropped iris image. Contact lens detection is not a single problem; it is two fundamentally different problems sharing the same output space. Patterned lenses introduce a strong global texture anomaly detectable from a compact iris crop, while clear lenses leave only a faint annular trace in the vicinity of the limbal boundary. Treating both with a single three-class classifier forces the network to simultaneously learn coarse texture suppression and subtle annular contrast differences, which compete for representational capacity and degrade performance on the harder class. We address this asymmetry with a two-stage method that routes each image through two specialized binary decisions: y^=patterned,if âf1â(I^)=patterned,clear,if f1â(I^)=non-patternedâ§f2â(I,M)=clear,normal,otherwise. y= cases patterned,&if f_1( I)= patterned,\\[3.0pt] clear,& aligned if &\ f_1( I)= non-patterned\\ & \ f_2(I,M)= clear, aligned\\[6.0pt] normal,&otherwise. cases (1) where, I I is a localized iris crop fed to Stage 1, I is the full periocular image fed to Stage 2, and M is an automatically derived anatomical ROI mask. f1:I^âpatterned,non-patternedf_1: Iâ patterned,\ non-patterned is the Stage 1 D-NetPAD binary classifier, and f2:(I,M)âclear,normalf_2:(I,M)â clear,\ normal is the Stage 2 ConvNeXt-Base + MGSA classifier. Positive Stage 1 detections exit immediately; only non-patterned images reach the computationally heavier Stage 2, keeping the clear-lens sub-problem free from easy cases that would otherwise dominate the gradient. Fig. 2 illustrates the complete pipeline. 3.2 Iris Masks The fundamental spatial prior for clear-lens detection is that all evidence is confined to the iris and the narrow band immediately surrounding it at the limbus. We encode this prior as a binary mask Mâ0,1HĂWMâ\0,1\^HĂ W derived entirely from classical iris geometry so that no learned segmentation and no pixel-level annotation are required. For pupil and limbus localization, we apply the circular Hough transform [32] twice to each greyscale image. The circular Hough transform [32, 28] is a voting-based method that finds circles in an edge image by accumulating evidence for all (x,y,r)(x,y,r) triples consistent with detected edge points. For the pupil, the image is first median-blurred and thresholded to isolate the dark pupil disk; Canny edges are extracted and fed to HoughCircles; the accumulator threshold is lowered iteratively until at least ten candidate circles are returned. The final pupil circle (cxp,cyp,rp)(c_x^p,\,c_y^p,\,r_p)âwhere cxp,cypc_x^p,c_y^p are the horizontal and vertical coordinates of the pupil centre and rpr_p is its radiusâis the coordinate-wise mean of the accepted candidates, which suppresses outliers from eyelash occlusion. The limbus circle (cxâ,cyâ,râ)(c_x ,\,c_y ,\,r_ ) is found with the same multi-scale search, constrained so that its centre lies within 0.25ârp0.25\,r_p of the pupil centre and its radius exceeds 1.5ârp1.5\,r_p. Three-region labelling. Given the two detected circles, every pixel =(x,y)p=(x,y) is assigned one of three anatomical labels: ââ()=pupil,if âââpâ2â€rp,iris,if ârp<ââââ2â€râ,outer_arc,if râ<ââââ2â€râ+ÎŽâ§Îžâ()âÎ,0,otherwise. (p)= cases pupil,&if \|p-c^p\|_2†r_p,\\[-1.0pt] iris,&if r_p<\|p-c \|_2†r_ ,\\[-1.0pt] outer\_arc,& aligned if &\ r_ <\|p-c \|_2†r_ +ÎŽ\\ & \ Ξ(p)â , aligned\\[-1.0pt] 0,&otherwise. cases (2) The outer_arc region captures the periocular tissue immediately beyond the limbus, where the lens edge creates its most visible reflectance deviation. It is restricted to the inferior and superior quadrants Î=[45â,135â]âȘ[225â,315â] =[45 ,135 ]âȘ[225 ,315 ] (measured from the horizontal), since eyelid occlusion dominates the nasal and temporal sectors and provides no lens cue. The radial offset ÎŽ=35ÎŽ=35 px beyond the detected limbus accommodates localization uncertainty. The final binary mask is the union of all three labelled regions: Mâ()=â[ââ()âpupil,iris,outer_arc].M(p)=1\! [ (p)â\ pupil,\; iris,\; outer\_arc\ ]. (3) Fig. 1 shows representative masks for each lens class. However, the generated masks are not always robust, leaving room for further improvement. Since mask quality directly affects the final classification decision, accurate mask generation is crucial to the overall performance of the proposed method. 3.3 Stage 1 â Patterned Lens Screening Stage 1 solves a binary classification problem: does the image contain a patterned (cosmetic) lens? Patterned lenses impose a bold, periodic synthetic texture over the iris that is immediately apparent in a well-aligned iris crop I I of size 256Ă256256Ă 256. We adopt D-NetPAD [33], a DenseNet 121 [19] fine-tuned specifically for iris presentation attack detection, as our Stage 1 classifier, f1f_1. D-NetPAD is pre-trained on a large pool of iris PAD data and generalizes well to the cosmetic lens texture anomaly; we fine-tune only its classification head on each target dataset. Images predicted as patterned exit the pipeline immediately with label patterned. Images not predicted as patterned are passed to Stage 2, which now operates on a cleaner distribution free of the dominant texture cue. 3.4 Stage 2 â Mask-Guided Attention Network Stage 2 addresses the truly challenging sub-problem: distinguishing a bare iris (normal) from one wearing a clear lens (clear). Unlike patterned lenses, which alter the global iris texture, clear lenses preserve it for the most part. The only physical evidence of their presence is a faint reflectance discontinuity confined to the narrow annular band at the iris limbus. Two design decisions follow directly from this observation. First, we use the full periocular image I at its native resolution (480Ă640480Ă 640) rather than a tight iris crop, to preserve the spatial extent of the limbal region. Second, we introduce a Mask-Guided Spatial Attention (MGSA) module that injects the anatomical mask M directly into the backboneâs intermediate feature hierarchy, preventing the model from attending to periocular background structures that carry no lens signal. 3.4.1 Backbone We instantiate Stage 2 with ConvNeXt-Base [22] pre-trained on ImageNet-1K [10]. The ConvNeXt architecture consists of a stem followed by four stages with progressively downsampled feature maps; channel dimensions for the Base variant are [128,256,512,1024][128,256,512,1024] at stages 0â3 respectively. We inject MGSA after Stage 2 of the backbone (C=512C=512 channels), where the feature map retains sufficient spatial resolution to resolve the limbal cues while encoding semantically meaningful mid-level representations. Following the attention module, a global average pool and a dropout layer (p=0.4p=0.4) precede the two-class linear head. We compare ConvNeXt-Base against ResNet-101 [16] and EfficientNet-B4 [34] under identical MGSA conditions in our ablation study (Sec. 4). 3.4.2 Mask-Guided Spatial Attention (MGSA) Let ââBĂCĂHâČĂWâČ F ^BĂ CĂ H Ă W be the feature map produced after the chosen backbone stage. The core insight driving MGSA is that two complementary signals are needed to localize the lens cue: (i) where to look, encoded by the anatomical mask M, and (i) what to look for, encoded by a learned spatial saliency gate. Neither signal alone is sufficient: the mask without learning provides no discriminative weighting within the iris, and a learned gate without the mask may attend to eyelid or periocular skin features that correlate with the label only due to dataset biases. The mask is first binarized and upsampled by nearest-neighbour interpolation to match the feature-map resolution: M~=Upsampleâ(M,(HâČ,WâČ))â0,1HâČĂWâČ. M=Upsample(M,\,(H ,W ))\;â\;\0,1\^H Ă W . (4) A learned saliency gate is produced by a single 1Ă11Ă1 convolution that collapses the channel dimension to a scalar confidence map: G=Ïâ(gâ)â(0,1)1ĂHâČĂWâČ,gââ1ĂCĂ1Ă1.G=Ï\! (w_g F )\;â\;(0,1)^1Ă H Ă W , _g ^1Ă CĂ 1Ă 1. (5) The MGSA output is the element-wise triple product: â=âGâM~, F^*= F\; \;G\; \; M, (6) where â denotes spatially broadcast element-wise multiplication. M~ M acts as a hard anatomical gate: it unconditionally zeros all activations outside the iris region, regardless of what the network has learned. G provides a soft learned re-weighting of positions within that constrained region, learning to further up-weight the outer annulus and limbal arc where the lens boundary appears. SE Channel Recalibration. After spatial gating, a Squeeze-and-Excitation (SE) block [18] recalibrates â F^* channel-wise with reduction ratio r=16r=16: â=ââÏâ(2âReLUâ(1âGAPâ(â))), F^**= F^*\; \;Ï\! (W_2\,ReLU\! (W_1\,GAP( F^*) ) ), (7) where GAPGAP denotes global average pooling. MGSA gates the spatial (HĂW) dimensions of the C-channel feature map, zeroing activations outside the iris annulus uniformly across all channels; SE then reweights those same C channels globally, selecting which feature types encode the lens signal most strongly. The two are complementary: MGSA determines where to look; SE determines what to look for within that region. 3.5 Training and Testing Subject-stratified cross-validation. Biometric datasets exhibit strong intra-subject correlation: all images of the same identity share iris texture, colour, and periocular appearance independent of lens type. Naive random splitting allows the same subject to appear on both sides of a fold boundary, enabling the model to exploit identity-specific cues rather than lens appearanceâa form of identity leakage that inflates validation accuracy and masks poor generalization. We eliminate this with subject-stratified 5-fold cross-validation. (i) identity disjointnessâall images of the same subject are confined to a single fold; and (i) class-balance preservationâthe three-class distribution is maintained across folds. The held-out test split (20% of subjects, never seen during training or model selection) is fixed at the dataset level. The best model checkpoint for each fold is selected by balanced accuracy on that foldâs held-out training subjects. Here âcheckpointâ denotes the saved model weights from the best-performing epoch during training. âHeld-out training subjectsâ are the subjects kept aside from the training fold as the validation data to decide which epoch is best â this is different from the final test set. âBalanced accuracyâ is used instead of just âaccuracyâ because the two classes (clear vs. normal) may not have equal numbers of images. Balanced accuracy averages the per-class recall, so a model that ignores the minority class cannot score well. Ensemble inference. At test time, the five independently trained Stage-2 models fkk=15\f_k\_k=1^5 are combined by averaging their softmax posterior distributions: p^â(câŁI,M)=15ââk=15pkâ(câŁI,M),y^=argâmaxcâĄp^â(câŁI,M). p(c I,M)= 15 _k=1^5p_k(c I,M),\; y= *arg\,max_c p(c I,M). (8) Since each fold is trained on a disjoint subject population, the five models are complementary in the identities they have observed; averaging their posteriors reduces decision-boundary variance without any additional computational cost at training time. For the optimization, stage 2 is trained with AdamW [23] (λ=10â4λ=10^-4), cosine annealing (η0=2Ă10â4 _0=2Ă10^-4, ηmin=10â6 _ =10^-6, T=80T=80 epochs), cross-entropy loss with label smoothing (Δ=0.05 =0.05), and dropout (p=0.4p=0.4) before the classification head. Stage 1 is fine-tuned with SGD (η0=5Ă10â4 _0=5Ă10^-4, step decay Ă0.1Ă 0.1 every 10 epochs, 50 epochs total, weight decay 10â410^-4). 3.6 Lens-informed Score Calibration Once the two-stage classifier assigns a label to every enrolled and probe image, we apply a per-condition z-score calibration to the raw VeriEye verification scores. For each pair type tâN,C,CNtâ\ N,\, C,\, CN\, we estimate the mean ÎŒt _t and standard deviation Ït _t from the genuine pairs of that type in the training set. Every score is then rescaled onto the Normal-Normal (N) reference distribution: scal=sâÎŒtÏtâÏN+ÎŒN,s_cal= s- _t _t\, _ N+ _ N, (9) where, (ÎŒN,ÏN)( _ N,\, _ N) are the parameters of the N reference distribution of the training set. This mapping is a structure-preserving affine transform: it shifts the mean of each non-N distribution onto the N mean and matches its spread, closing the score gap introduced by the lens without altering the rank ordering within any single pair-type group. Crucially, the pair type for each verification pair is determined entirely by the classifierâs predicted labelsâno ground-truth lens information is used at inference time.333We also performed calibration using ground-truth lens labels and obtained the same accuracy, confirming that the predicted labels were sufficient for this evaluation. 4 Experiments and Results 4.1 Contact Lens Classification Table 1: Two-stage contact-lens detection results across all four datasets. Stage 1 (D-NetPAD) classifies patterned vs. non-patterned irises. Stage 2 (ConvNeXt-Base + MGSA) classifies normal vs. clear irises on the non-patterned subset; therefore, Stage 2 accuracy is computed only on non-patterned test images. Accuracy ± values denote bootstrap standard errors from 10,000 resamples. Train/test splits are subject-stratified with no identity overlap. *Approximately 50 images with apparent labeling inconsistencies were excluded from the UND and IITD datasets during preprocessing; however, these inconsistencies were not exhaustively verified across the full datasets. bBaseline CCR (%) reproduced from ContlensNet [29] using intra-sensor validation from Table 1. ContlensNet was trained and tested using its own IIITD/ND protocol and a different number of images; therefore, its results are not directly comparable to the subject-stratified splits used here. Dataset Dataset Statistics* Accuracy (%) Subject-Stratified Split Imgs IDs Pat. Nor. Clr. S1 S2 Pipe. ContlensNetb Tr. IDs Tr. N Te. IDs Te. N UND-AD100 [13] 300 41 100 100 100 100.00100.00 87.50±5.2887.50± 5.28 90.00±4.2390.00± 4.23 95.00 32 250 9 50 UND-LG4000 [13] 1200 76 400 400 400 100.00100.00 92.00±2.2292.00± 2.22 95.20±1.3595.20± 1.35 96.91 61 950 15 250 IITD-Cogent [21, 38] 3508 101 1202 1163 1143 100.00100.00 96.15±0.8796.15± 0.87 97.50±0.5697.50± 0.56 86.73 79 2747 22 761 IITD-Vista [21, 38] 2906 101 1005 940 961 100.00100.00 98.18±0.6998.18± 0.69 98.80±0.4598.80± 0.45 87.33 81 2322 20 584 Average â â â â â 100.00 93.4693.46 95.3895.38 91.4991.49 â â â â S1: Stage 1; S2: Stage 2 evaluated on the non-patterned subset only; Pat.: patterned; Nor.: normal; Clr.: clear; Pipe.: complete pipeline; Tr.: training split; Te.: test split. Table 2: Per-class correct predictions on a test split (correct / total). âTotalâ is the number of test images, and âcorrectâ is the correctly classified test images. Dataset Patterned Normal Clear Overall UND-AD100 10/10 (100.0%) 16/20 (80.0%) 19/20 (95.0%) 45/50 (90.00%) UND-LG4000 100/100 (100.0%) 80/80 (100.0%) 58/70 (82.9%) 238/250 (95.20%) IITD-Cogent 268/268 (100.0%) 240/243 (98.8%) 234/250 (93.6%) 742/761 (97.50%) IITD-Vista 200/200 (100.0%) 192/192 (100.0%) 185/192 (96.4%) 577/584 (98.80%) Table 3: VeriEye EER (%) before and after score calibration. Calibration parameters are estimated from genuine and impostor pairs; the presence or absence of a lens is determined by the proposed two-stage classifier (real-world protocol, no oracle labels used). Dataset Before (%) After (%) Rel. Red. UND-AD100 0.49 0.47 â4.1%-4.1\% UND-LG4000 0.75 0.64 â14.7%-14.7\% IITD-Cogent 0.76 0.70 â7.9%-7.9\% IITD-Vista 2.26 1.62 â28.3%-28.3\% Table 1 reports Stage-1, Stage-2, and full pipeline accuracy together with per-class scores across all four datasets. Table 2 provides the per-class correct-prediction breakdown. Stage-1 achieves perfect patterned detection. Across all four datasets, Stage-1 (D-NetPAD) correctly classifies every patterned iris in the test set â 100% accuracy with zero bootstrap SD (Table 2). Patterned lenses introduce salient colour and periodic texture artefacts that are easily separable from natural iris textures, confirming that binary patterned/non-patterned classification is a well-conditioned problem for a standard convolutional network fine-tuned from ImageNet weights. Clear lens detection is the harder sub-problem. The main classification challenge lies in Stage-2. Across all datasets, the Clear lens accuracy is the lowest of the three classes, confirming that clear lens detection is harder than patterned detection and also harder than normal iris detection. As shown in Table 1, the gap is most pronounced on AD100 (S2Acc._Acc. = 87.50%) and LG4000 (S2Acc._Acc. = 92.00), which are the smallest datasets: with fewer training subjects, the model has less opportunity to learn the subtle reflectance deviation at the limbal boundary that distinguishes a clear lens from a bare iris. MGSA drives good Stage-2 accuracy on large datasets. On the two large IITD datasets, the MGSA-equipped ConvNeXt-Base achieves 96.15% (Cogent) and 98.18% (Vista) Stage-2 accuracy. By hard-zeroing all activations outside the iris and periocular arc via the anatomical mask M~ M, the network cannot exploit background structures; the soft learned gate G then refines attention within that constrained region to the specific annular sub-areas where lens evidence physically appears. For cross-sensor analysis, despite substantial differences in sensor type, image resolution, and subject demographics across the four datasets, the pipeline achieves 90.0%â98.8% accuracy with a bootstrap SD between 0.45 and 4.23. This cross-sensor consistency supports the generalization of the proposed two-stage method. 4.2 Score and Calibration Figure 3: VeriEye genuine-pair score distributions before (dashed/faded) and after (solid) lens-informed score calibration, for all four datasets. Blue: NormalâNormal (reference, unchanged). Amber: ClearâClear before â after. Red: ClearâNormal before â after. The Î annotation marks the mean gap between NormalâNormal and ClearâClear before calibration. Pair types are determined by the two-stage classifier predictions (no ground-truth labels used). Figure 4: VeriEye ROC curves before (dashed) and after (solid) score calibration for all four datasets. EER values are annotated in the legend. Calibration uses pair types predicted by the proposed two-stage classifier under a real-world protocol (no oracle labels, 100% pair coverage). Clear lenses cause systematic score degradation. Fig. 3 visualizes the VeriEye genuine-pair score distributions for each pair type before calibration. In every dataset, Clear-Clear (C) genuine pairs score measurably lower than Normal-Normal (N) pairs on the same subjects. The score gap ranges from 36 points on AD100 to 39 points on IITD-Vista, consistent with the systematic bias introduced by the refractive index mismatch at the lens boundary. ClearâNormal (CN) cross-condition pairs exhibit an even larger displacement on IITD-Vista (ÎŒCN _CN = 650 vs. ÎŒN _N = 830), pushing genuine match scores closer to the impostor distribution (grey curve, concentrated near zero). This confirms that clear lenses introduce a predictable, condition-dependent bias in iris matchingânot random noiseâand that the bias is present across all four sensors and cohorts tested. Calibration closes the gap. After applying the z-score calibration of Eq. (9) with pair types predicted by the two-stage classifier, the C and CN distributions shift to align with the N reference (Fig. 3, solid curves). On IITD-Vista, the CN mean rises from 650 to 830, completely closing the 180-point gap. On AD100, the C distribution (amber solid) overlaps almost perfectly with the N reference (blue), confirming that the calibration parameters estimated from genuine pairs generalise correctly to all pairs in the verification database. Crucially, this improvement is achieved using only the predicted lens types from the two-stage classifierâno ground-truth label information is available at inference time. Verification EER improves across all datasets. Table 3 and Fig. 4 report EER before and after calibration. The improvement is consistent across all four datasets, ranging from a modest 4.1% relative reduction on AD100 (0.49% â 0.47%) to a substantial 28.30% on IITD-Vista (2.26% â 1.62%). The magnitude of the EER improvement correlates with the severity of the initial score degradation: IITD-Vista has the largest gap (39 points) and the largest CN displacement, and accordingly shows the greatest benefit from calibration. AD100, with a smaller Î (36 points) and a small number of non-patterned pairs, shows the most modest improvement. The ROC curves in Fig. 4 confirm that the improvement is not confined to the EER operating point: the calibrated curve (solid) lies above the raw curve (dashed) across the full range of False Match Rates, from 10â410^-4 to 10010^0. This means that regardless of the security threshold chosen by the operator, calibrated scores yield a higher TMR than raw scores. 4.3 Backbone Comparison Table 4: Stage-2 backbone comparison: full pipeline accuracy (%) with standard deviations from 10,000 resamples. Stage 1 is fixed. Dataset EfficientNet-V2 + MGSA ResNet-101 + MGSA ConvNeXt-B + MGSA UND-AD100 58.00±6.9958.00± 6.99 90.00±4.2690.00± 4.26 90.00±4.2690.00± 4.26 UND-LG4000 72.00±2.8372.00± 2.83 95.60±1.2995.60± 1.29 95.20±1.3895.20± 1.38 IITD-Cogent 68.20±1.6768.20± 1.67 97.24±0.6097.24± 0.60 97.50±0.5697.50± 0.56 IITD-Vista 65.41±1.9665.41± 1.96 98.29±0.5498.29± 0.54 98.80±0.4598.80± 0.45 Table 4 compares three Stage-2 backbone architectures. ResNet-101, EfficientNet-V2-S, and ConvNeXt-Base all incorporate MGSA spatial attention. MGSA substantially outperforms the mask-free baseline. We conducted experiments on EfficientNet-V2-S without mask guidance, which degraded by up to 30% on AD100 (60.00% vs. 90.00%) and by 10 points on IITD-Cogent (87.25% vs. 97.50%), confirming that the anatomical mask is the primary driver of Stage-2 performance, not backbone capacity alone. ConvNeXt-Base is the best MGSA backbone. Under identical mask guidance, ConvNeXt-Base matches or exceeds ResNet-101 on three of four datasets; ResNet-101 holds a marginal edge on LG4000 (+0.40 percentage points (p)). On the larger IITD datasets, ConvNeXt-Base leads by +0.26 p on Cogent and +0.51 p on Vista, indicating that its 7Ă77Ă7 depthwise kernels better capture the subtle limbal texture evidence of clear lens wear. The 7Ă77Ă7 depthwise kernels of ConvNeXt capture the spatially extended limbal arc more effectively than ResNetâs 3Ă33Ă3 convolutions, and this advantage is most pronounced where there is sufficient data to exploit the larger receptive field. 4.4 Cross-Dataset Generalization Table 5: Cross-dataset pipeline accuracy (%). Rows = train dataset; columns = test dataset. â Cross-sensor only (shared subjects). Train â Test â UND-AD100 UND-LG4000 IITD-Cogent IITD-Vista UND-AD100 â 87.60 69.12 36.99 UND-LG4000 68.00 â 61.63 43.66 IITD-Cogent 78.00 97.20 â â 97.26 IITD-Vista 84.00 73.20 â 85.41 â Table 5 reports pipeline accuracy for all 12 sourceâ combinations using the same checkpoints as Table 1, with no fine-tuning on the target dataset. The results show that the proposed method generalizes well within the same sensor family and moderately across related sensors, with the IITD-Cogent model achieving 97.20% on UND-LG4000 and 97.26% on IITD-Vista, suggesting that the anatomical mask provides a robust, sensor-invariant spatial prior. Even under cross-family transfer, IITD-trained models remain well above chance, indicating that MGSA captures meaningful limbal cues rather than purely sensor-specific patterns. In contrast, models trained on the smaller UND datasets generalize poorly to IITD sensors, highlighting an asymmetric transfer gap likely caused by limited training diversity and substantial differences in image resolution and depth of field, and motivating future work on domain adaptation and stronger data augmentation. 5 Conclusion We presented a two-stage method for iris contact lens detection that exploits the visual asymmetry between patterned and clear lenses. Stage 1 addresses the easier problem of patterned lens detection from cropped iris images with high accuracy across all four datasets, while Stage 2 addresses the harder clear vs. normal classification problem using a ConvNeXt-Base model with the proposed MGSA, which combines an anatomical ROI mask, spatial gating, and channel recalibration to focus on the most informative iris and periocular regions. Across UND-AD100, UND-LG4000, IITD-Cogent, and IITD-Vista, the full pipeline achieves 90.0%â98.8% accuracy and bootstrap SD scores, with ConvNeXt-Base consistently outperforming EfficientNet and ResNet baselines. We further showed that clear lenses systematically degrade commercial iris matcher scores, pushing genuine scores closer to the impostor distribution, and that condition-aware z-score calibration based on the predicted lens labels can recover much of this loss, reducing EER by 4.1%â28.3% across all datasets. Although the current system relies on Hough-based anatomical masks and has been evaluated only on near-infrared data, the results demonstrate that lens-aware detection and calibration provide a practical path toward more robust real-world iris recognition. Acknowledgments: This work was supported by NSF CITeR under Awards #1841517 and #2413309. The code supporting the findings of this study is publicly available on GitHub [https://github.com/iPRoBe-lab/MGSA_ContactLens_Detection]. References [1] A. Agarwal, N. Ratha, A. Noore, R. Singh, and M. Vatsa (2023) Misclassifications of contact lens iris PAD algorithms: is it gender bias or environmental conditions?. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, p. 961â970. Cited by: §2. [2] S. Baker, K. W. Bowyer, and P. J. Flynn (2010) Degradation of iris recognition performance due to non-cosmetic prescription contact lenses. Computer Vision and Image Understanding 114 (9), p. 1030â1044. Cited by: §1, §2. [3] A. Boyd, Z. Fang, A. Czajka, and K. W. Bowyer (2020) Iris presentation attack detection: where are we now?. Pattern Recognition Letters 138, p. 483â489. Cited by: §1. [4] A. Boyd, J. Speth, L. Parzianello, K. W. Bowyer, and A. Czajka (2023) Comprehensive study in open-set iris presentation attack detection. IEEE Transactions on Information Forensics and Security 18, p. 3238â3250. Cited by: §1. [5] C. Chen and A. Ross (2021) An explainable attention-guided iris presentation attack detector. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, p. 97â106. Cited by: §2. [6] A. Czajka, Z. Fang, and K. Bowyer (2019) Iris presentation attack detection based on photometric stereo features. In IEEE Winter Conference on Applications of Computer Vision (WACV), p. 877â885. Cited by: §2. [7] P. Das, J. McGrath, Z. Fang, L. Parzianello, and A. Czajka (2020) Iris liveness detection competition (LivDet-Iris) â the 2020 edition. In Proceedings of the IEEE International Joint Conference on Biometrics (IJCB), p. 1â10. External Links: Document Cited by: §2. [8] J. Daugman and C. Downing (2001) Epigenetic randomness, complexity and singularity of human iris patterns. Proceedings of the Royal Society of London. Series B: Biological Sciences 268 (1477), p. 1737â1740. Cited by: §1. [9] J. Daugman (2004) How iris recognition works. IEEE Transactions on Circuits and Systems for Video Technology 14 (1), p. 21â30. External Links: Document Cited by: §1, §1. [10] J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei (2009) ImageNet: a large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, p. 248â255. Cited by: §3.4.1. [11] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby (2021) An image is worth 16x16 words: transformers for image recognition at scale. International Conference on Learning Representations (ICLR). Cited by: §2. [12] J. S. Doyle, K. W. Bowyer, and P. J. Flynn (2013) Variation in accuracy of textured contact lens detection based on sensor and lens pattern. In IEEE Sixth International Conference on Biometrics: Theory, applications and Systems (BTAS), p. 1â7. Cited by: §1, §2. [13] J. Doyle and K. W. Bowyer (2014) Notre dame image database for contact lens detection in iris recognitionâ2013. Technical report University of Notre Dame, Computer Vision Research Laboratory. Note: Dataset documentation/README Cited by: Table 1, Table 1. [14] G. Erdogan and A. Ross (2013) Automatic detection of non-cosmetic soft contact lenses in ocular images. In Biometric and Surveillance Technology for Human and Activity Identification X, Vol. 8712, p. 62â76. Cited by: footnote 1. [15] M. Gupta, V. Singh, A. Agarwal, M. Vatsa, and R. Singh (2021) Generalized iris presentation attack detection algorithm under cross-database settings. In 25th International Conference on Pattern Recognition (ICPR), p. 5318â5325. Cited by: §1. [16] K. He, X. Zhang, S. Ren, and J. Sun (2016) Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), p. 770â778. External Links: Document Cited by: §1, §2, §3.4.1. [17] S. Hoffman, R. Sharma, and A. Ross (2018) Convolutional neural networks for iris presentation attack detection: toward cross-dataset and cross-sensor generalization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Cited by: §1. [18] J. Hu, L. Shen, and G. Sun (2018) Squeeze-and-excitation networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 7132â7141. External Links: Document Cited by: §1, §2, §3.4.2. [19] G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger (2017) Densely connected convolutional networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 4700â4708. External Links: Document Cited by: §2, §3.3. [20] B. Jafrasteh, S. P. LubiĂĄn-LĂłpez, E. Trimarco, M. R. Ruiz, C. R. Barrios, Y. M. Almagro, and I. Benavente-FernĂĄndez (2024) MGA-net: a novel mask-guided attention neural network for precision neonatal brain imaging. NeuroImage 300, p. 120872. Cited by: §2. [21] N. Kohli, D. Yadav, M. Vatsa, and R. Singh (2013) Revisiting iris recognition with color cosmetic contact lenses. In Proceedings of the International Conference on Biometrics, p. 1â7. External Links: Document Cited by: Table 1, Table 1. [22] Z. Liu, H. Mao, C. Wu, C. Feichtenhofer, T. Darrell, and S. Xie (2022) A ConvNet for the 2020s. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 11976â11986. External Links: Document Cited by: §1, §2, §3.4.1. [23] I. Loshchilov and F. Hutter (2019) Decoupled weight decay regularization. In International Conference on Learning Representations, Cited by: §3.5. [24] D. Menotti, G. Chiachia, A. Pinto, W. R. Schwartz, H. Pedrini, A. X. FalcĂŁo, and A. Rocha (2015) Deep representations for iris, face, and fingerprint spoofing detection. IEEE Transactions on Information Forensics and Security 10 (4), p. 864â879. External Links: Document Cited by: §1, §2. [25] M. Mitcheff, A. Hossain, S. Webster, S. Khan, K. Roszczewska, J. E. Tapia, F. Stockhardt, L. J. GonzĂĄlez-Soler, J. Lim, M. Pollok, et al. (2025) Iris liveness detection competition (livdet-iris)âthe 2025 edition. In IEEE International Joint Conference on Biometrics (IJCB), p. 1â10. Cited by: §2. [26] Neurotechnology (2026) VeriEye SDK. Note: https://w.neurotechnology.com/verieye.htmlIris identification technology and software development kit Cited by: §1. [27] K. Nguyen, H. Proença, and F. Alonso-Fernandez (2024) Deep learning for iris recognition: a survey. ACM Computing Surveys 56 (9), p. 1â35. Cited by: §1. [28] G. W. Quinn (2024) An open source iris segmentation algorithm for non-ideal images. NIST. Cited by: §3.2. [29] R. Raghavendra, K. B. Raja, and C. Busch Contlensnet: robust iris contact lens detection using deep convolutional neural networks. In 2017 IEEE winter conference on applications of computer vision (WACV), p. 1160â1167. Cited by: §1, §2, Table 1, Table 1. [30] A. Ross, S. Banerjee, C. Chen, A. Chowdhury, V. Mirjalili, R. Sharma, T. Swearingen, and S. Yadav (2019) Some research problems in biometrics: the future beckons. In International Conference on Biometrics (ICB), p. 1â8. Cited by: §2. [31] S&S Insider (2026) Contact Lens Market Size, Share & Growth Report 2035. Note: https://w.snsinsider.com/reports/contact-lens-market-8566Last updated: March 2, 2026. Accessed: April 30, 2026 Cited by: §1. [32] S. Shah and A. Ross (2009) Iris segmentation using geodesic active contours. IEEE Transactions on Information Forensics and Security 4 (4), p. 824â836. Cited by: §3.2. [33] R. Sharma and A. Ross (2020) D-NetPAD: an explainable and interpretable iris presentation attack detector. In IEEE International Joint Conference on Biometrics (IJCB), p. 1â10. Cited by: §1, §2, §3.3. [34] M. Tan and Q. V. Le (2019) EfficientNet: rethinking model scaling for convolutional neural networks. In Proceedings of the International Conference on Machine Learning (ICML), p. 6105â6114. Cited by: §1, §2, §3.4.1. [35] P. Tinsley, S. Purnapatra, M. Mitcheff, A. Boyd, C. Crum, K. Bowyer, P. Flynn, S. Schuckers, A. Czajka, M. Fang, et al. (2023) Iris liveness detection competition (livdet-iris)âthe 2023 edition. In IEEE International Joint Conference on Biometrics (IJCB), p. 1â10. Cited by: §2. [36] X. Wang, R. Girshick, A. Gupta, and K. He (2018) Non-local neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, p. 7794â7803. Cited by: §2. [37] S. Woo, J. Park, J. Lee, and I. S. Kweon (2018) CBAM: convolutional block attention module. In Proceedings of the European Conference on Computer Vision (ECCV), p. 3â19. External Links: Document Cited by: §2. [38] D. Yadav, N. Kohli, J. S. Doyle, R. Singh, M. Vatsa, and K. W. Bowyer (2014) Unraveling the effect of textured contact lenses on iris recognition. IEEE Transactions on Information Forensics and Security 9 (5), p. 851â862. External Links: Document Cited by: §1, Table 1, Table 1. [39] D. Yadav, N. Kohli, J. S. Doyle, R. Singh, M. Vatsa, and K. W. Bowyer (2014) Unraveling the effect of textured contact lenses on iris recognition. IEEE Transactions on Information Forensics and Security 9 (5), p. 851â862. External Links: Document Cited by: §1, §2. [40] S. Yadav and A. Ross (2021) CIT-GAN: cyclic image translation generative adversarial network with application in iris presentation attack detection. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, p. 2412â2421. Cited by: §1. [41] S. Yadav and A. Ross (2025) A multi-domain image translative diffusion stylegan for iris presentation attack detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshop, p. 3688â3697. Cited by: §1. [42] D. Yambay, B. Becker, N. Kohli, D. Yadav, A. Czajka, K. W. Bowyer, S. Schuckers, R. Singh, M. Vatsa, A. Noore, D. Gragnaniello, C. Sansone, and L. Verdoliva (2017) LivDet iris 2017 â iris liveness detection competition 2017. In Proceedings of the IEEE International Joint Conference on Biometrics (IJCB), p. 733â741. External Links: Document Cited by: §2. [43] D. Yambay, P. Das, A. Boyd, J. McGrath, Z. Fang, A. Czajka, S. Schuckers, K. Bowyer, M. Vatsa, R. Singh, et al. (2023) Review of iris presentation attack detection competitions. In Handbook of Biometric Anti-Spoofing: Presentation Attack Detection and Vulnerability Assessment, p. 149â169. Cited by: §1.