Abstract:
Owing to its uniqueness, stability, and high accuracy, iris recognition has been widely applied in fields such as financial payments and public security. However, driven by continuous technological advancements, attackers increasingly forge iris features to bypass identity-verification systems. Iris presentation attack detection (IPAD) aims to distinguish bona fide iris samples from various attack iris images. Existing methods for IPAD fail to generalize or remain robust under complex environmental conditions and diverse attack techniques, thus rendering them unsuitable for practical applications. To address the limitations of existing IPAD methods, which rely solely on visual feature modeling and exhibit insufficient generalization in scenarios involving both physical and digital attacks, this paper proposes an IPAD method that integrates a text-driven attention mechanism with a vision foundation model, named text-guided attention for IPAD. This method fuses information from both visual and textual modalities by leveraging the rich semantic knowledge of pretrained language models to guide the separation of bona fide and attack iris samples in the representation space. The learning process of the model is explicitly constrained and guided by utilizing textual semantics as a form of weak supervision, thus resulting in more robust and discriminative iris feature representations. This architecture is designed to overcome the challenges of adapting generic vision-language models to the fine-grained domain of IPAD. To enhance the capability of the method to perceive iris features, a fine-grained visual-text alignment framework based on textual prompts is introduced. This enables precise matching between iris image features and semantic information, thereby directing the model’s attention to subtle, attack-relevant regions, such as texture irregularities and printing artifacts. Moreover, multiple learnable textual prompts are employed to avoid bias from a single description, and their mean prototype features serve as semantic anchors for alignment. Furthermore, the method employs text-conditioned constraints based on slot attention to decouple bona fide and spoof features in the latent space, thereby improving generalization across diverse attack types. Experimental results on the LivDet-Iris 2017 and LivDet-Iris 2023 competition datasets reveal the superior performance of the proposed method in evaluating both cross-dataset textured contact lens and physical–digital joint attacks, thus confirming its strong generalizability in open-world scenarios. Specifically, compared with the respective best-performing baseline methods in the two evaluation settings, the proposed method reduces the average classification error rate (ACER) by 1.27 percentage points in the cross-dataset textured contact lens attack evaluation and lowers the overall average classification error rates (ACER1 and ACER2) by 12.03 and 6.46 percentage points, respectively, in the challenging physical–digital joint attack setting. Beyond quantitative evaluations, qualitative experiments utilizing textual prompts demonstrate that the model effectively captures discriminative features specific to bona fide and attack irises, confirming that the integration of textual guidance offers a more generalizable understanding of IPAD. This study offers a high-performance solution for IPAD and pioneers a promising technical direction by effectively leveraging vision-language foundational models to significantly enhance the security and generalizability of biometric systems.