基于视觉–语言基础模型的可泛化虹膜呈现攻击检测

Generalizable iris presentation attack detection based on vision-language model

  • 摘要: 虹膜呈现攻击检测是保障虹膜识别系统可靠性的核心技术之一. 针对现有方法依赖纯视觉特征建模,在复杂攻击场景下泛化能力不足的问题,本文提出一种结合文本驱动注意力机制与视觉基础模型的虹膜呈现攻击检测方法,旨在利用预训练语言模型的世界知识,显式引导模型学习过程,从而学习到更具判别性的虹膜特征表示. 为了增强方法对虹膜特征的感知能力,引入基于文本提示的细粒度视觉文本对齐框架,实现虹膜图像特征与语义信息之间的精细匹配. 同时,为了克服单一文本描述可能带来的偏见,采用可学习文本提示平均原型作为语义锚点. 在此基础上,进一步通过基于槽注意力的文本条件约束对真假特征进行解耦,以提升方法在多样化攻击类型下的泛化性能. 在虹膜呈现攻击检测比赛LivDet-Iris 2017和LivDet-Iris 2023数据集上的实验表明,所提方法在跨数据集纹理隐形眼镜攻击评估和物理–数字联合攻击评估中均取得了较优性能,验证了其在开放场景下的泛化能力. 尤其在物理与数字联合攻击场景中,所提方法相比现有方法,综合平均分类错误率ACER1和ACER2分别降低了12.03和6.46个百分点. 除定量评估外,本文还开展了基于文本提示的定性实验,结果表明该方法能够捕捉到真假虹膜的特定特征,相较于纯视觉方法具有更强的泛化能力.

     

    Abstract: Owing to its uniqueness, stability, and high accuracy, iris recognition has been widely applied in fields such as financial payments and public security. However, driven by continuous technological advancements, attackers increasingly forge iris features to bypass identity-verification systems. Iris presentation attack detection (IPAD) aims to distinguish bona fide iris samples from various attack iris images. Existing methods for IPAD fail to generalize or remain robust under complex environmental conditions and diverse attack techniques, thus rendering them unsuitable for practical applications. To address the limitations of existing IPAD methods, which rely solely on visual feature modeling and exhibit insufficient generalization in scenarios involving both physical and digital attacks, this paper proposes an IPAD method that integrates a text-driven attention mechanism with a vision foundation model, named text-guided attention for IPAD. This method fuses information from both visual and textual modalities by leveraging the rich semantic knowledge of pretrained language models to guide the separation of bona fide and attack iris samples in the representation space. The learning process of the model is explicitly constrained and guided by utilizing textual semantics as a form of weak supervision, thus resulting in more robust and discriminative iris feature representations. This architecture is designed to overcome the challenges of adapting generic vision-language models to the fine-grained domain of IPAD. To enhance the capability of the method to perceive iris features, a fine-grained visual-text alignment framework based on textual prompts is introduced. This enables precise matching between iris image features and semantic information, thereby directing the model’s attention to subtle, attack-relevant regions, such as texture irregularities and printing artifacts. Moreover, multiple learnable textual prompts are employed to avoid bias from a single description, and their mean prototype features serve as semantic anchors for alignment. Furthermore, the method employs text-conditioned constraints based on slot attention to decouple bona fide and spoof features in the latent space, thereby improving generalization across diverse attack types. Experimental results on the LivDet-Iris 2017 and LivDet-Iris 2023 competition datasets reveal the superior performance of the proposed method in evaluating both cross-dataset textured contact lens and physical–digital joint attacks, thus confirming its strong generalizability in open-world scenarios. Specifically, compared with the respective best-performing baseline methods in the two evaluation settings, the proposed method reduces the average classification error rate (ACER) by 1.27 percentage points in the cross-dataset textured contact lens attack evaluation and lowers the overall average classification error rates (ACER1 and ACER2) by 12.03 and 6.46 percentage points, respectively, in the challenging physical–digital joint attack setting. Beyond quantitative evaluations, qualitative experiments utilizing textual prompts demonstrate that the model effectively captures discriminative features specific to bona fide and attack irises, confirming that the integration of textual guidance offers a more generalizable understanding of IPAD. This study offers a high-performance solution for IPAD and pioneers a promising technical direction by effectively leveraging vision-language foundational models to significantly enhance the security and generalizability of biometric systems.

     

/

返回文章
返回