FaceShield: Explainable Face Anti-Spoofing with Multimodal Large Language Models
Hongyang Wang, Yichen Shi, Zhuofu Tao, Yuhao Gao, Liepiao Zhang, Xun Lin, Jun Feng, Xiaochen Yuan
Abstract
Face anti-spoofing (FAS) is crucial for protecting facial recognition systems from presentation attacks. Previous methods approached this task as a classification problem, lacking interpretability and reasoning behind the predicted results. Recently, multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and decision-making in visual tasks. However, there is currently no universal and comprehensive MLLM and dataset specifically designed for FAS task. To address this gap, we propose FaceShield, a MLLM for FAS, along with the corresponding pre-training and supervised fine-tuning (SFT) datasets, FaceShield-pre10K and FaceShield-sft45K. FaceShield is capable of determining the authenticity of faces, identifying types of spoofing attacks, providing reasoning for its judgments, and detecting attack areas. Specifically, we employ spoof-aware vision perception (SAVP) that incorporates both the original image and auxiliary information based on prior knowledge. We then use an prompt-guided vision token masking (PVTM) strategy to random mask vision tokens, thereby improving the model
BibTeX
@inproceedings{aaai2026_faceshieldexplai,
title = {FaceShield: Explainable Face Anti-Spoofing with Multimodal Large Language Models},
author = {Hongyang Wang and Yichen Shi and Zhuofu Tao and Yuhao Gao and Liepiao Zhang and Xun Lin and Jun Feng and Xiaochen Yuan and Zitong Yu and Xiaochun Cao},
booktitle = {AAAI 2026},
year = {2026}
}