paper-with-me

홈 › Papers

FM-ViT: Flexible Modal Vision Transformers for Face Anti-Spoofing

2023-05-05 · Ajian Liu, Zichang Tan, Zitong Yu, Chenxu Zhao, Jun Wan, Yanyan Liang, Zhen Lei, Du Zhang, Stan Z. Li, Guodong Guo

The availability of handy multi-modal (i.e., RGB-D) sensors has brought about a surge of face anti-spoofing research. However, the current multi-modal face presentation attack detection (PAD) has two defects: (1) The framework based on multi-modal fusion requires providing modalities consistent with the training input, which seriously limits the deployment scenario. (2) The performance of ConvNet-based model on high fidelity datasets is increasingly limited. In this work, we present a pure transformer-based framework, dubbed the Flexible Modal Vision Transformer (FM-ViT), for face anti-spoofing to flexibly target any single-modal (i.e., RGB) attack scenarios with the help of available multi-modal data. Specifically, FM-ViT retains a specific branch for each modality to capture different modal information and introduces the Cross-Modal Transformer Block (CMTB), which consists of two cascaded attentions named Multi-headed Mutual-Attention (MMA) and Fusion-Attention (MFA) to guide each modal branch to mine potential features from informative patch tokens, and to learn modality-agnostic liveness features by enriching the modal information of own CLS token, respectively. Experiments demonstrate that the single model trained based on FM-ViT can not only flexibly evaluate different modal samples, but also outperforms existing single-modal frameworks by a large margin, and approaches the multi-modal frameworks introduced with smaller FLOPs and model parameters.

📄 PDF Abstract BibTeX arXiv:2305.03277

Code (0)

등록된 구현이 없습니다.

Tasks

Face Anti-SpoofingFace Presentation Attack Detection

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Multi-Head Attention 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…

Similar Papers 제목 키워드 기반

Visual Prompt Flexible-Modal Face Anti-Spoofing

2023-07-26 · Zitong Yu, Rizhao Cai, Yawen Cui, Ajian Liu 외

Recently, vision transformer based multimodal learning methods have been proposed to improve the robustness of face anti-spoofing (FAS) systems. However, multimodal face data collected from the real world is often imperf…

Face Anti-SpoofingPrompt Learning

PaLI: A Jointly-Scaled Multilingual Language-Image Model

2022-09-14 · Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni 외

Effective scaling and a flexible task interface enable large language models to excel at many tasks. We present PaLI (Pathways Language and Image model), a model that extends this approach to the joint modeling of langua…

DecoderFew-Shot Image ClassificationImage CaptioningImage Classification+7

MA-ViT: Modality-Agnostic Vision Transformers for Face Anti-Spoofing

2023-04-15 · Ajian Liu, Yanyan Liang

The existing multi-modal face anti-spoofing (FAS) frameworks are designed based on two strategies: halfway and late fusion. However, the former requires test modalities consistent with the training input, which seriously…

Face Anti-Spoofing

Surface Vision Transformers: Flexible Attention-Based Modelling of Biomedical Surfaces

2022-04-07 · Simon Dahan, Hao Xu, Logan Z. J. Williams, Abdulah Fawaz 외

Recent state-of-the-art performances of Vision Transformers (ViT) in computer vision tasks demonstrate that a general-purpose architecture, which implements long-range self-attention, could replace the local feature lear…

ClassificationData Augmentation

Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation

2025-08-30 · Xuechao Zou, Shun Zhang, Xing Fu, Yue Li 외 arxiv

Controllable face generation poses critical challenges in generative modeling due to the intricate balance required between semantic controllability and photorealism. While existing approaches struggle with disentangling…

Zero-shot Generalization