Fast and Interpretable Face Identification for Out-Of-Distribution Data Using Vision Transformers
Most face identification approaches employ a Siamese neural network to compare two images at the image embedding level. Yet, this technique can be subject to occlusion (e.g. faces with masks or sunglasses) and out-of-distribution data. DeepFace-EMD (Phan et al. 2022) reaches state-of-the-art accuracy on out-of-distribution data by first comparing two images at the image level, and then at the patch level. Yet, its later patch-wise re-ranking stage admits a large $O(n^3 \log n)$ time complexity (for $n$ patches in an image) due to the optimal transport optimization. In this paper, we propose a novel, 2-image Vision Transformers (ViTs) that compares two images at the patch level using cross-attention. After training on 2M pairs of images on CASIA Webface (Yi et al. 2014), our model performs at a comparable accuracy as DeepFace-EMD on out-of-distribution data, yet at an inference speed more than twice as fast as DeepFace-EMD (Phan et al. 2022). In addition, via a human study, our model shows promising explainability through the visualization of cross-attention. We believe our work can inspire more explorations in using ViTs for face identification.
Code (1)
Tasks
Face IdentificationRe-RankingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Robust and Low-Rank Representation for Fast Face Identification with Occlusions
In this paper we propose an iterative method to address the face identification problem with block occlusions. Our approach utilizes a robust representation based on two characteristics in order to model contiguous error…
Face IdentificationFast refacing of MR images with a generative neural network lowers re-identification risk and preserves volumetric consistency
With the rise of open data, identifiability of individuals based on 3D renderings obtained from routine structural magnetic resonance imaging (MRI) scans of the head has become a growing privacy concern. To protect subje…
Brain MorphometryDe-identificationFace GenerationGenerative Adversarial NetworkA Note on Identification of Match Fixed Effects as Interpretable Unobserved Match Affinity
We highlight that match fixed effects, represented by the coefficients of interaction terms involving dummy variables for two elements, lack identification without specific restrictions on parameters. Consequently, the c…
IFAST: Weakly Supervised Interpretable Face Anti-spoofing from Single-shot Binocular NIR Images
Single-shot face anti-spoofing (FAS) is a key technique for securing face recognition systems, and it requires only static images as input. However, single-shot FAS remains a challenging and under-explored problem due to…
Depth EstimationDisparity EstimationFace Anti-SpoofingFace RecognitionA Fast and Accurate System for Face Detection, Identification, and Verification
The availability of large annotated datasets and affordable computation power have led to impressive improvements in the performance of CNNs on various object detection and recognition benchmarks. These, along with a bet…
Face DetectionFace IdentificationFace Recognitionobject-detection+2