paper-with-me

홈 › Papers

Attention Control with Metric Learning Alignment for Image Set-based Recognition

2019-08-05 · Xiaofeng Liu, Zhenhua Guo, Jane You, B. V. K. Vijaya Kumar

This paper considers the problem of image set-based face verification and identification. Unlike traditional single sample (an image or a video) setting, this situation assumes the availability of a set of heterogeneous collection of orderless images and videos. The samples can be taken at different check points, different identity documents $etc$. The importance of each image is usually considered either equal or based on a quality assessment of that image independent of other images and/or videos in that image set. How to model the relationship of orderless images within a set remains a challenge. We address this problem by formulating it as a Markov Decision Process (MDP) in a latent space. Specifically, we first propose a dependency-aware attention control (DAC) network, which uses actor-critic reinforcement learning for attention decision of each image to exploit the correlations among the unordered images. An off-policy experience replay is introduced to speed up the learning process. Moreover, the DAC is combined with a temporal model for videos using divide and conquer strategies. We also introduce a pose-guided representation (PGR) scheme that can further boost the performance at extreme poses. We propose a parameter-free PGR without the need for training as well as a novel metric learning-based PGR for pose alignment without the need for pose detection in testing stage. Extensive evaluations on IJB-A/B/C, YTF, Celebrity-1000 datasets demonstrate that our method outperforms many state-of-art approaches on the set-based as well as video-based face recognition databases.

📄 PDF Abstract BibTeX arXiv:1908.01872

Code (0)

등록된 구현이 없습니다.

Tasks

Face RecognitionFace VerificationMetric LearningReinforcement Learning

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Experience Replay Experience Replay is a replay memory technique used in reinforcement learning where we store the agent’s experiences at each time-step, $e\_{t} = \left(s\_{t}, a\_{t}, r\_{t},…

Similar Papers 제목 키워드 기반

Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models

2021-07-05 · Khazar Khorrami, Okko Räsänen

Systems that can find correspondences between multiple modalities, such as between speech and images, have great potential to solve different recognition and data analysis tasks in an unsupervised manner. This work studi…

Cross-Modal RetrievalObject LocalizationRetrievalSemantic Retrieval

StableEmit: Selection Probability Discount for Reducing Emission Latency of Streaming Monotonic Attention ASR

2021-07-01 · Hirofumi Inaguma, Tatsuya Kawahara

While attention-based encoder-decoder (AED) models have been successfully extended to the online variants for streaming automatic speech recognition (ASR), such as monotonic chunkwise attention (MoChA), the models still …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Boundary DetectionDecoder+2

Decoupled Attention Network for Text Recognition

2019-12-21 · Tianwei Wang, Yuanzhi Zhu, Lianwen Jin, Canjie Luo 외

Text recognition has attracted considerable research interests because of its various applications. The cutting-edge text recognition methods are based on attention mechanisms. However, most of attention methods usually …

DecoderHandwritten Text RecognitionScene Text Recognition

Accurate Recognition of Pneumonia and COVID-19 by Geometric Shape Normalization of Lung Region using Automatic Landmark Detection and Piecewise Affine Warping

2026-06-29 · Salvador E. Ayala-Raggi, Rafael Alejandro Cruz-Ovando, Lauro Reyes-Cocoletzi, Aldrin Barreto-Flores arxiv

This paper presents an automatic system for recognizing pulmonary diseases in chest X-rays using geometric normalization of the lung region. The method combines three modules: (1) a ResNet-18 landmark detector with coord…

Transfer Learning

Adaptive Embedding Gate for Attention-Based Scene Text Recognition

2019-08-26 · Xiaoxue Chen, Tianwei Wang, Yuanzhi Zhu, Lianwen Jin 외

Scene text recognition has attracted particular research interest because it is a very challenging problem and has various applications. The most cutting-edge methods are attentional encoder-decoder frameworks that learn…

DecoderScene Text Recognition