paper-with-me

홈 › Papers

Text-Guided Face Recognition using Multi-Granularity Cross-Modal Contrastive Learning

2023-12-14 · Md Mahedi Hasan, Shoaib Meraj Sami, Nasser Nasrabadi

State-of-the-art face recognition (FR) models often experience a significant performance drop when dealing with facial images in surveillance scenarios where images are in low quality and often corrupted with noise. Leveraging facial characteristics, such as freckles, scars, gender, and ethnicity, becomes highly beneficial in improving FR performance in such scenarios. In this paper, we introduce text-guided face recognition (TGFR) to analyze the impact of integrating facial attributes in the form of natural language descriptions. We hypothesize that adding semantic information into the loop can significantly improve the image understanding capability of an FR algorithm compared to other soft biometrics. However, learning a discriminative joint embedding within the multimodal space poses a considerable challenge due to the semantic gap in the unaligned image-text representations, along with the complexities arising from ambiguous and incoherent textual descriptions of the face. To address these challenges, we introduce a face-caption alignment module (FCAM), which incorporates cross-modal contrastive losses across multiple granularities to maximize the mutual information between local and global features of the face-caption pair. Within FCAM, we refine both facial and textual features for learning aligned and discriminative features. We also design a face-caption fusion module (FCFM) that applies fine-grained interactions and coarse-grained associations among cross-modal features. Through extensive experiments conducted on three face-caption datasets, proposed TGFR demonstrates remarkable improvements, particularly on low-quality images, over existing FR models and outperforms other related methods and benchmarks.

📄 PDF Abstract BibTeX arXiv:2312.09367

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningFace Recognition

Similar Papers 제목 키워드 기반

Multi-Granularity Reasoning for Social Relation Recognition from Images

2019-01-10 · Meng Zhang, Xinchen Liu, Wu Liu, Anfu Zhou 외

Discovering social relations in images can make machines better interpret the behavior of human beings. However, automatically recognizing social relations in images is a challenging task due to the significant gap betwe…

RelationVisual Social Relationship Recognition

MGHFT: Multi-Granularity Hierarchical Fusion Transformer for Cross-Modal Sticker Emotion Recognition

2025-07-25 · Jian Chen, Yuxuan Hu, Haifeng Lu, Wei Wang 외 arxiv

Although pre-trained visual models with text have demonstrated strong capabilities in visual feature extraction, sticker emotion understanding remains challenging due to its reliance on multi-view information, such as ba…

Contrastive LearningEmotion Recognition

A Supervised Information Enhanced Multi-Granularity Contrastive Learning Framework for EEG Based Emotion Recognition

2024-05-12 · Xiang Li, Jian Song, Zhigang Zhao, Chunxiao Wang 외

This study introduces a novel Supervised Info-enhanced Contrastive Learning framework for EEG based Emotion Recognition (SICLEER). SI-CLEER employs multi-granularity contrastive learning to create robust EEG contextual r…

Contrastive LearningEEGEmotion Recognition

Multi-Granularity Guided Fusion-in-Decoder

2024-04-03 · Eunseong Choi, Hyeri Lee, Jongwuk Lee

In Open-domain Question Answering (ODQA), it is essential to discern relevant contexts as evidence and avoid spurious ones among retrieved results. The model architecture that uses concatenated multiple contexts in the d…

DecoderMulti-Task LearningNatural QuestionsOpen-Domain Question Answering+6

Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity

2023-06-28 · Zhenlin Xu, Yi Zhu, Tiffany Deng, Abhay Mittal 외

This paper presents novel benchmarks for evaluating vision-language models (VLMs) in zero-shot recognition, focusing on granularity and specificity. Although VLMs excel in tasks like image captioning, they face challenge…

BenchmarkingImage CaptioningSpecificityZero-Shot Learning