paper-with-me

홈 › Papers

NAPReg: Nouns As Proxies Regularization for Semantically Aware Cross-Modal Embeddings

2023-01-07 · IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2023 1 · Bhavin Jawade, Deen Dayal Mohan, Naji Mohamed Ali, Srirangaraj Setlur, Venu Govindaraju

Cross-modal retrieval is a fundamental vision-language task with a broad range of practical applications. Text-to-image matching is the most common form of cross-modal retrieval where, given a large database of images and a textual query, the task is to retrieve the most relevant set of images. Existing methods utilize dual encoders with an attention mechanism and a ranking loss for learning embeddings that can be used for retrieval based on cosine similarity. Despite the fact that these methods attempt to perform semantic alignment across visual regions and textual words using tailored attention mechanisms, there is no explicit supervision from the training objective to enforce such alignment. To address this, we propose NAPReg, a novel regularization formulation that projects high-level semantic entities i.e Nouns into the embedding space as shared learnable proxies. We show that using such a formulation allows the attention mechanism to learn better word-region alignment while also utilizing region information from other samples to build a more generalized latent representation for semantic concepts. Experiments on three benchmark datasets i.e. MS-COCO, Flickr30k and Flickr8k demonstrate that our method achieves state-of-the-art results in cross-modal metric learning for text-image and image-text retrieval tasks. Code: https://github.com/bhavinjawade/NAPReq

📄 PDF Abstract BibTeX

Code (1)

bhavinjawade/NAPReq

Tasks

Cross-Modal RetrievalImage-text RetrievalMetric LearningRetrievalText Retrieval

Similar Papers 제목 키워드 기반

Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement

2025-02-24 · Suchae Jeong, Inseong Choi, Youngsik Yun, Jihie Kim

Text-to-Image models, including Stable Diffusion, have significantly improved in generating images that are highly semantically aligned with the given prompts. However, existing models may fail to produce appropriate ima…

Image GenerationText to Image GenerationText-to-Image Generation

Exophoric Pronoun Resolution in Dialogues with Topic Regularization

2021-09-10 · EMNLP 2021 11 · Xintong Yu, Hongming Zhang, Yangqiu Song, ChangShui Zhang 외

Resolving pronouns to their referents has long been studied as a fundamental natural language understanding problem. Previous works on pronoun coreference resolution (PCR) mostly focus on resolving pronouns to mentions i…

coreference-resolutionCoreference ResolutionNatural Language Understanding

Using Noun Similarity to Adapt an Acceptability Measure for Persian Light Verb Constructions

2012-05-01 · LREC 2012 5 · Shiva Taslimipoor, Afsaneh Fazly, Ali Hamzeh

Light verb constructions (LVCs), such as take a walk and make a decision, are a common subclass of multiword expressions (MWEs), whose distinct syntactic and semantic properties call for a special treatment within a comp…

Machine Translation

Anchor-aware Deep Metric Learning for Audio-visual Retrieval

2024-04-21 · Donghuo Zeng, Yanan Wang, Kazushi Ikeda, Yi Yu

Metric learning minimizes the gap between similar (positive) pairs of data points and increases the separation of dissimilar (negative) pairs, aiming at capturing the underlying data structure and enhancing the performan…

Cross-Modal RetrievalMetric LearningRetrieval

HIER: Metric Learning Beyond Class Labels via Hierarchical Regularization

2022-12-29 · CVPR 2023 1 · Sungyeon Kim, Boseung Jeong, Suha Kwak

Supervision for metric learning has long been given in the form of equivalence between human-labeled classes. Although this type of supervision has been a basis of metric learning for decades, we argue that it hinders fu…

Metric Learning