paper-with-me

Papers

Uncertainty-based Cross-Modal Retrieval with Probabilistic Representations

2022-04-20 · Leila Pishdad, Ran Zhang, Konstantinos G. Derpanis, Allan Jepson, Afsaneh Fazly

Probabilistic embeddings have proven useful for capturing polysemous word meanings, as well as ambiguity in image matching. In this paper, we study the advantages of probabilistic embeddings in a cross-modal setting (i.e., text and images), and propose a simple approach that replaces the standard vector point embeddings in extant image-text matching models with probabilistic distributions that are parametrically learned. Our guiding hypothesis is that the uncertainty encoded in the probabilistic embeddings captures the cross-modal ambiguity in the input instances, and that it is through capturing this uncertainty that the probabilistic models can perform better at downstream tasks, such as image-to-text or text-to-image retrieval. Through extensive experiments on standard and new benchmarks, we show a consistent advantage for probabilistic representations in cross-modal retrieval, and validate the ability of our embeddings to capture uncertainty.

📄 PDF Abstract BibTeX arXiv:2204.09268

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalImage RetrievalImage-text matchingImage to textRetrievalText Matching

Similar Papers 제목 키워드 기반

Probabilistic Multimodal Representation Learning

2021-01-01 · Leila Pishdad, Ran Zhang, Afsaneh Fazly, Allan Jepson

Learning multimodal representations is a requirement for many tasks such as image--caption retrieval. Previous work on this problem has only focused on finding good vector representations without any explicit measure of …

Representation LearningRetrieval

Probabilistic Embeddings for Frozen Vision-Language Models: Uncertainty Quantification with Gaussian Process Latent Variable Models

2025-05-08 · Aishwarya Venkataramanan, Paul Bodesheim, Joachim Denzler

Vision-Language Models (VLMs) learn joint representations by mapping images and text into a shared latent space. However, recent research highlights that deterministic embeddings from standard VLMs often struggle to capt…

Active Learningcross-modal alignmentCross-Modal RetrievalEmbeddings Evaluation+3

Probabilistic Embeddings for Cross-Modal Retrieval

2021-01-13 · CVPR 2021 1 · Sanghyuk Chun, Seong Joon Oh, Rafael Sampaio de Rezende, Yannis Kalantidis 외

Cross-modal retrieval methods build a common representation space for samples from multiple modalities, typically from the vision and the language domains. For images and their captions, the multiplicity of the correspon…

Cross-Modal RetrievalRetrieval

Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning

2025-09-01 · Long Zhang, Peipei Song, Jianfeng Dong, Kun Li 외 arxiv

Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos partially relevant to a given query. The core challenge lies in learning robust query-video alignment against spurious semantic correlations ari…

Partially Relevant Video RetrievalVideo Alignment

Probabilistic framework for solving Visual Dialog

2019-09-11 · Badri N. Patro, Anupriy, Vinay P. Namboodiri

In this paper, we propose a probabilistic framework for solving the task of `Visual Dialog'. Solving this task requires reasoning and understanding of visual modality, language modality, and common sense knowledge to ans…

Common Sense ReasoningVisual Dialog