paper-with-me

홈 › Papers

Revising Image-Text Retrieval via Multi-Modal Entailment

2022-08-22 · Xu Yan, Chunhui Ai, Ziqiang Cao, Min Cao, Sujian Li, Wenjie Li, Guohong Fu

An outstanding image-text retrieval model depends on high-quality labeled data. While the builders of existing image-text retrieval datasets strive to ensure that the caption matches the linked image, they cannot prevent a caption from fitting other images. We observe that such a many-to-many matching phenomenon is quite common in the widely-used retrieval datasets, where one caption can describe up to 178 images. These large matching-lost data not only confuse the model in training but also weaken the evaluation accuracy. Inspired by visual and textual entailment tasks, we propose a multi-modal entailment classifier to determine whether a sentence is entailed by an image plus its linked captions. Subsequently, we revise the image-text retrieval datasets by adding these entailed captions as additional weak labels of an image and develop a universal variable learning rate strategy to teach a retrieval model to distinguish the entailed captions from other negative samples. In experiments, we manually annotate an entailment-corrected image-text retrieval dataset for evaluation. The results demonstrate that the proposed entailment classifier achieves about 78% accuracy and consistently improves the performance of image-text retrieval baselines.

📄 PDF Abstract BibTeX arXiv:2208.10126

Code (0)

등록된 구현이 없습니다.

Tasks

Image-text RetrievalNatural Language InferenceRetrievalSentenceText Retrieval

Similar Papers 제목 키워드 기반

Are All Combinations Equal? Combining Textual and Visual Features with Multiple Space Learning for Text-Based Video Retrieval

2022-11-21 · Damianos Galanopoulos, Vasileios Mezaris

In this paper we tackle the cross-modal video retrieval problem and, more specifically, we focus on text-to-video retrieval. We investigate how to optimally combine multiple diverse textual and visual features into featu…

AllRetrievalText to Video RetrievalVideo Retrieval

Mr. Right: Multimodal Retrieval on Representation of ImaGe witH Text

2022-09-28 · Cheng-An Hsieh, Cheng-Ping Hsieh, Pu-Jen Cheng

Multimodal learning is a recent challenge that extends unimodal learning by generalizing its domain to diverse modalities, such as texts, images, or speech. This extension requires models to process and relate informatio…

Image CaptioningImage RetrievalImage-text RetrievalInformation Retrieval+2

Learning Cross-Modal Deep Embeddings for Multi-Object Image Retrieval using Text and Sketch

2018-04-28 · Sounak Dey, Anjan Dutta, Suman K. Ghosh, Ernest Valveny 외

In this work we introduce a cross modal image retrieval system that allows both text and sketch as input modalities for the query. A cross-modal deep network architecture is formulated to jointly model the sketch and tex…

Image RetrievalRetrieval

Universal Vision-Language Dense Retrieval: Learning A Unified Representation Space for Multi-Modal Retrieval

2022-09-01 · Zhenghao Liu, Chenyan Xiong, Yuanhuiyi Lv, Zhiyuan Liu 외

This paper presents Universal Vision-Language Dense Retrieval (UniVL-DR), which builds a unified model for multi-modal retrieval. UniVL-DR encodes queries and multi-modality resources in an embedding space for searching …

Image RetrievalOpen-Domain Question AnsweringQuestion AnsweringRetrieval+1

Revisiting Cross Modal Retrieval

2018-07-19 · Shah Nawaz, Muhammad Kamran Janjua, Alessandro Calefati, Ignazio Gallo

This paper proposes a cross-modal retrieval system that leverages on image and text encoding. Most multimodal architectures employ separate networks for each modality to capture the semantic relationship between them. Ho…

Cross-Modal RetrievalRetrieval