paper-with-me

Papers

Flickr30K-CFQ: A Compact and Fragmented Query Dataset for Text-image Retrieval

2024-03-20 · Haoyu Liu, Yaoxian Song, Xuwu Wang, Zhu Xiangru, Zhixu Li, Wei Song, Tiefeng Li

With the explosive growth of multi-modal information on the Internet, unimodal search cannot satisfy the requirement of Internet applications. Text-image retrieval research is needed to realize high-quality and efficient retrieval between different modalities. Existing text-image retrieval research is mostly based on general vision-language datasets (e.g. MS-COCO, Flickr30K), in which the query utterance is rigid and unnatural (i.e. verbosity and formality). To overcome the shortcoming, we construct a new Compact and Fragmented Query challenge dataset (named Flickr30K-CFQ) to model text-image retrieval task considering multiple query content and style, including compact and fine-grained entity-relation corpus. We propose a novel query-enhanced text-image retrieval method using prompt engineering based on LLM. Experiments show that our proposed Flickr30-CFQ reveals the insufficiency of existing vision-language datasets in realistic text-image tasks. Our LLM-based Query-enhanced method applied on different existing text-image retrieval models improves query understanding performance both on public dataset and our challenge set Flickr30-CFQ with over 0.9% and 2.4% respectively. Our project can be available anonymously in https://sites.google.com/view/Flickr30K-cfq.

📄 PDF Abstract BibTeX arXiv:2403.13317

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalPrompt EngineeringRetrieval

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Compact and Effective Representations for Sketch-based Image Retrieval

2021-04-20 · Pablo Torres, Jose M. Saavedra

Sketch-based image retrieval (SBIR) has undergone an increasing interest in the community of computer vision bringing high impact in real applications. For instance, SBIR brings an increased benefit to eCommerce search e…

Dimensionality ReductionImage RetrievalRetrievalSketch-Based Image Retrieval

Multimedia Semantic Integrity Assessment Using Joint Embedding Of Images And Text

2017-07-06 · Ayush Jaiswal, Ekraam Sabir, Wael Abd-Almageed, Premkumar Natarajan

Real world multimedia data is often composed of multiple modalities such as an image or a video with associated text (e.g. captions, user comments, etc.) and metadata. Such multimodal data packages are prone to manipulat…

Representation Learning

Assessing Brittleness of Image-Text Retrieval Benchmarks from Vision-Language Models Perspective

2024-07-21 · Mariya Hendriksen, Shuo Zhang, Ridho Reinanda, Mohamed Yahya 외

We examine the brittleness of the image-text retrieval (ITR) evaluation pipeline with a focus on concept granularity. We start by analyzing two common benchmarks, MS-COCO and Flickr30k, and compare them with augmented, f…

Image-text RetrievalInformation RetrievalRetrievalText Retrieval

Stacked Cross Attention for Image-Text Matching

2018-03-21 · ECCV 2018 9 · Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu 외

In this paper, we study the problem of image-text matching. Inferring the latent semantic alignment between objects or other salient stuff (e.g. snow, sky, lawn) and the corresponding words in sentences allows to capture…

Cross-Modal RetrievalImage RetrievalImage-text matchingRetrieval+5

Query-guided Regression Network with Context Policy for Phrase Grounding

2017-08-04 · ICCV 2017 10 · Kan Chen, Rama Kovvuri, Ram Nevatia

Given a textual description of an image, phrase grounding localizes objects in the image referred by query phrases in the description. State-of-the-art methods address the problem by ranking a set of proposals based on t…

Phrase GroundingregressionReinforcement Learning