paper-with-me

Papers

Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval

2025-04-10 · Zehong Ma, Hao Chen, Wei Zeng, Limin Su, Shiliang Zhang

Fine-grained text-to-image retrieval aims to retrieve a fine-grained target image with a given text query. Existing methods typically assume that each training image is accurately depicted by its textual descriptions. However, textual descriptions can be ambiguous and fail to depict discriminative visual details in images, leading to inaccurate representation learning. To alleviate the effects of text ambiguity, we propose a Multi-Modal Reference learning framework to learn robust representations. We first propose a multi-modal reference construction module to aggregate all visual and textual details of the same object into a comprehensive multi-modal reference. The multi-modal reference hence facilitates the subsequent representation learning and retrieval similarity computation. Specifically, a reference-guided representation learning module is proposed to use multi-modal references to learn more accurate visual and textual representations. Additionally, we introduce a reference-based refinement method that employs the object references to compute a reference-based similarity that refines the initial retrieval results. Extensive experiments are conducted on five fine-grained text-to-image retrieval datasets for different text-to-image retrieval tasks. The proposed method has achieved superior performance over state-of-the-art methods. For instance, on the text-to-person image retrieval dataset RSTPReid, our method achieves the Rank1 accuracy of 56.2\%, surpassing the recent CFine by 5.6\%.

📄 PDF Abstract BibTeX arXiv:2504.07718

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalRepresentation LearningRetrieval

Similar Papers 제목 키워드 기반

Reliability-Prioritized Fine-Grained Generation in Multimodal Large

2026-06-28 · Xiaomeng Fan, Wei Wu, Yuwei Wu, Zhi Gao 외 arxiv

Multimodal large language models (MLLMs) are increasingly expected to generate fine-grained descriptions of visual content. However, we observe and theoretically show that generating fine-grained responses poses a reliab…

Fine-grained and Explainable Factuality Evaluation for Multimodal Summarization

2024-02-18 · Yue Zhang, Jingxuan Zuo, Liqiang Jing

Multimodal summarization aims to generate a concise summary based on the input text and image. However, the existing methods potentially suffer from unfactual output. To evaluate the factuality of multimodal summarizatio…

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning

2025-05-25 · Yeyuan Wang, Dehong Gao, Rujiao Long, Lei Yi 외

Direct Preference Optimization (DPO) has gained significant attention for its simplicity and computational efficiency in aligning large language models (LLMs). Recent advancements have extended DPO to multimodal scenario…

Computational EfficiencyMultimodal ReasoningSentence

Semi-supervised multimodal coreference resolution in image narrations

2023-10-20 · Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen

In this paper, we study multimodal coreference resolution, specifically where a longer descriptive text, i.e., a narration is paired with an image. This poses significant challenges due to fine-grained image-text alignme…

coreference-resolutionCoreference ResolutionDescriptive

FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval

2025-03-27 · Zixu Li, Zhiheng Fu, Yupeng Hu, Zhiwei Chen 외

Composed Image Retrieval (CIR) facilitates image retrieval through a multimodal query consisting of a reference image and modification text. The reference image defines the retrieval context, while the modification text …

Image RetrievalRetrieval