paper-with-me

홈 › Papers

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation

2024-12-14 · Yang Yang, Wenjuan Xi, Luping Zhou, Jinhui Tang

Vision-language retrieval aims to search for similar instances in one modality based on queries from another modality. The primary objective is to learn cross-modal matching representations in a latent common space. Actually, the assumption underlying cross-modal matching is modal balance, where each modality contains sufficient information to represent the others. However, noise interference and modality insufficiency often lead to modal imbalance, making it a common phenomenon in practice. The impact of imbalance on retrieval performance remains an open question. In this paper, we first demonstrate that ultimate cross-modal matching is generally sub-optimal for cross-modal retrieval when imbalanced modalities exist. The structure of instances in the common space is inherently influenced when facing imbalanced modalities, posing a challenge to cross-modal similarity measurement. To address this issue, we emphasize the importance of meaningful structure-preserved matching. Accordingly, we propose a simple yet effective method to rebalance cross-modal matching by learning structure-preserved matching representations. Specifically, we design a novel multi-granularity cross-modal matching that incorporates structure-aware distillation alongside the cross-modal matching loss. While the cross-modal matching loss constraints instance-level matching, the structure-aware distillation further regularizes the geometric consistency between learned matching representations and intra-modal representations through the developed relational matching. Extensive experiments on different datasets affirm the superior cross-modal retrieval performance of our approach, simultaneously enhancing single-modal retrieval capabilities compared to the baseline models.

📄 PDF Abstract BibTeX arXiv:2412.10761

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalRetrieval

Similar Papers 제목 키워드 기반

A Note on Universal Bilinear Portfolios

2019-07-23 · Alex Garivaltis

This note provides a neat and enjoyable expansion and application of the magnificent Ordentlich-Cover theory of "universal portfolios." I generalize Cover's benchmark of the best constant-rebalanced portfolio (or 1-linea…

BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval

2024-06-14 · Imanol Miranda, Ander Salaberria, Eneko Agirre, Gorka Azkune

Existing Vision-Language Compositionality (VLC) benchmarks like SugarCrepe are formulated as image-to-text retrieval problems, where, given an image, the models need to select between the correct textual description and …

Image RetrievalImage to textImage-to-Text RetrievalRetrieval+1

Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval

2024-12-18 · Rui Cai, Zhiyu Dong, Jianfeng Dong, Xun Wang

Existing cross-modal retrieval methods typically rely on large-scale vision-language pair data. This makes it challenging to efficiently develop a cross-modal retrieval model for under-resourced languages of interest. Th…

Cross-Modal RetrievalRetrieval

Asymptotic Normality of Infinite Centered Random Forests -Application to Imbalanced Classification

2025-06-10 · Moria Mayala, Erwan Scornet, Charles Tillier, Olivier Wintenberger

Many classification tasks involve imbalanced data, in which a class is largely underrepresented. Several techniques consists in creating a rebalanced dataset on which a classifier is trained. In this paper, we study theo…

imbalanced classificationvalid

Universal portfolios in continuous time: an approach in pathwise Itô calculus

2025-04-16 · Xiyue Han, Alexander Schied

We provide a simple and straightforward approach to a continuous-time version of Cover's universal portfolio strategies within the model-free context of F\"ollmer's pathwise It\^o calculus. We establish the existence of …