paper-with-me

홈 › Papers

Cross-Modal Retrieval with Cauchy-Schwarz Divergence

2025-09-15 · Jiahao Zhang, Wenzhe Yin, Shujian Yu arxiv

Effective cross-modal retrieval requires robust alignment of heterogeneous data types. Most existing methods focus on bi-modal retrieval tasks and rely on distributional alignment techniques such as Kullback-Leibler divergence, Maximum Mean Discrepancy, and correlation alignment. However, these methods often suffer from critical limitations, including numerical instability, sensitivity to hyperparameters, and their inability to capture the full structure of the underlying distributions. In this paper, we introduce the Cauchy-Schwarz (CS) divergence, a hyperparameter-free measure that improves both training stability and retrieval performance. We further propose a novel Generalized CS (GCS) divergence inspired by Hölder's inequality. This extension enables direct alignment of three or more modalities within a unified mathematical framework through a bidirectional circular comparison scheme, eliminating the need for exhaustive pairwise comparisons. Extensive experiments on six benchmark datasets demonstrate the effectiveness of our method in both bi-modal and tri-modal retrieval tasks. The code of our CS/GCS divergence is publicly available at https://github.com/JiahaoZhang666/CSD.

📄 PDF Abstract BibTeX arXiv:2509.21339

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal Retrieval

Similar Papers 제목 키워드 기반

Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence

2025-02-24 · Wenzhe Yin, Zehao Xiao, Pan Zhou, Shujian Yu 외

Multimodal alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize mutual information, primarily aligning pairw…

Image GenerationRetrievalText to Image GenerationText-to-Image Generation

LAMB: LLM-based Audio Captioning with Modality Gap Bridging via Cauchy-Schwarz Divergence

2026-01-08 · Hyeongkeun Lee, Jongmin Choi, KiHyun Nam, Joon Son Chung arxiv

Automated Audio Captioning aims to describe the semantic content of input audio. Recent works have employed large language models (LLMs) as a text decoder to leverage their reasoning capabilities. However, prior approach…

Audio captioning

Efficient Sensor Management for Multitarget Tracking in Passive Sensor Networks via Cauchy-Schwarz Divergence

2020-11-03 · Yun Zhu

This paper presents an efficient sensor management approach for multi-target tracking in passive sensor networks. Compared with active sensor networks, passive sensor networks have larger uncertainty due to the nature of…

Management

On Hölder projective divergences

2017-01-14 · Frank Nielsen, Ke Sun, Stéphane Marchand-Maillet

We describe a framework to build distances by measuring the tightness of inequalities, and introduce the notion of proper statistical divergences and improper pseudo-divergences. We then consider the H\"older ordinary an…

Clustering

The Conditional Cauchy-Schwarz Divergence with Applications to Time-Series Data and Sequential Decision Making

2023-01-21 · Shujian Yu, Hongming Li, Sigurd Løkse, Robert Jenssen 외

The Cauchy-Schwarz (CS) divergence was developed by Pr\'{i}ncipe et al. in 2000. In this paper, we extend the classic CS divergence to quantify the closeness between two conditional distributions and show that the develo…

Decision MakingSequential Decision MakingTime SeriesTime Series Analysis+1