paper-with-me

홈 › Papers

Unbiased Sliced Wasserstein Kernels for High-Quality Audio Captioning

2025-02-08 · Manh Luong, Khai Nguyen, Dinh Phung, Gholamreza Haffari, Lizhen Qu

Teacher-forcing training for audio captioning usually leads to exposure bias due to training and inference mismatch. Prior works propose the contrastive method to deal with caption degeneration. However, the contrastive method ignores the temporal information when measuring similarity across acoustic and linguistic modalities, leading to inferior performance. In this work, we develop the temporal-similarity score by introducing the unbiased sliced Wasserstein RBF (USW-RBF) kernel equipped with rotary positional embedding to account for temporal information across modalities. In contrast to the conventional sliced Wasserstein RBF kernel, we can form an unbiased estimation of USW-RBF kernel via Monte Carlo estimation. Therefore, it is well-suited to stochastic gradient optimization algorithms, and its approximation error decreases at a parametric rate of $\mathcal{O}(L^{-1/2})$ with $L$ Monte Carlo samples. Additionally, we introduce an audio captioning framework based on the unbiased sliced Wasserstein kernel, incorporating stochastic decoding methods to mitigate caption degeneration during the generation process. We conduct extensive quantitative and qualitative experiments on two datasets, AudioCaps and Clotho, to illustrate the capability of generating high-quality audio captions. Experimental results show that our framework is able to increase caption length, lexical diversity, and text-to-audio self-retrieval accuracy.

📄 PDF Abstract BibTeX arXiv:2502.05435

Code (0)

등록된 구현이 없습니다.

Tasks

AudioCapsAudio captioning

Methods 이 논문이 사용한 방법론

RBF 설명 없음

Similar Papers 제목 키워드 기반

Sliced Wasserstein Kernels for Probability Distributions

2015-11-10 · CVPR 2016 6 · Soheil Kolouri, Yang Zou, Gustavo K. Rohde

Optimal transport distances, otherwise known as Wasserstein distances, have recently drawn ample attention in computer vision and machine learning as a powerful discrepancy measure for probability distributions. The rece…

BIG-bench Machine Learning

Quasi-Monte Carlo for 3D Sliced Wasserstein

2023-09-21 · Khai Nguyen, Nicola Bariletto, Nhat Ho

Monte Carlo (MC) integration has been employed as the standard approximation method for the Sliced Wasserstein (SW) distance, whose analytical expression involves an intractable expectation. However, MC integration is no…

Stochastic OptimizationStyle Transfer

ReSWD: ReSTIR'd, not shaken. Combining Reservoir Sampling and Sliced Wasserstein Distance for Variance Reduction

2025-10-01 · Mark Boss, Andreas Engelhardt, Simon Donné, Varun Jampani arxiv

Distribution matching is central to many vision and graphics tasks, where the widely used Wasserstein distance is too costly to compute for high dimensional distributions. The Sliced Wasserstein Distance (SWD) offers a s…

Sliced Wasserstein Kernel for Persistence Diagrams

2017-06-11 · ICML 2017 8 · Mathieu Carrière, Marco Cuturi, Steve Oudot

Persistence diagrams (PDs) play a key role in topological data analysis (TDA), in which they are routinely used to describe topological properties of complicated shapes. PDs enjoy strong stability properties and have pro…

Graph ClassificationTopological Data Analysis

Distance-Matrix Wasserstein Statistics for Scalable Gromov--Wasserstein Learning

2026-05-14 · Ao Xu, Tieru Wu arxiv

Gromov--Wasserstein (GW) distances compare graphs, shapes, and point clouds through internal distances, without requiring a common coordinate system. This invariance is powerful, but discrete GW is a nonconvex quadratic …

Graph ClassificationTwo-sample testingPoint Clouds