paper-with-me

홈 › Papers

Supervised contrastive learning from weakly-labeled audio segments for musical version matching

2025-02-24 · Joan Serrà, R. Oguz Araz, Dmitry Bogdanov, Yuki Mitsufuji

Detecting musical versions (different renditions of the same piece) is a challenging task with important applications. Because of the ground truth nature, existing approaches match musical versions at the track level (e.g., whole song). However, most applications require to match them at the segment level (e.g., 20s chunks). In addition, existing approaches resort to classification and triplet losses, disregarding more recent losses that could bring meaningful improvements. In this paper, we propose a method to learn from weakly annotated segments, together with a contrastive loss variant that outperforms well-studied alternatives. The former is based on pairwise segment distance reductions, while the latter modifies an existing loss following decoupling, hyper-parameter, and geometric considerations. With these two elements, we do not only achieve state-of-the-art results in the standard track-level evaluation, but we also obtain a breakthrough performance in a segment-level evaluation. We believe that, due to the generality of the challenges addressed here, the proposed methods may find utility in domains beyond audio or musical version matching.

📄 PDF Abstract BibTeX arXiv:2502.16936

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningTriplet

Similar Papers 제목 키워드 기반

Weakly-Supervised Audio-Visual Video Parsing with Prototype-based Pseudo-Labeling

2024-01-01 · CVPR 2024 1 · Kranthi Kumar Rachavarapu, Kalyan Ramakrishnan, Rajagopalan A. N.

In this paper we address the weakly-supervised Audio-Visual Video Parsing (AVVP) problem which aims at labeling events in a video as audible visible or both and temporally localizing and classifying them into known c…

Contrastive LearningMultiple Instance Learning

Revisit Weakly-Supervised Audio-Visual Video Parsing from the Language Perspective

2023-09-21 · NeurIPS 2023 11

We focus on the weakly-supervised audio-visual video parsing task (AVVP), which aims to identify and locate all the events in audio/visual modalities. Previous works only concentrate on video-level overall label denoisin…

Weakly Supervised Representation Learning for Unsynchronized Audio-Visual Events

2018-04-19 · Sanjeel Parekh, Slim Essid, Alexey Ozerov, Ngoc Q. K. Duong 외

Audio-visual representation learning is an important task from the perspective of designing machines with the ability to understand complex events. To this end, we propose a novel multimodal framework that instantiates m…

Multiple Instance LearningRepresentation Learning

Audio Event and Scene Recognition: A Unified Approach using Strongly and Weakly Labeled Data

2016-11-12 · Anurag Kumar, Bhiksha Raj

In this paper we propose a novel learning framework called Supervised and Weakly Supervised Learning where the goal is to learn simultaneously from weakly and strongly labeled data. Strongly labeled data can be simply un…

Scene RecognitionWeakly-supervised Learning

Self-supervised Attention Model for Weakly Labeled Audio Event Classification

2019-08-07 · Bongjun Kim, Shabnam Ghaffarzadegan

We describe a novel weakly labeled Audio Event Classification approach based on a self-supervised attention model. The weakly labeled framework is used to eliminate the need for expensive data labeling procedure and self…

ClassificationGeneral Classification