paper-with-me

Papers

Multi-Format Contrastive Learning of Audio Representations

2021-03-11 · Luyu Wang, Aaron van den Oord

Recent advances suggest the advantage of multi-modal training in comparison with single-modal methods. In contrast to this view, in our work we find that similar gain can be obtained from training with different formats of a single modality. In particular, we investigate the use of the contrastive learning framework to learn audio representations by maximizing the agreement between the raw audio and its spectral representation. We find a significant gain using this multi-format strategy against the single-format counterparts. Moreover, on the downstream AudioSet and ESC-50 classification task, our audio-only approach achieves new state-of-the-art results with a mean average precision of 0.376 and an accuracy of 90.5%, respectively.

📄 PDF Abstract BibTeX arXiv:2103.06508

Code (0)

등록된 구현이 없습니다.

Tasks

Audio ClassificationContrastive Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

CALM: Contrastive Aligned Audio-Language Multirate and Multimodal Representations

2022-02-08 · Vin Sachidananda, Shao-Yen Tseng, Erik Marchi, Sachin Kajarekar 외

Deriving multimodal representations of audio and lexical inputs is a central problem in Natural Language Understanding (NLU). In this paper, we present Contrastive Aligned Audio-Language Multirate and Multimodal Represen…

Emotion RecognitionNatural Language Understanding

Enriched Music Representations with Multiple Cross-modal Contrastive Learning

2021-04-01 · Andres Ferraro, Xavier Favory, Konstantinos Drossos, Yuntae Kim 외

Modeling various aspects that make a music piece unique is a challenging task, requiring the combination of multiple sources of information. Deep learning is commonly used to obtain representations using various sources …

Contrastive LearningGenre classification

Joint-Centric Dual Contrastive Alignment with Structure-Preserving and Information-Balanced Regularization

2026-04-17 · Habibeh Naderi, Behrouz Haji Soleimani, Stan Matwin arxiv

We propose HILBERT (HIerarchical Long-sequence Balanced Embedding with Reciprocal contrastive Training), a cross-attentive multimodal framework for learning document-level audio-text representations from long, segmented …

FLAP: Fast Language-Audio Pre-training

2023-11-02 · Ching-Feng Yeh, Po-Yao Huang, Vasu Sharma, Shang-Wen Li 외

We propose Fast Language-Audio Pre-training (FLAP), a self-supervised approach that efficiently and effectively learns aligned audio and language representations through masking, contrastive learning and reconstruction. …

AudioCapsContrastive LearningRetrievalText Retrieval

Contrastive Self-Supervised Learning of Global-Local Audio-Visual Representations

2021-01-01 · Shuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale Song

Contrastive self-supervised learning has delivered impressive results in many audio-visual recognition tasks. However, existing approaches optimize for learning either global representations useful for high-level underst…

ClassificationDeepFake DetectionFace SwappingGeneral Classification+4