paper-with-me

홈 › Papers

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

2025-06-13 · ICLR 2025 4 · Tony Alex, Sara Ahmed, Armin Mustafa, Muhammad Awais, Philip JB Jackson

Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the SSL pre-training has sufficiently equipped them to handle real-world audio. However, a critical question remains: how well do these models actually perform in real-world conditions, where audio is typically polyphonic and complex, involving multiple overlapping sound sources? Current audio SSL methods are often benchmarked on datasets predominantly featuring monophonic audio, such as environmental sounds, and speech. As a result, the ability of SSL models to generalize to polyphonic audio, a common characteristic in natural scenarios, remains underexplored. This limitation raises concerns about the practical robustness of SSL models in more realistic audio settings. To address this gap, we introduce Self-Supervised Learning from Audio Mixtures (SSLAM), a novel direction in audio SSL research, designed to improve, designed to improve the model's ability to learn from polyphonic data while maintaining strong performance on monophonic data. We thoroughly evaluate SSLAM on standard audio SSL benchmark datasets which are predominantly monophonic and conduct a comprehensive comparative analysis against SOTA methods using a range of high-quality, publicly available polyphonic datasets. SSLAM not only improves model performance on polyphonic audio, but also maintains or exceeds performance on standard audio SSL benchmarks. Notably, it achieves up to a 3.9\% improvement on the AudioSet-2M (AS-2M), reaching a mean average precision (mAP) of 50.2. For polyphonic datasets, SSLAM sets new SOTA in both linear evaluation and fine-tuning regimes with performance improvements of up to 9.1\% (mAP).

📄 PDF Abstract BibTeX arXiv:2506.12222

Code (1)

ta012/SSLAM pytorch

Tasks

Linear evaluationSelf-Supervised Learning

Similar Papers 제목 키워드 기반

T-VSL: Text-Guided Visual Sound Source Localization in Mixtures

2024-04-02 · CVPR 2024 1 · Tanvir Mahmud, Yapeng Tian, Diana Marculescu

Visual sound source localization poses a significant challenge in identifying the semantic region of each sounding source within a video. Existing self-supervised and weakly supervised source localization methods struggl…

Sound Source Localization

Self-Supervised Learning from Automatically Separated Sound Scenes

2021-05-05 · Eduardo Fonseca, Aren Jansen, Daniel P. W. Ellis, Scott Wisdom 외

Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The association of these constituent sound events wit…

Contrastive LearningSelf-Supervised Learning

Toward Fully Self-Supervised Multi-Pitch Estimation

2024-02-23 · Frank Cwitkowitz, Zhiyao Duan

Multi-pitch estimation is a decades-long research problem involving the detection of pitch activity associated with concurrent musical events within multi-instrument mixtures. Supervised learning techniques have demonstr…

Self-Supervised Learning

Universal Sound Separation with Self-Supervised Audio Masked Autoencoder

2024-07-16 · Junqi Zhao, Xubo Liu, Jinzheng Zhao, Yi Yuan 외

Universal sound separation (USS) is a task of separating mixtures of arbitrary sound sources. Typically, universal separation models are trained from scratch in a supervised manner, using labeled data. Self-supervised le…

Self-Supervised Learning

Weakly-supervised Audio Separation via Bi-modal Semantic Similarity

2024-04-02 · Tanvir Mahmud, Saeed Amizadeh, Kazuhito Koishida, Diana Marculescu

Conditional sound separation in multi-source audio mixtures without having access to single source sound data during training is a long standing challenge. Existing mix-and-separate based methods suffer from significant …

Semantic SimilaritySemantic Textual Similarity