paper-with-me

Papers

GASS: Generalizing Audio Source Separation with Large-scale Data

2023-09-29 · Jordi Pons, Xiaoyu Liu, Santiago Pascual, Joan Serrà

Universal source separation targets at separating the audio sources of an arbitrary mix, removing the constraint to operate on a specific domain like speech or music. Yet, the potential of universal source separation is limited because most existing works focus on mixes with predominantly sound events, and small training datasets also limit its potential for supervised learning. Here, we study a single general audio source separation (GASS) model trained to separate speech, music, and sound events in a supervised fashion with a large-scale dataset. We assess GASS models on a diverse set of tasks. Our strong in-distribution results show the feasibility of GASS models, and the competitive out-of-distribution performance in sound event and speech separation shows its generalization abilities. Yet, it is challenging for GASS models to generalize for separating out-of-distribution cinematic and music content. We also fine-tune GASS models on each dataset and consistently outperform the ones without pre-training. All fine-tuned models (except the music separation one) obtain state-of-the-art results in their respective benchmarks.

📄 PDF Abstract BibTeX arXiv:2310.00140

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Source SeparationSpeech Separation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Zero-shot Audio Source Separation through Query-based Learningfrom Weakly-labeled Data

2021-12-15 · AAAI 2021 12 · Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma 외

Deep learning techniques for separating audio into different sound sources face several challenges. Standard architectures require training separate models for different types of audio sources. Although some universal se…

Audio Source SeparationEvent DetectionSound Event DetectionZero-shot Generalization

Zero-shot Audio Source Separation through Query-based Learning from Weakly-labeled Data

2021-12-15 · Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma 외

Deep learning techniques for separating audio into different sound sources face several challenges. Standard architectures require training separate models for different types of audio sources. Although some universal se…

Audio Source SeparationAudio TaggingEvent DetectionSound Event Detection+1

A Generalized Bandsplit Neural Network for Cinematic Audio Source Separation

2023-09-05 · Karn N. Watcharasupat, Chih-Wei Wu, Yiwei Ding, Iroro Orife 외

Cinematic audio source separation is a relatively new subtask of audio source separation, with the aim of extracting the dialogue, music, and effects stems from their mixture. In this work, we developed a model generaliz…

Audio Source Separation

Separate Anything You Describe

2023-08-09 · Xubo Liu, Qiuqiang Kong, Yan Zhao, Haohe Liu 외

Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given a natural language query, which provide…

Audio Source SeparationNatural Language QueriesSpeech EnhancementZero-shot Generalization

OpenSep: Leveraging Large Language Models with Textual Inversion for Open World Audio Separation

2024-09-28 · Tanvir Mahmud, Diana Marculescu

Audio separation in real-world scenarios, where mixtures contain a variable number of sources, presents significant challenges due to limitations of existing models, such as over-separation, under-separation, and depende…

Audio captioning