paper-with-me

Papers

Multi-Task Audio Source Separation

2021-07-14 · Lu Zhang, Chenxing Li, Feng Deng, Xiaorui Wang

The audio source separation tasks, such as speech enhancement, speech separation, and music source separation, have achieved impressive performance in recent studies. The powerful modeling capabilities of deep neural networks give us hope for more challenging tasks. This paper launches a new multi-task audio source separation (MTASS) challenge to separate the speech, music, and noise signals from the monaural mixture. First, we introduce the details of this task and generate a dataset of mixtures containing speech, music, and background noises. Then, we propose an MTASS model in the complex domain to fully utilize the differences in spectral characteristics of the three audio signals. In detail, the proposed model follows a two-stage pipeline, which separates the three types of audio signals and then performs signal compensation separately. After comparing different training targets, the complex ratio mask is selected as a more suitable target for the MTASS. The experimental results also indicate that the residual signal compensation module helps to recover the signals further. The proposed model shows significant advantages in separation performance over several well-known separation models.

📄 PDF Abstract BibTeX arXiv:2107.06467

Code (1)

Windstudent/Complex-MTASSNet 공식 구현 pytorch

Tasks

Audio Source SeparationMulti-task Audio Source SeperationMusic Source SeparationSpeech EnhancementSpeech Separation

Similar Papers 제목 키워드 기반

Separate Anything You Describe

2023-08-09 · Xubo Liu, Qiuqiang Kong, Yan Zhao, Haohe Liu 외

Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given a natural language query, which provide…

Audio Source SeparationNatural Language QueriesSpeech EnhancementZero-shot Generalization

Visual Scene Graphs for Audio Source Separation

2021-09-24 · ICCV 2021 10 · Moitreya Chatterjee, Jonathan Le Roux, Narendra Ahuja, Anoop Cherian

State-of-the-art approaches for visually-guided audio source separation typically assume sources that have characteristic sounds, such as musical instruments. These approaches often ignore the visual context of these sou…

Audio Source SeparationVisually Guided Sound Source Separation

Training-Free Multi-Step Audio Source Separation

2025-05-26 · Yongyi Zang, Jingyi Li, Qiuqiang Kong

Audio source separation aims to separate a mixture into target sources. Previous audio source separation systems usually conduct one-step inference, which does not fully explore the separation ability of models. In this …

Audio Source SeparationDenoisingMusic Source SeparationSpeech Enhancement

Learning to Separate Object Sounds by Watching Unlabeled Video

2018-04-05 · ECCV 2018 9 · Ruohan Gao, Rogerio Feris, Kristen Grauman

Perceiving a scene most fully requires all the senses. Yet modeling how objects look and sound is challenging: most natural scenes and events contain multiple objects, and the audio track mixes all the sound sources toge…

Audio DenoisingAudio Source SeparationDenoisingMulti-Label Learning

ZeroSep: Separate Anything in Audio with Zero Training

2025-05-29 · Chao Huang, Yuesheng Ma, Junxuan Huang, Susan Liang 외

Audio source separation is fundamental for machines to understand complex acoustic environments and underpins numerous audio applications. Current supervised deep learning approaches, while powerful, are limited by the n…

Audio Source SeparationDenoising