paper-with-me

홈 › Papers

Audio Source Separation with Discriminative Scattering Networks

2014-12-22 · Pablo Sprechmann, Joan Bruna, Yann Lecun

In this report we describe an ongoing line of research for solving single-channel source separation problems. Many monaural signal decomposition techniques proposed in the literature operate on a feature space consisting of a time-frequency representation of the input data. A challenge faced by these approaches is to effectively exploit the temporal dependencies of the signals at scales larger than the duration of a time-frame. In this work we propose to tackle this problem by modeling the signals using a time-frequency representation with multiple temporal resolutions. The proposed representation consists of a pyramid of wavelet scattering operators, which generalizes Constant Q Transforms (CQT) with extra layers of convolution and complex modulus. We first show that learning standard models with this multi-resolution setting improves source separation results over fixed-resolution methods. As study case, we use Non-Negative Matrix Factorizations (NMF) that has been widely considered in many audio application. Then, we investigate the inclusion of the proposed multi-resolution setting into a discriminative training regime. We discuss several alternatives using different deep neural network architectures.

📄 PDF Abstract BibTeX arXiv:1412.7022

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Source Separation

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Problems using deep generative models for probabilistic audio source separation

2020-11-03 · NeurIPS Workshop ICBINB 2020 12 · Maurice Frank, Maximilian Ilse

Recent advancements in deep generative modeling make it possible to learn prior distributions from complex data that subsequently can be used for Bayesian inference. However, we find that distributions learned by deep ge…

Audio Source SeparationBayesian Inference

High-Quality Visually-Guided Sound Separation from Diverse Categories

2023-07-31 · Chao Huang, Susan Liang, Yapeng Tian, Anurag Kumar 외

We propose DAVIS, a Diffusion-based Audio-VIsual Separation framework that solves the audio-visual sound source separation task through generative learning. Existing methods typically frame sound separation as a mask-bas…

ZeroSep: Separate Anything in Audio with Zero Training

2025-05-29 · Chao Huang, Yuesheng Ma, Junxuan Huang, Susan Liang 외

Audio source separation is fundamental for machines to understand complex acoustic environments and underpins numerous audio applications. Current supervised deep learning approaches, while powerful, are limited by the n…

Audio Source SeparationDenoising

Towards Reliable Objective Evaluation Metrics for Generative Singing Voice Separation Models

2025-07-15 · Paul A. Bereuter, Benjamin Stahl, Mark D. Plumbley, Alois Sontacchi

Traditional Blind Source Separation Evaluation (BSS-Eval) metrics were originally designed to evaluate linear audio source separation models based on methods such as time-frequency masking. However, recent generative mod…

Audio Source Separationblind source separation

TISDiSS: A Training-Time and Inference-Time Scalable Framework for Discriminative Source Separation

2025-09-19 · Yongsheng Feng, Yuetonghui Xu, Jiehui Luo, Hongjia Liu 외 arxiv

Source separation is a fundamental task in speech, music, and audio processing, and it also provides cleaner and larger data for training generative models. However, improving separation performance in practice often dep…

Speech Separation