Speech Separation
19개 벤치마크 · 논문 384편 · 이 태스크의 논문 보기 →
Benchmarks
WSJ0-2mix
WHAMR!
Libri2Mix
WSJ0-3mix
LRS2
WHAM!
WSJ0-5mix
LRS3
VoxCeleb2
WSJ0-4mix
Libri5Mix
Libri10Mix
Libri20Mix
LibriCSS
Libri15Mix
WSJ0-2mix-16k
iKala
Most implemented
Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation
Dual-path RNN: efficient long sequence modeling for time-domain single-channel speech separation
Deep clustering: Discriminative embeddings for segmentation and separation
WaveCRN: An Efficient Convolutional Recurrent Neural Network for End-to-end Speech Enhancement
Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
Attention is All You Need in Speech Separation
Papers
DuplexChat: Constructing Speaker-Separated Full-Duplex Dialogue Speech at Scale for Spoken Dialogue Language Modeling
Full-duplex spoken dialogue models are trained on conversational speech in which each speaker is represented as a separate stream, but existing large-scale public speech corpora are mostly monaural, making them unsuited …
Speech SeparationTF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation
Recent advances in speech separation (SS) have led to compact front-end models with small parameter sizes, yet their high computational cost remains a major barrier for deployment on edge devices. To address this, we pro…
Speech SeparationMeCo: One-Step MeanFlow-based Corrector for Multi-Channel Speech Separation
While discriminative models for multi-channel speech separation excel in reference-based metrics, they often exhibit suboptimal human listening quality. To address this, we propose a novel MeanFlow-based one-step generat…
Speech SeparationPredictive-Generative Drift Decomposition for Speech Enhancement and Separation
We propose a plug-and-play framework for speech enhancement and separation that augments predictive methods with a generative speech prior. Our approach, termed Stochastic Interpolant Prior for Speech (SIPS), builds on s…
Speech EnhancementSpeech SeparationA Brain-Inspired Deep Separation Network for Single Channel Raman Spectra Unmixing
Raman spectra obtained in real world applications are often a noisy combination of several spectra of various substances in a tested sample. Unmixing such spectra into individual components corresponding to each of the s…
Speech SeparationSSNAPS: Audio-Visual Separation of Speech and Background Noise with Diffusion Inverse Sampling
This paper addresses the challenge of audio-visual single-microphone speech separation and enhancement in the presence of real-world environmental noise. Our approach is based on generative inverse sampling, where we mod…
Speech Separation