paper-with-me

Papers

Mixture to Mixture: Leveraging Close-talk Mixtures as Weak-supervision for Speech Separation

2024-02-14 · Zhong-Qiu Wang

We propose mixture to mixture (M2M) training, a weakly-supervised neural speech separation algorithm that leverages close-talk mixtures as a weak supervision for training discriminative models to separate far-field mixtures. Our idea is that, for a target speaker, its close-talk mixture has a much higher signal-to-noise ratio (SNR) of the target speaker than any far-field mixtures, and hence could be utilized to design a weak supervision for separation. To realize this, at each training step we feed a far-field mixture to a deep neural network (DNN) to produce an intermediate estimate for each speaker, and, for each of considered close-talk and far-field microphones, we linearly filter the DNN estimates and optimize a loss so that the filtered estimates of all the speakers can sum up to the mixture captured by each of the considered microphones. Evaluation results on a 2-speaker separation task in simulated reverberant conditions show that M2M can effectively leverage close-talk mixtures as a weak supervision for separating far-field mixtures.

📄 PDF Abstract BibTeX arXiv:2402.09313

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker SeparationSpeech Separation

Similar Papers 제목 키워드 기반

ctPuLSE: Close-Talk, and Pseudo-Label Based Far-Field, Speech Enhancement

2024-07-28 · Zhong-Qiu Wang

The current dominant approach for neural speech enhancement is via purely-supervised deep learning on simulated pairs of far-field noisy-reverberant speech (i.e., mixtures) and clean speech. The trained models, however, …

Pseudo LabelSpeech Enhancement

SuperM2M: Supervised and Mixture-to-Mixture Co-Learning for Speech Enhancement and Noise-Robust ASR

2024-03-15 · Zhong-Qiu Wang

The current dominant approach for neural speech enhancement is based on supervised learning by using simulated training data. The trained models, however, often exhibit limited generalizability to real-recorded data. To …

Speaker SeparationSpeech Enhancement

Cross-Talk Reduction

2024-05-30 · Zhong-Qiu Wang, Anurag Kumar, Shinji Watanabe

While far-field multi-talker mixtures are recorded, each speaker can wear a close-talk microphone so that close-talk mixtures can be recorded at the same time. Although each close-talk mixture has a high signal-to-noise …

Speech Separation

USEV: Universal Speaker Extraction with Visual Cue

2021-09-30 · Zexu Pan, Meng Ge, Haizhou Li

A speaker extraction algorithm seeks to extract the target speaker's speech from a multi-talker speech mixture. The prior studies focus mostly on speaker extraction from a highly overlapped multi-talker speech mixture. H…

Knowledge Distillation for Neural Transducer-based Target-Speaker ASR: Exploiting Parallel Mixture/Single-Talker Speech Data

2023-05-25 · Takafumi Moriya, Hiroshi Sato, Tsubasa Ochiai, Marc Delcroix 외

Neural transducer (RNNT)-based target-speaker speech recognition (TS-RNNT) directly transcribes a target speaker's voice from a multi-talker mixture. It is a promising approach for streaming applications because it does …

Knowledge DistillationSpeech Extractionspeech-recognitionSpeech Recognition