paper-with-me

Papers

Single-Channel Speech Separation with Auxiliary Speaker Embeddings

2019-06-24 · Shuo Liu, Gil Keren, Björn Schuller

We present a novel source separation model to decompose asingle-channel speech signal into two speech segments belonging to two different speakers. The proposed model is a neural network based on residual blocks, and uses learnt speaker embeddings created from additional clean context recordings of the two speakers as input to assist in attributing the different time-frequency bins to the two speakers. In experiments, we show that the proposed model yields good performance in the source separation task, and outperforms the state-of-the-art baselines. Specifically, separating speech from the challenging VoxCeleb dataset, the proposed model yields 4.79dB signal-to-distortion ratio, 8.44dB signal-to-artifacts ratio and 7.11dB signal-to-interference ratio.

📄 PDF Abstract BibTeX arXiv:1906.09997

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Separation

Similar Papers 제목 키워드 기반

FaceFilter: Audio-visual speech separation using still images

2020-05-14 · Soo-Whan Chung, Soyeon Choe, Joon Son Chung, Hong-Goo Kang

The objective of this paper is to separate a target speaker's speech from a mixture of two speakers using a deep audio-visual speech separation network. Unlike previous works that used lip movement on video clips or pre-…

Speech Separation

New Insights on Target Speaker Extraction

2022-02-01 · Mohamed Elminshawi, Wolfgang Mack, Srikanth Raj Chetupalli, Soumitro Chakrabarty 외

Speaker extraction (SE) aims to segregate the speech of a target speaker from a mixture of interfering speakers with the help of auxiliary information. Several forms of auxiliary information have been employed in single-…

Speaker SeparationTarget Speaker Extraction

Many-Speakers Single Channel Speech Separation with Optimal Permutation Training

2021-04-18 · Shaked Dovrat, Eliya Nachmani, Lior Wolf

Single channel speech separation has experienced great progress in the last few years. However, training neural speech separation for a large number of speakers (e.g., more than 10 speakers) is out of reach for the curre…

Speech Separation

Neural Blind Source Separation and Diarization for Distant Speech Recognition

2024-06-12 · Yoshiaki Bando, Tomohiko Nakamura, Shinji Watanabe

This paper presents a neural method for distant speech recognition (DSR) that jointly separates and diarizes speech mixtures without supervision by isolated signals. A standard separation method for multi-talker DSR is a…

blind source separationDistant Speech Recognitionspeaker-diarizationSpeaker Diarization+2

Joint Sound Source Separation and Speaker Recognition

2016-04-29 · Jeroen Zegers, Hugo Van hamme

Non-negative Matrix Factorization (NMF) has already been applied to learn speaker characterizations from single or non-simultaneous speech for speaker recognition applications. It is also known for its good performance i…

blind source separationSpeaker Recognition