paper-with-me

홈 › Papers

Speaker-independent Speech Separation with Deep Attractor Network

2017-07-12 · Yi Luo, Zhuo Chen, Nima Mesgarani

Despite the recent success of deep learning for many speech processing tasks, single-microphone, speaker-independent speech separation remains challenging for two main reasons. The first reason is the arbitrary order of the target and masker speakers in the mixture permutation problem, and the second is the unknown number of speakers in the mixture output dimension problem. We propose a novel deep learning framework for speech separation that addresses both of these issues. We use a neural network to project the time-frequency representation of the mixture signal into a high-dimensional embedding space. A reference point attractor is created in the embedding space to represent each speaker which is defined as the centroid of the speaker in the embedding space. The time-frequency embeddings of each speaker are then forced to cluster around the corresponding attractor point which is used to determine the time-frequency assignment of the speaker. We propose three methods for finding the attractors for each source in the embedding space and compare their advantages and limitations. The objective function for the network is standard signal reconstruction error which enables end-to-end operation during both training and test phases. We evaluated our system using the Wall Street Journal dataset WSJ0 on two and three speaker mixtures and report comparable or better performance than other state-of-the-art deep learning methods for speech separation.

📄 PDF Abstract BibTeX arXiv:1707.03634

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningSpeech Separation

Similar Papers 제목 키워드 기반

Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers

2025-05-22 · Yuzhu Wang, Archontis Politis, Konstantinos Drossos, Tuomas Virtanen

This paper addresses the problem of single-channel speech separation, where the number of speakers is unknown, and each speaker may speak multiple utterances. We propose a speech separation model that simultaneously perf…

Speech Separation

Exploring the time-domain deep attractor network with two-stream architectures in a reverberant environment

2020-07-01 · Hangting Chen, Pengyuan Zhang

Deep attractor networks (DANs) perform speech separation with discriminative embeddings and speaker attractors. Compared with methods based on the permutation invariant training (PIT), DANs define a deep embedding space …

Speech Separation

SATTS: Speaker Attractor Text to Speech, Learning to Speak by Learning to Separate

2022-07-13 · Nabarun Goswami, Tatsuya Harada

The mapping of text to speech (TTS) is non-deterministic, letters may be pronounced differently based on context, or phonemes can vary depending on various physiological and stylistic factors like gender, age, accent, em…

Speech Separationtext-to-speechText to Speech

Multi-talker Speech Separation with Utterance-level Permutation Invariant Training of Deep Recurrent Neural Networks

2017-03-18 · Morten Kolbæk, Dong Yu, Zheng-Hua Tan, Jesper Jensen

In this paper we propose the utterance-level Permutation Invariant Training (uPIT) technique. uPIT is a practically applicable, end-to-end, deep learning based solution for speaker independent multi-talker speech separat…

ClusteringDeep ClusteringSpeech Separation

EEND-SS: Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakers

2022-03-31 · Soumi Maiti, Yushi Ueda, Shinji Watanabe, Chunlei Zhang 외

In this paper, we present a novel framework that jointly performs three tasks: speaker diarization, speech separation, and speaker counting. Our proposed framework integrates speaker diarization based on end-to-end neura…

Decoderspeaker-diarizationSpeaker DiarizationSpeech Separation