paper-with-me

홈 › Papers

SATTS: Speaker Attractor Text to Speech, Learning to Speak by Learning to Separate

2022-07-13 · Nabarun Goswami, Tatsuya Harada

The mapping of text to speech (TTS) is non-deterministic, letters may be pronounced differently based on context, or phonemes can vary depending on various physiological and stylistic factors like gender, age, accent, emotions, etc. Neural speaker embeddings, trained to identify or verify speakers are typically used to represent and transfer such characteristics from reference speech to synthesized speech. Speech separation on the other hand is the challenging task of separating individual speakers from an overlapping mixed signal of various speakers. Speaker attractors are high-dimensional embedding vectors that pull the time-frequency bins of each speaker's speech towards themselves while repelling those belonging to other speakers. In this work, we explore the possibility of using these powerful speaker attractors for zero-shot speaker adaptation in multi-speaker TTS synthesis and propose speaker attractor text to speech (SATTS). Through various experiments, we show that SATTS can synthesize natural speech from text from an unseen target speaker's reference signal which might have less than ideal recording conditions, i.e. reverberations or mixed with other speakers.

📄 PDF Abstract BibTeX arXiv:2207.06011

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Separationtext-to-speechText to Speech

Similar Papers 제목 키워드 기반

End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors

2020-05-20 · Shota Horiguchi, Yusuke Fujita, Shinji Watanabe, Yawen Xue 외

End-to-end speaker diarization for an unknown number of speakers is addressed in this paper. Recently proposed end-to-end speaker diarization outperformed conventional clustering-based speaker diarization, but it has one…

ClusteringDecoderspeaker-diarizationSpeaker Diarization

Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers

2025-05-22 · Yuzhu Wang, Archontis Politis, Konstantinos Drossos, Tuomas Virtanen

This paper addresses the problem of single-channel speech separation, where the number of speakers is unknown, and each speaker may speak multiple utterances. We propose a speech separation model that simultaneously perf…

Speech Separation

SAMO: Speaker Attractor Multi-Center One-Class Learning for Voice Anti-Spoofing

2022-11-04 · Siwen Ding, You Zhang, Zhiyao Duan

Voice anti-spoofing systems are crucial auxiliaries for automatic speaker verification (ASV) systems. A major challenge is caused by unseen attacks empowered by advanced speech synthesis technologies. Our previous resear…

DiversitySpeaker VerificationSpeech SynthesisVoice Anti-spoofing

Speaker-independent Speech Separation with Deep Attractor Network

2017-07-12 · Yi Luo, Zhuo Chen, Nima Mesgarani

Despite the recent success of deep learning for many speech processing tasks, single-microphone, speaker-independent speech separation remains challenging for two main reasons. The first reason is the arbitrary order of …

Deep LearningSpeech Separation

Selective Listening by Synchronizing Speech with Lips

2021-06-14 · Zexu Pan, Ruijie Tao, Chenglin Xu, Haizhou Li

A speaker extraction algorithm seeks to extract the speech of a target speaker from a multi-talker speech mixture when given a cue that represents the target speaker, such as a pre-enrolled speech utterance, or an accomp…

Lip ReadingTarget Speaker Extraction