paper-with-me

Papers

Streaming Voice Conversion Via Intermediate Bottleneck Features And Non-streaming Teacher Guidance

2022-10-27 · Yuanzhe Chen, Ming Tu, Tang Li, Xin Li, Qiuqiang Kong, Jiaxin Li, Zhichao Wang, Qiao Tian, Yuping Wang, Yuxuan Wang

Streaming voice conversion (VC) is the task of converting the voice of one person to another in real-time. Previous streaming VC methods use phonetic posteriorgrams (PPGs) extracted from automatic speech recognition (ASR) systems to represent speaker-independent information. However, PPGs lack the prosody and vocalization information of the source speaker, and streaming PPGs contain undesired leaked timbre of the source speaker. In this paper, we propose to use intermediate bottleneck features (IBFs) to replace PPGs. VC systems trained with IBFs retain more prosody and vocalization information of the source speaker. Furthermore, we propose a non-streaming teacher guidance (TG) framework that addresses the timbre leakage problem. Experiments show that our proposed IBFs and the TG framework achieve a state-of-the-art streaming VC naturalness of 3.85, a content consistency of 3.77, and a timbre similarity of 3.77 under a future receptive field of 160 ms which significantly outperform previous streaming VC systems.

📄 PDF Abstract BibTeX arXiv:2210.15158

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionVoice Conversion

Similar Papers 제목 키워드 기반

StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion

2024-08-05 · Zhichao Wang, Yuanzhe Chen, Xinsheng Wang, Lei Xie 외

StreamVoice has recently pushed the boundaries of zero-shot voice conversion (VC) in the streaming domain. It uses a streamable language model (LM) with a context-aware approach to convert semantic features from automati…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Modellingspeech-recognition+2

StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion

2024-01-19 · Zhichao Wang, Yuanzhe Chen, Xinsheng Wang, Lei Xie 외

Recent language model (LM) advancements have showcased impressive zero-shot voice conversion (VC) performance. However, existing LM-based VC models usually apply offline conversion from source semantics to acoustic featu…

Language ModelingLanguage ModellingVoice Conversion

DualVC: Dual-mode Voice Conversion using Intra-model Knowledge Distillation and Hybrid Predictive Coding

2023-05-21 · Ziqian Ning, Yuepeng Jiang, Pengcheng Zhu, Jixun Yao 외

Voice conversion is an increasingly popular technology, and the growing number of real-time applications requires models with streaming conversion capabilities. Unlike typical (non-streaming) voice conversion, which can …

Data AugmentationDecoderKnowledge DistillationVoice Conversion

Effects of Convolutional Autoencoder Bottleneck Width on StarGAN-based Singing Technique Conversion

2023-08-19 · Tung-Cheng Su, Yung-Chuan Chang, Yi-Wen Liu

Singing technique conversion (STC) refers to the task of converting from one voice technique to another while leaving the original singer identity, melody, and linguistic components intact. Previous STC studies, as well …

Voice Conversion

End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions

2022-05-19 · Wonjune Kang, Mark Hasegawa-Johnson, Deb Roy

Zero-shot voice conversion is becoming an increasingly popular research topic, as it promises the ability to transform speech to sound like any speaker. However, relatively little work has been done on end-to-end methods…

Speech SynthesisStyle TransferVoice Conversion