paper-with-me

홈 › Papers

Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training

2024-06-03 · Jan Melechovsky, Ambuj Mehrish, Berrak Sisman, Dorien Herremans

With rapid globalization, the need to build inclusive and representative speech technology cannot be overstated. Accent is an important aspect of speech that needs to be taken into consideration while building inclusive speech synthesizers. Inclusive speech technology aims to erase any biases towards specific groups, such as people of certain accent. We note that state-of-the-art Text-to-Speech (TTS) systems may currently not be suitable for all people, regardless of their background, as they are designed to generate high-quality voices without focusing on accent. In this paper, we propose a TTS model that utilizes a Multi-Level Variational Autoencoder with adversarial learning to address accented speech synthesis and conversion in TTS, with a vision for more inclusive systems in the future. We evaluate the performance through both objective metrics and subjective listening tests. The results show an improvement in accent conversion ability compared to the baseline.

📄 PDF Abstract BibTeX arXiv:2406.01018

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Convert and Speak: Zero-shot Accent Conversion with Minimum Supervision

2024-08-19 · Zhijun Jia, Huaying Xue, Xiulian Peng, Yan Lu

Low resource of parallel data is the key challenge of accent conversion(AC) problem in which both the pronunciation units and prosody pattern need to be converted. We propose a two-stage generative framework "convert-and…

TTS-Guided Training for Accent Conversion Without Parallel Data

2022-12-20 · Yi Zhou, Zhizheng Wu, Mingyang Zhang, Xiaohai Tian 외

Accent Conversion (AC) seeks to change the accent of speech from one (source) to another (target) while preserving the speech content and speaker identity. However, many AC approaches rely on source-target parallel speec…

Decodertext-to-speechText to Speech

Improving Accent Conversion with Reference Encoder and End-To-End Text-To-Speech

2020-05-19 · Wenjie Li, Benlai Tang, Xiang Yin, Yushi Zhao 외

Accent conversion (AC) transforms a non-native speaker's accent into a native accent while maintaining the speaker's voice timbre. In this paper, we propose approaches to improving accent conversion applicability, as wel…

text-to-speechText to Speech

Non-autoregressive real-time Accent Conversion model with voice cloning

2024-05-21 · Vladimir Nechaev, Sergey Kosyakov

Currently, the development of Foreign Accent Conversion (FAC) models utilizes deep neural network architectures, as well as ensembles of neural networks for speech recognition and speech generation. The use of these mode…

Speech Enhancementspeech-recognitionSpeech RecognitionVoice Cloning

Accent conversion using discrete units with parallel data synthesized from controllable accented TTS

2024-09-30 · Tuan Nam Nguyen, Ngoc Quan Pham, Alexander Waibel

The goal of accent conversion (AC) is to convert speech accents while preserving content and speaker identity. Previous methods either required reference utterances during inference, did not preserve speaker identity wel…

Data AugmentationSpeech Synthesistext-to-speechText to Speech+1