paper-with-me

Papers

Explicit Intensity Control for Accented Text-to-speech

2022-10-27 · Rui Liu, Haolin Zuo, De Hu, Guanglai Gao, Haizhou Li

Accented text-to-speech (TTS) synthesis seeks to generate speech with an accent (L2) as a variant of the standard version (L1). How to control the intensity of accent in the process of TTS is a very interesting research direction, and has attracted more and more attention. Recent work design a speaker-adversarial loss to disentangle the speaker and accent information, and then adjust the loss weight to control the accent intensity. However, such a control method lacks interpretability, and there is no direct correlation between the controlling factor and natural accent intensity. To this end, this paper propose a new intuitive and explicit accent intensity control scheme for accented TTS. Specifically, we first extract the posterior probability, called as ``goodness of pronunciation (GoP)'' from the L1 speech recognition model to quantify the phoneme accent intensity for accented speech, then design a FastSpeech2 based TTS model, named Ai-TTS, to take the accent intensity expression into account during speech generation. Experiments show that the our method outperforms the baseline model in terms of accent rendering and intensity control.

📄 PDF Abstract BibTeX arXiv:2210.15364

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognitiontext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Controllable Accented Text-to-Speech Synthesis

2022-09-22 · Rui Liu, Berrak Sisman, Guanglai Gao, Haizhou Li

Accented text-to-speech (TTS) synthesis seeks to generate speech with an accent (L2) as a variant of the standard version (L1). Accented TTS synthesis is challenging as L2 is different from L1 in both in terms of phoneti…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Learning-free L2-Accented Speech Generation using Phonological Rules

2026-03-08 · Thanathai Lertpetchpun, Yoonjeong Lee, Jihwan Lee, Tiantian Feng 외 arxiv

Accent plays a crucial role in speaker identity and inclusivity in speech technologies. Existing accented text-to-speech (TTS) systems either require large-scale accented datasets or lack fine-grained phoneme-level contr…

Accented Text-to-Speech Synthesis with Limited Data

2023-05-08 · Xuehao Zhou, Mingyang Zhang, Yi Zhou, Zhizheng Wu 외

This paper presents an accented text-to-speech (TTS) synthesis framework with limited training data. We study two aspects concerning accent rendering: phonetic (phoneme difference) and prosodic (pitch pattern and phoneme…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data

2026-03-08 · Thanathai Lertpetchpun, Thanapat Trachu, Jihwan Lee, Tiantian Feng 외 arxiv

Accent is an integral part of society, reflecting multiculturalism and shaping how individuals express identity. The majority of English speakers are non-native (L2) speakers, yet current Text-To-Speech (TTS) systems pri…

Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis

2024-07-04 · Cong-Thanh Do, Shuhei Imai, Rama Doddipatla, Thomas Hain

This paper investigates the use of unsupervised text-to-speech synthesis (TTS) as a data augmentation method to improve accented speech recognition. TTS systems are trained with a small amount of accented speech training…

Accented Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentation+7