paper-with-me

Papers

A Study of Modeling Rising Intonation in Cantonese Neural Speech Synthesis

2022-08-03 · Qibing Bai, Tom Ko, Yu Zhang

In human speech, the attitude of a speaker cannot be fully expressed only by the textual content. It has to come along with the intonation. Declarative questions are commonly used in daily Cantonese conversations, and they are usually uttered with rising intonation. Vanilla neural text-to-speech (TTS) systems are not capable of synthesizing rising intonation for these sentences due to the loss of semantic information. Though it has become more common to complement the systems with extra language models, their performance in modeling rising intonation is not well studied. In this paper, we propose to complement the Cantonese TTS model with a BERT-based statement/question classifier. We design different training strategies and compare their performance. We conduct our experiments on a Cantonese corpus named CanTTS. Empirical results show that the separate training approach obtains the best generalization performance and feasibility.

📄 PDF Abstract BibTeX arXiv:2208.02189

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis

2024-12-16 · Xiangheng He, Junjie Chen, Zixing Zhang, Björn W. Schuller

Prosody contains rich information beyond the literal meaning of words, which is crucial for the intelligibility of speech. Current models still fall short in phrasing and intonation; they not only miss or misplace breaks…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Syllable based DNN-HMM Cantonese Speech to Text System

2024-02-13 · LREC 2016 5 · Timothy Wong, Claire Li, Sam Lam, Billy Chiu 외

This paper reports our work on building up a Cantonese Speech-to-Text (STT) system with a syllable based acoustic model. This is a part of an effort in building a STT system to aid dyslexic students who have cognitive de…

speech-recognitionSpeech RecognitionSpeech-to-Text

Into-TTS : Intonation Template Based Prosody Control System

2022-04-04 · JIhwan Lee, Joun Yeop Lee, Heejin Choi, Seongkyu Mun 외

Intonations play an important role in delivering the intention of a speaker. However, current end-to-end TTS systems often fail to model proper intonations. To alleviate this problem, we propose a novel, intuitive method…

Language ModelingLanguage Modelling

Word-wise intonation model for cross-language TTS systems

2024-09-30 · Tomilov A. A., Gromova A. Y., Svischev A. N

In this paper we propose a word-wise intonation model for Russian language and show how it can be generalized for other languages. The proposed model is suitable for automatic data markup and its extended application to …

Dynamic Time WarpingProsody Predictiontext-to-speechText to Speech

QI-TTS: Questioning Intonation Control for Emotional Speech Synthesis

2023-03-14 · Haobin Tang, xulong Zhang, Jianzong Wang, Ning Cheng 외

Recent expressive text to speech (TTS) models focus on synthesizing emotional speech, but some fine-grained styles such as intonation are neglected. In this paper, we propose QI-TTS which aims to better transfer and cont…

Emotional Speech SynthesisSentenceSpeech Synthesistext-to-speech+1