paper-with-me

홈 › Papers

F0 Modeling In Hmm-Based Speech Synthesis System Using Deep Belief Network

2015-02-18 · Sankar Mukherjee, Shyamal Kumar Das Mandal

In recent years multilayer perceptrons (MLPs) with many hid- den layers Deep Neural Network (DNN) has performed sur- prisingly well in many speech tasks, i.e. speech recognition, speaker verification, speech synthesis etc. Although in the context of F0 modeling these techniques has not been ex- ploited properly. In this paper, Deep Belief Network (DBN), a class of DNN family has been employed and applied to model the F0 contour of synthesized speech which was generated by HMM-based speech synthesis system. The experiment was done on Bengali language. Several DBN-DNN architectures ranging from four to seven hidden layers and up to 200 hid- den units per hidden layer was presented and evaluated. The results were compared against clustering tree techniques pop- ularly found in statistical parametric speech synthesis. We show that from textual inputs DBN-DNN learns a high level structure which in turn improves F0 contour in terms of ob- jective and subjective tests.

📄 PDF Abstract BibTeX arXiv:1502.05213

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringSpeaker Verificationspeech-recognitionSpeech RecognitionSpeech Synthesis

Methods 이 논문이 사용한 방법론

Deep Belief Network A Deep Belief Network (DBN) is a multi-layer generative graphical model. DBNs have bi-directional connections…

Similar Papers 제목 키워드 기반

Fine-grained Style Modeling, Transfer and Prediction in Text-to-Speech Synthesis via Phone-Level Content-Style Disentanglement

2020-11-08 · Daxin Tan, Tan Lee

This paper presents a novel design of neural network system for fine-grained style modeling, transfer and prediction in expressive text-to-speech (TTS) synthesis. Fine-grained modeling is realized by extracting style emb…

DisentanglementSpeech SynthesisStyle Transfertext-to-speech+2

A Waveform Representation Framework for High-quality Statistical Parametric Speech Synthesis

2015-10-06 · Bo Fan, Siu Wa Lee, Xiaohai Tian, Lei Xie 외

State-of-the-art statistical parametric speech synthesis (SPSS) generally uses a vocoder to represent speech signals and parameterize them into features for subsequent modeling. Magnitude spectrum has been a dominant fea…

Speech SynthesisVocal Bursts Intensity Prediction

The Theory behind Controllable Expressive Speech Synthesis: a Cross-disciplinary Approach

2019-10-14 · Noé Tits, Kevin El Haddad, Thierry Dutoit

As part of the Human-Computer Interaction field, Expressive speech synthesis is a very rich domain as it requires knowledge in areas such as machine learning, signal processing, sociology, psychology. In this Chapter, we…

Expressive Speech SynthesisSociologySpeech Synthesistext-to-speech+2

Laughter Synthesis: Combining Seq2seq modeling with Transfer Learning

2020-08-20 · Noé Tits, Kevin El Haddad, Thierry Dutoit

Despite the growing interest for expressive speech synthesis, synthesis of nonverbal expressions is an under-explored area. In this paper we propose an audio laughter synthesis system based on a sequence-to-sequence TTS …

Expressive Speech SynthesisSpeech SynthesisTransfer Learning

PauseSpeech: Natural Speech Synthesis via Pre-trained Language Model and Pause-based Prosody Modeling

2023-06-13 · Ji-Sang Hwang, Sang-Hoon Lee, Seong-Whan Lee

Although text-to-speech (TTS) systems have significantly improved, most TTS systems still have limitations in synthesizing speech with appropriate phrasing. For natural speech synthesis, it is important to synthesize the…

Language ModelingLanguage ModellingSpeech Synthesistext-to-speech+1