paper-with-me

Papers

Deep Denoising Auto-encoder for Statistical Speech Synthesis

2015-06-17 · Zhenzhou Wu, Shinji Takaki, Junichi Yamagishi

This paper proposes a deep denoising auto-encoder technique to extract better acoustic features for speech synthesis. The technique allows us to automatically extract low-dimensional features from high dimensional spectral features in a non-linear, data-driven, unsupervised way. We compared the new stochastic feature extractor with conventional mel-cepstral analysis in analysis-by-synthesis and text-to-speech experiments. Our results confirm that the proposed method increases the quality of synthetic speech in both experiments.

📄 PDF Abstract BibTeX arXiv:1506.05268

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingSpeech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

A Deep Learning Approach to Data-driven Parameterizations for Statistical Parametric Speech Synthesis

2014-09-30 · Prasanna Kumar Muthukumar, Alan W. black

Nearly all Statistical Parametric Speech Synthesizers today use Mel Cepstral coefficients as the vocal tract parameterization of the speech signal. Mel Cepstral coefficients were never intended to work in a parametric sp…

DenoisingSpeech Synthesis

DiffMotion: Speech-Driven Gesture Synthesis Using Denoising Diffusion Model

2023-01-24 · Fan Zhang, Naye Ji, Fuxing Gao, Yongping Li

Speech-driven gesture synthesis is a field of growing interest in virtual human creation. However, a critical challenge is the inherent intricate one-to-many mapping between speech and gestures. Previous studies have exp…

Denoising

Statistical Parametric Speech Synthesis Using Bottleneck Representation From Sequence Auto-encoder

2016-06-19 · Sivanand Achanta, KNRK Raju Alluri, Suryakanth V. Gangashetty

In this paper, we describe a statistical parametric speech synthesis approach with unit-level acoustic representation. In conventional deep neural network based speech synthesis, the input text features are repeated for …

Speech Synthesis

ARCHI-TTS: A flow-matching-based Text-to-Speech Model with Self-supervised Semantic Aligner and Accelerated Inference

2026-02-05 · Chunyat Wu, Jiajun Deng, Zhengxi Liu, Zheqi Dai 외 arxiv

Although diffusion-based, non-autoregressive text-to-speech (TTS) systems have demonstrated impressive zero-shot synthesis capabilities, their efficacy is still hindered by two key challenges: the difficulty of text-spee…

Deep Encoder-Decoder Models for Unsupervised Learning of Controllable Speech Synthesis

2018-07-30 · Gustav Eje Henter, Jaime Lorenzo-Trueba, Xin Wang, Junichi Yamagishi

Generating versatile and appropriate synthetic speech requires control over the output expression separate from the spoken text. Important non-textual speech variation is seldom annotated, in which case output control mu…

Acoustic ModellingDecoderEmotional Speech SynthesisSpeech Synthesis+1