paper-with-me

Papers

Automatic Prosody Prediction for Chinese Speech Synthesis using BLSTM-RNN and Embedding Features

2015-11-02 · Chuang Ding, Lei Xie, Jie Yan, Weini Zhang, Yang Liu

Prosody affects the naturalness and intelligibility of speech. However, automatic prosody prediction from text for Chinese speech synthesis is still a great challenge and the traditional conditional random fields (CRF) based method always heavily relies on feature engineering. In this paper, we propose to use neural networks to predict prosodic boundary labels directly from Chinese characters without any feature engineering. Experimental results show that stacking feed-forward and bidirectional long short-term memory (BLSTM) recurrent network layers achieves superior performance over the CRF-based method. The embedding features learned from raw text further enhance the performance.

📄 PDF Abstract BibTeX arXiv:1511.00360

Code (0)

등록된 구현이 없습니다.

Tasks

Feature EngineeringProsody PredictionSpeech Synthesis

Similar Papers 제목 키워드 기반

Improving Prosody for Unseen Texts in Speech Synthesis by Utilizing Linguistic Information and Noisy Data

2021-11-15 · Zhu Li, Yuqing Zhang, Mengxi Nie, Ming Yan 외

Recent advancements in end-to-end speech synthesis have made it possible to generate highly natural speech. However, training these models typically requires a large amount of high-fidelity speech data, and for unseen te…

Chinese Word SegmentationMulti-Task LearningPart-Of-Speech TaggingPolyphone disambiguation+3

GraphPB: Graphical Representations of Prosody Boundary in Speech Synthesis

2020-12-03 · Aolan Sun, Jianzong Wang, Ning Cheng, Huayi Peng 외

This paper introduces a graphical representation approach of prosody boundary (GraphPB) in the task of Chinese speech synthesis, intending to parse the semantic and syntactic relationship of input sequences in a graphica…

DecoderGraph EmbeddingGraph Neural NetworkGraph-to-Sequence+4

Ensemble prosody prediction for expressive speech synthesis

2023-04-03 · Tian Huey Teh, Vivian Hu, Devang S Ram Mohan, Zack Hodari 외

Generating expressive speech with rich and varied prosody continues to be a challenge for Text-to-Speech. Most efforts have focused on sophisticated neural architectures intended to better model the data distribution. Ye…

DiversityEnsemble LearningExpressive Speech SynthesisPrediction+4

Voice Conversion by Cascading Automatic Speech Recognition and Text-to-Speech Synthesis with Prosody Transfer

2020-09-03 · Jing-Xuan Zhang, Li-Juan Liu, Yan-Nian Chen, Ya-Jun Hu 외

With the development of automatic speech recognition (ASR) and text-to-speech synthesis (TTS) technique, it's intuitive to construct a voice conversion system by cascading an ASR and TTS system. In this paper, we present…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+5

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

2025-05-23 · Zhihao Du, Changfeng Gao, Yuxuan Wang, Fan Yu 외

In our prior works, we introduced a scalable streaming speech synthesis model, CosyVoice 2, which integrates a large language model (LLM) and a chunk-aware flow matching (FM) model, and achieves low-latency bi-streaming …

Automatic Speech RecognitionEmotion RecognitionEvent DetectionLanguage Identification+5