paper-with-me

Papers

EMPHASIS: An Emotional Phoneme-based Acoustic Model for Speech Synthesis System

2018-06-26

We present EMPHASIS, an emotional phoneme-based acoustic model for speech synthesis system. EMPHASIS includes a phoneme duration prediction model and an acoustic parameter prediction model. It uses a CBHG-based regression network to model the dependencies between linguistic features and acoustic features. We modify the input and output layer structures of the network to improve the performance. For the linguistic features, we apply a feature grouping strategy to enhance emotional and prosodic features. The acoustic parameters are designed to be suitable for the regression task and waveform reconstruction. EMPHASIS can synthesize speech in real-time and generate expressive interrogative and exclamatory speech with high audio quality. EMPHASIS is designed to be a multi-lingual model and can synthesize Mandarin-English speech for now. In the experiment of emotional speech synthesis, it achieves better subjective results than other real-time speech synthesis systems.

📄 PDF Abstract BibTeX arXiv:1806.09276

Code (0)

등록된 구현이 없습니다.

Tasks

Emotional Speech SynthesisParameter PredictionregressionSpeech Synthesis

Similar Papers 제목 키워드 기반

Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings

2026-05-04 · Vamshi Nallaguntla, Shruti Kshirsagar, Anderson R. Avila arxiv

Recent advances in emotional voice conversion (EVC) have enabled the generation of expressive synthetic speech, raising new concerns in audio deepfake detection. Existing approaches treat speech as a homogeneous signal a…

Audio Deepfake DetectionVoice Conversion

Audiovisual Speech Synthesis using Tacotron2

2020-08-03 · Ahmed Hussen Abdelaziz, Anushree Prasanna Kumar, Chloe Seivwright, Gabriele Fanelli 외

Audiovisual speech synthesis is the problem of synthesizing a talking face while maximizing the coherency of the acoustic and visual speech. In this paper, we propose and compare two audiovisual speech synthesis systems …

Face ModelSentenceSpeech Synthesis

Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition

2026-05-20 · Vinicius Ribeiro, Yves Laprie arxiv

Recent advances in machine learning and the availability of articulatory datasets allow vocal tract synthesis to be conditioned on phonetic sequences, a primary task of articulatory speech synthesis. However, quality ass…

Speech Synthesis

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

2026-06-08 · Minghui Wu, Ganjun Liu, Zikun Fang, Ting Meng 외 arxiv

Instruction-based controllable speech synthesis enables users to specify emotions through natural language. However, existing approaches often rely on coarse emotion labels and lack explicit modeling of fine-grained inte…

Speech Synthesis

SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

2026-05-01 · Priyam Mazumdar, Yurii Halychanskyi, Steven Guo, Mark Hasegawa-Johnson 외 arxiv

Recent advances in speech synthesis have shifted from phoneme representations to direct grapheme modeling. While phonemes address the one-to-many mapping between text and acoustics, they rely on grapheme-to-phoneme (G2P)…

Speech Synthesis