paper-with-me

홈 › Papers

Style Modeling for Multi-Speaker Articulation-to-Speech

2023-12-21 · Miseul Kim, Zhenyu Piao, Jihyun Lee, Hong-Goo Kang

In this paper, we propose a neural articulation-to-speech (ATS) framework that synthesizes high-quality speech from articulatory signal in a multi-speaker situation. Most conventional ATS approaches only focus on modeling contextual information of speech from a single speaker's articulatory features. To explicitly represent each speaker's speaking style as well as the contextual information, our proposed model estimates style embeddings, guided from the essential speech style attributes such as pitch and energy. We adopt convolutional layers and transformer-based attention layers for our model to fully utilize both local and global information of articulatory signals, measured by electromagnetic articulography (EMA). Our model significantly improves the quality of synthesized speech compared to the baseline in terms of objective and subjective measurements in the Haskins dataset.

📄 PDF Abstract BibTeX arXiv:2312.13603

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Creating Personalized Synthetic Voices from Post-Glossectomy Speech with Guided Diffusion Models

2023-05-27 · Yusheng Tian, Guangyan Zhang, Tan Lee

This paper is about developing personalized speech synthesis systems with recordings of mildly impaired speech. In particular, we consider consonant and vowel alterations resulted from partial glossectomy, the surgical r…

Speech SynthesisVoice Conversion

Creating Personalized Synthetic Voices from Articulation Impaired Speech Using Augmented Reconstruction Loss

2024-01-08 · Yusheng Tian, Jingyu Li, Tan Lee

This research is about the creation of personalized synthetic voices for head and neck cancer survivors. It is focused particularly on tongue cancer patients whose speech might exhibit severe articulation impairment. Our…

Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation

2026-04-17 · Federico Nocentini, Kwanggyoon Seo, Qingju Liu, Claudio Ferrari 외 arxiv

Speech-Driven Facial Animation (SDFA) has gained significant attention due to its applications in movies, video games, and virtual reality. However, most existing models are trained on single-language data, limiting thei…

Self-Supervised Learning

Debatts: Zero-Shot Debating Text-to-Speech Synthesis

2024-11-10 · Yiqiao Huang, Yuancheng Wang, Jiaqi Li, Haotian Guo 외

In debating, rebuttal is one of the most critical stages, where a speaker addresses the arguments presented by the opposing side. During this process, the speaker synthesizes their own persuasive articulation given the c…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Improving the quality of neural TTS using long-form content and multi-speaker multi-style modeling

2022-12-20 · Tuomo Raitio, Javier Latorre, Andrea Davis, Tuuli Morrill 외

Neural text-to-speech (TTS) can provide quality close to natural speech if an adequate amount of high-quality speech material is available for training. However, acquiring speech data for TTS training is costly and time-…

Formtext-to-speechText to Speech