paper-with-me

홈 › Papers

EMA2S: An End-to-End Multimodal Articulatory-to-Speech System

2021-02-07 · Yu-Wen Chen, Kuo-Hsuan Hung, Shang-Yi Chuang, Jonathan Sherman, Wen-Chin Huang, Xugang Lu, Yu Tsao

Synthesized speech from articulatory movements can have real-world use for patients with vocal cord disorders, situations requiring silent speech, or in high-noise environments. In this work, we present EMA2S, an end-to-end multimodal articulatory-to-speech system that directly converts articulatory movements to speech signals. We use a neural-network-based vocoder combined with multimodal joint-training, incorporating spectrogram, mel-spectrogram, and deep features. The experimental results confirm that the multimodal approach of EMA2S outperforms the baseline system in terms of both objective evaluation and subjective evaluation metrics. Moreover, results demonstrate that joint mel-spectrogram and deep feature loss training can effectively improve system performance.

📄 PDF Abstract BibTeX arXiv:2102.03786

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Deep Speech Synthesis from Multimodal Articulatory Representations

2024-12-17 · Peter Wu, Bohan Yu, Kevin Scheck, Alan W Black 외

The amount of articulatory data available for training deep learning models is much less compared to acoustic speech data. In order to improve articulatory-to-acoustic synthesis performance in these low-resource settings…

Speech SynthesisTransfer Learning

Audio-Vision Contrastive Learning for Phonological Class Recognition

2025-07-23 · Daiqi Liu, Tomás Arias-Vergara, Jana Hutter, Andreas Maier 외 arxiv

Accurate classification of articulatory-phonological features plays a vital role in understanding human speech production and developing robust speech technologies, particularly in clinical contexts where targeted phonem…

Multimodal Deep LearningRepresentation LearningContrastive Learning

A Study of Incorporating Articulatory Movement Information in Speech Enhancement

2020-11-03 · Yu-Wen Chen, Kuo-Hsuan Hung, Shang-Yi Chuang, Jonathan Sherman 외

Although deep learning algorithms are widely used for improving speech enhancement (SE) performance, the performance remains limited under highly challenging conditions, such as unseen noise or noise signals having low s…

Speech Enhancement

Coding Speech through Vocal Tract Kinematics

2024-06-18 · Cheol Jun Cho, Peter Wu, Tejas S. Prabhune, Dhruv Agarwal 외

Vocal tract articulation is a natural, grounded control space of speech production. The spatiotemporal coordination of articulators combined with the vocal source shapes intelligible speech sounds to enable effective spo…

Voice Conversion

Acoustic-to-Articulatory Speech Inversion Features for Mispronunciation Detection of /r/ in Child Speech Sound Disorders

2023-05-25 · Nina R Benway, Yashish M Siriwardena, Jonathan L Preston, Elaine Hitchcock 외

Acoustic-to-articulatory speech inversion could enhance automated clinical mispronunciation detection to provide detailed articulatory feedback unattainable by formant-based mispronunciation detection algorithms; however…