paper-with-me

홈 › Papers

Articulation GAN: Unsupervised modeling of articulatory learning

2022-10-27 · Gašper Beguš, Alan Zhou, Peter Wu, Gopala K Anumanchipalli

Generative deep neural networks are widely used for speech synthesis, but most existing models directly generate waveforms or spectral outputs. Humans, however, produce speech by controlling articulators, which results in the production of speech sounds through physical properties of sound propagation. We introduce the Articulatory Generator to the Generative Adversarial Network paradigm, a new unsupervised generative model of speech production/synthesis. The Articulatory Generator more closely mimics human speech production by learning to generate articulatory representations (electromagnetic articulography or EMA) in a fully unsupervised manner. A separate pre-trained physical model (ema2wav) then transforms the generated EMA representations to speech waveforms, which get sent to the Discriminator for evaluation. Articulatory analysis suggests that the network learns to control articulators in a similar manner to humans during speech production. Acoustic analysis of the outputs suggests that the network learns to generate words that are both present and absent in the training distribution. We additionally discuss implications of articulatory representations for cognitive models of human language and speech technology in general.

📄 PDF Abstract BibTeX arXiv:2210.15173

Code (1)

gbegus/articulationgan 공식 구현 pytorch

Tasks

Generative Adversarial NetworkSpeech Synthesis

Similar Papers 제목 키워드 기반

DYNARTmo: A Dynamic Articulatory Model for Visualization of Speech Movement Patterns

2025-07-27 · Bernd J. Kröger arxiv

We present DYNARTmo, a dynamic articulatory model designed to visualize speech articulation processes in a two-dimensional midsagittal plane. The model builds upon the UK-DYNAMO framework and integrates principles of art…

Style Modeling for Multi-Speaker Articulation-to-Speech

2023-12-21 · Miseul Kim, Zhenyu Piao, Jihyun Lee, Hong-Goo Kang

In this paper, we propose a neural articulation-to-speech (ATS) framework that synthesizes high-quality speech from articulatory signal in a multi-speaker situation. Most conventional ATS approaches only focus on modelin…

Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition

2026-05-20 · Vinicius Ribeiro, Yves Laprie arxiv

Recent advances in machine learning and the availability of articulatory datasets allow vocal tract synthesis to be conditioned on phonetic sequences, a primary task of articulatory speech synthesis. However, quality ass…

Speech Synthesis

Self-supervised speech unit discovery from articulatory and acoustic features using VQ-VAE

2022-06-17 · Marc-Antoine Georges, Jean-Luc Schwartz, Thomas Hueber

The human perception system is often assumed to recruit motor knowledge when processing auditory speech inputs. Using articulatory modeling and deep learning, this study examines how this articulatory information can be …

Why can big.bi be changed to bi.gbi? A mathematical model of syllabification and articulatory synthesis

2023-07-05 · Frédéric Berthommier

A simplified model of articulatory synthesis involving four stages is presented. The planning of articulatory gestures is based on syllable graphs with arcs and nodes that are implemented in a complex representation. Thi…

Trajectory Planning