paper-with-me

홈 › Papers

Learning to Compute the Articulatory Representations of Speech with the MIRRORNET

2022-10-29 · Yashish M. Siriwardena, Carol Espy-Wilson, Shihab Shamma

Most organisms including humans function by coordinating and integrating sensory signals with motor actions to survive and accomplish desired tasks. Learning these complex sensorimotor mappings proceeds simultaneously and often in an unsupervised or semi-supervised fashion. An autoencoder architecture (MirrorNet) inspired by this sensorimotor learning paradigm is explored in this work to control an articulatory synthesizer, with minimal exposure to ground-truth articulatory data. The articulatory synthesizer takes as input a set of six vocal Tract Variables (TVs) and source features (voicing indicators and pitch) and is able to synthesize continuous speech for unseen speakers. We show that the MirrorNet, once initialized (with ~30 mins of articulatory data) and further trained in unsupervised fashion (`learning phase'), can learn meaningful articulatory representations with comparable accuracy to articulatory speech-inversion systems trained in a completely supervised fashion.

📄 PDF Abstract BibTeX arXiv:2210.16454

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Accent Conversion with Articulatory Representations

2024-06-10 · Yashish M. Siriwardena, Nathan Swedlow, Audrey Howard, Evan Gitterman 외

Conversion of non-native accented speech to native (American) English has a wide range of applications such as improving intelligibility of non-native speech. Previous work on this domain has used phonetic posteriograms …

Multi-Task Learning

The Mirrornet : Learning Audio Synthesizer Controls Inspired by Sensorimotor Interaction

2021-10-12 · Yashish M. Siriwardena, Guilhem Marion, Shihab Shamma

Experiments to understand the sensorimotor neural interactions in the human cortical speech system support the existence of a bidirectional flow of interactions between the auditory and motor regions. Their key function …

Autonomous Vehicles

Articulation GAN: Unsupervised modeling of articulatory learning

2022-10-27 · Gašper Beguš, Alan Zhou, Peter Wu, Gopala K Anumanchipalli

Generative deep neural networks are widely used for speech synthesis, but most existing models directly generate waveforms or spectral outputs. Humans, however, produce speech by controlling articulators, which results i…

Generative Adversarial NetworkSpeech Synthesis

Self-supervised speech unit discovery from articulatory and acoustic features using VQ-VAE

2022-06-17 · Marc-Antoine Georges, Jean-Luc Schwartz, Thomas Hueber

The human perception system is often assumed to recruit motor knowledge when processing auditory speech inputs. Using articulatory modeling and deep learning, this study examines how this articulatory information can be …

Probing in the Wild: A Case Study of Self-Supervised Speech Representations on Mandarin Sub-dialects with Unsupervised Articulatory Analysis

2026-06-24 · Shu Shang, Fuliang Weng, Zeqian Hu, Yaqian Zhou arxiv

While self-supervised speech models have achieved strong performance across speech tasks, relatively little is known about how their internal phonetic representations behave under fine-grained dialect variation. Existing…