paper-with-me

홈 › Papers

Articulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention Fusion

2025-10-01 · Ahmed Adel Attia, Jing Liu, Carol Espy Wilson arxiv

Prior works have investigated the use of articulatory features as complementary representations for automatic speech recognition (ASR), but their use was largely confined to shallow acoustic models. In this work, we revisit articulatory information in the era of deep learning and propose a framework that leverages articulatory representations both as an auxiliary task and as a pseudo-input to the recognition model. Specifically, we employ speech inversion as an auxiliary prediction task, and the predicted articulatory features are injected into the model as a query stream in a cross-attention module with acoustic embeddings as keys and values. Experiments on LibriSpeech demonstrate that our approach yields consistent improvements over strong transformer-based baselines, particularly under low-resource conditions. These findings suggest that articulatory features, once sidelined in ASR research, can provide meaningful benefits when reintroduced with modern architectures.

📄 PDF Abstract BibTeX arXiv:2510.08585

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition

2026-05-20 · Vinicius Ribeiro, Yves Laprie arxiv

Recent advances in machine learning and the availability of articulatory datasets allow vocal tract synthesis to be conditioned on phonetic sequences, a primary task of articulatory speech synthesis. However, quality ass…

Speech Synthesis

Articulatory Phonetics Informed Controllable Expressive Speech Synthesis

2024-06-15 · Zehua Kcriss Li, Meiying Melissa Chen, Yi Zhong, Pinxin Liu 외

Expressive speech synthesis aims to generate speech that captures a wide range of para-linguistic features, including emotion and articulation, though current research primarily emphasizes emotional aspects over the nuan…

Expressive Speech SynthesisSpeech Synthesis

Coding Speech through Vocal Tract Kinematics

2024-06-18 · Cheol Jun Cho, Peter Wu, Tejas S. Prabhune, Dhruv Agarwal 외

Vocal tract articulation is a natural, grounded control space of speech production. The spatiotemporal coordination of articulators combined with the vocal source shapes intelligible speech sounds to enable effective spo…

Voice Conversion

Towards a Quantitative Analysis of Coarticulation with a Phoneme-to-Articulatory Model

2024-08-10 · Chaofei Fan, Jaimie M. Henderson, Chris Manning, Francis R. Willett

Prior coarticulation studies focus mainly on limited phonemic sequences and specific articulators, providing only approximate descriptions of the temporal extent and magnitude of coarticulation. This paper is an initial …

DYNARTmo: A Dynamic Articulatory Model for Visualization of Speech Movement Patterns

2025-07-27 · Bernd J. Kröger arxiv

We present DYNARTmo, a dynamic articulatory model designed to visualize speech articulation processes in a two-dimensional midsagittal plane. The model builds upon the UK-DYNAMO framework and integrates principles of art…