paper-with-me

Papers

Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech

2024-09-23 · Hong Nguyen, Sean Foley, Kevin Huang, Xuan Shi, Tiantian Feng, Shrikanth Narayanan

Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven method to visually represent articulator motion in Magnetic Resonance Imaging (MRI) videos of the human vocal tract during speech based on arbitrary audio or speech input. We leverage large pre-trained speech models, which are embedded with prior knowledge, to generalize the visual domain to unseen data using a speech-to-video diffusion model. Our findings demonstrate that the visual generation significantly benefits from the pre-trained speech representations. We also observed that evaluating phonemes in isolation is challenging but becomes more straightforward when assessed within the context of spoken words. Limitations of the current results include the presence of unsmooth tongue motion and video distortion when the tongue contacts the palate.

📄 PDF Abstract BibTeX arXiv:2409.15525

Code (1)

hong7cong/span-rtmri 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI

2024-12-25 · Neil Shah, Ayan Kashyap, Shirish Karande, Vineet Gandhi

Previous real-time MRI (rtMRI)-based speech synthesis models depend heavily on noisy ground-truth speech. Applying loss directly over ground truth mel-spectrograms entangles speech content with MRI noise, resulting in po…

DecoderSpeech Synthesis

SIREM: Speech-Informed MRI Reconstruction with Learned Sampling

2026-05-18 · Md Hasan, Nyvenn Castro, Daiqi Liu, Lukas Mulzer 외 arxiv

Real-time magnetic resonance imaging (rtMRI) of speech production enables non-invasive visualization of dynamic vocal-tract motion and is valuable for speech science and clinical assessment. However, rtMRI is fundamental…

MRI Reconstruction

Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis

2026-09-09 · Hong Nguyen, Sean Foley, Christina Hagedorn, Yijing Lu 외 arxiv

Real-time MRI (rtMRI) captures the dynamics of the entire vocal tract during speech, but labeled data are scarce and the modality - single-slice, grayscale, low-resolution - differs substantially from the natural videos …

Domain Adaptation

Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI

2026-05-18 · Daiqi Liu, Lukas Mulzer, Md Hasan, Nyvenn de Castro 외 arxiv

Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid motion, and limited spatial resolution. However, while rtMRI acquisiti…

Image Segmentation

Real-Time MRI Video synthesis from time aligned phonemes with sequence-to-sequence networks

2022-10-30 · Sathvik Udupa, Prasanta Kumar Ghosh

Real-Time Magnetic resonance imaging (rtMRI) of the midsagittal plane of the mouth is of interest for speech production research. In this work, we focus on estimating utterance level rtMRI video from the spoken phoneme s…

Decoder