paper-with-me

홈 › Papers

VocSegMRI: Multimodal Learning for Precise Vocal Tract Segmentation in Real-time MRI

2025-09-17 · Daiqi Liu, Johannes Enk, Maureen Stone, Fangxu Xing, Tomás Arias-Vergara, Jerry L. Prince, Jana Hutter, Jonghye Woo, Andreas Maier, Paula Andrea Pérez-Toro arxiv

Accurate segmentation of articulatory structures in real-time MRI (rtMRI) remains challenging, as existing methods rely primarily on visual cues and overlook complementary information from synchronized speech signals. We propose VocSegMRI, a multimodal framework integrating video, audio, and phonological inputs via cross-attention fusion and a contrastive learning objective that improves cross-modal alignment and segmentation precision. Evaluated on USC-75 and further validated via zero-shot transfer on USC-TIMIT, VocSegMRI outperforms unimodal and multimodal baselines, with ablations confirming the contribution of each component.

📄 PDF Abstract BibTeX arXiv:2509.13767

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learning

Similar Papers 제목 키워드 기반

Multimodal Segmentation for Vocal Tract Modeling

2024-06-22 · Rishi Jain, Bohan Yu, Peter Wu, Tejas Prabhune 외

Accurate modeling of the vocal tract is necessary to construct articulatory representations for interpretable speech processing and linguistics. However, vocal tract modeling is challenging because many internal articula…

SegmentationVideo SegmentationVideo Semantic Segmentation

Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI

2026-05-18 · Daiqi Liu, Lukas Mulzer, Md Hasan, Nyvenn de Castro 외 arxiv

Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid motion, and limited spatial resolution. However, while rtMRI acquisiti…

Image Segmentation

Multimodal Laryngoscopic Video Analysis for Assisted Diagnosis of Vocal Fold Paralysis

2024-09-05 · Yucong Zhang, Xin Zou, Jinshan Yang, Wenjun Chen 외

This paper presents the Multimodal Laryngoscopic Video Analyzing System (MLVAS), a novel system that leverages both audio and video data to automatically extract key video segments and metrics from raw laryngeal videostr…

Keyword SpottingSegmentation

Open-Source Manually Annotated Vocal Tract Database for Automatic Segmentation from 3D MRI Using Deep Learning: Benchmarking 2D and 3D Convolutional and Transformer Networks

2025-01-08 · Subin Erattakulangara, Karthika Kelat, Katie Burnham, Rachel Balbi 외

Accurate segmentation of the vocal tract from magnetic resonance imaging (MRI) data is essential for various voice and speech applications. Manual segmentation is time intensive and susceptible to errors. This study aime…

BenchmarkingDeep LearningSegmentation

DEEPBEAS3D: Deep Learning and B-Spline Explicit Active Surfaces

2023-09-05 · Helena Williams, João Pedrosa, Muhammad Asad, Laura Cattani 외

Deep learning-based automatic segmentation methods have become state-of-the-art. However, they are often not robust enough for direct clinical application, as domain shifts between training and testing data affect their …

Deep LearningInteractive SegmentationSegmentation