paper-with-me

홈 › Papers

Multi-modal Speech Enhancement with Limited Electromyography Channels

2025-01-11 · Fuyuan Feng, Longting Xu, Rohan Kumar Das

Speech enhancement (SE) aims to improve the clarity, intelligibility, and quality of speech signals for various speech enabled applications. However, air-conducted (AC) speech is highly susceptible to ambient noise, particularly in low signal-to-noise ratio (SNR) and non-stationary noise environments. Incorporating multi-modal information has shown promise in enhancing speech in such challenging scenarios. Electromyography (EMG) signals, which capture muscle activity during speech production, offer noise-resistant properties beneficial for SE in adverse conditions. Most previous EMG-based SE methods required 35 EMG channels, limiting their practicality. To address this, we propose a novel method that considers only 8-channel EMG signals with acoustic signals using a modified SEMamba network with added cross-modality modules. Our experiments demonstrate substantial improvements in speech quality and intelligibility over traditional approaches, especially in extremely low SNR settings. Notably, compared to the SE (AC) approach, our method achieves a significant PESQ gain of 0.235 under matched low SNR conditions and 0.527 under mismatched conditions, highlighting its robustness.

📄 PDF Abstract BibTeX arXiv:2501.06530

Code (0)

등록된 구현이 없습니다.

Tasks

Electromyography (EMG)Speech Enhancement

Similar Papers 제목 키워드 기반

EMGSE: Acoustic/EMG Fusion for Multimodal Speech Enhancement

2022-02-14 · Kuan-Chen Wang, Kai-Chun Liu, Hsin-Min Wang, Yu Tsao

Multimodal learning has been proven to be an effective method to improve speech enhancement (SE) performance, especially in challenging situations such as low signal-to-noise ratios, speech noise, or unseen noise types. …

Electromyography (EMG)Speech Enhancement

Deep Speech Synthesis from Multimodal Articulatory Representations

2024-12-17 · Peter Wu, Bohan Yu, Kevin Scheck, Alan W Black 외

The amount of articulatory data available for training deep learning models is much less compared to acoustic speech data. In order to improve articulatory-to-acoustic synthesis performance in these low-resource settings…

Speech SynthesisTransfer Learning

Bone-conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models

2026-01-18 · Sina Khanagha, Bunlong Lay, Timo Gerkmann arxiv

Single-channel speech enhancement models face significant performance degradation in extremely noisy environments. While prior work has shown that complementary bone-conducted speech can guide enhancement, effective inte…

Speech Enhancement

A Study of Incorporating Articulatory Movement Information in Speech Enhancement

2020-11-03 · Yu-Wen Chen, Kuo-Hsuan Hung, Shang-Yi Chuang, Jonathan Sherman 외

Although deep learning algorithms are widely used for improving speech enhancement (SE) performance, the performance remains limited under highly challenging conditions, such as unseen noise or noise signals having low s…

Speech Enhancement

Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading

2026-06-08 · Eder del Blanco, David Gimeno-Gómez, Eva Navas, Carlos-D. Martínez-Hinarejos 외 arxiv

Speech restoration through silent speech interfaces (SSIs) has emerged as a promising assistive technology for individuals with impaired or absent laryngeal voice production. Among non-invasive SSI modalities, surface el…

Speech Synthesis