Non-invasive electromyographic speech neuroprosthesis: a geometric perspective
In this article, we present a high-bandwidth egocentric neuromuscular speech interface for translating silently voiced speech articulations into textand audio. Specifically, we collect electromyogram (EMG) signals from multiple articulatorysites on the face and neck as individuals articulate speech in an alaryngeal manner to perform EMG-to-text or EMG-to-audio translation. Such an interface is useful for restoring audible speech in individuals who have lost the ability to speak intelligibly due to laryngectomy, neuromuscular disease, stroke, or trauma-induced damage (e.g., radiotherapy toxicity) to speech articulators. Previous works have focused on training text or speech synthesis models using EMG collected during audible speech articulations or by transferring audio targets from EMG collected during audible articulation to EMG collected during silent articulation. However, such paradigms are not suited for individuals who have already lost the ability to audibly articulate speech. We are the first to present an alignment-free EMG-to-text and EMG-to-audio conversion using only EMG collected during silently articulated speech in an open-sourced manner. On a limited vocabulary corpora, our approach achieves almost 2.4x improvement in word error rate with a model that is 25x smaller by leveraging the inherent geometry of EMG.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech SynthesisSimilar Papers 제목 키워드 기반
Towards an End-to-End Framework for Invasive Brain Signal Decoding with Large Language Models
In this paper, we introduce a groundbreaking end-to-end (E2E) framework for decoding invasive brain signals, marking a significant advancement in the field of speech neuroprosthesis. Our methodology leverages the compreh…
Towards Dynamic Neural Communication and Speech Neuroprosthesis Based on Viseme Decoding
Decoding text, speech, or images from human neural signals holds promising potential both as neuroprosthesis for patients and as innovative communication tools for general users. Although neural signals contain various i…
Face ReconstructionA Conversational Brain-Artificial Intelligence Interface
We introduce Brain-Artificial Intelligence Interfaces (BAIs) as a new class of Brain-Computer Interfaces (BCIs). Unlike conventional BCIs, which rely on intact cognitive capabilities, BAIs leverage the power of artificia…
AI AgentBrain DecodingEEGMoDAl: Self-Supervised Neural Modality Discovery via Decorrelation for Speech Neuroprosthesis
Speech neuroprosthesis systems decode intended speech from neural activity in the absence of audible output, offering a path to restoring communication for individuals with speech-impairing conditions. Current approaches…
emg2qwerty: A Large Dataset with Baselines for Touch Typing using Surface Electromyography
Surface electromyography (sEMG) non-invasively measures signals generated by muscle activity with sufficient sensitivity to detect individual spinal neurons and richness to identify dozens of gestures and their nuances. …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Frictionspeech-recognition+1