paper-with-me

Papers

MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors

2026-05-29 · Guangyin Bao, Taiping Zeng, Jianfeng Feng, Xiangyang Xue arxiv

Reconstructing continuous speech from non-invasive neural recordings is a fundamental problem for probing human auditory perception and building safe, scalable speech brain-computer interfaces. Despite recent progress, intelligible reconstruction remains elusive, as non-invasive recordings are inherently noisy, spatially blurred, and only partially preserve information about perceived speech. Existing methods directly map neural activity to entangled speech representations before synthesizing waveforms with neural vocoders, resulting in spectral-similar but unintelligible results. To overcome these limitations, we introduce MindVoice, a neuro-to-speech reconstruction framework that uses pretrained models to compensate for the incomplete semantic and acoustic information in neural recordings. MindVoice disentangles reconstruction into two complementary pathways: one recovers high-level semantic content, while the other estimates fine-grained acoustic attributes. These inferred representations are then fused with powerful speech generation models and in-context voice cloning to synthesize natural and intelligible utterances. Extensive experiments on EEG and MEG demonstrate that MindVoice substantially outperforms existing methods on various metrics. These results show that pretrained priors provide a principled way to bridge the gap between noisy neural recordings and natural speech, highlighting a promising attempt for auditory neuroscience research and non-invasive speech brain-computer interfaces.

📄 PDF Abstract BibTeX arXiv:2605.31173

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Voice Reconstruction from EEG during Imagined Speech

2023-01-02 · Young-Eun Lee, Seo-Hyun Lee, Sang-Ho Kim, Seong-Whan Lee

Translating imagined speech from human brain activity into voice is a challenging and absorbing research issue that can provide new means of human communication via brain signals. Endeavors toward reconstructing speech f…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderEEG+4

Lip2AudSpec: Speech reconstruction from silent lip movements video

2017-10-26 · Hassan Akbari, Himani Arora, Liangliang Cao, Nima Mesgarani

In this study, we propose a deep neural network for reconstructing intelligible speech from silent lip movement videos. We use auditory spectrogram as spectral representation of speech and its corresponding sound generat…

Lip Reading

SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG

2025-01-03 · Cunhang Fan, Sheng Zhang, Jingjing Zhang, Zexu Pan 외

Decoding speech from brain signals is a challenging research problem that holds significant importance for studying speech processing in the brain. Although breakthroughs have been made in reconstructing the mel spectrog…

EEGMamba

Reconstructing Unseen Sentences from Speech-related Biosignals for Open-vocabulary Neural Communication

2025-10-31 · Deok-Seon Kim, Seo-Hyun Lee, Kang Yin, Seong-Whan Lee arxiv

Brain-to-speech (BTS) systems represent a groundbreaking approach to human communication by enabling the direct transformation of neural activity into linguistic expressions. While recent non-invasive BTS studies have la…

Speech SynthesisEeg Decoding

Improved Speech Reconstruction from Silent Video

2017-08-01 · Ariel Ephrat, Tavi Halperin, Shmuel Peleg

Speechreading is the task of inferring phonetic information from visually observed articulatory facial movements, and is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model ba…