paper-with-me

홈 › Papers

BrainTalker: Low-Resource Brain-to-Speech Synthesis with Transfer Learning using Wav2Vec 2.0

2023-12-21 · Miseul Kim, Zhenyu Piao, Jihyun Lee, Hong-Goo Kang

Decoding spoken speech from neural activity in the brain is a fast-emerging research topic, as it could enable communication for people who have difficulties with producing audible speech. For this task, electrocorticography (ECoG) is a common method for recording brain activity with high temporal resolution and high spatial precision. However, due to the risky surgical procedure required for obtaining ECoG recordings, relatively little of this data has been collected, and the amount is insufficient to train a neural network-based Brain-to-Speech (BTS) system. To address this problem, we propose BrainTalker-a novel BTS framework that generates intelligible spoken speech from ECoG signals under extremely low-resource scenarios. We apply a transfer learning approach utilizing a pre-trained self supervised model, Wav2Vec 2.0. Specifically, we train an encoder module to map ECoG signals to latent embeddings that match Wav2Vec 2.0 representations of the corresponding spoken speech. These embeddings are then transformed into mel-spectrograms using stacked convolutional and transformer-based layers, which are fed into a neural vocoder to synthesize speech waveform. Experimental results demonstrate our proposed framework achieves outstanding performance in terms of subjective and objective metrics, including a Pearson correlation coefficient of 0.9 between generated and ground truth mel spectrograms. We share publicly available Demos and Code.

📄 PDF Abstract BibTeX arXiv:2312.13600

Code (0)

등록된 구현이 없습니다.

Tasks

Speech SynthesisTransfer Learning

Similar Papers 제목 키워드 기반

Speech Synthesis for Low Resource Languages using Transliteration Enabled Transfer Learning

2021-11-16 · ACL ARR November 2021 11 · Anonymous

In the area of Human Computer Interaction (HCI), Text To Speech (TTS) synthesis has received a significant boost in recent years, especially with the development of various deep learning techniques capable of generating …

speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speech+3

Exploring Transfer Learning for Urdu Speech Synthesis

2022-06-01 · EURALI (LREC) 2022 6 · Sahar Jamal, Sadaf Abdul Rauf, Quratulain Majid

Neural methods in Text to Speech synthesis (TTS) have demonstrated momentous advancement in terms of the naturalness and intelligibility of the synthesized speech. In this paper we present neural speech synthesis system …

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis+1

Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks

2023-12-10 · Seo-Hyun Lee, Young-Eun Lee, Soowon Kim, Byung-Kwan Ko 외

Brain-to-speech technology represents a fusion of interdisciplinary applications encompassing fields of artificial intelligence, brain-computer interfaces, and speech synthesis. Neural representation learning based inten…

Representation LearningSpeech Synthesis

Cross-lingual Transfer for Speech Processing using Acoustic Language Similarity

2021-11-02 · Peter Wu, Jiatong Shi, Yifan Zhong, Shinji Watanabe 외

Speech processing systems currently do not support the vast majority of languages, in part due to the lack of data in low-resource languages. Cross-lingual transfer offers a compelling way to help bridge this digital div…

Cross-Lingual Transferspeech-recognitionSpeech RecognitionSpeech Synthesis

Byakto Speech: Real-time long speech synthesis with convolutional neural network: Transfer learning from English to Bangla

2021-05-31 · Zabir Al Nazi, Sayed Mohammed Tasmimul Huda

Speech synthesis is one of the challenging tasks to automate by deep learning, also being a low-resource language there are very few attempts at Bangla speech synthesis. Most of the existing works can't work with anythin…

Deep Learningspeech-recognitionSpeech RecognitionSpeech Synthesis+1