paper-with-me

홈 › Papers

Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models

2024-07-26 · Neil Shah, Shirish Karande, Vineet Gandhi

We propose a novel approach to significantly improve the intelligibility in the Non-Audible Murmur (NAM)-to-speech conversion task, leveraging self-supervision and sequence-to-sequence (Seq2Seq) learning techniques. Unlike conventional methods that explicitly record ground-truth speech, our methodology relies on self-supervision and speech-to-speech synthesis to simulate ground-truth speech. Despite utilizing simulated speech, our method surpasses the current state-of-the-art (SOTA) by 29.08% improvement in the Mel-Cepstral Distortion (MCD) metric. Additionally, we present error rates and demonstrate our model's proficiency to synthesize speech in novel voices of interest. Moreover, we present a methodology for augmenting the existing CSTR NAM TIMIT Plus corpus, setting a benchmark with a Word Error Rate (WER) of 42.57% to gauge the intelligibility of the synthesized speech. Speech samples can be found at https://nam2speech.github.io/NAM2Speech/

📄 PDF Abstract BibTeX arXiv:2407.18541

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Methods 이 논문이 사용한 방법론

NAM Neural Additive Models (NAMs) make restrictions on the structure of neural networks, which yields a family of models that are inherently interpretable while suffering little…

Similar Papers 제목 키워드 기반

Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

2021-04-01 · Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov 외

We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representations for speech content, prosodic inform…

DisentanglementRepresentation LearningResynthesisSpeaker Identification+1

Simple and Effective Unsupervised Speech Synthesis

2022-04-06 · Alexander H. Liu, Cheng-I Jeff Lai, Wei-Ning Hsu, Michael Auli 외

We introduce the first unsupervised speech synthesis system based on a simple, yet effective recipe. The framework leverages recent work in unsupervised speech recognition as well as existing neural-based speech synthesi…

speech-recognitionSpeech RecognitionSpeech SynthesisUnsupervised Speech Recognition

Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks

2023-09-14 · Soumi Maiti, Yifan Peng, Shukjae Choi, Jee-weon Jung 외

We propose a decoder-only language model, VoxtLM, that can perform four tasks: speech recognition, speech synthesis, text generation, and speech continuation. VoxtLM integrates text vocabulary with discrete speech tokens…

DecoderLanguage ModelingLanguage Modellingspeech-recognition+3

ReVISE: Self-Supervised Speech Resynthesis with Visual Input for Universal and Generalized Speech Enhancement

2022-12-21 · Wei-Ning Hsu, Tal Remez, Bowen Shi, Jacob Donley 외

Prior works on improving speech quality with visual input typically study each type of auditory distortion separately (e.g., separation, inpainting, video-to-speech) and present tailored algorithms. This paper proposes t…

Audio-Visual Speech RecognitionResynthesisSpeech Enhancementspeech-recognition+7

LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning

2025-09-30 · Kang Yang, Yifan Liang, Fangkun Liu, Zhenping Xie 외 arxiv

Lip-to-speech (L2S) synthesis for Mandarin is a significant challenge, hindered by complex viseme-to-phoneme mappings and the critical role of lexical tones in intelligibility. To address this issue, we propose Lexical T…

Self-Supervised LearningCross-Lingual TransferSpeech Synthesis