paper-with-me

홈 › Papers

On the Use of Self-Supervised Speech Representations in Spontaneous Speech Synthesis

2023-07-11 · Siyang Wang, Gustav Eje Henter, Joakim Gustafson, Éva Székely

Self-supervised learning (SSL) speech representations learned from large amounts of diverse, mixed-quality speech data without transcriptions are gaining ground in many speech technology applications. Prior work has shown that SSL is an effective intermediate representation in two-stage text-to-speech (TTS) for both read and spontaneous speech. However, it is still not clear which SSL and which layer from each SSL model is most suited for spontaneous TTS. We address this shortcoming by extending the scope of comparison for SSL in spontaneous TTS to 6 different SSLs and 3 layers within each SSL. Furthermore, SSL has also shown potential in predicting the mean opinion scores (MOS) of synthesized speech, but this has only been done in read-speech MOS prediction. We extend an SSL-based MOS prediction framework previously developed for scoring read speech synthesis and evaluate its performance on synthesized spontaneous speech. All experiments are conducted twice on two different spontaneous corpora in order to find generalizable trends. Overall, we present comprehensive experimental results on the use of SSL in spontaneous TTS and MOS prediction to further quantify and understand how SSL can be used in spontaneous TTS. Audios samples: https://www.speech.kth.se/tts-demos/sp_ssl_tts

📄 PDF Abstract BibTeX arXiv:2307.05132

Code (0)

등록된 구현이 없습니다.

Tasks

PredictionSelf-Supervised LearningSpeech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

A Comparative Study of Self-Supervised Speech Representations in Read and Spontaneous TTS

2023-03-05 · Siyang Wang, Gustav Eje Henter, Joakim Gustafson, Éva Székely

Recent work has explored using self-supervised learning (SSL) speech representations such as wav2vec2.0 as the representation medium in standard two-stage TTS, in place of conventionally used mel-spectrograms. It is howe…

Self-Supervised Learning

Analyzing the Robustness of Unsupervised Speech Recognition

2021-10-07 · Guan-Ting Lin, Chan-Jan Hsu, Da-Rong Liu, Hung-Yi Lee 외

Unsupervised speech recognition (unsupervised ASR) aims to learn the ASR system with non-parallel speech and text corpus only. Wav2vec-U has shown promising results in unsupervised ASR by self-supervised speech represent…

Generative Adversarial Networkspeech-recognitionSpeech RecognitionUnsupervised Speech Recognition

Towards Spontaneous Style Modeling with Semi-supervised Pre-training for Conversational Text-to-Speech Synthesis

2023-08-31 · Weiqin Li, Shun Lei, Qiaochu Huang, Yixuan Zhou 외

The spontaneous behavior that often occurs in conversations makes speech more human-like compared to reading-style. However, synthesizing spontaneous-style speech is challenging due to the lack of high-quality spontaneou…

Expressive Speech SynthesisSentenceSpeech Synthesistext-to-speech+2

Unveiling Interpretability in Self-Supervised Speech Representations for Parkinson's Diagnosis

2024-12-02 · David Gimeno-Gómez, Catarina Botelho, Anna Pompili, Alberto Abad 외

Recent works in pathological speech analysis have increasingly relied on powerful self-supervised speech representations, leading to promising results. However, the complex, black-box nature of these embeddings and the l…

BEA-Base: A Benchmark for ASR of Spontaneous Hungarian

2022-02-01 · P. Mihajlik, A. Balog, T. E. Gráczi, A. Kohári 외

Hungarian is spoken by 15 million people, still, easily accessible Automatic Speech Recognition (ASR) benchmark datasets - especially for spontaneous speech - have been practically unavailable. In this paper, we introduc…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3