Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
The Lombard effect plays a key role in natural communication, particularly in noisy environments or when addressing hearing-impaired listeners. We present a controllable text-to-speech (TTS) system capable of synthesizing Lombard speech for any speaker without requiring explicit Lombard data during training. Our approach leverages style embeddings learned from a large, prosodically diverse dataset and analyzes their correlation with Lombard attributes using principal component analysis (PCA). By shifting the relevant PCA components, we manipulate the style embeddings and incorporate them into our TTS model to generate speech at desired Lombard levels. Evaluations demonstrate that our method preserves naturalness and speaker identity, enhances intelligibility under noise, and provides fine-grained control over prosody, offering a robust solution for controllable Lombard TTS for any speaker.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech SynthesisSimilar Papers 제목 키워드 기반
Whispered and Lombard Neural Speech Synthesis
It is desirable for a text-to-speech system to take into account the environment where synthetic speech is presented, and provide appropriate context-dependent output to the user. In this paper, we present and compare va…
Speaker VerificationSpeech Synthesistext-to-speechText to SpeechVoice Conversion for Lombard Speaking Style with Implicit and Explicit Acoustic Feature Conditioning
Text-to-Speech (TTS) systems in Lombard speaking style can improve the overall intelligibility of speech, useful for hearing loss and noisy conditions. However, training those models requires a large amount of data and t…
Voice ConversionStyle TransferSpeaking style adaptation in Text-To-Speech synthesis using Sequence-to-sequence models with attention
Currently, there are increasing interests in text-to-speech (TTS) synthesis to use sequence-to-sequence models with attention. These models are end-to-end meaning that they learn both co-articulation and duration propert…
Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis+1Enhancing Speech Intelligibility in Text-To-Speech Synthesis using Speaking Style Conversion
The increased adoption of digital assistants makes text-to-speech (TTS) synthesis systems an indispensable feature of modern mobile devices. It is hence desirable to build a system capable of generating highly intelligib…
Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis+1Expressive Neural Voice Cloning
Voice cloning is the task of learning to synthesize the voice of an unseen speaker from a few samples. While current voice cloning methods achieve promising results in Text-to-Speech (TTS) synthesis for a new voice, thes…
Speech SynthesisStyle Transfertext-to-speechText to Speech+1