Papers Expressive Speech Synthesis
“Expressive Speech Synthesis” 태그가 달린 논문 47편 · 필터 해제
NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech
Current expressive speech synthesis models are constrained by the limited availability of open-source datasets containing diverse nonverbal vocalizations (NVs). In this work, we introduce NonverbalTTS (NVTTS), a 17-hour …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion ClassificationExpressive Speech Synthesis+5Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse…
Expressive Speech SynthesisPrompt LearningSpeech Synthesistext-to-speech+1RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations
We introduce RASMALAI, a large-scale speech dataset with rich text descriptions, designed to advance controllable and expressive text-to-speech (TTS) synthesis for 23 Indian languages and English. It comprises 13,000 hou…
Expressive Speech SynthesisSpeech Synthesistext-to-speechText to SpeechGender Bias in Instruction-Guided Speech Synthesis Models
Recent advancements in controllable expressive speech synthesis, especially in text-to-speech (TTS) models, have allowed for the generation of speech with specific styles guided by textual descriptions, known as style pr…
Expressive Speech SynthesisSpeech Synthesistext-to-speechText to SpeechSpeech Synthesis along Perceptual Voice Quality Dimensions
While expressive speech synthesis or voice conversion systems mainly focus on controlling or manipulating abstract prosodic characteristics of speech, such as emotion or accent, we here address the control of perceptual …
Expressive Speech SynthesisSpeech Synthesistext-to-speechText to Speech+1Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
We release Rasa, the first multilingual expressive TTS dataset for any Indian language, which contains 10 hours of neutral speech and 1-3 hours of expressive speech for each of the 6 Ekman emotions covering 3 languages: …
Expressive Speech SynthesisSpeech SynthesisMSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
We introduce an open source high-quality Mandarin TTS dataset MSceneSpeech (Multiple Scene Speech Dataset), which is intended to provide resources for expressive speech synthesis. MSceneSpeech comprises numerous audio re…
Expressive Speech SynthesisSpeech SynthesisArticulatory Phonetics Informed Controllable Expressive Speech Synthesis
Expressive speech synthesis aims to generate speech that captures a wide range of para-linguistic features, including emotion and articulation, though current research primarily emphasizes emotional aspects over the nuan…
Expressive Speech SynthesisSpeech SynthesisExpressivity and Speech Synthesis
Imbuing machines with the ability to talk has been a longtime pursuit of artificial intelligence (AI) research. From the very beginning, the community has not only aimed to synthesise high-fidelity speech that accurately…
Expressive Speech SynthesisSpeech SynthesisBoosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
This paper aims to build a multi-speaker expressive TTS system, synthesizing a target speaker's speech with multiple styles and emotions. To this end, we propose a novel contrastive learning-based TTS approach to transfe…
Contrastive LearningExpressive Speech SynthesisSpeech SynthesisTowards Spontaneous Style Modeling with Semi-supervised Pre-training for Conversational Text-to-Speech Synthesis
The spontaneous behavior that often occurs in conversations makes speech more human-like compared to reading-style. However, synthesizing spontaneous-style speech is challenging due to the lack of high-quality spontaneou…
Expressive Speech SynthesisSentenceSpeech Synthesistext-to-speech+2DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
Expressive text-to-speech systems have undergone significant advancements owing to prosody modeling, but conventional methods can still be improved. Traditional approaches have relied on the autoregressive method to pred…
DenoisingExpressive Speech SynthesisSpeech Synthesistext-to-speech+1SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer
Expressive speech synthesis models are trained by adding corpora with diverse speakers, various emotions, and different speaking styles to the dataset, in order to control various characteristics of speech and generate t…
Expressive Speech SynthesisLanguage ModellingSpeech Synthesistext-to-speech+1Cross-lingual Prosody Transfer for Expressive Machine Dubbing
Prosody transfer is well-studied in the context of expressive speech synthesis. Cross-lingual prosody transfer, however, is challenging and has been under-explored to date. In this paper, we present a novel solution to l…
Expressive Speech SynthesisSpeech SynthesisEMNS /Imz/ Corpus: An emotive single-speaker dataset for narrative storytelling in games, television and graphic novels
The increasing adoption of text-to-speech technologies has led to a growing demand for natural and emotive voices that adapt to a conversation's context and emotional tone. The Emotive Narrative Storytelling (EMNS) corpu…
Expressive Speech SynthesisSpeech Synthesistext-to-speechText to SpeechEnhancing Suno's Bark Text-to-Speech Model: Addressing Limitations Through Meta's Encodec and Pre-Trained Hubert
Bark, a transformer-based text-to-audio model by Suno, generates highly realistic, multilingual speech as well as other audio, including music, background noise, and simple sound effects. While this model has shown promi…
Audio GenerationExpressive Speech SynthesisSpeech Synthesistext-to-speech+3Ensemble prosody prediction for expressive speech synthesis
Generating expressive speech with rich and varied prosody continues to be a challenge for Text-to-Speech. Most efforts have focused on sophisticated neural architectures intended to better model the data distribution. Ye…
DiversityEnsemble LearningExpressive Speech SynthesisPrediction+4On granularity of prosodic representations in expressive text-to-speech
In expressive speech synthesis it is widely adopted to use latent prosody representations to deal with variability of the data during training. Same text may correspond to various acoustic realizations, which is known as…
Expressive Speech SynthesisSpeech Synthesistext-to-speechText to SpeechMulti-Speaker Expressive Speech Synthesis via Multiple Factors Decoupling
This paper aims to synthesize the target speaker's speech with desired speaking style and emotion by transferring the style and emotion from reference speech recorded by other speakers. We address this challenging proble…
Expressive Speech SynthesisSpeech SynthesisPredicting phoneme-level prosody latents using AR and flow-based Prior Networks for expressive speech synthesis
A large part of the expressive speech synthesis literature focuses on learning prosodic representations of the speech signal which are then modeled by a prior distribution during inference. In this paper, we compare diff…
Expressive Speech SynthesisSpeech Synthesis