Text to Speech
2개 벤치마크 · 논문 1,433편 · 이 태스크의 논문 보기 →
Benchmarks
.
^(#$!@#$)(()))******
Most implemented
WaveNet: A Generative Model for Raw Audio
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Tacotron: Towards End-to-End Speech Synthesis
FastSpeech: Fast, Robust and Controllable Text to Speech
Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention
Efficient Neural Audio Synthesis
Papers
Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?
In recent years, several Speech Language Models (SLMs) that represent speech and written text jointly have been presented. The question then emerges about how model-internal mechanisms are similar and different when oper…
Text to SpeechPreferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages
Crowdsourced pairwise evaluation has emerged as a scalable approach for assessing foundation models. However, applying it to Text to Speech(TTS) introduces high variance due to linguistic diversity and multidimensional n…
Text to SpeechLLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models
Despite the advances in neural text to speech (TTS), many Arabic dialectal varieties remain marginally addressed, with most resources concentrated on Modern Spoken Arabic (MSA) and Gulf dialects, leaving Egyptian Arabic …
Synthetic Data GenerationSpeaker DiarizationSpeech SynthesisText to SpeechManchuTTS: Towards High-Quality Manchu Speech Synthesis via Flow Matching and Hierarchical Text Representation
As an endangered language, Manchu presents unique challenges for speech synthesis, including severe data scarcity and strong phonological agglutination. This paper proposes ManchuTTS(Manchu Text to Speech), a novel appro…
Data AugmentationSpeech SynthesisText to SpeechEmotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization
Emotional text-to-speech seeks to convey affect while preserving intelligibility and prosody, yet existing methods rely on coarse labels or proxy classifiers and receive only utterance-level feedback. We introduce Emotio…
Text to SpeechUniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
The ultimate goal of expressive speech-to-speech translation (S2ST) is to accurately translate spoken content while preserving the speaker identity and emotional style. However, progress in this field is largely hindered…
Speech-to-Speech TranslationText to Speech