Speech Synthesis
5개 벤치마크 · 논문 1,383편 · 이 태스크의 논문 보기 →
Benchmarks
LibriTTS
North American English
LJSpeech
Mandarin Chinese
Blizzard Challenge 2013
Most implemented
WaveNet: A Generative Model for Raw Audio
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions
Tacotron: Towards End-to-End Speech Synthesis
FastSpeech: Fast, Robust and Controllable Text to Speech
MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
Papers
What Did I Just Say? Self-Listening for Full-Duplex Speech Models
Full-duplex spoken language models can listen and speak simultaneously, enabling them to handle interruptions and backchannels in human conversation. However, text generation, speech synthesis, and audio playback proceed…
Speech SynthesisText GenerationSpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis
Recent advances in AI have revolutionized speech processing, yet effective speech understanding requires discerning not just what is said, but how it is said. Speech Sentiment Analysis plays a critical role in decoding t…
Sentiment AnalysisSpeech RecognitionSpeech SynthesisS2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling
Multimodal dialogue retrieval aims to retrieve dialogues from multimodal dialogue banks that are similar to a target dialogue in terms of both textual semantics and acoustic conversational styles. Such dialogue-level ret…
Emotion Recognition in ConversationContrastive LearningSpeech SynthesisVoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents
Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-spe…
Speech SynthesisCuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents
Zero-shot text-to-speech (TTS) now supports interactive assistants, personalized media, and accessibility tools. All TTS systems require faithful linguistic rendering, consistent speaker identity, and low-latency respons…
Speech SynthesisTeffic-Audio: Tell Fact from Fiction
Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoof…
DeepFake DetectionSpeech SynthesisVoice Conversion