Text-To-Speech Synthesis
6개 벤치마크 · 논문 351편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Tacotron: Towards End-to-End Speech Synthesis
FastSpeech: Fast, Robust and Controllable Text to Speech
Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention
Efficient Neural Audio Synthesis
Papers
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs
Autoregressive text-to-speech models achieve strong naturalness but suffer from slow inference due to sequential token generation, limiting their deployment in production applications that require low latency. IndexTTS-2…
Text-To-Speech SynthesisDialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information
Conversational data collected in domains such as healthcare or social sciences is a valuable resource for research and automated analysis. However, responsible data sharing requires the detection and removal of personall…
Text-To-Speech SynthesisDziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect
Automatic speech and language technologies are still heavily biased toward high-resource languages, limiting their applicability to dialectal and low-resource settings such as Algerian Dialect. This language presents add…
Natural Language UnderstandingText-To-Speech SynthesisIntent ClassificationResponse GenerationStreaming T5-based Text-to-Speech Synthesis with Limited Lookahead
Streaming text-to-speech synthesis in cascaded LLM-TTS systems still faces latency challenges as most TTS models require full context before initiating generation. We present S5-TTS, a streaming variant of T5-TTS that en…
Text-To-Speech SynthesisToward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization
Recent advances in voice cloning and text-to-speech synthesis have made partial speech manipulation - where an adversary replaces a few words within an utterance to alter its meaning while preserving the speaker's identi…
Text-To-Speech SynthesisAudio Deepfake DetectionBinary ClassificationAudio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use
Voice assistants increasingly rely on Speech Language Models (SpeechLMs) to interpret spoken queries and execute complex tasks, yet existing benchmarks lack domain breadth, acoustic diversity, and compositional reasoning…
Text-To-Speech Synthesis