paper-with-me

Text-To-Speech Synthesis

6개 벤치마크 · 논문 351편 · 이 태스크의 논문 보기 →

Benchmarks

LJSpeech

결과 16개

20000 utterances

결과 1개

CMUDict 0.7b

결과 1개

HUI speech corpus

결과 1개

Most implemented

Efficient Neural Audio Synthesis

2018-02-23 · 구현 16개

Papers

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs

2026-07-23 · Muyang Du, Shuang Yu, Junjie Lai arxiv

Autoregressive text-to-speech models achieve strong naturalness but suffer from slow inference due to sequential token generation, limiting their deployment in production applications that require low latency. IndexTTS-2…

Text-To-Speech Synthesis

DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information

2026-06-29 · Roland Roller, Vera Czehmann, Derya Erman, Luke Flanagan 외 arxiv

Conversational data collected in domains such as healthcare or social sciences is a valuable resource for research and automated analysis. However, responsible data sharing requires the detection and removal of personall…

Text-To-Speech Synthesis

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

2026-06-24 · Dihia Lanasri, Rebeh Imane Ammar Aouchiche, Abdelkarim Remmide, Fairouz Taki 외 arxiv

Automatic speech and language technologies are still heavily biased toward high-resource languages, limiting their applicability to dialectal and low-resource settings such as Algerian Dialect. This language presents add…

Natural Language UnderstandingText-To-Speech SynthesisIntent ClassificationResponse Generation

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead

2026-06-20 · Muyang Du, Jason Roche, Junjie Lai arxiv

Streaming text-to-speech synthesis in cascaded LLM-TTS systems still faces latency challenges as most TTS models require full context before initiating generation. We present S5-TTS, a streaming variant of T5-TTS that en…

Text-To-Speech Synthesis

Toward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization

2026-05-04 · Tung Vu, Yen Nguyen, Hai Nguyen, Cuong Pham 외 arxiv

Recent advances in voice cloning and text-to-speech synthesis have made partial speech manipulation - where an adversary replaces a few words within an utterance to alter its meaning while preserving the speaker's identi…

Text-To-Speech SynthesisAudio Deepfake DetectionBinary Classification

Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use

2026-04-17 · Ramit Pahwa, Apoorva Beedu, Parivesh Priye, Rutu Gandhi 외 arxiv

Voice assistants increasingly rely on Speech Language Models (SpeechLMs) to interpret spoken queries and execute complex tasks, yet existing benchmarks lack domain breadth, acoustic diversity, and compositional reasoning…

Text-To-Speech Synthesis

전체 351편 보기 →