Papers Text-To-Speech Synthesis
“Text-To-Speech Synthesis” 태그가 달린 논문 351편 · 필터 해제
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs
Autoregressive text-to-speech models achieve strong naturalness but suffer from slow inference due to sequential token generation, limiting their deployment in production applications that require low latency. IndexTTS-2…
Text-To-Speech SynthesisDialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information
Conversational data collected in domains such as healthcare or social sciences is a valuable resource for research and automated analysis. However, responsible data sharing requires the detection and removal of personall…
Text-To-Speech SynthesisDziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect
Automatic speech and language technologies are still heavily biased toward high-resource languages, limiting their applicability to dialectal and low-resource settings such as Algerian Dialect. This language presents add…
Natural Language UnderstandingText-To-Speech SynthesisIntent ClassificationResponse GenerationStreaming T5-based Text-to-Speech Synthesis with Limited Lookahead
Streaming text-to-speech synthesis in cascaded LLM-TTS systems still faces latency challenges as most TTS models require full context before initiating generation. We present S5-TTS, a streaming variant of T5-TTS that en…
Text-To-Speech SynthesisToward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization
Recent advances in voice cloning and text-to-speech synthesis have made partial speech manipulation - where an adversary replaces a few words within an utterance to alter its meaning while preserving the speaker's identi…
Text-To-Speech SynthesisAudio Deepfake DetectionBinary ClassificationAudio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use
Voice assistants increasingly rely on Speech Language Models (SpeechLMs) to interpret spoken queries and execute complex tasks, yet existing benchmarks lack domain breadth, acoustic diversity, and compositional reasoning…
Text-To-Speech SynthesisSpeech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
SpeechLLMs typically combine ASR-trained encoders with text-based LLM backbones, leading them to inherit written-style output patterns unsuitable for text-to-speech synthesis. This mismatch is particularly pronounced in …
Text-To-Speech SynthesisDSFlow: Dual Supervision and Step-Aware Architecture for One-Step Flow Matching Speech Synthesis
Flow-matching models have enabled high-quality text-to-speech synthesis, but their iterative sampling process during inference incurs substantial computational cost. Although distillation is widely used to reduce the num…
Text-To-Speech SynthesisQuran-MD: A Fine-Grained Multilingual Multimodal Dataset of the Quran
We present Quran MD, a comprehensive multimodal dataset of the Quran that integrates textual, linguistic, and audio dimensions at the verse and word levels. For each verse (ayah), the dataset provides its original Arabic…
Text-To-Speech SynthesisSemantic RetrievalSpeech RecognitionStyle TransferStuttering-Aware Automatic Speech Recognition for Indonesian Language
Automatic speech recognition systems have achieved remarkable performance on fluent speech but continue to degrade significantly when processing stuttered speech, a limitation that is particularly acute for low-resource …
Text-To-Speech SynthesisSpeech RecognitionData AugmentationTransfer LearningSpeech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
Recent advances in text-to-speech (TTS) have enabled models to clone arbitrary unseen speakers and synthesize high-quality, natural-sounding speech. However, evaluation methods lag behind: typical mean opinion score (MOS…
Text-To-Speech SynthesisSpeech RecognitionParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech-language modelling, and low-resource speech processing. We introduce ParsVoice…
Text-To-Speech SynthesisSpeaker IdentificationLanguage ModellingKAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI
Real-time speech-to-speech (S2S) models excel at generating natural, low-latency conversational responses but often lack deep knowledge and semantic understanding. Conversely, cascaded systems combining automatic speech …
Text-To-Speech SynthesisSpeech RecognitionOLaPh: Optimal Language Phonemizer
Phonemization is a critical component in text-to-speech synthesis. Traditional approaches rely on deterministic transformations and lexica, while neural methods offer potential for higher generalization on out-of-vocabul…
Text-To-Speech SynthesisAn AI-Based Shopping Assistant System to Support the Visually Impaired
Shopping plays a significant role in shaping consumer identity and social integration. However, for individuals with visual impairments, navigating in supermarkets and identifying products can be an overwhelming and chal…
Text-To-Speech SynthesisSpeech RecognitionChipChat: Low-Latency Cascaded Conversational Agent in MLX
The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question. While end-to-end approaches promise theoret…
Text-To-Speech SynthesisSpeech RecognitionReal-Time Sign Language Gestures to Speech Transcription using Deep Learning
Communication barriers pose significant challenges for individuals with hearing and speech impairments, often limiting their ability to effectively interact in everyday environments. This project introduces a real-time a…
Text-To-Speech SynthesisMahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis
Current Text-to-Speech models pose a multilingual challenge, where most of the models traditionally focus on English and European languages, thereby hurting the potential to provide access to information to many more peo…
Text-To-Speech SynthesisMulti-interaction TTS toward professional recording reproduction
Voice directors often iteratively refine voice actors' performances by providing feedback to achieve the desired outcome. While this iterative feedback-based refinement process is important in actual recordings, it has b…
Text-To-Speech SynthesisZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
Existing large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-qualit…
DecoderSpeech Synthesistext-to-speechText to Speech+1