paper-with-me

Papers Text-To-Speech Synthesis

“Text-To-Speech Synthesis” 태그가 달린 논문 351편 · 필터 해제

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs

2026-07-23 · Muyang Du, Shuang Yu, Junjie Lai arxiv

Autoregressive text-to-speech models achieve strong naturalness but suffer from slow inference due to sequential token generation, limiting their deployment in production applications that require low latency. IndexTTS-2…

Text-To-Speech Synthesis

DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information

2026-06-29 · Roland Roller, Vera Czehmann, Derya Erman, Luke Flanagan 외 arxiv

Conversational data collected in domains such as healthcare or social sciences is a valuable resource for research and automated analysis. However, responsible data sharing requires the detection and removal of personall…

Text-To-Speech Synthesis

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

2026-06-24 · Dihia Lanasri, Rebeh Imane Ammar Aouchiche, Abdelkarim Remmide, Fairouz Taki 외 arxiv

Automatic speech and language technologies are still heavily biased toward high-resource languages, limiting their applicability to dialectal and low-resource settings such as Algerian Dialect. This language presents add…

Natural Language UnderstandingText-To-Speech SynthesisIntent ClassificationResponse Generation

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead

2026-06-20 · Muyang Du, Jason Roche, Junjie Lai arxiv

Streaming text-to-speech synthesis in cascaded LLM-TTS systems still faces latency challenges as most TTS models require full context before initiating generation. We present S5-TTS, a streaming variant of T5-TTS that en…

Text-To-Speech Synthesis

Toward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization

2026-05-04 · Tung Vu, Yen Nguyen, Hai Nguyen, Cuong Pham 외 arxiv

Recent advances in voice cloning and text-to-speech synthesis have made partial speech manipulation - where an adversary replaces a few words within an utterance to alter its meaning while preserving the speaker's identi…

Text-To-Speech SynthesisAudio Deepfake DetectionBinary Classification

Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use

2026-04-17 · Ramit Pahwa, Apoorva Beedu, Parivesh Priye, Rutu Gandhi 외 arxiv

Voice assistants increasingly rely on Speech Language Models (SpeechLMs) to interpret spoken queries and execute complex tasks, yet existing benchmarks lack domain breadth, acoustic diversity, and compositional reasoning…

Text-To-Speech Synthesis

Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization

2026-03-13 · Mengjie Zhao, Lianbo Liu, Yusuke Fujita, Hao Shi 외 arxiv

SpeechLLMs typically combine ASR-trained encoders with text-based LLM backbones, leading them to inherit written-style output patterns unsuitable for text-to-speech synthesis. This mismatch is particularly pronounced in …

Text-To-Speech Synthesis

DSFlow: Dual Supervision and Step-Aware Architecture for One-Step Flow Matching Speech Synthesis

2026-02-03 · Bin Lin, Peng Yang, Chao Yan, Xiaochen Liu 외 arxiv

Flow-matching models have enabled high-quality text-to-speech synthesis, but their iterative sampling process during inference incurs substantial computational cost. Although distillation is widely used to reduce the num…

Text-To-Speech Synthesis

Quran-MD: A Fine-Grained Multilingual Multimodal Dataset of the Quran

2026-01-25 · Muhammad Umar Salman, Mohammad Areeb Qazi, Mohammed Talha Alam arxiv

We present Quran MD, a comprehensive multimodal dataset of the Quran that integrates textual, linguistic, and audio dimensions at the verse and word levels. For each verse (ayah), the dataset provides its original Arabic…

Text-To-Speech SynthesisSemantic RetrievalSpeech RecognitionStyle Transfer

Stuttering-Aware Automatic Speech Recognition for Indonesian Language

2026-01-07 · Fadhil Muhammad, Alwin Djuliansah, Adrian Aryaputra Hamzah, Kurniawati Azizah arxiv

Automatic speech recognition systems have achieved remarkable performance on fluent speech but continue to degrade significantly when processing stuttered speech, a limitation that is particularly acute for low-resource …

Text-To-Speech SynthesisSpeech RecognitionData AugmentationTransfer Learning

Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward

2025-11-12 · Guansu Wang, Peijie Sun arxiv

Recent advances in text-to-speech (TTS) have enabled models to clone arbitrary unseen speakers and synthesize high-quality, natural-sounding speech. However, evaluation methods lag behind: typical mean opinion score (MOS…

Text-To-Speech SynthesisSpeech Recognition

ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis

2025-10-12 · Mohammad Javad Ranjbar Kalahroodi, Heshaam Faili, Azadeh Shakery arxiv

Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech-language modelling, and low-resource speech processing. We introduce ParsVoice…

Text-To-Speech SynthesisSpeaker IdentificationLanguage Modelling

KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI

2025-09-26 · So Kuroki, Yotaro Kubo, Takuya Akiba, Yujin Tang arxiv

Real-time speech-to-speech (S2S) models excel at generating natural, low-latency conversational responses but often lack deep knowledge and semantic understanding. Conversely, cascaded systems combining automatic speech …

Text-To-Speech SynthesisSpeech Recognition

OLaPh: Optimal Language Phonemizer

2025-09-24 · Johannes Wirth arxiv

Phonemization is a critical component in text-to-speech synthesis. Traditional approaches rely on deterministic transformations and lexica, while neural methods offer potential for higher generalization on out-of-vocabul…

Text-To-Speech Synthesis

An AI-Based Shopping Assistant System to Support the Visually Impaired

2025-09-01 · Larissa R. de S. Shibata, Ankit A. Ravankar, Jose Victorio Salazar Luces, Yasuhisa Hirata arxiv

Shopping plays a significant role in shaping consumer identity and social integration. However, for individuals with visual impairments, navigating in supermarkets and identifying products can be an overwhelming and chal…

Text-To-Speech SynthesisSpeech Recognition

ChipChat: Low-Latency Cascaded Conversational Agent in MLX

2025-08-26 · Tatiana Likhomanenko, Luke Carlson, Richard He Bai, Zijin Gu 외 arxiv

The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question. While end-to-end approaches promise theoret…

Text-To-Speech SynthesisSpeech Recognition

Real-Time Sign Language Gestures to Speech Transcription using Deep Learning

2025-08-18 · Brandone Fonya, Clarence Worrell arxiv

Communication barriers pose significant challenges for individuals with hearing and speech impairments, often limiting their ability to effectively interact in everyday environments. This project introduces a real-time a…

Text-To-Speech Synthesis

MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis

2025-08-05 · Jaskaran Singh, Amartya Roy Chowdhury, Raghav Prabhakar, Varshul C. W arxiv

Current Text-to-Speech models pose a multilingual challenge, where most of the models traditionally focus on English and European languages, thereby hurting the potential to provide access to information to many more peo…

Text-To-Speech Synthesis

Multi-interaction TTS toward professional recording reproduction

2025-07-01 · Hiroki Kanagawa, Kenichi Fujita, Aya Watanabe, Yusuke Ijima arxiv

Voice directors often iteratively refine voice actors' performances by providing feedback to achieve the desired outcome. While this iterative feedback-based refinement process is important in actual recordings, it has b…

Text-To-Speech Synthesis

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

2025-06-16 · Han Zhu, Wei Kang, Zengwei Yao, Liyong Guo 외

Existing large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-qualit…

DecoderSpeech Synthesistext-to-speechText to Speech+1
1–20 / 351 다음 →