paper-with-me

홈 › Papers

Efficient Incremental Text-to-Speech on GPUs

2022-11-25 · Muyang Du, Chuan Liu, Jiaxing Qi, Junjie Lai

Incremental text-to-speech, also known as streaming TTS, has been increasingly applied to online speech applications that require ultra-low response latency to provide an optimal user experience. However, most of the existing speech synthesis pipelines deployed on GPU are still non-incremental, which uncovers limitations in high-concurrency scenarios, especially when the pipeline is built with end-to-end neural network models. To address this issue, we present a highly efficient approach to perform real-time incremental TTS on GPUs with Instant Request Pooling and Module-wise Dynamic Batching. Experimental results demonstrate that the proposed method is capable of producing high-quality speech with a first-chunk latency lower than 80ms under 100 QPS on a single NVIDIA A10 GPU and significantly outperforms the non-incremental twin in both concurrency and latency. Our work reveals the effectiveness of high-performance incremental TTS on GPUs.

📄 PDF Abstract BibTeX arXiv:2211.13939

Code (0)

등록된 구현이 없습니다.

Tasks

GPUSpeech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Incremental FastPitch: Chunk-based High Quality Text to Speech

2024-01-03 · Muyang Du, Chuan Liu, Junjie Lai

Parallel text-to-speech models have been widely applied for real-time speech synthesis, and they offer more controllability and a much faster synthesis process compared with conventional auto-regressive models. Although …

Speech Synthesistext-to-speechText to Speech

Incremental Machine Speech Chain Towards Enabling Listening while Speaking in Real-time

2020-11-04 · Sashi Novitasari, Andros Tjandra, Tomoya Yanagita, Sakriani Sakti 외

Inspired by a human speech chain mechanism, a machine speech chain framework based on deep learning was recently proposed for the semi-supervised development of automatic speech recognition (ASR) and text-to-speech synth…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+4

Low-Latency Incremental Text-to-Speech Synthesis with Distilled Context Prediction Network

2021-09-22 · Takaaki Saeki, Shinnosuke Takamichi, Hiroshi Saruwatari

Incremental text-to-speech (TTS) synthesis generates utterances in small linguistic units for the sake of real-time and low-latency applications. We previously proposed an incremental TTS method that leverages a large pr…

Knowledge DistillationLanguage ModelingLanguage ModellingSpeech Synthesis+3

Simultaneous Speech-to-Speech Translation System with Neural Incremental ASR, MT, and TTS

2020-11-10 · Katsuhito Sudoh, Takatomo Kano, Sashi Novitasari, Tomoya Yanagita 외

This paper presents a newly developed, simultaneous neural speech-to-speech translation system and its evaluation. The system consists of three fully-incremental neural processing modules for automatic speech recognition…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSimultaneous Speech-to-Speech Translation+8

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs

2026-07-23 · Muyang Du, Shuang Yu, Junjie Lai arxiv

Autoregressive text-to-speech models achieve strong naturalness but suffer from slow inference due to sequential token generation, limiting their deployment in production applications that require low latency. IndexTTS-2…

Text-To-Speech Synthesis