paper-with-me

Papers

AutoTTS: End-to-End Text-to-Speech Synthesis through Differentiable Duration Modeling

2022-03-21 · Bac Nguyen, Fabien Cardinaux, Stefan Uhlich

Parallel text-to-speech (TTS) models have recently enabled fast and highly-natural speech synthesis. However, they typically require external alignment models, which are not necessarily optimized for the decoder as they are not jointly trained. In this paper, we propose a differentiable duration method for learning monotonic alignments between input and output sequences. Our method is based on a soft-duration mechanism that optimizes a stochastic process in expectation. Using this differentiable duration method, we introduce AutoTTS, a direct text-to-waveform speech synthesis model. AutoTTS enables high-fidelity speech synthesis through a combination of adversarial training and matching the total ground-truth duration. Experimental results show that our model obtains competitive results while enjoying a much simpler training pipeline. Audio samples are available online.

📄 PDF Abstract BibTeX arXiv:2203.11049

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling

2026-05-08 · Tong Zheng, Haolin Liu, Chengsong Huang, Huiwen Bao 외 arxiv

Test-time scaling (TTS) has become an effective approach for improving large language model performance by allocating additional computation during inference. However, existing TTS strategies are largely hand-crafted: re…

Mathematical Reasoning

DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis

2024-10-14 · Yingahao Aaron Li, Rithesh Kumar, Zeyu Jin

Diffusion models have demonstrated significant potential in speech synthesis tasks, including text-to-speech (TTS) and voice cloning. However, their iterative denoising processes are computationally intensive, and previo…

DenoisingSpeaker VerificationSpeech Synthesistext-to-speech+3

Embedding a Differentiable Mel-cepstral Synthesis Filter to a Neural Speech Synthesis System

2022-11-21 · Takenori Yoshimura, Shinji Takaki, Kazuhiro Nakamura, Keiichiro Oura 외

This paper integrates a classic mel-cepstral synthesis filter into a modern neural speech synthesis system towards end-to-end controllable speech synthesis. Since the mel-cepstral synthesis filter is explicitly embedded …

GPUSpeech Synthesis

StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

2023-06-13 · NeurIPS 2023 11 · Yinghao Aaron Li, Cong Han, Vinay S. Raghavan, Gavin Mischler 외

In this paper, we present StyleTTS 2, a text-to-speech (TTS) model that leverages style diffusion and adversarial training with large speech language models (SLMs) to achieve human-level TTS synthesis. StyleTTS 2 differs…

Speech Synthesistext-to-speechText to Speech

NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing

2025-02-17 · Yifan Liang, Fangkun Liu, Andong Li, XiaoDong Li 외

Recent advancements in visual speech recognition (VSR) have promoted progress in lip-to-speech synthesis, where pre-trained VSR models enhance the intelligibility of synthesized speech by providing valuable semantic info…

Lip to Speech Synthesisspeech-recognitionSpeech RecognitionSpeech Synthesis+3