paper-with-me

Papers

Environment Aware Text-to-Speech Synthesis

2021-10-08 · Daxin Tan, Guangyan Zhang, Tan Lee

This study aims at designing an environment-aware text-to-speech (TTS) system that can generate speech to suit specific acoustic environments. It is also motivated by the desire to leverage massive data of speech audio from heterogeneous sources in TTS system development. The key idea is to model the acoustic environment in speech audio as a factor of data variability and incorporate it as a condition in the process of neural network based speech synthesis. Two embedding extractors are trained with two purposely constructed datasets for characterization and disentanglement of speaker and environment factors in speech. A neural network model is trained to generate speech from extracted speaker and environment embeddings. Objective and subjective evaluation results demonstrate that the proposed TTS system is able to effectively disentangle speaker and environment factors and synthesize speech audio that carries designated speaker characteristics and environment attribute. Audio samples are available online for demonstration https://daxintan-cuhk.github.io/Environment-Aware-TTS/ .

📄 PDF Abstract BibTeX arXiv:2110.03887

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeDisentanglementSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

2025-06-11 · Neta Glazer, Aviv Navon, Yael Segal, Aviv Shamsian 외

Recent advances in Text-to-Speech (TTS) have enabled highly natural speech synthesis, yet integrating speech with complex background environments remains challenging. We introduce UmbraTTS, a flow-matching based TTS mode…

Speech Synthesistext-to-speechText to Speech

Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis

2024-12-22 · Ye-Xin Lu, Hui-Peng Du, Zheng-Yan Sheng, Yang Ai 외

This paper proposes an Incremental Disentanglement-based Environment-Aware zero-shot text-to-speech (TTS) method, dubbed IDEA-TTS, that can synthesize speech for unseen speakers while preserving the acoustic characterist…

DecoderDisentanglementSpeech Synthesistext-to-speech+2

Text-aware and Context-aware Expressive Audiobook Speech Synthesis

2024-06-09 · Dake Guo, Xinfa Zhu, Liumeng Xue, Yongmao Zhang 외

Recent advances in text-to-speech have significantly improved the expressiveness of synthetic speech. However, a major challenge remains in generating speech that captures the diverse styles exhibited by professional nar…

Contrastive LearningLanguage ModelingLanguage ModellingSentence+3

VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis

2024-12-26 · Jaemin Jung, Junseok Ahn, Chaeyoung Jung, Tan Dat Nguyen 외

We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignm…

Audio GenerationSpeech Synthesis

AMAVA: Adaptive Motion-Aware Video-to-Audio Framework for Visually-Impaired Assistance

2026-04-26 · Benjamin Klein, Kazi Ruslan Rahman, Sanchita Ghose arxiv

Navigational aids for blind and low vision individuals struggle conveying dynamic real-world environments, leading to cognitive overload from continuous, undifferentiated feedback. We present AMAVA, a novel real-time vid…