paper-with-me

Papers

Humane Speech Synthesis through Zero-Shot Emotion and Disfluency Generation

2024-03-31 · Rohan Chaudhury, Mihir Godbole, Aakash Garg, Jinsil Hwaryoung Seo

Contemporary conversational systems often present a significant limitation: their responses lack the emotional depth and disfluent characteristic of human interactions. This absence becomes particularly noticeable when users seek more personalized and empathetic interactions. Consequently, this makes them seem mechanical and less relatable to human users. Recognizing this gap, we embarked on a journey to humanize machine communication, to ensure AI systems not only comprehend but also resonate. To address this shortcoming, we have designed an innovative speech synthesis pipeline. Within this framework, a cutting-edge language model introduces both human-like emotion and disfluencies in a zero-shot setting. These intricacies are seamlessly integrated into the generated text by the language model during text generation, allowing the system to mirror human speech patterns better, promoting more intuitive and natural user interactions. These generated elements are then adeptly transformed into corresponding speech patterns and emotive sounds using a rule-based approach during the text-to-speech phase. Based on our experiments, our novel system produces synthesized speech that's almost indistinguishable from genuine human communication, making each interaction feel more personal and authentic.

📄 PDF Abstract BibTeX arXiv:2404.01339

Code (1)

rohan-chaudhury/humane-speech-synthesis-through-zero-shot-emotion-and-disfluency-generation 공식 구현

Tasks

Language ModelingLanguage ModellingSpeech SynthesisText Generationtext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters

2024-01-10 · Kenichi Fujita, Hiroshi Sato, Takanori Ashihara, Hiroki Kanagawa 외

The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce speaker characteristics very accurately. H…

Self-Supervised LearningSpeech EnhancementSpeech Synthesistext-to-speech+2

HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

2023-11-21 · Sang-Hoon Lee, Ha-Yeong Choi, Seung-bin Kim, Seong-Whan Lee

Large language models (LLM)-based speech synthesis has been widely adopted in zero-shot speech synthesis. However, they require a large-scale data and possess the same limitations as previous autoregressive speech models…

Speech SynthesisSuper-Resolutiontext-to-speechText to Speech+2

FlashSpeech: Efficient Zero-Shot Speech Synthesis

2024-04-23 · Zhen Ye, Zeqian Ju, Haohe Liu, Xu Tan 외

Recent progress in large-scale zero-shot speech synthesis has been significantly advanced by language models and diffusion models. However, the generation process of both methods is slow and computationally intensive. Ef…

RhythmSpeech SynthesisVoice Conversion

StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis

2024-09-24 · Zhiyong Chen, Xinnuo Li, Zhiqi Ai, Shugong Xu

We introduce StyleFusion-TTS, a prompt and/or audio referenced, style and speaker-controllable, zero-shot text-to-speech (TTS) synthesis system designed to enhance the editability and naturalness of current research lite…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

HumanEval on Latest GPT Models -- 2024

2024-02-20 · Daniel Li, Lincoln Murr

In 2023, we are using the latest models of GPT-4 to advance program synthesis. The large language models have significantly improved the state-of-the-art for this purpose. To make these advancements more accessible, we h…

Code GenerationHumanEvalLanguage ModelingLanguage Modelling+1