paper-with-me

Papers

PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion

2024-03-03 · Tianhua Qi, Wenming Zheng, Cheng Lu, Yuan Zong, Hailun Lian

In this paper, we propose Prosody-aware VITS (PAVITS) for emotional voice conversion (EVC), aiming to achieve two major objectives of EVC: high content naturalness and high emotional naturalness, which are crucial for meeting the demands of human perception. To improve the content naturalness of converted audio, we have developed an end-to-end EVC architecture inspired by the high audio quality of VITS. By seamlessly integrating an acoustic converter and vocoder, we effectively address the common issue of mismatch between emotional prosody training and run-time conversion that is prevalent in existing EVC models. To further enhance the emotional naturalness, we introduce an emotion descriptor to model the subtle prosody variations of different speech emotions. Additionally, we propose a prosody predictor, which predicts prosody features from text based on the provided emotion label. Notably, we introduce a prosody alignment loss to establish a connection between latent prosody features from two distinct modalities, ensuring effective training. Experimental results show that the performance of PAVITS is superior to the state-of-the-art EVC methods. Speech Samples are available at https://jeremychee4.github.io/pavits4EVC/ .

📄 PDF Abstract BibTeX arXiv:2403.01494

Code (0)

등록된 구현이 없습니다.

Tasks

Voice Conversion

Similar Papers 제목 키워드 기반

Exploring VQ-VAE with Prosody Parameters for Speaker Anonymization

2024-09-24 · Sotheara Leang, Anderson Augusma, Eric Castelli, Frédérique Letué 외

Human speech conveys prosody, linguistic content, and speaker identity. This article investigates a novel speaker anonymization approach using an end-to-end network based on a Vector-Quantized Variational Auto-Encoder (V…

DecoderSpeaker anonymizationSpeaker Identification

Exploring speech style spaces with language models: Emotional TTS without emotion labels

2024-05-18 · Shreeram Suresh Chandra, Zongyang Du, Berrak Sisman

Many frameworks for emotional text-to-speech (E-TTS) rely on human-annotated emotion labels that are often inaccurate and difficult to obtain. Learning emotional prosody implicitly presents a tough challenge due to the s…

text-to-speechText to SpeechTransfer Learning

HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech

2025-09-25 · Aurosweta Mahapatra, Ismail Rasim Ulgen, Berrak Sisman arxiv

Current anti-spoofing systems remain vulnerable to expressive and emotional synthetic speech, since they rarely leverage prosody as a discriminative cue. Prosody is central to human expressiveness and emotion, and humans…

Self-Supervised LearningMulti-Task LearningSpoof Detection

Reading the Mood Behind Words: Integrating Prosody-Derived Emotional Context into Socially Responsive VR Agents

2026-03-10 · SangYeop Jeong, Yeongseo Na, Seung Gyu Jeong, Jin-Woo Jeong 외 arxiv

In VR interactions with embodied conversational agents, users' emotional intent is often conveyed more by how something is said than by what is said. However, most VR agent pipelines rely on speech-to-text processing, di…

Speech Emotion Recognition

Period VITS: Variational Inference with Explicit Pitch Modeling for End-to-end Emotional Speech Synthesis

2022-10-28 · Yuma Shirahata, Ryuichi Yamamoto, Eunwoo Song, Ryo Terashima 외

Several fully end-to-end text-to-speech (TTS) models have been proposed that have shown better performance compared to cascade models (i.e., training acoustic and vocoder models separately). However, they often generate …

DecoderDiversityEmotional Speech SynthesisSpeech Synthesis+3