paper-with-me

홈 › Papers

Emotional Face-to-Speech

2025-02-03 · Jiaxin Ye, Boyuan Cao, Hongming Shan

How much can we infer about an emotional voice solely from an expressive face? This intriguing question holds great potential for applications such as virtual character dubbing and aiding individuals with expressive language disorders. Existing face-to-speech methods offer great promise in capturing identity characteristics but struggle to generate diverse vocal styles with emotional expression. In this paper, we explore a new task, termed emotional face-to-speech, aiming to synthesize emotional speech directly from expressive facial cues. To that end, we introduce DEmoFace, a novel generative framework that leverages a discrete diffusion transformer (DiT) with curriculum learning, built upon a multi-level neural audio codec. Specifically, we propose multimodal DiT blocks to dynamically align text and speech while tailoring vocal styles based on facial emotion and identity. To enhance training efficiency and generation quality, we further introduce a coarse-to-fine curriculum learning algorithm for multi-level token processing. In addition, we develop an enhanced predictor-free guidance to handle diverse conditioning scenarios, enabling multi-conditional generation and disentangling complex attributes effectively. Extensive experimental results demonstrate that DEmoFace generates more natural and consistent speech compared to baselines, even surpassing speech-driven methods. Demos are shown at https://demoface-ai.github.io/.

📄 PDF Abstract BibTeX arXiv:2502.01046

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation

2023-03-20 · ICCV 2023 1 · Ziqiao Peng, HaoYu Wu, Zhenbo Song, Hao Xu 외

Speech-driven 3D face animation aims to generate realistic facial expressions that match the speech content and emotion. However, existing methods often neglect emotional facial expressions or fail to disentangle them fr…

3D Face AnimationDecoderDisentanglement

EmoDiffusion: Enhancing Emotional 3D Facial Animation with Latent Diffusion Models

2025-03-14 · Yixuan Zhang, Qing Chang, Yuxi Wang, Guang Chen 외

Speech-driven 3D facial animation seeks to produce lifelike facial expressions that are synchronized with the speech content and its emotional nuances, finding applications in various multimedia fields. However, previous…

Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning

2026-05-30 · Sukru Samet Dindar, Riki Shimizu, Xilin Jiang, Nima Mesgarani arxiv

Empathetic spoken dialogue systems must infer a user's emotional state to respond appropriately, yet everyday speech often carries weak, neutral, or ambiguous affective cues. To address this, we introduce Sympatheia, a s…

Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions

2024-09-25 · Kun Zhou, You Zhang, Shengkui Zhao, Hao Wang 외

Current emotional text-to-speech systems face challenges in conveying the full spectrum of human emotions, largely due to the inherent complexity of human emotions and the limited range of emotional labels in existing sp…

AttributeDimensionality ReductionDiversityLanguage Modeling+4

CAMEO: Collection of Multilingual Emotional Speech Corpora

2025-05-16 · Iwona Christop, Maciej Czajka

This paper presents CAMEO -- a curated collection of multilingual emotional speech datasets designed to facilitate research in emotion recognition and other speech-related tasks. The main objectives were to ensure easy a…

Emotion RecognitionSpeech Emotion Recognition