paper-with-me

홈 › Papers

MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation

2025-01-03 · CVPR 2025 1 · Huaize Liu, Wenzhang Sun, Donglin Di, Shibo Sun, Jiahui Yang, Changqing Zou, Hujun Bao

The generation of talking avatars has achieved significant advancements in precise audio synchronization. However, crafting lifelike talking head videos requires capturing a broad spectrum of emotions and subtle facial expressions. Current methods face fundamental challenges: a) the absence of frameworks for modeling single basic emotional expressions, which restricts the generation of complex emotions such as compound emotions; b) the lack of comprehensive datasets rich in human emotional expressions, which limits the potential of models. To address these challenges, we propose the following innovations: 1) the Mixture of Emotion Experts (MoEE) model, which decouples six fundamental emotions to enable the precise synthesis of both singular and compound emotional states; 2) the DH-FaceEmoVid-150 dataset, specifically curated to include six prevalent human emotional expressions as well as four types of compound emotions, thereby expanding the training potential of emotion-driven models. Furthermore, to enhance the flexibility of emotion control, we propose an emotion-to-latents module that leverages multimodal inputs, aligning diverse control signals-such as audio, text, and labels-to ensure more varied control inputs as well as the ability to control emotions using audio alone. Through extensive quantitative and qualitative evaluations, we demonstrate that the MoEE framework, in conjunction with the DH-FaceEmoVid-150 dataset, excels in generating complex emotional expressions and nuanced facial details, setting a new benchmark in the field. These datasets will be publicly released.

📄 PDF Abstract BibTeX arXiv:2501.01808

Code (0)

등록된 구현이 없습니다.

Tasks

Portrait Animation

Similar Papers 제목 키워드 기반

MoEEdit: Efficient and Routing-Stable Knowledge Editing for Mixture-of-Experts LLMs

2026-02-11 · Yupu Gu, Rongzhe Wei, Andy Zhu, Pan Li arxiv

Knowledge editing (KE) enables precise modifications to factual content in large language models (LLMs). Existing KE methods are largely designed for dense architectures, limiting their applicability to the increasingly …

knowledge editing

Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free

2024-10-14 · Ziyue Li, Tianyi Zhou

While large language models (LLMs) excel on generation tasks, their decoder-only architecture often limits their potential as embedding models if no further representation finetuning is applied. Does this contradict thei…

Mixture-of-Experts

Unimodal-driven Distillation in Multimodal Emotion Recognition with Dynamic Fusion

2025-03-31 · Jiagen Li, Rui Yu, Huihao Huang, Huaicheng Yan

Multimodal Emotion Recognition in Conversations (MERC) identifies emotional states across text, audio and video, which is essential for intelligent dialogue systems and opinion analysis. Existing methods emphasize hetero…

Emotion RecognitionKnowledge DistillationMixture-of-ExpertsMultimodal Emotion Recognition

Hypertext Entity Extraction in Webpage

2024-03-04 · Yifei Yang, Tianqiao Liu, Bo Shao, Hai Zhao 외

Webpage entity extraction is a fundamental natural language processing task in both research and applications. Nowadays, the majority of webpage entity extraction models are trained on structured datasets which strive to…

Mixture-of-Experts

Talk2Event: Grounded Understanding of Dynamic Scenes from Event Cameras

2025-07-23 · Lingdong Kong, Dongyue Lu, Ao Liang, Rong Li 외 arxiv

Event cameras offer microsecond-level latency and robustness to motion blur, making them ideal for understanding dynamic environments. Yet, connecting these asynchronous streams to human language remains an open challeng…

Relational Reasoning