paper-with-me

Papers

InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation

2024-05-24 · Yuchi Wang, Junliang Guo, Jianhong Bai, Runyi Yu, Tianyu He, Xu Tan, Xu sun, Jiang Bian

Recent talking avatar generation models have made strides in achieving realistic and accurate lip synchronization with the audio, but often fall short in controlling and conveying detailed expressions and emotions of the avatar, making the generated video less vivid and controllable. In this paper, we propose a novel text-guided approach for generating emotionally expressive 2D avatars, offering fine-grained control, improved interactivity, and generalizability to the resulting video. Our framework, named InstructAvatar, leverages a natural language interface to control the emotion as well as the facial motion of avatars. Technically, we design an automatic annotation pipeline to construct an instruction-video paired training dataset, equipped with a novel two-branch diffusion-based generator to predict avatars with audio and text instructions at the same time. Experimental results demonstrate that InstructAvatar produces results that align well with both conditions, and outperforms existing methods in fine-grained emotion control, lip-sync quality, and naturalness. Our project page is https://wangyuchi369.github.io/InstructAvatar/.

📄 PDF Abstract BibTeX arXiv:2405.15758

Code (1)

wangyuchi369/InstructAvatar-Project 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

EmoDiff: Intensity Controllable Emotional Text-to-Speech with Soft-Label Guidance

2022-11-17 · Yiwei Guo, Chenpeng Du, Xie Chen, Kai Yu

Although current neural text-to-speech (TTS) models are able to generate high-quality speech, intensity controllable emotional TTS is still a challenging task. Most existing methods need external optimizations for intens…

Denoisingtext-to-speechText to Speech

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation

2026-08-01 · Chenggong Hu, Shaoyin Ma, Yi Wang, Li Sun 외 arxiv

Audio-driven emotional talking face generation aims to synthesize realistic videos with expressive facial dynamics. However, existing methods struggle to balance controllability and visual fidelity. Although implicit rep…

Talking Face GenerationContinuous Control

MVPortrait: Text-Guided Motion and Emotion Control for Multi-view Vivid Portrait Animation

2025-01-01 · CVPR 2025 1 · Yukang Lin, Hokit Fung, Jianjin Xu, Zeping Ren 외

Recent portrait animation methods have made significant strides in generating realistic lip synchronization. However, they often lack explicit control over head movements and facial expressions, and cannot produce vi…

Portrait AnimationVideo Generation

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

2026-06-08 · Minghui Wu, Ganjun Liu, Zikun Fang, Ting Meng 외 arxiv

Instruction-based controllable speech synthesis enables users to specify emotions through natural language. However, existing approaches often rely on coarse emotion labels and lack explicit modeling of fine-grained inte…

Speech Synthesis

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models

2025-10-15 · Yizhou Peng, Yukun Ma, Chong Zhang, Yi-Wen Chao 외 arxiv

While Text-to-Speech (TTS) systems enable emotional control via natural-language instructions, expressiveness, naturalness, and speech quality degrade when the target emotion conflicts with the textual semantics. We prop…