paper-with-me

Papers

SingingBot: An Avatar-Driven System for Robotic Face Singing Performance

2026-01-05 · Zhuoxiong Xu, Xuanchen Li, Yuhao Cheng, Fei Xu, Yichao Yan, Xiaokang Yang arxiv

Equipping robotic faces with singing capabilities is crucial for empathetic Human-Robot Interaction. However, existing robotic face driving research primarily focuses on conversations or mimicking static expressions, struggling to meet the high demands for continuous emotional expression and coherence in singing. To address this, we propose a novel avatar-driven framework for appealing robotic singing. We first leverage portrait video generation models embedded with extensive human priors to synthesize vivid singing avatars, providing reliable expression and emotion guidance. Subsequently, these facial features are transferred to the robot via semantic-oriented mapping functions that span a wide expression space. Furthermore, to quantitatively evaluate the emotional richness of robotic singing, we propose the Emotion Dynamic Range metric to measure the emotional breadth within the Valence-Arousal space, revealing that a broad emotional spectrum is crucial for appealing performances. Comprehensive experiments prove that our method achieves rich emotional expressions while maintaining lip-audio synchronization, significantly outperforming existing approaches.

📄 PDF Abstract BibTeX arXiv:2601.02125

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital Avatars

2025-08-22 · NVIDIA, :, Chaeyeon Chung, Ilya Fedorov 외 arxiv

Audio-driven facial animation presents an effective solution for animating digital avatars. In this paper, we detail the technical aspects of NVIDIA Audio2Face-3D, including data acquisition, network architecture, retarg…

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens

2026-05-29 · Qingcheng Zhao, Yifang Pan, Karan Singh arxiv

Recent advances in Audio-LLMs like GPT-4o have ushered in an era of conversational interaction with language models. Conversational avatars however, still seem robotic in facial expression and conversational flow, in par…

Speech RecognitionSpeech SynthesisText Generation

Attention-Based VR Facial Animation with Visual Mouth Camera Guidance for Immersive Telepresence Avatars

2023-12-15 · Andre Rochow, Max Schwarz, Sven Behnke

Facial animation in virtual reality environments is essential for applications that necessitate clear visibility of the user's face and the ability to convey emotional signals. In our scenario, we animate the face of an …

AvatarBooth: High-Quality and Customizable 3D Human Avatar Generation

2023-06-16 · Yifei Zeng, Yuanxun Lu, Xinya Ji, Yao Yao 외

We introduce AvatarBooth, a novel method for generating high-quality 3D avatars using text prompts or specific images. Unlike previous approaches that can only synthesize avatars based on simple text descriptions, our me…

Text to 3D

AvatarDynamizer: From Static to Dynamic Human Avatars via Generative Dynamic Textures

2026-08-20 · Guoxing Sun, Heming Zhu, Linjie Lyu, Pascal Fua 외 arxiv

For full-body avatars, modeling surface dynamics is crucial for overcoming the uncanny valley and achieving perceptual realism. Person-agnostic methods recover static 3D avatars from monocular images, videos, or text pro…