paper-with-me

홈 › Papers

SingingHead: A Large-scale 4D Dataset for Singing Head Animation

2023-12-07 · Sijing Wu, Yunhao Li, Weitian Zhang, Jun Jia, Yucheng Zhu, Yichao Yan, Guangtao Zhai, Xiaokang Yang

Singing, as a common facial movement second only to talking, can be regarded as a universal language across ethnicities and cultures, plays an important role in emotional communication, art, and entertainment. However, it is often overlooked in the field of audio-driven facial animation due to the lack of singing head datasets and the domain gap between singing and talking in rhythm and amplitude. To this end, we collect a high-quality large-scale singing head dataset, SingingHead, which consists of more than 27 hours of synchronized singing video, 3D facial motion, singing audio, and background music from 76 individuals and 8 types of music. Along with the SingingHead dataset, we benchmark existing audio-driven 3D facial animation methods and 2D talking head methods on the singing task. Furthermore, we argue that 3D and 2D facial animation tasks can be solved together, and propose a unified singing head animation framework named UniSinger to achieve both singing audio-driven 3D singing head animation and 2D singing portrait video synthesis, which achieves competitive results on both 3D and 2D benchmarks. Extensive experiments demonstrate the significance of the proposed singing-specific dataset in promoting the development of singing head animation tasks, as well as the promising performance of our unified facial animation framework.

📄 PDF Abstract BibTeX arXiv:2312.04369

Code (0)

등록된 구현이 없습니다.

Tasks

Portrait AnimationRhythm

Similar Papers 제목 키워드 기반

Let's Chorus: Partner-aware Hybrid Song-Driven 3D Head Animation

2025-01-01 · CVPR 2025 1 · Xiumei Xie, Zikai Huang, Wenhao Xu, Peng Xiao 외

Singing is a vital form of human emotional expression and social interaction, distinguished from speech by its richer emotional nuances and freer expressive style. Thus, investigating 3D facial animation driven by si…

Automatic Estimation of Singing Voice Musical Dynamics

2024-10-27 · Jyoti Narang, Nazif Can Tamer, Viviana De La Vega, Xavier Serra

Musical dynamics form a core part of expressive singing voice performances. However, automatic analysis of musical dynamics for singing voice has received limited attention partly due to the scarcity of suitable datasets…

Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation

2025-09-02 · Zikai Huang, Yihan Zhou, Xuemiao Xu, Cheng Xu 외 arxiv

Singing-driven 3D head animation is a challenging yet promising task with applications in virtual avatars, entertainment, and education. Unlike speech, singing involves richer emotional nuance, dynamic prosody, and lyric…

Resource-constrained stereo singing voice cancellation

2024-01-22 · Clara Borrelli, James Rae, Dogac Basaran, Matt Mcvicar 외

We study the problem of stereo singing voice cancellation, a subtask of music source separation, whose goal is to estimate an instrumental background from a stereo mix. We explore how to achieve performance similar to la…

Music Source SeparationSpeech Separation

A Comparative Study of Voice Conversion Models with Large-Scale Speech and Singing Data: The T13 Systems for the Singing Voice Conversion Challenge 2023

2023-10-08 · Ryuichi Yamamoto, Reo Yoneyama, Lester Phillip Violeta, Wen-Chin Huang 외

This paper presents our systems (denoted as T13) for the singing voice conversion challenge (SVCC) 2023. For both in-domain and cross-domain English singing voice conversion (SVC) tasks (Task 1 and Task 2), we adopt a re…

Self-Supervised LearningTask 2Voice Conversion