paper-with-me

홈 › Papers

Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation

2025-09-02 · Zikai Huang, Yihan Zhou, Xuemiao Xu, Cheng Xu, Xiaofen Xing, Jing Qin, Shengfeng He arxiv

Singing-driven 3D head animation is a challenging yet promising task with applications in virtual avatars, entertainment, and education. Unlike speech, singing involves richer emotional nuance, dynamic prosody, and lyric-based semantics, requiring the synthesis of fine-grained, temporally coherent facial motion. Existing speech-driven approaches often produce oversimplified, emotionally flat, and semantically inconsistent results, which are insufficient for singing animation. To address this, we propose Think2Sing, a diffusion-based framework that leverages pretrained large language models to generate semantically coherent and temporally consistent 3D head animations, conditioned on both lyrics and acoustics. A key innovation is the introduction of motion subtitles, an auxiliary semantic representation derived through a novel Singing Chain-of-Thought reasoning process combined with acoustic-guided retrieval. These subtitles contain precise timestamps and region-specific motion descriptions, serving as interpretable motion priors. We frame the task as a motion intensity prediction problem, enabling finer control over facial regions and improving the modeling of expressive motion. To support this, we create a multimodal singing dataset with synchronized video, acoustic descriptors, and motion subtitles, enabling diverse and expressive motion learning. Extensive experiments show that Think2Sing outperforms state-of-the-art methods in realism, expressiveness, and emotional fidelity, while also offering flexible, user-controllable animation editing.

📄 PDF Abstract BibTeX arXiv:2509.02278

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fine-grained Emotion and Intent Learning in Movie Dialogues

2020-12-25 · Anuradha Welivita, Yubo Xie, Pearl Pu

We propose a novel large-scale emotional dialogue dataset, consisting of 1M dialogues retrieved from the OpenSubtitles corpus and annotated with 32 emotions and 9 empathetic response intents using a BERT-based fine-grain…

Post-editing in Automatic Subtitling: A Subtitlers’ perspective

2022-06-01 · EAMT 2022 6 · Alina Karakanta, Luisa Bentivogli, Mauro Cettolo, Matteo Negri 외

Recent developments in machine translation and speech translation are opening up opportunities for computer-assisted translation tools with extended automation functions. Subtitling tools are recently being adapted for p…

Machine TranslationTranslation

Automatic Comic Generation with Stylistic Multi-page Layouts and Emotion-driven Text Balloon Generation

2021-01-26 · Xin Yang, Zongliang Ma, Letian Yu, Ying Cao 외

In this paper, we propose a fully automatic system for generating comic books from videos without any human intervention. Given an input video along with its subtitles, our approach first extracts informative keyframes b…

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space

2026-03-14 · Quoc-Huy Trinh, Xi Ding, Yang Liu, Zhenyue Qin 외 arxiv

Visual spatial intelligence is critical for medical image interpretation, yet remains largely unexplored in Multimodal Large Language Models (MLLMs) for 3D imaging. This gap persists due to a systemic lack of datasets fe…

Spatial Reasoning

NeutronOrch: Rethinking Sample-based GNN Training under CPU-GPU Heterogeneous Environments

2023-11-22 · Xin Ai, Qiange Wang, Chunyu Cao, Yanfeng Zhang 외

Graph Neural Networks (GNNs) have demonstrated outstanding performance in various applications. Existing frameworks utilize CPU-GPU heterogeneous environments to train GNN models and integrate mini-batch and sampling tec…

CPUGPU