paper-with-me

홈 › Papers

DanceCrafter: Fine-Grained Text-Driven Controllable Dance Generation via Choreographic Syntax

2026-04-20 · Hang Yuan, Xiaolin Hu, Yan Wan, Menglin Gao, Wenzhe Yu, Cong Huang, Fei Xu, Qing Li, Christina Dan Wang, Zhou Yu, Kai Chen arxiv

Text-driven controllable dance generation remains under-explored, primarily due to the severe scarcity of high-quality datasets and the inherent difficulty of articulating complex choreographies. Characterizing dance is particularly challenging owing to its intricate spatial dynamics, strong directionality, and the highly decoupled movements of distinct body parts. To overcome these bottlenecks, we bridge principles from dance studies, human anatomy, and biomechanics to propose \textit{Choreographic Syntax}, a novel theoretical framework with a tailored annotation system. Grounded in this syntax, we combine professional dance archives with high-fidelity motion capture data to construct \textbf{DanceFlow}, the most fine-grained dance dataset to date. It encompasses 41 hours of high-quality motions paired with 6.34 million words of detailed descriptions. At the model level, we introduce \textbf{DanceCrafter}, a tailored motion transformer built upon the Momentum Human Rig. To circumvent optimization instabilities, we construct a continuous manifold motion representation paired with a hybrid normalization strategy. Furthermore, we design an anatomy-aware loss to explicitly regulate the decoupled nature of body parts. Together, these adaptations empower DanceCrafter to achieve the high-fidelity and stable generation of complex dance sequences. Extensive evaluations and user studies demonstrate our state-of-the-art performance in motion quality, fine-grained controllability, and generation naturalness.

📄 PDF Abstract BibTeX arXiv:2604.18648

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ConFusion: Continuous Fusion Space Learning for Fine-Grained Controllable Infrared and Visible Image Fusion

2026-07-26 · Guo Yurong, He Yufei, Li Yonghao, Chang Dongliang 외 arxiv

Controllable infrared-visible image fusion aims to integrate complementary thermal and structural information with flexible region-aware modulation, producing fused images that adapt to diverse user requirements and down…

3DStyle-Diffusion: Pursuing Fine-grained Text-driven 3D Stylization with 2D Diffusion Models

2023-11-09 · Haibo Yang, Yang Chen, Yingwei Pan, Ting Yao 외

3D content creation via text-driven stylization has played a fundamental challenge to multimedia and graphics community. Recent advances of cross-modal foundation models (e.g., CLIP) have made this problem feasible. Thos…

Image Generation

OpenDance: Multimodal Controllable 3D Dance Generation Using Large-scale Internet Data

2025-06-09 · Jinlu Zhang, Zixi Kang, Yizhou Wang

Music-driven dance generation offers significant creative potential yet faces considerable challenges. The absence of fine-grained multimodal data and the difficulty of flexible multi-conditional generation limit previou…

Diversity

FineXtrol: Controllable Motion Generation via Fine-Grained Text

2025-11-24 · Keming Shen, Bizhu Wu, Junliang Chen, Xiaoqin Wang 외 arxiv

Recent works have sought to enhance the controllability and precision of text-driven motion generation. Some approaches leverage large language models (LLMs) to produce more detailed texts, while others incorporate globa…

Contrastive Learning

Text2Traffic: A Text-to-Image Generation and Editing Method for Traffic Scenes

2025-11-17 · Feng Lv, Haoxuan Feng, Zilu Zhang, Chunlong Xia 외 arxiv

With the rapid advancement of intelligent transportation systems, text-driven image generation and editing techniques have demonstrated significant potential in providing rich, controllable visual scene data for applicat…

Text-to-Image GenerationAutonomous Driving