paper-with-me

홈 › Papers

ParTY: Part-Guidance for Expressive Text-to-Motion Synthesis

2026-03-10 · KunHo Heo, SuYeon Kim, Yonghyun Gwon, Youngbin Kim, MyeongAh Cho arxiv

Text-to-motion synthesis aims to generate natural and expressive human motions from textual descriptions. While existing approaches primarily focus on generating holistic motions from text descriptions, they struggle to accurately reflect actions involving specific body parts. Recent part-wise motion generation methods attempt to resolve this but face two critical limitations: (i) they lack explicit mechanisms for aligning textual semantics with individual body parts, and (ii) they often generate incoherent full-body motions due to integrating independently generated part motions. To overcome these issues and resolve the fundamental trade-off in existing methods, we propose ParTY, a novel framework that enhances part expressiveness while generating coherent full-body motions. ParTY comprises: (1) Part-Guided Network, which first generates part motions to obtain part guidance, then uses it to generate holistic motions; (2) Part-aware Text Grounding, which diversely transforms text embeddings and appropriately aligns them with each body part; and (3) Holistic-Part Fusion, which adaptively fuses holistic motions and part motions. Extensive experiments, including part-level and coherence-level evaluations, demonstrate that ParTY achieves substantial improvements over previous methods.

📄 PDF Abstract BibTeX arXiv:2603.09611

Code (0)

등록된 구현이 없습니다.

Tasks

Motion Synthesis

Similar Papers 제목 키워드 기반

DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue

2025-04-20 · Xiang Li, Duyi Pan, Hongru Xiao, Jiale Han 외

Speech synthesis is crucial for human-computer interaction, enabling natural and intuitive communication. However, existing datasets involve high construction costs due to manual annotation and suffer from limited charac…

DiversitySpeech Synthesis

MuCDN: Mutual Conversational Detachment Network for Emotion Recognition in Multi-Party Conversations

2022-10-01 · COLING 2022 10 · Weixiang Zhao, Yanyan Zhao, Bing Qin

As an emerging research topic in natural language processing community, emotion recognition in multi-party conversations has attained increasing interest. Previous approaches that focus either on dyadic or multi-party sc…

Emotion Recognition

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models

2025-10-15 · Yizhou Peng, Yukun Ma, Chong Zhang, Yi-Wen Chao 외 arxiv

While Text-to-Speech (TTS) systems enable emotional control via natural-language instructions, expressiveness, naturalness, and speech quality degrade when the target emotion conflicts with the textual semantics. We prop…

ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance

2024-10-12 · Yongkang Cheng, Mingjiang Liang, Shaoli Huang, Jifeng Ning 외

Existing gesture generation methods primarily focus on upper body gestures based on audio features, neglecting speech content, emotion, and locomotion. These limitations result in stiff, mechanical gestures that fail to …

Gesture Generation

DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance

2025-04-02 · Yuxuan Luo, Zhengkun Rong, Lizhen Wang, Longhao Zhang 외

While recent image-based human animation methods achieve realistic body and facial motion synthesis, critical gaps remain in fine-grained holistic controllability, multi-scale adaptability, and long-term temporal coheren…

Human AnimationImage AnimationMotion Synthesis