paper-with-me

홈 › Papers

Fine-grained text-driven dual-human motion generation via dynamic hierarchical interaction

2025-10-09 · Mu Li, Yin Wang, Zhiying Leng, Jiapeng Liu, Frederick W. B. Li, Xiaohui Liang arxiv

Human interaction is inherently dynamic and hierarchical, where the dynamic refers to the motion changes with distance, and the hierarchy is from individual to inter-individual and ultimately to overall motion. Exploiting these properties is vital for dual-human motion generation, while existing methods almost model human interaction temporally invariantly, ignoring distance and hierarchy. To address it, we propose a fine-grained dual-human motion generation method, namely FineDual, a tri-stage method to model the dynamic hierarchical interaction from individual to inter-individual. The first stage, Self-Learning Stage, divides the dual-human overall text into individual texts through a Large Language Model, aligning text features and motion features at the individual level. The second stage, Adaptive Adjustment Stage, predicts interaction distance by an interaction distance predictor, modeling human interactions dynamically at the inter-individual level by an interaction-aware graph network. The last stage, Teacher-Guided Refinement Stage, utilizes overall text features as guidance to refine motion features at the overall level, generating fine-grained and high-quality dual-human motion. Extensive quantitative and qualitative evaluations on dual-human motion datasets demonstrate that our proposed FineDual outperforms existing approaches, effectively modeling dynamic hierarchical human interaction.

📄 PDF Abstract BibTeX arXiv:2510.08260

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FinePOSE: Fine-Grained Prompt-Driven 3D Human Pose Estimation via Diffusion Models

2024-05-08 · CVPR 2024 1 · Jinglin Xu, Yijie Guo, Yuxin Peng

The 3D Human Pose Estimation (3D HPE) task uses 2D images or videos to predict human joint coordinates in 3D space. Despite recent advancements in deep learning-based methods, they mostly ignore the capability of couplin…

3D Human Pose EstimationDenoisingPose EstimationPrompt Learning

Dual Diffusion Models for Multi-modal Guided 3D Avatar Generation

2026-03-04 · Hong Li, Yutang Feng, Minqi Meng, Yichen Yang 외 arxiv

Generating high-fidelity 3D avatars from text or image prompts is highly sought after in virtual reality and human-computer interaction. However, existing text-driven methods often rely on iterative Score Distillation Sa…

Computational Efficiency

Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion Model

2023-09-12 · ICCV 2023 1 · Yin Wang, Zhiying Leng, Frederick W. B. Li, Shun-Cheng Wu 외

Text-driven human motion generation in computer vision is both significant and challenging. However, current methods are limited to producing either deterministic or imprecise motion sequences, failing to effectively con…

Motion GenerationMotion Synthesis

Enhanced Fine-grained Motion Diffusion for Text-driven Human Motion Synthesis

2023-05-23 · Dong Wei, Xiaoning Sun, Huaijiang Sun, Bin Li 외

The emergence of text-driven motion synthesis technique provides animators with great potential to create efficiently. However, in most cases, textual expressions only contain general and qualitative motion descriptions,…

Motion Synthesisvalid

TextToucher: Fine-Grained Text-to-Touch Generation

2024-09-09 · Jiahang Tu, Hao Fu, Fengyu Yang, Hanbin Zhao 외

Tactile sensation plays a crucial role in the development of multi-modal large models and embodied intelligence. To collect tactile data with minimal cost as possible, a series of studies have attempted to generate tacti…

Language ModellingLarge Language ModelMultimodal Large Language Model