paper-with-me

홈 › Papers

PackDiT: Joint Human Motion and Text Generation via Mutual Prompting

2025-01-27 · Zhongyu Jiang, Wenhao Chai, Zhuoran Zhou, Cheng-Yen Yang, Hsiang-Wei Huang, Jenq-Neng Hwang

Human motion generation has advanced markedly with the advent of diffusion models. Most recent studies have concentrated on generating motion sequences based on text prompts, commonly referred to as text-to-motion generation. However, the bidirectional generation of motion and text, enabling tasks such as motion-to-text alongside text-to-motion, has been largely unexplored. This capability is essential for aligning diverse modalities and supports unconditional generation. In this paper, we introduce PackDiT, the first diffusion-based generative model capable of performing various tasks simultaneously, including motion generation, motion prediction, text generation, text-to-motion, motion-to-text, and joint motion-text generation. Our core innovation leverages mutual blocks to integrate multiple diffusion transformers (DiTs) across different modalities seamlessly. We train PackDiT on the HumanML3D dataset, achieving state-of-the-art text-to-motion performance with an FID score of 0.106, along with superior results in motion prediction and in-between tasks. Our experiments further demonstrate that diffusion models are effective for motion-to-text generation, achieving performance comparable to that of autoregressive models.

📄 PDF Abstract BibTeX arXiv:2501.16551

Code (0)

등록된 구현이 없습니다.

Tasks

Motion Generationmotion predictionText Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Jointly Understand Your Command and Intention:Reciprocal Co-Evolution between Scene-Aware 3D Human Motion Synthesis and Analysis

2025-03-01 · Xuehao Gao, Yang Yang, Shaoyi Du, Guo-Jun Qi 외

As two intimate reciprocal tasks, scene-aware human motion synthesis and analysis require a joint understanding between multiple modalities, including 3D body motions, 3D scenes, and textual descriptions. In this paper, …

DiversityMotion GenerationMotion Synthesis

Uni-HOI:A Unified framework for Learning the Joint distribution of Text and Human-Object Interaction

2026-04-30 · Mengfei Zhang, Jinlu Zhang, Zhigang Tu arxiv

Modeling 4D human-object interaction (HOI) is a compelling challenge in computer vision and an essential technology powering virtual and mixed-reality applications. While existing works have achieved promising results on…

Multi-Task Learning

Strong and Controllable 3D Motion Generation

2025-01-30 · Canxuan Gang

Human motion generation is a significant pursuit in generative computer vision with widespread applications in film-making, video games, AR/VR, and human-robot interaction. Current methods mainly utilize either diffusion…

Motion GenerationRobot Manipulation

Pulp Motion: Framing-aware multimodal camera and human motion generation

2025-10-06 · Robin Courant, Xi Wang, David Loiseaux, Marc Christie 외 arxiv

Treating human motion and camera trajectory generation separately overlooks a core principle of cinematography: the tight interplay between actor performance and camera work in the screen space. In this paper, we are the…

Motion-2-to-3: Leveraging 2D Motion Data to Boost 3D Motion Generation

2024-12-17 · Huaijin Pi, Ruoxi Guo, Zehong Shen, Qing Shuai 외

Text-driven human motion synthesis is capturing significant attention for its ability to effortlessly generate intricate movements from abstract text cues, showcasing its potential for revolutionizing motion design not o…

Motion GenerationMotion Synthesis