paper-with-me

홈 › Papers

Generating Human Interaction Motions in Scenes with Text Control

2024-04-16 · Hongwei Yi, Justus Thies, Michael J. Black, Xue Bin Peng, Davis Rempe

We present TeSMo, a method for text-controlled scene-aware motion generation based on denoising diffusion models. Previous text-to-motion methods focus on characters in isolation without considering scenes due to the limited availability of datasets that include motion, text descriptions, and interactive scenes. Our approach begins with pre-training a scene-agnostic text-to-motion diffusion model, emphasizing goal-reaching constraints on large-scale motion-capture datasets. We then enhance this model with a scene-aware component, fine-tuned using data augmented with detailed scene information, including ground plane and object shapes. To facilitate training, we embed annotated navigation and interaction motions within scenes. The proposed method produces realistic and diverse human-object interactions, such as navigation and sitting, in different scenes with various object shapes, orientations, initial body positions, and poses. Extensive experiments demonstrate that our approach surpasses prior techniques in terms of the plausibility of human-scene interactions, as well as the realism and variety of the generated motions. Code will be released upon publication of this work at https://research.nvidia.com/labs/toronto-ai/tesmo.

📄 PDF Abstract BibTeX arXiv:2404.10685

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingHuman-Object Interaction DetectionMotion GenerationObject

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Generating Human Motion in 3D Scenes from Text Descriptions

2024-05-13 · CVPR 2024 1 · Zhi Cen, Huaijin Pi, Sida Peng, Zehong Shen 외

Generating human motions from textual descriptions has gained growing research interest due to its wide range of applications. However, only a few works consider human-scene interactions together with text conditions, wh…

Motion GenerationObjectSpatial Reasoning

MOGRAS: Human Motion with Grasping in 3D Scenes

2025-10-25 · Kunal Bhosikar, Siddharth Katageri, Vivek Madhavaram, Kai Han 외 arxiv

Generating realistic full-body motion interacting with objects is critical for applications in robotics, virtual reality, and human-computer interaction. While existing methods can generate full-body motion within 3D sce…

InterPose: Learning to Generate Human-Object Interactions from Large-Scale Web Videos

2025-08-31 · Yangsong Zhang, Abdul Ahad Butt, Gül Varol, Ivan Laptev arxiv

Human motion generation has shown great advances thanks to the recent diffusion models trained on large-scale motion capture data. Most of existing works, however, currently target animation of isolated people in empty s…

Dynamic Worlds, Dynamic Humans: Generating Virtual Human-Scene Interaction Motion in Dynamic Scenes

2026-01-27 · Yin Wang, Zhiying Leng, Haitian Liu, Frederick W. B. Li 외 arxiv

Scenes are continuously undergoing dynamic changes in the real world. However, existing human-scene interaction generation methods typically treat the scene as static, which deviates from reality. Inspired by world model…

Autonomous Character-Scene Interaction Synthesis from Text Instruction

2024-10-04 · Nan Jiang, Zimo He, Zi Wang, Hongjie Li 외

Synthesizing human motions in 3D environments, particularly those with complex activities such as locomotion, hand-reaching, and human-object interaction, presents substantial demands for user-defined waypoints and stage…

Human-Object Interaction Detection