IMos: Intent-Driven Full-Body Motion Synthesis for Human-Object Interactions
Can we make virtual characters in a scene interact with their surrounding objects through simple instructions? Is it possible to synthesize such motion plausibly with a diverse set of objects and instructions? Inspired by these questions, we present the first framework to synthesize the full-body motion of virtual human characters performing specified actions with 3D objects placed within their reach. Our system takes textual instructions specifying the objects and the associated intentions of the virtual characters as input and outputs diverse sequences of full-body motions. This contrasts existing works, where full-body action synthesis methods generally do not consider object interactions, and human-object interaction methods focus mainly on synthesizing hand or finger movements for grasping objects. We accomplish our objective by designing an intent-driven fullbody motion generator, which uses a pair of decoupled conditional variational auto-regressors to learn the motion of the body parts in an autoregressive manner. We also optimize the 6-DoF pose of the objects such that they plausibly fit within the hands of the synthesized characters. We compare our proposed method with the existing methods of motion synthesis and establish a new and stronger state-of-the-art for the task of intent-driven motion synthesis.
Code (1)
Tasks
Human-Object Interaction DetectionMotion SynthesisSimilar Papers 제목 키워드 기반
Synthesizing Diverse Human Motions in 3D Indoor Scenes
We present a novel method for populating 3D indoor scenes with virtual humans that can navigate in the environment and interact with objects in a realistic manner. Existing approaches rely on training sequences that cont…
Collision AvoidanceHuman-Object Interaction DetectionNavigateRecognizing Individuals and Their Emotions Using Plants as Bio-Sensors through Electro-static Discharge
By measuring the electrostatic discharge of human bodies together with Mimosa Pudica and other plants in response to the human movement, we have been able to recognize (a) individuals based on their distinctive pattern o…
TextOp: Real-time Interactive Text-Driven Humanoid Robot Motion Generation and Control
Recent advances in humanoid whole-body motion tracking have enabled the execution of diverse and highly coordinated motions on real hardware. However, existing controllers are commonly driven either by predefined motion …
Mimosa Framework: Toward Evolving Multi-Agent Systems for Scientific Research
Current Autonomous Scientific Research (ASR) systems, despite leveraging large language models (LLMs) and agentic architectures, remain constrained by fixed workflows and toolsets that prevent adaptation to evolving task…
TextIM: Part-aware Interactive Motion Synthesis from Text
In this work, we propose TextIM, a novel framework for synthesizing TEXT-driven human Interactive Motions, with a focus on the precise alignment of part-level semantics. Existing methods often overlook the critical roles…
Motion Synthesis