Learn to Predict How Humans Manipulate Large-sized Objects from Interactive Motions
Understanding human intentions during interactions has been a long-lasting theme, that has applications in human-robot interaction, virtual reality and surveillance. In this study, we focus on full-body human interactions with large-sized daily objects and aim to predict the future states of objects and humans given a sequential observation of human-object interaction. As there is no such dataset dedicated to full-body human interactions with large-sized daily objects, we collected a large-scale dataset containing thousands of interactions for training and evaluation purposes. We also observe that an object's intrinsic physical properties are useful for the object motion prediction, and thus design a set of object dynamic descriptors to encode such intrinsic properties. We treat the object dynamic descriptors as a new modality and propose a graph neural network, HO-GCN, to fuse motion data and dynamic descriptors for the prediction task. We show the proposed network that consumes dynamic descriptors can achieve state-of-the-art prediction results and help the network better generalize to unseen objects. We also demonstrate the predicted results are useful for human-robot collaborations.
Code (0)
등록된 구현이 없습니다.
Tasks
Graph Neural NetworkHuman-Object Interaction Detectionmotion predictionObjectPredictionSimilar Papers 제목 키워드 기반
Object Motion Guided Human Motion Synthesis
Modeling human behaviors in contextual environments has a wide range of applications in character animation, embodied AI, VR/AR, and robotics. In real-world scenarios, humans frequently interact with the environment and …
DenoisingHuman-Object Interaction DetectionMotion SynthesisObjectM3D-VTON: A Monocular-to-3D Virtual Try-On Network
Virtual 3D try-on can provide an intuitive and realistic view for online shopping and has a huge potential commercial value. However, existing 3D virtual try-on methods mainly rely on annotated 3D human shapes and garmen…
Virtual Try-onAssessing Gender Bias in Predictive Algorithms using eXplainable AI
Predictive algorithms have a powerful potential to offer benefits in areas as varied as medicine or education. However, these algorithms and the data they use are built by humans, consequently, they can inherit the bias …
Facial Expression RecognitionFacial Expression Recognition (FER)Understanding Human Hands in Contact at Internet Scale
Hands are the central means by which humans manipulate their world and being able to reliably extract hand state information from Internet videos of humans engaged in their hands has the potential to pave the way to syst…
MVP-Bench: Can Large Vision--Language Models Conduct Multi-level Visual Perception Like Humans?
Humans perform visual perception at multiple levels, including low-level object recognition and high-level semantic interpretation such as behavior understanding. Subtle differences in low-level details can lead to subst…
Object Recognition