paper-with-me

Papers

Learning Asynchronous and Sparse Human-Object Interaction in Videos

2021-03-03 · CVPR 2021 1 · Romero Morais, Vuong Le, Svetha Venkatesh, Truyen Tran

Human activities can be learned from video. With effective modeling it is possible to discover not only the action labels but also the temporal structures of the activities such as the progression of the sub-activities. Automatically recognizing such structure from raw video signal is a new capability that promises authentic modeling and successful recognition of human-object interactions. Toward this goal, we introduce Asynchronous-Sparse Interaction Graph Networks (ASSIGN), a recurrent graph network that is able to automatically detect the structure of interaction events associated with entities in a video scene. ASSIGN pioneers learning of autonomous behavior of video entities including their dynamic structure and their interaction with the coexisting neighbors. Entities' lives in our model are asynchronous to those of others therefore more flexible in adaptation to complex scenarios. Their interactions are sparse in time hence more faithful to the true underlying nature and more robust in inference and learning. ASSIGN is tested on human-object interaction recognition and shows superior performance in segmenting and labeling of human sub-activities and object affordances from raw videos. The native ability for discovering temporal structures of the model also eliminates the dependence on external segmentation that was previously mandatory for this task.

📄 PDF Abstract BibTeX arXiv:2103.02758

Code (0)

등록된 구현이 없습니다.

Tasks

Human-Object Interaction DetectionObject

Similar Papers 제목 키워드 기반

SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation

2026-07-07 · Shenbo Xie, Mingrui Cai, Xu Yang, Yifei Liu 외 arxiv

Human-Object Interaction (HOI) video generation aims to synthesize realistic videos of humans manipulating diverse objects, serving as a promising avenue for AI-driven live streaming e-commerce. A primary obstacle in thi…

Motion SynthesisVideo Generationhand-object pose

DISPLAY: Directable Human-Object Interaction Video Generation via Sparse Motion Guidance and Multi-Task Auxiliary

2026-03-10 · Jiazhi Guan, Quanwei Yang, Luying Huang, Junhao Liang 외 arxiv

Human-centric video generation has advanced rapidly, yet existing methods struggle to produce controllable and physically consistent Human-Object Interaction (HOI) videos. Existing works rely on dense control signals, te…

Video Generation

Efficient and Scalable Monocular Human-Object Interaction Motion Reconstruction

2025-11-30 · Boran Wen, Ye Lu, Sirui Wang, Keyan Wan 외 arxiv

Generalized robots must learn from diverse, large-scale human-object interactions (HOI) to operate robustly in the real world. Monocular internet videos offer a nearly limitless and readily available source of data, capt…

HDMI: Learning Interactive Humanoid Whole-Body Control from Human Videos

2025-09-20 · Haoyang Weng, Yitang Li, Nikhil Sobanbabu, Zihan Wang 외 arxiv

Enabling robust whole-body humanoid-object interaction (HOI) remains challenging due to motion data scarcity and the contact-rich nature. We present HDMI (HumanoiD iMitation for Interaction), a simple and general framewo…

Reinforcement Learning

The MECCANO Dataset: Understanding Human-Object Interactions from Egocentric Videos in an Industrial-like Domain

2020-10-12 · Francesco Ragusa, Antonino Furnari, Salvatore Livatino, Giovanni Maria Farinella

Wearable cameras allow to collect images and videos of humans interacting with the world. While human-object interactions have been thoroughly investigated in third person vision, the problem has been understudied in ego…

Action RecognitionActive Object DetectionHuman-Object Interaction DetectionObject+3