paper-with-me

홈 › Papers

Interact with me: Joint Egocentric Forecasting of Intent to Interact, Attitude and Social Actions

2024-12-21 · Tongfei Bian, Yiming Ma, Mathieu Chollet, Victor Sanchez, Tanaya Guha

For efficient human-agent interaction, an agent should proactively recognize their target user and prepare for upcoming interactions. We formulate this challenging problem as the novel task of jointly forecasting a person's intent to interact with the agent, their attitude towards the agent and the action they will perform, from the agent's (egocentric) perspective. So we propose \emph{SocialEgoNet} - a graph-based spatiotemporal framework that exploits task dependencies through a hierarchical multitask learning approach. SocialEgoNet uses whole-body skeletons (keypoints from face, hands and body) extracted from only 1 second of video input for high inference speed. For evaluation, we augment an existing egocentric human-agent interaction dataset with new class labels and bounding box annotations. Extensive experiments on this augmented dataset, named JPL-Social, demonstrate \emph{real-time} inference and superior performance (average accuracy across all tasks: 83.15\%) of our model outperforming several competitive baselines. The additional annotations and code will be available upon acceptance.

📄 PDF Abstract BibTeX arXiv:2412.16698

Code (1)

biantongfei/SocialEgoNet 공식 구현 pytorch

Similar Papers 제목 키워드 기반

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting

2026-05-08 · Jaeyoung Choi, Hyeondong Kim, Yujin Kim, Daehee Park arxiv

Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such as AR/VR assistance and human-robot interaction. However, this task r…

EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control

2026-06-07 · Haoyang Ge, Peng Ren, Yukun Shi, Cong Huang 외 arxiv

Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Motion tracking reproduces specified trajectories, and humanoid vision-language-action systems provide semantic …

Forecasting Human-Object Interaction: Joint Prediction of Motor Attention and Actions in First Person Video

2019-11-25 · ECCV 2020 8 · Miao Liu, Siyu Tang, Yin Li, James Rehg

We address the challenging task of anticipating human-object interaction in first person videos. Most existing methods ignore how the camera wearer interacts with the objects, or simply consider body motion as a separate…

Action AnticipationHuman-Object Interaction Detection

Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span

2025-11-23 · Heeseung Yun, Joonil Na, Jaeyeon Kim, Calvin Murdock 외 arxiv

People continuously perceive and interact with their surroundings based on underlying intentions that drive their exploration and behaviors. While research in egocentric user and scene understanding has focused primarily…

Scene Understanding

EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views

2024-05-22 · Yuhang Yang, Wei Zhai, Chengfeng Wang, Chengjun Yu 외

Understanding egocentric human-object interaction (HOI) is a fundamental aspect of human-centric perception, facilitating applications like AR/VR and embodied AI. For the egocentric HOI, in addition to perceiving semanti…

Human-Object Interaction DetectionObject