paper-with-me

홈 › Papers

Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?

2026-07-09 · Masatoshi Tateno, Alexandros Stergiou, Risa Shinoda, Yoichi Sato, Dima Damen arxiv

Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models often fall into shortcuts by relying on spurious correlations among hands, objects, or environmental context, rather than reasoning from the appearance and dynamics of hands and objects themselves. To address this limitation, we propose a new learning paradigm that combines (i) hand-object masked training, which enables robust reasoning from partial hand or object observations, and (ii) an HOI-dynamics-aware decoder that explicitly learns hand- and object-centric embeddings through auxiliary predictions of their locations and semantics, enhancing sensitivity to both cues. To systematically evaluate such cue-specific reasoning, we introduce Cue-Isolated HOI (CI-HOI), a new evaluation that assesses models' ability to predict actions from hand- and object-related cues independently. To enable CI-HOI, we curate the DEHOI testbed, which separates hand- and object-related observations for disentangled HOI evaluation through inpainting. Using DEHOI, we demonstrate both quantitatively and qualitatively that our training strategy exploits hand- and object-centric information more effectively than existing models. Our approach improves over existing models on DEHOI, standard action recognition, object state recognition, and even robot manipulation action recognition, leading to more robust HOI understanding.

📄 PDF Abstract BibTeX arXiv:2607.08514

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionRobot Manipulation

Similar Papers 제목 키워드 기반

Developing Vision-Language-Action Model from Egocentric Videos

2025-09-26 · Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, Shinsuke Mori arxiv

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-La…

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting

2026-05-08 · Jaeyoung Choi, Hyeondong Kim, Yujin Kim, Daehee Park arxiv

Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such as AR/VR assistance and human-robot interaction. However, this task r…

Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning

2024-08-07 · Zi-Yi Dou, Xitong Yang, Tushar Nagarajan, Huiyu Wang 외

We present EMBED (Egocentric Models Built with Exocentric Data), a method designed to transform exocentric video-language data for egocentric video representation learning. Large-scale exocentric data covers diverse acti…

Multi-Instance RetrievalRepresentation LearningStyle Transfer

Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning

2025-03-02 · Baoqi Pei, Yifei HUANG, Jilan Xu, Guo Chen 외

In egocentric video understanding, the motion of hands and objects as well as their interactions play a significant role by nature. However, existing egocentric video representation learning methods mainly focus on align…

Large Language ModelMulti-Instance RetrievalObjectRepresentation Learning+2

EgoSurgery-HTS: A Dataset for Egocentric Hand-Tool Segmentation in Open Surgery Videos

2025-03-24 · Nathan Darjana, Ryo Fujii, Hideo Saito, Hiroki Kajita

Egocentric open-surgery videos capture rich, fine-grained details essential for accurately modeling surgical procedures and human behavior in the operating room. A detailed, pixel-level understanding of hands and surgica…

Instance SegmentationSegmentationSemantic Segmentation