paper-with-me

Papers

Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning

2025-03-02 · Baoqi Pei, Yifei HUANG, Jilan Xu, Guo Chen, Yuping He, Lijin Yang, Yali Wang, Weidi Xie, Yu Qiao, Fei Wu, LiMin Wang

In egocentric video understanding, the motion of hands and objects as well as their interactions play a significant role by nature. However, existing egocentric video representation learning methods mainly focus on aligning video representation with high-level narrations, overlooking the intricate dynamics between hands and objects. In this work, we aim to integrate the modeling of fine-grained hand-object dynamics into the video representation learning process. Since no suitable data is available, we introduce HOD, a novel pipeline employing a hand-object detector and a large language model to generate high-quality narrations with detailed descriptions of hand-object dynamics. To learn these fine-grained dynamics, we propose EgoVideo, a model with a new lightweight motion adapter to capture fine-grained hand-object motion information. Through our co-training strategy, EgoVideo effectively and efficiently leverages the fine-grained hand-object dynamics in the HOD data. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple egocentric downstream tasks, including improvements of 6.3% in EK-100 multi-instance retrieval, 5.7% in EK-100 classification, and 16.3% in EGTEA classification in zero-shot settings. Furthermore, our model exhibits robust generalization capabilities in hand-object interaction and robot manipulation tasks. Code and data are available at https://github.com/OpenRobotLab/EgoHOD/.

📄 PDF Abstract BibTeX arXiv:2503.00986

Code (1)

openrobotlab/egohod 공식 구현 pytorch

Tasks

Large Language ModelMulti-Instance RetrievalObjectRepresentation LearningRobot ManipulationVideo Understanding

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Adapter 설명 없음

Similar Papers 제목 키워드 기반

HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics

2025-11-30 · Masatoshi Tateno, Gido Kato, Hirokatsu Kataoka, Yoichi Sato 외 arxiv

Hand-object interaction (HOI) inherently involves dynamics where human manipulations produce distinct spatio-temporal effects on objects. However, existing semantic HOI benchmarks focused either on manipulation or on the…

Video Object Segmentation

Real-Time Human-Centric World Modeling for Upper-Body Human-Object Interaction

2026-07-26 · Chaonan Ji, Jinwei Qi, Peng Zhang, Bang Zhang arxiv

We present a real-time human-centric world model for upper-body interactive generation, aiming to synthesize coherent local world dynamics centered on a person, where coordinated body, hand, and facial motions evolve joi…

DyG$^2$T: Modeling Object Dynamics with 3D Gaussian Temporal-Spatial Particle Graph Transformer

2026-08-19 · Yansong Wang, Zhaobo Qi, Xinyan Liu, Beichen Zhang 외 arxiv

Modeling object dynamics from limited visual observations is a fundamental problem for enabling accurate motion trajectory prediction in embodied interaction scenarios. Existing dynamics modeling methods first compress r…

Trajectory Prediction

Fine-Grained Egocentric Hand-Object Segmentation: Dataset, Model, and Applications

2022-08-07 · Lingzhi Zhang, Shenghao Zhou, Simon Stent, Jianbo Shi

Egocentric videos offer fine-grained information for high-fidelity modeling of human behaviors. Hands and interacting objects are one crucial aspect of understanding a viewer's behaviors and intentions. We provide a labe…

Activity RecognitionData AugmentationObjectSegmentation+2

Multiple Granularity Analysis for Fine-grained Action Detection

2014-06-01 · CVPR 2014 6 · Bingbing Ni, Vignesh R. Paramathayalan, Pierre Moulin

We propose to decompose the fine-grained human activity analysis problem into two sequential tasks with increasing granularity. Firstly, we infer the coarse interaction status, i.e., which object is being manipulated and…

Action DetectionFine-Grained Action DetectionObjectPosition