paper-with-me

홈 › Papers

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction

2024-12-17 · Chen Bao, Jiarui Xu, Xiaolong Wang, Abhinav Gupta, Homanga Bharadhwaj

How can we predict future interaction trajectories of human hands in a scene given high-level colloquial task specifications in the form of natural language? In this paper, we extend the classic hand trajectory prediction task to two tasks involving explicit or implicit language queries. Our proposed tasks require extensive understanding of human daily activities and reasoning abilities about what should be happening next given cues from the current scene. We also develop new benchmarks to evaluate the proposed two tasks, Vanilla Hand Prediction (VHP) and Reasoning-Based Hand Prediction (RBHP). We enable solving these tasks by integrating high-level world knowledge and reasoning capabilities of Vision-Language Models (VLMs) with the auto-regressive nature of low-level ego-centric hand trajectories. Our model, HandsOnVLM is a novel VLM that can generate textual responses and produce future hand trajectories through natural-language conversations. Our experiments show that HandsOnVLM outperforms existing task-specific methods and other VLM baselines on proposed tasks, and demonstrates its ability to effectively utilize world knowledge for reasoning about low-level human hand trajectories based on the provided context. Our website contains code and detailed video results https://www.chenbao.tech/handsonvlm/

📄 PDF Abstract BibTeX arXiv:2412.13187

Code (0)

등록된 구현이 없습니다.

Tasks

PredictionTrajectory PredictionWorld Knowledge

Similar Papers 제목 키워드 기반

HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision

2024-04-15 · Siddhant Bansal, Michael Wray, Dima Damen

Large Vision Language Models (VLMs) are now the de facto state-of-the-art for a number of tasks including visual question answering, recognising objects, and spatial referral. In this work, we propose the HOI-Ref task fo…

ObjectQuestion AnsweringVisual Question Answering

Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI

2026-08-19 · Mohammad Zamani, Fatemeh Ziaeetabar arxiv

Egocentric video captures activities from the wearer's perspective, providing a direct view of human attention, hand--object interaction, and goal-directed behavior. This perspective is increasingly important for wearabl…

Representation LearningDomain GeneralizationDecision Making

EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning

2026-01-27 · Binzhu Xie, Shi Qiu, Sicheng Zhang, Yinqiao Wang 외 arxiv

Robust 3D hand reconstruction in egocentric vision is challenging due to depth ambiguity, self-occlusion, and complex hand-object interactions. Prior methods mitigate these issues by scaling training data or adding auxil…

UniHM: Unified Dexterous Hand Manipulation with Vision Language Model

2026-02-28 · Zhenhao Zhang, Jiaxin Liu, Ye Shi, Jingya Wang arxiv

Planning physically feasible dexterous hand manipulation is a central challenge in robotic manipulation and Embodied AI. Prior work typically relies on object-centric cues or precise hand-object interaction sequences, fo…

FlowHOI: Flow-based Semantics-Grounded Generation of Hand-Object Interactions for Dexterous Robot Manipulation

2026-02-13 · Huajian Zeng, Lingyun Chen, Jiaqi Yang, Yuantai Zhang 외 arxiv

Recent vision-language-action (VLA) models can generate plausible end-effector motions, yet they often fail in long-horizon, contact-rich tasks because the underlying hand-object interaction (HOI) structure is not explic…

Robot ManipulationAction Recognition