paper-with-me

Papers

POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-View World

2024-03-09 · Boshen Xu, Sipeng Zheng, Qin Jin

We humans are good at translating third-person observations of hand-object interactions (HOI) into an egocentric view. However, current methods struggle to replicate this ability of view adaptation from third-person to first-person. Although some approaches attempt to learn view-agnostic representation from large-scale video datasets, they ignore the relationships among multiple third-person views. To this end, we propose a Prompt-Oriented View-agnostic learning (POV) framework in this paper, which enables this view adaptation with few egocentric videos. Specifically, We introduce interactive masking prompts at the frame level to capture fine-grained action information, and view-aware prompts at the token level to learn view-agnostic representation. To verify our method, we establish two benchmarks for transferring from multiple third-person views to the egocentric view. Our extensive experiments on these benchmarks demonstrate the efficiency and effectiveness of our POV framework and prompt tuning techniques in terms of view adaptation and view generalization. Our code is available at \url{https://github.com/xuboshen/pov_acmmm2023}.

📄 PDF Abstract BibTeX arXiv:2403.05856

Code (1)

xuboshen/pov_acmmm2023 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory Forecasting

2023-07-17 · ICCV 2023 1 · Wentao Bao, Lele Chen, Libing Zeng, Zhong Li 외

Hand trajectory forecasting from egocentric views is crucial for enabling a prompt understanding of human intentions when interacting with AR/VR systems. However, existing methods handle this problem in a 2D image space …

3D Human Pose TrackingTrajectory ForecastingTrajectory PredictionVisual Prompt Tuning

Opening the Vocabulary of Egocentric Actions

2023-08-22 · NeurIPS 2023 11 · Dibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela Yao

Human actions in egocentric videos are often hand-object interactions composed from a verb (performed by the hand) applied to an object. Despite their extensive scaling up, egocentric datasets still face two limitations …

Action RecognitionObjectOpen Vocabulary Action Recognition

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting

2026-05-08 · Jaeyoung Choi, Hyeondong Kim, Yujin Kim, Daehee Park arxiv

Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such as AR/VR assistance and human-robot interaction. However, this task r…

Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects

2024-03-25 · Zicong Fan, Takehiko Ohkawa, Linlin Yang, Nie Lin 외

We interact with the world with our hands and see it through our own (egocentric) perspective. A holistic 3Dunderstanding of such interactions from egocentric views is important for tasks in robotics, AR/VR, action recog…

Action RecognitionMotion GenerationObjectObject Reconstruction+1

1st Place Solution of Egocentric 3D Hand Pose Estimation Challenge 2023 Technical Report:A Concise Pipeline for Egocentric Hand Pose Reconstruction

2023-10-07 · Zhishan Zhou, Zhi Lv, Shihao Zhou, Minqiang Zou 외

This report introduce our work on Egocentric 3D Hand Pose Estimation workshop. Using AssemblyHands, this challenge focuses on egocentric 3D hand pose estimation from a single-view image. In the competition, we adopt ViT …

3D Hand Pose EstimationHand Pose EstimationPose Estimation