paper-with-me

Papers

Hand-Object Interaction Reasoning

2022-01-13 · Jian Ma, Dima Damen

This paper proposes an interaction reasoning network for modelling spatio-temporal relationships between hands and objects in video. The proposed interaction unit utilises a Transformer module to reason about each acting hand, and its spatio-temporal relation to the other hand as well as objects being interacted with. We show that modelling two-handed interactions are critical for action recognition in egocentric video, and demonstrate that by using positionally-encoded trajectories, the network can better recognise observed interactions. We evaluate our proposal on EPIC-KITCHENS and Something-Else datasets, with an ablation study.

📄 PDF Abstract BibTeX arXiv:2201.04906

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionObject

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Position-Wise Feed-Forward Layer 설명 없음
Multi-Head Attention 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…

Similar Papers 제목 키워드 기반

HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching

2026-04-12 · Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, Jiankang Deng 외 arxiv

Generating realistic 3D hand-object interactions (HOI) is a fundamental challenge in computer vision and robotics, requiring both temporal coherence and high-fidelity physical plausibility. Existing methods remain limite…

Motion Synthesis

EgoPHI: Estimating Contact and Force from Egocentric Vision

2026-08-13 · Andela Ilic, Rachel Schuchert, Yijing Jiang, Christian Holz arxiv

Understanding hand-object interaction from egocentric vision is essential for modeling how people physically engage with the surrounding world. Yet reasoning about physically grounded interaction requires estimating the …

GenHOI: Generalized Hand-Object Pose Estimation with Occlusion Awareness

2026-03-19 · Hui Yang, Wei Sun, Jian Liu, Jian Xiao 외 arxiv

Generalized 3D hand-object pose estimation from a single RGB image remains challenging due to the large variations in object appearances and interaction patterns, especially under heavy occlusion. We propose GenHOI, a fr…

hand-object posePose EstimationPoint Clouds

TOCH: Spatio-Temporal Object-to-Hand Correspondence for Motion Refinement

2022-05-16 · Keyang Zhou, Bharat Lal Bhatnagar, Jan Eric Lenssen, Gerard Pons-Moll

We present TOCH, a method for refining incorrect 3D hand-object interaction sequences using a data prior. Existing hand trackers, especially those that rely on very few cameras, often produce visually unrealistic results…

DenoisingObjectObject Reconstruction

HOReeNet: 3D-aware Hand-Object Grasping Reenactment

2022-11-11 · Changhwa Lee, Junuk Cha, Hansol Lee, Seongyeong Lee 외

We present HOReeNet, which tackles the novel task of manipulating images involving hands, objects, and their interactions. Especially, we are interested in transferring objects of source images to target images and manip…

3D ReconstructionObject