paper-with-me

홈 › Papers

Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics

2025-01-13 · Tze Ho Elden Tse, Runyang Feng, Linfang Zheng, Jiho Park, Yixing Gao, Jihie Kim, Ales Leonardis, Hyung Jin Chang

With the availability of egocentric 3D hand-object interaction datasets, there is increasing interest in developing unified models for hand-object pose estimation and action recognition. However, existing methods still struggle to recognise seen actions on unseen objects due to the limitations in representing object shape and movement using 3D bounding boxes. Additionally, the reliance on object templates at test time limits their generalisability to unseen objects. To address these challenges, we propose to leverage superquadrics as an alternative 3D object representation to bounding boxes and demonstrate their effectiveness on both template-free object reconstruction and action recognition tasks. Moreover, as we find that pure appearance-based methods can outperform the unified methods, the potential benefits from 3D geometric information remain unclear. Therefore, we study the compositionality of actions by considering a more challenging task where the training combinations of verbs and nouns do not overlap with the testing split. We extend H2O and FPHA datasets with compositional splits and design a novel collaborative learning framework that can explicitly reason about the geometric relations between hands and the manipulated object. Through extensive quantitative and qualitative evaluations, we demonstrate significant improvements over the state-of-the-arts in (compositional) action recognition.

📄 PDF Abstract BibTeX arXiv:2501.07100

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognitionhand-object poseObjectObject ReconstructionPose Estimation

Similar Papers 제목 키워드 기반

HOLD: Category-agnostic 3D Reconstruction of Interacting Hands and Objects from Video

2023-11-30 · CVPR 2024 1 · Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Muhammed Kocabas 외

Since humans interact with diverse objects every day, the holistic 3D capture of these interactions is important to understand and model human behaviour. However, most existing methods for hand-object reconstruction from…

3D ReconstructionObjectObject Reconstruction

Collaborative Learning for Hand and Object Reconstruction with Attention-guided Graph Convolution

2022-04-27 · CVPR 2022 1 · Tze Ho Elden Tse, Kwang In Kim, Ales Leonardis, Hyung Jin Chang

Estimating the pose and shape of hands and objects under interaction finds numerous applications including augmented and virtual reality. Existing approaches for hand and object reconstruction require explicitly defined …

3D Hand Pose Estimation3D Pose Estimationhand-object poseObject+2

Learning Partonomic 3D Reconstruction from Image Collections

2025-01-01 · CVPR 2025 1 · Xiaoqian Ruan, Pei Yu, Dian Jia, Hyeonjeong Park 외

Reconstructing the 3D shape of an object from a single-view image is a fundamental task in computer vision. Recent advances in differentiable rendering have enabled 3D reconstruction from image collections using only…

3D ReconstructionImage GenerationObjectObject Reconstruction

Fine-Grained Egocentric Hand-Object Segmentation: Dataset, Model, and Applications

2022-08-07 · Lingzhi Zhang, Shenghao Zhou, Simon Stent, Jianbo Shi

Egocentric videos offer fine-grained information for high-fidelity modeling of human behaviors. Hands and interacting objects are one crucial aspect of understanding a viewer's behaviors and intentions. We provide a labe…

Activity RecognitionData AugmentationObjectSegmentation+2

Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos

2026-07-12 · Hao Zheng, Jinyi Huang, Tiantian Zheng, Xun Xu 외 arxiv

Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand-object interactions. We adapt vision-language models…

Hyperparameter OptimizationAction UnderstandingMulti-Task LearningAction Recognition