paper-with-me

Papers

Reconstructing Hand-Held Objects in 3D from Images and Videos

2024-04-09 · Jane Wu, Georgios Pavlakos, Georgia Gkioxari, Jitendra Malik

Objects manipulated by the hand (i.e., manipulanda) are particularly challenging to reconstruct from Internet videos. Not only does the hand occlude much of the object, but also the object is often only visible in a small number of image pixels. At the same time, two strong anchors emerge in this setting: (1) estimated 3D hands help disambiguate the location and scale of the object, and (2) the set of manipulanda is small relative to all possible objects. With these insights in mind, we present a scalable paradigm for hand-held object reconstruction that builds on recent breakthroughs in large language/vision models and 3D object datasets. Given a monocular RGB video, we aim to reconstruct hand-held object geometry in 3D, over time. In order to obtain the best performing single frame model, we first present MCC-Hand-Object (MCC-HO), which jointly reconstructs hand and object geometry given a single RGB image and inferred 3D hand as inputs. Subsequently, we prompt a text-to-3D generative model using GPT-4(V) to retrieve a 3D object model that matches the object in the image(s); we call this alignment Retrieval-Augmented Reconstruction (RAR). RAR provides unified object geometry across all frames, and the result is rigidly aligned with both the input images and 3D MCC-HO observations in a temporally consistent manner. Experiments demonstrate that our approach achieves state-of-the-art performance on lab and Internet image/video datasets. We make our code and models available on the project website: https://janehwu.github.io/mcc-ho

📄 PDF Abstract BibTeX arXiv:2404.06507

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectObject ReconstructionText to 3D

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

3D Reconstruction of Objects in Hands without Real World 3D Supervision

2023-05-04 · Aditya Prakash, Matthew Chang, Matthew Jin, Ruisen Tu 외

Prior works for reconstructing hand-held objects from a single image train models on images paired with 3D shapes. Such data is challenging to gather in the real world at scale. Consequently, these approaches do not gene…

3D ReconstructionObjectObject Reconstruction

HORT: Monocular Hand-held Objects Reconstruction with Transformers

2025-03-27 · Zerui Chen, Rolandos Alexandros Potamias, ShiZhe Chen, Cordelia Schmid

Reconstructing hand-held objects in 3D from monocular images remains a significant challenge in computer vision. Most existing approaches rely on implicit 3D representations, which produce overly smooth reconstructions a…

Denoising

Learning Explicit Contact for Implicit Reconstruction of Hand-held Objects from Monocular Images

2023-05-31 · Junxing Hu, Hongwen Zhang, Zerui Chen, Mengcheng Li 외

Reconstructing hand-held objects from monocular RGB images is an appealing yet challenging task. In this task, contacts between hands and objects provide important cues for recovering the 3D geometry of the hand-held obj…

3D geometryObject

Reconstructing Objects along Hand Interaction Timelines in Egocentric Video

2025-12-08 · Zhifan Zhu, Siddhant Bansal, Shashank Tripathi, Dima Damen arxiv

We introduce the task of Reconstructing Objects along Hand Interaction Timelines (ROHIT). We first define the Hand Interaction Timeline (HIT) from a rigid object's perspective. In a HIT, an object is first static relativ…

HOIST-Former: Hand-held Objects Identification Segmentation and Tracking in the Wild

2024-01-01 · CVPR 2024 1 · Supreeth Narasimhaswamy, Huy Anh Nguyen, Lihan Huang, Minh Hoai

We address the challenging task of identifying segmenting and tracking hand-held objects which is crucial for applications such as human action segmentation and performance evaluation. This task is particularly chall…

Action SegmentationSegmentation