paper-with-me

홈 › Papers

Multi-view Hand Reconstruction with a Point-Embedded Transformer

2024-08-20 · Lixin Yang, Licheng Zhong, Pengxiang Zhu, Xinyu Zhan, Junxiao Kong, Jian Xu, Cewu Lu

This work introduces a novel and generalizable multi-view Hand Mesh Reconstruction (HMR) model, named POEM, designed for practical use in real-world hand motion capture scenarios. The advances of the POEM model consist of two main aspects. First, concerning the modeling of the problem, we propose embedding a static basis point within the multi-view stereo space. A point represents a natural form of 3D information and serves as an ideal medium for fusing features across different views, given its varied projections across these views. Consequently, our method harnesses a simple yet effective idea: a complex 3D hand mesh can be represented by a set of 3D basis points that 1) are embedded in the multi-view stereo, 2) carry features from the multi-view images, and 3) encompass the hand in it. The second advance lies in the training strategy. We utilize a combination of five large-scale multi-view datasets and employ randomization in the number, order, and poses of the cameras. By processing such a vast amount of data and a diverse array of camera configurations, our model demonstrates notable generalizability in the real-world applications. As a result, POEM presents a highly practical, plug-and-play solution that enables user-friendly, cost-effective multi-view motion capture for both left and right hands. The model and source codes are available at https://github.com/JubSteven/POEM-v2.

📄 PDF Abstract BibTeX arXiv:2408.10581

Code (1)

jubsteven/poem-v2 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

POEM: Reconstructing Hand in a Point Embedded Multi-view Stereo

2023-04-08 · CVPR 2023 1 · Lixin Yang, Jian Xu, Licheng Zhong, Xinyu Zhan 외

Enable neural networks to capture 3D geometrical-aware features is essential in multi-view based vision tasks. Previous methods usually encode the 3D information of multi-view stereo into the 2D features. In contrast, we…

DiffS-NOCS: 3D Point Cloud Reconstruction through Coloring Sketches to NOCS Maps Using Diffusion Models

2025-06-15 · Di Kong, Qianhui Wan

Reconstructing a 3D point cloud from a given conditional sketch is challenging. Existing methods often work directly in 3D space, but domain variability and difficulty in reconstructing accurate 3D structures from 2D ske…

3D Point Cloud ReconstructionDecoderDenoisingPoint cloud reconstruction

HOSt3R: Keypoint-free Hand-Object 3D Reconstruction from RGB images

2025-08-22 · Anilkumar Swamy, Vincent Leroy, Philippe Weinzaepfel, Jean-Sébastien Franco 외 arxiv

Hand-object 3D reconstruction has become increasingly important for applications in human-robot interaction and immersive AR/VR experiences. A common approach for object-agnostic hand-object reconstruction from RGB seque…

Multi-View 3D ReconstructionKeypoint Detection

Local and Global Point Cloud Reconstruction for 3D Hand Pose Estimation

2021-12-13 · Ziwei Yu, Linlin Yang, Shicheng Chen, Angela Yao

This paper addresses the 3D point cloud reconstruction and 3D pose estimation of the human hand from a single RGB image. To that end, we present a novel pipeline for local and global point cloud reconstruction using a 3D…

3D Hand Pose Estimation3D Point Cloud Reconstruction3D Pose EstimationHand Pose Estimation+2

LaVR: Scene Latent Conditioned Generative Video Trajectory Re-Rendering using Large 4D Reconstruction Models

2026-01-21 · Mingyang Xie, Numair Khan, Tianfu Wang, Naina Dhingra 외 arxiv

Given a monocular video, the goal of video re-rendering is to generate views of the scene from a novel camera trajectory. Existing methods face two distinct challenges. Geometrically unconditioned models lack spatial awa…

Video Generation