paper-with-me

Papers

OmniHands: Towards Robust 4D Hand Mesh Recovery via A Versatile Transformer

2024-05-30 · Dixuan Lin, Yuxiang Zhang, Mengcheng Li, Yebin Liu, Wei Jing, Qi Yan, Qianying Wang, Hongwen Zhang

In this paper, we introduce OmniHands, a universal approach to recovering interactive hand meshes and their relative movement from monocular or multi-view inputs. Our approach addresses two major limitations of previous methods: lacking a unified solution for handling various hand image inputs and neglecting the positional relationship of two hands within images. To overcome these challenges, we develop a universal architecture with novel tokenization and contextual feature fusion strategies, capable of adapting to a variety of tasks. Specifically, we propose a Relation-aware Two-Hand Tokenization (RAT) method to embed positional relation information into the hand tokens. In this way, our network can handle both single-hand and two-hand inputs and explicitly leverage relative hand positions, facilitating the reconstruction of intricate hand interactions in real-world scenarios. As such tokenization indicates the relative relationship of two hands, it also supports more effective feature fusion. To this end, we further develop a 4D Interaction Reasoning (FIR) module to fuse hand tokens in 4D with attention and decode them into 3D hand meshes and relative temporal movements. The efficacy of our approach is validated on several benchmark datasets. The results on in-the-wild videos and real-world scenarios demonstrate the superior performances of our approach for interactive hand reconstruction. More video results can be found on the project page: https://OmniHand.github.io.

📄 PDF Abstract BibTeX arXiv:2405.20330

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DiffSurf: A Transformer-based Diffusion Model for Generating and Reconstructing 3D Surfaces in Pose

2024-08-27 · Yusuke Yoshiyasu, Leyuan Sun

This paper presents DiffSurf, a transformer-based denoising diffusion model for generating and reconstructing 3D surfaces. Specifically, we design a diffusion transformer architecture that predicts noise from noisy 3D su…

DenoisingDiversityHuman Mesh Recovery

TORE: Token Reduction for Efficient Human Mesh Recovery with Transformer

2022-11-19 · ICCV 2023 1 · Zhiyang Dou, Qingxuan Wu, Cheng Lin, Zeyu Cao 외

In this paper, we introduce a set of simple yet effective TOken REduction (TORE) strategies for Transformer-based Human Mesh Recovery from monocular images. Current SOTA performance is achieved by Transformer-based struc…

3D geometryHuman Mesh RecoveryToken Reduction

Camera-Space Hand Mesh Recovery via Semantic Aggregation and Adaptive 2D-1D Registration

2021-03-04 · CVPR 2021 1 · Xingyu Chen, Yufeng Liu, Chongyang Ma, Jianlong Chang 외

Recent years have witnessed significant progress in 3D hand mesh recovery. Nevertheless, because of the intrinsic 2D-to-3D ambiguity, recovering camera-space 3D information from a single RGB image remains challenging. To…

3D Hand Pose EstimationPosition

MMHMR: Generative Masked Modeling for Hand Mesh Recovery

2024-12-18 · Muhammad Usama Saleem, Ekkasit Pinyoanuntapong, Mayur Jagdishbhai Patel, Hongfei Xue 외

Reconstructing a 3D hand mesh from a single RGB image is challenging due to complex articulations, self-occlusions, and depth ambiguities. Traditional discriminative methods, which learn a deterministic mapping from a 2D…

3D Hand Pose Estimation

HuMMan: Multi-Modal 4D Human Dataset for Versatile Sensing and Modeling

2022-04-28 · Zhongang Cai, Daxuan Ren, Ailing Zeng, Zhengyu Lin 외

4D human sensing and modeling are fundamental tasks in vision and graphics with numerous applications. With the advances of new sensors and algorithms, there is an increasing demand for more versatile datasets. In this w…

Action RecognitionFine-grained Action RecognitionPose Estimation