paper-with-me

홈 › Papers

Multi-View Video-Based 3D Hand Pose Estimation

2021-09-24 · Leyla Khaleghi, Alireza Sepas Moghaddam, Joshua Marshall, Ali Etemad

Hand pose estimation (HPE) can be used for a variety of human-computer interaction applications such as gesture-based control for physical or virtual/augmented reality devices. Recent works have shown that videos or multi-view images carry rich information regarding the hand, allowing for the development of more robust HPE systems. In this paper, we present the Multi-View Video-Based 3D Hand (MuViHand) dataset, consisting of multi-view videos of the hand along with ground-truth 3D pose labels. Our dataset includes more than 402,000 synthetic hand images available in 4,560 videos. The videos have been simultaneously captured from six different angles with complex backgrounds and random levels of dynamic lighting. The data has been captured from 10 distinct animated subjects using 12 cameras in a semi-circle topology where six tracking cameras only focus on the hand and the other six fixed cameras capture the entire body. Next, we implement MuViHandNet, a neural pipeline consisting of image encoders for obtaining visual embeddings of the hand, recurrent learners to learn both temporal and angular sequential information, and graph networks with U-Net architectures to estimate the final 3D pose information. We perform extensive experiments and show the challenging nature of this new dataset as well as the effectiveness of our proposed method. Ablation studies show the added value of each component in MuViHandNet, as well as the benefit of having temporal and sequential information in the dataset.

📄 PDF Abstract BibTeX arXiv:2109.11747

Code (1)

leylakhaleghi/muvihand 공식 구현

Tasks

3D Hand Pose EstimationHand Pose EstimationPose Estimation

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
U-Net 설명 없음

Similar Papers 제목 키워드 기반

EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos

2026-01-03 · Hongming Fu, Wenjia Wang, Xiaozhen Qiao, Rolandos Alexandros Potamias 외 arxiv

We propose EgoGrasp, the first method to reconstruct world-space hand-object interactions (W-HOI) from dynamic egoview videos, supporting open-vocabulary objects. Accurate W-HOI reconstruction is critical for embodied in…

Hand Pose Estimation

Body2Hands: Learning to Infer 3D Hands from Conversational Gesture Body Dynamics

2020-07-23 · CVPR 2021 1 · Evonne Ng, Shiry Ginosar, Trevor Darrell, Hanbyul Joo

We propose a novel learned deep prior of body motion for 3D hand shape synthesis and estimation in the domain of conversational gestures. Our model builds upon the insight that body motion and hand gestures are strongly …

3D Hand Pose EstimationHand Pose EstimationPose Estimation

A Multi-scale Video Denoising Algorithm for Raw Image

2022-09-05 · Bin Ma, Yueli Hu, Xianxian Lv, Kai Li

Video denoising for raw image has always been the difficulty of camera image processing. On the one hand, image denoising performance largely determines the image quality, moreover denoising effect in raw image will affe…

DenoisingImage DenoisingMotion EstimationVideo Denoising

Reconstructing Hand-Held Objects from Monocular Video

2022-11-30 · Di Huang, Xiaopeng Ji, Xingyi He, Jiaming Sun 외

This paper presents an approach that reconstructs a hand-held object from a monocular video. In contrast to many recent methods that directly predict object geometry by a trained network, the proposed approach does not r…

Hand Pose EstimationObjectPose Estimation

3DFS: Deformable Dense Depth Fusion and Segmentation for Object Reconstruction from a Handheld Camera

2016-06-15 · Tanmay Gupta, Daeyun Shin, Naren Sivagnanadasan, Derek Hoiem

We propose an approach for 3D reconstruction and segmentation of a single object placed on a flat surface from an input video. Our approach is to perform dense depth map estimation for multiple views using a proposed obj…

3D ReconstructionDepth EstimationObjectObject Reconstruction+2