paper-with-me

홈 › Papers

IPN Hand: A Video Dataset and Benchmark for Real-Time Continuous Hand Gesture Recognition

2020-04-20 · Gibran Benitez-Garcia, Jesus Olivares-Mercado, Gabriel Sanchez-Perez, Keiji Yanai

In this paper, we introduce a new benchmark dataset named IPN Hand with sufficient size, variety, and real-world elements able to train and evaluate deep neural networks. This dataset contains more than 4,000 gesture samples and 800,000 RGB frames from 50 distinct subjects. We design 13 different static and dynamic gestures focused on interaction with touchless screens. We especially consider the scenario when continuous gestures are performed without transition states, and when subjects perform natural movements with their hands as non-gesture actions. Gestures were collected from about 30 diverse scenes, with real-world variation in background and illumination. With our dataset, the performance of three 3D-CNN models is evaluated on the tasks of isolated and continuous real-time HGR. Furthermore, we analyze the possibility of increasing the recognition accuracy by adding multiple modalities derived from RGB frames, i.e., optical flow and semantic segmentation, while keeping the real-time performance of the 3D-CNN model. Our empirical study also provides a comparison with the publicly available nvGesture (NVIDIA) dataset. The experimental results show that the state-of-the-art ResNext-101 model decreases about 30% accuracy when using our real-world dataset, demonstrating that the IPN Hand dataset can be used as a benchmark, and may help the community to step forward in the continuous HGR. Our dataset and pre-trained models used in the evaluation are publicly available at https://github.com/GibranBenitez/IPN-hand.

📄 PDF Abstract BibTeX arXiv:2005.02134

Code (1)

GibranBenitez/IPN-hand 공식 구현 pytorch

Tasks

Gesture RecognitionHand Gesture RecognitionHand-Gesture RecognitionOptical Flow EstimationSemantic Segmentation

Similar Papers 제목 키워드 기반

EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing

2025-12-05 · Runjia Li, Moayed Haji-Ali, Ashkan Mirzaei, Chaoyang Wang 외 arxiv

We study instruction-guided editing of egocentric videos for interactive AR applications. While recent AI video editors perform well on third-person footage, egocentric views present unique challenges - including rapid e…

Real-time Action Recognition for Fine-Grained Actions and The Hand Wash Dataset

2022-10-13 · Akash Nagaraj, Mukund Sood, Chetna Sureka, Gowri Srinivasa

In this paper we present a three-stream algorithm for real-time action recognition and a new dataset of handwash videos, with the intent of aligning action recognition with real-world constraints to yield effective concl…

Action RecognitionFine-grained Action RecognitionOptical Flow Estimation

Benchmark Dataset and Effective Inter-Frame Alignment for Real-World Video Super-Resolution

2022-12-10 · Ruohao Wang, Xiaohui Liu, Zhilu Zhang, Xiaohe Wu 외

Video super-resolution (VSR) aiming to reconstruct a high-resolution (HR) video from its low-resolution (LR) counterpart has made tremendous progress in recent years. However, it remains challenging to deploy existing VS…

Optical Flow EstimationSuper-ResolutionVideo Super-Resolution

WildGHand: Learning Anti-Perturbation Gaussian Hand Avatars from Monocular In-the-Wild Videos

2026-02-24 · Hanhui Li, Xuan Huang, Wanquan Liu, Yuhao Cheng 외 arxiv

Despite recent progress in 3D hand reconstruction from monocular videos, most existing methods rely on data captured in well-controlled environments and therefore degrade in real-world settings with severe perturbations,…

Robust Promptable Video Object Segmentation

2026-05-12 · Sohyun Lee, Yeho Gwon, Lukas Hoyer, Konrad Schindler 외 arxiv

The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deployment in safety-critical domains. This paper offers the first comprehensive s…

Video Object Segmentation