paper-with-me

홈 › Papers

Deep Hand: How to Train a CNN on 1 Million Hand Images When Your Data Is Continuous and Weakly Labelled

2016-06-01 · CVPR 2016 6 · Oscar Koller, Hermann Ney, Richard Bowden

This work presents a new approach to learning a frame-based classifier on weakly labelled sequence data by embedding a CNN within an iterative EM algorithm. This allows the CNN to be trained on a vast number of example images when only loose sequence level information is available for the source videos. Although we demonstrate this in the context of hand shape recognition, the approach has wider application to any video recognition task where frame level labelling is not available. The iterative EM algorithm leverages the discriminative ability of the CNN to iteratively refine the frame level annotation and subsequent training of the CNN. By embedding the classifier within an EM framework the CNN can easily be trained on 1 million hand images. We demonstrate that the final classifier generalises over both individuals and data sets. The algorithm is evaluated on over 3000 manually labelled hand shape images of 60 different classes which will be released to the community. Furthermore, we demonstrate its use in continuous sign language recognition on two publicly available large sign language data sets, where it outperforms the current state-of-the-art by a large margin. To our knowledge no previous work has explored expectation maximization without Gaussian mixture models to exploit weak sequence labels for sign language recognition.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sign Language RecognitionVideo Recognition

Similar Papers 제목 키워드 기반

LS-HDIB: A Large Scale Handwritten Document Image Binarization Dataset

2021-01-27 · Kaustubh Sadekar, Ashish Tiwari, Prajwal Singh, Shanmuganathan Raman

Handwritten document image binarization is challenging due to high variability in the written content and complex background attributes such as page style, paper quality, stains, shadow gradients, and non-uniform illumin…

BinarizationDiversity

AnimeDL-2M: Million-Scale AI-Generated Anime Image Detection and Localization in Diffusion Era

2025-04-15 · Chenyang Zhu, Xing Zhang, Yuyang Sun, Ching-Chun Chang 외

Recent advances in image generation, particularly diffusion models, have significantly lowered the barrier for creating sophisticated forgeries, making image manipulation detection and localization (IMDL) increasingly ch…

Image GenerationImage ManipulationImage Manipulation Detection

DARE: A large-scale handwritten date recognition system

2022-10-02 · Christian M. Dahl, Torben S. D. Johansen, Emil N. Sørensen, Christian E. Westermann 외

Handwritten text recognition for historical documents is an important task but it remains difficult due to a lack of sufficient training data in combination with a large variability of writing styles and degradation of h…

Handwritten Text RecognitionTransfer Learning

MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts

2024-10-30 · Jie Zhu, Yixiong Chen, Mingyu Ding, Ping Luo 외

Text-to-image diffusion has attracted vast attention due to its impressive image-generation capabilities. However, when it comes to human-centric text-to-image generation, particularly in the context of faces and hands, …

Image GenerationText to Image GenerationText-to-Image Generation

DeepHPS: End-to-end Estimation of 3D Hand Pose and Shape by Learning from Synthetic Depth

2018-08-28 · Jameel Malik, Ahmed Elhayek, Fabrizio Nunnari, Kiran varanasi 외

Articulated hand pose and shape estimation is an important problem for vision-based applications such as augmented reality and animation. In contrast to the existing methods which optimize only for joint positions, we pr…