Multitask training with unlabeled data for end-to-end sign language fingerspelling recognition
We address the problem of automatic American Sign Language fingerspelling recognition from video. Prior work has largely relied on frame-level labels, hand-crafted features, or other constraints, and has been hampered by the scarcity of data for this task. We introduce a model for fingerspelling recognition that addresses these issues. The model consists of an auto-encoder-based feature extractor and an attention-based neural encoder-decoder, which are trained jointly. The model receives a sequence of image frames and outputs the fingerspelled word, without relying on any frame-level training labels or hand-crafted features. In addition, the auto-encoder subcomponent makes it possible to leverage unlabeled data to improve the feature learning. The model achieves 11.6% and 4.4% absolute letter accuracy improvement respectively in signer-independent and signer-adapted fingerspelling recognition over previous approaches that required frame-level training labels.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Unsupervised Online Multitask Learning of Behavioral Sentence Embeddings
Unsupervised learning has been an attractive method for easily deriving meaningful data representations from vast amounts of unlabeled data. These representations, or embeddings, often yield superior results in many task…
Domain AdaptationSentenceSentence EmbeddingsLabel-efficient audio classification through multitask learning and self-supervision
While deep learning has been incredibly successful in modeling tasks with large, carefully curated labeled datasets, its application to problems with limited labeled data remains a challenge. The aim of the present work …
Audio ClassificationClassificationData AugmentationGeneral Classification+1Zemi: Learning Zero-Shot Semi-Parametric Language Models from Multiple Tasks
Although large language models have achieved impressive zero-shot ability, the huge model size generally incurs high cost. Recently, semi-parametric language models, which augment a smaller language model with an externa…
Language ModelingLanguage ModellingRetrievalText Augmentation+1Improving label efficiency through multi-task learning on auditory data
Collecting high-quality, large scale datasets typically requires significant resources. The aim of the present work is to improve the label efficiency of large neural networks operating on audio data through multitask le…
Data AugmentationMulti-Task LearningSelf-Supervised LearningMultiMix: Sparingly Supervised, Extreme Multitask Learning From Medical Images
Semi-supervised learning via learning from limited quantities of labeled data has been investigated as an alternative to supervised counterparts. Maximizing knowledge gains from copious unlabeled data benefit semi-superv…
General ClassificationSegmentation