Zero-Shot Sign Language Recognition: Can Textual Data Uncover Sign Languages?
We introduce the problem of zero-shot sign language recognition (ZSSLR), where the goal is to leverage models learned over the seen sign class examples to recognize the instances of unseen signs. To this end, we propose to utilize the readily available descriptions in sign language dictionaries as an intermediate-level semantic representation for knowledge transfer. We introduce a new benchmark dataset called ASL-Text that consists of 250 sign language classes and their accompanying textual descriptions. Compared to the ZSL datasets in other domains (such as object recognition), our dataset consists of limited number of training examples for a large number of classes, which imposes a significant challenge. We propose a framework that operates over the body and hand regions by means of 3D-CNNs, and models longer temporal relationships via bidirectional LSTMs. By leveraging the descriptive text embeddings along with these spatio-temporal representations within a zero-shot learning framework, we show that textual data can indeed be useful in uncovering sign languages. We anticipate that the introduced approach and the accompanying dataset will provide a basis for further exploration of this new zero-shot learning problem.
Code (0)
등록된 구현이 없습니다.
Tasks
DescriptiveObject RecognitionSign Language RecognitionTransfer LearningZero-Shot LearningSimilar Papers 제목 키워드 기반
Towards Zero-shot Sign Language Recognition
This paper tackles the problem of zero-shot sign language recognition (ZSSLR), where the goal is to leverage models learned over the seen sign classes to recognize the instances of unseen sign classes. In this context, r…
AttributeDescriptiveSign Language RecognitionTransfer Learning+1Text-Enhanced Zero-Shot Action Recognition: A training-free approach
Vision-language models (VLMs) have demonstrated remarkable performance across various visual tasks, leveraging joint learning of visual and textual representations. While these models excel in zero-shot image tasks, thei…
Action RecognitionTemporal Action LocalizationZero-Shot Action RecognitionMulti-Modal Zero-Shot Sign Language Recognition
Zero-Shot Learning (ZSL) has rapidly advanced in recent years. Towards overcoming the annotation bottleneck in the Sign Language Recognition (SLR), we explore the idea of Zero-Shot Sign Language Recognition (ZS-SLR) with…
Hand DetectionSign Language RecognitionZero-Shot LearningTelling Stories for Common Sense Zero-Shot Action Recognition
Video understanding has long suffered from reliance on large labeled datasets, motivating research into zero-shot learning. Recent progress in language modeling presents opportunities to advance zero-shot video analysis,…
Action RecognitionArticlesCommon Sense ReasoningLanguage Modeling+6Language Models as Zero-shot Visual Semantic Learners
Visual Semantic Embedding (VSE) models, which map images into a rich semantic embedding space, have been a milestone in object recognition and zero-shot learning. Current approaches to VSE heavily rely on static word em-…
ObjectObject RecognitionWord EmbeddingsZero-Shot Learning