paper-with-me

Papers

Cross-Language Transfer Learning using Visual Information for Automatic Sign Gesture Recognition

2023-05-12 · The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences 2023 5 · Dmitry Ryumin, Denis Ivanko, Alexandr Axyonov

Automatic sign gesture recognition (GR) plays a critical role in facilitating communication between hearing-impaired individuals and the rest of society. However, recognizing sign gestures accurately and efficiently remains a challenging task due to the diversity of sign languages (SLs) and their limited availability of labeled data. This scientific paper proposes a new approach to improving the accuracy of automatic sign GR using cross-language transfer learning with visual information. Two large-scale multimodal SL corpora are utilized as the basic SLs for this study: the Ankara University Turkish Sign Language Dataset (AUTSL) and the Thesaurus Russian Sign Language (TheRusLan). Experimental studies were conducted, resulting in an accuracy of 93.33% for 18 different gestures, including the Russian target SL gestures. This result exceeds the previous state-of-the-art accuracy by 2.19%, demonstrating the effectiveness of the proposed approach. The study highlights the potential of the proposed approach to enhance the accuracy and robustness of machine SL translation, improve the naturalness of human-computer interaction, and facilitate the social adaptation of people with hearing impairments. This paper proposes a promising direction for future research to explore the application of the proposed approach to other SLs and to investigate the impact of individual and cultural differences on GR.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Gesture RecognitionSign Language RecognitionTransfer Learning

Similar Papers 제목 키워드 기반

Leveraging Visual Knowledge in Language Tasks: An Empirical Study on Intermediate Pre-training for Cross-modal Knowledge Transfer

2022-03-14 · ACL 2022 5 · Woojeong Jin, Dong-Ho Lee, Chenguang Zhu, Jay Pujara 외

Pre-trained language models are still far from human performance in tasks that need understanding of properties (e.g. appearance, measurable quantity) and affordances of everyday objects in the real world since the text …

Image CaptioningLanguage ModelingLanguage ModellingTransfer Learning

Leveraging Visual Knowledge in Language Tasks: An Empirical Study on Intermediate Pre-training for Cross-Modal Knowledge Transfer

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Pre-trained language models are still far from human performance in tasks that need understanding of properties (e.g. appearance, measurable quantity) and affordances of everyday objects in the real world since the text …

Image CaptioningLanguage ModelingLanguage ModellingTransfer Learning

SWIFT Aligner, A Multifunctional Tool for Parallel Corpora: Visualization, Word Alignment, and (Morpho)-Syntactic Cross-Language Transfer

2014-05-01 · LREC 2014 5 · Timur Gilmanov, Olga Scrivner, S K{\"u}bler, ra

It is well known that word aligned parallel corpora are valuable linguistic resources. Since many factors affect automatic alignment quality, manual post-editing may be required in some applications. While there are seve…

Machine TranslationWord AlignmentWord Sense Disambiguation

Cross-lingual Transfer of Abstractive Summarizer to Less-resource Language

2020-12-08 · Aleš Žagar, Marko Robnik-Šikonja

Automatic text summarization extracts important information from texts and presents the information in the form of a summary. Abstractive summarization approaches progressed significantly by switching to deep neural netw…

Abstractive Text SummarizationArticlesCross-Lingual TransferDecoder+3

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

2019-08-06 · NeurIPS 2019 12 · Jiasen Lu, Dhruv Batra, Devi Parikh, Stefan Lee

We present ViLBERT (short for Vision-and-Language BERT), a model for learning task-agnostic joint representations of image content and natural language. We extend the popular BERT architecture to a multi-modal two-stream…

Image RetrievalQuestion AnsweringReferring Expression ComprehensionRetrieval+5