paper-with-me

Papers

Word-level Sign Language Recognition with Multi-stream Neural Networks Focusing on Local Regions and Skeletal Information

2021-06-30 · Mizuki Maruyama, Shrey Singh, Katsufumi Inoue, Partha Pratim Roy, Masakazu Iwamura, Michifumi Yoshioka

Word-level sign language recognition (WSLR) has attracted attention because it is expected to overcome the communication barrier between people with speech impairment and those who can hear. In the WSLR problem, a method designed for action recognition has achieved the state-of-the-art accuracy. Indeed, it sounds reasonable for an action recognition method to perform well on WSLR because sign language is regarded as an action. However, a careful evaluation of the tasks reveals that the tasks of action recognition and WSLR are inherently different. Hence, in this paper, we propose a novel WSLR method that takes into account information specifically useful for the WSLR problem. We realize it as a multi-stream neural network (MSNN), which consist of three streams: 1) base stream, 2) local image stream, and 3) skeleton stream. Each stream is designed to handle different types of information. The base stream deals with quick and detailed movements of the hands and body, the local image stream focuses on handshapes and facial expressions, and the skeleton stream captures the relative positions of the body and both hands. This approach allows us to combine various types of data for more comprehensive gesture analysis. Experimental results on the WLASL and MS-ASL datasets show the effectiveness of the proposed method; it achieved an improvement of approximately 10\%--15\% in Top-1 accuracy when compared with conventional methods.

📄 PDF Abstract BibTeX arXiv:2106.15989

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionSign Language Recognition

Similar Papers 제목 키워드 기반

Improving Mandarin End-to-End Speech Recognition with Word N-gram Language Model

2022-01-06 · Jinchuan Tian, Jianwei Yu, Chao Weng, Yuexian Zou 외

Despite the rapid progress of end-to-end (E2E) automatic speech recognition (ASR), it has been shown that incorporating external language models (LMs) into the decoding can further improve the recognition performance of …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

Learning Multi-level Dependencies for Robust Word Recognition

2019-11-22 · Zhiwei Wang, Hui Liu, Jiliang Tang, Songfan Yang 외

Robust language processing systems are becoming increasingly important given the recent awareness of dangerous situations where brittle machine learning models can be easily broken with the presence of noises. In this pa…

Hierarchical Meta-Embeddings for Code-Switching Named Entity Recognition

2019-09-18 · IJCNLP 2019 11 · Genta Indra Winata, Zhaojiang Lin, Jamin Shin, Zihan Liu 외

In countries that speak multiple main languages, mixing up different languages within a conversation is commonly called code-switching. Previous works addressing this challenge mainly focused on word-level aspects such a…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Word Embeddings

Named Entity Recognition in Multi-level Contexts

2020-12-01 · Asian Chapter of the Association for Computational Linguistics 2020 · Yubo Chen, Chuhan Wu, Tao Qi, Zhigang Yuan 외

Named entity recognition is a critical task in the natural language processing field. Most existing methods for this task can only exploit contextual information within a sentence. However, their performance on recognizi…

Multi-Task Learningnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models

2022-10-13 · Haoyu Wang, Wei-Qiang Zhang, Hongbin Suo, Yulong Wan

Labeled audio data is insufficient to build satisfying speech recognition systems for most of the languages in the world. There have been some zero-resource methods trying to perform phoneme or word-level speech recognit…

Language ModelingLanguage ModellingPhoneme Recognitionspeech-recognition+1