Improving Sign Recognition with Phonology
We use insights from research on American Sign Language (ASL) phonology to train models for isolated sign language recognition (ISLR), a step towards automatic sign language understanding. Our key insight is to explicitly recognize the role of phonology in sign production to achieve more accurate ISLR than existing work which does not consider sign language phonology. We train ISLR models that take in pose estimations of a signer producing a single sign to predict not only the sign but additionally its phonological characteristics, such as the handshape. These auxiliary predictions lead to a nearly 9% absolute gain in sign recognition accuracy on the WLASL benchmark, with consistent improvements in ISLR regardless of the underlying prediction model architecture. This work has the potential to accelerate linguistic research in the domain of signed languages and reduce communication barriers between deaf and hearing people.
Code (2)
Tasks
Sign Language RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Improve Bilingual TTS Using Dynamic Language and Phonology Embedding
In most cases, bilingual TTS needs to handle three types of input scripts: first language only, second language only, and second language embedded in the first language. In the latter two situations, the pronunciation an…
Multilingual and crosslingual speech recognition using phonological-vector based phone embeddings
The use of phonological features (PFs) potentially allows language-specific phones to remain linked in training, which is highly desirable for information sharing for multilingual and crosslingual speech recognition meth…
speech-recognitionSpeech RecognitionModeling Markedness with a Split-and-Merger Model of Sound Change
The concept of {`}markedness{'} has been influential in phonology for almost a century. Theoretical phonology has found it useful to describe some segments as more {`}marked{'} than others, referring to a cluster of lang…
PhonologyBench: Evaluating Phonological Skills of Large Language Models
Phonology, the study of speech's structure and pronunciation rules, is a critical yet often overlooked component in Large Language Model (LLM) research. LLMs are widely used in various downstream applications that levera…
DiagnosticGrapheme-to-Phoneme ConversionLanguage ModelingLanguage Modelling+1Massively Multilingual Adversarial Speech Recognition
We report on adaptation of multilingual end-to-end speech recognition models trained on as many as 100 languages. Our findings shed light on the relative importance of similarity between the target and pretraining langua…
General Classificationspeech-recognitionSpeech Recognition