paper-with-me

홈 › Papers

The Sem-Lex Benchmark: Modeling ASL Signs and Their Phonemes

2023-09-30 · Lee Kezar, Elana Pontecorvo, Adele Daniels, Connor Baer, Ruth Ferster, Lauren Berger, Jesse Thomason, Zed Sevcikova Sehyr, Naomi Caselli

Sign language recognition and translation technologies have the potential to increase access and inclusion of deaf signing communities, but research progress is bottlenecked by a lack of representative data. We introduce a new resource for American Sign Language (ASL) modeling, the Sem-Lex Benchmark. The Benchmark is the current largest of its kind, consisting of over 84k videos of isolated sign productions from deaf ASL signers who gave informed consent and received compensation. Human experts aligned these videos with other sign language resources including ASL-LEX, SignBank, and ASL Citizen, enabling useful expansions for sign and phonological feature recognition. We present a suite of experiments which make use of the linguistic information in ASL-LEX, evaluating the practicality and fairness of the Sem-Lex Benchmark for isolated sign recognition (ISR). We use an SL-GCN model to show that the phonological features are recognizable with 85% accuracy, and that they are effective as an auxiliary target to ISR. Learning to recognize phonological features alongside gloss results in a 6% improvement for few-shot ISR accuracy and a 2% improvement for ISR accuracy overall. Instructions for downloading the data can be found at https://github.com/leekezar/SemLex.

📄 PDF Abstract BibTeX arXiv:2310.00196

Code (1)

leekezar/semlex 공식 구현

Tasks

FairnessSign Language Recognition

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

Exploring Strategies for Modeling Sign Language Phonology

2023-09-30 · Lee Kezar, Riley Carlin, Tejas Srinivasan, Zed Sehyr 외

Like speech, signs are composed of discrete, recombinable features called phonemes. Prior work shows that models which can recognize phonemes are better at sign recognition, motivating deeper exploration into strategies …

A Comparison of Modeling Units in Sequence-to-Sequence Speech Recognition with the Transformer on Mandarin Chinese

2018-05-16 · Shiyu Zhou, Linhao Dong, Shuang Xu, Bo Xu

The choice of modeling units is critical to automatic speech recognition (ASR) tasks. Conventional ASR systems typically choose context-dependent states (CD-states) or context-dependent phonemes (CD-phonemes) as their mo…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderLanguage Modeling+4

Brain-to-Text Decoding with Context-Aware Neural Representations and Large Language Models

2024-11-16 · Jingyuan Li, Trung Le, Chaofei Fan, Mingfei Chen 외

Decoding attempted speech from neural activity offers a promising avenue for restoring communication abilities in individuals with speech impairments. Previous studies have focused on mapping neural activity to text usin…

Talking With Your Hands: Scaling Hand Gestures and Recognition With CNNs

2019-05-10 · Okan Köpüklü, Yao Rong, Gerhard Rigoll

The use of hand gestures provides a natural alternative to cumbersome interface devices for Human-Computer Interaction (HCI) systems. As the technology advances and communication between humans and machines becomes more …

From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes

2024-10-30 · Zébulon Goriely, Richard Diehl Martinez, Andrew Caines, Lisa Beinborn 외

Language models are typically trained on large corpora of text in their default orthographic form. However, this is not the only option; representing data as streams of phonemes can offer unique advantages, from deeper i…

Language Acquisition