paper-with-me

홈 › Papers

Breaking the Barriers: Video Vision Transformers for Word-Level Sign Language Recognition

2025-04-10 · Alexander Brettmann, Jakob Grävinghoff, Marlene Rüschoff, Marie Westhues

Sign language is a fundamental means of communication for the deaf and hard-of-hearing (DHH) community, enabling nuanced expression through gestures, facial expressions, and body movements. Despite its critical role in facilitating interaction within the DHH population, significant barriers persist due to the limited fluency in sign language among the hearing population. Overcoming this communication gap through automatic sign language recognition (SLR) remains a challenge, particularly at a dynamic word-level, where temporal and spatial dependencies must be effectively recognized. While Convolutional Neural Networks have shown potential in SLR, they are computationally intensive and have difficulties in capturing global temporal dependencies between video sequences. To address these limitations, we propose a Video Vision Transformer (ViViT) model for word-level American Sign Language (ASL) recognition. Transformer models make use of self-attention mechanisms to effectively capture global relationships across spatial and temporal dimensions, which makes them suitable for complex gesture recognition tasks. The VideoMAE model achieves a Top-1 accuracy of 75.58% on the WLASL100 dataset, highlighting its strong performance compared to traditional CNNs with 65.89%. Our study demonstrates that transformer-based architectures have great potential to advance SLR, overcome communication barriers and promote the inclusion of DHH individuals.

📄 PDF Abstract BibTeX arXiv:2504.07792

Code (0)

등록된 구현이 없습니다.

Tasks

Gesture RecognitionSign Language Recognition

Methods 이 논문이 사용한 방법론

Attention 설명 없음
American 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Position-Wise Feed-Forward Layer 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Multi-Head Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…

Similar Papers 제목 키워드 기반

Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video

2025-04-28 · Sonia Joseph, Praneet Suresh, Lorenz Hufe, Edward Stevinson 외

Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision mechanistic interpretability has been hi…

From Federated Learning to X-Learning: Breaking the Barriers of Decentrality Through Random Walks

2025-09-03 · Allan Salihovic, Payam Abdisarabshali, Michael Langberg, Seyyedali Hosseinalipour arxiv

We provide our perspective on X-Learning (XL), a novel distributed learning architecture that generalizes and extends the concept of decentralization. Our goal is to present a vision for XL, introducing its unexplored de…

Federated Learning

Exploring Transformers in Natural Language Generation: GPT, BERT, and XLNet

2021-02-16 · M. Onat Topal, Anil Bas, Imke van Heerden

Recent years have seen a proliferation of attention mechanisms and the rise of Transformers in Natural Language Generation (NLG). Previously, state-of-the-art NLG architectures such as RNN and LSTM ran into vanishing gra…

Text Generation

Learning Self-Regularized Adversarial Views for Self-Supervised Vision Transformers

2022-10-16 · Tao Tang, Changlin Li, Guangrun Wang, Kaicheng Yu 외

Automatic data augmentation (AutoAugment) strategies are indispensable in supervised data-efficient training protocols of vision transformers, and have led to state-of-the-art results in supervised learning. Despite the …

Data AugmentationImage RetrievalRetrievalSelf-Supervised Learning+1

Scaling Linear Mode Connectivity and Merging to Billion Parameter Pretrained Transformers

2026-06-22 · Tianyi Li, Zhiqiang Shen arxiv

Linear mode connectivity (LMC) provides a promising foundation for understanding and merging independently trained neural networks, but existing methods typically optimize the interpolation path from only one model endpo…