A Spatio-Temporal Representation Learning as an Alternative to Traditional Glosses in Sign Language Translation and Production
This work addresses the challenges associated with the use of glosses in both Sign Language Translation (SLT) and Sign Language Production (SLP). While glosses have long been used as a bridge between sign language and spoken language, they come with two major limitations that impede the advancement of sign language systems. First, annotating the glosses is a labor-intensive and time-consuming process, which limits the scalability of datasets. Second, the glosses oversimplify sign language by stripping away its spatio-temporal dynamics, reducing complex signs to basic labels and missing the subtle movements essential for precise interpretation. To address these limitations, we introduce Universal Gloss-level Representation (UniGloR), a framework designed to capture the spatio-temporal features inherent in sign language, providing a more dynamic and detailed alternative to the use of the glosses. The core idea of UniGloR is simple yet effective: We derive dense spatio-temporal representations from sign keypoint sequences using self-supervised learning and seamlessly integrate them into SLT and SLP tasks. Our experiments in a keypoint-based setting demonstrate that UniGloR either outperforms or matches the performance of previous SLT and SLP methods on two widely-used datasets: PHOENIX14T and How2Sign.
Code (0)
등록된 구현이 없습니다.
Tasks
Gloss-free Sign Language TranslationRepresentation LearningSelf-Supervised LearningSign Language ProductionSign Language RecognitionSign Language TranslationTranslationSimilar Papers 제목 키워드 기반
STLight: a Fully Convolutional Approach for Efficient Predictive Learning by Spatio-Temporal joint Processing
Spatio-Temporal predictive Learning is a self-supervised learning paradigm that enables models to identify spatial and temporal patterns by predicting future frames based on past frames. Traditional methods, which use re…
Computational EfficiencySelf-Supervised LearningAutoSign: Direct Pose-to-Text Translation for Continuous Sign Language Recognition
Continuously recognizing sign gestures and converting them to glosses plays a key role in bridging the gap between the hearing and hearing-impaired communities. This involves recognizing and interpreting the hands, face,…
Sign Language RecognitionC2ST: Cross-Modal Contextualized Sequence Transduction for Continuous Sign Language Recognition
Continuous Sign Language Recognition (CSLR) aims to transcribe the signs of an untrimmed video into written words or glosses. The mainstream framework for CSLR consists of a spatial module for visual representation l…
Language ModellingRepresentation LearningSign Language RecognitionEnhancing Maritime Trajectory Forecasting via H3 Index and Causal Language Modelling (CLM)
The prediction of ship trajectories is a growing field of study in artificial intelligence. Traditional methods rely on the use of LSTM, GRU networks, and even Transformer architectures for the prediction of spatio-tempo…
Language ModellingTrajectory ForecastingMultilingual eXtended WordNet Knowledge Base: Semantic Parsing and Translation of Glosses
This paper presents a method to create WordNet-like lexical resources for different languages. Instead of directly translating glosses from one language to another, we perform first semantic parsing of WordNet glosses an…
Machine TranslationSemantic ParsingTranslation