Spatio-temporal Sign Language Representation and Translation
This paper describes the DFKI-MLT submission to the WMT-SLT 2022 sign language translation (SLT) task from Swiss German Sign Language (video) into German (text). State-of-the-art techniques for SLT use a generic seq2seq architecture with customized input embeddings. Instead of word embeddings as used in textual machine translation, SLT systems use features extracted from video frames. Standard approaches often do not benefit from temporal features. In our participation, we present a system that learns spatio-temporal feature representations and translation in a single model, resulting in a real end-to-end architecture expected to better generalize to new data sets. Our best system achieved 5 ± 1 BLEU points on the development set, but the performance on the test dropped to 0.11 ± 0.06 BLEU points.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationSign Language TranslationTranslationWord EmbeddingsSimilar Papers 제목 키워드 기반
Spatio-temporal Sign Language Representation and Translation
This paper describes the DFKI-MLT submission to the WMT-SLT 2022 sign language translation (SLT) task from Swiss German Sign Language (video) into German (text). State-of-the-art techniques for SLT use a generic seq2seq …
Sign Language TranslationMachine TranslationA Spatio-Temporal Representation Learning as an Alternative to Traditional Glosses in Sign Language Translation and Production
This work addresses the challenges associated with the use of glosses in both Sign Language Translation (SLT) and Sign Language Production (SLP). While glosses have long been used as a bridge between sign language and sp…
Gloss-free Sign Language TranslationRepresentation LearningSelf-Supervised LearningSign Language Production+3Sign Language Translation with Hierarchical Spatio-TemporalGraph Neural Network
Sign language translation (SLT), which generates text in a spoken language from visual content in a sign language, is important to assist the hard-of-hearing community for their communications. Inspired by neural machine…
Graph Neural NetworkMachine TranslationNMTSign Language Translation+1A multitask transformer to sign language translation using motion gesture primitives
The absence of effective communication the deaf population represents the main social gap in this community. Furthermore, the sign language, main deaf communication tool, is unlettered, i.e., there is no formal written r…
Sign Language TranslationSMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action Recognition
Event cameras action recognition (EAR) offers compelling privacy-protecting and efficiency advantages, where temporal motion dynamics is of great importance. Existing spatiotemporal multi-view representation learning (SM…
Representation LearningAction RecognitionObject Recognition