SignMusketeers: An Efficient Multi-Stream Approach for Sign Language Translation at Scale
A persistent challenge in sign language video processing, including the task of sign language to written language translation, is how we learn representations of sign language in an effective and efficient way that can preserve the important attributes of these languages, while remaining invariant to irrelevant visual differences. Informed by the nature and linguistics of signed languages, our proposed method focuses on just the most relevant parts in a signing video: the face, hands and body posture of the signer. However, instead of using pose estimation coordinates from off-the-shelf pose tracking models, which have inconsistent performance for hands and faces, we propose to learn the complex handshapes and rich facial expressions of sign languages in a self-supervised fashion. Our approach is based on learning from individual frames (rather than video sequences) and is therefore much more efficient than prior work on sign language pre-training. Compared to a recent model that established a new state of the art in sign language translation on the How2Sign dataset, our approach yields similar translation performance, using less than 3\% of the compute.
Code (0)
등록된 구현이 없습니다.
Tasks
Pose EstimationPose TrackingSign Language TranslationTranslationSimilar Papers 제목 키워드 기반
StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model
Streaming speech translation (StreamST) requires determining appropriate timing, known as policy, to generate translations while continuously receiving source speech inputs, balancing low latency with high translation qu…
Large-Scale Streaming End-to-End Speech Translation with Neural Transducers
Neural transducers have been widely used in automatic speech recognition (ASR). In this paper, we introduce it to streaming end-to-end speech translation (ST), which aims to convert audio signals to texts in other langua…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+2PHRASED: Phrase Dictionary Biasing for Speech Translation
Phrases are essential to understand the core concepts in conversations. However, due to their rare occurrence in training data, correct translation of phrases is challenging in speech translation tasks. In this paper, we…
Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+1Two-Stream Network for Sign Language Recognition and Translation
Sign languages are visual languages using manual articulations and non-manual elements to convey information. For sign language recognition and translation, the majority of existing approaches directly encode RGB videos …
Sign Language RecognitionSign Language TranslationTranslationVocal Bursts Valence PredictionTextless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
Cascaded speech-to-speech translation systems often suffer from the error accumulation problem and high latency, which is a result of cascaded modules whose inference delays accumulate. In this paper, we propose a transd…
Language ModelingLanguage ModellingMachine TranslationSpeech-to-Speech Translation+3