paper-with-me

Papers

Mediapi-RGB: Enabling Technological Breakthroughs in French Sign Language (LSF) Research through an Extensive Video-Text Corpus

2024-02-29 · VISAPP 2024 2 · Yanis Ouakrim, Hannah Bull, Michèle Gouiffès, Denis Beautemps, Thomas Hueber, Annelies Braffort

We introduce Mediapi-RGB, a new dataset of French Sign Language (LSF) along with the first LSF-to-French machine translation model. With 86 hours of video, it the largest LSF corpora with translation. The corpus consists of original content in French Sign Language produced by deaf journalists, and has subtitles in written French aligned to the signing. The current release of Mediapi-RGB is available at the Ortolang corpus repository, and can be used for academic research purposes. The test and validation sets contain 13 and 7 hours of video respectively. The training set contains 66 hours of video that will be released progressively until December 2024. Additionally, the current release contains skeleton keypoints, sign temporal segmentation, spatio-temporal features and subtitles for all the videos in the train, validation and test sets, as well as a suggested vocabulary of nouns for evaluation purposes. In addition, we present the results obtained on this corpus with the first LSF-to-French translation baseline to give an overview of the possibilities offered by this corpus of unprecedented caliber for LSF. Finally, we suggest potential technological and linguistic applications for this new video-text dataset.

📄 PDF Abstract BibTeX

Code (1)

youakrim/sign_language_translation_Mediapi-RGB tf

Tasks

Machine TranslationSign Language TranslationTranslation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

MEDIAPI-SKEL - A 2D-Skeleton Video Database of French Sign Language With Aligned French Subtitles

2020-05-01 · LREC 2020 5 · Hannah Bull, Annelies Braffort, Mich{\`e}le Gouiff{\`e}s

This paper presents MEDIAPI-SKEL, a 2D-skeleton database of French Sign Language videos aligned with French subtitles. The corpus contains 27 hours of video of body, face and hand keypoints, aligned to subtitles with a v…

Cross-Modal RetrievalRetrievalSemantic SegmentationVideo Semantic Segmentation

EMPATH: MediaPipe-Aided Ensemble Learning with Attention-Based Transformers for Accurate Recognition of Bangla Word-Level Sign Language

2024-12-04 · International Conference on Pattern Recognition 2024 12 · Kazi Reyazul Hasan, Muhammad Abdullah Adnan

In this paper, we introduce EMPATH, an advanced computational framework developed to substantially enhance the recognition of Bangla Sign Language (BdSL). By integrating Ensemble Learning, MediaPipe Holistic for gesture …

Ensemble LearningSign Language Recognition

FQuAD2.0: French Question Answering and Learning When You Don’t Know

2022-06-01 · LREC 2022 6 · Quentin Heinrich, Gautier Viaud, Wacim Belblidia

Question Answering, including Reading Comprehension, is one of the NLP research areas that has seen significant scientific breakthroughs over the past few years, thanks to the concomitant advances in Language Modeling. M…

ArticlesFQuADLanguage ModelingLanguage Modelling+2

FQuAD2.0: French Question Answering and knowing that you know nothing

2021-09-27 · Quentin Heinrich, Gautier Viaud, Wacim Belblidia

Question Answering, including Reading Comprehension, is one of the NLP research areas that has seen significant scientific breakthroughs over the past few years, thanks to the concomitant advances in Language Modeling. M…

ArticlesFQuADLanguage ModelingLanguage Modelling+2

MediaPipe Hands: On-device Real-time Hand Tracking

2020-06-18 · Fan Zhang, Valentin Bazarevsky, Andrey Vakunov, Andrei Tkachenka 외

We present a real-time on-device hand tracking pipeline that predicts hand skeleton from single RGB camera for AR/VR applications. The pipeline consists of two models: 1) a palm detector, 2) a hand landmark model. It's i…