paper-with-me

Papers

MM-WLAuslan: Multi-View Multi-Modal Word-Level Australian Sign Language Recognition Dataset

2024-10-25 · Xin Shen, Heming Du, Hongwei Sheng, Shuyun Wang, Hui Chen, Huiqiang Chen, Zhuojie Wu, Xiaobiao Du, Jiaying Ying, Ruihan Lu, Qingzheng Xu, Xin Yu

Isolated Sign Language Recognition (ISLR) focuses on identifying individual sign language glosses. Considering the diversity of sign languages across geographical regions, developing region-specific ISLR datasets is crucial for supporting communication and research. Auslan, as a sign language specific to Australia, still lacks a dedicated large-scale word-level dataset for the ISLR task. To fill this gap, we curate \underline{\textbf{the first}} large-scale Multi-view Multi-modal Word-Level Australian Sign Language recognition dataset, dubbed MM-WLAuslan. Compared to other publicly available datasets, MM-WLAuslan exhibits three significant advantages: (1) the largest amount of data, (2) the most extensive vocabulary, and (3) the most diverse of multi-modal camera views. Specifically, we record 282K+ sign videos covering 3,215 commonly used Auslan glosses presented by 73 signers in a studio environment. Moreover, our filming system includes two different types of cameras, i.e., three Kinect-V2 cameras and a RealSense camera. We position cameras hemispherically around the front half of the model and simultaneously record videos using all four cameras. Furthermore, we benchmark results with state-of-the-art methods for various multi-modal ISLR settings on MM-WLAuslan, including multi-view, cross-camera, and cross-view. Experiment results indicate that MM-WLAuslan is a challenging ISLR dataset, and we hope this dataset will contribute to the development of Auslan and the advancement of sign languages worldwide. All datasets and benchmarks are available at MM-WLAuslan.

📄 PDF Abstract BibTeX arXiv:2410.19488

Code (0)

등록된 구현이 없습니다.

Tasks

Sign Language Recognition

Similar Papers 제목 키워드 기반

Spectral Graph-Based Method of Multimodal Word Embedding

2017-08-01 · WS 2017 8 · Kazuki Fukui, Takamasa Oshikiri, Hidetoshi Shimodaira

In this paper, we propose a novel method for multimodal word embedding, which exploit a generalized framework of multi-view spectral graph embedding to take into account visual appearances or scenes denoted by words in a…

Graph EmbeddingImage RetrievalMachine TranslationPart-Of-Speech Tagging+4

MuVAM: A Multi-View Attention-based Model for Medical Visual Question Answering

2021-07-07 · Haiwei Pan, Shuning He, Kejia Zhang, Bo Qu 외

Medical Visual Question Answering (VQA) is a multi-modal challenging task widely considered by research communities of the computer vision and natural language processing. Since most current medical VQA models focus on v…

Medical Visual Question AnsweringMissing LabelsQuestion AnsweringVisual Question Answering+1

EmoCo: Visual Analysis of Emotion Coherence in Presentation Videos

2019-07-29 · Haipeng Zeng, Xingbo Wang, Aoyu Wu, Yong Wang 외

Emotions play a key role in human communication and public presentations. Human emotions are usually expressed through multiple modalities. Therefore, exploring multimodal emotions and their coherence is of great value f…

ClusteringSentence

Bridging Lexical Ambiguity and Vision: A Mini Review on Visual Word Sense Disambiguation

2026-02-01 · Shashini Nilukshi, Deshan Sumanathilaka arxiv

This paper offers a mini review of Visual Word Sense Disambiguation (VWSD), which is a multimodal extension of traditional Word Sense Disambiguation (WSD). VWSD helps tackle lexical ambiguity in vision-language tasks. Wh…

Word Sense DisambiguationText-to-Image GenerationPrompt Engineering

Symbol Emergence as Inter-personal Categorization with Head-to-head Latent Word

2022-05-24 · Kazuma Furukawa, Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi

In this study, we propose a head-to-head type (H2H-type) inter-personal multimodal Dirichlet mixture (Inter-MDM) by modifying the original Inter-MDM, which is a probabilistic generative model that represents the symbol e…

Vocal Bursts Type Prediction