paper-with-me

홈 › Papers

Exploring Attention Mechanisms in Integration of Multi-Modal Information for Sign Language Recognition and Translation

2023-09-04 · Zaber Ibn Abdul Hakim, Rasman Mubtasim Swargo, Muhammad Abdullah Adnan

Understanding intricate and fast-paced movements of body parts is essential for the recognition and translation of sign language. The inclusion of additional information intended to identify and locate the moving body parts has been an interesting research topic recently. However, previous works on using multi-modal information raise concerns such as sub-optimal multi-modal feature merging method, or the model itself being too computationally heavy. In our work, we have addressed such issues and used a plugin module based on cross-attention to properly attend to each modality with another. Moreover, we utilized 2-stage training to remove the dependency of separate feature extractors for additional modalities in an end-to-end approach, which reduces the concern about computational complexity. Besides, our additional cross-attention plugin module is very lightweight which doesn't add significant computational overhead on top of the original baseline. We have evaluated the performance of our approaches on the RWTH-PHOENIX-2014 dataset for sign language recognition and the RWTH-PHOENIX-2014T dataset for the sign language translation task. Our approach reduced the WER by 0.9 on the recognition task and increased the BLEU-4 scores by 0.8 on the translation task.

📄 PDF Abstract BibTeX arXiv:2309.01860

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Flow EstimationSign Language RecognitionSign Language TranslationTranslation

Similar Papers 제목 키워드 기반

Exploring Attention Mechanisms for Multimodal Emotion Recognition in an Emergency Call Center Corpus

2023-06-12 · Théo Deschamps-Berger, Lori Lamel, Laurence Devillers

The emotion detection technology to enhance human decision-making is an important research issue for real-world applications, but real-life emotion datasets are relatively rare and small. The experiments conducted in thi…

Decision MakingEmotion RecognitionMultimodal Emotion RecognitionSpeech Emotion Recognition

MoRE-GNN: Multi-omics Data Integration with a Heterogeneous Graph Autoencoder

2025-10-08 · Zhiyu Wang, Sonia Koszut, Pietro Liò, Francesco Ceccarelli arxiv

The integration of multi-omics single-cell data remains challenging due to high-dimensionality and complex inter-modality relationships. To address this, we introduce MoRE-GNN (Multi-omics Relational Edge Graph Neural Ne…

Graph Neural Network

Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition

2025-12-04 · Novanto Yudistira arxiv

This study introduces a pioneering methodology for human action recognition by harnessing deep neural network techniques and adaptive fusion strategies across multiple modalities, including RGB, optical flows, audio, and…

Self-Supervised LearningAction RecognitionAction Detection

Exploring Diverse Methods in Visual Question Answering

2024-04-21 · Panfeng Li, Qikai Yang, Xieming Geng, Wenjing Zhou 외

This study explores innovative methods for improving Visual Question Answering (VQA) using Generative Adversarial Networks (GANs), autoencoders, and attention mechanisms. Leveraging a balanced VQA dataset, we investigate…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

GAME: Generalized deep learning model towards multimodal data integration for early screening of adolescent mental disorders

2023-09-18 · Zhicheng Du, Chenyao Jiang, Xi Yuan, Shiyao Zhai 외

The timely identification of mental disorders in adolescents is a global public health challenge.Single factor is difficult to detect the abnormality due to its complex and subtle nature. Additionally, the generalized mu…

Data Integration