paper-with-me

홈 › Papers

A Multimodal Fusion Network For Student Emotion Recognition Based on Transformer and Tensor Product

2024-03-13 · Ao Xiang, Zongqing Qi, Han Wang, Qin Yang, Danqing Ma

This paper introduces a new multi-modal model based on the Transformer architecture and tensor product fusion strategy, combining BERT's text vectors and ViT's image vectors to classify students' psychological conditions, with an accuracy of 93.65%. The purpose of the study is to accurately analyze the mental health status of students from various data sources. This paper discusses modal fusion methods, including early, late and intermediate fusion, to overcome the challenges of integrating multi-modal information. Ablation studies compare the performance of different models and fusion techniques, showing that the proposed model outperforms existing methods such as CLIP and ViLBERT in terms of accuracy and inference speed. Conclusions indicate that while this model has significant advantages in emotion recognition, its potential to incorporate other data modalities provides areas for future research.

📄 PDF Abstract BibTeX arXiv:2403.08511

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion Recognitionobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Position-Wise Feed-Forward Layer 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation

2024-01-16 · Taeyang Yun, Hyunkuk Lim, Jeonghwan Lee, Min Song

Emotion Recognition in Conversation (ERC) plays a crucial role in enabling dialogue systems to effectively respond to user requests. The emotions in a conversation can be identified by the representations from various mo…

Emotion RecognitionEmotion Recognition in ConversationKnowledge DistillationLanguage Modeling+1

Multi-Modal Emotion Recognition by Text, Speech and Video Using Pretrained Transformers

2024-02-11 · Minoo Shayaninasab, Bagher BabaAli

Due to the complex nature of human emotions and the diversity of emotion representation methods in humans, emotion recognition is a challenging field. In this research, three input modalities, namely text, audio (speech)…

DiversityEmotion RecognitionMultimodal Emotion RecognitionSelf-Supervised Learning+1

Multimodal Emotion Recognition with Transformer-Based Self Supervised Feature Fusion

2020-10-27 · Shamane Siriwardhana ; Tharindu Kaluarachchi ; Mark Billinghurst ; Suranga Nanayakkara

Emotion Recognition is a challenging research area given its complex nature, and humans express emotional cues across various modalities such as language, facial expressions, and speech. Representation and fusion of feat…

Emotion RecognitionMultimodal Deep LearningMultimodal Emotion RecognitionMultimodal Sentiment Analysis+2

Multimodal Emotion Recognition using Audio-Video Transformer Fusion with Cross Attention

2024-07-26 · Joe Dhanith P R, Shravan Venkatraman, Vigya Sharma, Santhosh Malarvannan 외

Understanding emotions is a fundamental aspect of human communication. Integrating audio and video signals offers a more comprehensive understanding of emotional states compared to traditional methods that rely on a sing…

Emotion RecognitionMultimodal Emotion Recognition

MDEAW: A Multimodal Dataset for Emotion Analysis through EDA and PPG signals from wireless wearable low-cost off-the-shelf Devices

2022-07-14 · Arijit Nandi, Fatos Xhafa, Laia Subirats, Santi Fort

We present MDEAW, a multimodal database consisting of Electrodermal Activity (EDA) and Photoplethysmography (PPG) signals recorded during the exams for the course taught by the teacher at Eurecat Academy, Sabadell, Barce…

Emotion RecognitionPhotoplethysmography (PPG)