Multi-modal embeddings using multi-task learning for emotion recognition
General embeddings like word2vec, GloVe and ELMo have shown a lot of success in natural language tasks. The embeddings are typically extracted from models that are built on general tasks such as skip-gram models and natural language generation. In this paper, we extend the work from natural language understanding to multi-modal architectures that use audio, visual and textual information for machine learning tasks. The embeddings in our network are extracted using the encoder of a transformer model trained using multi-task training. We use person identification and automatic speech recognition as the tasks in our embedding generation framework. We tune and evaluate the embeddings on the downstream task of emotion recognition and demonstrate that on the CMU-MOSEI dataset, the embeddings can be used to improve over previous state of the art results.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionMulti-Task LearningNatural Language UnderstandingPerson Identificationspeech-recognitionSpeech RecognitionText GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Modality-Transferable Emotion Embeddings for Low-Resource Multimodal Emotion Recognition
Despite the recent achievements made in the multi-modal emotion recognition task, two problems still exist and have not been well investigated: 1) the relationship between different emotion categories are not utilized, w…
Emotion RecognitionMultimodal Emotion RecognitionWord EmbeddingsMulti-level Fusion of Wav2vec 2.0 and BERT for Multimodal Emotion Recognition
The research and applications of multimodal emotion recognition have become increasingly popular recently. However, multimodal emotion recognition faces the challenge of lack of data. To solve this problem, we propose to…
Emotion RecognitionMultimodal Emotion RecognitionTransfer LearningQuality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion
This paper addresses data quality issues in multimodal emotion recognition in conversation (MERC) through systematic quality control and multi-stage transfer learning. We implement a quality control pipeline for MELD and…
Multimodal Emotion RecognitionTransfer LearningFace RecognitionFace DetectionHCAM -- Hierarchical Cross Attention Model for Multi-modal Emotion Recognition
Emotion recognition in conversations is challenging due to the multi-modal nature of the emotion expression. We propose a hierarchical cross-attention model (HCAM) approach to multi-modal emotion recognition using a comb…
Emotion ClassificationEmotion RecognitionEmotion Recognition in ConversationMultimodal Emotion RecognitionGaze-enhanced Crossmodal Embeddings for Emotion Recognition
Emotional expressions are inherently multimodal -- integrating facial behavior, speech, and gaze -- but their automatic recognition is often limited to a single modality, e.g. speech during a phone call. While previous w…
Emotion ClassificationEmotion Recognition