Multi-Task Learning with Auxiliary Speaker Identification for Conversational Emotion Recognition
Conversational emotion recognition (CER) has attracted increasing interests in the natural language processing (NLP) community. Different from the vanilla emotion recognition, effective speaker-sensitive utterance representation is one major challenge for CER. In this paper, we exploit speaker identification (SI) as an auxiliary task to enhance the utterance representation in conversations. By this method, we can learn better speaker-aware contextual representations from the additional SI corpus. Experiments on two benchmark datasets demonstrate that the proposed architecture is highly effective for CER, obtaining new state-of-the-art results on two datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion Recognition in ConversationMulti-Task LearningSpeaker IdentificationSimilar Papers 제목 키워드 기반
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
This paper presents a novel framework for multi-talker automatic speech recognition without the need for auxiliary information. Serialized Output Training (SOT), a widely used approach, suffers from recognition errors du…
Automatic Speech RecognitionMulti-Task Learningspeech-recognitionSpeech RecognitionSpeech Enhancement using Self-Adaptation and Multi-Head Self-Attention
This paper investigates a self-adaptation method for speech enhancement using auxiliary speaker-aware features; we extract a speaker representation used for adaptation directly from the test utterance. Conventional studi…
Multi-Task LearningSpeaker IdentificationSpeech Enhancementspeech-recognition+1Towards Making the Most of Dialogue Characteristics for Neural Chat Translation
Neural Chat Translation (NCT) aims to translate conversational text between speakers of different languages. Despite the promising performance of sentence-level and context-aware neural machine translation models, there …
Machine TranslationResponse GenerationSentenceSpeaker Identification+1Reusing Preprocessing Data as Auxiliary Supervision in Conversational Analysis
Conversational analysis systems are trained using noisy human labels and often require heavy preprocessing during multi-modal feature extraction. Using noisy labels in single-task learning increases the risk of over-fitt…
Feature EngineeringMulti-Task LearningTransfer Learning in Conversational Analysis through Reusing Preprocessing Data as Supervisors
Conversational analysis systems are trained using noisy human labels and often require heavy preprocessing during multi-modal feature extraction. Using noisy labels in single-task learning increases the risk of over-fitt…
Feature EngineeringMulti-Task LearningTransfer Learning