S+PAGE: A Speaker and Position-Aware Graph Neural Network Model for Emotion Recognition in Conversation
Emotion recognition in conversation (ERC) has attracted much attention in recent years for its necessity in widespread applications. Existing ERC methods mostly model the self and inter-speaker context separately, posing a major issue for lacking enough interaction between them. In this paper, we propose a novel Speaker and Position-Aware Graph neural network model for ERC (S+PAGE), which contains three stages to combine the benefits of both Transformer and relational graph convolution network (R-GCN) for better contextual modeling. Firstly, a two-stream conversational Transformer is presented to extract the coarse self and inter-speaker contextual features for each utterance. Then, a speaker and position-aware conversation graph is constructed, and we propose an enhanced R-GCN model, called PAG, to refine the coarse features guided by a relative positional encoding. Finally, both of the features from the former two stages are input into a conditional random field layer to model the emotion transfer.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion RecognitionEmotion Recognition in ConversationGraph Neural NetworkPositionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
PAGE: A Position-Aware Graph-Based Model for Emotion Cause Entailment in Conversation
Conversational Causal Emotion Entailment (C2E2) is a task that aims at recognizing the causes corresponding to a target emotion in a conversation. The order of utterances in the conversation affects the causal inference.…
Causal Emotion EntailmentCausal InferencePositionRelation-aware Graph Attention Networks with Relational Position Encodings for Emotion Recognition in Conversations
Interest in emotion recognition in conversations (ERC) has been increasing in various fields, because it can be used to analyze user behaviors and detect fake news. Many recent ERC methods use graph-based neural networks…
Emotion RecognitionEmotion Recognition in ConversationGraph AttentionPosition+1Speaker Style-Aware Phoneme Anchoring for Improved Cross-Lingual Speech Emotion Recognition
Cross-lingual speech emotion recognition (SER) remains a challenging task due to differences in phonetic variability and speaker-specific expressive styles across languages. Effectively capturing emotion under such diver…
Speech Emotion RecognitionAutomatic Comic Generation with Stylistic Multi-page Layouts and Emotion-driven Text Balloon Generation
In this paper, we propose a fully automatic system for generating comic books from videos without any human intervention. Given an input video along with its subtitles, our approach first extracts informative keyframes b…
Multi-speaker Emotional Text-to-speech Synthesizer
We present a methodology to train our multi-speaker emotional text-to-speech synthesizer that can express speech for 10 speakers' 7 different emotions. All silences from audio samples are removed prior to learning. This …
Alltext-to-speechText to Speech