paper-with-me

Papers

Speaker-Guided Encoder-Decoder Framework for Emotion Recognition in Conversation

2022-06-07 · Yinan Bao, Qianwen Ma, Lingwei Wei, Wei Zhou, Songlin Hu

The emotion recognition in conversation (ERC) task aims to predict the emotion label of an utterance in a conversation. Since the dependencies between speakers are complex and dynamic, which consist of intra- and inter-speaker dependencies, the modeling of speaker-specific information is a vital role in ERC. Although existing researchers have proposed various methods of speaker interaction modeling, they cannot explore dynamic intra- and inter-speaker dependencies jointly, leading to the insufficient comprehension of context and further hindering emotion prediction. To this end, we design a novel speaker modeling scheme that explores intra- and inter-speaker dependencies jointly in a dynamic manner. Besides, we propose a Speaker-Guided Encoder-Decoder (SGED) framework for ERC, which fully exploits speaker information for the decoding of emotion. We use different existing methods as the conversational context encoder of our framework, showing the high scalability and flexibility of the proposed framework. Experimental results demonstrate the superiority and effectiveness of SGED.

📄 PDF Abstract BibTeX arXiv:2206.03173

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderEmotion RecognitionEmotion Recognition in Conversation

Similar Papers 제목 키워드 기반

Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion

2020-05-13 · Kun Zhou, Berrak Sisman, Mingyang Zhang, Haizhou Li

Emotional voice conversion aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. The prior studies on emotional voice conversion are mostly carried …

DecoderVoice Conversion

Exploring speaker enrolment for few-shot personalisation in emotional vocalisation prediction

2022-06-14 · Andreas Triantafyllopoulos, Meishu Song, Zijiang Yang, Xin Jing 외

In this work, we explore a novel few-shot personalisation architecture for emotional vocalisation prediction. The core contribution is an `enrolment' encoder which utilises two unlabelled samples of the target speaker to…

Decoupling Speaker-Independent Emotions for Voice Conversion Via Source-Filter Networks

2021-10-04 · Zhaojie Luo, Shoufeng Lin, Rui Liu, Jun Baba 외

Emotional voice conversion (VC) aims to convert a neutral voice to an emotional (e.g. happy) one while retaining the linguistic information and speaker identity. We note that the decoupling of emotional features from oth…

DecoderVoice Conversion

Multi-speaker Emotion Conversion via Latent Variable Regularization and a Chained Encoder-Decoder-Predictor Network

2020-07-25 · Ravi Shankar, Hsi-Wei Hsieh, Nicolas Charon, Archana Venkataraman

We propose a novel method for emotion conversion in speech based on a chained encoder-decoder-predictor neural network architecture. The encoder constructs a latent embedding of the fundamental frequency (F0) contour and…

Decoder

Seen and Unseen emotional style transfer for voice conversion with a new emotional speech dataset

2020-10-28 · Kun Zhou, Berrak Sisman, Rui Liu, Haizhou Li

Emotional voice conversion aims to transform emotional prosody in speech while preserving the linguistic content and speaker identity. Prior studies show that it is possible to disentangle emotional prosody using an enco…

DecoderEmotion RecognitionGenerative Adversarial NetworkSpeech Emotion Recognition+2