paper-with-me

홈 › Papers

Integrating Emotion Recognition with Speech Recognition and Speaker Diarisation for Conversations

2023-08-14 · Wen Wu, Chao Zhang, Philip C. Woodland

Although automatic emotion recognition (AER) has recently drawn significant research interest, most current AER studies use manually segmented utterances, which are usually unavailable for dialogue systems. This paper proposes integrating AER with automatic speech recognition (ASR) and speaker diarisation (SD) in a jointly-trained system. Distinct output layers are built for four sub-tasks including AER, ASR, voice activity detection and speaker classification based on a shared encoder. Taking the audio of a conversation as input, the integrated system finds all speech segments and transcribes the corresponding emotion classes, word sequences, and speaker identities. Two metrics are proposed to evaluate AER performance with automatic segmentation based on time-weighted emotion and speaker classification errors. Results on the IEMOCAP dataset show that the proposed system consistently outperforms two baselines with separately trained single-task systems on AER, ASR and SD.

📄 PDF Abstract BibTeX arXiv:2308.07145

Code (1)

w-wu/steer 공식 구현 pytorch

Tasks

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

A Survey on Paralinguistics in Tamil Speech Processing

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Anosha Ignatius, Uthayasanker Thayasivam

Speech carries not only the semantic content but also the paralinguistic information which captures the speaking style. Speaker traits and emotional states affect how words are being spoken. The research on paralinguisti…

Emotion RecognitionSpeaker Identificationspeech-recognitionSpeech Recognition+1

Emotion Impacts Speech Recognition Performance

2019-06-01 · NAACL 2019 6 · Rushab Munot, Ani Nenkova

It has been established that the performance of speech recognition systems depends on multiple factors including the lexical content, speaker identity and dialect. Here we use three English datasets of acted emotion to d…

speech-recognitionSpeech Recognition

Speaker Attentive Speech Emotion Recognition

2021-04-15 · Clément Le Moine, Nicolas Obin, Axel Roebel

Speech Emotion Recognition (SER) task has known significant improvements over the last years with the advent of Deep Neural Networks (DNNs). However, even the most successful methods are still rather failing when adaptat…

Emotion RecognitionSpeech Emotion Recognition

Speaker Normalization for Self-supervised Speech Emotion Recognition

2022-02-02 · Itai Gat, Hagai Aronowitz, Weizhong Zhu, Edmilson Morais 외

Large speech emotion recognition datasets are hard to obtain, and small datasets may contain biases. Deep-net-based classifiers, in turn, are prone to exploit those biases and find shortcuts such as speaker characteristi…

Emotion RecognitionSpeech Emotion Recognition

Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition

2024-01-19 · Ismail Rasim Ulgen, Zongyang Du, Carlos Busso, Berrak Sisman

Speaker embeddings carry valuable emotion-related information, which makes them a promising resource for enhancing speech emotion recognition (SER), especially with limited labeled data. Traditionally, it has been assume…

Contrastive LearningEmotion RecognitionSpeech Emotion Recognition