paper-with-me

홈 › Papers

Incorporating End-to-End Speech Recognition Models for Sentiment Analysis

2019-02-28 · Egor Lakomkin, Mohammad Ali Zamani, Cornelius Weber, Sven Magg, Stefan Wermter

Previous work on emotion recognition demonstrated a synergistic effect of combining several modalities such as auditory, visual, and transcribed text to estimate the affective state of a speaker. Among these, the linguistic modality is crucial for the evaluation of an expressed emotion. However, manually transcribed spoken text cannot be given as input to a system practically. We argue that using ground-truth transcriptions during training and evaluation phases leads to a significant discrepancy in performance compared to real-world conditions, as the spoken text has to be recognized on the fly and can contain speech recognition mistakes. In this paper, we propose a method of integrating an automatic speech recognition (ASR) output with a character-level recurrent neural network for sentiment recognition. In addition, we conduct several experiments investigating sentiment recognition for human-robot interaction in a noise-realistic scenario which is challenging for the ASR systems. We quantify the improvement compared to using only the acoustic modality in sentiment recognition. We demonstrate the effectiveness of this approach on the Multimodal Corpus of Sentiment Intensity (MOSI) by achieving 73,6% accuracy in a binary sentiment classification task, exceeding previously reported results that use only acoustic input. In addition, we set a new state-of-the-art performance on the MOSI dataset (80.4% accuracy, 2% absolute improvement).

📄 PDF Abstract BibTeX arXiv:1902.11245

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSentiment AnalysisSentiment Classificationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Zara: A Virtual Interactive Dialogue System Incorporating Emotion, Sentiment and Personality Recognition

2016-12-01 · COLING 2016 12 · Pascale Fung, Anik Dey, Farhad Bin Siddique, Ruixi Lin 외

Zara, or {`}Zara the Supergirl{'} is a virtual robot, that can exhibit empathy while interacting with an user, with the aid of its built in facial and emotion recognition, sentiment analysis, and speech module. At the en…

Emotion RecognitionFeature EngineeringSentiment Analysisspeech-recognition+1

Sentiment Reasoning for Healthcare

2024-07-24 · Khai-Nguyen Nguyen, Khai Le-Duc, Bach Phan Tat, Duy Le 외

Transparency in AI healthcare decision-making is crucial. By incorporating rationales to explain reason for each predicted label, users could understand Large Language Models (LLMs)'s reasoning to make better decision. I…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decision MakingMultimodal Sentiment Analysis+4

Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts

2025-03-09 · Aref Farhadipour, Hossein Ranjbar, Masoumeh Chapariniya, Teodora Vukovic 외

Emotion recognition and sentiment analysis are pivotal tasks in speech and language processing, particularly in real-world scenarios involving multi-party, conversational data. This paper presents a multimodal approach t…

Emotion RecognitionMultimodal Emotion RecognitionSentiment Analysis

Do Multi-Sense Embeddings Improve Natural Language Understanding?

2015-06-02 · EMNLP 2015 9 · Jiwei Li, Dan Jurafsky

Learning a distinct representation for each sense of an ambiguous word could lead to more powerful and fine-grained models of vector-space representations. Yet while `multi-sense' methods have been proposed and tested on…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Natural Language Understanding+4

Leveraging Pre-trained Language Model for Speech Sentiment Analysis

2021-06-11 · Suwon Shon, Pablo Brusco, Jing Pan, Kyu J. Han 외

In this paper, we explore the use of pre-trained language models to learn sentiment information of written texts for speech sentiment analysis. First, we investigate how useful a pre-trained language model would be in a …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+4