paper-with-me

홈 › Papers

Advancing Audio Emotion and Intent Recognition with Large Pre-Trained Models and Bayesian Inference

2023-10-16 · Dejan Porjazovski, Yaroslav Getman, Tamás Grósz, Mikko Kurimo

Large pre-trained models are essential in paralinguistic systems, demonstrating effectiveness in tasks like emotion recognition and stuttering detection. In this paper, we employ large pre-trained models for the ACM Multimedia Computational Paralinguistics Challenge, addressing the Requests and Emotion Share tasks. We explore audio-only and hybrid solutions leveraging audio and text modalities. Our empirical results consistently show the superiority of the hybrid approaches over the audio-only models. Moreover, we introduce a Bayesian layer as an alternative to the standard linear output layer. The multimodal fusion approach achieves an 85.4% UAR on HC-Requests and 60.2% on HC-Complaints. The ensemble model for the Emotion Share task yields the best rho value of .614. The Bayesian wav2vec2 approach, explored in this study, allows us to easily build ensembles, at the cost of fine-tuning only one model. Moreover, we can have usable confidence values instead of the usual overconfident posterior probabilities.

📄 PDF Abstract BibTeX arXiv:2310.10179

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian InferenceEmotion RecognitionIntent Recognition

Similar Papers 제목 키워드 기반

Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?

2024-10-12 · Ryotaro Nagase, Takashi Sumiyoshi, Natsuo Yamashita, Kota Dohi 외

This paper proposes a zero-shot speech emotion recognition (SER) method that estimates emotions not previously defined in the SER model training. Conventional methods are limited to recognizing emotions defined by a sing…

Emotion RecognitionSpeech Emotion Recognition

CLARA: Multilingual Contrastive Learning for Audio Representation Acquisition

2023-10-18 · Kari A Noriy, Xiaosong Yang, Marcin Budka, Jian Jun Zhang

Multilingual speech processing requires understanding emotions, a task made difficult by limited labelled data. CLARA, minimizes reliance on labelled data, enhancing generalization across languages. It excels at fosterin…

Audio ClassificationContrastive LearningCross-Lingual TransferData Augmentation+6

Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare

2024-06-15 · Nishargo Nigar

The process of identifying human emotion and affective states from speech is known as speech emotion recognition (SER). This is based on the observation that tone and pitch in the voice frequently convey underlying emoti…

Emotion RecognitionSpeech Emotion Recognitionspeech-recognitionSpeech Recognition

Song Emotion Recognition: a Performance Comparison Between Audio Features and Artificial Neural Networks

2022-09-24 · Karen Rosero, Arthur Nicholas dos Santos, Pedro Benevenuto Valadares, Bruno Sanches Masiero

When songs are composed or performed, there is often an intent by the singer/songwriter of expressing feelings or emotions through it. For humans, matching the emotiveness in a musical composition or performance with the…

Emotion Recognition

Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey

2025-09-29 · Yuntao Shou, Tao Meng, Wei Ai, Keqin Li arxiv

In recent years, large language models (LLMs) have driven major advances in language understanding, marking a significant step toward artificial general intelligence (AGI). With increasing demands for higher-level semant…

Multimodal Emotion Recognition