paper-with-me

홈 › Papers

Describing emotions with acoustic property prompts for speech emotion recognition

2022-11-14 · Hira Dhamyal, Benjamin Elizalde, Soham Deshmukh, Huaming Wang, Bhiksha Raj, Rita Singh

Emotions lie on a broad continuum and treating emotions as a discrete number of classes limits the ability of a model to capture the nuances in the continuum. The challenge is how to describe the nuances of emotions and how to enable a model to learn the descriptions. In this work, we devise a method to automatically create a description (or prompt) for a given audio by computing acoustic properties, such as pitch, loudness, speech rate, and articulation rate. We pair a prompt with its corresponding audio using 5 different emotion datasets. We trained a neural network model using these audio-text pairs. Then, we evaluate the model using one more dataset. We investigate how the model can learn to associate the audio with the descriptions, resulting in performance improvement of Speech Emotion Recognition and Speech Audio Retrieval. We expect our findings to motivate research describing the broad continuum of emotion

📄 PDF Abstract BibTeX arXiv:2211.07737

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionRetrievalSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

Prompting Audios Using Acoustic Properties For Emotion Representation

2023-10-03 · Hira Dhamyal, Benjamin Elizalde, Soham Deshmukh, Huaming Wang 외

Emotions lie on a continuum, but current models treat emotions as a finite valued discrete variable. This representation does not capture the diversity in the expression of emotion. To better represent emotions we propos…

Contrastive LearningDiversityEmotion RecognitionRetrieval+1

Classification of Emotions and Evaluation of Customer Satisfaction from Speech in Real World Acoustic Environments

2021-08-26 · Luis Felipe Parra-Gallego, Juan Rafael Orozco-Arroyave

This paper focuses on finding suitable features to robustly recognize emotions and evaluate customer satisfaction from speech in real acoustic scenarios. The classification of emotions is based on standard and well-known…

Beyond saliency: enhancing explanation of speech emotion recognition with expert-referenced acoustic cues

2025-11-12 · Seham Nasr, Zhao Ren, David Johnson arxiv

Explainable AI (XAI) for Speech Emotion Recognition (SER) is critical for building transparent, trustworthy models. Current saliency-based methods, adapted from vision, highlight spectrogram regions but fail to show whet…

Speech Emotion Recognition

Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations

2024-09-26 · Yujia Sun, Zeyu Zhao, Korin Richmond, Yuanchao Li

Emotion recognition from speech and music shares similarities due to their acoustic overlap, which has led to interest in transferring knowledge between these domains. However, the shared acoustic cues between speech and…

Domain AdaptationDomain GeneralizationEmotion RecognitionMusic Emotion Recognition+3

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style

2025-08-15 · Wonjune Kang, Deb Roy arxiv

We introduce the task of expressive speech retrieval, where the goal is to retrieve speech utterances spoken in a given style based on a natural language description of that style. While prior work has primarily focused …