paper-with-me

Papers

Speech Emotion: Investigating Model Representations, Multi-Task Learning and Knowledge Distillation

2022-07-02 · Vikramjit Mitra, Hsiang-Yun Sherry Chien, Vasudha Kowtha, Joseph Yitan Cheng, Erdrin Azemi

Estimating dimensional emotions, such as activation, valence and dominance, from acoustic speech signals has been widely explored over the past few years. While accurate estimation of activation and dominance from speech seem to be possible, the same for valence remains challenging. Previous research has shown that the use of lexical information can improve valence estimation performance. Lexical information can be obtained from pre-trained acoustic models, where the learned representations can improve valence estimation from speech. We investigate the use of pre-trained model representations to improve valence estimation from acoustic speech signal. We also explore fusion of representations to improve emotion estimation across all three emotion dimensions: activation, valence and dominance. Additionally, we investigate if representations from pre-trained models can be distilled into models trained with low-level features, resulting in models with a less number of parameters. We show that fusion of pre-trained model embeddings result in a 79% relative improvement in concordance correlation coefficient CCC on valence estimation compared to standard acoustic feature baseline (mel-filterbank energies), while distillation from pre-trained model embeddings to lower-dimensional representations yielded a relative 12% improvement. Such performance gains were observed over two evaluation sets, indicating that our proposed architecture generalizes across those evaluation sets. We report new state-of-the-art "text-free" acoustic-only dimensional emotion estimation $CCC$ values on two MSP-Podcast evaluation sets.

📄 PDF Abstract BibTeX arXiv:2207.03334

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationMulti-Task LearningValence Estimation

Similar Papers 제목 키워드 기반

Investigating salient representations and label Variance in Dimensional Speech Emotion Analysis

2023-12-17 · Vikramjit Mitra, Jingping Nie, Erdrin Azemi

Representations derived from models such as BERT (Bidirectional Encoder Representations from Transformers) and HuBERT (Hidden units BERT), have helped to achieve state-of-the-art performance in dimensional speech emotion…

Emotion RecognitionSpeech Emotion Recognition

Pre-Finetuning for Few-Shot Emotional Speech Recognition

2023-02-24 · Maximillian Chen, Zhou Yu

Speech models have long been known to overfit individual speakers for many classification tasks. This leads to poor generalization in settings where the speakers are out-of-domain or out-of-distribution, as is common in …

Few-Shot Learningspeech-recognitionSpeech RecognitionTransfer Learning

Investigating the Impact of Word Informativeness on Speech Emotion Recognition

2025-06-02 · Sofoklis Kakouros

In emotion recognition from speech, a key challenge lies in identifying speech signal segments that carry the most relevant acoustic variations for discerning specific emotions. Traditional approaches compute functionals…

Emotion RecognitionInformativenessLanguage ModelingLanguage Modelling+1

EmoTale: An Enacted Speech-emotion Dataset in Danish

2025-08-20 · Maja J. Hjuler, Harald V. Skat-Rørdam, Line H. Clemmensen, Sneha Das arxiv

While multiple emotional speech corpora exist for commonly spoken languages, there is a lack of functional datasets for smaller (spoken) languages, such as Danish. To our knowledge, Danish Emotional Speech (DES), publish…

Speech Emotion Recognition

Filter-based multi-task cross-corpus feature learning for speech emotion recognition

2024-02-20 · Signal, Image and Video Processing 2024 2 · Behzad Bakhtiari, Elham Kalhor, Seyed Hossein Ghafarian

Speech emotion recognition is a highly active field of research in human–machine interaction. A primary challenge faced by researchers in this area is how to tackle the problem of changing data distribution. In the last…

Cross-corpusEmotion Recognitionfeature selectionMulti-Task Learning+1