paper-with-me

Papers

Transfer Learning for Improving Speech Emotion Classification Accuracy

2018-01-19 · Siddique Latif, Rajib Rana, Shahzad Younis, Junaid Qadir, Julien Epps

The majority of existing speech emotion recognition research focuses on automatic emotion detection using training and testing data from same corpus collected under the same conditions. The performance of such systems has been shown to drop significantly in cross-corpus and cross-language scenarios. To address the problem, this paper exploits a transfer learning technique to improve the performance of speech emotion recognition systems that is novel in cross-language and cross-corpus scenarios. Evaluations on five different corpora in three different languages show that Deep Belief Networks (DBNs) offer better accuracy than previous approaches on cross-corpus emotion recognition, relative to a Sparse Autoencoder and SVM baseline system. Results also suggest that using a large number of languages for training and using a small fraction of the target data in training can significantly boost accuracy compared with baseline also for the corpus with limited training examples.

📄 PDF Abstract BibTeX arXiv:1801.06353

Code (1)

raulsteleac/Speech_Emotion_Recognition tf

Tasks

ClassificationCross-corpusEmotion ClassificationEmotion RecognitionGeneral ClassificationSpeech Emotion RecognitionTransfer Learning

Methods 이 논문이 사용한 방법론

Sparse Autoencoder A Sparse Autoencoder is a type of autoencoder that employs sparsity to achieve an information bottleneck. Specifically the loss function is constructed so that activations are…
Solana Customer Service Number +1-833-534-1729 설명 없음
SVM A Support Vector Machine, or SVM, is a non-parametric supervised learning model. For non-linear classification and regression, they utilise the kernel trick to map inputs…

Similar Papers 제목 키워드 기반

Continuous Metric Learning For Transferable Speech Emotion Recognition and Embedding Across Low-resource Languages

2022-03-28 · Sneha Das, Nicklas Leander Lund, Nicole Nadine Lønfeldt, Anne Katrine Pagsberg 외

Speech emotion recognition~(SER) refers to the technique of inferring the emotional state of an individual from speech signals. SERs continue to garner interest due to their wide applicability. Although the domain is mai…

DenoisingEmotion ClassificationEmotion RecognitionMetric Learning+1

Towards Transferable Speech Emotion Representation: On loss functions for cross-lingual latent representations

2022-03-28 · Sneha Das, Nicole Nadine Lønfeldt, Anne Katrine Pagsberg, Line H. Clemmensen

In recent years, speech emotion recognition (SER) has been used in wide ranging applications, from healthcare to the commercial sector. In addition to signal processing approaches, methods for SER now also use deep learn…

ClassificationDenoisingEmotion ClassificationEmotion Recognition+2

A Transfer Learning Method for Speech Emotion Recognition from Automatic Speech Recognition

2020-08-06 · Sitong Zhou, Homayoon Beigi

This paper presents a transfer learning method in speech emotion recognition based on a Time-Delay Neural Network (TDNN) architecture. A major challenge in the current speech-based emotion detection research is data scar…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSpeech Emotion Recognition+3

Emotion Recognition in Speech using Cross-Modal Transfer in the Wild

2018-08-16 · Samuel Albanie, Arsha Nagrani, Andrea Vedaldi, Andrew Zisserman

Obtaining large, human labelled speech datasets to train models for emotion recognition is a notoriously challenging task, hindered by annotation cost and label ambiguity. In this work, we consider the task of learning e…

Emotion RecognitionFacial Emotion RecognitionFacial Expression Recognition (FER)Speech Emotion Recognition

Multi-Modal Emotion Recognition by Text, Speech and Video Using Pretrained Transformers

2024-02-11 · Minoo Shayaninasab, Bagher BabaAli

Due to the complex nature of human emotions and the diversity of emotion representation methods in humans, emotion recognition is a challenging field. In this research, three input modalities, namely text, audio (speech)…

DiversityEmotion RecognitionMultimodal Emotion RecognitionSelf-Supervised Learning+1