A study on cross-corpus speech emotion recognition and data augmentation
Models that can handle a wide range of speakers and acoustic conditions are essential in speech emotion recognition (SER). Often, these models tend to show mixed results when presented with speakers or acoustic conditions that were not visible during training. This paper investigates the impact of cross-corpus data complementation and data augmentation on the performance of SER models in matched (test-set from same corpus) and mismatched (test-set from different corpus) conditions. Investigations using six emotional speech corpora that include single and multiple speakers as well as variations in emotion style (acted, elicited, natural) and recording conditions are presented. Observations show that, as expected, models trained on single corpora perform best in matched conditions while performance decreases between 10-40% in mismatched conditions, depending on corpus specific features. Models trained on mixed corpora can be more stable in mismatched contexts, and the performance reductions range from 1 to 8% when compared with single corpus models in matched conditions. Data augmentation yields additional gains up to 4% and seem to benefit mismatched conditions more than matched ones.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-corpusData AugmentationEmotion RecognitionSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
Filter-based multi-task cross-corpus feature learning for speech emotion recognition
Speech emotion recognition is a highly active field of research in human–machine interaction. A primary challenge faced by researchers in this area is how to tackle the problem of changing data distribution. In the last…
Cross-corpusEmotion Recognitionfeature selectionMulti-Task Learning+1Cross Lingual Cross Corpus Speech Emotion Recognition
The majority of existing speech emotion recognition models are trained and evaluated on a single corpus and a single language setting. These systems do not perform as well when applied in a cross-corpus and cross-languag…
Cross-corpusEmotion RecognitionMulti-Task LearningSpeech Emotion RecognitionCross-Corpus Validation of Speech Emotion Recognition in Urdu using Domain-Knowledge Acoustic Features
Speech Emotion Recognition (SER) is a key affective computing technology that enables emotionally intelligent artificial intelligence. While SER is challenging in general, it is particularly difficult for low-resource la…
Speech Emotion RecognitionCross Lingual Speech Emotion Recognition: Urdu vs. Western Languages
Cross-lingual speech emotion recognition is an important task for practical applications. The performance of automatic speech emotion recognition systems degrades in cross-corpus scenarios, particularly in scenarios invo…
Cross-corpusEmotion RecognitionSpeech Emotion RecognitionMouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition
Cross-corpus speech emotion recognition (SER) plays a vital role in numerous practical applications. Traditional approaches to cross-corpus emotion transfer often concentrate on adapting acoustic features to align with d…
Cross-corpusEmotion RecognitionSpeech Emotion RecognitionTransfer Learning