CopyPaste: An Augmentation Method for Speech Emotion Recognition
Data augmentation is a widely used strategy for training robust machine learning models. It partially alleviates the problem of limited data for tasks like speech emotion recognition (SER), where collecting data is expensive and challenging. This study proposes CopyPaste, a perceptually motivated novel augmentation procedure for SER. Assuming that the presence of emotions other than neutral dictates a speaker's overall perceived emotion in a recording, concatenation of an emotional (emotion E) and a neutral utterance can still be labeled with emotion E. We hypothesize that SER performance can be improved using these concatenated utterances in model training. To verify this, three CopyPaste schemes are tested on two deep learning models: one trained independently and another using transfer learning from an x-vector model, a speaker recognition model. We observed that all three CopyPaste schemes improve SER performance on all the three datasets considered: MSP-Podcast, Crema-D, and IEMOCAP. Additionally, CopyPaste performs better than noise augmentation and, using them together improves the SER performance further. Our experiments on noisy test sets suggested that CopyPaste is effective even in noisy test conditions.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationEmotion RecognitionSpeaker RecognitionSpeech Emotion RecognitionTransfer LearningSimilar Papers 제목 키워드 기반
Multi-Window Data Augmentation Approach for Speech Emotion Recognition
We present a Multi-Window Data Augmentation (MWA-SER) approach for speech emotion recognition. MWA-SER is a unimodal approach that focuses on two key concepts; designing the speech augmentation method and building the de…
Data AugmentationEmotion RecognitionSpeech Emotion RecognitionBest Practices for Noise-Based Augmentation to Improve the Performance of Deployable Speech-Based Emotion Recognition Systems
Speech emotion recognition is an important component of any human centered system. But speech characteristics produced and perceived by a person can be influenced by a multitude of reasons, both desirable such as emotion…
Adversarial AttackAutomatic Speech RecognitionData AugmentationEmotion Recognition+4Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
In this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and speech synthesis technique, Azure TTS. Fi…
Data AugmentationEmotion RecognitionLanguage ModelingLanguage Modelling+7Cross Lingual Speech Emotion Recognition: Urdu vs. Western Languages
Cross-lingual speech emotion recognition is an important task for practical applications. The performance of automatic speech emotion recognition systems degrades in cross-corpus scenarios, particularly in scenarios invo…
Cross-corpusEmotion RecognitionSpeech Emotion RecognitionMulti-stream Attention-based BLSTM with Feature Segmentation for Speech Emotion Recognition
This paper proposes a speech emotion recognition technique that considers the suprasegmental characteristics and temporal change of individual speech parameters. In recent years, speech emotion recognition using Bidir…
Data AugmentationEmotional Speech SynthesisEmotion RecognitionSpeech Emotion Recognition+1