Best Practices for Noise-Based Augmentation to Improve the Performance of Deployable Speech-Based Emotion Recognition Systems
Emotion recognition models are a key component of several downstream applications, such as mental health assessments. These models are usually trained on small, clean, and synthetically controlled datasets, which leads to high failure rates in presence of unseen' background noises, promoting noise-overlay based adversarial attacks. Noisy data augmentation has aided robustness of speech recognition and classification models, wherein, the ground truth label remains consistent even in the presence of noise which, isn't always true for subjectively perceived emotion labels. In this work, we create realistic noisy samples of IEMOCAP, using multiple categories of environmental and synthetic noise. We evaluate how ground truth labels (human) and predicted labels (model) change as a function of these noise source introductions. We show that some commonly used noisy augmentation techniques, impact human perception of emotion, thus, falsifying the clean ground truth label. Our experiments show that the performance of both, baseline, and even denoised emotion recognition models significantly declines on noisy samples as compared to that on the clean set. This performance degradation prevails when model is trained on a combination of clean and test set mismatched noisy samples. We investigate how using the above found `human-perceptible noise overlays can lead to inaccurate metrics when testing the model for robustness or vulnerability to adversarial attacks. Finally, we present a set of recommendations for noise-based augmentation of speech emotion datasets and for deploying the models trained using those datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationEmotion Recognitionspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Best Practices for Noise-Based Augmentation to Improve the Performance of Deployable Speech-Based Emotion Recognition Systems
Speech emotion recognition is an important component of any human centered system. But speech characteristics produced and perceived by a person can be influenced by a multitude of reasons, both desirable such as emotion…
Adversarial AttackAutomatic Speech RecognitionData AugmentationEmotion Recognition+4ADA: A Game-Theoretic Perspective on Data Augmentation for Object Detection
The use of random perturbations of ground truth data, such as random translation or scaling of bounding boxes, is a common heuristic used for data augmentation that has been shown to prevent overfitting and improve gener…
Data AugmentationObjectobject-detectionObject Detection+1Evaluating LLM Prompts for Data Augmentation in Multi-label Classification of Ecological Texts
Large language models (LLMs) play a crucial role in natural language processing (NLP) tasks, improving the understanding, generation, and manipulation of human language across domains such as translating, summarizing, an…
Data AugmentationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONData Augmentation for Electrocardiograms
Neural network models have demonstrated impressive performance in predicting pathologies and outcomes from the 12-lead electrocardiogram (ECG). However, these models often need to be trained with large, labelled datasets…
Data AugmentationThe effect of data augmentation and 3D-CNN depth on Alzheimer's Disease detection
Machine Learning (ML) has emerged as a promising approach in healthcare, outperforming traditional statistical techniques. However, to establish ML as a reliable tool in clinical practice, adherence to best practices reg…
Alzheimer's Disease DetectionData AugmentationExperimental Design