Best Practices for Noise-Based Augmentation to Improve the Performance of Deployable Speech-Based Emotion Recognition Systems
Speech emotion recognition is an important component of any human centered system. But speech characteristics produced and perceived by a person can be influenced by a multitude of reasons, both desirable such as emotion, and undesirable such as noise. To train robust emotion recognition models, we need a large, yet realistic data distribution, but emotion datasets are often small and hence are augmented with noise. Often noise augmentation makes one important assumption, that the prediction label should remain the same in presence or absence of noise, which is true for automatic speech recognition but not necessarily true for perception based tasks. In this paper we make three novel contributions. We validate through crowdsourcing that the presence of noise does change the annotation label and hence may alter the original ground truth label. We then show how disregarding this knowledge and assuming consistency in ground truth labels propagates to downstream evaluation of ML models, both for performance evaluation and robustness testing. We end the paper with a set of recommendations for noise augmentations in speech emotion recognition datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Adversarial AttackAutomatic Speech RecognitionData AugmentationEmotion RecognitionSpeaker VerificationSpeech Emotion Recognitionspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Best Practices for Noise-Based Augmentation to Improve the Performance of Deployable Speech-Based Emotion Recognition Systems
Emotion recognition models are a key component of several downstream applications, such as mental health assessments. These models are usually trained on small, clean, and synthetically controlled datasets, which leads t…
Data AugmentationEmotion Recognitionspeech-recognitionSpeech RecognitionADA: A Game-Theoretic Perspective on Data Augmentation for Object Detection
The use of random perturbations of ground truth data, such as random translation or scaling of bounding boxes, is a common heuristic used for data augmentation that has been shown to prevent overfitting and improve gener…
Data AugmentationObjectobject-detectionObject Detection+1Evaluating LLM Prompts for Data Augmentation in Multi-label Classification of Ecological Texts
Large language models (LLMs) play a crucial role in natural language processing (NLP) tasks, improving the understanding, generation, and manipulation of human language across domains such as translating, summarizing, an…
Data AugmentationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONData Augmentation for Electrocardiograms
Neural network models have demonstrated impressive performance in predicting pathologies and outcomes from the 12-lead electrocardiogram (ECG). However, these models often need to be trained with large, labelled datasets…
Data AugmentationThe effect of data augmentation and 3D-CNN depth on Alzheimer's Disease detection
Machine Learning (ML) has emerged as a promising approach in healthcare, outperforming traditional statistical techniques. However, to establish ML as a reliable tool in clinical practice, adherence to best practices reg…
Alzheimer's Disease DetectionData AugmentationExperimental Design