Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
Machine learning models for speech emotion recognition (SER) can be trained for different tasks and are usually evaluated based on a few available datasets per task. Tasks could include arousal, valence, dominance, emotional categories, or tone of voice. Those models are mainly evaluated in terms of correlation or recall, and always show some errors in their predictions. The errors manifest themselves in model behaviour, which can be very different along different dimensions even if the same recall or correlation is achieved by the model. This paper introduces a testing framework to investigate behaviour of speech emotion recognition models, by requiring different metrics to reach a certain threshold in order to pass a test. The test metrics can be grouped in terms of correctness, fairness, and robustness. It also provides a method for automatically specifying test thresholds for fairness tests, based on the datasets used, and recommendations on how to select the remaining test thresholds. We evaluated a xLSTM-based and nine transformer-based acoustic foundation models against a convolutional baseline model, testing their performance on arousal, valence, dominance, and emotional category classification. The test results highlight, that models with high correlation or recall might rely on shortcuts -- such as text sentiment --, and differ in terms of fairness.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion RecognitionFairnessSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
Machine Learning Testing: Survey, Landscapes and Horizons
This paper provides a comprehensive survey of Machine Learning Testing (ML testing) research. It covers 144 papers on testing properties (e.g., correctness, robustness, and fairness), testing components (e.g., the data, …
Autonomous DrivingBIG-bench Machine LearningFairnessMachine Translation+2Robust Federated Learning Against Adversarial Attacks for Speech Emotion Recognition
Due to the development of machine learning and speech processing, speech emotion recognition has been a popular research topic in recent years. However, the speech data cannot be protected when it is uploaded and process…
Emotion RecognitionFederated LearningSpeech Emotion RecognitionBest Practices for Noise-Based Augmentation to Improve the Performance of Deployable Speech-Based Emotion Recognition Systems
Speech emotion recognition is an important component of any human centered system. But speech characteristics produced and perceived by a person can be influenced by a multitude of reasons, both desirable such as emotion…
Adversarial AttackAutomatic Speech RecognitionData AugmentationEmotion Recognition+4Is It Still Fair? Investigating Gender Fairness in Cross-Corpus Speech Emotion Recognition
Speech emotion recognition (SER) is a vital component in various everyday applications. Cross-corpus SER models are increasingly recognized for their ability to generalize performance. However, concerns arise regarding f…
Cross-corpusEmotion RecognitionFairnessSpeech Emotion Recognition+1Enrolment-based personalisation for improving individual-level fairness in speech emotion recognition
The expression of emotion is highly individualistic. However, contemporary speech emotion recognition (SER) systems typically rely on population-level models that adopt a `one-size-fits-all' approach for predicting emoti…
Emotion RecognitionFairnessSpeech Emotion Recognition