CO-VADA: A Confidence-Oriented Voice Augmentation Debiasing Approach for Fair Speech Emotion Recognition
Bias in speech emotion recognition (SER) systems often stems from spurious correlations between speaker characteristics and emotional labels, leading to unfair predictions across demographic groups. Many existing debiasing methods require model-specific changes or demographic annotations, limiting their practical use. We present CO-VADA, a Confidence-Oriented Voice Augmentation Debiasing Approach that mitigates bias without modifying model architecture or relying on demographic information. CO-VADA identifies training samples that reflect bias patterns present in the training data and then applies voice conversion to alter irrelevant attributes and generate samples. These augmented samples introduce speaker variations that differ from dominant patterns in the data, guiding the model to focus more on emotion-relevant features. Our framework is compatible with various SER models and voice conversion tools, making it a scalable and practical solution for improving fairness in SER systems.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion RecognitionFairnessSpeech Emotion RecognitionVoice ConversionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Target-oriented Multimodal Sentiment Classification with Counterfactual-enhanced Debiasing
Target-oriented multimodal sentiment classification seeks to predict sentiment polarity for specific targets from image-text pairs. While existing works achieve competitive performance, they often over-rely on textual co…
Contrastive LearningData AugmentationSwift Markov Logic for Probabilistic Reasoning on Knowledge Graphs
We provide a framework for probabilistic reasoning in Vadalog-based Knowledge Graphs (KGs), satisfying the requirements of ontological reasoning: full recursion, powerful existential quantification, expression of inducti…
Knowledge GraphsManagementRelational ReasoningExploring Robust Face-Voice Matching in Multilingual Environments
This paper presents Team Xaiofei's innovative approach to exploring Face-Voice Association in Multilingual Environments (FAME) at ACM Multimedia 2024. We focus on the impact of different languages in face-voice matching …
Data AugmentationHow Smart Are `Water Smart Landscapes'?
Understanding the effectiveness of alternative approaches to water conservation is crucially important for ensuring the security and reliability of water services for urban residents. We analyze data from one of the long…
MVAdapt: Zero-Shot Multi-Vehicle Adaptation for End-to-End Autonomous Driving
End-to-End (E2E) autonomous driving models are usually trained and evaluated with a fixed ego-vehicle, even though their driving policy is implicitly tied to vehicle dynamics. When such a model is deployed on a vehicle w…
Autonomous Driving