paper-with-me

홈 › Papers

Enrolment-based personalisation for improving individual-level fairness in speech emotion recognition

2024-06-10 · Andreas Triantafyllopoulos, Björn Schuller

The expression of emotion is highly individualistic. However, contemporary speech emotion recognition (SER) systems typically rely on population-level models that adopt a `one-size-fits-all' approach for predicting emotion. Moreover, standard evaluation practices measure performance also on the population level, thus failing to characterise how models work across different speakers. In the present contribution, we present a new method for capitalising on individual differences to adapt an SER model to each new speaker using a minimal set of enrolment utterances. In addition, we present novel evaluation schemes for measuring fairness across different speakers. Our findings show that aggregated evaluation metrics may obfuscate fairness issues on the individual-level, which are uncovered by our evaluation, and that our proposed method can improve performance both in aggregated and disaggregated terms.

📄 PDF Abstract BibTeX arXiv:2406.06665

Code (1)

ATriantafyllopoulos/enrollment-personalization 공식 구현 pytorch

Tasks

Emotion RecognitionFairnessSpeech Emotion Recognition

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Exploring speaker enrolment for few-shot personalisation in emotional vocalisation prediction

2022-06-14 · Andreas Triantafyllopoulos, Meishu Song, Zijiang Yang, Xin Jing 외

In this work, we explore a novel few-shot personalisation architecture for emotional vocalisation prediction. The core contribution is an `enrolment' encoder which utilises two unlabelled samples of the target speaker to…

Consistency Based Unsupervised Self-training For ASR Personalisation

2024-01-22 · Jisi Zhang, Vandana Rajan, Haaris Mehmood, David Tuckey 외

On-device Automatic Speech Recognition (ASR) models trained on speech data of a large population might underperform for individuals unseen during training. This is due to a domain shift between user data and the original…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Cross-Attention is all you need: Real-Time Streaming Transformers for Personalised Speech Enhancement

2022-11-08 · Shucong Zhang, Malcolm Chadwick, Alberto Gil C. P. Ramos, Sourav Bhattacharya

Personalised speech enhancement (PSE), which extracts only the speech of a target user and removes everything else from a recorded audio clip, can potentially improve users' experiences of audio AI modules deployed in th…

AllSpeech Enhancement

Personalisation or Prejudice? Addressing Geographic Bias in Hate Speech Detection using Debias Tuning in Large Language Models

2025-05-04 · Paloma Piot, Patricia Martín-Rodilla, Javier Parapar

Commercial Large Language Models (LLMs) have recently incorporated memory features to deliver personalised responses. This memory retains details such as user demographics and individual characteristics, allowing LLMs to…

Hate Speech Detection

Improving Personalisation in Valence and Arousal Prediction using Data Augmentation

2024-04-13 · Munachiso Nwadike, Jialin Li, Hanan Salam

In the field of emotion recognition and Human-Machine Interaction (HMI), personalised approaches have exhibited their efficacy in capturing individual-specific characteristics and enhancing affective prediction accuracy.…

Data AugmentationEmotion Recognition