Meta-PerSER: Few-Shot Listener Personalized Speech Emotion Recognition via Meta-learning
This paper introduces Meta-PerSER, a novel meta-learning framework that personalizes Speech Emotion Recognition (SER) by adapting to each listener's unique way of interpreting emotion. Conventional SER systems rely on aggregated annotations, which often overlook individual subtleties and lead to inconsistent predictions. In contrast, Meta-PerSER leverages a Model-Agnostic Meta-Learning (MAML) approach enhanced with Combined-Set Meta-Training, Derivative Annealing, and per-layer per-step learning rates, enabling rapid adaptation with only a few labeled examples. By integrating robust representations from pre-trained self-supervised models, our framework first captures general emotional cues and then fine-tunes itself to personal annotation styles. Experiments on the IEMOCAP corpus demonstrate that Meta-PerSER significantly outperforms baseline methods in both seen and unseen data scenarios, highlighting its promise for personalized emotion recognition.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion RecognitionMeta-LearningSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
No Audiogram: Leveraging Existing Scores for Personalized Speech Intelligibility Prediction
Personalized speech intelligibility prediction is challenging. Previous approaches have mainly relied on audiograms, which are inherently limited in accuracy as they only capture a listener's hearing threshold for pure t…
Predictionspeech-recognitionSpeech RecognitionPerceptual Implications of Automatic Anonymization in Pathological Speech
Automatic anonymization techniques are essential for ethical sharing of pathological speech data, yet their perceptual consequences remain understudied. This study presents the first comprehensive human-centered analysis…
DiagnosticComparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study
In our goal to develop personalised dysarthric speech recognition (DSR) models, this study compared the recognition performances of human listeners and those of three state-of-the-art, off-the-shelf ASR systems (Whisper-…
Speech RecognitionSingDance: Compositional Zero-Shot Singing-and-Dancing Video Generation with Role-Aware Audio Conditioning
Generating personalized dance videos from a reference image, text prompt, and audio track requires music-conditioned body motion. Singing-and-dancing adds a second requirement: the visible subject must also articulate th…
Video GenerationWhen Denoising Hinders: Revisiting Zero-Shot ASR with SAM-Audio and Whisper
Recent advances in automatic speech recognition (ASR) and speech enhancement have led to a widespread assumption that improving perceptual audio quality should directly benefit recognition accuracy. In this work, we rigo…
Speech RecognitionSpeech Enhancement