paper-with-me

홈 › Papers

Meta-PerSER: Few-Shot Listener Personalized Speech Emotion Recognition via Meta-learning

2025-05-22 · Liang-Yeh Shen, Shi-Xin Fang, Yi-Cheng Lin, Huang-Cheng Chou, Hung-Yi Lee

This paper introduces Meta-PerSER, a novel meta-learning framework that personalizes Speech Emotion Recognition (SER) by adapting to each listener's unique way of interpreting emotion. Conventional SER systems rely on aggregated annotations, which often overlook individual subtleties and lead to inconsistent predictions. In contrast, Meta-PerSER leverages a Model-Agnostic Meta-Learning (MAML) approach enhanced with Combined-Set Meta-Training, Derivative Annealing, and per-layer per-step learning rates, enabling rapid adaptation with only a few labeled examples. By integrating robust representations from pre-trained self-supervised models, our framework first captures general emotional cues and then fine-tunes itself to personal annotation styles. Experiments on the IEMOCAP corpus demonstrate that Meta-PerSER significantly outperforms baseline methods in both seen and unseen data scenarios, highlighting its promise for personalized emotion recognition.

📄 PDF Abstract BibTeX arXiv:2505.16220

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionMeta-LearningSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

No Audiogram: Leveraging Existing Scores for Personalized Speech Intelligibility Prediction

2025-05-31 · Haoshuai Zhou, Changgeng Mo, Boxuan Cao, Linkai Li 외

Personalized speech intelligibility prediction is challenging. Previous approaches have mainly relied on audiograms, which are inherently limited in accuracy as they only capture a listener's hearing threshold for pure t…

Predictionspeech-recognitionSpeech Recognition

Perceptual Implications of Automatic Anonymization in Pathological Speech

2025-05-01 · Soroosh Tayebi Arasteh, Saba Afza, Tri-Thien Nguyen, Lukas Buess 외

Automatic anonymization techniques are essential for ethical sharing of pathological speech data, yet their perceptual consequences remain understudied. This study presents the first comprehensive human-centered analysis…

Diagnostic

Comparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study

2026-06-29 · Yuanyuan Zhang, Dimme de Groot, Jorge Martinez, Odette Scharenborg arxiv

In our goal to develop personalised dysarthric speech recognition (DSR) models, this study compared the recognition performances of human listeners and those of three state-of-the-art, off-the-shelf ASR systems (Whisper-…

Speech Recognition

SingDance: Compositional Zero-Shot Singing-and-Dancing Video Generation with Role-Aware Audio Conditioning

2026-08-17 · Tao Feng, Xu Li, Xiangyang Luo, Ming Wen 외 arxiv

Generating personalized dance videos from a reference image, text prompt, and audio track requires music-conditioned body motion. Singing-and-dancing adds a second requirement: the visible subject must also articulate th…

Video Generation

When Denoising Hinders: Revisiting Zero-Shot ASR with SAM-Audio and Whisper

2026-03-05 · Akif Islam, Raufun Nahar, Md. Ekramul Hamid arxiv

Recent advances in automatic speech recognition (ASR) and speech enhancement have led to a widespread assumption that improving perceptual audio quality should directly benefit recognition accuracy. In this work, we rigo…

Speech RecognitionSpeech Enhancement