paper-with-me

Papers

Learning More with Less: Self-Supervised Approaches for Low-Resource Speech Emotion Recognition

2025-06-01 · Ziwei Gong, Pengyuan Shi, Kaan Donbekci, Lin Ai, Run Chen, David Sasu, Zehui Wu, Julia Hirschberg

Speech Emotion Recognition (SER) has seen significant progress with deep learning, yet remains challenging for Low-Resource Languages (LRLs) due to the scarcity of annotated data. In this work, we explore unsupervised learning to improve SER in low-resource settings. Specifically, we investigate contrastive learning (CL) and Bootstrap Your Own Latent (BYOL) as self-supervised approaches to enhance cross-lingual generalization. Our methods achieve notable F1 score improvements of 10.6% in Urdu, 15.2% in German, and 13.9% in Bangla, demonstrating their effectiveness in LRLs. Additionally, we analyze model behavior to provide insights on key factors influencing performance across languages, and also highlighting challenges in low-resource SER. This work provides a foundation for developing more inclusive, explainable, and robust emotion recognition systems for underrepresented languages.

📄 PDF Abstract BibTeX arXiv:2506.02059

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningEmotion RecognitionSpeech Emotion Recognition

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training

2024-07-18 · Lukuan Dong, Donghong Qin, Fengbo Bai, Fanhua Song 외

The mainstream automatic speech recognition (ASR) technology usually requires hundreds to thousands of hours of annotated speech data. Three approaches to low-resourced ASR are phoneme or subword based supervised pre-tra…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Self-supervised Contrastive Zero to Few-shot Learning from Small, Long-tailed Text data

2020-09-28 · Nils Rethmeier, Isabelle Augenstein

For natural language processing (NLP) ‘text-to-text’ tasks, prevailing approaches heavily rely on pretraining large self-supervised models on massive external datasources. However, this methodology is being critiqued fo…

Few-Shot LearningMulti Label Text ClassificationMulti-Label Text Classificationtext-classification+1

Self-supervised language learning from raw audio: Lessons from the Zero Resource Speech Challenge

2022-10-27 · Ewan Dunbar, Nicolas Hamilakis, Emmanuel Dupoux

Recent progress in self-supervised or unsupervised machine learning has opened the possibility of building a full speech processing system from raw audio without using any textual representations or expert labels such as…

Acoustic Unit DiscoveryLanguage ModelingLanguage ModellingResynthesis

Self-DANA: A Resource-Efficient Channel-Adaptive Self-Supervised Approach for ECG Foundation Models

2025-07-03 · Giuliana Monachino, Nicolò La Porta, Beatrice Zanchi, Luigi Fiorillo 외 arxiv

Foundation Models (FMs) are large-scale machine learning models trained on extensive, diverse datasets that can be adapted to a wide range of downstream tasks with minimal fine-tuning. In the last two years, interest in …

Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations

2024-06-30 · Salah Zaiem, Titouan Parcollet, Slim Essid

Despite being trained on massive and diverse datasets, speech self-supervised encoders are generally used for downstream purposes as mere frozen feature extractors or model initializers before fine-tuning. The former sev…

Continual LearningDomain Generalizationspeech-recognitionSpeech Recognition