paper-with-me

Papers

Speech Emotion Recognition Leveraging OpenAI's Whisper Representations and Attentive Pooling Methods

2026-02-05 · Ali Shendabadi, Parnia Izadirad, Mostafa Salehi, Mahmoud Bijankhan arxiv

Speech Emotion Recognition (SER) research has faced limitations due to the lack of standard and sufficiently large datasets. Recent studies have leveraged pre-trained models to extract features for downstream tasks such as SER. This work explores the capabilities of Whisper, a pre-trained ASR system, in speech emotion recognition by proposing two attention-based pooling methods, Multi-head Attentive Average Pooling and QKV Pooling, designed to efficiently reduce the dimensionality of Whisper representations while preserving emotional features. We experiment on English and Persian, using the IEMOCAP and ShEMO datasets respectively, with Whisper Tiny and Small. Our multi-head QKV architecture achieves state-of-the-art results on the ShEMO dataset, with a 2.47% improvement in unweighted accuracy. We further compare the performance of different Whisper encoder layers and find that intermediate layers often perform better for SER on the Persian dataset, providing a lightweight and efficient alternative to much larger models such as HuBERT X-Large. Our findings highlight the potential of Whisper as a representation extractor for SER and demonstrate the effectiveness of attention-based pooling for dimension reduction.

📄 PDF Abstract BibTeX arXiv:2602.06000

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Emotion Recognition

Similar Papers 제목 키워드 기반

Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition

2025-05-23 · Leonora Vesterbacka, Faton Rekathati, Robin Kurtz, Justyna Sikora 외

This work presents a suite of fine-tuned Whisper models for Swedish, trained on a dataset of unprecedented size and variability for this mid-resourced language. As languages of smaller sizes are often underrepresented in…

speech-recognitionSpeech Recognition

Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition

2024-07-30 · Aref Farhadipour, Homa Asadi, Volker Dellwo

In automatic speech recognition, any factor that alters the acoustic properties of speech can pose a challenge to the system's performance. This paper presents a novel approach for automatic whispered speech recognition …

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Empowering Dysarthric Speech: Leveraging Advanced LLMs for Accurate Speech Correction and Multimodal Emotion Analysis

2024-10-13 · Kaushal Attaluri, Anirudh CHVS, Sireesha Chittepu

Dysarthria is a motor speech disorder caused by neurological damage that affects the muscles used for speech production, leading to slurred, slow, or difficult-to-understand speech. It affects millions of individuals wor…

Emotion RecognitionSentence

Whispering in Norwegian: Navigating Orthographic and Dialectic Challenges

2024-02-02 · Per E Kummervold, Javier de la Rosa, Freddy Wetjen, Rolv-Arild Braaten 외

This article introduces NB-Whisper, an adaptation of OpenAI's Whisper, specifically fine-tuned for Norwegian language Automatic Speech Recognition (ASR). We highlight its key contributions and summarise the results achie…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Adapting OpenAI's Whisper for Speech Recognition on Code-Switch Mandarin-English SEAME and ASRU2019 Datasets

2023-11-29 · Yuhang Yang, Yizhou Peng, Xionghu Zhong, Hao Huang 외

This paper details the experimental results of adapting the OpenAI's Whisper model for Code-Switch Mandarin-English Speech Recognition (ASR) on the SEAME and ASRU2019 corpora. We conducted 2 experiments: a) using adaptat…

speech-recognitionSpeech Recognition