paper-with-me

홈 › Papers

Arabic Speech Emotion Recognition Employing Wav2vec2.0 and HuBERT Based on BAVED Dataset

2021-10-09 · Omar Mohamed, Salah A. Aly

Recently, there have been tremendous research outcomes in the fields of speech recognition and natural language processing. This is due to the well-developed multi-layers deep learning paradigms such as wav2vec2.0, Wav2vecU, WavBERT, and HuBERT that provide better representation learning and high information capturing. Such paradigms run on hundreds of unlabeled data, then fine-tuned on a small dataset for specific tasks. This paper introduces a deep learning constructed emotional recognition model for Arabic speech dialogues. The developed model employs the state of the art audio representations include wav2vec2.0 and HuBERT. The experiment and performance results of our model overcome the previous known outcomes.

📄 PDF Abstract BibTeX arXiv:2110.04425

Code (1)

OmarMohammed88/AR-Emotion-Recognition 공식 구현

Tasks

Deep LearningEmotion RecognitionRepresentation LearningSpeech Emotion Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion Recognition

2025-09-01 · Ali Abouzeid, Bilal Elbouardi, Mohamed Maged, Shady Shehata arxiv

Speech emotion recognition is vital for human-computer interaction, particularly for low-resource languages like Arabic, which face challenges due to limited data and research. We introduce ArabEmoNet, a lightweight arch…

Speech Emotion Recognition

HARNESS: Lightweight Distilled Arabic Speech Foundation Models

2026-03-31 · Vrunda N. Sukhadia, Shammur Absar Chowdhury arxiv

Large self-supervised speech (SSL) models achieve strong downstream performance, but their size limits deployment in resource-constrained settings. We present HArnESS, an Arabic-centric self-supervised speech model famil…

Speech Emotion RecognitionSpeech Recognition

HARNESS: Lightweight Distilled Arabic Speech Foundation Models

2025-09-18 · Vrunda N. sukhadia, Shammur Absar Chowdhury arxiv

Large pre-trained speech models excel in downstream tasks but their deployment is impractical for resource-limited environments. In this paper, we introduce HArnESS, the first Arabic-centric self-supervised speech model …

Emotion Recognition

A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

2021-11-04 · Yingzhi Wang, Abdelmoumene Boumadane, Abdelwahab Heba

Speech self-supervised models such as wav2vec 2.0 and HuBERT are making revolutionary progress in Automatic Speech Recognition (ASR). However, they have not been totally proven to produce better performance on tasks othe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion Recognitionintent-classification+8

ExHuBERT: Enhancing HuBERT Through Block Extension and Fine-Tuning on 37 Emotion Datasets

2024-06-11 · Shahin Amiriparian, Filip Packań, Maurice Gerczuk, Björn W. Schuller

Foundation models have shown great promise in speech emotion recognition (SER) by leveraging their pre-trained representations to capture emotion patterns in speech signals. To further enhance SER performance across vari…

Emotion RecognitionSpeech Emotion Recognition