paper-with-me

Papers

A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

2021-11-04 · Yingzhi Wang, Abdelmoumene Boumadane, Abdelwahab Heba

Speech self-supervised models such as wav2vec 2.0 and HuBERT are making revolutionary progress in Automatic Speech Recognition (ASR). However, they have not been totally proven to produce better performance on tasks other than ASR. In this work, we explored partial fine-tuning and entire fine-tuning on wav2vec 2.0 and HuBERT pre-trained models for three non-ASR speech tasks: Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding. With simple proposed downstream frameworks, the best scores reached 79.58% weighted accuracy on speaker-dependent setting and 73.01% weighted accuracy on speaker-independent setting for Speech Emotion Recognition on IEMOCAP, 2.36% equal error rate for Speaker Verification on VoxCeleb1, 89.38% accuracy for Intent Classification and 78.92% F1 for Slot Filling on SLURP, showing the strength of fine-tuned wav2vec 2.0 and HuBERT on learning prosodic, voice-print and semantic representations.

📄 PDF Abstract BibTeX arXiv:2111.02735

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion Recognitionintent-classificationIntent Classificationslot-fillingSlot FillingSpeaker VerificationSpeech Emotion Recognitionspeech-recognitionSpeech RecognitionSpoken Language Understanding

Similar Papers 제목 키워드 기반

Arabic Speech Emotion Recognition Employing Wav2vec2.0 and HuBERT Based on BAVED Dataset

2021-10-09 · Omar Mohamed, Salah A. Aly

Recently, there have been tremendous research outcomes in the fields of speech recognition and natural language processing. This is due to the well-developed multi-layers deep learning paradigms such as wav2vec2.0, Wav2v…

Deep LearningEmotion RecognitionRepresentation LearningSpeech Emotion Recognition+2

ExHuBERT: Enhancing HuBERT Through Block Extension and Fine-Tuning on 37 Emotion Datasets

2024-06-11 · Shahin Amiriparian, Filip Packań, Maurice Gerczuk, Björn W. Schuller

Foundation models have shown great promise in speech emotion recognition (SER) by leveraging their pre-trained representations to capture emotion patterns in speech signals. To further enhance SER performance across vari…

Emotion RecognitionSpeech Emotion Recognition

Frame-level emotional state alignment method for speech emotion recognition

2023-12-27 · Qifei Li, Yingming Gao, Cong Wang, Yayue Deng 외

Speech emotion recognition (SER) systems aim to recognize human emotional state during human-computer interaction. Most existing SER systems are trained based on utterance-level labels. However, not all frames in an audi…

Emotion RecognitionSpeech Emotion Recognition

Explainable Speech Emotion Recognition: Weighted Attribute Fairness to Model Demographic Contributions to Social Bias

2026-03-26 · Tomisin Ogunnubi, Yupei Li, Björn Schuller arxiv

Speech Emotion Recognition (SER) systems have growing applications in sensitive domains such as mental health and education, where biased predictions can cause harm. Traditional fairness metrics, such as Equalised Odds a…

Speech Emotion Recognition

Emotion Detection in Speech Using Lightweight and Transformer-Based Models: A Comparative and Ablation Study

2025-11-01 · Lucky Onyekwelu-Udoka, Md Shafiqul Islam, Md Shahedul Hasan arxiv

Emotion recognition from speech plays a vital role in the development of empathetic human-computer interaction systems. This paper presents a comparative analysis of lightweight transformer-based models, DistilHuBERT and…

Speech Emotion Recognition