paper-with-me

Papers

Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques

2024-06-12 · Yuanchao Li, Peter Bell, Catherine Lai

Text data is commonly utilized as a primary input to enhance Speech Emotion Recognition (SER) performance and reliability. However, the reliance on human-transcribed text in most studies impedes the development of practical SER systems, creating a gap between in-lab research and real-world scenarios where Automatic Speech Recognition (ASR) serves as the text source. Hence, this study benchmarks SER performance using ASR transcripts with varying Word Error Rates (WERs) from eleven models on three well-known corpora: IEMOCAP, CMU-MOSI, and MSP-Podcast. Our evaluation includes both text-only and bimodal SER with six fusion techniques, aiming for a comprehensive analysis that uncovers novel findings and challenges faced by current SER research. Additionally, we propose a unified ASR error-robust framework integrating ASR error correction and modality-gated fusion, achieving lower WER and higher SER results compared to the best-performing ASR transcript. These findings provide insights into SER with ASR assistance, especially for real-world applications.

📄 PDF Abstract BibTeX arXiv:2406.08353

Code (1)

yc-li20/SER-on-WER-and-Fusion 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSpeech Emotion Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

ASR and Emotional Speech: A Word-Level Investigation of the Mutual Impact of Speech and Emotion Recognition

2023-05-25 · Yuanchao Li, Zeyu Zhao, Ondrej Klejch, Peter Bell 외

In Speech Emotion Recognition (SER), textual data is often used alongside audio signals to address their inherent variability. However, the reliance on human annotated text in most research hinders the development of pra…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSpeech Emotion Recognition+2

LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations

2025-01-20 · Soumya Dutta, Sriram Ganapathy

Emotion recognition in conversations (ERC) is challenging due to the multimodal nature of the emotion expression. In this paper, we propose to pretrain a text-based recognition model from unsupervised speech transcripts …

Emotion RecognitionMultimodal Emotion Recognition

Emotion Impacts Speech Recognition Performance

2019-06-01 · NAACL 2019 6 · Rushab Munot, Ani Nenkova

It has been established that the performance of speech recognition systems depends on multiple factors including the lexical content, speaker identity and dialect. Here we use three English datasets of acted emotion to d…

speech-recognitionSpeech Recognition

Fusing ASR Outputs in Joint Training for Speech Emotion Recognition

2021-10-29 · Yuanchao Li, Peter Bell, Catherine Lai

Alongside acoustic information, linguistic features based on speech transcripts have been proven useful in Speech Emotion Recognition (SER). However, due to the scarcity of emotion labelled data and the difficulty of rec…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSpeech Emotion Recognition+2

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization

2025-07-25 · Hsuan-Yu Wang, Pei-Ying Lee, Berlin Chen arxiv

In this paper, we investigate the impact of incorporating timestamp-based alignment between Automatic Speech Recognition (ASR) transcripts and Speaker Diarization (SD) outputs on Speech Emotion Recognition (SER) accuracy…

Multimodal Emotion RecognitionSpeech Emotion RecognitionSpeaker DiarizationSpeech Recognition