paper-with-me

Papers

Multi-Window Data Augmentation Approach for Speech Emotion Recognition

2020-10-19 · Sarala Padi, Dinesh Manocha, Ram D. Sriram

We present a Multi-Window Data Augmentation (MWA-SER) approach for speech emotion recognition. MWA-SER is a unimodal approach that focuses on two key concepts; designing the speech augmentation method and building the deep learning model to recognize the underlying emotion of an audio signal. Our proposed multi-window augmentation approach generates additional data samples from the speech signal by employing multiple window sizes in the audio feature extraction process. We show that our augmentation method, combined with a deep learning model, improves speech emotion recognition performance. We evaluate the performance of our approach on three benchmark datasets: IEMOCAP, SAVEE, and RAVDESS. We show that the multi-window model improves the SER performance and outperforms a single-window model. The notion of finding the best window size is an essential step in audio feature extraction. We perform extensive experimental evaluations to find the best window choice and explore the windowing effect for SER analysis.

📄 PDF Abstract BibTeX arXiv:2010.09895

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationEmotion RecognitionSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion Recognition

2024-01-19 · Yong Wang, Cheng Lu, Hailun Lian, Yan Zhao 외

Swin-Transformer has demonstrated remarkable success in computer vision by leveraging its hierarchical feature representation based on Transformer. In speech signals, emotional information is distributed across different…

Emotion RecognitionSpeech Emotion Recognition

DWFormer: Dynamic Window transFormer for Speech Emotion Recognition

2023-03-03 · Shuaiqi Chen, Xiaofen Xing, Weibin Zhang, Weidong Chen 외

Speech emotion recognition is crucial to human-computer interaction. The temporal regions that represent different emotions scatter in different parts of the speech locally. Moreover, the temporal scales of important inf…

Emotion RecognitionSpeech Emotion Recognition

Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition

2023-09-19 · Ziyang Ma, Wen Wu, Zhisheng Zheng, Yiwei Guo 외

In this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and speech synthesis technique, Azure TTS. Fi…

Data AugmentationEmotion RecognitionLanguage ModelingLanguage Modelling+7

Best Practices for Noise-Based Augmentation to Improve the Performance of Deployable Speech-Based Emotion Recognition Systems

2021-04-18 · Mimansa Jaiswal, Emily Mower Provost

Speech emotion recognition is an important component of any human centered system. But speech characteristics produced and perceived by a person can be influenced by a multitude of reasons, both desirable such as emotion…

Adversarial AttackAutomatic Speech RecognitionData AugmentationEmotion Recognition+4

Dynamic Parameter Memory: Temporary LoRA-Enhanced LLM for Long-Sequence Emotion Recognition in Conversation

2025-07-11 · Jialong Mai, Xiaofen Xing, Yawei Li, Zhipeng Li 외

Recent research has focused on applying speech large language model (SLLM) to improve speech emotion recognition (SER). However, the inherently high frame rate in speech modality severely limits the signal processing and…

4kEmotion RecognitionEmotion Recognition in ConversationLarge Language Model+2