paper-with-me

홈 › Papers

Speech Emotion Recognition Using Quaternion Convolutional Neural Networks

2021-10-31 · Aneesh Muppidi, Martin Radfar

Although speech recognition has become a widespread technology, inferring emotion from speech signals still remains a challenge. To address this problem, this paper proposes a quaternion convolutional neural network (QCNN) based speech emotion recognition (SER) model in which Mel-spectrogram features of speech signals are encoded in an RGB quaternion domain. We show that our QCNN based SER model outperforms other real-valued methods in the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS, 8-classes) dataset, achieving, to the best of our knowledge, state-of-the-art results. The QCNN also achieves comparable results with the state-of-the-art methods in the Interactive Emotional Dyadic Motion Capture (IEMOCAP 4-classes) and Berlin EMO-DB (7-classes) datasets. Specifically, the model achieves an accuracy of 77.87\%, 70.46\%, and 88.78\% for the RAVDESS, IEMOCAP, and EMO-DB datasets, respectively. In addition, our results show that the quaternion unit structure is better able to encode internal dependencies to reduce its model size significantly compared to other methods.

📄 PDF Abstract BibTeX arXiv:2111.00404

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionSpeech Emotion Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Learning Speech Emotion Representations in the Quaternion Domain

2022-04-05 · Eric Guizzo, Tillman Weyde, Simone Scardapane, Danilo Comminiello

The modeling of human emotion expression in speech signals is an important, yet challenging task. The high resource demand of speech emotion recognition models, combined with the the general scarcity of emotion-labelled …

Emotion RecognitionSpeech Emotion Recognition

Quaternion Convolutional Neural Networks for End-to-End Automatic Speech Recognition

2018-06-20 · Titouan Parcollet, Ying Zhang, Mohamed Morchid, Chiheb Trabelsi 외

Recently, the connectionist temporal classification (CTC) model coupled with recurrent (RNN) or convolutional neural networks (CNN), made it easier to train speech recognition systems in an end-to-end fashion. However in…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Phoneme Recognitionspeech-recognition+1

Compressing Quaternion Convolutional Neural Networks for Audio Classification

2025-10-24 · Arshdeep Singh, Vinayak Abrol, Mark D. Plumbley arxiv

Conventional Convolutional Neural Networks (CNNs) in the real domain have been widely used for audio classification. However, their convolution operations process multi-channel inputs independently, limiting the ability …

Environmental Sound ClassificationSpeech Emotion RecognitionMusic Genre RecognitionKnowledge Distillation

Speech recognition with quaternion neural networks

2018-11-21 · Titouan Parcollet, Mirco Ravanelli, Mohamed Morchid, Georges Linarès 외

Neural network architectures are at the core of powerful automatic speech recognition systems (ASR). However, while recent researches focus on novel model architectures, the acoustic input features remain almost unchange…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Speech Recognition with Quaternion Neural Networks

2018-10-15 · NIPS Workshop IRASL 2018 · Anonymous

Neural network architectures are at the core of powerful automatic speech recognition systems (ASR). However, while recent researches focus on novel model architectures, the acoustic input features remain almost unchange…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition