paper-with-me

홈 › Papers

Novel Dual-Channel Long Short-Term Memory Compressed Capsule Networks for Emotion Recognition

2021-12-26 · Ismail Shahin, Noor Hindawi, Ali Bou Nassif, Adi Alhudhaif, Kemal Polat

Recent analysis on speech emotion recognition has made considerable advances with the use of MFCCs spectrogram features and the implementation of neural network approaches such as convolutional neural networks (CNNs). Capsule networks (CapsNet) have gained gratitude as alternatives to CNNs with their larger capacities for hierarchical representation. To address these issues, this research introduces a text-independent and speaker-independent SER novel architecture, where a dual-channel long short-term memory compressed-CapsNet (DC-LSTM COMP-CapsNet) algorithm is proposed based on the structural features of CapsNet. Our proposed novel classifier can ensure the energy efficiency of the model and adequate compression method in speech emotion recognition, which is not delivered through the original structure of a CapsNet. Moreover, the grid search approach is used to attain optimal solutions. Results witnessed an improved performance and reduction in the training and testing running time. The speech datasets used to evaluate our algorithm are: Arabic Emirati-accented corpus, English speech under simulated and actual stress corpus, English Ryerson audio-visual database of emotional speech and song corpus, and crowd-sourced emotional multimodal actors dataset. This work reveals that the optimum feature extraction method compared to other known methods is MFCCs delta-delta. Using the four datasets and the MFCCs delta-delta, DC-LSTM COMP-CapsNet surpasses all the state-of-the-art systems, classical classifiers, CNN, and the original CapsNet. Using the Arabic Emirati-accented corpus, our results demonstrate that the proposed work yields average emotion recognition accuracy of 89.3% compared to 84.7%, 82.2%, 69.8%, 69.2%, 53.8%, 42.6%, and 31.9% based on CapsNet, CNN, support vector machine, multi-layer perceptron, k-nearest neighbor, radial basis function, and naive Bayes, respectively.

📄 PDF Abstract BibTeX arXiv:2112.13350

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionSpeech Emotion Recognition

Methods 이 논문이 사용한 방법론

CapsNet Capsule Network is a machine learning system that is a type of artificial neural network that can be used to better model hierarchical relationships. The approach is an…

Similar Papers 제목 키워드 기반

Short-time detection of QRS complexes using dual channels 1 based on U-Net and bidirectional long short-term memory

2019-12-18

Cardiovascular disease is associated with high rates of morbidity and mortality, and can be reflected by 19 abnormal features of electrocardiogram (ECG). Detecting changes in the QRS complexes in ECG 20 signals is regard…

Compensation of Fiber Nonlinearities in Digital Coherent Systems Leveraging Long Short-Term Memory Neural Networks

2020-01-31 · Stavros Deligiannidis, Adonis Bogris, Charis Mesaritakis, Yannis Kopsinis

We introduce for the first time the utilization of Long short-term memory (LSTM) neural network architectures for the compensation of fiber nonlinearities in digital coherent systems. We conduct numerical simulations con…

LSTM-CF: Unifying Context Modeling and Fusion with LSTMs for RGB-D Scene Labeling

2016-04-18 · Zhen Li, Yukang Gan, Xiaodan Liang, Yizhou Yu 외

Semantic labeling of RGB-D scenes is crucial to many intelligent applications including perceptual robotics. It generates pixelwise and fine-grained label maps from simultaneously sensed photometric (RGB) and depth chann…

Scene Labeling

Classifying Relations via Long Short Term Memory Networks along Shortest Dependency Path

2015-08-15 · Xu Yan, Lili Mou, Ge Li, Yunchuan Chen 외

Relation classification is an important research arena in the field of natural language processing (NLP). In this paper, we present SDP-LSTM, a novel neural network to classify the relation of two entities in a sentence.…

ClassificationGeneral ClassificationRelationRelation Classification+1

Inplace Gated Convolutional Recurrent Neural Network For Dual-channel Speech Enhancement

2021-07-26 · Jinjiang Liu, Xueliang Zhang

For dual-channel speech enhancement, it is a promising idea to design an end-to-end model based on the traditional array signal processing guideline and the manifold space of multi-channel signals. We found that the idea…

Speech Enhancement