paper-with-me

홈 › Papers

EmoAugNet: A Signal-Augmented Hybrid CNN-LSTM Framework for Speech Emotion Recognition

2025-08-06 · Durjoy Chandra Paul, Gaurob Saha, Md Amjad Hossain arxiv

Recognizing emotional signals in speech has a significant impact on enhancing the effectiveness of human-computer interaction (HCI). This study introduces EmoAugNet, a hybrid deep learning framework, that incorporates Long Short-Term Memory (LSTM) layers with one-dimensional Convolutional Neural Networks (1D-CNN) to enable reliable Speech Emotion Recognition (SER). The quality and variety of the features that are taken from speech signals have a significant impact on how well SER systems perform. A comprehensive speech data augmentation strategy was used to combine both traditional methods, such as noise addition, pitch shifting, and time stretching, with a novel combination-based augmentation pipeline to enhance generalization and reduce overfitting. Each audio sample was transformed into a high-dimensional feature vector using root mean square energy (RMSE), Mel-frequency Cepstral Coefficient (MFCC), and zero-crossing rate (ZCR). Our model with ReLU activation has a weighted accuracy of 95.78\% and unweighted accuracy of 92.52\% on the IEMOCAP dataset and, with ELU activation, has a weighted accuracy of 96.75\% and unweighted accuracy of 91.28\%. On the RAVDESS dataset, we get a weighted accuracy of 94.53\% and 94.98\% unweighted accuracy for ReLU activation and 93.72\% weighted accuracy and 94.64\% unweighted accuracy for ELU activation. These results highlight EmoAugNet's effectiveness in improving the robustness and performance of SER systems through integated data augmentation and hybrid modeling.

📄 PDF Abstract BibTeX arXiv:2508.06321

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Emotion RecognitionData Augmentation

Results from the Paper

RankTaskDatasetModelMetrics
#1 Emotion Recognition RAVDESS EmoAugNet Accuracy: 91.28
#2 Speech Emotion Recognition RAVDESS EmoAugNet Accuracy: 91.28

Similar Papers 제목 키워드 기반

Hybrid LSTM-UKF Framework: Ankle Angle and Ground Reaction Force Estimation

2026-01-10 · Mundla Narasimhappa, Praveen Kumar arxiv

Accurate prediction of joint kinematics and kinetics is essential for advancing gait analysis and developing intelligent assistive systems such as prosthetics and exoskeletons. This study presents a hybrid LSTM-UKF frame…

Classifying Objects in 3D Point Clouds Using Recurrent Neural Network: A GRU LSTM Hybrid Approach

2024-03-09 · Ramin Mousa, Mitra Khezli, Mohamadreza Azadi, Vahid Nikoofard 외

Accurate classification of objects in 3D point clouds is a significant problem in several applications, such as autonomous navigation and augmented/virtual reality scenarios, which has become a research hot spot. In this…

3D Object ClassificationAutonomous NavigationPoint Cloud Classification

LLM-Augmented Traffic Signal Control with LSTM-Based Traffic State Prediction and Safety-Constrained Decision Support

2026-04-26 · Jiazhao Shi arxiv

Traffic signal control is a critical task in intelligent transportation systems, yet conventional fixed-time and rule-based methods often struggle to adapt to dynamic traffic demand and provide limited decision interpret…

An Attention-Augmented VAE-BiLSTM Framework for Anomaly Detection in 12-Lead ECG Signals

2025-10-07 · Marc Garreta Basora, Mehmet Oguz Mulayim arxiv

Anomaly detection in 12-lead electrocardiograms (ECGs) is critical for identifying deviations associated with cardiovascular disease. This work presents a comparative analysis of three autoencoder-based architectures: co…

Unsupervised Anomaly Detection

Multi-Modal Hybrid Deep Neural Network for Speech Enhancement

2016-06-15 · Zhenzhou Wu, Sunil Sivadas, Yong Kiam Tan, Ma Bin 외

Deep Neural Networks (DNN) have been successful in en- hancing noisy speech signals. Enhancement is achieved by learning a nonlinear mapping function from the features of the corrupted speech signal to that of the refere…

Speech Enhancement