paper-with-me

홈 › Papers

Leveraged Mel spectrograms using Harmonic and Percussive Components in Speech Emotion Recognition

2023-12-18 · Pacific-Asia Conference on Knowledge Discovery and Data Mining 2022 5 · David Hason Rudd, Huan Huo, Guandong Xu

Speech Emotion Recognition (SER) affective technology enables the intelligent embedded devices to interact with sensitivity. Similarly, call centre employees recognise customers' emotions from their pitch, energy, and tone of voice so as to modify their speech for a high-quality interaction with customers. This work explores, for the first time, the effects of the harmonic and percussive components of Mel spectrograms in SER. We attempt to leverage the Mel spectrogram by decomposing distinguishable acoustic features for exploitation in our proposed architecture, which includes a novel feature map generator algorithm, a CNN-based network feature extractor and a multi-layer perceptron (MLP) classifier. This study specifically focuses on effective data augmentation techniques for building an enriched hybrid-based feature map. This process results in a function that outputs a 2D image so that it can be used as input data for a pre-trained CNN-VGG16 feature extractor. Furthermore, we also investigate other acoustic features such as MFCCs, chromagram, spectral contrast, and the tonnetz to assess our proposed framework. A test accuracy of 92.79% on the Berlin EMO-DB database is achieved. Our result is higher than previous works using CNN-VGG16.

📄 PDF Abstract BibTeX arXiv:2312.10949

Code (1)

DavidHason/ser 공식 구현 tf

Tasks

Data AugmentationEmotion RecognitionSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

MS-SincResNet: Joint learning of 1D and 2D kernels using multi-scale SincNet and ResNet for music genre classification

2021-09-18 · Pei-Chun Chang, Yong-Sheng Chen, Chang-Hsing Lee

In this study, we proposed a new end-to-end convolutional neural network, called MS-SincResNet, for music genre classification. MS-SincResNet appends 1D multi-scale SincNet (MS-SincNet) to 2D ResNet as the first convolut…

Genre classificationMusic Genre Classification

A Case Study of Deep-Learned Activations via Hand-Crafted Audio Features

2019-07-03 · Olga Slizovskaia, Emilia Gómez, Gloria Haro

The explainability of Convolutional Neural Networks (CNNs) is a particularly challenging task in all areas of application, and it is notably under-researched in music and audio domain. In this paper, we approach explaina…

HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transform

2023-09-18 · Yinghao Aaron Li, Cong Han, Xilin Jiang, Nima Mesgarani

Recent advancements in speech synthesis have leveraged GAN-based networks like HiFi-GAN and BigVGAN to produce high-fidelity waveforms from mel-spectrograms. However, these networks are computationally expensive and para…

Speech Synthesis

Acoustic Scene Classification Using Bilinear Pooling on Time-liked and Frequency-liked Convolution Neural Network

2020-02-14 · Xing Yong Kek, Cheng Siong Chin, Ye Li

The current methodology in tackling Acoustic Scene Classification (ASC) task can be described in two steps, preprocessing of the audio waveform into log-mel spectrogram and then using it as the input representation for C…

Acoustic Scene ClassificationGeneral ClassificationInformation RetrievalMusic Information Retrieval+2

Adversarial Unsupervised Domain Adaptation for Harmonic-Percussive Source Separation

2021-01-03 · Carlos Lordelo, Emmanouil Benetos, Simon Dixon, Sven Ahlbäck 외

This paper addresses the problem of domain adaptation for the task of music source separation. Using datasets from two different domains, we compare the performance of a deep learning-based harmonic-percussive source sep…

Domain AdaptationMusic Source SeparationUnsupervised Domain Adaptation