paper-with-me

홈 › Papers

MFCC based Enlargement of the Training Set for Emotion Recognition in Speech

2014-03-19 · Inma Mohino-Herranz, Roberto Gil-Pita, Sagrario Alonso-Diaz, Manuel Rosa-Zurera

Emotional state recognition through speech is being a very interesting research topic nowadays. Using subliminal information of speech, denominated as prosody, it is possible to recognize the emotional state of the person. One of the main problems in the design of automatic emotion recognition systems is the small number of available patterns. This fact makes the learning process more difficult, due to the generalization problems that arise under these conditions. In this work we propose a solution to this problem consisting in enlarging the training set through the creation the new virtual patterns. In the case of emotional speech, most of the emotional information is included in speed and pitch variations. So, a change in the average pitch that does not modify neither the speed nor the pitch variations does not affect the expressed emotion. Thus, we use this prior information in order to create new patterns applying a gender dependent pitch shift modification in the feature extraction process of the classification system. For this purpose, we propose a frequency scaling modification of the Mel Frequency Cepstral Coefficients, used to classify the emotion. For this purpose, we propose a gender dependent frequency scaling modification. This proposed process allows us to synthetically increase the number of available patterns in the training set, thus increasing the generalization capability of the system and reducing the test error. Results carried out with two different classifiers with different degree of generalization capability demonstrate the suitability of the proposal.

📄 PDF Abstract BibTeX arXiv:1403.4777

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion Recognition

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Evaluating Gammatone Frequency Cepstral Coefficients with Neural Networks for Emotion Recognition from Speech

2018-06-23 · Gabrielle K. Liu

Current approaches to speech emotion recognition focus on speech features that can capture the emotional content of a speech signal. Mel Frequency Cepstral Coefficients (MFCCs) are one of the most commonly used represent…

ClassificationEmotion RecognitionGeneral ClassificationSpeech Emotion Recognition+2

Speech Emotion Recognition Using MFCC Features and LSTM-Based Deep Learning Model

2026-04-16 · Adelekun Oluwademilade, Ademola Adedamola, Abiola Abdulhakeem, Akinpelu Azeezat 외 arxiv

Speech Emotion Recognition (SER) is the use of machines to detect the emotional state of humans based on the speech, which is gaining importance in natural human-computer interaction. Speech is a very valuable source of …

Speech Emotion Recognition

DNN-HMM based Speaker Adaptive Emotion Recognition using Proposed Epoch and MFCC Features

2018-06-04 · Md. Shah Fahad, Jainath Yadav, Gyadhar Pradhan, Akshay Deepak

Speech is produced when time varying vocal tract system is excited with time varying excitation source. Therefore, the information present in a speech such as message, emotion, language, speaker is due to the combined ef…

Emotion Recognition

Deep Learning based Emotion Recognition System Using Speech Features and Transcriptions

2019-06-11 · Suraj Tripathi, Abhay Kumar, Abhiram Ramesh, Chirag Singh 외

This paper proposes a speech emotion recognition method based on speech features and speech transcriptions (text). Speech features such as Spectrogram and Mel-frequency Cepstral Coefficients (MFCC) help retain emotion-re…

Emotion RecognitionSpeech Emotion Recognition

Deep scattering network for speech emotion recognition

2021-05-11 · Premjeet Singh, Goutam Saha, Md Sahidullah

This paper introduces scattering transform for speech emotion recognition (SER). Scattering transform generates feature representations which remain stable to deformations and shifting in time and frequency without much …

Emotion RecognitionSpeech Emotion Recognition