Biologically inspired speech emotion recognition
Conventional feature-based classification methods do not apply well to automatic recognition of speech emotions, mostly because the precise set of spectral and prosodic features that is required to identify the emotional state of a speaker has not been determined yet. This paper presents a method that operates directly on the speech signal, thus avoiding the problematic step of feature extraction. Furthermore, this method combines the strengths of the classical source-filter model of human speech production with those of the recently introduced liquid state machine (LSM), a biologically-inspired spiking neural network (SNN). The source and vocal tract components of the speech signal are first separated and converted into perceptually relevant spectral representations. These representations are then processed separately by two reservoirs of neurons. The output of each reservoir is reduced in dimensionality and fed to a final classifier. This method is shown to provide very good classification performance on the Berlin Database of Emotional Speech (Emo-DB). This seems a very promising framework for solving efficiently many other problems in speech processing.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion RecognitionSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
Multimodal Emotion Recognition with High-level Speech and Text Features
Automatic emotion recognition is one of the central concerns of the Human-Computer Interaction field as it can bridge the gap between humans and machines. Current works train deep learning models on low-level data repres…
DisentanglementEmotion RecognitionMultimodal Emotion RecognitionRepresentation Learning+1Towards efficient end-to-end speech recognition with biologically-inspired neural networks
Automatic speech recognition (ASR) is a capability which enables a program to process human speech into a written form. Recent developments in artificial intelligence (AI) have led to high-accuracy ASR systems based on d…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Deep Learningspeech-recognition+1Multi-Classifier Interactive Learning for Ambiguous Speech Emotion Recognition
In recent years, speech emotion recognition technology is of great significance in industrial applications such as call centers, social robots and health care. The combination of speech recognition and speech emotion rec…
Emotion RecognitionSpeech Emotion Recognitionspeech-recognitionSpeech RecognitionCochCeps-Augment: A Novel Self-Supervised Contrastive Learning Using Cochlear Cepstrum-based Masking for Speech Emotion Recognition
Self-supervised learning (SSL) for automated speech recognition in terms of its emotional content, can be heavily degraded by the presence noise, affecting the efficiency of modeling the intricate temporal and spectral i…
Contrastive LearningEmotion RecognitionImage AugmentationSelf-Supervised Learning+4BioNIC: Biologically Inspired Neural Network for Image Classification Using Connectomics Principles
We present BioNIC, a multi-layer feedforward neural network for emotion classification, inspired by detailed synaptic connectivity graphs from the MICrONs dataset. At a structural level, we incorporate architectural cons…
Facial Emotion RecognitionEmotion ClassificationImage ClassificationData Augmentation