EmoDiarize: Speaker Diarization and Emotion Identification from Speech Signals using Convolutional Neural Networks
In the era of advanced artificial intelligence and human-computer interaction, identifying emotions in spoken language is paramount. This research explores the integration of deep learning techniques in speech emotion recognition, offering a comprehensive solution to the challenges associated with speaker diarization and emotion identification. It introduces a framework that combines a pre-existing speaker diarization pipeline and an emotion identification model built on a Convolutional Neural Network (CNN) to achieve higher precision. The proposed model was trained on data from five speech emotion datasets, namely, RAVDESS, CREMA-D, SAVEE, TESS, and Movie Clips, out of which the latter is a speech emotion dataset created specifically for this research. The features extracted from each sample include Mel Frequency Cepstral Coefficients (MFCC), Zero Crossing Rate (ZCR), Root Mean Square (RMS), and various data augmentation algorithms like pitch, noise, stretch, and shift. This feature extraction approach aims to enhance prediction accuracy while reducing computational complexity. The proposed model yields an unweighted accuracy of 63%, demonstrating remarkable efficiency in accurately identifying emotional states within speech signals.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationEmotion Recognitionspeaker-diarizationSpeaker DiarizationSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
Speech Emotion Diarization: Which Emotion Appears When?
Speech Emotion Recognition (SER) typically relies on utterance-level solutions. However, emotions conveyed through speech should be considered as discrete speech events with definite temporal boundaries, rather than attr…
Emotion Recognitionspeaker-diarizationSpeaker DiarizationSpeech Emotion RecognitionEnhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
In this paper, we investigate the impact of incorporating timestamp-based alignment between Automatic Speech Recognition (ASR) transcripts and Speaker Diarization (SD) outputs on Speech Emotion Recognition (SER) accuracy…
Multimodal Emotion RecognitionSpeech Emotion RecognitionSpeaker DiarizationSpeech RecognitionCompositional embedding models for speaker identification and diarization with simultaneous speech from 2+ speakers
We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compos…
speaker-diarizationSpeaker DiarizationSpeaker IdentificationMitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speaker to be misclassified as different indi…
Speaker DiarizationLarge-Scale Speaker Diarization of Radio Broadcast Archives
This paper describes our initial efforts to build a large-scale speaker diarization (SD) and identification system on a recently digitized radio broadcast archive from the Netherlands which has more than 6500 audio tapes…
speaker-diarizationSpeaker DiarizationSpeaker Identification