Extending RNN-T-based speech recognition systems with emotion and language classification
Speech transcription, emotion recognition, and language identification are usually considered to be three different tasks. Each one requires a different model with a different architecture and training process. We propose using a recurrent neural network transducer (RNN-T)-based speech-to-text (STT) system as a common component that can be used for emotion recognition and language identification as well as for speech recognition. Our work extends the STT system for emotion classification through minimal changes, and shows successful results on the IEMOCAP and MELD datasets. In addition, we demonstrate that by adding a lightweight component to the RNN-T module, it can also be used for language identification. In our evaluations, this new classifier demonstrates state-of-the-art accuracy for the NIST-LRE-07 dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion ClassificationEmotion RecognitionLanguage Identificationspeech-recognitionSpeech RecognitionSpeech-to-TextSimilar Papers 제목 키워드 기반
Cross Lingual Speech Emotion Recognition: Urdu vs. Western Languages
Cross-lingual speech emotion recognition is an important task for practical applications. The performance of automatic speech emotion recognition systems degrades in cross-corpus scenarios, particularly in scenarios invo…
Cross-corpusEmotion RecognitionSpeech Emotion RecognitionCross Lingual Cross Corpus Speech Emotion Recognition
The majority of existing speech emotion recognition models are trained and evaluated on a single corpus and a single language setting. These systems do not perform as well when applied in a cross-corpus and cross-languag…
Cross-corpusEmotion RecognitionMulti-Task LearningSpeech Emotion RecognitionDecoding Emotions: A comprehensive Multilingual Study of Speech Models for Speech Emotion Recognition
Recent advancements in transformer-based speech representation models have greatly transformed speech processing. However, there has been limited research conducted on evaluating these models for speech emotion recogniti…
Emotion RecognitionSpeech Emotion RecognitionMulti-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
Speech Emotion Recognition (SER) is crucial for improving human-computer interaction. Despite strides in monolingual SER, extending them to build a multilingual system remains challenging. Our goal is to train a single m…
Emotion RecognitionKnowledge DistillationSpeech Emotion RecognitionTransfer Learning for Improving Speech Emotion Classification Accuracy
The majority of existing speech emotion recognition research focuses on automatic emotion detection using training and testing data from same corpus collected under the same conditions. The performance of such systems ha…
ClassificationCross-corpusEmotion ClassificationEmotion Recognition+3