Unsupervised Representation Learning with Future Observation Prediction for Speech Emotion Recognition
Prior works on speech emotion recognition utilize various unsupervised learning approaches to deal with low-resource samples. However, these methods pay less attention to modeling the long-term dynamic dependency, which is important for speech emotion recognition. To deal with this problem, this paper combines the unsupervised representation learning strategy -- Future Observation Prediction (FOP), with transfer learning approaches (such as Fine-tuning and Hypercolumns). To verify the effectiveness of the proposed method, we conduct experiments on the IEMOCAP database. Experimental results demonstrate that our method is superior to currently advanced unsupervised learning strategies.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion RecognitionRepresentation LearningSpeech Emotion RecognitionTransfer LearningSimilar Papers 제목 키워드 기반
Unsupervised Video Representation Learning by Bidirectional Feature Prediction
This paper introduces a novel method for self-supervised video representation learning via feature prediction. In contrast to the previous methods that focus on future feature prediction, we argue that a supervisory sign…
Action RecognitionContrastive LearningPredictionRepresentation LearningAcquiring language from speech by learning to remember and predict
Classical accounts of child language learning invoke memory limits as a pressure to discover sparse, language-like representations of speech, while more recent proposals stress the importance of prediction for language l…
Predictive Learning: Using Future Representation Learning Variantial Autoencoder for Human Action Prediction
The unsupervised Pretraining method has been widely used in aiding human action recognition. However, existing methods focus on reconstructing the already present frames rather than generating frames which happen in futu…
Action RecognitionRepresentation LearningTemporal Action LocalizationRegularizing Contrastive Predictive Coding for Speech Applications
Self-supervised methods such as Contrastive predictive Coding (CPC) have greatly improved the quality of the unsupervised representations. These representations significantly reduce the amount of labeled data needed for …
Acoustic Unit DiscoveryAutomatic Speech RecognitionData Augmentationspeech-recognition+1Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders
We present Mockingjay as a new speech representation learning approach, where bidirectional Transformer encoders are pre-trained on a large amount of unlabeled speech. Previous speech representation methods learn through…
General ClassificationRepresentation LearningSentiment AnalysisSentiment Classification+2