Integrating Contrastive Learning into a Multitask Transformer Model for Effective Domain Adaptation
While speech emotion recognition (SER) research has made significant progress, achieving generalization across various corpora continues to pose a problem. We propose a novel domain adaptation technique that embodies a multitask framework with SER as the primary task, and contrastive learning and information maximisation loss as auxiliary tasks, underpinned by fine-tuning of transformers pre-trained on large language models. Empirical results obtained through experiments on well-established datasets like IEMOCAP and MSP-IMPROV, illustrate that our proposed model achieves state-of-the-art performance in SER within cross-corpus scenarios.
Code (0)
등록된 구현이 없습니다.
Tasks
Contrastive LearningCross-corpusDomain AdaptationEmotion RecognitionSpeech Emotion RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Contrastive and Multi-Task Learning on Noisy Brain Signals with Nonlinear Dynamical Signatures
We introduce a two-stage multitask learning framework for analyzing Electroencephalography (EEG) signals that integrates denoising, dynamical modeling, and representation learning. In the first stage, a denoising autoenc…
Self-Supervised LearningRepresentation LearningMulti-Task LearningEeg DecodingMulT: An End-to-End Multitask Learning Transformer
We propose an end-to-end Multitask Learning Transformer framework, named MulT, to simultaneously learn multiple high-level vision tasks, including depth estimation, semantic segmentation, reshading, surface normal estima…
DecoderDepth EstimationEdge DetectionKeypoint Detection+2Vision Transformer Adapters for Generalizable Multitask Learning
We introduce the first multitasking vision transformer adapters that learn generalizable task affinities which can be applied to novel tasks and domains. Integrated into an off-the-shelf vision transformer backbone, our …
Domain AdaptationUnsupervised Domain AdaptationOmniVec2 - A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning
We present a novel multimodal multitask network and associated training algorithm. The method is capable of ingesting data from approximately 12 different modalities namely image video audio text depth point cloud ti…
3D Point Cloud ClassificationAction ClassificationAction RecognitionAudio Classification+6OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning
We present a novel multimodal multitask network and associated training algorithm. The method is capable of ingesting data from approximately 12 different modalities namely image, video, audio, text, depth, point cloud, …