Towards Generalizable SER: Soft Labeling and Data Augmentation for Modeling Temporal Emotion Shifts in Large-Scale Multilingual Speech
Recognizing emotions in spoken communication is crucial for advanced human-machine interaction. Current emotion detection methodologies often display biases when applied cross-corpus. To address this, our study amalgamates 16 diverse datasets, resulting in 375 hours of data across languages like English, Chinese, and Japanese. We propose a soft labeling system to capture gradational emotional intensities. Using the Whisper encoder and data augmentation methods inspired by contrastive learning, our method emphasizes the temporal dynamics of emotions. Our validation on four multilingual datasets demonstrates notable zero-shot generalization. We publish our open source model weights and initial promising results after fine-tuning on Hume-Prosody.
Code (1)
Tasks
Contrastive LearningCross-corpusData AugmentationZero-shot GeneralizationSimilar Papers 제목 키워드 기반
Soft-labeling Strategies for Rapid Sub-Typing
The challenge of labeling large example datasets for computer vision continues to limit the availability and scope of image repositories. This research provides a new method for automated data collection, curation, label…
object-detectionObject DetectionDetecting Deepfakes with Multivariate Soft Blending and CLIP-based Image-Text Alignment
The proliferation of highly realistic facial forgeries necessitates robust detection methods. However, existing approaches often suffer from limited accuracy and poor generalization due to significant distribution shifts…
DeepFake DetectionE-Stitchup: Data Augmentation for Pre-Trained Embeddings
In this work, we propose data augmentation methods for embeddings from pre-trained deep learning models that take a weighted combination of a pair of input embeddings, as inspired by Mixup, and combine such augmentation …
Data AugmentationGeneral ClassificationTransfer LearningMAST: Masked Augmentation Subspace Training for Generalizable Self-Supervised Priors
Recent Self-Supervised Learning (SSL) methods are able to learn feature representations that are invariant to different data augmentations, which can then be transferred to downstream tasks of interest. However, differen…
Instance SegmentationSelf-Supervised LearningSemantic SegmentationGeneralizable Cone Beam CT Esophagus Segmentation Using Physics-Based Data Augmentation
Automated segmentation of esophagus is critical in image guided/adaptive radiotherapy of lung cancer to minimize radiation-induced toxicities such as acute esophagitis. We developed a semantic physics-based data augmenta…
Data AugmentationDomain Adaptation