paper-with-me

Papers

Deep Implicit Distribution Alignment Networks for Cross-Corpus Speech Emotion Recognition

2023-02-17 · Yan Zhao, Jincen Wang, Yuan Zong, Wenming Zheng, Hailun Lian, Li Zhao

In this paper, we propose a novel deep transfer learning method called deep implicit distribution alignment networks (DIDAN) to deal with cross-corpus speech emotion recognition (SER) problem, in which the labeled training (source) and unlabeled testing (target) speech signals come from different corpora. Specifically, DIDAN first adopts a simple deep regression network consisting of a set of convolutional and fully connected layers to directly regress the source speech spectrums into the emotional labels such that the proposed DIDAN can own the emotion discriminative ability. Then, such ability is transferred to be also applicable to the target speech samples regardless of corpus variance by resorting to a well-designed regularization term called implicit distribution alignment (IDA). Unlike widely-used maximum mean discrepancy (MMD) and its variants, the proposed IDA absorbs the idea of sample reconstruction to implicitly align the distribution gap, which enables DIDAN to learn both emotion discriminative and corpus invariant features from speech spectrums. To evaluate the proposed DIDAN, extensive cross-corpus SER experiments on widely-used speech emotion corpora are carried out. Experimental results show that the proposed DIDAN can outperform lots of recent state-of-the-art methods in coping with the cross-corpus SER tasks.

📄 PDF Abstract BibTeX arXiv:2302.08921

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-corpusEmotion RecognitionSpeech Emotion RecognitionTransfer Learning

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Emo-DNA: Emotion Decoupling and Alignment Learning for Cross-Corpus Speech Emotion Recognition

2023-08-04 · Jiaxin Ye, Yujie Wei, Xin-Cheng Wen, Chenglong Ma 외

Cross-corpus speech emotion recognition (SER) seeks to generalize the ability of inferring speech emotion from a well-labeled corpus to an unlabeled one, which is a rather challenging task due to the significant discrepa…

Cross-corpusDomain AdaptationEmotion RecognitionSpeech Emotion Recognition+1

Filter-based multi-task cross-corpus feature learning for speech emotion recognition

2024-02-20 · Signal, Image and Video Processing 2024 2 · Behzad Bakhtiari, Elham Kalhor, Seyed Hossein Ghafarian

Speech emotion recognition is a highly active field of research in human–machine interaction. A primary challenge faced by researchers in this area is how to tackle the problem of changing data distribution. In the last…

Cross-corpusEmotion Recognitionfeature selectionMulti-Task Learning+1

Unsupervised Discovery of Recurring Speech Patterns Using Probabilistic Adaptive Metrics

2020-08-03 · Okko Räsänen, María Andrea Cruz Blandón

Unsupervised spoken term discovery (UTD) aims at finding recurring segments of speech from a corpus of acoustic speech data. One potential approach to this problem is to use dynamic time warping (DTW) to find well-aligni…

Dynamic Time Warping

LibriVoxDeEn: A Corpus for German-to-English Speech Translation and German Speech Recognition

2019-10-17 · LREC 2020 5 · Benjamin Beilharz, Xin Sun, Sariya Karimova, Stefan Riezler

We present a corpus of sentence-aligned triples of German audio, German text, and English translation, based on German audiobooks. The speech translation data consist of 110 hours of audio material aligned to over 50k pa…

Sentencespeech-recognitionSpeech RecognitionTranslation

HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation

2023-06-20 · Cihan Xiao, Henry Li Xinyuan, Jinyi Yang, Dongji Gao 외

We introduce HK-LegiCoST, a new three-way parallel corpus of Cantonese-English translations, containing 600+ hours of Cantonese audio, its standard traditional Chinese transcript, and English translation, segmented and a…

Cross-corpusSentencespeech-recognitionSpeech Recognition+1