paper-with-me

Papers

SCaLa: Supervised Contrastive Learning for End-to-End Speech Recognition

2021-10-08 · Li Fu, Xiaoxiao Li, Runyu Wang, Lu Fan, Zhengchen Zhang, Meng Chen, Youzheng Wu, Xiaodong He

End-to-end Automatic Speech Recognition (ASR) models are usually trained to optimize the loss of the whole token sequence, while neglecting explicit phonemic-granularity supervision. This could result in recognition errors due to similar-phoneme confusion or phoneme reduction. To alleviate this problem, we propose a novel framework based on Supervised Contrastive Learning (SCaLa) to enhance phonemic representation learning for end-to-end ASR systems. Specifically, we extend the self-supervised Masked Contrastive Predictive Coding (MCPC) to a fully-supervised setting, where the supervision is applied in the following way. First, SCaLa masks variable-length encoder features according to phoneme boundaries given phoneme forced-alignment extracted from a pre-trained acoustic model; it then predicts the masked features via contrastive learning. The forced-alignment can provide phoneme labels to mitigate the noise introduced by positive-negative pairs in self-supervised MCPC. Experiments on reading and spontaneous speech datasets show that our proposed approach achieves 2.8 and 1.4 points Character Error Rate (CER) absolute reductions compared to the baseline, respectively.

📄 PDF Abstract BibTeX arXiv:2110.04187

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive LearningRepresentation Learningspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음
InfoNCE 설명 없음
Contrastive Predictive Coding Contrastive Predictive Coding (CPC) learns self-supervised representations by predicting the future in latent space by using powerful autoregressive models. The model uses a…

Similar Papers 제목 키워드 기반

Supervised Contrastive Learning for Accented Speech Recognition

2021-07-02 · Tao Han, Hantao Huang, Ziang Yang, Wei Han

Neural network based speech recognition systems suffer from performance degradation due to accented speech, especially unfamiliar accents. In this paper, we study the supervised contrastive learning framework for accente…

Accented Speech RecognitionContrastive LearningData AugmentationSentence+2

A Cross-Corpus Speech Emotion Recognition Method Based on Supervised Contrastive Learning

2024-11-25 · Xiang minjie

Research on Speech Emotion Recognition (SER) often faces challenges such as the lack of large-scale public datasets and limited generalization capability when dealing with data from different distributions. To solve this…

Contrastive LearningCross-corpusEmotion RecognitionSpeech Emotion Recognition

Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning

2020-10-27 · Dongwei Jiang, Wubo Li, Miao Cao, Wei Zou 외

Self-supervised visual pretraining has shown significant progress recently. Among those methods, SimCLR greatly advanced the state of the art in self-supervised and semi-supervised learning on ImageNet. The input feature…

Emotion RecognitionRepresentation LearningSpeech Emotion Recognitionspeech-recognition+2

Non-Contrastive Self-Supervised Learning of Utterance-Level Speech Representations

2022-08-10 · Jaejin Cho, Raghavendra Pappagari, Piotr Żelasko, Laureano Moro-Velazquez 외

Considering the abundance of unlabeled speech data and the high labeling costs, unsupervised learning methods can be essential for better system development. One of the most successful methods is contrastive self-supervi…

Emotion RecognitionSelf-Supervised LearningSpeaker Verification

CochCeps-Augment: A Novel Self-Supervised Contrastive Learning Using Cochlear Cepstrum-based Masking for Speech Emotion Recognition

2024-02-10 · Ioannis Ziogas, Hessa Alfalahi, Ahsan H. Khandoker, Leontios J. Hadjileontiadis

Self-supervised learning (SSL) for automated speech recognition in terms of its emotional content, can be heavily degraded by the presence noise, affecting the efficiency of modeling the intricate temporal and spectral i…

Contrastive LearningEmotion RecognitionImage AugmentationSelf-Supervised Learning+4