paper-with-me

Papers

Intermediate-layer output Regularization for Attention-based Speech Recognition with Shared Decoder

2022-07-09 · Jicheng Zhang, Yizhou Peng, HaiHua Xu, Yi He, Eng Siong Chng, Hao Huang

Intermediate layer output (ILO) regularization by means of multitask training on encoder side has been shown to be an effective approach to yielding improved results on a wide range of end-to-end ASR frameworks. In this paper, we propose a novel method to do ILO regularized training differently. Instead of using conventional multitask methods that entail more training overhead, we directly make the intermediate layer output as input to the decoder, that is, our decoder not only accepts the output of the final encoder layer as input, it also takes the output of the encoder ILO as input during training. With the proposed method, as both encoder and decoder are simultaneously "regularized", the network is more sufficiently trained, consistently leading to improved results, over the ILO-based CTC method, as well as over the original attention-based modeling method without the proposed method employed.

📄 PDF Abstract BibTeX arXiv:2207.04177

Code (0)

등록된 구현이 없습니다.

Tasks

Decoderspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Focus on the present: a regularization method for the ASR source-target attention layer

2020-11-02 · Nanxin Chen, Piotr Żelasko, Jesús Villalba, Najim Dehak

This paper introduces a novel method to diagnose the source-target attention in state-of-the-art end-to-end speech recognition models with joint connectionist temporal classification (CTC) and attention training. Our met…

Decoderspeech-recognitionSpeech Recognition

Cross-Layer Distillation with Semantic Calibration

2020-12-06 · Defang Chen, Jian-Ping Mei, Yuan Zhang, Can Wang 외

Knowledge distillation is a technique to enhance the generalization ability of a student model by exploiting outputs from a teacher model. Recently, feature-map based variants explore knowledge transfer between manually …

Knowledge DistillationTransfer Learning

Deja-vu: Double Feature Presentation and Iterated Loss in Deep Transformer Networks

2019-10-23 · Andros Tjandra, Chunxi Liu, Frank Zhang, Xiaohui Zhang 외

Deep acoustic models typically receive features in the first layer of the network, and process increasingly abstract representations in the subsequent layers. Here, we propose to feed the input features at multiple depth…

Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss

2024-06-23 · Muhammad Shakeel, Yui Sudo, Yifan Peng, Shinji Watanabe

Contextualized end-to-end automatic speech recognition has been an active research area, with recent efforts focusing on the implicit learning of contextual phrases based on the final loss objective. However, these appro…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Gujarati-English Code-Switching Speech Recognition using ensemble prediction of spoken language

2024-03-12 · Yash Sharma, Basil Abraham, Preethi Jyothi

An important and difficult task in code-switched speech recognition is to recognize the language, as lots of words in two languages can sound similar, especially in some accents. We focus on improving performance of end-…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition