paper-with-me

홈 › Papers

Improving the fusion of acoustic and text representations in RNN-T

2022-01-25 · Chao Zhang, Bo Li, Zhiyun Lu, Tara N. Sainath, Shuo-Yiin Chang

The recurrent neural network transducer (RNN-T) has recently become the mainstream end-to-end approach for streaming automatic speech recognition (ASR). To estimate the output distributions over subword units, RNN-T uses a fully connected layer as the joint network to fuse the acoustic representations extracted using the acoustic encoder with the text representations obtained using the prediction network based on the previous subword units. In this paper, we propose to use gating, bilinear pooling, and a combination of them in the joint network to produce more expressive representations to feed into the output layer. A regularisation method is also proposed to enable better acoustic encoder training by reducing the gradients back-propagated into the prediction network at the beginning of RNN-T training. Experimental results on a multilingual ASR setting for voice search over nine languages show that the joint use of the proposed methods can result in 4%--5% relative word error rate reductions with only a few million extra parameters.

📄 PDF Abstract BibTeX arXiv:2201.10240

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Confusion2Vec: Towards Enriching Vector Space Word Representations with Representational Ambiguities

2018-11-08 · Prashanth Gurunath Shivakumar, Panayiotis Georgiou

Word vector representations are a crucial part of Natural Language Processing (NLP) and Human Computer Interaction. In this paper, we propose a novel word vector representation, Confusion2Vec, motivated from the human sp…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+2

Detecting Emotion Carriers by Combining Acoustic and Lexical Representations

2021-12-13 · Sebastian P. Bayerl, Aniruddha Tammewar, Korbinian Riedhammer, Giuseppe Riccardi

Personal narratives (PN) - spoken or written - are recollections of facts, people, events, and thoughts from one's own experience. Emotion recognition and sentiment analysis tasks are usually defined at the utterance or …

Emotion RecognitionNatural Language UnderstandingSentiment Analysis

A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

2026-05-25 · Loukas Ilias, Dimitris Askounis arxiv

Alzheimer's disease (AD) is a progressive neurodegenerative disorder and the leading cause of dementia, affecting memory, reasoning, communication, and daily functioning. Early diagnosis is particularly important, as tim…

Multimodal Deep LearningRepresentation Learning

Speech Emotion: Investigating Model Representations, Multi-Task Learning and Knowledge Distillation

2022-07-02 · Vikramjit Mitra, Hsiang-Yun Sherry Chien, Vasudha Kowtha, Joseph Yitan Cheng 외

Estimating dimensional emotions, such as activation, valence and dominance, from acoustic speech signals has been widely explored over the past few years. While accurate estimation of activation and dominance from speech…

Knowledge DistillationMulti-Task LearningValence Estimation

High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion Models

2023-09-27 · Chunyu Qiang, Hao Li, Yixin Tian, Yi Zhao 외

Text-to-speech (TTS) methods have shown promising results in voice cloning, but they require a large number of labeled text-speech pairs. Minimally-supervised speech synthesis decouples TTS by combining two types of disc…

AllSpeech Synthesistext-to-speechText to Speech+1