paper-with-me

홈 › Papers

Transformer-based End-to-End Speech Recognition with Local Dense Synthesizer Attention

2020-10-23 · Menglong Xu, Shengqiang Li, Xiao-Lei Zhang

Recently, several studies reported that dot-product selfattention (SA) may not be indispensable to the state-of-theart Transformer models. Motivated by the fact that dense synthesizer attention (DSA), which dispenses with dot products and pairwise interactions, achieved competitive results in many language processing tasks, in this paper, we first propose a DSA-based speech recognition, as an alternative to SA. To reduce the computational complexity and improve the performance, we further propose local DSA (LDSA) to restrict the attention scope of DSA to a local range around the current central frame for speech recognition. Finally, we combine LDSA with SA to extract the local and global information simultaneously. Experimental results on the Ai-shell1 Mandarine speech recognition corpus show that the proposed LDSA-Transformer achieves a character error rate (CER) of 6.49%, which is slightly better than that of the SA-Transformer. Meanwhile, the LDSA-Transformer requires less computation than the SATransformer. The proposed combination method not only achieves a CER of 6.18%, which significantly outperforms the SA-Transformer, but also has roughly the same number of parameters and computational complexity as the latter. The implementation of the multi-head LDSA is available at https://github.com/mlxu995/multihead-LDSA.

📄 PDF Abstract BibTeX arXiv:2010.12155

Code (1)

mlxu995/multihead-LDSA 공식 구현 pytorch

Tasks

speech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Adam 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…

Similar Papers 제목 키워드 기반

Transformer-Based Speech Synthesizer Attribution in an Open Set Scenario

2022-10-14 · Emily R. Bartusiak, Edward J. Delp

Speech synthesis methods can create realistic-sounding speech, which may be used for fraud, spoofing, and misinformation campaigns. Forensic methods that detect synthesized speech are important for protection against suc…

AttributeMisinformationMulti-class ClassificationSpeech Synthesis

D²Net: A Denoising and Dereverberation Network Based on Two-branch Encoder and Dual-path Transformer

2022-11-21 · APSIPA ASC 2022 11 · Liusong Wang, Wenbing Wei, Yadong Chen, and Ying Hu

The simultaneous denoising and dereverberation for single-channel mixture speech under the complicated acoustic environment is considered to be a challengeable task. In this paper, we propose a denoising and dereverberat…

DenoisingSpeech Enhancement

Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation

2024-06-14 · Nameer Hirschkind, Xiao Yu, Mahesh Kumar Nandwana, Joseph Liu 외

We introduce DiffuseST, a low-latency, direct speech-to-speech translation system capable of preserving the input speaker's voice zero-shot while translating from multiple source languages into English. We experiment wit…

Speech-to-Speech TranslationTranslation

Transferring neural speech waveform synthesizers to musical instrument sounds generation

2019-10-27 · Yi Zhao, Xin Wang, Lauri Juvela, Junichi Yamagishi

Recent neural waveform synthesizers such as WaveNet, WaveGlow, and the neural-source-filter (NSF) model have shown good performance in speech synthesis despite their different methods of waveform generation. The similari…

Audio GenerationAudio SynthesisSpeech SynthesisZero-Shot Learning

Multi-modal Dense Video Captioning

2020-03-17 · Vladimir Iashin, Esa Rahtu

Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are so…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Dense Video Captioningspeech-recognition+1