paper-with-me

홈 › Papers

MLP-based architecture with variable length input for automatic speech recognition

2021-09-29 · Jin Sakuma, Tatsuya Komatsu, Robin Scheibler

We propose multi-layer perceptron (MLP)-based architectures suitable for variable length input. Recently, several such architectures that do not rely on self-attention have been proposed for image classification. They achieve performance competitive with that of transformer-based architectures, albeit with a simpler structure and low computational cost. They split an image into patches and mix information by applying MLPs within and across patches alternately. Due to the use of MLPs, one such model can only be used for inputs of a fixed, pre-defined size. However, many types of data are naturally variable in length, for example acoustic signals. We propose three approaches to extend MLP-based architectures for use with sequences of arbitrary length. In all of them, we start by splitting the signal into contiguous tokens of fixed size (equivalent to patches in images). Naturally, the number of tokens is variable. The two first approaches use a gating mechanism that mixes local information across tokens in a shift-invariant and length-agnostic way. One uses a depthwise convolution to derive the gate values, while the other rely on shifting tokens. The final approach explores non-gated mixing using a circular convolution applied in the Fourier domain. We evaluate the proposed architectures on an automatic speech recognition task with the Librispeech and Tedlium2 corpora. Compared to Transformer, our proposed architecture reduces the WER by \SI{1.9 / 3.4}{\percent} on Librispeech test-clean/test-other set, and 1.8 / 1.6 % on Tedlium2 dev/test set, using only 75.3 % of the parameters.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)image-classificationImage Classificationspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…

Similar Papers 제목 키워드 기반

MLP-ASR: Sequence-length agnostic all-MLP architectures for speech recognition

2022-02-17 · Jin Sakuma, Tatsuya Komatsu, Robin Scheibler

We propose multi-layer perceptron (MLP)-based architectures suitable for variable length input. MLP-based architectures, recently proposed for image classification, can only be used for inputs of a fixed, pre-defined siz…

AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)image-classification+3

Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks

2023-09-14 · Sizhou Chen, Songyang Gao, Sen Fang

The Transformer architecture has proven to be highly effective for Automatic Speech Recognition (ASR) tasks, becoming a foundational component for a plethora of research in the domain. Historically, many approaches have …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Non-Autoregressive Chinese ASR Error Correction with Phonological Training

2022-07-01 · NAACL 2022 7 · Zheng Fang, Ruiqing Zhang, Zhongjun He, Hua Wu 외

Automatic Speech Recognition (ASR) is an efficient and widely used input method that transcribes speech signals into text. As the errors introduced by ASR systems will impair the performance of downstream tasks, we intro…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Sentencespeech-recognition+1

Variable-rate discrete representation learning

2021-03-10 · Sander Dieleman, Charlie Nash, Jesse Engel, Karen Simonyan

Semantically meaningful information content in perceptual signals is usually unevenly distributed. In speech signals for example, there are often many silences, and the speed of pronunciation can vary considerably. In th…

Representation Learning

Skipformer: A Skip-and-Recover Strategy for Efficient Speech Recognition

2024-03-13 · Wenjing Zhu, Sining Sun, Changhao Shan, Peng Fan 외

Conformer-based attention models have become the de facto backbone model for Automatic Speech Recognition tasks. A blank symbol is usually introduced to align the input and output sequences for CTC or RNN-T models. Unfor…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition