paper-with-me

홈 › Papers

MLP-ASR: Sequence-length agnostic all-MLP architectures for speech recognition

2022-02-17 · Jin Sakuma, Tatsuya Komatsu, Robin Scheibler

We propose multi-layer perceptron (MLP)-based architectures suitable for variable length input. MLP-based architectures, recently proposed for image classification, can only be used for inputs of a fixed, pre-defined size. However, many types of data are naturally variable in length, for example, acoustic signals. We propose three approaches to extend MLP-based architectures for use with sequences of arbitrary length. The first one uses a circular convolution applied in the Fourier domain, the second applies a depthwise convolution, and the final relies on a shift operation. We evaluate the proposed architectures on an automatic speech recognition task with the Librispeech and Tedlium2 corpora. The best proposed MLP-based architectures improves WER by 1.0 / 0.9%, 0.9 / 0.5% on Librispeech dev-clean/dev-other, test-clean/test-other set, and 0.8 / 1.1% on Tedlium2 dev/test set using 86.4% the size of self-attention-based architecture.

📄 PDF Abstract BibTeX arXiv:2202.08456

Code (0)

등록된 구현이 없습니다.

Tasks

AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)image-classificationImage Classificationspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

MLP-based architecture with variable length input for automatic speech recognition

2021-09-29 · Jin Sakuma, Tatsuya Komatsu, Robin Scheibler

We propose multi-layer perceptron (MLP)-based architectures suitable for variable length input. Recently, several such architectures that do not rely on self-attention have been proposed for image classification. They a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)image-classificationImage Classification+2

Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models

2024-09-27 · Xiaoxue Gao, Nancy F. Chen

Current automatic speech recognition systems struggle with modeling long speech sequences due to high quadratic complexity of Transformer-based models. Selective state space models such as Mamba has performed well on lon…

Automatic Speech RecognitionMambaspeech-recognitionSpeech Recognition+1

Spike-Triggered Non-Autoregressive Transformer for End-to-End Speech Recognition

2020-05-16 · Zhengkun Tian, Jiangyan Yi, Jian-Hua Tao, Ye Bai 외

Non-autoregressive transformer models have achieved extremely fast inference speed and comparable performance with autoregressive sequence-to-sequence models in neural machine translation. Most of the non-autoregressive …

Machine Translationspeech-recognitionSpeech RecognitionTranslation

A Comparison of Transformer, Convolutional, and Recurrent Neural Networks on Phoneme Recognition

2022-10-01 · Kyuhong Shim, Wonyong Sung

Phoneme recognition is a very important part of speech recognition that requires the ability to extract phonetic features from multiple frames. In this paper, we compare and analyze CNN, RNN, Transformer, and Conformer m…

Phoneme Recognitionspeech-recognitionSpeech Recognition

BBPE16: UTF-16-based byte-level byte-pair encoding for improved multilingual speech recognition

2026-02-02 · Hyunsik Kim, Haeri Kim, Munhak Lee, Kyungmin Lee arxiv

Multilingual automatic speech recognition (ASR) requires tokenization that efficiently covers many writing systems. Byte-level BPE (BBPE) using UTF-8 is widely adopted for its language-agnostic design and full Unicode co…

Speech Recognition