Augmenting conformers with structured state-space sequence models for online speech recognition
Online speech recognition, where the model only accesses context to the left, is an important and challenging use case for ASR systems. In this work, we investigate augmenting neural encoders for online ASR by incorporating structured state-space sequence models (S4), a family of models that provide a parameter-efficient way of accessing arbitrarily long left context. We performed systematic ablation studies to compare variants of S4 models and propose two novel approaches that combine them with convolutions. We found that the most effective design is to stack a small S4 using real-valued recurrent weights with a local convolution, allowing them to work complementarily. Our best model achieves WERs of 4.01%/8.53% on test sets from Librispeech, outperforming Conformers with extensively tuned convolution.
Code (0)
등록된 구현이 없습니다.
Tasks
speech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Normal-mode driven exploration of protein domain motions
Domain motions involved in the function of proteins can often be well described as a combination of motions along a handfull of low-frequency modes, that is, with the values of a few normal coordinates. This means that, …
First-principles molecular structure search with a genetic algorithm
The identification of low-energy conformers for a given molecule is a fundamental problem in computational chemistry and cheminformatics. We assess here a conformer search that employs a genetic algorithm for sampling th…
Computational chemistryMolMix: A Simple Yet Effective Baseline for Multimodal Molecular Representation Learning
In this work, we propose a simple transformer-based baseline for multimodal molecular representation learning, integrating three distinct modalities: SMILES strings, 2D graph representations, and 3D conformers of molecul…
molecular representationRepresentation LearningPre-Training Protein Encoder via Siamese Sequence-Structure Diffusion Trajectory Prediction
Self-supervised pre-training methods on proteins have recently gained attention, with most approaches focusing on either protein sequences or structures, neglecting the exploration of their joint distribution, which is c…
DenoisingTrajectory PredictionOn Time Domain Conformer Models for Monaural Speech Separation in Noisy Reverberant Acoustic Environments
Speech separation remains an important topic for multi-speaker technology researchers. Convolution augmented transformers (conformers) have performed well for many speech processing tasks but have been under-researched f…
Computational EfficiencySpeech Separation