Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
Current automatic speech recognition systems struggle with modeling long speech sequences due to high quadratic complexity of Transformer-based models. Selective state space models such as Mamba has performed well on long-sequence modeling in natural language processing and computer vision tasks. However, research endeavors in speech technology tasks has been under-explored. We propose Speech-Mamba, which incorporates selective state space modeling in Transformer neural architectures. Long sequence representations with selective state space models in Speech-Mamba is complemented with lower-level representations from Transformer-based modeling. Speech-mamba achieves better capacity to model long-range dependencies, as it scales near-linearly with sequence length.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionMambaspeech-recognitionSpeech RecognitionState Space ModelsMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
It is too early to conclude that Mamba is a better alternative to transformers for speech before comparing Mamba with transformers in terms of both performance and efficiency in multiple speech-related tasks. To reach th…
Mambaspeech-recognitionSpeech RecognitionSpeech Separation+1Rethinking Mamba in Speech Processing by Self-Supervised Models
The Mamba-based model has demonstrated outstanding performance across tasks in computer vision, natural language processing, and speech processing. However, in the realm of speech processing, the Mamba-based model's perf…
MambaSpeech Enhancementspeech-recognitionSpeech RecognitionMLMA: Towards Multilingual ASR With Mamba-based Architectures
Multilingual automatic speech recognition (ASR) remains a challenging task, especially when balancing performance across high- and low-resource languages. Recent advances in sequence modeling suggest that architectures b…
Speech RecognitionSPMamba: State-space model is all you need in speech separation
Existing CNN-based speech separation models face local receptive field limitations and cannot effectively capture long time dependencies. Although LSTM and Transformer-based speech separation models can avoid this proble…
AllMambaSpeech SeparationMamba in Speech: Towards an Alternative to Self-Attention
Transformer and its derivatives have achieved success in diverse tasks across computer vision, natural language processing, and speech processing. To reduce the complexity of computations within the multi-head self-atten…
MambaSpeech Enhancementspeech-recognitionSpeech Recognition+1