ROBUST SPEECH COMMAND RECOGNITION USING LABEL-DRIVEN TIME-FREQUENCY MASKING
Speech enhancement driven robust Automatic Speech Recognition (ASR) systems typically require parallel corpus with noisy and clean speech utterances for training. Moreover, many studies have reported that such front-ends, even though improve speech quality, do not always improve the recognition performance. On the other hand, the multi-condition training of ASR systems have little visualization or interpretability capabilities of how these systems achieve robustness. In this paper, we propose a novel neural architecture with unified enhancement and sequence classification block, that is trained in an end-to-end manner only using noisy speech without having information of clean speech. The enhancement block is a fully convolutional network that is designed to perform Time Frequency (T-F) masking like operation, followed by an LSTM sequence classification block. The T-F masking formulation enables visualization of learned mask and helps us to visualize the T-F points important for classification of a speech command. Experiments performed on Google Speech Command dataset show that our proposed network achieves better results than the baseline model without an enhancement front-end.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClassificationSpeech Enhancementspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
Spiking neural networks (SNNs) offer a promising path toward energy-efficient speech command recognition (SCR) by leveraging their event-driven processing paradigm. However, existing SNN-based SCR methods often struggle …
Efficient Speech Command Recognition Leveraging Spiking Neural Network and Curriculum Learning-based Knowledge Distillation
The intrinsic dynamics and event-driven nature of spiking neural networks (SNNs) make them excel in processing temporal information by naturally utilizing embedded time sequences as time steps. Recent studies adopting th…
Edge-computingKnowledge DistillationRepresentation LearningImproving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning
This work addresses the need for enhanced accuracy and efficiency in speech command recognition systems, a critical component for improving user interaction in various smart applications. Leveraging the robust pretrained…
Transfer LearningA neural attention model for speech command recognition
This paper introduces a convolutional recurrent network with attention for speech command recognition. Attention models are powerful tools to improve performance on natural language, image captioning and speech tasks. Th…
Image CaptioningmodelAdvancing Airport Tower Command Recognition: Integrating Squeeze-and-Excitation and Broadcasted Residual Learning
Accurate recognition of aviation commands is vital for flight safety and efficiency, as pilots must follow air traffic control instructions precisely. This paper addresses challenges in speech command recognition, such a…
Keyword Spotting