On the Compression of Recurrent Neural Networks with an Application to LVCSR acoustic modeling for Embedded Speech Recognition
We study the problem of compressing recurrent neural networks (RNNs). In particular, we focus on the compression of RNN acoustic models, which are motivated by the goal of building compact and accurate speech recognition systems which can be run efficiently on mobile devices. In this work, we present a technique for general recurrent model compression that jointly compresses both recurrent and non-recurrent inter-layer weight matrices. We find that the proposed technique allows us to reduce the size of our Long Short-Term Memory (LSTM) acoustic model to a third of its original size with negligible loss in accuracy.
Code (0)
등록된 구현이 없습니다.
Tasks
Model Compressionspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
The IBM 2016 English Conversational Telephone Speech Recognition System
We describe a collection of acoustic and language modeling techniques that lowered the word error rate of our English conversational telephone LVCSR system to a record 6.6% on the Switchboard subset of the Hub5 2000 eval…
Language ModelingLanguage Modellingspeech-recognitionSpeech RecognitionLinguistic Search Optimization for Deep Learning Based LVCSR
Recent advances in deep learning based large vocabulary con- tinuous speech recognition (LVCSR) invoke growing demands in large scale speech transcription. The inference process of a speech recognizer is to find a sequen…
Deep Learningspeech-recognitionSpeech RecognitionTrace norm regularization and faster inference for embedded speech recognition RNNs
We propose and evaluate new techniques for compressing and speeding up dense matrix multiplications as found in the fully connected and recurrent layers of neural networks for embedded large vocabulary continuous speech …
speech-recognitionSpeech RecognitionEffects of Number of Filters of Convolutional Layers on Speech Recognition Model Accuracy
Inspired by the progress of the End-to-End approach [1], this paper systematically studies the effects of Number of Filters of convolutional layers on the model prediction accuracy of CNN+RNN (Convolutional Neural Networ…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2The RWTH Aachen LVCSR system for IWSLT-2016 German Skype conversation recognition task
In this paper the RWTH large vocabulary continuous speech recognition (LVCSR) systems developed for the IWSLT-2016 evaluation campaign are described. This evaluation campaign focuses on transcribing spontaneous speech fr…
Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition