paper-with-me

홈 › Papers

Fully Convolutional Speech Recognition

2018-12-17 · Neil Zeghidour, Qiantong Xu, Vitaliy Liptchinsky, Nicolas Usunier, Gabriel Synnaeve, Ronan Collobert

Current state-of-the-art speech recognition systems build on recurrent neural networks for acoustic and/or language modeling, and rely on feature extraction pipelines to extract mel-filterbanks or cepstral coefficients. In this paper we present an alternative approach based solely on convolutional neural networks, leveraging recent advances in acoustic models from the raw waveform and language modeling. This fully convolutional approach is trained end-to-end to predict characters from the raw waveform, removing the feature extraction step altogether. An external convolutional language model is used to decode words. On Wall Street Journal, our model matches the current state-of-the-art. On Librispeech, we report state-of-the-art performance among end-to-end models, including Deep Speech 2 trained with 12 times more acoustic data and significantly more linguistic data.

📄 PDF Abstract BibTeX arXiv:1812.06864

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Attention Based Fully Convolutional Network for Speech Emotion Recognition

2018-06-05 · Yuanyuan Zhang, Jun Du, Zi-Rui Wang, Jianshu Zhang

Speech emotion recognition is a challenging task for three main reasons: 1) human emotion is abstract, which means it is hard to distinguish; 2) in general, human emotion can only be detected in some specific moments dur…

Emotion RecognitionSpeech Emotion RecognitionTransfer Learning

Light-SERNet: A lightweight fully convolutional neural network for speech emotion recognition

2021-10-07 · Arya Aftab, Alireza Morsali, Shahrokh Ghaemmaghami, Benoit Champagne

Detecting emotions directly from a speech signal plays an important role in effective human-computer interactions. Existing speech emotion recognition models require massive computational and storage resources, making th…

Emotion RecognitionSpeech Emotion Recognition

Performance Evaluation of Deep Convolutional Maxout Neural Network in Speech Recognition

2021-05-04 · Arash Dehghani, Seyyed Ali Seyyedsalehi

In this paper, various structures and methods of Deep Artificial Neural Networks (DNN) will be evaluated and compared for the purpose of continuous Persian speech recognition. One of the first models of neural networks u…

speech-recognitionSpeech Recognition

Which French speech recognition system for assistant robots?

2022-03-04 · IEEE 2022 3 · Wiam FADEL, Imane ARAF, Toumi BOUCHENTOUF, Pierre-André BUVET 외

Artificial intelligence-based speech recognition systems are already available and capable of recognizing the French language. Still, it is quite time-consuming to compare which one will be effective for an assistant rob…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Persian Signature Verification using Fully Convolutional Networks

2019-09-20 · Mohammad Rezaei, Nader Naderi

Fully convolutional networks (FCNs) have been recently used for feature extraction and classification in image and speech recognition, where their inputs have been raw signal or other complicated features. Persian signat…

speech-recognitionSpeech Recognition