paper-with-me

홈 › Papers

How do Hyenas deal with Human Speech? Speech Recognition and Translation with ConfHyena

2024-02-20 · Marco Gaido, Sara Papi, Matteo Negri, Luisa Bentivogli

The attention mechanism, a cornerstone of state-of-the-art neural models, faces computational hurdles in processing long sequences due to its quadratic complexity. Consequently, research efforts in the last few years focused on finding more efficient alternatives. Among them, Hyena (Poli et al., 2023) stands out for achieving competitive results in both language modeling and image classification, while offering sub-quadratic memory and computational complexity. Building on these promising results, we propose ConfHyena, a Conformer whose encoder self-attentions are replaced with an adaptation of Hyena for speech processing, where the long input sequences cause high computational costs. Through experiments in automatic speech recognition (for English) and translation (from English into 8 target languages), we show that our best ConfHyena model significantly reduces the training time by 27%, at the cost of minimal quality degradation (~1%), which, in most cases, is not statistically significant.

📄 PDF Abstract BibTeX arXiv:2402.13208

Code (1)

hlt-mt/fbk-fairseq 공식 구현 pytorch

Tasks

Automatic Speech Recognitionimage-classificationImage ClassificationLanguage ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Voice based self help System: User Experience Vs Accuracy

2015-04-07 · Sunil Kumar Kopparapu

In general, self help systems are being increasingly deployed by service based industries because they are capable of delivering better customer service and increasingly the switch is to voice based self help systems bec…

speech-recognitionSpeech RecognitionSpeech-to-Text

Visual Speech Recognition

2014-09-03 · Ahmad B. A. Hassanat

Lip reading is used to understand or interpret speech without hearing it, a technique especially mastered by people with hearing difficulties. The ability to lip read enables a person with a hearing impairment to communi…

Audio-Visual Speech RecognitionLip Readingobject-detectionObject Detection+5

Opportunities & Challenges In Automatic Speech Recognition

2013-05-09 · Rashmi Makhijani, Urmila Shrawankar, V. M. Thakare

Automatic speech recognition enables a wide range of current and emerging applications such as automatic transcription, multimedia content analysis, and natural human-computer interfaces. This paper provides a glimpse of…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

From Audio to Symbolic Encoding

2023-02-26 · Shenli Yuan, Lingjie Kong, Jiushuang Guo

Automatic music transcription (AMT) aims to convert raw audio to symbolic music representation. As a fundamental problem of music information retrieval (MIR), AMT is considered a difficult task even for trained human exp…

Information RetrievalMusic Information RetrievalMusic TranscriptionRetrieval+2

Batch-normalized joint training for DNN-based distant speech recognition

2017-03-24 · Mirco Ravanelli, Philemon Brakel, Maurizio Omologo, Yoshua Bengio

Improving distant speech recognition is a crucial step towards flexible human-machine interfaces. Current technology, however, still exhibits a lack of robustness, especially when adverse acoustic conditions are met. Des…

Distant Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition