paper-with-me

Papers

Speech Command Recognition in Computationally Constrained Environments with a Quadratic Self-organized Operational Layer

2020-11-23 · Mohammad Soltanian, Junaid Malik, Jenni Raitoharju, Alexandros Iosifidis, Serkan Kiranyaz, Moncef Gabbouj

Automatic classification of speech commands has revolutionized human computer interactions in robotic applications. However, employed recognition models usually follow the methodology of deep learning with complicated networks which are memory and energy hungry. So, there is a need to either squeeze these complicated models or use more efficient light-weight models in order to be able to implement the resulting classifiers on embedded devices. In this paper, we pick the second approach and propose a network layer to enhance the speech command recognition capability of a lightweight network and demonstrate the result via experiments. The employed method borrows the ideas of Taylor expansion and quadratic forms to construct a better representation of features in both input and hidden layers. This richer representation results in recognition accuracy improvement as shown by extensive experiments on Google speech commands (GSC) and synthetic speech commands (SSC) datasets.

📄 PDF Abstract BibTeX arXiv:2011.11436

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Advancing Airport Tower Command Recognition: Integrating Squeeze-and-Excitation and Broadcasted Residual Learning

2024-06-26 · Yuanxi Lin, Tonglin Zhou, Yang Xiao

Accurate recognition of aviation commands is vital for flight safety and efficiency, as pilots must follow air traffic control instructions precisely. This paper addresses challenges in speech command recognition, such a…

Keyword Spotting

Moonshine: Speech Recognition for Live Transcription and Voice Commands

2024-10-21 · Nat Jeffries, Evan King, Manjunath Kudlur, Guy Nicholson 외

This paper introduces Moonshine, a family of speech recognition models optimized for live transcription and voice command processing. Moonshine is based on an encoder-decoder transformer architecture and employs Rotary P…

DecoderPositionspeech-recognitionSpeech Recognition

Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning

2025-04-26 · Sidahmed Lachenani, Hamza Kheddar, Mohamed Ouldzmirli

This work addresses the need for enhanced accuracy and efficiency in speech command recognition systems, a critical component for improving user interaction in various smart applications. Leveraging the robust pretrained…

Transfer Learning

CI-AVSR: A Cantonese Audio-Visual Speech Datasetfor In-car Command Recognition

2022-06-01 · LREC 2022 6 · Wenliang Dai, Samuel Cahyawijaya, Tiezheng Yu, Elham J. Barezi 외

With the rise of deep learning and intelligent vehicles, the smart assistant has become an essential in-car component to facilitate driving and provide extra functionalities. In-car smart assistants should be able to pro…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

CI-AVSR: A Cantonese Audio-Visual Speech Dataset for In-car Command Recognition

2022-01-11 · Wenliang Dai, Samuel Cahyawijaya, Tiezheng Yu, Elham J. Barezi 외

With the rise of deep learning and intelligent vehicle, the smart assistant has become an essential in-car component to facilitate driving and provide extra functionalities. In-car smart assistants should be able to proc…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition