paper-with-me

Papers

Frequency & Channel Attention Network for Small Footprint Noisy Spoken Keyword Spotting

2024-07-29 · Yuanxi Lin, Yuriy Evgenyevich Gapanyuk

In this paper, we aim to improve the robustness of Keyword Spotting (KWS) systems in noisy environments while keeping a small memory footprint. We propose a new convolutional neural network (CNN) called FCA-Net, which combines mixer unit-based feature interaction with a two-dimensional convolution-based attention module. First, we introduce and compare lightweight attention methods to enhance noise robustness in CNN. Then, we propose an attention module that creates fine-grained attention weights to capture channel and frequency-specific information, boosting the model's ability to handle noisy conditions. By combining the mixer unit-based feature interaction with the attention module, we enhance performance. Additionally, we use a curriculum-based multi-condition training strategy. Our experiments show that our system outperforms current state-of-the-art solutions for small-footprint KWS in noisy environments, making it reliable for real-world use.

📄 PDF Abstract BibTeX arXiv:2407.19834

Code (0)

등록된 구현이 없습니다.

Tasks

Keyword Spotting

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Improving Deep Learning-based Respiratory Sound Analysis with Frequency Selection and Attention Mechanism

2025-07-26 · Nouhaila Fraihi, Ouassim Karrakchou, Mounir Ghogho arxiv

Accurate classification of respiratory sounds requires deep learning models that effectively capture fine-grained acoustic features and long-range temporal dependencies. Convolutional Neural Networks (CNNs) are well-suit…

PilotWiMAE: Pilot-Native Representation Learning for Wireless Channels

2026-05-19 · Berkay Guler, Giovanni Geraci, Hamid Jafarkhani arxiv

Channel foundation models assume access to fully observed channels, an assumption that fails in deployment. We introduce PilotWiMAE, a self-supervised framework whose encoder ingests noisy pilot observations directly and…

Representation Learning

Versatile and Efficient Medical Image Super-Resolution Via Frequency-Gated Mamba

2025-10-31 · Wenfeng Huang, Xiangyun Liao, Wei Cao, Wenjing Jia 외 arxiv

Medical image super-resolution (SR) is essential for enhancing diagnostic accuracy while reducing acquisition cost and scanning time. However, modeling both long-range anatomical structures and fine-grained frequency det…

Medical Image EnhancementImage Super-Resolution

Filterbank Learning for Noise-Robust Small-Footprint Keyword Spotting

2022-11-19 · Iván López-Espejo, Ram C. M. C. Shekar, Zheng-Hua Tan, Jesper Jensen 외

In the context of keyword spotting (KWS), the replacement of handcrafted speech features by learnable features has not yielded superior KWS performance. In this study, we demonstrate that filterbank learning outperforms …

Keyword SpottingSmall-Footprint Keyword Spotting

Cleanformer: A multichannel array configuration-invariant neural enhancement frontend for ASR in smart speakers

2022-04-25 · Joseph Caroselli, Arun Narayanan, Nathan Howard, Tom O'Malley

This work introduces the Cleanformer, a streaming multichannel neural based enhancement frontend for automatic speech recognition (ASR). This model has a conformer-based architecture which takes as inputs a single channe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition