paper-with-me

홈 › Papers

Noise-Robust Keyword Spotting through Self-supervised Pretraining

2024-03-27 · Jacob Mørk, Holger Severin Bovbjerg, Gergely Kiss, Zheng-Hua Tan

Voice assistants are now widely available, and to activate them a keyword spotting (KWS) algorithm is used. Modern KWS systems are mainly trained using supervised learning methods and require a large amount of labelled data to achieve a good performance. Leveraging unlabelled data through self-supervised learning (SSL) has been shown to increase the accuracy in clean conditions. This paper explores how SSL pretraining such as Data2Vec can be used to enhance the robustness of KWS models in noisy conditions, which is under-explored. Models of three different sizes are pretrained using different pretraining approaches and then fine-tuned for KWS. These models are then tested and compared to models trained using two baseline supervised learning methods, one being standard training using clean data and the other one being multi-style training (MTR). The results show that pretraining and fine-tuning on clean data is superior to supervised learning on clean data across all testing conditions, and superior to supervised MTR for testing conditions of SNR above 5 dB. This indicates that pretraining alone can increase the model's robustness. Finally, it is found that using noisy data for pretraining models, especially with the Data2Vec-denoising approach, significantly enhances the robustness of KWS models in noisy conditions.

📄 PDF Abstract BibTeX arXiv:2403.18560

Code (1)

aau-es-ml/ssl_noise-robust_kws 공식 구현 pytorch

Tasks

DenoisingKeyword SpottingSelf-Supervised Learning

Similar Papers 제목 키워드 기반

On-Device Constrained Self-Supervised Speech Representation Learning for Keyword Spotting via Knowledge Distillation

2023-07-06 · Gene-Ping Yang, Yue Gu, Qingming Tang, Dongsu Du 외

Large self-supervised models are effective feature extractors, but their application is challenging under on-device budget constraints and biased dataset collection, especially in keyword spotting. To address this, we pr…

Keyword SpottingKnowledge DistillationRepresentation LearningSpeech Representation Learning

On the Efficiency of Integrating Self-supervised Learning and Meta-learning for User-defined Few-shot Keyword Spotting

2022-04-01 · Wei-Tsung Kao, Yuan-Kuei Wu, Chia-Ping Chen, Zhi-Sheng Chen 외

User-defined keyword spotting is a task to detect new spoken terms defined by users. This can be viewed as a few-shot learning problem since it is unreasonable for users to define their desired keywords by providing many…

Few-Shot LearningKeyword SpottingMeta-LearningSelf-Supervised Learning

Toward noise-robust whisper keyword spotting on headphones with in-earcup microphone and curriculum learning

2025-02-01 · Qiaoyu Yang, Shuo Zhang, Chuan-Che Huang

The expanding feature set of modern headphones puts a challenge on the design of their control interface. Users may want to separately control each feature or quickly switch between modes that activate different features…

Keyword Spotting

Enhancing Few-shot Keyword Spotting Performance through Pre-Trained Self-supervised Speech Models

2025-06-21 · Alican Gok, Oguzhan Buyuksolak, Osman Erman Okman, Murat Saraclar

Keyword Spotting plays a critical role in enabling hands-free interaction for battery-powered edge devices. Few-Shot Keyword Spotting (FS-KWS) addresses the scalability and adaptability challenges of traditional systems …

Dimensionality ReductionKeyword SpottingKnowledge DistillationSelf-Supervised Learning

Low-resource keyword spotting using contrastively trained transformer acoustic word embeddings

2025-06-21 · Julian Herreilers, Christiaan Jacobs, Thomas Niesler

We introduce a new approach, the ContrastiveTransformer, that produces acoustic word embeddings (AWEs) for the purpose of very low-resource keyword spotting. The ContrastiveTransformer, an encoder-only model, directly op…

Keyword SpottingWord Embeddings