paper-with-me

Papers

Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech

2024-01-12 · Yu Xi, Baochen Yang, Hao Li, Jiaqi Guo, Kai Yu

Customizable keyword spotting (KWS) in continuous speech has attracted increasing attention due to its real-world application potential. While contrastive learning (CL) has been widely used to extract keyword representations, previous CL approaches all operate on pre-segmented isolated words and employ only audio-text representations matching strategy. However, for KWS in continuous speech, co-articulation and streaming word segmentation can easily yield similar audio patterns for different texts, which may consequently trigger false alarms. To address this issue, we propose a novel CL with Audio Discrimination (CLAD) approach to learning keyword representation with both audio-text matching and audio-audio discrimination ability. Here, an InfoNCE loss considering both audio-audio and audio-text CL data pairs is employed for each sliding window during training. Evaluations on the open-source LibriPhrase dataset show that the use of sliding-window level InfoNCE loss yields comparable performance compared to previous CL approaches. Furthermore, experiments on the continuous speech dataset LibriSpeech demonstrate that, by incorporating audio discrimination, CLAD achieves significant performance gain over CL without audio discrimination. Meanwhile, compared to two-stage KWS approaches, the end-to-end KWS with CLAD achieves not only better performance, but also significant speed-up.

📄 PDF Abstract BibTeX arXiv:2401.06485

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningKeyword SpottingText Matching

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음
InfoNCE 설명 없음

Similar Papers 제목 키워드 기반

Joint Multimodal Contrastive Learning for Robust Spoken Term Detection and Keyword Spotting

2025-12-16 · Ramesh Gundluru, Shubham Gupta, Sri Rama Murty K arxiv

Acoustic Word Embeddings (AWEs) improve the efficiency of speech retrieval tasks such as Spoken Term Detection (STD) and Keyword Spotting (KWS). However, existing approaches suffer from limitations, including unimodal su…

Contrastive LearningKeyword Spotting

DCCRN-KWS: an audio bias based model for noise robust small-footprint keyword spotting

2023-05-21 · Shubo Lv, Xiong Wang, Sining Sun, Long Ma 외

Real-world complex acoustic environments especially the ones with a low signal-to-noise ratio (SNR) will bring tremendous challenges to a keyword spotting (KWS) system. Inspired by the recent advances of neural speech en…

DenoisingKeyword SpottingMulti-Task LearningSmall-Footprint Keyword Spotting+3

LLM-Synth4KWS: Scalable Automatic Generation and Synthesis of Confusable Data for Custom Keyword Spotting

2025-05-29 · Pai Zhu, Quan Wang, Dhruuv Agarwal, Kurt Partridge

Custom keyword spotting (KWS) allows detecting user-defined spoken keywords from streaming audio. This is achieved by comparing the embeddings from voice enrollments and input audio. State-of-the-art custom KWS models ar…

Keyword Spottingtext-to-speechText to Speech

Predicting detection filters for small footprint open-vocabulary keyword spotting

2019-12-16 · Theodore Bluche, Thibault Gisselbrecht

In this paper, we propose a fully-neural approach to open-vocabulary keyword spotting, that allows the users to include a customizable voice interface to their device and that does not require task-specific data. We pres…

Keyword Spotting

PCOV-KWS: Multi-task Learning for Personalized Customizable Open Vocabulary Keyword Spotting

2026-03-05 · Jianan Pan, Kejie Huang arxiv

As advancements in technologies like Internet of Things (IoT), Automatic Speech Recognition (ASR), Speaker Verification (SV), and Text-to-Speech (TTS) lead to increased usage of intelligent voice assistants, the demand f…

Multi-class ClassificationSpeaker VerificationMulti-Task LearningSpeech Recognition