paper-with-me

홈 › Papers

Plug-and-Play Multilingual Few-shot Spoken Words Recognition

2023-05-03 · Aaqib Saeed, Vasileios Tsouvalas

As technology advances and digital devices become prevalent, seamless human-machine communication is increasingly gaining significance. The growing adoption of mobile, wearable, and other Internet of Things (IoT) devices has changed how we interact with these smart devices, making accurate spoken words recognition a crucial component for effective interaction. However, building robust spoken words detection system that can handle novel keywords remains challenging, especially for low-resource languages with limited training data. Here, we propose PLiX, a multilingual and plug-and-play keyword spotting system that leverages few-shot learning to harness massive real-world data and enable the recognition of unseen spoken words at test-time. Our few-shot deep models are learned with millions of one-second audio clips across 20 languages, achieving state-of-the-art performance while being highly efficient. Extensive evaluations show that PLiX can generalize to novel spoken words given as few as just one support example and performs well on unseen languages out of the box. We release models and inference code to serve as a foundation for future research and voice-enabled user interface development for emerging devices.

📄 PDF Abstract BibTeX arXiv:2305.03058

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot LearningKeyword Spotting

Similar Papers 제목 키워드 기반

Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition

2024-06-04 · Hao Yen, Pin-Jui Ku, Sabato Marco Siniscalchi, Chin-Hui Lee

We propose a novel language-universal approach to end-to-end automatic spoken keyword recognition (SKR) leveraging upon (i) a self-supervised pre-trained model, and (ii) a set of universal speech attributes (manner and p…

Attribute

ADIMA: Abuse Detection In Multilingual Audio

2022-02-16 · Vikram Gupta, Rini Sharon, Ramit Sawhney, Debdoot Mukherjee

Abusive content detection in spoken text can be addressed by performing Automatic Speech Recognition (ASR) and leveraging advancements in natural language processing. However, ASR models introduce latency and often perfo…

Abuse DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1

Enhancing Few-shot Keyword Spotting Performance through Pre-Trained Self-supervised Speech Models

2025-06-21 · Alican Gok, Oguzhan Buyuksolak, Osman Erman Okman, Murat Saraclar

Keyword Spotting plays a critical role in enabling hands-free interaction for battery-powered edge devices. Few-Shot Keyword Spotting (FS-KWS) addresses the scalability and adaptability challenges of traditional systems …

Dimensionality ReductionKeyword SpottingKnowledge DistillationSelf-Supervised Learning

Multilingual and Cross-Lingual Intent Detection from Spoken Data

2021-04-17 · EMNLP 2021 11 · Daniela Gerz, Pei-Hao Su, Razvan Kusztos, Avishek Mondal 외

We present a systematic study on multilingual and cross-lingual intent detection from spoken data. The study leverages a new resource put forth in this work, termed MInDS-14, a first training and evaluation resource for …

Few-Shot LearningIntent DetectionMachine TranslationSentence+3

Does language matter for spoken word classification? A multilingual generative meta-learning approach

2026-05-13 · Batsirayi Mupamhi Ziki, Louise Beyers, Ruan van der Merwe arxiv

Meta-learning has been shown to have better performance than supervised learning for few-shot monolingual spoken word classification. However, the meta-learning approach remains under-explored in multilingual spoken word…

Continual Learning