paper-with-me

Papers

QbyE-MLPMixer: Query-by-Example Open-Vocabulary Keyword Spotting using MLPMixer

2022-06-23 · Jinmiao Huang, Waseem Gharbieh, Qianhui Wan, Han Suk Shim, Chul Lee

Current keyword spotting systems are typically trained with a large amount of pre-defined keywords. Recognizing keywords in an open-vocabulary setting is essential for personalizing smart device interaction. Towards this goal, we propose a pure MLP-based neural network that is based on MLPMixer - an MLP model architecture that effectively replaces the attention mechanism in Vision Transformers. We investigate different ways of adapting the MLPMixer architecture to the QbyE open-vocabulary keyword spotting task. Comparisons with the state-of-the-art RNN and CNN models show that our method achieves better performance in challenging situations (10dB and 6dB environments) on both the publicly available Hey-Snips dataset and a larger scale internal dataset with 400 speakers. Our proposed model also has a smaller number of parameters and MACs compared to the baseline models.

📄 PDF Abstract BibTeX arXiv:2206.13231

Code (0)

등록된 구현이 없습니다.

Tasks

Keyword Spotting

Similar Papers 제목 키워드 기반

Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning

2024-08-27 · Zhenyu Wang, Shuyu Kong, Li Wan, Biqiao Zhang 외

Existing keyword spotting (KWS) systems primarily rely on predefined keyword phrases. However, the ability to recognize customized keywords is crucial for tailoring interactions with intelligent devices. In this paper, w…

Keyword SpottingMulti-Task Learning

Few-Shot Open-Vocabulary Remote Sensing Segmentation via Textual Inversion

2026-07-28 · Junhyuk Heo, Junghwan Park arxiv

Open-vocabulary segmentation labels arbitrary categories from a text query without per-class training, yet on remote sensing imagery it underperforms on categories it handles reliably elsewhere. We find that much of this…

OpenScene: 3D Scene Understanding with Open Vocabularies

2022-11-28 · CVPR 2023 1 · Songyou Peng, Kyle Genova, Chiyu "Max" Jiang, Andrea Tagliasacchi 외

Traditional 3D scene understanding approaches rely on labeled 3D datasets to train a model for a single task with supervision. We propose OpenScene, an alternative approach where a model predicts dense features for 3D sc…

3D Open-Vocabulary Instance Segmentation3D Semantic SegmentationScene UnderstandingSemantic Segmentation

FlowOVD: Learning Generative Latent Flows for Zero-shot Open-vocabulary Detection

2026-05-30 · Yao Wei, Andrea Cavallaro, Changjae Oh arxiv

Open-vocabulary object detection (OVD) has achieved remarkable progress through large-scale vision-language pre-training. Existing methods, however, typically formulate OVD as a discriminative prediction problem, where d…

Object Detection

Object-Centric Open-Vocabulary Image-Retrieval with Aggregated Features

2023-09-26 · Hila Levi, Guy Heller, Dan Levi, Ethan Fetaya

The task of open-vocabulary object-centric image retrieval involves the retrieval of images containing a specified object of interest, delineated by an open-set text query. As working on large image datasets becomes stan…

Image RetrievalObjectRetrieval