paper-with-me

홈 › Papers

Visually grounded cross-lingual keyword spotting in speech

2018-06-13 · Herman Kamper, Michael Roth

Recent work considered how images paired with speech can be used as supervision for building speech systems when transcriptions are not available. We ask whether visual grounding can be used for cross-lingual keyword spotting: given a text keyword in one language, the task is to retrieve spoken utterances containing that keyword in another language. This could enable searching through speech in a low-resource language using text queries in a high-resource language. As a proof-of-concept, we use English speech with German queries: we use a German visual tagger to add keyword labels to each training image, and then train a neural network to map English speech to German keywords. Without seeing parallel speech-transcriptions or translations, the model achieves a precision at ten of 58%. We show that most erroneous retrievals contain equivalent or semantically relevant keywords; excluding these would improve P@10 to 91%.

📄 PDF Abstract BibTeX arXiv:1806.05030

Code (0)

등록된 구현이 없습니다.

Tasks

Keyword SpottingVisual Grounding

Similar Papers 제목 키워드 기반

Visually Grounded Keyword Detection and Localisation for Low-Resource Languages

2023-02-01 · Kayode Kolawole Olaleye

This study investigates the use of Visually Grounded Speech (VGS) models for keyword localisation in speech. The study focusses on two main research questions: (1) Is keyword localisation possible with VGS models and (2)…

Few-Shot Keyword Spotting in Any Language

2021-04-03 · Mark Mazumder, Colby Banbury, Josh Meyer, Pete Warden 외

We introduce a few-shot transfer learning method for keyword spotting in any language. Leveraging open speech corpora in nine languages, we automate the extraction of a large multilingual keyword bank and use it to train…

Keyword SpottingTransfer Learning

Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling

2024-09-03 · Leanne Nortje

This dissertation examines visually grounded speech (VGS) models that learn from unlabelled speech paired with images. It focuses on applications for low-resource languages and understanding human language acquisition. W…

Few-Shot LearningLanguage Acquisition

Mapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding

2026-08-27 · Gabriel Pirlogeanu, Dan Oneata, Horia Cucu, Herman Kamper arxiv

In many low-resource settings, even just eliciting speech for data collection is difficult. One promising approach has been to ask speakers to describe images. But how do we build models from such visually grounded speec…

Visual GroundingImage CaptioningKeyword Spotting

ASR-free CNN-DTW keyword spotting using multilingual bottleneck features for almost zero-resource languages

2018-07-23 · Raghav Menon, Herman Kamper, Emre Yilmaz, John Quinn 외

We consider multilingual bottleneck features (BNFs) for nearly zero-resource keyword spotting. This forms part of a United Nations effort using keyword spotting to support humanitarian relief programmes in parts of Afric…

Dynamic Time WarpingHumanitarianKeyword SpottingTemplate Matching