paper-with-me

Papers

Attention-Based Keyword Localisation in Speech using Visual Grounding

2021-06-16 · Kayode Olaleye, Herman Kamper

Visually grounded speech models learn from images paired with spoken captions. By tagging images with soft text labels using a trained visual classifier with a fixed vocabulary, previous work has shown that it is possible to train a model that can detect whether a particular text keyword occurs in speech utterances or not. Here we investigate whether visually grounded speech models can also do keyword localisation: predicting where, within an utterance, a given textual keyword occurs without any explicit text-based or alignment supervision. We specifically consider whether incorporating attention into a convolutional model is beneficial for localisation. Although absolute localisation performance with visually supervised models is still modest (compared to using unordered bag-of-word text labels for supervision), we show that attention provides a large gain in performance over previous visually grounded models. As in many other speech-image studies, we find that many of the incorrect localisations are due to semantic confusions, e.g. locating the word 'backstroke' for the query keyword 'swimming'.

📄 PDF Abstract BibTeX arXiv:2106.08859

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

Towards visually prompted keyword localisation for zero-resource spoken languages

2022-10-12 · Leanne Nortje, Herman Kamper

Imagine being able to show a system a visual depiction of a keyword and finding spoken utterances that contain this keyword from a zero-resource speech corpus. We formalise this task and call it visually prompted keyword…

YFACC: A Yorùbá speech-image dataset for cross-lingual keyword localisation through visual grounding

2022-10-10 · Kayode Olaleye, Dan Oneata, Herman Kamper

Visually grounded speech (VGS) models are trained on images paired with unlabelled spoken captions. Such models could be used to build speech systems in settings where it is impossible to get labelled data, e.g. for docu…

Visual Grounding

Keyword localisation in untranscribed speech using visually grounded speech models

2022-02-02 · Kayode Olaleye, Dan Oneata, Herman Kamper

Keyword localisation is the task of finding where in a speech utterance a given query keyword occurs. We investigate to what extent keyword localisation is possible using a visually grounded speech (VGS) model. VGS model…

Keyword SpottingTAG

Visually Grounded Keyword Detection and Localisation for Low-Resource Languages

2023-02-01 · Kayode Kolawole Olaleye

This study investigates the use of Visually Grounded Speech (VGS) models for keyword localisation in speech. The study focusses on two main research questions: (1) Is keyword localisation possible with VGS models and (2)…

Towards localisation of keywords in speech using weak supervision

2020-12-14 · Kayode Olaleye, Benjamin van Niekerk, Herman Kamper

Developments in weakly supervised and self-supervised models could enable speech technology in low-resource settings where full transcriptions are not available. We consider whether keyword localisation is possible using…