paper-with-me

Papers

Visually Grounded Keyword Detection and Localisation for Low-Resource Languages

2023-02-01 · Kayode Kolawole Olaleye

This study investigates the use of Visually Grounded Speech (VGS) models for keyword localisation in speech. The study focusses on two main research questions: (1) Is keyword localisation possible with VGS models and (2) Can keyword localisation be done cross-lingually in a real low-resource setting? Four methods for localisation are proposed and evaluated on an English dataset, with the best-performing method achieving an accuracy of 57%. A new dataset containing spoken captions in Yoruba language is also collected and released for cross-lingual keyword localisation. The cross-lingual model obtains a precision of 16% in actual keyword localisation and this performance can be improved by initialising from a model pretrained on English data. The study presents a detailed analysis of the model's success and failure modes and highlights the challenges of using VGS models for keyword localisation in low-resource settings.

📄 PDF Abstract BibTeX arXiv:2302.00765

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Attention-Based Keyword Localisation in Speech using Visual Grounding

2021-06-16 · Kayode Olaleye, Herman Kamper

Visually grounded speech models learn from images paired with spoken captions. By tagging images with soft text labels using a trained visual classifier with a fixed vocabulary, previous work has shown that it is possibl…

Visual Grounding

Towards visually prompted keyword localisation for zero-resource spoken languages

2022-10-12 · Leanne Nortje, Herman Kamper

Imagine being able to show a system a visual depiction of a keyword and finding spoken utterances that contain this keyword from a zero-resource speech corpus. We formalise this task and call it visually prompted keyword…

Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling

2024-09-03 · Leanne Nortje

This dissertation examines visually grounded speech (VGS) models that learn from unlabelled speech paired with images. It focuses on applications for low-resource languages and understanding human language acquisition. W…

Few-Shot LearningLanguage Acquisition

Keyword localisation in untranscribed speech using visually grounded speech models

2022-02-02 · Kayode Olaleye, Dan Oneata, Herman Kamper

Keyword localisation is the task of finding where in a speech utterance a given query keyword occurs. We investigate to what extent keyword localisation is possible using a visually grounded speech (VGS) model. VGS model…

Keyword SpottingTAG

Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings

2024-09-09 · Leanne Nortje, Dan Oneata, Herman Kamper

Given an image query, visually prompted keyword localisation (VPKL) aims to find occurrences of the depicted word in a speech collection. This can be useful when transcriptions are not available for a low-resource langua…

Few-Shot Learning