paper-with-me

홈 › Papers

Visually grounded few-shot word learning in low-resource settings

2023-06-20 · Leanne Nortje, Dan Oneata, Herman Kamper

We propose a visually grounded speech model that learns new words and their visual depictions from just a few word-image example pairs. Given a set of test images and a spoken query, we ask the model which image depicts the query word. Previous work has simplified this few-shot learning problem by either using an artificial setting with digit word-image pairs or by using a large number of examples per class. Moreover, all previous studies were performed using English speech-image data. We propose an approach that can work on natural word-image pairs but with less examples, i.e. fewer shots, and then illustrate how this approach can be applied for multimodal few-shot learning in a real low-resource language, Yor\ub\'a. Our approach involves using the given word-image example pairs to mine new unsupervised word-image training pairs from large collections of unlabelled speech and images. Additionally, we use a word-to-image attention mechanism to determine word-image similarity. With this new model, we achieve better performance with fewer shots than previous approaches on an existing English benchmark. Many of the model's mistakes are due to confusion between visual concepts co-occurring in similar contexts. The experiments on Yor\ub\'a show the benefit of transferring knowledge from a multimodal model trained on a larger set of English speech-image data.

📄 PDF Abstract BibTeX arXiv:2306.11371

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot Learning

Similar Papers 제목 키워드 기반

Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings

2024-09-09 · Leanne Nortje, Dan Oneata, Herman Kamper

Given an image query, visually prompted keyword localisation (VPKL) aims to find occurrences of the depicted word in a speech collection. This can be useful when transcriptions are not available for a low-resource langua…

Few-Shot Learning

Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling

2024-09-03 · Leanne Nortje

This dissertation examines visually grounded speech (VGS) models that learn from unlabelled speech paired with images. It focuses on applications for low-resource languages and understanding human language acquisition. W…

Few-Shot LearningLanguage Acquisition

Visually Grounded Keyword Detection and Localisation for Low-Resource Languages

2023-02-01 · Kayode Kolawole Olaleye

This study investigates the use of Visually Grounded Speech (VGS) models for keyword localisation in speech. The study focusses on two main research questions: (1) Is keyword localisation possible with VGS models and (2)…

Visually grounded few-shot word acquisition with fewer shots

2023-05-25 · Leanne Nortje, Benjamin van Niekerk, Herman Kamper

We propose a visually grounded speech model that acquires new words and their visual depictions from just a few word-image example pairs. Given a set of test images and a spoken query, we ask the model which image depict…

VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot Learning

2022-03-20 · CVPR 2022 1 · Wenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele 외

Human-annotated attributes serve as powerful semantic embeddings in zero-shot learning. However, their annotation process is labor-intensive and needs expert supervision. Current unsupervised semantic embeddings, i.e., w…

Transfer LearningWord EmbeddingsZero-Shot Learning