paper-with-me

홈 › Papers

Visually grounded few-shot word acquisition with fewer shots

2023-05-25 · Leanne Nortje, Benjamin van Niekerk, Herman Kamper

We propose a visually grounded speech model that acquires new words and their visual depictions from just a few word-image example pairs. Given a set of test images and a spoken query, we ask the model which image depicts the query word. Previous work has simplified this problem by either using an artificial setting with digit word-image pairs or by using a large number of examples per class. We propose an approach that can work on natural word-image pairs but with less examples, i.e. fewer shots. Our approach involves using the given word-image example pairs to mine new unsupervised word-image training pairs from large collections of unlabelled speech and images. Additionally, we use a word-to-image attention mechanism to determine word-image similarity. With this new model, we achieve better performance with fewer shots than any existing approach.

📄 PDF Abstract BibTeX arXiv:2305.15937

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Visually grounded few-shot word learning in low-resource settings

2023-06-20 · Leanne Nortje, Dan Oneata, Herman Kamper

We propose a visually grounded speech model that learns new words and their visual depictions from just a few word-image example pairs. Given a set of test images and a spoken query, we ask the model which image depicts …

Few-Shot Learning

Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling

2024-09-03 · Leanne Nortje

This dissertation examines visually grounded speech (VGS) models that learn from unlabelled speech paired with images. It focuses on applications for low-resource languages and understanding human language acquisition. W…

Few-Shot LearningLanguage Acquisition

World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models

2023-06-14 · Ziqiao Ma, Jiayi Pan, Joyce Chai

The ability to connect language units to their referents in the physical world, referred to as grounding, is crucial to learning and understanding grounded meanings of words. While humans demonstrate fast mapping in new …

Grounded Open Vocabulary AcquisitionLanguage ModelingLanguage Modelling

VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot Learning

2022-03-20 · CVPR 2022 1 · Wenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele 외

Human-annotated attributes serve as powerful semantic embeddings in zero-shot learning. However, their annotation process is labor-intensive and needs expert supervision. Current unsupervised semantic embeddings, i.e., w…

Transfer LearningWord EmbeddingsZero-Shot Learning

Learning the meanings of function words from grounded language using a visual question answering model

2023-08-16 · Eva Portelance, Michael C. Frank, Dan Jurafsky

Interpreting a seemingly-simple function word like "or", "behind", or "more" can require logical, numerical, and relational reasoning. How are such words learned by children? Prior acquisition theories have often relied …

Logical ReasoningQuestion AnsweringRelational ReasoningVisual Question Answering