paper-with-me

홈 › Papers

Direct multimodal few-shot learning of speech and images

2020-12-10 · Leanne Nortje, Herman Kamper

We propose direct multimodal few-shot models that learn a shared embedding space of spoken words and images from only a few paired examples. Imagine an agent is shown an image along with a spoken word describing the object in the picture, e.g. pen, book and eraser. After observing a few paired examples of each class, the model is asked to identify the "book" in a set of unseen pictures. Previous work used a two-step indirect approach relying on learned unimodal representations: speech-speech and image-image comparisons are performed across the support set of given speech-image pairs. We propose two direct models which instead learn a single multimodal space where inputs from different modalities are directly comparable: a multimodal triplet network (MTriplet) and a multimodal correspondence autoencoder (MCAE). To train these direct models, we mine speech-image pairs: the support set is used to pair up unlabelled in-domain speech and images. In a speech-to-image digit matching task, direct models outperform indirect models, with the MTriplet achieving the best multimodal five-shot accuracy. We show that the improvements are due to the combination of unsupervised and transfer learning in the direct models, and the absence of two-step compounding errors.

📄 PDF Abstract BibTeX arXiv:2012.05680

Code (1)

LeanneNortje/direct_multimodal_few-shot_learning 공식 구현 tf

Tasks

Few-Shot LearningTransfer LearningTriplet

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

CrisisHateMM: Multimodal Analysis of Directed and Undirected Hate Speech in Text-Embedded Images from Russia-Ukraine Conflict

2023-06-01 · IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshop 2023 6 · Aashish Bhandari, Siddhant B. Shah, Surendrabikram Thapa, Usman Naseem 외

Text-embedded images are frequently used on social media to convey opinions and emotions, but they can also be a medium for disseminating hate speech, propaganda, and extremist ideologies. During the Russia-Ukraine war, …

Hate Speech Detection CrisisHateMM Benchmark

Multimodal One-Shot Learning of Speech and Images

2018-11-09 · Ryan Eloff, Herman A. Engelbrecht, Herman Kamper

Imagine a robot is shown new concepts visually together with spoken tags, e.g. "milk", "eggs", "butter". After seeing one paired audio-visual example per class, it is shown a new set of unseen instances of these objects,…

Dynamic Time WarpingOne-Shot Learning

T-Modules: Translation Modules for Zero-Shot Cross-Modal Machine Translation

2022-05-24 · Paul-Ambroise Duquenne, Hongyu Gong, Benoît Sagot, Holger Schwenk

We present a new approach to perform zero-shot cross-modal transfer between speech and text for translation tasks. Multilingual speech and text are encoded in a joint fixed-size representation space. Then, we compare dif…

DecoderMachine Translationtext-to-speechText to Speech+2

Visually grounded few-shot word learning in low-resource settings

2023-06-20 · Leanne Nortje, Dan Oneata, Herman Kamper

We propose a visually grounded speech model that learns new words and their visual depictions from just a few word-image example pairs. Given a set of test images and a spoken query, we ask the model which image depicts …

Few-Shot Learning

Unsupervised vs. transfer learning for multimodal one-shot matching of speech and images

2020-08-14 · Leanne Nortje, Herman Kamper

We consider the task of multimodal one-shot speech-image matching. An agent is shown a picture along with a spoken word describing the object in the picture, e.g. cookie, broccoli and ice-cream. After observing one paire…

Transfer Learning