paper-with-me

홈 › Papers

Unsupervised vs. transfer learning for multimodal one-shot matching of speech and images

2020-08-14 · Leanne Nortje, Herman Kamper

We consider the task of multimodal one-shot speech-image matching. An agent is shown a picture along with a spoken word describing the object in the picture, e.g. cookie, broccoli and ice-cream. After observing one paired speech-image example per class, it is shown a new set of unseen pictures, and asked to pick the "ice-cream". Previous work attempted to tackle this problem using transfer learning: supervised models are trained on labelled background data not containing any of the one-shot classes. Here we compare transfer learning to unsupervised models trained on unlabelled in-domain data. On a dataset of paired isolated spoken and visual digits, we specifically compare unsupervised autoencoder-like models to supervised classifier and Siamese neural networks. In both unimodal and multimodal few-shot matching experiments, we find that transfer learning outperforms unsupervised training. We also present experiments towards combining the two methodologies, but find that transfer learning still performs best (despite idealised experiments showing the benefits of unsupervised learning).

📄 PDF Abstract BibTeX arXiv:2008.06258

Code (1)

LeanneNortje/multimodal_speech-image_matching 공식 구현 tf

Tasks

Transfer Learning

Similar Papers 제목 키워드 기반

Direct multimodal few-shot learning of speech and images

2020-12-10 · Leanne Nortje, Herman Kamper

We propose direct multimodal few-shot models that learn a shared embedding space of spoken words and images from only a few paired examples. Imagine an agent is shown an image along with a spoken word describing the obje…

Few-Shot LearningTransfer LearningTriplet

Transferring Pre-trained Multimodal Representations with Cross-modal Similarity Matching

2023-01-07 · Byoungjip Kim, Sungik Choi, Dasol Hwang, Moontae Lee 외

Despite surprising performance on zero-shot transfer, pre-training a large-scale multimodal model is often prohibitive as it requires a huge amount of data and computing resources. In this paper, we propose a method (Bea…

Language ModelingLanguage ModellingSelf-Supervised Learning

Visually grounded few-shot word learning in low-resource settings

2023-06-20 · Leanne Nortje, Dan Oneata, Herman Kamper

We propose a visually grounded speech model that learns new words and their visual depictions from just a few word-image example pairs. Given a set of test images and a spoken query, we ask the model which image depicts …

Few-Shot Learning

An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS

2024-06-09 · Xiaofei Wang, Sefik Emre Eskimez, Manthan Thakker, Hemin Yang 외

Recently, zero-shot text-to-speech (TTS) systems, capable of synthesizing any speaker's voice from a short audio prompt, have made rapid advancements. However, the quality of the generated speech significantly deteriorat…

DenoisingSpeech DenoisingSpeech Enhancementtext-to-speech+2

SLMGAN: Exploiting Speech Language Model Representations for Unsupervised Zero-Shot Voice Conversion in GANs

2023-07-18 · Yinghao Aaron Li, Cong Han, Nima Mesgarani

In recent years, large-scale pre-trained speech language models (SLMs) have demonstrated remarkable advancements in various generative speech modeling applications, such as text-to-speech synthesis, voice conversion, and…

Generative Adversarial NetworkLanguage ModelingLanguage ModellingSpeech Enhancement+5