paper-with-me

Papers

Large-scale representation learning from visually grounded untranscribed speech

2019-09-19 · CONLL 2019 11 · Gabriel Ilharco, Yuan Zhang, Jason Baldridge

Systems that can associate images with their spoken audio captions are an important step towards visually grounded language learning. We describe a scalable method to automatically generate diverse audio for image captioning datasets. This supports pretraining deep networks for encoding both audio and images, which we do via a dual encoder that learns to align latent representations from both modalities. We show that a masked margin softmax loss for such models is superior to the standard triplet loss. We fine-tune these models on the Flickr8k Audio Captions Corpus and obtain state-of-the-art results---improving recall in the top 10 from 29.6% to 49.5%. We also obtain human ratings on retrieval outputs to better assess the impact of incidentally matching image-caption pairs that were not associated in the data, finding that automatic evaluation substantially underestimates the quality of the retrieved results.

📄 PDF Abstract BibTeX arXiv:1909.08782

Code (0)

등록된 구현이 없습니다.

Tasks

Grounded language learningImage CaptioningRepresentation LearningRetrievalTriplet

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Semantic speech retrieval with a visually grounded model of untranscribed speech

2017-10-05 · Herman Kamper, Gregory Shakhnarovich, Karen Livescu

There is growing interest in models that can learn from unlabelled speech paired with visual context. This setting is relevant for low-resource speech processing, robotics, and human language acquisition research. Here w…

Language AcquisitionRetrieval

Visually grounded learning of keyword prediction from untranscribed speech

2017-03-23 · Herman Kamper, Shane Settle, Gregory Shakhnarovich, Karen Livescu

During language acquisition, infants have the benefit of visual cues to ground spoken language. Robots similarly have access to audio and visual sensors. Recent work has shown that images and spoken captions can be mappe…

Language AcquisitionTAG

Semantic query-by-example speech search using visual grounding

2019-04-15 · Herman Kamper, Aristotelis Anastassiou, Karen Livescu

A number of recent studies have started to investigate how speech systems can be trained on untranscribed speech by leveraging accompanying images at training time. Examples of tasks include keyword prediction and within…

RetrievalSemantic RetrievalVisual Grounding

Guided-TTS 2: A Diffusion Model for High-quality Adaptive Text-to-Speech with Untranscribed Data

2022-05-30 · Sungwon Kim, Heeseung Kim, Sungroh Yoon

We propose Guided-TTS 2, a diffusion-based generative model for high-quality adaptive TTS using untranscribed data. Guided-TTS 2 combines a speaker-conditional diffusion model with a speaker-dependent phoneme classifier …

text-to-speechText to Speech

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP

2026-06-25 · Sicheng Zhang, Muzammal Naseer, Binzhu Xie, Naufal Suryanto 외 arxiv

CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applications increasingly demand visually ground…

Continual Pretraining