paper-with-me

홈 › Papers

Learning Visual Representations with Caption Annotations

2020-08-04 · ECCV 2020 8 · Mert Bulent Sariyildiz, Julien Perez, Diane Larlus

Pretraining general-purpose visual features has become a crucial part of tackling many computer vision tasks. While one can learn such features on the extensively-annotated ImageNet dataset, recent approaches have looked at ways to allow for noisy, fewer, or even no annotations to perform such pretraining. Starting from the observation that captioned images are easily crawlable, we argue that this overlooked source of information can be exploited to supervise the training of visual representations. To do so, motivated by the recent progresses in language models, we introduce {\em image-conditioned masked language modeling} (ICMLM) -- a proxy task to learn visual representations over image-caption pairs. ICMLM consists in predicting masked words in captions by relying on visual cues. To tackle this task, we propose hybrid models, with dedicated visual and textual encoders, and we show that the visual representations learned as a by-product of solving this task transfer well to a variety of target tasks. Our experiments confirm that image captions can be leveraged to inject global and localized semantic information into visual representations. Project website: https://europe.naverlabs.com/icmlm.

📄 PDF Abstract BibTeX arXiv:2008.01392

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningLanguage ModelingLanguage ModellingMasked Language Modeling

Similar Papers 제목 키워드 기반

VirTex: Learning Visual Representations from Textual Annotations

2020-06-11 · CVPR 2021 1 · Karan Desai, Justin Johnson

The de-facto approach to many vision tasks is to start from pretrained visual representations, typically learned via supervised training on ImageNet. Recent methods have explored unsupervised pretraining to scale to vast…

General ClassificationImage Captioningimage-classificationImage Classification+4

Exploring Explicit and Implicit Visual Relationships for Image Captioning

2021-05-06 · Zeliang Song, Xiaofei Zhou

Image captioning is one of the most challenging tasks in AI, which aims to automatically generate textual sentences for an image. Recent methods for image captioning follow encoder-decoder framework that transforms the s…

DecoderImage Captioning

Zero-shot Building Attribute Extraction from Large-Scale Vision and Language Models

2023-12-19 · Fei Pan, Sangryul Jeon, Brian Wang, Frank Mckenna 외

Existing building recognition methods, exemplified by BRAILS, utilize supervised learning to extract information from satellite and street-view images for classification and segmentation. However, each task module requir…

AttributeAttribute ExtractionDescriptive

Fine-Grained Video Captioning through Scene Graph Consolidation

2025-02-23 · Sanghyeok Chu, Seonguk Seo, Bohyung Han

Recent advances in visual language models (VLMs) have significantly improved image captioning, but extending these gains to video understanding remains challenging due to the scarcity of fine-grained video captioning dat…

Caption GenerationImage CaptioningVideo CaptioningVideo Understanding

What Is a Good Caption? A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness

2025-02-19 · Zhihang Liu, Chen-Wei Xie, Bin Wen, Feiwu Yu 외

Visual captioning benchmarks have become outdated with the emergence of modern multimodal large language models (MLLMs), as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effect…

Image CaptioningKeyword Extraction