paper-with-me

홈 › Papers

Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

2021-02-17 · CVPR 2021 1 · Soravit Changpinyo, Piyush Sharma, Nan Ding, Radu Soricut

The availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pre-training. However, these datasets are often collected with overrestrictive requirements inherited from their original target tasks (e.g., image caption generation), which limit the resulting dataset scale and diversity. We take a step further in pushing the limits of vision-and-language pre-training data by relaxing the data collection pipeline used in Conceptual Captions 3M (CC3M) [Sharma et al. 2018] and introduce the Conceptual 12M (CC12M), a dataset with 12 million image-text pairs specifically meant to be used for vision-and-language pre-training. We perform an analysis of this dataset and benchmark its effectiveness against CC3M on multiple downstream tasks with an emphasis on long-tail visual recognition. Our results clearly illustrate the benefit of scaling up pre-training data for vision-and-language tasks, as indicated by the new state-of-the-art results on both the nocaps and Conceptual Captions benchmarks.

📄 PDF Abstract BibTeX arXiv:2102.08981

Code (3)

google-research-datasets/conceptual-12m 공식 구현
facebookresearch/meru pytorch
gicheonkang/gst-visdial pytorch

Tasks

Caption GenerationDiversityImage CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Ambiguity Helps: Classification With Disagreements in Crowdsourced Annotations

2016-06-01 · CVPR 2016 6 · Viktoriia Sharmanska, Daniel Hernandez-Lobato, Jose Miguel Hernandez-Lobato, Novi Quadrianto

Imagine we show an image to a person and ask her/him to decide whether the scene in the image is warm or not warm, and whether it is easy or not to spot a squirrel in the image. For exactly the same image, the answers to…

ClassificationGeneral Classification

PnPNet: Pull-and-Push Networks for Volumetric Segmentation with Boundary Confusion

2023-12-13 · Xin You, Ming Ding, Minghui Zhang, Hanxiao Zhang 외

Precise boundary segmentation of volumetric images is a critical task for image-guided diagnosis and computer-assisted intervention, especially for boundary confusion in clinical practice. However, U-shape networks canno…

Visual Conceptual Blending with Large-scale Language and Vision Models

2021-06-27 · Songwei Ge, Devi Parikh

We ask the question: to what extent can recent large-scale language and image generation models blend visual concepts? Given an arbitrary object, we identify a relevant object and generate a single-sentence description o…

Image GenerationLanguage ModelingLanguage ModellingObject+1

ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

2020-01-22 · Di Qi, Lin Su, Jia Song, Edward Cui 외

In this paper, we introduce a new vision-language pre-trained model -- ImageBERT -- for image-text joint embedding. Our model is a Transformer-based model, which takes different modalities as input and models the relatio…

Image RetrievalImage-text matchingLanguage ModelingLanguage Modelling+5

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

2018-10-11 · NAACL 2019 6 · Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova

We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers. Unlike recent language representation models, BERT is designed to pre-train deep bid…

Citation Intent ClassificationCommon Sense ReasoningConversational Response SelectionCoreference Resolution+17