paper-with-me

홈 › Papers

Scaling Up Vision-Language Pre-training for Image Captioning

2021-11-24 · CVPR 2022 1 · Xiaowei Hu, Zhe Gan, JianFeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, Lijuan Wang

In recent years, we have witnessed significant performance boost in the image captioning task based on vision-language pre-training (VLP). Scale is believed to be an important factor for this advance. However, most existing work only focuses on pre-training transformers with moderate sizes (e.g., 12 or 24 layers) on roughly 4 million images. In this paper, we present LEMON, a LargE-scale iMage captiONer, and provide the first empirical study on the scaling behavior of VLP for image captioning. We use the state-of-the-art VinVL model as our reference model, which consists of an image feature extractor and a transformer model, and scale the transformer both up and down, with model sizes ranging from 13 to 675 million parameters. In terms of data, we conduct experiments with up to 200 million image-text pairs which are automatically collected from web based on the alt attribute of the image (dubbed as ALT200M). Extensive analysis helps to characterize the performance trend as the model size and the pre-training data size increase. We also compare different training recipes, especially for training on large-scale noisy data. As a result, LEMON achieves new state of the arts on several major image captioning benchmarks, including COCO Caption, nocaps, and Conceptual Captions. We also show LEMON can generate captions with long-tail visual concepts when used in a zero-shot manner.

📄 PDF Abstract BibTeX arXiv:2111.12233

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeImage Captioning

Similar Papers 제목 키워드 기반

PaLI-X: On Scaling up a Multilingual Vision and Language Model

2023-05-29 · Xi Chen, Josip Djolonga, Piotr Padlewski, Basil Mustafa 외

We present the training recipe and results of scaling up PaLI-X, a multilingual vision and language model, both in terms of size of the components and the breadth of its training task mixture. Our model achieves new leve…

Chart Question Answeringdocument understandingFine-Grained Image RecognitionIn-Context Learning+10

On Scaling Up a Multilingual Vision and Language Model

2024-01-01 · CVPR 2024 1 · Xi Chen, Josip Djolonga, Piotr Padlewski, Basil Mustafa 외

We explore the boundaries of scaling up a multilingual vision and language model both in terms of size of the components and the breadth of its training task mixture. Our model achieves new levels of performance on a…

document understandingIn-Context LearningLanguage ModelingLanguage Modelling+6

Image Captioners Are Scalable Vision Learners Too

2023-06-13 · NeurIPS 2023 11 · Michael Tschannen, Manoj Kumar, Andreas Steiner, Xiaohua Zhai 외

Contrastive pretraining on image-text pairs from the web is one of the most popular large-scale pretraining strategies for vision backbones, especially in the context of large multimodal models. At the same time, image c…

DecoderImage Captioning

I-Tuning: Tuning Frozen Language Models with Image for Lightweight Image Captioning

2022-02-14 · Ziyang Luo, Zhipeng Hu, Yadong Xi, Rongsheng Zhang 외

Image Captioning is a traditional vision-and-language task that aims to generate the language description of an image. Recent studies focus on scaling up the model size and the number of training data, which significantl…

DecoderImage CaptioningLanguage Modelling

Florenz: Scaling Laws for Systematic Generalization in Vision-Language Models

2025-03-12 · Julian Spravil, Sebastian Houben, Sven Behnke

Cross-lingual transfer enables vision-language models (VLMs) to perform vision tasks in various languages with training data only in one language. Current approaches rely on large pre-trained multilingual language models…

Cross-Lingual TransferImage CaptioningLarge Language ModelMachine Translation+3