paper-with-me

Papers

Pre-training image-language transformers for open-vocabulary tasks

2022-09-09 · AJ Piergiovanni, Weicheng Kuo, Anelia Angelova

We present a pre-training approach for vision and language transformer models, which is based on a mixture of diverse tasks. We explore both the use of image-text captioning data in pre-training, which does not need additional supervision, as well as object-aware strategies to pre-train the model. We evaluate the method on a number of textgenerative vision+language tasks, such as Visual Question Answering, visual entailment and captioning, and demonstrate large gains over standard pre-training methods.

📄 PDF Abstract BibTeX arXiv:2209.04372

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual EntailmentVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Region-Aware Pretraining for Open-Vocabulary Object Detection with Vision Transformers

2023-05-11 · CVPR 2023 1 · Dahun Kim, Anelia Angelova, Weicheng Kuo

We present Region-aware Open-vocabulary Vision Transformers (RO-ViT) - a contrastive image-text pretraining recipe to bridge the gap between image-level pretraining and open-vocabulary object detection. At the pretrainin…

Contrastive LearningImage-text Retrievalobject-detectionObject Detection+5

CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

2023-10-02 · Size Wu, Wenwei Zhang, Lumin Xu, Sheng Jin 외

Open-vocabulary dense prediction tasks including object detection and image segmentation have been advanced by the success of Contrastive Language-Image Pre-training (CLIP). CLIP models, particularly those incorporating …

image-classificationImage ClassificationImage Segmentationobject-detection+10

Simple Open-Vocabulary Object Detection with Vision Transformers

2022-05-12 · Matthias Minderer, Alexey Gritsenko, Austin Stone, Maxim Neumann 외

Combining simple architectures with large-scale pre-training has led to massive improvements in image classification. For object detection, pre-training and scaling approaches are less well established, especially in the…

Described Object Detectionimage-classificationImage ClassificationObject+4

Task Addition and Weight Disentanglement in Closed-Vocabulary Models

2025-11-18 · Adam Hazimeh, Alessandro Favero, Pascal Frossard arxiv

Task arithmetic has recently emerged as a promising method for editing pre-trained \textit{open-vocabulary} models, offering a cost-effective alternative to standard multi-task fine-tuning. However, despite the abundance…

Image Classification

Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers

2025-05-07 · Divyansh Srivastava, Xiang Zhang, He Wen, Chenru Wen 외

We present Lay-Your-Scene (shorthand LayouSyn), a novel text-to-layout generation pipeline for natural scenes. Prior scene layout generation methods are either closed-vocabulary or use proprietary large language models f…

Image GenerationLayout Generation