12-in-1: Multi-Task Vision and Language Representation Learning
Much of vision-and-language research focuses on a small but diverse set of independent tasks and supporting datasets often studied in isolation; however, the visually-grounded language understanding skills required for success at these tasks overlap significantly. In this work, we investigate these relationships between vision-and-language tasks by developing a large-scale, multi-task training regime. Our approach culminates in a single model on 12 datasets from four broad categories of task including visual question answering, caption-based image retrieval, grounding referring expressions, and multi-modal verification. Compared to independently trained single-task models, this represents a reduction from approximately 3 billion parameters to 270 million while simultaneously improving performance by 2.05 points on average across tasks. We use our multi-task framework to perform in-depth analysis of the effect of joint training diverse tasks. Further, we show that finetuning task-specific models from our single multi-task model can lead to further improvements, achieving performance at or above the state-of-the-art.
Code (5)
Tasks
10-shot image generationImage RetrievalQuestion AnsweringRepresentation LearningRetrievalVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
Is Multimodal Vision Supervision Beneficial to Language?
Vision (image and video) - Language (VL) pre-training is the recent popular paradigm that achieved state-of-the-art results on multi-modal tasks like image-retrieval, video-retrieval, visual question answering etc. These…
Image RetrievalNatural Language UnderstandingQuestion AnsweringRetrieval+3DiMBERT: Learning Vision-Language Grounded Representations with Disentangled Multimodal-Attention
Vision-and-language (V-L) tasks require the system to understand both vision content and natural language, thus learning fine-grained joint representations of vision and language (a.k.a. V-L representations) is of paramo…
Image CaptioningLanguage ModelingLanguage ModellingSentence+2ViL-Sum: Enhancing Vision and Language Representations via Multi-task Learning for Multi-modal Summarization
With the advance of multimedia on the Internet, multi-modal summarization has drawn much attention. Most current methods follow a pipeline strategy, where an off-the-shelf object detector is used to extract visual featur…
DecoderMulti-Task LearningMAMO: Masked Multimodal Modeling for Fine-Grained Vision-Language Representation Learning
Multimodal representation learning has shown promising improvements on various vision-language tasks. Most existing methods excel at building global-level alignment between vision and language while lacking effective fin…
Image-text Retrievalmultimodal interactionQuestion AnsweringRepresentation Learning+6Learning to Scale Multilingual Representations for Vision-Language Tasks
Current multilingual vision-language models either require a large number of additional parameters for each supported language, or suffer performance degradation as languages are added. In this paper, we propose a Scalab…
Language ModelingLanguage ModellingMachine TranslationRetrieval+3