paper-with-me

홈 › Papers

12-in-1: Multi-Task Vision and Language Representation Learning

2019-12-05 · CVPR 2020 6 · Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, Stefan Lee

Much of vision-and-language research focuses on a small but diverse set of independent tasks and supporting datasets often studied in isolation; however, the visually-grounded language understanding skills required for success at these tasks overlap significantly. In this work, we investigate these relationships between vision-and-language tasks by developing a large-scale, multi-task training regime. Our approach culminates in a single model on 12 datasets from four broad categories of task including visual question answering, caption-based image retrieval, grounding referring expressions, and multi-modal verification. Compared to independently trained single-task models, this represents a reduction from approximately 3 billion parameters to 270 million while simultaneously improving performance by 2.05 points on average across tasks. We use our multi-task framework to perform in-depth analysis of the effect of joint training diverse tasks. Further, we show that finetuning task-specific models from our single multi-task model can lead to further improvements, achieving performance at or above the state-of-the-art.

📄 PDF Abstract BibTeX arXiv:1912.02315

Code (5)

facebookresearch/vilbert-multi-task 공식 구현 pytorch
Cloud-CV/vilbert-multi-task
jialinwu17/tmpimgs pytorch
jiasenlu/vilbert_beta pytorch
johntiger1/multitask_multimodal pytorch

Tasks

10-shot image generationImage RetrievalQuestion AnsweringRepresentation LearningRetrievalVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Is Multimodal Vision Supervision Beneficial to Language?

2023-02-10 · Avinash Madasu, Vasudev Lal

Vision (image and video) - Language (VL) pre-training is the recent popular paradigm that achieved state-of-the-art results on multi-modal tasks like image-retrieval, video-retrieval, visual question answering etc. These…

Image RetrievalNatural Language UnderstandingQuestion AnsweringRetrieval+3

DiMBERT: Learning Vision-Language Grounded Representations with Disentangled Multimodal-Attention

2022-10-28 · Fenglin Liu, Xian Wu, Shen Ge, Xuancheng Ren 외

Vision-and-language (V-L) tasks require the system to understand both vision content and natural language, thus learning fine-grained joint representations of vision and language (a.k.a. V-L representations) is of paramo…

Image CaptioningLanguage ModelingLanguage ModellingSentence+2

ViL-Sum: Enhancing Vision and Language Representations via Multi-task Learning for Multi-modal Summarization

2022-01-16 · ACL ARR January 2022 1 · Anonymous

With the advance of multimedia on the Internet, multi-modal summarization has drawn much attention. Most current methods follow a pipeline strategy, where an off-the-shelf object detector is used to extract visual featur…

DecoderMulti-Task Learning

MAMO: Masked Multimodal Modeling for Fine-Grained Vision-Language Representation Learning

2022-10-09 · Zijia Zhao, Longteng Guo, Xingjian He, Shuai Shao 외

Multimodal representation learning has shown promising improvements on various vision-language tasks. Most existing methods excel at building global-level alignment between vision and language while lacking effective fin…

Image-text Retrievalmultimodal interactionQuestion AnsweringRepresentation Learning+6

Learning to Scale Multilingual Representations for Vision-Language Tasks

2020-04-09 · ECCV 2020 8 · Andrea Burns, Donghyun Kim, Derry Wijaya, Kate Saenko 외

Current multilingual vision-language models either require a large number of additional parameters for each supported language, or suffer performance degradation as languages are added. In this paper, we propose a Scalab…

Language ModelingLanguage ModellingMachine TranslationRetrieval+3