paper-with-me

Papers

UNITER: Learning UNiversal Image-TExt Representations

2019-09-25 · Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, Jingjing Liu

Joint image-text embedding is the bedrock for most Vision-and-Language (V+L) tasks, where multimodality inputs are jointly processed for visual and textual understanding. In this paper, we introduce UNITER, a UNiversal Image-TExt Representation, learned through large-scale pre-training over four image-text datasets (COCO, Visual Genome, Conceptual Captions, and SBU Captions), which can power heterogeneous downstream V+L tasks with joint multimodal embeddings. We design three pre-training tasks: Masked Language Modeling (MLM), Image-Text Matching (ITM), and Masked Region Modeling (MRM, with three variants). Different from concurrent work on multimodal pre-training that apply joint random masking to both modalities, we use Conditioned Masking on pre-training tasks (i.e., masked language/region modeling is conditioned on full observation of image/text). Comprehensive analysis shows that conditioned masking yields better performance than unconditioned masking. We also conduct a thorough ablation study to find an optimal combination of pre-training tasks for UNITER. Extensive experiments show that UNITER achieves new state of the art across six V+L tasks over nine datasets, including Visual Question Answering, Image-Text Retrieval, Referring Expression Comprehension, Visual Commonsense Reasoning, Visual Entailment, and NLVR2.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image-text matchingImage-text RetrievalLanguage ModelingLanguage ModellingMasked Language ModelingQuestion AnsweringReferring ExpressionReferring Expression ComprehensionRetrievalText MatchingText RetrievalVisual Commonsense ReasoningVisual EntailmentVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

UNITER UNITER or UNiversal Image-TExt Representation model is a large-scale pre-trained model for joint multimodal embedding. It is pre-trained using four image-text datasets COCO,…

Similar Papers 제목 키워드 기반

UNITER: UNiversal Image-TExt Representation Learning

2019-09-25 · ECCV 2020 8 · Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy 외

Joint image-text embedding is the bedrock for most Vision-and-Language (V+L) tasks, where multimodality inputs are simultaneously processed for joint visual and textual understanding. In this paper, we introduce UNITER, …

Image-text matchingImage-text RetrievalLanguage ModelingLanguage Modelling+14

X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers

2020-09-23 · EMNLP 2020 11 · Jaemin Cho, Jiasen Lu, Dustin Schwenk, Hannaneh Hajishirzi 외

Mirroring the success of masked language models, vision-and-language counterparts like ViLBERT, LXMERT and UNITER have achieved state of the art performance on a variety of multimodal discriminative tasks like visual que…

Image CaptioningImage GenerationQuestion AnsweringVisual Grounding+2

Playing Lottery Tickets with Vision and Language

2021-04-23 · Zhe Gan, Yen-Chun Chen, Linjie Li, Tianlong Chen 외

Large-scale pre-training has recently revolutionized vision-and-language (VL) research. Models such as LXMERT and UNITER have significantly lifted the state of the art over a wide range of VL tasks. However, the large nu…

Image-text RetrievalQuestion AnsweringReferring ExpressionReferring Expression Comprehension+6

Target-dependent UNITER: A Transformer-Based Multimodal Language Comprehension Model for Domestic Service Robots

2021-07-02 · Shintaro Ishikawa, Komei Sugiura

Currently, domestic service robots have an insufficient ability to interact naturally through language. This is because understanding human instructions is complicated by various ambiguities and missing information. In e…

UNITER-Based Situated Coreference Resolution with Rich Multimodal Input

2021-12-07 · Yichen Huang, Yuchen Wang, Yik-Cheung Tam

We present our work on the multimodal coreference resolution task of the Situated and Interactive Multimodal Conversation 2.0 (SIMMC 2.0) dataset as a part of the tenth Dialog System Technology Challenge (DSTC10). We pro…

coreference-resolutionCoreference ResolutionObjectVisual Dialog