paper-with-me

홈 › Papers

A Closer Look at the Robustness of Vision-and-Language Pre-trained Models

2020-12-15 · Linjie Li, Zhe Gan, Jingjing Liu

Large-scale pre-trained multimodal transformers, such as ViLBERT and UNITER, have propelled the state of the art in vision-and-language (V+L) research to a new level. Although achieving impressive performance on standard tasks, to date, it still remains unclear how robust these pre-trained models are. To investigate, we conduct a host of thorough evaluations on existing pre-trained models over 4 different types of V+L specific model robustness: (i) Linguistic Variation; (ii) Logical Reasoning; (iii) Visual Content Manipulation; and (iv) Answer Distribution Shift. Interestingly, by standard model finetuning, pre-trained V+L models already exhibit better robustness than many task-specific state-of-the-art methods. To further enhance model robustness, we propose Mango, a generic and efficient approach that learns a Multimodal Adversarial Noise GeneratOr in the embedding space to fool pre-trained V+L models. Differing from previous studies focused on one specific type of robustness, Mango is task-agnostic, and enables universal performance lift for pre-trained models over diverse tasks designed to evaluate broad aspects of robustness. Comprehensive experiments demonstrate that Mango achieves new state of the art on 7 out of 9 robustness benchmarks, surpassing existing methods by a significant margin. As the first comprehensive study on V+L robustness, this work puts robustness of pre-trained models into sharper focus, pointing new directions for future study.

📄 PDF Abstract BibTeX arXiv:2012.08673

Code (0)

등록된 구현이 없습니다.

Tasks

Logical Reasoning

Methods 이 논문이 사용한 방법론

UNITER UNITER or UNiversal Image-TExt Representation model is a large-scale pre-trained model for joint multimodal embedding. It is pre-trained using four image-text datasets COCO,…
ViLBERT Vision-and-Language BERT (ViLBERT) is a BERT-based model for learning task-agnostic joint representations of image content and…

Similar Papers 제목 키워드 기반

A Closer Look at the Adversarial Robustness of Information Bottleneck Models

2021-07-12 · ICML Workshop AML 2021 7 · Iryna Korshunova, David Stutz, Alexander A. Alemi, Olivia Wiles 외

We study the adversarial robustness of information bottleneck models for classification. Previous works showed that the robustness of models trained with information bottlenecks can improve upon adversarial training. Our…

Adversarial Robustness

A Closer Look into the Robustness of Neural Dependency Parsers Using Better Adversarial Examples

2021-08-01 · Findings (ACL) 2021 8 · Yuxuan Wang, Wanxiang Che, Ivan Titov, Shay B. Cohen 외

Robustness and Regularization in Hierarchical Re-Basin

2025-10-10 · Benedikt Franke, Florian Heinrich, Markus Lange, Arne Raulf arxiv

This paper takes a closer look at Git Re-Basin, an interesting new approach to merge trained models. We propose a hierarchical model merging scheme that significantly outperforms the standard MergeMany algorithm. With ou…

A Closer Look at Accuracy vs. Robustness

2020-03-05 · NeurIPS 2020 12 · Yao-Yuan Yang, Cyrus Rashtchian, Hongyang Zhang, Ruslan Salakhutdinov 외

Current methods for training robust networks lead to a drop in test accuracy, which has led prior works to posit that a robustness-accuracy tradeoff may be inevitable in deep learning. We take a closer look at this pheno…

A Closer Look at Invariances in Self-supervised Pre-training for 3D Vision

2022-07-11 · Lanxiao Li, Michael Heizmann

Self-supervised pre-training for 3D vision has drawn increasing research interest in recent years. In order to learn informative representations, a lot of previous works exploit invariances of 3D features, e.g., perspect…

Contrastive Learningobject-detectionObject Detection