paper-with-me

홈 › Papers

Delving Deep into the Generalization of Vision Transformers under Distribution Shifts

2021-06-14 · CVPR 2022 1 · Chongzhi Zhang, Mingyuan Zhang, Shanghang Zhang, Daisheng Jin, Qiang Zhou, Zhongang Cai, Haiyu Zhao, Xianglong Liu, Ziwei Liu

Vision Transformers (ViTs) have achieved impressive performance on various vision tasks, yet their generalization under distribution shifts (DS) is rarely understood. In this work, we comprehensively study the out-of-distribution (OOD) generalization of ViTs. For systematic investigation, we first present a taxonomy of DS. We then perform extensive evaluations of ViT variants under different DS and compare their generalization with Convolutional Neural Network (CNN) models. Important observations are obtained: 1) ViTs learn weaker biases on backgrounds and textures, while they are equipped with stronger inductive biases towards shapes and structures, which is more consistent with human cognitive traits. Therefore, ViTs generalize better than CNNs under DS. With the same or less amount of parameters, ViTs are ahead of corresponding CNNs by more than 5% in top-1 accuracy under most types of DS. 2) As the model scale increases, ViTs strengthen these biases and thus gradually narrow the in-distribution and OOD performance gap. To further improve the generalization of ViTs, we design the Generalization-Enhanced ViTs (GE-ViTs) from the perspectives of adversarial learning, information theory, and self-supervised learning. By comprehensively investigating these GE-ViTs and comparing with their corresponding CNN models, we observe: 1) For the enhanced model, larger ViTs still benefit more for the OOD generalization. 2) GE-ViTs are more sensitive to the hyper-parameters than their corresponding CNN models. We design a smoother learning strategy to achieve a stable training process and obtain performance improvements on OOD data by 4% from vanilla ViTs. We hope our comprehensive study could shed light on the design of more generalizable learning architectures.

📄 PDF Abstract BibTeX arXiv:2106.07617

Code (1)

Phoenix1153/ViT_OOD_generalization 공식 구현 pytorch

Tasks

Out-of-Distribution GeneralizationSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Delving Deep into Semantic Relation Distillation

2025-03-27 · Zhaoyi Yan, KangJun Liu, Qixiang Ye

Knowledge distillation has become a cornerstone technique in deep learning, facilitating the transfer of knowledge from complex models to lightweight counterparts. Traditional distillation approaches focus on transferrin…

Knowledge DistillationModel CompressionRelationSuperpixels

Delving Deeper Into Astromorphic Transformers

2023-12-18 · Md Zesun Ahmed Mia, Malyaban Bal, Abhronil Sengupta

Preliminary attempts at incorporating the critical role of astrocytes - cells that constitute more than 50\% of human brain cells - in brain-inspired neuromorphic computing remain in infancy. This paper seeks to delve de…

image-classificationImage ClassificationText Generation

Delving Deeper into Cross-lingual Visual Question Answering

2022-02-15 · Chen Liu, Jonas Pfeiffer, Anna Korhonen, Ivan Vulić 외

Visual question answering (VQA) is one of the crucial vision-and-language tasks. Yet, existing VQA research has mostly focused on the English language, due to a lack of suitable evaluation resources. Previous work on cro…

Inductive BiasQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Vision transformers in domain adaptation and domain generalization: a study of robustness

2024-04-05 · Shadi Alijani, Jamil Fayyad, Homayoun Najjaran

Deep learning models are often evaluated in scenarios where the data distribution is different from those used in the training and validation phases. The discrepancy presents a challenge for accurately predicting the per…

Data AugmentationDomain AdaptationDomain GeneralizationMeta-Learning

Transformer for Object Re-Identification: A Survey

2024-01-13 · Mang Ye, Shuoyi Chen, Chenyue Li, Wei-Shi Zheng 외

Object Re-identification (Re-ID) aims to identify specific objects across different times and scenes, which is a widely researched task in computer vision. For a prolonged period, this field has been predominantly driven…

ObjectSurvey