paper-with-me

홈 › Papers

Vision Transformer Neural Architecture Search for Out-of-Distribution Generalization: Benchmark and Insights

2025-01-07 · Sy-Tuyen Ho, Tuan Van Vo, Somayeh Ebrahimkhani, Ngai-Man Cheung

While ViTs have achieved across machine learning tasks, deploying them in real-world scenarios faces a critical challenge: generalizing under OoD shifts. A crucial research gap exists in understanding how to design ViT architectures, both manually and automatically, for better OoD generalization. To this end, we introduce OoD-ViT-NAS, the first systematic benchmark for ViTs NAS focused on OoD generalization. This benchmark includes 3000 ViT architectures of varying computational budgets evaluated on 8 common OoD datasets. Using this benchmark, we analyze factors contributing to OoD generalization. Our findings reveal key insights. First, ViT architecture designs significantly affect OoD generalization. Second, ID accuracy is often a poor indicator of OoD accuracy, highlighting the risk of optimizing ViT architectures solely for ID performance. Third, we perform the first study of NAS for ViTs OoD robustness, analyzing 9 Training-free NAS methods. We find that existing Training-free NAS methods are largely ineffective in predicting OoD accuracy despite excelling at ID accuracy. Simple proxies like Param or Flop surprisingly outperform complex Training-free NAS methods in predicting OoD accuracy. Finally, we study how ViT architectural attributes impact OoD generalization and discover that increasing embedding dimensions generally enhances performance. Our benchmark shows that ViT architectures exhibit a wide range of OoD accuracy, with up to 11.85% improvement for some OoD shifts. This underscores the importance of studying ViT architecture design for OoD. We believe OoD-ViT-NAS can catalyze further research into how ViT designs influence OoD generalization.

📄 PDF Abstract BibTeX arXiv:2501.03782

Code (1)

vovantuan1999a/OoD-ViT-NAS 공식 구현 pytorch

Tasks

Neural Architecture SearchOut-of-Distribution Generalization

Similar Papers 제목 키워드 기반

Vision transformers in domain adaptation and domain generalization: a study of robustness

2024-04-05 · Shadi Alijani, Jamil Fayyad, Homayoun Najjaran

Deep learning models are often evaluated in scenarios where the data distribution is different from those used in the training and validation phases. The discrepancy presents a challenge for accurately predicting the per…

Data AugmentationDomain AdaptationDomain GeneralizationMeta-Learning

Domain Generalisation with Bidirectional Encoder Representations from Vision Transformers

2023-07-16 · Hamza Riaz, Alan F. Smeaton

Domain generalisation involves pooling knowledge from source domain(s) into a single model that can generalise to unseen target domain(s). Recent research in domain generalisation has faced challenges when using deep lea…

Compact Vision Transformer by Reduction of Kernel Complexity

2025-07-17 · Yancheng Wang, Yingzhen Yang arxiv

Self-attention and transformer architectures have become foundational components in modern deep learning. Recent efforts have integrated transformer blocks into compact neural architectures for computer vision, giving ri…

Multi-Dimensional Hyena for Spatial Inductive Bias

2023-09-24 · Itamar Zimerman, Lior Wolf

In recent years, Vision Transformers have attracted increasing interest from computer vision researchers. However, the advantage of these transformers over CNNs is only fully manifested when trained over a large dataset,…

Inductive Bias

Delving Deep into the Generalization of Vision Transformers under Distribution Shifts

2021-06-14 · CVPR 2022 1 · Chongzhi Zhang, Mingyuan Zhang, Shanghang Zhang, Daisheng Jin 외

Vision Transformers (ViTs) have achieved impressive performance on various vision tasks, yet their generalization under distribution shifts (DS) is rarely understood. In this work, we comprehensively study the out-of-dis…

Out-of-Distribution GeneralizationSelf-Supervised Learning