paper-with-me

홈 › Papers

What Happens During Finetuning of Vision Transformers: An Invariance Based Investigation

2023-07-12 · Gabriele Merlin, Vedant Nanda, Ruchit Rawal, Mariya Toneva

The pretrain-finetune paradigm usually improves downstream performance over training a model from scratch on the same task, becoming commonplace across many areas of machine learning. While pretraining is empirically observed to be beneficial for a range of tasks, there is not a clear understanding yet of the reasons for this effect. In this work, we examine the relationship between pretrained vision transformers and the corresponding finetuned versions on several benchmark datasets and tasks. We present new metrics that specifically investigate the degree to which invariances learned by a pretrained model are retained or forgotten during finetuning. Using these metrics, we present a suite of empirical findings, including that pretraining induces transferable invariances in shallow layers and that invariances from deeper pretrained layers are compressed towards shallower layers during finetuning. Together, these findings contribute to understanding some of the reasons for the successes of pretrained models and the changes that a pretrained model undergoes when finetuned on a downstream task.

📄 PDF Abstract BibTeX arXiv:2307.06006

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SparseMAE: Sparse Training Meets Masked Autoencoders

2023-01-01 · ICCV 2023 1 · Aojun Zhou, Yang Li, Zipeng Qin, Jianbo Liu 외

Masked Autoencoders (MAE) and its variants have proven to be effective for pretraining large-scale Vision Transformers (ViTs). However, small-scale models do not benefit from the pretraining mechanisms due to limited…

Continual SFT Matches Multimodal RLHF with Negative Supervision

2024-11-22 · CVPR 2025 1 · Ke Zhu, Yu Wang, Yanpeng Sun, Qiang Chen 외

Multimodal RLHF usually happens after supervised finetuning (SFT) stage to continually improve vision-language models' (VLMs) comprehension. Conventional wisdom holds its superiority over continual SFT during this prefer…

What Happens During the Loss Plateau? Understanding Abrupt Learning in Transformers

2025-06-16 · Pulkit Gopalani, Wei Hu

Training Transformers on algorithmic tasks frequently demonstrates an intriguing abrupt learning phenomenon: an extended performance plateau followed by a sudden, sharp improvement. This work investigates the underlying …

Vision Transformer Finetuning Benefits from Non-Smooth Components

2026-02-06 · Ambroise Odonnat, Laetitia Chapel, Romain Tavenard, Ievgen Redko arxiv

The smoothness of the transformer architecture has been extensively studied in the context of generalization, training stability, and adversarial robustness. However, its role in transfer learning remains poorly understo…

Adversarial RobustnessTransfer Learning

Which Pretrain Samples to Rehearse when Finetuning Pretrained Models?

2024-02-12 · Andrew Bai, Chih-Kuan Yeh, Cho-Jui Hsieh, Ankur Taly

Fine-tuning pretrained foundational models on specific tasks is now the de facto approach for text and vision tasks. A known pitfall of this approach is the forgetting of pretraining knowledge that happens during finetun…