paper-with-me

홈 › Papers

Revisiting Mixout: An Overlooked Path to Robust Finetuning

2025-10-08 · Masih Aminbeidokhti, Heitor Rapela Medeiros, Eric Granger, Marco Pedersoli arxiv

Finetuning vision foundation models often improves in-domain accuracy but comes at the cost of robustness under distribution shift. We revisit Mixout, a stochastic regularizer that intermittently replaces finetuned weights with their pretrained reference, through the lens of a single-run, weight-sharing implicit ensemble. This perspective reveals three key levers that govern robustness: the \emph{masking anchor}, \emph{resampling frequency}, and \emph{mask sparsity}. Guided by this analysis, we introduce GMixout, which (i) replaces the fixed anchor with an exponential moving-average snapshot that adapts during training, and (ii) regulates masking period via an explicit resampling-frequency hyperparameter. Our sparse-kernel implementation updates only a small fraction of parameters with no inference-time overhead, enabling training on consumer-grade GPUs. Experiments on benchmarks covering covariate shift, corruption, and class imbalance, ImageNet / ImageNet-LT, DomainNet, iWildCam, and CIFAR100-C, GMixout consistently improves in-domain accuracy beyond zero-shot performance while surpassing both Model Soups and strong parameter-efficient finetuning baselines under distribution shift.

📄 PDF Abstract BibTeX arXiv:2510.06982

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mixout: Effective Regularization to Finetune Large-scale Pretrained Language Models

2019-09-25 · ICLR 2020 1 · Cheolhyoung Lee, Kyunghyun Cho, Wanmo Kang

In natural language processing, it has been observed recently that generalization could be greatly improved by finetuning a large-scale language model pretrained on a large unlabeled corpus. Despite its recent success an…

Language ModelingLanguage Modelling

High-Rate Mixout: Revisiting Mixout for Robust Domain Generalization

2025-10-08 · Masih Aminbeidokhti, Heitor Rapela Medeiros, Srikanth Muralidharan, Eric Granger 외 arxiv

Ensembling fine-tuned models initialized from powerful pre-trained weights is a common strategy to improve robustness under distribution shifts, but it comes with substantial computational costs due to the need to train …

Domain Generalization

The Group Robustness is in the Details: Revisiting Finetuning under Spurious Correlations

2024-07-19 · Tyler LaBonte, John C. Hill, Xinchen Zhang, Vidya Muthukumar 외

Modern machine learning models are prone to over-reliance on spurious correlations, which can often lead to poor performance on minority groups. In this paper, we identify surprising and nuanced behavior of finetuned mod…

Revisiting Silhouette Aggregation

2024-01-11 · John Pavlopoulos, Georgios Vardakas, Aristidis Likas

Silhouette coefficient is an established internal clustering evaluation measure that produces a score per data point, assessing the quality of its clustering assignment. To assess the quality of the clustering of the who…

Clustering

Pre-Finetuning with Impact Duration Awareness for Stock Movement Prediction

2024-09-25 · Chr-Jr Chiu, Chung-Chi Chen, Hen-Hsen Huang, Hsin-Hsi Chen

Understanding the duration of news events' impact on the stock market is crucial for effective time-series forecasting, yet this facet is largely overlooked in current research. This paper addresses this research gap by …

Sentiment AnalysisTime SeriesTime Series Forecasting