Clustering-Based Validation Splits for Model Selection under Domain Shift
This paper considers the problem of model selection under domain shift. Motivated by principles from distributionally robust optimisation (DRO) and domain adaptation theory, it is proposed that the training-validation split should maximise the distribution mismatch between the two sets. By adopting the maximum mean discrepancy (MMD) as the measure of mismatch, it is shown that the partitioning problem reduces to kernel k-means clustering. A constrained clustering algorithm, which leverages linear programming to control the size, label, and (optionally) group distributions of the splits, is presented. The algorithm does not require additional metadata, and comes with convergence guarantees. In experiments, the technique consistently outperforms alternative splitting strategies across a range of datasets and training algorithms, for both domain generalisation (DG) and unsupervised domain adaptation (UDA) tasks. Analysis also shows the MMD between the training and validation sets to be strongly rank-correlated ($\rho=0.63$) with test domain accuracy, further substantiating the validity of this approach.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringConstrained ClusteringDomain AdaptationModel SelectionUnsupervised Domain AdaptationSimilar Papers 제목 키워드 기반
Variational Resampling Based Assessment of Deep Neural Networks under Distribution Shift
A novel variational inference based resampling framework is proposed to evaluate the robustness and generalization capability of deep learning models with respect to distribution shift. We use Auto Encoding Variational B…
Domain AdaptationDomain GeneralizationGeneral Classificationimage-classification+3A new type of federated clustering: A non-model-sharing approach
In recent years, the growing need to leverage sensitive data across institutions has led to increased attention on federated learning (FL), a decentralized machine learning paradigm that enables model training without sh…
ClusteringFederated LearningA review of systematic selection of clustering algorithms and their evaluation
Data analysis plays an indispensable role for value creation in industry. Cluster analysis in this context is able to explore given datasets with little or no prior knowledge and to identify unknown patterns. As (big) da…
ClusteringThe Three Ensemble Clustering (3EC) Algorithm for Pattern Discovery in Unsupervised Learning
This paper presents a multiple learner algorithm called the 'Three Ensemble Clustering 3EC' algorithm that classifies unlabeled data into quality clusters as a part of unsupervised learning. It offers the flexibility to …
ClusteringBeyond the Performance Illusion: Structure-Aware Stratified Partitioning and Curriculum Distributionally Robust Optimization for Spatially Correlated Domains
Performance evaluation in AI systems commonly assumes that random dataset splits produce independent and identically distributed (i.i.d.) subsets. We show that this assumption often breaks down in spatiotemporally correl…