paper-with-me

홈 › Papers

WILDS: A Benchmark of in-the-Wild Distribution Shifts

2020-12-14 · Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton A. Earnshaw, Imran S. Haque, Sara Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, Percy Liang

Distribution shifts -- where the training distribution differs from the test distribution -- can substantially degrade the accuracy of machine learning (ML) systems deployed in the wild. Despite their ubiquity in the real-world deployments, these distribution shifts are under-represented in the datasets widely used in the ML community today. To address this gap, we present WILDS, a curated benchmark of 10 datasets reflecting a diverse range of distribution shifts that naturally arise in real-world applications, such as shifts across hospitals for tumor identification; across camera traps for wildlife monitoring; and across time and location in satellite imaging and poverty mapping. On each dataset, we show that standard training yields substantially lower out-of-distribution than in-distribution performance. This gap remains even with models trained by existing methods for tackling distribution shifts, underscoring the need for new methods for training models that are more robust to the types of distribution shifts that arise in practice. To facilitate method development, we provide an open-source package that automates dataset loading, contains default model architectures and hyperparameters, and standardizes evaluations. Code and leaderboards are available at https://wilds.stanford.edu.

📄 PDF Abstract BibTeX arXiv:2012.07421

Code (6)

p-lambda/wilds 공식 구현 pytorch
facebookresearch/DomainBed pytorch
hlzhang109/ddg pytorch
qiaoruiyt/noiserobustdg pytorch
skyve2012/DBA pytorch
tigrangalstyan/wilds pytorch

Tasks

Image Classification

Similar Papers 제목 키워드 기반

Extending the WILDS Benchmark for Unsupervised Adaptation

2021-12-09 · ICLR 2022 4 · Shiori Sagawa, Pang Wei Koh, Tony Lee, Irena Gao 외

Machine learning systems deployed in the wild are often trained on a source distribution but deployed on a different target distribution. Unlabeled data can be a powerful point of leverage for mitigating these distributi…

Accuracy on the Line: On the Strong Correlation Between Out-of-Distribution and In-Distribution Generalization

2021-07-09 · John Miller, Rohan Taori, aditi raghunathan, Shiori Sagawa 외

For machine learning systems to be reliable, we must understand their performance in unseen, out-of-distribution environments. In this paper, we empirically show that out-of-distribution performance is strongly correlate…

ClassificationDomain AdaptationPose Estimation

WildSVG: Towards Reliable SVG Generation Under Real-Word Conditions

2026-02-24 · Marco Terral, Haotian Zhang, Tianyang Zhang, Meng Lin 외 arxiv

We introduce the task of SVG extraction, which consists in translating specific visual inputs from an image into scalable vector graphics. Existing multimodal models achieve strong results when generating SVGs from clean…

Finetune like you pretrain: Improved finetuning of zero-shot vision models

2022-12-01 · CVPR 2023 1 · Sachin Goyal, Ananya Kumar, Sankalp Garg, Zico Kolter 외

Finetuning image-text models such as CLIP achieves state-of-the-art accuracies on a variety of benchmarks. However, recent works like WiseFT (Wortsman et al., 2021) and LP-FT (Kumar et al., 2022) have shown that even sub…

DescriptiveFew-Shot LearningTransfer Learning

Dual Risk Minimization: Towards Next-Level Robustness in Fine-tuning Zero-Shot Models

2024-11-29 · Kaican Li, Weiyan Xie, Yongxiang Huang, Didan Deng 외

Fine-tuning foundation models often compromises their robustness to distribution shifts. To remedy this, most robust fine-tuning methods aim to preserve the pre-trained features. However, not all pre-trained features are…