paper-with-me

홈 › Papers

Learning With Multi-Group Guarantees For Clusterable Subpopulations

2024-10-18 · Jessica Dai, Nika Haghtalab, Eric Zhao

A canonical desideratum for prediction problems is that performance guarantees should hold not just on average over the population, but also for meaningful subpopulations within the overall population. But what constitutes a meaningful subpopulation? In this work, we take the perspective that relevant subpopulations should be defined with respect to the clusters that naturally emerge from the distribution of individuals for which predictions are being made. In this view, a population refers to a mixture model whose components constitute the relevant subpopulations. We suggest two formalisms for capturing per-subgroup guarantees: first, by attributing each individual to the component from which they were most likely drawn, given their features; and second, by attributing each individual to all components in proportion to their relative likelihood of having been drawn from each component. Using online calibration as a case study, we study a multi-objective algorithm that provides guarantees for each of these formalisms by handling all plausible underlying subpopulation structures simultaneously, and achieve an $O(T^{1/2})$ rate even when the subpopulations are not well-separated. In comparison, the more natural cluster-then-predict approach that first recovers the structure of the subpopulations and then makes predictions suffers from a $O(T^{2/3})$ rate and requires the subpopulations to be separable. Along the way, we prove that providing per-subgroup calibration guarantees for underlying clusters can be easier than learning the clusters: separation between median subgroup features is required for the latter but not the former.

📄 PDF Abstract BibTeX arXiv:2410.14588

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multigroup Robustness

2024-05-01 · Lunjia Hu, Charlotte Peale, Judy Hanwen Shen

To address the shortcomings of real-world datasets, robust learning algorithms have been designed to overcome arbitrary and indiscriminate data corruption. However, practical processes of gathering data may lead to patte…

Fairness

Statistical Inference for Fairness Auditing

2023-05-05 · John J. Cherian, Emmanuel J. Candès

Before deploying a black-box model in high-stakes problems, it is important to evaluate the model's performance on sensitive subpopulations. For example, in a recidivism prediction task, we may wish to identify demograph…

Fairness

Distributionally Robust Losses for Latent Covariate Mixtures

2020-07-28 · John Duchi, Tatsunori Hashimoto, Hongseok Namkoong

While modern large-scale datasets often consist of heterogeneous subpopulations -- for example, multiple demographic groups or multiple text corpora -- the standard practice of minimizing average loss fails to guarantee …

Maximin Relative Improvement: Fair Learning as a Bargaining Problem

2026-02-04 · Jiwoo Han, Moulinath Banerjee, Yuekai Sun arxiv

When deploying a single predictor across multiple subpopulations, we propose a fundamentally different approach: interpreting group fairness as a bargaining problem among subpopulations. This game-theoretic perspective r…

Clusterability in Neural Networks

2021-03-04 · Daniel Filan, Stephen Casper, Shlomi Hod, Cody Wild 외

The learned weights of a neural network have often been considered devoid of scrutable internal structure. In this paper, however, we look for structure in the form of clusterability: how well a network can be divided in…