paper-with-me

Papers

Estimating the unseen from multiple populations

2017-07-12 · ICML 2017 8 · Aditi Raghunathan, Greg Valiant, James Zou

Given samples from a distribution, how many new elements should we expect to find if we continue sampling this distribution? This is an important and actively studied problem, with many applications ranging from unseen species estimation to genomics. We generalize this extrapolation and related unseen estimation problems to the multiple population setting, where population $j$ has an unknown distribution $D_j$ from which we observe $n_j$ samples. We derive an optimal estimator for the total number of elements we expect to find among new samples across the populations. Surprisingly, we prove that our estimator's accuracy is independent of the number of populations. We also develop an efficient optimization algorithm to solve the more general problem of estimating multi-population frequency distributions. We validate our methods and theory through extensive experiments. Finally, on a real dataset of human genomes across multiple ancestries, we demonstrate how our approach for unseen estimation can enable cohort designs that can discover interesting mutations with greater efficiency.

📄 PDF Abstract BibTeX arXiv:1707.03854

Code (2)

roydeb/unseen_estimator
siddarthhari95/unseen_estimator

Similar Papers 제목 키워드 기반

Distributionally Robust Losses for Latent Covariate Mixtures

2020-07-28 · John Duchi, Tatsunori Hashimoto, Hongseok Namkoong

While modern large-scale datasets often consist of heterogeneous subpopulations -- for example, multiple demographic groups or multiple text corpora -- the standard practice of minimizing average loss fails to guarantee …

Estimating Probabilities of Causation with Machine Learning Models

2025-02-13 · Shuai Wang, Ang Li

Probabilities of causation play a crucial role in modern decision-making. This paper addresses the challenge of predicting probabilities of causation for subpopulations with insufficient data using machine learning model…

Estimating Individual Treatment Effects through Causal Populations Identification

2020-04-10 · Céline Beji, Michaël Bon, Florian Yger, Jamal Atif

Estimating the Individual Treatment Effect from observational data, defined as the difference between outcomes with and without treatment or intervention, while observing just one of both, is a challenging problems in ca…

Discriminative Subspace Emersion from learning feature relevances across different populations

2025-03-31 · Marco Canducci, Lida Abdi, Alessandro Prete, Roland J. Veen 외

In a given classification task, the accuracy of the learner is often hampered by finiteness of the training set, high-dimensionality of the feature space and severe overlap between classes. In the context of interpretabl…

Classification

Reinforcement Learning with Heterogeneous Data: Estimation and Inference

2022-01-31 · Elynn Y. Chen, Rui Song, Michael I. Jordan

Reinforcement Learning (RL) has the promise of providing data-driven support for decision-making in a wide range of problems in healthcare, education, business, and other domains. Classical RL methods focus on the mean o…

Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)