Effect Inference from Two-Group Data with Sampling Bias
In many applications, different populations are compared using data that are sampled in a biased manner. Under sampling biases, standard methods that estimate the difference between the population means yield unreliable inferences. Here we develop an inference method that is resilient to sampling biases and is able to control the false positive errors under moderate bias levels in contrast to the standard approach. We demonstrate the method using synthetic and real biomarker data.
Code (1)
Tasks
Vocal Bursts Valence PredictionSimilar Papers 제목 키워드 기반
Homogeneity Bias as Differential Sampling Uncertainty in Language Models
Prior research show that Large Language Models (LLMs) and Vision-Language Models (VLMs) represent marginalized groups more homogeneously than dominant groups. However, the mechanisms underlying this homogeneity bias rema…
Does a Rising Tide Lift All Boats? Bias Mitigation for AI-based CMR Segmentation
Artificial intelligence (AI) is increasingly being used for medical imaging tasks. However, there can be biases in the resulting models, particularly when they were trained using imbalanced training datasets. One such ex…
AllImage SegmentationSemantic SegmentationShedding light on underrepresentation and Sampling Bias in machine learning
Accurately measuring discrimination is crucial to faithfully assessing fairness of trained machine learning (ML) models. Any bias in measuring discrimination leads to either amplification or underestimation of the existi…
FairnessPushing the Accuracy-Group Robustness Frontier with Introspective Self-play
Standard empirical risk minimization (ERM) training can produce deep neural network (DNN) models that are accurate on average but under-perform in under-represented population subgroups, especially when there are imbalan…
Active LearningFairnessRandom Silicon Sampling: Simulating Human Sub-Population Opinion Using a Large Language Model Based on Group-Level Demographic Information
Large language models exhibit societal biases associated with demographic information, including race, gender, and others. Endowing such language models with personalities based on demographic data can enable generating …
Language ModelingLanguage ModellingLarge Language Model