Papers Selection bias
“Selection bias” 태그가 달린 논문 365편 · 필터 해제
Metric-DST: Mitigating Selection Bias Through Diversity-Guided Semi-Supervised Metric Learning
Selection bias poses a critical challenge for fairness in machine learning, as models trained on data that is less representative of the population might exhibit undesirable behavior for underrepresented profiles. Semi-s…
DiversityFairnessMetric LearningSelection biasA Flexible Defense Against the Winner's Curse
Across science and policy, decision-makers often need to draw conclusions about the best candidate among competing alternatives. For instance, researchers may seek to infer the effectiveness of the most successful treatm…
Selection biasvalidLee Bounds with a Continuous Treatment in Sample Selection
We study causal inference in sample selection models where a continuous or multivalued treatment affects both outcomes and their observability (e.g., employment or survey responses). We generalized the widely used Lee (2…
Causal InferenceSelection biasA Directional Rockafellar-Uryasev Regression
Most ost Big Data datasets suffer from selection bias. For example, X (Twitter) training observations differ largely from the testing offline observations as individuals on Twitter are generally more educated, democratic…
FairnessregressionSelection biasDiversidade linguística e inclusão digital: desafios para uma ia brasileira
Linguistic diversity is a human attribute which, with the advance of generative AIs, is coming under threat. This paper, based on the contributions of sociolinguistics, examines the consequences of the variety selection …
AttributeDiversitySelection biasNot All Languages are Equal: Insights into Multilingual Retrieval-Augmented Generation
RALMs (Retrieval-Augmented Language Models) broaden their knowledge scope by incorporating external textual resources. However, the multilingual nature of global knowledge necessitates RALMs to handle diverse languages, …
AllRetrievalRetrieval-augmented GenerationSelection bias+1Contextual Representation Anchor Network to Alleviate Selection Bias in Few-Shot Drug Discovery
In the drug discovery process, the low success rate of drug candidate screening often leads to insufficient labeled data, causing the few-shot learning problem in molecular property prediction. Existing methods for few-s…
Drug DiscoveryFew-Shot LearningMolecular Property PredictionProperty Prediction+1RecFlow: An Industrial Full Flow Recommendation Dataset
Industrial recommendation systems (RS) rely on the multi-stage pipeline to balance effectiveness and efficiency when delivering items from a vast corpus to users. Existing RS benchmark datasets primarily focus on the exp…
Recommendation SystemsSelection biasA Systematic Review of Machine Learning Approaches for Detecting Deceptive Activities on Social Media: Methods, Challenges, and Biases
Social media platforms like Twitter, Facebook, and Instagram have facilitated the spread of misinformation, necessitating automated detection systems. This systematic review evaluates 36 studies that apply machine learni…
MisinformationSelection biasInference on High Dimensional Selective Labeling Models
A class of simultaneous equation models arise in the many domains where observed binary outcomes are themselves a consequence of the existing choices of of one of the agents in the model. These models are gaining increas…
Selection biasHeterogeneous Random Forest
Random forest (RF) stands out as a highly favored machine learning approach for classification problems. The effectiveness of RF hinges on two key factors: the accuracy of individual trees and the diversity among them. I…
DiversitySelection biasLeveraging CORAL-Correlation Consistency Network for Semi-Supervised Left Atrium MRI Segmentation
Semi-supervised learning (SSL) has been widely used to learn from both a few labeled images and many unlabeled images to overcome the scarcity of labeled samples in medical image segmentation. Most current SSL-based segm…
Image SegmentationMedical Image SegmentationMRI segmentationSelection bias+1CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges
The use of large language models (LLMs) as automated evaluation tools to assess the quality of generated natural language, known as LLMs-as-Judges, has demonstrated promising capabilities and is rapidly gaining widesprea…
FairnessPredictionSelection biasAddressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models
Evaluating Video Language Models (VLMs) is a challenging task. Due to its transparency, Multiple-Choice Question Answering (MCQA) is widely used to measure the performance of these models through accuracy. However, exist…
FairnessMultiple-choiceMultiple Choice Question Answering (MCQA)Question Answering+1DiffPO: A causal diffusion model for learning distributions of potential outcomes
Predicting potential outcomes of interventions from observational data is crucial for decision-making in medicine, but the task is challenging due to the fundamental problem of causal inference. Existing methods are larg…
Causal InferenceDecision MakingDenoisingSelection biasPartially Identified Heterogeneous Treatment Effect with Selection: An Application to Gender Gaps
This paper addresses the sample selection model within the context of the gender gap problem, where even random treatment assignment is affected by selection bias. By offering a robust alternative free from distributiona…
Selection biasDCAST: Diverse Class-Aware Self-Training Mitigates Selection Bias for Fairer Learning
Fairness in machine learning seeks to mitigate model bias against individuals based on sensitive features such as sex or age, often caused by an uneven representation of the population in the training data due to selecti…
DiversityDomain AdaptationFairnessMulti-class Classification+1Mitigating Selection Bias with Node Pruning and Auxiliary Options
Large language models (LLMs) often show unwarranted preference for certain choice options when responding to multiple-choice questions, posing significant reliability concerns in LLM-automated systems. To mitigate this s…
Multiple-choiceSelection biasNavigating the Maze of Explainable AI: A Systematic Approach to Evaluating Methods and Metrics
Explainable AI (XAI) is a rapidly growing domain with a myriad of proposed methods as well as metrics aiming to evaluate their efficacy. However, current studies are often of limited scope, examining only a handful of XA…
Selection biasDecolonising Data Systems: Using Jyutping or Pinyin as tonal representations of Chinese names for data linkage
Data linkage is increasingly used in health research and policy making and is relied on for understanding health inequalities. However, linked data is only as useful as the underlying data quality, and differential linka…
Selection bias