paper-with-me

Papers

Diagnosing failures of fairness transfer across distribution shift in real-world medical settings

2022-02-02 · Jessica Schrouff, Natalie Harris, Oluwasanmi Koyejo, Ibrahim Alabdulmohsin, Eva Schnider, Krista Opsahl-Ong, Alex Brown, Subhrajit Roy, Diana Mincu, Christina Chen, Awa Dieng, YuAn Liu, Vivek Natarajan, Alan Karthikesalingam, Katherine Heller, Silvia Chiappa, Alexander D'Amour

Diagnosing and mitigating changes in model fairness under distribution shift is an important component of the safe deployment of machine learning in healthcare settings. Importantly, the success of any mitigation strategy strongly depends on the structure of the shift. Despite this, there has been little discussion of how to empirically assess the structure of a distribution shift that one is encountering in practice. In this work, we adopt a causal framing to motivate conditional independence tests as a key tool for characterizing distribution shifts. Using our approach in two medical applications, we show that this knowledge can help diagnose failures of fairness transfer, including cases where real-world shifts are more complex than is often assumed in the literature. Based on these results, we discuss potential remedies at each step of the machine learning pipeline.

📄 PDF Abstract BibTeX arXiv:2202.01034

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningFairness

Similar Papers 제목 키워드 기반

Diagnosing Conformal Prediction Failures Under Distribution Shift: A COVID-19 Case Study

2026-01-01 · Chorok Lee arxiv

Conformal prediction provides distribution-free coverage guarantees, but these degrade under distribution shift - and practitioners lack tools to anticipate which deployed models will fail before observing test data. We …

Generalization of Long-Range Machine Learning Potentials in Complex Chemical Spaces

2025-12-08 · Michal Sanocki, Julija Zavadlav arxiv

The vastness of chemical space makes generalization a central challenge in the development of machine learning interatomic potentials (MLIPs). While MLIPs could enable large-scale atomistic simulations with near-quantum …

Long-range modeling

Diagnosing Generalization Failures from Representational Geometry Markers

2026-03-02 · Chi-Ning Chou, Artem Kirsanov, Yao-Yuan Yang, SueYeon Chung arxiv

Generalization, the ability to perform well beyond the training context, is a hallmark of biological and artificial intelligence, yet anticipating unseen failures remains a central challenge. Conventional approaches ofte…

Image ClassificationTransfer Learning

C2PO: Diagnosing and Disentangling Bias Shortcuts in LLMs

2025-12-29 · Xuan Feng, Bo An, Tianlong Gu, Liang Chang 외 arxiv

Bias in Large Language Models (LLMs) poses significant risks to trustworthiness, manifesting primarily as stereotypical biases (e.g., gender or racial stereotypes) and structural biases (e.g., lexical overlap or position…

LLM-Based Automated Diagnosis Of Integration Test Failures At Google

2026-04-13 · Celal Ziftci, Ray Liu, Spencer Greene, Livio Dalloro arxiv

Integration testing is critical for the quality and reliability of complex software systems. However, diagnosing their failures presents significant challenges due to the massive volume, unstructured nature, and heteroge…