paper-with-me

홈 › Papers

Rethinking Robustness in Machine Learning: A Posterior Agreement Approach

2025-03-20 · João Borges S. Carvalho, Alessandro Torcinovich, Victor Jimenez Rodriguez, Antonio E. Cinà, Carlos Cotrini, Lea Schönherr, Joachim M. Buhmann

The robustness of algorithms against covariate shifts is a fundamental problem with critical implications for the deployment of machine learning algorithms in the real world. Current evaluation methods predominantly match the robustness definition to that of standard generalization, relying on standard metrics like accuracy-based scores, which, while designed for performance assessment, lack a theoretical foundation encompassing their application in estimating robustness to distribution shifts. In this work, we set the desiderata for a robustness metric, and we propose a novel principled framework for the robustness assessment problem that directly follows the Posterior Agreement (PA) theory of model validation. Specifically, we extend the PA framework to the covariate shift setting by proposing a PA metric for robustness evaluation in supervised classification tasks. We assess the soundness of our metric in controlled environments and through an empirical robustness analysis in two different covariate shift scenarios: adversarial learning and domain generalization. We illustrate the suitability of PA by evaluating several models under different nature and magnitudes of shift, and proportion of affected observations. The results show that the PA metric provides a sensible and consistent analysis of the vulnerabilities in learning algorithms, even in the presence of few perturbed observations.

📄 PDF Abstract BibTeX arXiv:2503.16271

Code (1)

viictorjimenezzz/pa-covariate-shift 공식 구현 pytorch

Tasks

Domain Generalization

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Diagnostic Uncertainty Calibration: Towards Reliable Machine Predictions in Medical Domain

2020-07-03 · Takahiro Mimori, Keiko Sasada, Hirotaka Matsui, Issei Sato

We propose an evaluation framework for class probability estimates (CPEs) in the presence of label uncertainty, which is commonly observed as diagnosis disagreement between experts in the medical domain. We also formaliz…

Diagnostic

Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior

2025-01-31 · Tongda Xu, Xiyan Cai, Xinjie Zhang, Xingtong Ge 외

Recent advancements in diffusion models have been leveraged to address inverse problems without additional training, and Diffusion Posterior Sampling (DPS) (Chung et al., 2022a) is among the most popular approaches. Prev…

GPU

Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains

2026-04-19 · Finn Schmidt, Jan Philip Wahle, Terry Ruas, Bela Gipp arxiv

Automatic evaluation metrics are central to the development of machine translation systems, yet their robustness under domain shift remains unclear. Most metrics are developed on the Workshop on Machine Translation (WMT)…

Machine Translation

End-to-End Weak Supervision

2021-07-05 · NeurIPS 2021 12 · Salva Rühling Cachay, Benedikt Boecking, Artur Dubrawski

Aggregating multiple sources of weak supervision (WS) can ease the data-labeling bottleneck prevalent in many machine learning applications, by replacing the tedious manual collection of ground truth labels. Current stat…

Classification

Rethinking the Agreement in Human Evaluation Tasks

2018-08-01 · COLING 2018 8 · Jacopo Amidei, Paul Piwek, Alistair Willis

Human evaluations are broadly thought to be more valuable the higher the inter-annotator agreement. In this paper we examine this idea. We will describe our experiments and analysis within the area of Automatic Question …

Dialogue GenerationQuestion GenerationQuestion-GenerationText Generation