paper-with-me

Papers

Empirical Risk Minimization under Random Censorship: Theory and Practice

2019-06-05 · Guillaume Ausset, Stéphan Clémençon, François Portier

We consider the classic supervised learning problem, where a continuous non-negative random label $Y$ (i.e. a random duration) is to be predicted based upon observing a random vector $X$ valued in $\mathbb{R}^d$ with $d\geq 1$ by means of a regression rule with minimum least square error. In various applications, ranging from industrial quality control to public health through credit risk analysis for instance, training observations can be right censored, meaning that, rather than on independent copies of $(X,Y)$, statistical learning relies on a collection of $n\geq 1$ independent realizations of the triplet $(X, \; \min\{Y,\; C\},\; \delta)$, where $C$ is a nonnegative r.v. with unknown distribution, modeling censorship and $\delta=\mathbb{I}\{Y\leq C\}$ indicates whether the duration is right censored or not. As ignoring censorship in the risk computation may clearly lead to a severe underestimation of the target duration and jeopardize prediction, we propose to consider a plug-in estimate of the true risk based on a Kaplan-Meier estimator of the conditional survival function of the censorship $C$ given $X$, referred to as Kaplan-Meier risk, in order to perform empirical risk minimization. It is established, under mild conditions, that the learning rate of minimizers of this biased/weighted empirical risk functional is of order $O_{\mathbb{P}}(\sqrt{\log(n)/n})$ when ignoring model bias issues inherent to plug-in estimation, as can be attained in absence of censorship. Beyond theoretical results, numerical experiments are presented in order to illustrate the relevance of the approach developed.

📄 PDF Abstract BibTeX arXiv:1906.01908

Code (0)

등록된 구현이 없습니다.

Tasks

Triplet

Similar Papers 제목 키워드 기반

Disparate Censorship & Undertesting: A Source of Label Bias in Clinical Machine Learning

2022-08-01 · Trenton Chang, Michael W. Sjoding, Jenna Wiens

As machine learning (ML) models gain traction in clinical applications, understanding the impact of clinician and societal biases on ML models is increasingly important. While biases can arise in the labels used for mode…

BIG-bench Machine LearningDiagnostic

Optimal Empirical Risk Minimization under Temporal Distribution Shifts

2025-07-17 · Yujin Jeong, Ramesh Johari, Dominik Rothenhäusler, Emily Fox

Temporal distribution shifts pose a key challenge for machine learning models trained and deployed in dynamically evolving environments. This paper introduces RIDER (RIsk minimization under Dynamically Evolving Regimes) …

A Simple Analysis for Exp-concave Empirical Minimization with Arbitrary Convex Regularizer

2017-09-09 · Tianbao Yang, Zhe Li, Lijun Zhang

In this paper, we present a simple analysis of {\bf fast rates} with {\it high probability} of {\bf empirical minimization} for {\it stochastic composite optimization} over a finite-dimensional bounded convex set with ex…

LLM Censorship: A Machine Learning Challenge or a Computer Security Problem?

2023-07-20 · David Glukhov, Ilia Shumailov, Yarin Gal, Nicolas Papernot 외

Large language models (LLMs) have exhibited impressive capabilities in comprehending complex instructions. However, their blind adherence to provided instructions has led to concerns regarding risks of malicious use. Exi…

Computer SecurityInstruction Following

Learning with Noisy Labels

2013-12-01 · NeurIPS 2013 12 · Nagarajan Natarajan, Inderjit S. Dhillon, Pradeep K. Ravikumar, Ambuj Tewari

In this paper, we theoretically study the problem of binary classification in the presence of random classification noise --- the learner, instead of seeing the true labels, sees labels that have independently been flipp…

Binary ClassificationGeneral ClassificationLearning with noisy labels