paper-with-me

Papers

Propensity score estimation using classification and regression trees in the presence of missing covariate data

2018-07-25 · Bas B. L. Penning de Vries, Maarten van Smeden, Rolf H. H. Groenwold

Data mining and machine learning techniques such as classification and regression trees (CART) represent a promising alternative to conventional logistic regression for propensity score estimation. Whereas incomplete data preclude the fitting of a logistic regression on all subjects, CART is appealing in part because some implementations allow for incomplete records to be incorporated in the tree fitting and provide propensity score estimates for all subjects. Based on theoretical considerations, we argue that the automatic handling of missing data by CART may however not be appropriate. Using a series of simulation experiments, we examined the performance of different approaches to handling missing covariate data; (i) applying the CART algorithm directly to the (partially) incomplete data, (ii) complete case analysis, and (iii) multiple imputation. Performance was assessed in terms of bias in estimating exposure-outcome effects \add{among the exposed}, standard error, mean squared error and coverage. Applying the CART algorithm directly to incomplete data resulted in bias, even in scenarios where data were missing completely at random. Overall, multiple imputation followed by CART resulted in the best performance. Our study showed that automatic handling of missing data in CART can cause serious bias and does not outperform multiple imputation as a means to account for missing data.

📄 PDF Abstract BibTeX arXiv:1807.09462

Code (0)

등록된 구현이 없습니다.

Tasks

General ClassificationImputationregression

Methods 이 논문이 사용한 방법론

Logistic Regression Logistic Regression, despite its name, is a linear model for classification rather than regression. Logistic regression is also known in the literature as logit regression,…

Similar Papers 제목 키워드 기반

Statistical inference using Regularized M-estimation in the reproducing kernel Hilbert space for handling missing data

2021-07-15 · Hengfang Wang, Jae Kwang Kim

Imputation and propensity score weighting are two popular techniques for handling missing data. We address these problems using the regularized M-estimation techniques in the reproducing kernel Hilbert space. Specificall…

Imputationregression

Semiparametric Methods for Exposure Misclassification in Propensity Score-Based Time-to-Event Data Analysis

2019-03-19 · Yingrui Yang, Molin Wang

In epidemiology, identifying the effect of exposure variables in relation to a time-to-event outcome is a classical research area of practical importance. Incorporating propensity score in the Cox regression model, as a …

Epidemiology

Stabilized Inverse Probability Weighting via Isotonic Calibration

2024-11-10 · Lars van der Laan, Ziming Lin, Marco Carone, Alex Luedtke

Inverse weighting with an estimated propensity score is widely used by estimation methods in causal inference to adjust for confounding bias. However, directly inverting propensity score estimates can lead to instability…

Causal Inference

Semi-supervised learning and the question of true versus estimated propensity scores

2020-09-14 · Andrew Herren, P. Richard Hahn

A straightforward application of semi-supervised machine learning to the problem of treatment effect estimation would be to consider data as "unlabeled" if treatment assignment and covariates are observed but outcomes ar…

Causal Inference

Robust inference on the average treatment effect using the outcome highly adaptive lasso

2018-06-18 · Cheng Ju, David Benkeser, Mark J. Van Der Laan

Many estimators of the average effect of a treatment on an outcome require estimation of the propensity score, the outcome regression, or both. It is often beneficial to utilize flexible techniques such as semiparametric…

regression