paper-with-me

홈 › Papers

Assumption-Lean Post-Integrated Inference with Negative Control Outcomes

2024-10-07 · Jin-Hong Du, Kathryn Roeder, Larry Wasserman

Data integration methods aim to extract low-dimensional embeddings from high-dimensional outcomes to remove unwanted variations, such as batch effects and unmeasured covariates, across heterogeneous datasets. However, multiple hypothesis testing after integration can be biased due to data-dependent processes. We introduce a robust post-integrated inference (PII) method that adjusts for latent heterogeneity using negative control outcomes. Leveraging causal interpretations, we derive nonparametric identifiability of the direct effects, which motivates our semiparametric inference method. Our method extends to projected direct effect estimands, accounting for hidden mediators, confounders, and moderators. These estimands remain statistically meaningful under model misspecifications and with error-prone embeddings. We provide bias quantifications and finite-sample linear expansions with uniform concentration bounds. The proposed doubly robust estimators are consistent and efficient under minimal assumptions and potential misspecification, facilitating data-adaptive estimation with machine learning algorithms. Our proposal is evaluated with random forests through simulations and analysis of single-cell CRISPR perturbed datasets with potential unmeasured confounders.

📄 PDF Abstract BibTeX arXiv:2410.04996

Code (0)

등록된 구현이 없습니다.

Tasks

Data Integration

Similar Papers 제목 키워드 기반

Bayesian Nonparametric Boolean Factor Models

2019-06-28 · Tammo Rukat, Christopher Yau

We build upon probabilistic models for Boolean Matrix and Boolean Tensor factorisation that have recently been shown to solve these problems with unprecedented accuracy and to enable posterior inference to scale to Billi…

Assumption-Lean and Data-Adaptive Post-Prediction Inference

2023-11-23 · Jiacheng Miao, Xinran Miao, Yixuan Wu, Jiwei Zhao 외

A primary challenge facing modern scientific research is the limited availability of gold-standard data which can be costly, labor-intensive, or invasive to obtain. With the rapid development of machine learning (ML), sc…

Predictionvalid

Statistical Speech Enhancement Based on Probabilistic Integration of Variational Autoencoder and Non-Negative Matrix Factorization

2017-10-31 · Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama, Kazuyoshi Yoshii 외

This paper presents a statistical method of single-channel speech enhancement that uses a variational autoencoder (VAE) as a prior distribution on clean speech. A standard approach to speech enhancement is to train a dee…

Speech Enhancement

Prediction De-Correlated Inference: A safe approach for post-prediction inference

2023-12-11 · Feng Gan, Wanfeng Liang, Changliang Zou

In modern data analysis, it is common to use machine learning methods to predict outcomes on unlabeled datasets and then use these pseudo-outcomes in subsequent statistical inference. Inference in this setting is often c…

Prediction

Optimize the Unseen - Fast NeRF Cleanup with Free Space Prior

2024-12-17 · Leo Segre, Shai Avidan

Neural Radiance Fields (NeRF) have advanced photorealistic novel view synthesis, but their reliance on photometric reconstruction introduces artifacts, commonly known as "floaters". These artifacts degrade novel view qua…

NeRFNovel View Synthesis