paper-with-me

홈 › Papers

Statistical Inference for Feature Selection after Optimal Transport-based Domain Adaptation

2024-10-19 · Nguyen Thang Loi, Duong Tan Loc, Vo Nguyen Le Duy

Feature Selection (FS) under domain adaptation (DA) is a critical task in machine learning, especially when dealing with limited target data. However, existing methods lack the capability to guarantee the reliability of FS under DA. In this paper, we introduce a novel statistical method to statistically test FS reliability under DA, named SFS-DA (statistical FS-DA). The key strength of SFS-DA lies in its ability to control the false positive rate (FPR) below a pre-specified level $\alpha$ (e.g., 0.05) while maximizing the true positive rate. Compared to the literature on statistical FS, SFS-DA presents a unique challenge in addressing the effect of DA to ensure the validity of the inference on FS results. We overcome this challenge by leveraging the Selective Inference (SI) framework. Specifically, by carefully examining the FS process under DA whose operations can be characterized by linear and quadratic inequalities, we prove that achieving FPR control in SFS-DA is indeed possible. Furthermore, we enhance the true detection rate by introducing a more strategic approach. Experiments conducted on both synthetic and real-world datasets robustly support our theoretical results, showcasing the superior performance of the proposed SFS-DA method.

📄 PDF Abstract BibTeX arXiv:2410.15022

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Adaptationfeature selection

Similar Papers 제목 키워드 기반

Statistical Inference for Sequential Feature Selection after Domain Adaptation

2025-01-17 · Duong Tan Loc, Nguyen Thang Loi, Vo Nguyen Le Duy

In high-dimensional regression, feature selection methods, such as sequential feature selection (SeqFS), are commonly used to identify relevant features. When data is limited, domain adaptation (DA) becomes crucial for t…

Domain Adaptationfeature selectionModel Selection

Interval Estimation of Coefficients in Penalized Regression Models of Insurance Data

2024-10-01 · Alokesh Manna, Zijian Huang, Dipak K. Dey, Yuwen Gu 외

The Tweedie exponential dispersion family is a popular choice among many to model insurance losses that consist of zero-inflated semicontinuous data. In such data, it is often important to obtain credibility (inference) …

feature selectionregressionvalid

Statistical inference after variable selection in Cox models: A simulation study

2026-02-07 · Lena Schemet, Sarah Friedrich-Welz arxiv

Choosing relevant predictors is central to the analysis of biomedical time-to-event data. Classical frequentist inference, however, presumes that the set of covariates is fixed in advance and does not account for data-dr…

More Powerful Selective Kernel Tests for Feature Selection

2019-10-14 · Jen Ning Lim, Makoto Yamada, Wittawat Jitkrittum, Yoshikazu Terada 외

Refining one's hypotheses in the light of data is a common scientific practice; however, the dependency on the data introduces selection bias and can lead to specious statistical analysis. An approach for addressing this…

feature selectionSelection bias

Statistical Test for Feature Selection Pipelines by Selective Inference

2024-06-27 · Tomohiro Shiraishi, Tatsuya Matsukawa, Shuichi Nishino, Ichiro Takeuchi

A data analysis pipeline is a structured sequence of steps that transforms raw data into meaningful insights by integrating various analysis algorithms. In this paper, we propose a novel statistical test to assess the si…

feature selectionImputationOutlier Detectionvalid