paper-with-me

홈 › Papers

UpliftBench: Revealing Outcome-Regime and Objective Mismatch in Uplift Evaluation

2026-08-02 · Binshuang Li arxiv

Uplift modeling (conditional-average-treatment-effect estimation) drives personalized targeting, yet published uplift benchmarks frequently disagree on which estimator performs best; we show the disagreement is substantially about metrics, not models. UpliftBench evaluates 12 uplift estimators under an outer-test-isolated, multi-objective protocol across seven dataset families; its two findings are identified where a reference objective exists -- F1 on the standard continuous benchmark (IHDP), F2 in a within-sample case study on Jobs. On that benchmark, Qini shows no detectable alignment with effect accuracy -- across all 100 IHDP realizations its mean rank correlation with effect accuracy is +0.07, 95% CI [-0.03, +0.16] -- while AUUC is consistently more aligned (paired prefix-mean-AUUC-over-Qini gap +0.49 [+0.40, +0.59]; the shipped cumulative-gain AUUC aligns better still, +0.73). On Jobs, ranking metrics are structurally insufficient for a sign-threshold policy because they discard the score level; empirically, within the released split-rotation analysis direct policy-risk selection yields lower benchmark regret than random model selection while Qini, AUUC, and uplift-at-$k$ do not (14-15% regret). Calibrating the decision threshold removes 81% of the Qini-selection regret. Both findings are bounded, not universal: F1 is not detected on either validation family (the ACIC and Revenue-Synthetic gaps are both indistinguishable from zero), and F2 vanishes under a budgeted-value objective where rank suffices. UpliftBench releases versioned loaders, fixed protocols, result artifacts, and a reproducible living leaderboard; the public repository accompanies the paper.

📄 PDF Abstract BibTeX arXiv:2608.00915

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Optimized Beamforming for Joint Bistatic Positioning and Monostatic Sensing

2025-01-20 · Yuchen Zhang, Hui Chen, Pinjun Zheng, Boyu Ning 외

We investigate the performance tradeoff between \textit{bistatic positioning (BP)} and \textit{monostatic sensing (MS)} in a multi-input multi-output orthogonal frequency division multiplexing scenario. We derive the Cra…

EpiEvolve: Self-Evolving Agents for Streaming Pandemic Forecasting under Regime Shifts

2026-06-03 · Yiming Lu, Sihang Zeng, Zhengxu Tang, Max Lau 외 arxiv

Epidemic LLM forecasters are usually trained and evaluated as static supervised models, whereas operational pandemic forecasting is a streaming process in which labels arrive after predictions and disease regimes shift o…

Improving RCT-Based Treatment Effect Estimation Under Covariate Mismatch via Calibrated Alignment

2026-03-19 · Amir Asiaee, Samhita Pal arxiv

Randomized controlled trials (RCTs) are the gold standard for estimating treatment effects, yet they are often underpowered for detecting effect heterogeneity. Large observational studies (OS) can supplement RCTs for con…

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners

2026-06-28 · Chao Wang, Hongtao Tian, Tao Yang, Yunsheng Shi 외 arxiv

Group Relative Policy Optimization (GRPO) is a default recipe for process-supervised reinforcement learning of LLM reasoners, and dense process supervision -- via learned process reward models (PRMs) or on-policy-distill…

Multi-hop Question AnsweringReinforcement LearningMathematical Reasoning

Estimating Bayesian Optimal Treatment Regimes for Dichotomous Outcomes using Observational Data

2018-09-18 · Thomas Klausch, Peter van de Ven, Tim van de Brug, Mark A. van de Wiel 외

Optimal treatment regimes (OTR) are individualised treatment assignment strategies that identify a medical treatment as optimal given all background information available on the individual. We discuss Bayes optimal treat…

Bayesian Inference