paper-with-me

홈 › Papers

The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction

2026-05-28 · Shu Wan, Abhinav Gorantla, Huan Liu, K. Selçuk Candan arxiv

Under standard graphical assumptions, the Markov boundary of a target variable is the smallest set of features that renders every other feature redundant. Once the boundary is observed, the target is conditionally independent of the rest of the table. This is a tempting object for tabular prediction, since it names exactly the columns a model should need. Yet modern regressors are still trained on the full feature set. We ask whether the Markov boundary is genuinely useful for prediction on SCM3K, a 3,450-task synthetic SCM benchmark with feature counts from 40 to 1000 and six SCM families, evaluated with six regressors. The answer is more nuanced than the theory suggests. Restricting a regressor to the oracle boundary often improves prediction substantially, and the improvement grows as the feature space becomes larger and sparser. But the natural pipeline of recovering the boundary with causal discovery and training on the recovered mask does not deliver. Existing estimators exhaust the compute budget before reaching the regime where the boundary helps most, and even where they run they rarely beat the full feature set. We trace this to three causes. Discovery optimizes structural recovery rather than prediction. False negatives and false positives carry sharply asymmetric predictive cost. The exact boundary is only one of many feature sets that beat all features. We then develop what these facts imply for prediction-aligned feature selection and for tabular models that learn to use causal structure.

📄 PDF Abstract BibTeX arXiv:2605.29411

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Marketron games: Self-propelling stocks vs dumb money and metastable dynamics of the Good, Bad and Ugly markets

2025-01-22 · I. Halperin, A. Itkin

We present a model of price formation in an inelastic market whose dynamics are partially driven by both money flows and their impact on asset prices. The money flow to the market is viewed as an investment policy of out…

Assessing Good, Bad and Ugly Arguments Generated by ChatGPT: a New Dataset, its Methodology and Associated Tasks

2024-06-21 · Victor Hugo Nascimento Rocha, Igor Cataneo Silveira, Paulo Pirozelli, Denis Deratani Mauá 외

The recent success of Large Language Models (LLMs) has sparked concerns about their potential to spread misinformation. As a result, there is a pressing need for tools to identify ``fake arguments'' generated by such mod…

Misinformation

Nonlinearity in Dynamic Causal Effects: Making the Bad into the Good, and the Good into the Great?

2025-04-01 · Toru Kitagawa, Weining Wang, Mengshan Xu

This paper was prepared as a comment on "Dynamic Causal Effects in a Nonlinear World: the Good, the Bad, and the Ugly" by Michal Koles\'ar, Mikkel Plagborg-M{\o}ller. We make three comments, including a novel contributio…

AugLy: Data Augmentations for Robustness

2022-01-17 · Zoe Papakipos, Joanna Bitton

We introduce AugLy, a data augmentation library with a focus on adversarial robustness. AugLy provides a wide array of augmentations for multiple modalities (audio, image, text, & video). These augmentations were inspire…

Adversarial RobustnessData Augmentation

Neural topology optimization: the good, the bad, and the ugly

2024-07-19 · Suryanarayanan Manoj Sanu, Alejandro M. Aragon, Miguel A. Bessa

Neural networks (NNs) hold great promise for advancing inverse design via topology optimization (TO), yet misconceptions about their application persist. This article focuses on neural topology optimization (neural TO), …

GPUMisconceptions