paper-with-me

Papers

Fourier Feature Methods for Nonlinear Causal Discovery: FFML Scoring, TRFF Scoring, and FFCI Testing in Mixed Data

2026-05-07 · Joseph D. Ramsey arxiv

Gaussian process (GP) marginal likelihood scores and kernel conditional independence tests are theoretically appealing for nonlinear causal discovery but computationally prohibitive at scale. We present three complementary RFF-based methods forming a practical toolkit for score-based, constraint-based, and hybrid causal discovery. The Fourier Feature Marginal Likelihood (FFML) score approximates the exact GP marginal likelihood by replacing the $n x n$ kernel Gram matrix with a finite-dimensional feature representation, reducing cost to $O(nm^2 + m^3)$ while retaining the probabilistic interpretation and automatic complexity penalty of the exact score. FFML extends to mixed (continuous and discrete) parent sets via a product-kernel construction, with a Kronecker path for small discrete parent sets and a Hadamard-product path otherwise. The Tetrad Random Fourier Feature (TRFF) score is a complementary BIC-style alternative using penalized Student-t regression with random Fourier features. TRFF offers robustness to heavy-tailed noise and faster runtime than FFML. Empirically, TRFF and FFML exhibit a complementary precision-recall profile: TRFF achieves higher precision while FFML achieves better recall and lower SHD overall. The Fourier Feature Conditional Independence (FFCI) test is a fast nonparametric CI test for mixed data, using ridge residualization in feature space and a Frobenius-norm cross-covariance statistic approximated as a weighted sum of chi-squared variables. Empirically, BOSS+FFML achieves the lowest SHD on nonlinear data, while BOSS+TRFF offers the highest precision. When run through PC-Max, FFCI and RCIT exhibit complementary precision-recall profiles: RCIT is more precise while FFCI achieves better recall and substantially lower SHD, at approximately twice the runtime.

📄 PDF Abstract BibTeX arXiv:2605.05743

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Boosting Synthetic Data Generation with Effective Nonlinear Causal Discovery

2023-01-18 · Martina Cinquini, Fosca Giannotti, Riccardo Guidotti

Synthetic data generation has been widely adopted in software testing, data privacy, imbalanced learning, and artificial intelligence explanation. In all such contexts, it is crucial to generate plausible data samples. A…

Causal Discoverysoftware testingSynthetic Data Generation

Constraint- and Score-Based Nonlinear Granger Causality Discovery with Kernels

2026-01-14 · Fiona Murphy, Alessio Benavoli arxiv

Kernel-based methods are used in the context of Granger Causality to enable the identification of nonlinear causal relationships between time series variables. In this paper, we show that two state of the art kernel-base…

Score matching through the roof: linear, nonlinear, and latent variables causal discovery

2024-07-26 · Francesco Montagna, Philipp M. Faller, Patrick Bloebaum, Elke Kirschbaum 외

Causal discovery from observational data holds great promise, but existing methods rely on strong assumptions about the underlying causal structure, often requiring full observability of all relevant variables. We tackle…

Causal Discovery

MDL Meets Latent Confounders: LNML-based Causal Discovery

2026-07-05 · Zhongyi Que, Shin Matsushima, Kenji Yamanishi arxiv

Causal discovery with nonlinear mechanisms and latent confounders remains challenging. Existing methods often rely on either linear assumptions or causal sufficiency, limiting their applicability. We propose an MDL-based…

Nonlinear causal discovery with additive noise models

2008-12-01 · NeurIPS 2008 12 · Patrik O. Hoyer, Dominik Janzing, Joris M. Mooij, Jonas Peters 외

The discovery of causal relationships between a set of observed variables is a fundamental problem in science. For continuous-valued data linear acyclic causal models are often used because these models are well understo…

Causal Discovery