paper-with-me

Papers

Variable Selection in Maximum Mean Discrepancy for Interpretable Distribution Comparison

2023-11-02 · Kensuke Mitsuzawa, Motonobu Kanagawa, Stefano Bortoli, Margherita Grossi, Paolo Papotti

Two-sample testing decides whether two datasets are generated from the same distribution. This paper studies variable selection for two-sample testing, the task being to identify the variables (or dimensions) responsible for the discrepancies between the two distributions. This task is relevant to many problems of pattern analysis and machine learning, such as dataset shift adaptation, causal inference and model validation. Our approach is based on a two-sample test based on the Maximum Mean Discrepancy (MMD). We optimise the Automatic Relevance Detection (ARD) weights defined for individual variables to maximise the power of the MMD-based test. For this optimisation, we introduce sparse regularisation and propose two methods for dealing with the issue of selecting an appropriate regularisation parameter. One method determines the regularisation parameter in a data-driven way, and the other aggregates the results of different regularisation parameters. We confirm the validity of the proposed methods by systematic comparisons with baseline methods, and demonstrate their usefulness in exploratory analysis of high-dimensional traffic simulation data. Preliminary theoretical analyses are also provided, including a rigorous definition of variable selection for two-sample testing.

📄 PDF Abstract BibTeX arXiv:2311.01537

Code (0)

등록된 구현이 없습니다.

Tasks

Causal InferenceRelevance DetectionTwo-sample testingVariable Selection

Methods 이 논문이 사용한 방법론

Causal inference Causal inference is the process of drawing a conclusion about a causal connection based on the conditions of the occurrence of an effect. The main difference between causal…

Similar Papers 제목 키워드 기반

Post Selection Inference with Incomplete Maximum Mean Discrepancy Estimator

2018-02-17 · ICLR 2019 5 · Makoto Yamada, Denny Wu, Yao-Hung Hubert Tsai, Ichiro Takeuchi 외

Measuring divergence between two distributions is essential in machine learning and statistics and has various applications including binary classification, change point detection, and two-sample test. Furthermore, in th…

Binary ClassificationChange Point Detectionfeature selection

Conditional Generative Moment-Matching Networks

2016-06-14 · NeurIPS 2016 12 · Yong Ren, Jialian Li, Yucen Luo, Jun Zhu

Maximum mean discrepancy (MMD) has been successfully applied to learn deep generative models for characterizing a joint distribution of variables via kernel mean embedding. In this paper, we present conditional generativ…

Variable Selection for Kernel Two-Sample Tests

2023-02-15 · Jie Wang, Santanu S. Dey, Yao Xie

We consider the variable selection problem for two-sample tests, aiming to select the most informative variables to determine whether two collections of samples follow the same distribution. To address this, we propose a…

Variable SelectionVocal Bursts Valence Prediction

Variational Bayes for high-dimensional proportional hazards models with applications within gene expression

2021-12-19 · Michael Komodromos, Eric Aboagye, Marina Evangelou, Sarah Filippi 외

Few Bayesian methods for analyzing high-dimensional sparse survival data provide scalable variable selection, effect estimation and uncertainty quantification. Such methods often either sacrifice uncertainty quantificati…

Uncertainty QuantificationVariable Selection

A Kernel Stein Test for Comparing Latent Variable Models

2019-07-01 · Heishiro Kanagawa, Wittawat Jitkrittum, Lester Mackey, Kenji Fukumizu 외

We propose a kernel-based nonparametric test of relative goodness of fit, where the goal is to compare two models, both of which may have unobserved latent variables, such that the marginal distribution of the observed v…