paper-with-me

Papers

Fast Two-Sample Testing with Analytic Representations of Probability Measures

2015-06-15 · NeurIPS 2015 12 · Kacper Chwialkowski, Aaditya Ramdas, Dino Sejdinovic, Arthur Gretton

We propose a class of nonparametric two-sample tests with a cost linear in the sample size. Two tests are given, both based on an ensemble of distances between analytic functions representing each of the distributions. The first test uses smoothed empirical characteristic functions to represent the distributions, the second uses distribution embeddings in a reproducing kernel Hilbert space. Analyticity implies that differences in the distributions may be detected almost surely at a finite number of randomly chosen locations/frequencies. The new tests are consistent against a larger class of alternatives than the previous linear-time tests based on the (non-smoothed) empirical characteristic functions, while being much faster than the current state-of-the-art quadratic-time kernel-based or energy distance-based tests. Experiments on artificial benchmarks and on challenging real-world testing problems demonstrate that our tests give a better power/time tradeoff than competing approaches, and in some cases, better outright power than even the most expensive quadratic-time tests. This performance advantage is retained even in high dimensions, and in cases where the difference in distributions is not observable with low order statistics.

📄 PDF Abstract BibTeX arXiv:1506.04725

Code (1)

kacperChwialkowski/analyticMeanEmbeddings

Tasks

Two-sample testingVocal Bursts Valence Prediction

Similar Papers 제목 키워드 기반

Cross-validation in high-dimensional spaces: a lifeline for least-squares models and multi-class LDA

2018-03-27 · Matthias S. Treder

Least-squares models such as linear regression and Linear Discriminant Analysis (LDA) are amongst the most popular statistical learning techniques. However, since their computation time increases cubically with the numbe…

EEGElectroencephalogram (EEG)

Understanding the Behaviour of the Empirical Cross-Entropy Beyond the Training Distribution

2019-05-28 · Matias Vera, Pablo Piantanida, Leonardo Rey Vega

Machine learning theory has mostly focused on generalization to samples from the same distribution as the training data. Whereas a better understanding of generalization beyond the training distribution where the observe…

Learning Theory

Optimal Testing of Discrete Distributions with High Probability

2020-09-14 · Ilias Diakonikolas, Themis Gouleakis, Daniel M. Kane, John Peebles 외

We study the problem of testing discrete distributions with a focus on the high probability regime. Specifically, given samples from one or more discrete distributions, a property $\mathcal{P}$, and parameters $0< \epsil…

Vocal Bursts Intensity Prediction

A Differentially Private Kernel Two-Sample Test

2018-08-01 · Anant Raj, Ho Chung Leon Law, Dino Sejdinovic, Mijung Park

Kernel two-sample testing is a useful statistical tool in determining whether data samples arise from different distributions without imposing any parametric assumptions on those distributions. However, raw data samples …

Two-sample testingVocal Bursts Valence Prediction

Conditional independence testing based on a nearest-neighbor estimator of conditional mutual information

2017-09-05 · Jakob Runge

Conditional independence testing is a fundamental problem underlying causal discovery and a particularly challenging task in the presence of nonlinear and high-dimensional dependencies. Here a fully non-parametric test f…

Causal Discovery