paper-with-me

홈 › Papers

Training Guarantees of Neural Network Classification Two-Sample Tests by Kernel Analysis

2024-07-05 · Varun Khurana, Xiuyuan Cheng, Alexander Cloninger

We construct and analyze a neural network two-sample test to determine whether two datasets came from the same distribution (null hypothesis) or not (alternative hypothesis). We perform time-analysis on a neural tangent kernel (NTK) two-sample test. In particular, we derive the theoretical minimum training time needed to ensure the NTK two-sample test detects a deviation-level between the datasets. Similarly, we derive the theoretical maximum training time before the NTK two-sample test detects a deviation-level. By approximating the neural network dynamics with the NTK dynamics, we extend this time-analysis to the realistic neural network two-sample test generated from time-varying training dynamics and finite training samples. A similar extension is done for the neural network two-sample test generated from time-varying training dynamics but trained on the population. To give statistical guarantees, we show that the statistical power associated with the neural network two-sample test goes to 1 as the neural network training samples and test evaluation samples go to infinity. Additionally, we prove that the training times needed to detect the same deviation-level in the null and alternative hypothesis scenarios are well-separated. Finally, we run some experiments showcasing a two-layer neural network two-sample test on a hard two-sample test problem and plot a heatmap of the statistical power of the two-sample test in relation to training time and network complexity.

📄 PDF Abstract BibTeX arXiv:2407.04806

Code (1)

varunkhuran/NTK_Logit 공식 구현

Methods 이 논문이 사용한 방법론

Heatmap 설명 없음
NTK 설명 없음

Similar Papers 제목 키워드 기반

Minimax-Optimal Two-Sample Test with Sliced Wasserstein

2025-10-31 · Binh Thuan Tran, Nicolas Schreuder arxiv

We study the problem of nonparametric two-sample testing using the sliced Wasserstein (SW) distance. While prior theoretical and empirical work indicates that the SW distance offers a promising balance between strong sta…

Computational EfficiencyTwo-sample testing

Compress Then Test: Powerful Kernel Testing in Near-linear Time

2023-01-14 · Carles Domingo-Enrich, Raaz Dwivedi, Lester Mackey

Kernel two-sample testing provides a powerful framework for distinguishing any pair of distributions based on $n$ sample points. However, existing kernel tests either run in $n^2$ time or sacrifice undue power to improve…

Two-sample testing

Variable Selection for Kernel Two-Sample Tests

2023-02-15 · Jie Wang, Santanu S. Dey, Yao Xie

We consider the variable selection problem for two-sample tests, aiming to select the most informative variables to determine whether two collections of samples follow the same distribution. To address this, we propose a…

Variable SelectionVocal Bursts Valence Prediction

MMD Aggregated Two-Sample Test

2021-10-28 · NeurIPS 2023 11 · Antonin Schrab, Ilmun Kim, Mélisande Albert, Béatrice Laurent 외

We propose two novel nonparametric two-sample kernel tests based on the Maximum Mean Discrepancy (MMD). First, for a fixed kernel, we construct an MMD test using either permutations or a wild bootstrap, two popular numer…

TranslationTwo-sample testingVocal Bursts Valence Prediction

Exact Distribution-Free Hypothesis Tests for the Regression Function of Binary Classification via Conditional Kernel Mean Embeddings

2021-03-08 · Ambrus Tamás, Balázs Csanád Csáji

In this paper we suggest two statistical hypothesis tests for the regression function of binary classification based on conditional kernel mean embeddings. The regression function is a fundamental object in classificatio…

Binary ClassificationClassificationGeneral Classificationregression