Optimal Kernel for Kernel-Based Modal Statistical Methods
Kernel-based modal statistical methods include mode estimation, regression, and clustering. Estimation accuracy of these methods depends on the kernel used as well as the bandwidth. We study effect of the selection of the kernel function to the estimation accuracy of these methods. In particular, we theoretically show a (multivariate) optimal kernel that minimizes its analytically-obtained asymptotic error criterion when using an optimal bandwidth, among a certain kernel class defined via the number of its sign changes.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringregressionSimilar Papers 제목 키워드 기반
Kernel Selection for Modal Linear Regression: Optimal Kernel and IRLS Algorithm
Modal linear regression (MLR) is a method for obtaining a conditional mode predictor as a linear model. We study kernel selection for MLR from two perspectives: "which kernel achieves smaller error?" and "which kernel is…
regressionNearly Optimal Clustering Risk Bounds for Kernel K-Means
In this paper, we study the statistical properties of kernel $k$-means and obtain a nearly optimal excess clustering risk bound, substantially improving the state-of-art bounds in the existing clustering risk analyses. W…
ClusteringA short note on extension theorems and their connection to universal consistency in machine learning
Statistical machine learning plays an important role in modern statistics and computer science. One main goal of statistical machine learning is to provide universally consistent algorithms, i.e., the estimator converges…
BIG-bench Machine LearningGeneralization Guarantees for Sparse Kernel Approximation with Entropic Optimal Features
Despite their success, kernel methods suffer from a massive computational cost in practice. In this paper, in lieu of commonly used kernel expansion with respect to $N$ inputs, we develop a novel optimal design maximizin…
Approximate Kernel PCA Using Random Features: Computational vs. Statistical Trade-off
Kernel methods are powerful learning methodologies that allow to perform non-linear data analysis. Despite their popularity, they suffer from poor scalability in big data scenarios. Various approximation methods, includi…