paper-with-me

홈 › Papers

Towards a Theoretical Analysis of PCA for Heteroscedastic Data

2016-10-12 · David Hong, Laura Balzano, Jeffrey A. Fessler

Principal Component Analysis (PCA) is a method for estimating a subspace given noisy samples. It is useful in a variety of problems ranging from dimensionality reduction to anomaly detection and the visualization of high dimensional data. PCA performs well in the presence of moderate noise and even with missing data, but is also sensitive to outliers. PCA is also known to have a phase transition when noise is independent and identically distributed; recovery of the subspace sharply declines at a threshold noise variance. Effective use of PCA requires a rigorous understanding of these behaviors. This paper provides a step towards an analysis of PCA for samples with heteroscedastic noise, that is, samples that have non-uniform noise variances and so are no longer identically distributed. In particular, we provide a simple asymptotic prediction of the recovery of a one-dimensional subspace from noisy heteroscedastic samples. The prediction enables: a) easy and efficient calculation of the asymptotic performance, and b) qualitative reasoning to understand how PCA is impacted by heteroscedasticity (such as outliers).

📄 PDF Abstract BibTeX arXiv:1610.03595

Code (0)

등록된 구현이 없습니다.

Tasks

Anomaly DetectionDimensionality Reduction

Methods 이 논문이 사용한 방법론

PCA Principle Components Analysis (PCA) is an unsupervised method primary used for dimensionality reduction within machine learning. PCA is calculated via a singular value…

Similar Papers 제목 키워드 기반

Theoretical Analysis of Heteroscedastic Gaussian Processes with Posterior Distributions

2024-09-19 · Yuji Ito

This study introduces a novel theoretical framework for analyzing heteroscedastic Gaussian processes (HGPs) that identify unknown systems in a data-driven manner. Although HGPs effectively address the heteroscedasticity …

Gaussian Processes

Active Heteroscedastic Regression

2017-08-01 · ICML 2017 8 · Kamalika Chaudhuri, Prateek Jain, Nagarajan Natarajan

An active learner is given a model class $\Theta$, a large sample of unlabeled data drawn from an underlying distribution and access to a labeling oracle that can provide a label for any of the unlabeled instances. …

Active LearningBinary Classificationregression

Matrix Denoising with Doubly Heteroscedastic Noise: Fundamental Limits and Optimal Spectral Methods

2024-05-22 · Yihan Zhang, Marco Mondelli

We study the matrix denoising problem of estimating the singular vectors of a rank-$1$ signal corrupted by noise with both column and row correlations. Existing works are either unable to pinpoint the exact asymptotic es…

Denoising

HeMPPCAT: Mixtures of Probabilistic Principal Component Analysers for Data with Heteroscedastic Noise

2023-01-21 · Alec S. Xu, Laura Balzano, Jeffrey A. Fessler

Mixtures of probabilistic principal component analysis (MPPCA) is a well-known mixture model extension of principal component analysis (PCA). Similar to PCA, MPPCA assumes the data samples in each mixture contain homosce…

Clustering

A Skewness-Based Criterion for Addressing Heteroscedastic Noise in Causal Discovery

2024-10-08 · Yingyu Lin, Yuxing Huang, Wenqin Liu, Haoran Deng 외

Real-world data often violates the equal-variance assumption (homoscedasticity), making it essential to account for heteroscedastic noise in causal discovery. In this work, we explore heteroscedastic symmetric noise mode…

Causal Discovery