paper-with-me

홈 › Papers

Diffusion Models Learn Low-Dimensional Distributions via Subspace Clustering

2024-09-04 · Peng Wang, Huijie Zhang, Zekai Zhang, Siyi Chen, Yi Ma, Qing Qu

Recent empirical studies have demonstrated that diffusion models can effectively learn the image distribution and generate new samples. Remarkably, these models can achieve this even with a small number of training samples despite a large image dimension, circumventing the curse of dimensionality. In this work, we provide theoretical insights into this phenomenon by leveraging key empirical observations: (i) the low intrinsic dimensionality of image data, (ii) a union of manifold structure of image data, and (iii) the low-rank property of the denoising autoencoder in trained diffusion models. These observations motivate us to assume the underlying data distribution of image data as a mixture of low-rank Gaussians and to parameterize the denoising autoencoder as a low-rank model according to the score function of the assumed distribution. With these setups, we rigorously show that optimizing the training loss of diffusion models is equivalent to solving the canonical subspace clustering problem over the training samples. Based on this equivalence, we further show that the minimal number of samples required to learn the underlying distribution scales linearly with the intrinsic dimensions under the above data and model assumptions. This insight sheds light on why diffusion models can break the curse of dimensionality and exhibit the phase transition in learning distributions. Moreover, we empirically establish a correspondence between the subspaces and the semantic representations of image data, facilitating image editing. We validate these results with corroborated experimental results on both simulated distributions and image datasets.

📄 PDF Abstract BibTeX arXiv:2409.02426

Code (1)

huijieZH/Diffusion-Model-Generalizability 공식 구현 pytorch

Tasks

ClusteringDenoising

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Denoising Autoencoder A Denoising Autoencoder is a modification on the autoencoder to prevent the network learning the identity function.…

Similar Papers 제목 키워드 기반

Sparse Subspace Clustering via Diffusion Process

2016-08-05 · Qilin Li, Ling Li, Wanquan Liu

Subspace clustering refers to the problem of clustering high-dimensional data that lie in a union of low-dimensional subspaces. State-of-the-art subspace clustering methods are based on the idea of expressing each data p…

Clustering

Diffusion Models Are Statistically Optimal for Learning Low-Dimensional Multi-Modal Distributions

2026-05-28 · Jingda Wu, Changxiao Cai arxiv

Score-based diffusion models have demonstrated remarkable empirical success in learning high-dimensional distributions, particularly those exhibiting low-dimensional and multi-modal structures. However, theoretical under…

Subspace Clustering by Mixture of Gaussian Regression

2015-06-01 · CVPR 2015 6 · Baohua Li, Ying Zhang, Zhouchen Lin, Huchuan Lu

Subspace clustering is a problem of finding a multisubspace representation that best fits sample points drawn from a high-dimensional space. The existing clustering models generally adopt different norms to describe nois…

Clusteringregression

Learning Robust Subspace Clustering

2013-08-01 · Qiang Qiu, Guillermo Sapiro

We propose a low-rank transformation-learning framework to robustify subspace clustering. Many high-dimensional data, such as face images and motion sequences, lie in a union of low-dimensional subspaces. The subspace cl…

Clustering

Fast Adaptive K-Means Subspace Clustering for High-Dimensional Data

2019-03-01 · IEEE Access 2019 3 · Xiaodong Wang; Rungching Chen; Fei Yan; Zhiqiang Zeng; Chaoqun Hong

In many real-world applications, data are represented by high-dimensional features. Despite the simplicity, existing K-means subspace clustering algorithms often employ eigenvalue decomposition to generate an approximate…

Clusteringfeature selectionVocal Bursts Intensity Prediction