paper-with-me

홈 › Papers

Neural Network Layer Matrix Decomposition reveals Latent Manifold Encoding and Memory Capacity

2023-09-12 · Ng Shyh-Chang, A-Li Luo, Bo Qiu

We prove the converse of the universal approximation theorem, i.e. a neural network (NN) encoding theorem which shows that for every stably converged NN of continuous activation functions, its weight matrix actually encodes a continuous function that approximates its training dataset to within a finite margin of error over a bounded domain. We further show that using the Eckart-Young theorem for truncated singular value decomposition of the weight matrix for every NN layer, we can illuminate the nature of the latent space manifold of the training dataset encoded and represented by every NN layer, and the geometric nature of the mathematical operations performed by each NN layer. Our results have implications for understanding how NNs break the curse of dimensionality by harnessing memory capacity for expressivity, and that the two are complementary. This Layer Matrix Decomposition (LMD) further suggests a close relationship between eigen-decomposition of NN layers and the latest advances in conceptualizations of Hopfield networks and Transformer NN models.

📄 PDF Abstract BibTeX arXiv:2309.05968

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Adam 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

SOFARI: High-Dimensional Manifold-Based Inference

2023-09-26 · Zemin Zheng, Xin Zhou, Yingying Fan, Jinchi Lv

Multi-task learning is a widely used technique for harnessing information from various tasks. Recently, the sparse orthogonal factor regression (SOFAR) framework, based on the sparse singular value decomposition (SVD) wi…

Multi-Task Learning

An extrapolated and provably convergent algorithm for nonlinear matrix decomposition with the ReLU function

2025-03-31 · Nicolas Gillis, Margherita Porcelli, Giovanni Seraghiti

Nonlinear matrix decomposition (NMD) with the ReLU function, denoted ReLU-NMD, is the following problem: given a sparse, nonnegative matrix $X$ and a factorization rank $r$, identify a rank-$r$ matrix $\Theta$ such that …

Data CompressionMathMatrix Completion

SOFARI-R: High-Dimensional Manifold-Based Inference for Latent Responses

2025-04-24 · Zemin Zheng, Xin Zhou, Jinchi Lv

Data reduction with uncertainty quantification plays a key role in various multi-task learning applications, where large numbers of responses and features are present. To this end, a general framework of high-dimensional…

Multi-Task LearningUncertainty Quantification

Latent Semantic Manifolds in Large Language Models

2026-03-17 · Mohamed A. Mabrok arxiv

Large Language Models (LLMs) perform internal computations in continuous vector spaces yet produce discrete tokens -- a fundamental mismatch whose geometric consequences remain poorly understood. We develop a mathematica…

Model Compression

On the Connection Between Non-negative Matrix Factorization and Latent Dirichlet Allocation

2024-05-30 · Benedikt Geiger, Peter J. Park

Non-negative matrix factorization with the generalized Kullback-Leibler divergence (NMF) and latent Dirichlet allocation (LDA) are two popular approaches for dimensionality reduction of non-negative data. Here, we show t…

Dimensionality Reduction