paper-with-me

홈 › Papers

Network size and weights size for memorization with two-layers neural networks

2020-06-04 · Sébastien Bubeck, Ronen Eldan, Yin Tat Lee, Dan Mikulincer

In 1988, Eric B. Baum showed that two-layers neural networks with threshold activation function can perfectly memorize the binary labels of $n$ points in general position in $\mathbb{R}^d$ using only $\ulcorner n/d \urcorner$ neurons. We observe that with ReLU networks, using four times as many neurons one can fit arbitrary real labels. Moreover, for approximate memorization up to error $\epsilon$, the neural tangent kernel can also memorize with only $O\left(\frac{n}{d} \cdot \log(1/\epsilon) \right)$ neurons (assuming that the data is well dispersed too). We show however that these constructions give rise to networks where the magnitude of the neurons' weights are far from optimal. In contrast we propose a new training procedure for ReLU networks, based on complex (as opposed to real) recombination of the neurons, for which we show approximate memorization with both $O\left(\frac{n}{d} \cdot \frac{\log(1/\epsilon)}{\epsilon}\right)$ neurons, as well as nearly-optimal size of the weights.

📄 PDF Abstract BibTeX arXiv:2006.02855

Code (0)

등록된 구현이 없습니다.

Tasks

Memorization

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

Network size and size of the weights in memorization with two-layers neural networks

2020-12-01 · NeurIPS 2020 12 · Sebastien Bubeck, Ronen Eldan, Yin Tat Lee, Dan Mikulincer

In 1988, Eric B. Baum showed that two-layers neural networks with threshold activation function can perfectly memorize the binary labels of $n$ points in general position in $\R^d$ using only $\ulcorner n/d \urcorner$ ne…

Memorization

On the geometry of generalization and memorization in deep neural networks

2021-05-30 · ICLR 2021 1 · Cory Stephenson, Suchismita Padhy, Abhinav Ganesh, Yue Hui 외

Understanding how large neural networks avoid memorizing training data is key to explaining their high generalization performance. To examine the structure of when and where memorization occurs in a deep network, we use …

Memorization

Generative Modeling of Weights: Generalization or Memorization?

2025-06-09 · Boya Zeng, Yida Yin, Zhiqiu Xu, Zhuang Liu

Generative models, with their success in image and video generation, have recently been explored for synthesizing effective neural network weights. These approaches take trained neural network checkpoints as training dat…

MemorizationVideo Generation

Can Neural Network Memorization Be Localized?

2023-07-18 · Pratyush Maini, Michael C. Mozer, Hanie Sedghi, Zachary C. Lipton 외

Recent efforts at explaining the interplay of memorization and generalization in deep overparametrized networks have posited that neural networks $\textit{memorize}$ "hard" examples in the final few layers of the model. …

Memorization

Localizing Paragraph Memorization in Language Models

2024-03-28 · Niklas Stoehr, Mitchell Gordon, Chiyuan Zhang, Owen Lewis

Can we localize the weights and mechanisms used by a language model to memorize and recite entire paragraphs of its training data? In this paper, we show that while memorization is spread across multiple layers and model…

Language ModelingLanguage ModellingMemorization