paper-with-me

홈 › Papers

On the Origins of the Block Structure Phenomenon in Neural Network Representations

2022-02-15 · Thao Nguyen, Maithra Raghu, Simon Kornblith

Recent work has uncovered a striking phenomenon in large-capacity neural networks: they contain blocks of contiguous hidden layers with highly similar representations. This block structure has two seemingly contradictory properties: on the one hand, its constituent layers exhibit highly similar dominant first principal components (PCs), but on the other hand, their representations, and their common first PC, are highly dissimilar across different random seeds. Our work seeks to reconcile these discrepant properties by investigating the origin of the block structure in relation to the data and training methods. By analyzing properties of the dominant PCs, we find that the block structure arises from dominant datapoints - a small group of examples that share similar image statistics (e.g. background color). However, the set of dominant datapoints, and the precise shared image statistic, can vary across random seeds. Thus, the block structure reflects meaningful dataset statistics, but is simultaneously unique to each model. Through studying hidden layer activations and creating synthetic datapoints, we demonstrate that these simple image statistics dominate the representational geometry of the layers inside the block structure. We explore how the phenomenon evolves through training, finding that the block structure takes shape early in training, but the underlying representations and the corresponding dominant datapoints continue to change substantially. Finally, we study the interplay between the block structure and different training mechanisms, introducing a targeted intervention to eliminate the block structure, as well as examining the effects of pretraining and Shake-Shake regularization.

📄 PDF Abstract BibTeX arXiv:2202.07184

Code (1)

google-research/google-research/tree/master/do_wide_and_deep_networks_learn_the_same_things jax

Similar Papers 제목 키워드 기반

Dominant Datapoints and the Block Structure Phenomenon in Neural Network Hidden Representations

2021-09-29 · Thao Nguyen, Maithra Raghu, Simon Kornblith

Recent work has uncovered a striking phenomenon in large-capacity neural networks: they contain blocks of contiguous hidden layers with highly similar representations. This block structure has two seemingly contradictory…

"Big Data" and its Origins

2020-08-13 · Francis X. Diebold

Against the background of explosive growth in data volume, velocity, and variety, I investigate the origins of the term "Big Data". Its origins are a bit murky and hence intriguing, involving both academics and industry,…

BIG-bench Machine Learning

Origins of Life: A Problem for Physics

2017-05-23

The origins of life stands among the great open scientific questions of our time. While a number of proposals exist for possible starting points in the pathway from non-living to living matter, these have so far not achi…

On the Origins of Linear Representations in Large Language Models

2024-03-06 · Yibo Jiang, Goutham Rajendran, Pradeep Ravikumar, Bryon Aragam 외

Recent works have argued that high-level semantic concepts are encoded "linearly" in the representation space of large language models. In this work, we study the origins of such linear representations. To that end, we i…

Language ModelingLanguage ModellingLarge Language Model

From Tables to Signals: Revealing Spectral Adaptivity in TabPFN

2025-11-23 · Jianqiao Zheng, Cameron Gordon, Yiping Ji, Hemanth Saratchandran 외 arxiv

Task-agnostic tabular foundation models such as TabPFN have achieved impressive performance on tabular learning tasks, yet the origins of their inductive biases remain poorly understood. In this work, we study TabPFN thr…

Image Denoising