Contravariance Theory: Strong Alignment for Minimal Solutions to Hard Tasks
A series of results from the NeuroAI over the past fifteen years have raised core questions both about how to compare Deep Neural Network (DNN) models to the brain, and about how much convergent evolution to expect between artificial networks and real brain networks. Here, we show that for any two minimal DNN solutions to a sufficiently hard task: (i) "weak" alignment of network representations based on affine mappings guarantees "strong" alignment of privileged axes, and (ii) alignment "zippers" up the network hierarchy, causing the emergence of privileged axes from end-to-end task optimization. These results formalize the notion of contravariance from Cao and Yamins [2024], and illustrate important consequences for the theory of NeuroAI: with sufficiently strong tasks, choice of metric for inter-network comparison is not all that sensitive, and that convergent evolution is probably inevitable.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Measuring and Controlling Solution Degeneracy across Task-Trained Recurrent Neural Networks
Task-trained recurrent neural networks (RNNs) are widely used in neuroscience and machine learning to model dynamical computations. To gain mechanistic insight into how neural systems solve tasks, prior work often revers…
Subharmonic solutions for a class of predator-prey models with degenerate weights in periodic environments
This paper deals with the existence, multiplicity, minimal complexity and global structure of the subharmonic solutions to a class of planar Hamiltonian systems with periodic coefficients, being the classical predator-pr…
Why Alignment Must Precede Distillation: A Minimal Working Explanation
For efficiency, preference alignment is often performed on compact, knowledge-distilled (KD) models. We argue this common practice introduces a significant limitation by overlooking a key property of the alignment's refe…
Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models
We introduce a novel analysis that leverages linguistic minimal pairs to probe the internal linguistic representations of Large Language Models (LLMs). By measuring the similarity between LLM activation differences acros…
Semantic SimilaritySemantic Textual SimilarityHow are linear representations learned? Exact solutions to the dynamics of abstraction
In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space. In deep learning, this idea is known as the linear representation hypothesis and underpins…