paper-with-me

홈 › Papers

Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU Networks

2025-03-08 · Devon Jarvis, Richard Klein, Benjamin Rosman, Andrew M. Saxe

In spite of finite dimension ReLU neural networks being a consistent factor behind recent deep learning successes, a theory of feature learning in these models remains elusive. Currently, insightful theories still rely on assumptions including the linearity of the network computations, unstructured input data and architectural constraints such as infinite width or a single hidden layer. To begin to address this gap we establish an equivalence between ReLU networks and Gated Deep Linear Networks, and use their greater tractability to derive dynamics of learning. We then consider multiple variants of a core task reminiscent of multi-task learning or contextual control which requires both feature learning and nonlinearity. We make explicit that, for these tasks, the ReLU networks possess an inductive bias towards latent representations which are not strictly modular or disentangled but are still highly structured and reusable between contexts. This effect is amplified with the addition of more contexts and hidden layers. Thus, we take a step towards a theory of feature learning in finite ReLU networks and shed light on how structured mixed-selective latent representations can emerge due to a bias for node-reuse and learning speed.

📄 PDF Abstract BibTeX arXiv:2503.06181

Code (0)

등록된 구현이 없습니다.

Tasks

Inductive BiasMulti-Task Learning

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

Tension Space Analysis for Emergent Narrative

2020-04-22 · Ben Kybartas, Clark Verbrugge, Jonathan Lessard

Emergent narratives provide a unique and compelling approach to interactive storytelling through simulation, and have applications in games, narrative generation, and virtual agents. However the inherent complexity of si…

A novel approach to data generation in generative model

2025-02-14 · Jaehong Kim, Jaewon Shim

Variational Autoencoders (VAEs) and other generative models are widely employed in artificial intelligence to synthesize new data. However, current approaches rely on Euclidean geometric assumptions and statistical appro…

Metric Learning

Graceful forgetting: Memory as a process

2025-02-16 · Alain de Cheveigné

A rational theory of memory is proposed to explain how we can accommodate unbounded sensory input within bounded storage space. Memory is stored as statistics, organized into complex structures that are constantly summar…

Towards Estimating Transferability using Hard Subsets

2023-01-17 · Tarun Ram Menta, Surgan Jandial, Akash Patil, Vimal KB 외

As transfer learning techniques are increasingly used to transfer knowledge from the source model to the target task, it becomes important to quantify which source models are suitable for a given target task without perf…

Transfer Learning

PRISM: Festina Lente Proactivity -- Risk-Sensitive, Uncertainty-Aware Deliberation for Proactive Agents

2026-02-02 · Yuxuan Fu, Xiaoyu Tan, Teqi Hao, Chen Zhan 외 arxiv

Proactive agents must decide not only what to say but also whether and when to intervene. Many current systems rely on brittle heuristics or indiscriminate long reasoning, which offers little control over the benefit-bur…