paper-with-me

홈 › Papers

Dissecting a Small Artificial Neural Network

2025-01-03 · Xiguang Yang, Krish Arora, Michael Bachmann

We investigate the loss landscape and backpropagation dynamics of convergence for the simplest possible artificial neural network representing the logical exclusive-OR (XOR) gate. Cross-sections of the loss landscape in the nine-dimensional parameter space are found to exhibit distinct features, which help understand why backpropagation efficiently achieves convergence toward zero loss, whereas values of weights and biases keep drifting. Differences in shapes of cross-sections obtained by nonrandomized and randomized batches are discussed. In reference to statistical physics we introduce the microcanonical entropy as a unique quantity that allows to characterize the phase behavior of the network. Learning in neural networks can thus be thought of as an annealing process that experiences the analogue of phase transitions known from thermodynamic systems. It also reveals how the loss landscape simplifies as more hidden neurons are added to the network, eliminating entropic barriers caused by finite-size effects.

📄 PDF Abstract BibTeX arXiv:2501.08341

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dissecting the Practical Lexical Function Model for Compositional Distributional Semantics

2015-06-01 · SEMEVAL 2015 6 · Abhijeet Gupta, Jason Utt, Sebastian Pad{\'o}
Machine TranslationSentiment Analysis

MedSAE: Dissecting MedCLIP Representations with Sparse Autoencoders

2025-10-30 · Riccardo Renzulli, Colas Lepoutre, Enrico Cassano, Marco Grangetto arxiv

Artificial intelligence in healthcare requires models that are accurate and interpretable. We advance mechanistic interpretability in medical vision by applying Medical Sparse Autoencoders (MedSAEs) to the latent space o…

Dissecting Hessian: Understanding Common Structure of Hessian in Neural Networks

2020-10-08 · Yikai Wu, Xingyu Zhu, Chenwei Wu, Annie Wang 외

Hessian captures important properties of the deep neural network loss landscape. Previous works have observed low rank structure in the Hessians of neural networks. In this paper, we propose a decoupling conjecture that …

Generalization Bounds

Dissecting Catastrophic Forgetting in Continual Learning by Deep Visualization

2020-01-06 · Giang Nguyen, Shuan Chen, Thao Do, Tae Joon Jun 외

Interpreting the behaviors of Deep Neural Networks (usually considered as a black box) is critical especially when they are now being widely adopted over diverse aspects of human life. Taking the advancements from Explai…

Continual Learning

End-to-End Reverse Screening Identifies Protein Targets of Small Molecules Using HelixFold3

2026-01-20 · Shengjie Xu, Xianbin Ye, Mengran Zhu, Xiaonan Zhang 외 arxiv

Identifying protein targets for small molecules, or reverse screening, is essential for understanding drug action, guiding compound repurposing, predicting off-target effects, and elucidating the molecular mechanisms of …

Drug Discovery