paper-with-me

홈 › Papers

Mean Field Models for Neural Networks in Teacher-student Setting

2019-09-25 · Lexing Ying, Yuandong Tian

Mean field models have provided a convenient framework for understanding the training dynamics for certain neural networks in the infinite width limit. The resulting mean field equation characterizes the evolution of the time-dependent empirical distribution of the network parameters. Following this line of work, this paper first focuses on the teacher-student setting. For the two-layer networks, we derive the necessary condition of the stationary distributions of the mean field equation and explain an empirical phenomenon concerning training speed differences using the Wasserstein flow description. Second, we apply this approach to two extended ResNet models and characterize the necessary condition of stationary distributions in the teacher-student setting.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A new role for circuit expansion for learning in neural networks

2020-08-19 · Julia Steinberg, Madhu Advani, Haim Sompolinsky

Many sensory pathways in the brain rely on sparsely active populations of neurons downstream from the input stimuli. The biological reason for the occurrence of expanded structure in the brain is unclear, but may be beca…

Dense Hopfield Networks in the Teacher-Student Setting

2024-01-08 · Robin Thériault, Daniele Tantari

Dense Hopfield networks are known for their feature to prototype transition and adversarial robustness. However, previous theoretical studies have been mostly concerned with their storage capacity. We bridge this gap by …

Adversarial Robustness

Analysis of Random Sequential Message Passing Algorithms for Approximate Inference

2022-02-16 · Burak Çakmak, Yue M. Lu, Manfred Opper

We analyze the dynamics of a random sequential message passing algorithm for approximate inference with large Gaussian latent variable models in a student-teacher scenario. To model nontrivial dependencies between the la…

Better Teacher Better Student: Dynamic Prior Knowledge for Knowledge Distillation

2022-06-13 · Zengyu Qiu, Xinzhu Ma, Kunlin Yang, Chunya Liu 외

Knowledge distillation (KD) has shown very promising capabilities in transferring learning representations from large models (teachers) to small models (students). However, as the capacity gap between students and teache…

image-classificationImage ClassificationKnowledge DistillationModel Selection+2

Knowledge distillation via adaptive instance normalization

2020-03-09 · Jing Yang, Brais Martinez, Adrian Bulat, Georgios Tzimiropoulos

This paper addresses the problem of model compression via knowledge distillation. To this end, we propose a new knowledge distillation method based on transferring feature statistics, specifically the channel-wise mean a…

Knowledge DistillationModel Compression