paper-with-me

홈 › Papers

A Mean Field Ansatz for Zero-Shot Weight Transfer

2024-08-16 · Xingyuan Chen, Wenwei Kuang, Lei Deng, Wei Han, Bo Bai, Goncalo dos Reis

The pre-training cost of large language models (LLMs) is prohibitive. One cutting-edge approach to reduce the cost is zero-shot weight transfer, also known as model growth for some cases, which magically transfers the weights trained in a small model to a large model. However, there are still some theoretical mysteries behind the weight transfer. In this paper, inspired by prior applications of mean field theory to neural network dynamics, we introduce a mean field ansatz to provide a theoretical explanation for weight transfer. Specifically, we propose the row-column (RC) ansatz under the mean field point of view, which describes the measure structure of the weights in the neural network (NN) and admits a close measure dynamic. Thus, the weights of different sizes NN admit a common distribution under proper assumptions, and weight transfer methods can be viewed as sampling methods. We empirically validate the RC ansatz by exploring simple MLP examples and LLMs such as GPT-3 and Llama-3.1. We show the mean-field point of view is adequate under suitable assumptions which can provide theoretical support for zero-shot weight transfer.

📄 PDF Abstract BibTeX arXiv:2408.08681

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Weight Decay 설명 없음
Adam 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음

Similar Papers 제목 키워드 기반

Phase transformation and synchrony for a network of coupled Izhikevich neurons

2024-07-29 · Áine Byrne

A number of recent articles have employed the Lorentz ansatz to reduce a network of Izhikevich neurons to a tractable mean-field description. In this letter, we construct an equivalent phase model for the Izhikevich mode…

Articlesvalid

Convergence to the fixed-node limit in deep variational Monte Carlo

2020-10-11 · Zeno Schätzle, Jan Hermann, Frank Noé

Variational quantum Monte Carlo (QMC) is an ab-initio method for solving the electronic Schr\"odinger equation that is exact in principle, but limited by the flexibility of the available ansatzes in practice. The recentl…

Variational Monte Carlo

CICA: Content-Injected Contrastive Alignment for Zero-Shot Document Image Classification

2024-05-06 · Sankalp Sinha, Muhammad Saif Ullah Khan, Talha Uddin Sheikh, Didier Stricker 외

Zero-shot learning has been extensively investigated in the broader field of visual recognition, attracting significant interest recently. However, the current work on zero-shot learning in document image classification …

Document Classificationdocument-image-classificationDocument Image ClassificationGeneralized Zero-Shot Learning+3

Learning Enhanced Ensemble Filters

2025-04-24 · Eviatar Bach, Ricardo Baptista, Edoardo Calvello, Bohan Chen 외

The filtering distribution in hidden Markov models evolves according to the law of a mean-field model in state-observation space. The ensemble Kalman filter (EnKF) approximates this mean-field model with an ensemble of i…

A multiple timescales approach to bridging spiking- and population-level dynamics

2018-08-15

A rigorous bridge between spiking-level and macroscopic quantities is an on-going and well-developed story for asynchronously firing neurons, but focus has shifted to include neural populations exhibiting varying synchro…