paper-with-me

홈 › Papers

Matching domain experts by training from scratch on domain knowledge

2024-05-15 · Xiaoliang Luo, Guangzhi Sun, Bradley C. Love

Recently, large language models (LLMs) have outperformed human experts in predicting the results of neuroscience experiments (Luo et al., 2024). What is the basis for this performance? One possibility is that statistical patterns in that specific scientific literature, as opposed to emergent reasoning abilities arising from broader training, underlie LLMs' performance. To evaluate this possibility, we trained (next word prediction) a relatively small 124M-parameter GPT-2 model on 1.3 billion tokens of domain-specific knowledge. Despite being orders of magnitude smaller than larger LLMs trained on trillions of tokens, small models achieved expert-level performance in predicting neuroscience results. Small models trained on the neuroscience literature succeeded when they were trained from scratch using a tokenizer specifically trained on neuroscience text or when the neuroscience literature was used to finetune a pretrained GPT-2. Our results indicate that expert-level performance may be attained by even small LLMs through domain-specific, auto-regressive training approaches.

📄 PDF Abstract BibTeX arXiv:2405.09395

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Discriminative Fine-Tuning Discriminative Fine-Tuning is a fine-tuning strategy that is used for ULMFiT type models. Instead of using the same learning rate…
Multi-Head Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts

2026-04-20 · Jacob Morrison, Sanjay Adhikesaven, Akshita Bhagia, Matei Zaharia 외 arxiv

Extending a fully post-trained language model with new domain capabilities is fundamentally limited by monolithic training paradigms: retraining from scratch is expensive and scales poorly, while continued training often…

Reinforcement Learning

Domain Attention with an Ensemble of Experts

2017-07-01 · ACL 2017 7 · Young-Bum Kim, Karl Stratos, Dongchan Kim

An important problem in domain adaptation is to quickly generalize to a new domain with limited supervision given K existing domains. One approach is to retrain a global model across all K + 1 domains using standard tech…

Domain AdaptationSpoken Language Understanding

Expert Divergence Learning for MoE-based Language Models

2026-02-10 · Jiaang Li, Haibin Chen, Langming Liu, Yujin Yuan 외 arxiv

The Mixture-of-Experts (MoE) architecture is a powerful technique for scaling language models, yet it often suffers from expert homogenization, where experts learn redundant functionalities, thereby limiting MoE's full p…

Exploring Domain Robust Lightweight Reward Models based on Router Mechanism

2024-07-24 · Hyuk Namgoong, Jeesu Jung, SangKeun Jung, YoonHyung Roh

Recent advancements in large language models have heavily relied on the large reward model from reinforcement learning from human feedback for fine-tuning. However, the use of a single reward model across various domains…

Language ModelingLanguage ModellingMixture-of-ExpertsSmall Language Model

Automatic Creativity Measurement in Scratch Programs Across Modalities

2022-11-07 · Anastasia Kovalkov, Benjamin Paaßen, Avi Segal, Niels Pinkwart 외

Promoting creativity is considered an important goal of education, but creativity is notoriously hard to measure.In this paper, we make the journey fromdefining a formal measure of creativity that is efficientlycomputabl…