paper-with-me

Pythia

2000년 도입 · 논문 60편에서 사용

Pythia is a suite of decoder-only autoregressive language models all trained on public data seen in the exact same order and ranging in size from 70M to 12B parameters. The model architecture and hyperparameters largely follow GPT-3, with a few notable deviations based on recent advances in best practices for large scale language modeling.

출처: Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

소개 논문: Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Language Models · Natural Language Processing