paper-with-me

홈 › Papers

Revisiting Neural Scaling Laws in Language and Vision

2022-09-13 · Ibrahim Alabdulmohsin, Behnam Neyshabur, Xiaohua Zhai

The remarkable progress in deep learning in recent years is largely driven by improvements in scale, where bigger models are trained on larger datasets for longer schedules. To predict the benefit of scale empirically, we argue for a more rigorous methodology based on the extrapolation loss, instead of reporting the best-fitting (interpolating) parameters. We then present a recipe for estimating scaling law parameters reliably from learning curves. We demonstrate that it extrapolates more accurately than previous methods in a wide range of architecture families across several domains, including image classification, neural machine translation (NMT) and language modeling, in addition to tasks from the BIG-Bench evaluation benchmark. Finally, we release a benchmark dataset comprising of 90 evaluation tasks to facilitate research in this domain.

📄 PDF Abstract BibTeX arXiv:2209.06640

Code (1)

google-research/google-research 공식 구현 tf

Tasks

image-classificationImage ClassificationLanguage ModelingLanguage ModellingMachine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

Reproducible scaling laws for contrastive language-image learning

2022-12-14 · CVPR 2023 1 · Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman 외

Scaling up neural networks has led to remarkable performance across a wide range of tasks. Moreover, performance often follows reliable scaling laws as a function of training set size, model size, and compute, which offe…

Image ClassificationOpen Vocabulary Attribute DetectionRetrievalzero-shot-classification+3

Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets

2025-06-05 · Marianna Nezhurina, Tomer Porian, Giovanni Pucceti, Tommie Kerssies 외

In studies of transferable learning, scaling laws are obtained for various important foundation models to predict their properties and performance at larger scales. We show here how scaling law derivation can also be use…

On the Invariance and Generality of Neural Scaling Laws

2026-05-08 · Xing Han, Ziyin Liu, Suchi Saria, Paul Pu Liang arxiv

Neural scaling laws establish a predictable relationship between model performance and data or compute, offering crucial guidance for resource allocation in new domains and tasks. Yet such laws are most needed precisely …

Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training

2025-12-09 · Jakub Krajewski, Amitis Shidani, Dan Busbridge, Sam Wiseman 외 arxiv

While scaling laws for Large Language Models (LLMs) traditionally focus on proxy metrics like pretraining loss, predicting downstream task performance has been considered unreliable. This paper challenges that view by pr…

Scaling laws in wearable human activity recognition

2025-02-05 · Tom Hoddes, Alex Bijamov, Saket Joshi, Daniel Roggen 외

Many deep architectures and self-supervised pre-training techniques have been proposed for human activity recognition (HAR) from wearable multimodal sensors. Scaling laws have the potential to help move towards more prin…

Activity RecognitionHuman Activity Recognition