paper-with-me

홈 › Papers

Relative-Based Scaling Law for Neural Language Models

2025-10-23 · Baoqing Yue, Jinyuan Zhou, Zixi Wei, Jingtao Zhan, Qingyao Ai, Yiqun Liu arxiv

Scaling laws aim to accurately predict model performance across different scales. Existing scaling-law studies almost exclusively rely on cross-entropy as the evaluation metric. However, cross-entropy provides only a partial view of performance: it measures the absolute probability assigned to the correct token, but ignores the relative ordering between correct and incorrect tokens. Yet, relative ordering is crucial for language models, such as in greedy-sampling scenario. To address this limitation, we investigate scaling from the perspective of relative ordering. We first propose the Relative-Based Probability (RBP) metric, which quantifies the probability that the correct token is ranked among the top predictions. Building on this metric, we establish the Relative-Based Scaling Law, which characterizes how RBP improves with increasing model size. Through extensive experiments on four datasets and four model families spanning five orders of magnitude, we demonstrate the robustness and accuracy of this law. Finally, we illustrate the broad application of this law with two examples, namely providing a deeper explanation of emergence phenomena and facilitating finding fundamental theories of scaling laws. In summary, the Relative-Based Scaling Law complements the cross-entropy perspective and contributes to a more complete understanding of scaling large language models. Thus, it offers valuable insights for both practical development and theoretical exploration.

📄 PDF Abstract BibTeX arXiv:2510.20387

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scaling Recurrent Neural Network Language Models

2015-02-02 · Will Williams, Niranjani Prasad, David Mrva, Tom Ash 외

This paper investigates the scaling properties of Recurrent Neural Network Language Models (RNNLMs). We discuss how to train very large RNNs on GPUs and address the questions of how RNNLMs scale with respect to model siz…

Language ModellingMachine TranslationTranslation

Relative Scaling Laws for LLMs

2025-10-28 · William Held, David Hall, Percy Liang, Diyi Yang arxiv

Scaling laws describe how language models improve with additional data, parameters, and compute. While widely used, they are typically measured on aggregate test sets. Aggregate evaluations yield clean trends but average…

Algorithmic progress in language models

2024-03-09 · Anson Ho, Tamay Besiroglu, Ege Erdil, David Owen 외

We investigate the rate at which algorithms for pre-training language models have improved since the advent of deep learning. Using a dataset of over 200 language model evaluations on Wikitext and Penn Treebank spanning …

Language ModelingLanguage Modelling

Multilingual Test-Time Scaling via Initial Thought Transfer

2025-05-21 · Prasoon Bajpai, Tanmoy Chakraborty

Test-time scaling has emerged as a widely adopted inference-time strategy for boosting reasoning performance. However, its effectiveness has been studied almost exclusively in English, leaving its behavior in other langu…

A decentralized route to the origins of scaling in human language

2017-05-16 · Felipe Urbina, Javier Vera

The Zipf's law establishes that if the words of a (large) text are ordered by decreasing frequency, the frequency versus the rank decreases as a power law with exponent close to $-1$. Previous work has stressed that this…