paper-with-me

Papers

Uncovering Scaling Laws for Large Language Models via Inverse Problems

2025-09-09 · Arun Verma, Zhaoxuan Wu, Zijian Zhou, Xiaoqiang Lin, Zhiliang Chen, Rachael Hwee Ling Sim, Rui Qiao, Jingtan Wang, Nhung Bui, Xinyuan Niu, Wenyang Hu, Gregory Kang Ruey Lau, Zi-Yu Khoo, Zitong Zhao, Xinyi Xu, Apivich Hemachandra, See-Kiong Ng, Bryan Kian Hsiang Low arxiv

Large Language Models (LLMs) are large-scale pretrained models that have achieved remarkable success across diverse domains. These successes have been driven by unprecedented complexity and scale in both data and computations. However, due to the high costs of training such models, brute-force trial-and-error approaches to improve LLMs are not feasible. Inspired by the success of inverse problems in uncovering fundamental scientific laws, this position paper advocates that inverse problems can also efficiently uncover scaling laws that guide the building of LLMs to achieve the desirable performance with significantly better cost-effectiveness.

📄 PDF Abstract BibTeX arXiv:2509.07909

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AlphaZero Neural Scaling and Zipf's Law: a Tale of Board Games and Power Laws

2024-12-16 · Oren Neumann, Claudius Gros

Neural scaling laws are observed in a range of domains, to date with no clear understanding of why they occur. Recent theories suggest that loss power laws arise from Zipf's law, a power law observed in domains like natu…

Board GamesLanguage Modelling

Inverse Scaling: When Bigger Isn't Better

2023-06-15 · Ian R. McKenzie, Alexander Lyzhov, Michael Pieler, Alicia Parrish 외

Work on scaling laws has found that large language models (LMs) show predictable improvements to overall loss with increased scale (model size, training data, and compute). Here, we present evidence for the claim that LM…

Uncovering Neural Scaling Laws in Molecular Representation Learning

2023-09-15 · NeurIPS 2023 11 · Dingshuo Chen, Yanqiao Zhu, Jieyu Zhang, Yuanqi Du 외

Molecular Representation Learning (MRL) has emerged as a powerful tool for drug and materials discovery in a variety of tasks such as virtual screening and inverse design. While there has been a surge of interest in adva…

molecular representationRepresentation Learning

Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients

2026-06-23 · Yizhou Liu, Jeff Gore arxiv

Neural scaling laws describe how pre-training loss decays as power laws with training time, model size, and compute. This position paper argues that the exponents of these power laws are fixed by generic mechanisms: a on…

ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws

2024-08-15 · Ruihang Li, Yixuan Wei, Miaosen Zhang, Nenghai Yu 외

High-quality data is crucial for the pre-training performance of large language models. Unfortunately, existing quality filtering methods rely on a known high-quality dataset as reference, which can introduce potential b…

Diversity