paper-with-me

홈 › Papers

A Scaling Law for Token Efficiency in LLM Fine-Tuning Under Fixed Compute Budgets

2025-05-09 · Ryan Lagasse, Aidan Kiernans, Avijit Ghosh, Shiri Dori-Hacohen

We introduce a scaling law for fine-tuning large language models (LLMs) under fixed compute budgets that explicitly accounts for data composition. Conventional approaches measure training data solely by total tokens, yet the number of examples and their average token length -- what we term \emph{dataset volume} -- play a decisive role in model performance. Our formulation is tuned following established procedures. Experiments on the BRICC dataset \cite{salavati2024reducing} and subsets of the MMLU dataset \cite{hendrycks2021measuringmassivemultitasklanguage}, evaluated under multiple subsampling strategies, reveal that data composition significantly affects token efficiency. These results motivate refined scaling laws for practical LLM fine-tuning in resource-constrained settings.

📄 PDF Abstract BibTeX arXiv:2505.06150

Code (0)

등록된 구현이 없습니다.

Tasks

MMLU

Similar Papers 제목 키워드 기반

Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets

2025-12-16 · Estelle Zheng, Nathan Cerisara, Sébastien Warichet, Emmanuel Helbert 외 arxiv

Fine-tuning large language models (LLMs) is often limited by the memory available on commodity GPUs. Parameter-efficient fine-tuning (PEFT) methods such as QLoRA reduce the number of trainable parameters, yet still incur…

parameter-efficient fine-tuningNatural Language Understanding

Inference Compute-Optimal Video Vision Language Models

2025-05-24 · Peiqi Wang, Shengyun Peng, Xuewen Zhang, Hanchao Yu 외

This work investigates the optimal allocation of inference compute across three key scaling factors in video vision language models: language model size, frame count, and the number of visual tokens per frame. While prio…

Language ModelingLanguage Modelling

Single layer tiny Co$^4$ outpaces GPT-2 and GPT-BERT

2025-10-09 · Noor Ul Zain, Mohsin Raza, Ahsan Adeel arxiv

We show that a tiny Co$^4$ machine(Adeel,2025) with a single layer, two heads, and 8M parameters, operating at an approximate cost of $O(N)$ (where $N$ is the number of input tokens), outpaces the BabyLM Challenge baseli…

Training Large Language Models To Reason In Parallel With Global Forking Tokens

2025-10-01 · Sheng Jia, Xiao Wang, Shiva Prasad Kasiviswanathan arxiv

Although LLMs have demonstrated improved performance by scaling parallel test-time compute, doing so relies on generating reasoning paths that are both diverse and accurate. For challenging problems, the forking tokens t…

Code Generation

VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning Models

2026-02-27 · Soumya Suvra Ghosal, Youngeun Kim, Zhuowei Li, Ritwick Chaudhry 외 arxiv

Advances in large reasoning models have shown strong performance on complex reasoning tasks by scaling test-time compute through extended reasoning. However, recent studies observe that in vision-dependent tasks, extende…

Reinforcement LearningVisual Reasoning