paper-with-me

홈 › Papers

Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency

2025-08-19 · Rapheal Huang, Weilong Guo arxiv

Large Language Models (LLMs) excel at reasoning, traditionally requiring high-quality large-scale data and extensive training. Recent works reveal a very appealing Less-Is-More phenomenon where very small, carefully curated high-quality datasets match resource-intensive approaches. In this work, we further systematically relax their quality constraints by adding controlled noise via persona context relevance and comparing datasets of different qualities. Counterintuitively, we find that mixing relevant and irrelevant contexts consistently across training and inference stages yields optimal results -- a phenomenon we term training-testing co-design. Dataset quality comparisons show that high-quality data benefits weaker models on easy questions, while low-quality data achieves higher scores on hard questions with capable models. Across our experiments, reasoning performance is linked to reasoning efficiency. We, for the first time, found adding noisy and irrelevant contexts into queries can improve reasoning efficiency without any prices and targeted designs. Building on these insights, we propose Input-Time Scaling: applying small, low-quality data to capable models with training-testing co-design. This maintains Less-Is-More while further removing labor-intensive quality curation and improving reasoning effectiveness and efficiency, making the approach more applicable and affordable. Our method achieves 76.7% pass@1 on AIME24/25 using Qwen2.5-32B-Instruct, and 90.0%/80.0% with DeepSeek-R1-Distill-Qwen-32B -- state-of-the-art among Qwen2.5-32B variants. We are open-sourcing our datasets, pipelines, evaluation results, and checkpoints to facilitate reproducibility and further research.

📄 PDF Abstract BibTeX arXiv:2508.13654

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Training Data Size Induced Double Descent For Denoising Neural Networks and the Role of Training Noise Level

2021-09-29 · Rishi Sonthalia, Raj Rao Nadakuditi

When training a denoising neural network, we show that more data isn’t more beneficial. In fact the generalization error versus number of of training data points is a double descent curve. Training a network to denoise n…

Denoising

Data Scaling Laws in NMT: The Effect of Noise and Architecture

2022-02-04 · Yamini Bansal, Behrooz Ghorbani, Ankush Garg, Biao Zhang 외

In this work, we study the effect of varying the architecture and training data quality on the data scaling properties of Neural Machine Translation (NMT). First, we establish that the test loss of encoder-decoder transf…

DecoderLanguage ModelingLanguage ModellingMachine Translation+1

On the Complexity of Strong and Epistemic Credal Networks

2013-09-26 · Denis D. Maua, Cassio Polpo de Campos, Alessio Benavoli, Alessandro Antonucci

Credal networks are graph-based statistical models whose parameters take values in a set, instead of being sharply specified as in traditional statistical models (e.g., Bayesian networks). The computational complexity of…

Explaining Scaling Laws of Neural Network Generalization

2021-09-29 · Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee 외

The test loss of well-trained neural networks often follows precise power-law scaling relations with either the size of the training dataset or the number of parameters in the network. We propose a theory that explains a…

Off-policy Evaluation with Deeply-abstracted States

2024-06-27 · Meiling Hao, Pingfan Su, Liyuan Hu, Zoltan Szabo 외

Off-policy evaluation (OPE) is crucial for assessing a target policy's impact offline before its deployment. However, achieving accurate OPE in large state spaces remains challenging. This paper studies state abstraction…

Off-policy evaluation