paper-with-me

홈 › Papers

Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation

2025-10-10 · Sondos Mahmoud Bsharat, Zhiqiang Shen arxiv

Large language models (LLMs) have demonstrated impressive reasoning capabilities when provided with chain-of-thought exemplars, but curating large reasoning datasets remains laborious and resource-intensive. In this work, we introduce Prompting Test-Time Scaling (P-TTS), a simple yet effective inference-time data augmentation strategy for enhancing LLM reasoning through finetuning. Rather than collecting thousands or even millions of examples, P-TTS leverages a small pool of only 90 manually selected reasoning instances and systematically varies exemplar augmentation through principled instruction prompting intensities at test time to synthesize diverse reasoning trajectory contexts. Then we finetune the various sizes of Qwen-2.5 models on P-TTS data. Across a suite of mathematical reasoning AIME2024 & 25, MATH500, and GPQA-Diamond, our P-TTS-7B and 32B models outperform the prior competitive baselines like S1 and S1.1 (1K-shot), achieving absolute accuracy gains of +26.66% and +30.00% on AIME'24 (7B), and +13.34% and +6.67% on AIME'25 (7B); P-TTS-32B yields gains of +23.33% and +16.63% on AIME'24, and +26.63% and +3.33% on AIME'25 (vs. S1 and S1.1, respectively), with comparable or better performance on MATH500 and GPQA-Diamond. We further show that P-TTS enhances zero-shot generalization accuracy on out-of-domain reasoning benchmarks of Gaokao, Kaoyan, OlympiadBench, AMC23, GradeSchoolMath, and Minerva. Our analysis suggests that test-time scaling effectively explores the latent space of reasoning patterns, amplifying LLM problem-solving with minimal annotation overhead, and further unlocking the reasoning potential and capabilities of LLMs. Prompting Test-Time Scaling offers a practical, low-cost way to elicit LLM reasoning in resource-constrained or rapidly evolving domains.

📄 PDF Abstract BibTeX arXiv:2510.09599

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationMathematical ReasoningData Augmentation

Similar Papers 제목 키워드 기반

Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability Theory

2025-05-16 · Yexiang Liu, Zekun Li, Zhi Fang, Nan Xu 외

Recently, scaling test-time compute on Large Language Models (LLM) has garnered wide attention. However, there has been limited investigation of how various reasoning prompting strategies perform as scaling. In this pape…

Learning Generative Selection for Best-of-N

2026-02-02 · Shubham Toshniwal, Aleksander Ficek, Siddhartha Jain, Wei Du 외 arxiv

Scaling test-time compute via parallel sampling can substantially improve LLM reasoning, but is often limited by Best-of-N selection quality. Generative selection methods, such as GenSelect, address this bottleneck, yet …

Reinforcement Learning

Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space

2025-05-19 · Hengli Li, Chenxi Li, Tong Wu, Xuekai Zhu 외

Reasoning ability, a core component of human intelligence, continues to pose a significant challenge for Large Language Models (LLMs) in the pursuit of AGI. Although model performance has improved under the training scal…

GSM8KMath

Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLM

2025-05-20 · Zhen Xiong, Yujun Cai, Zhecheng Li, Yiwei Wang

Recent advances in test-time scaling have enabled Large Language Models (LLMs) to display sophisticated reasoning abilities via extended Chain-of-Thought (CoT) generation. Despite their potential, these Reasoning LLMs (R…

Prompt Engineering

$\nabla$-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space

2026-03-05 · Peihao Wang, Ruisi Cai, Zhen Wang, Hongyuan Mei 외 arxiv

Scaling inference-time compute for Large Language Models (LLMs) has unlocked unprecedented reasoning capabilities. However, existing inference-time scaling methods typically rely on inefficient and suboptimal discrete se…

Reinforcement LearningMathematical Reasoning