paper-with-me

홈 › Papers

Self-Training Elicits Concise Reasoning in Large Language Models

2025-02-27 · Tergel Munkhbat, Namgyu Ho, Seohyun Kim, Yongjin Yang, Yujin Kim, Se-Young Yun

Chain-of-thought (CoT) reasoning has enabled large language models (LLMs) to utilize additional computation through intermediate tokens to solve complex tasks. However, we posit that typical reasoning traces contain many redundant tokens, incurring extraneous inference costs. Upon examination of the output distribution of current LLMs, we find evidence on their latent ability to reason more concisely, relative to their default behavior. To elicit this capability, we propose simple fine-tuning methods which leverage self-generated concise reasoning paths obtained by best-of-N sampling and few-shot conditioning, in task-specific settings. Our combined method achieves a 30% reduction in output tokens on average, across five model families on GSM8K and MATH, while maintaining average accuracy. By exploiting the fundamental stochasticity and in-context learning capabilities of LLMs, our self-training approach robustly elicits concise reasoning on a wide range of models, including those with extensive post-training. Code is available at https://github.com/TergelMunkhbat/concise-reasoning

📄 PDF Abstract BibTeX arXiv:2502.20122

Code (1)

tergelmunkhbat/concise-reasoning 공식 구현 pytorch

Tasks

GSM8KIn-Context LearningMath

Similar Papers 제목 키워드 기반

Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning

2026-07-17 · Leichao Dong, Dongxu Zhang, Yiding Sun, Qirui Wang 외 arxiv

Large reasoning models often solve problems through long chain-of-thought (CoT) traces, yet much of this computation is spent on redundant derivations, repeated self-verification, and detours that do not improve the fina…

In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-Feedback

2025-11-13 · Mingye Zhu, Yi Liu, Zheren Fu, Quan Wang 외 arxiv

Training Large Language Models (LLMs) for chain-of-thought reasoning presents a significant challenge: supervised fine-tuning on a single "golden" rationale hurts generalization as it penalizes equally valid alternatives…

Reinforcement Learning

Self-Refine Instruction-Tuning for Aligning Reasoning in Language Models

2024-05-01 · Leonardo Ranaldi, Andrè Freitas

The alignments of reasoning abilities between smaller and larger Language Models are largely conducted via Supervised Fine-Tuning (SFT) using demonstrations generated from robust Large Language Models (LLMs). Although th…

Math

Diversity of Thought Elicits Stronger Reasoning Capabilities in Multi-Agent Debate Frameworks

2024-10-10 · Mahmood Hegazy

Large language models (LLMs) excel in natural language generation but often confidently produce incorrect responses, especially in tasks like mathematical reasoning. Chain-of-thought prompting, self-verification, and mul…

8kDiversityMathematical ReasoningText Generation

Self-Polish: Enhance Reasoning in Large Language Models via Problem Refinement

2023-05-23 · Zhiheng Xi, Senjie Jin, Yuhao Zhou, Rui Zheng 외

To enhance the multi-step reasoning capabilities of large language models, researchers have extensively explored prompting methods, notably the Chain-of-Thought (CoT) method which explicitly elicits human-like rationales…

GSM8K