paper-with-me

홈 › Papers

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

2025-03-06 · Pranjal Aggarwal, Sean Welleck

Reasoning language models have shown an uncanny ability to improve performance at test-time by ``thinking longer''-that is, by generating longer chain-of-thought sequences and hence using more compute. However, the length of their chain-of-thought reasoning is not controllable, making it impossible to allocate test-time compute to achieve a desired level of performance. We introduce Length Controlled Policy Optimization (LCPO), a simple reinforcement learning method that optimizes for accuracy and adherence to user-specified length constraints. We use LCPO to train L1, a reasoning language model that produces outputs satisfying a length constraint given in its prompt. L1's length control allows for smoothly trading off computational cost and accuracy on a wide range of tasks, and outperforms the state-of-the-art S1 method for length control. Furthermore, we uncover an unexpected short chain-of-thought capability in models trained with LCPO. For instance, our 1.5B L1 model surpasses GPT-4o at equal reasoning lengths. Overall, LCPO enables precise control over reasoning length, allowing for fine-grained allocation of test-time compute and accuracy. We release code and models at https://www.cmu-l3.github.io/l1

📄 PDF Abstract BibTeX arXiv:2503.04697

Code (1)

cmu-l3/l1

Similar Papers 제목 키워드 기반

THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

2026-01-30 · Seanie Lee, Sangwoo Park, Yumin Choi, Gyeongman Kim 외 arxiv

Large reasoning models (LRMs) achieve remarkable performance by leveraging reinforcement learning (RL) on reasoning tasks to generate long chain-of-thought (CoT) reasoning. However, this over-optimization often prioritiz…

Reinforcement Learning

ThinkSwitcher: When to Think Hard, When to Think Fast

2025-05-20 · Guosheng Liang, Longguang Zhong, ZiYi Yang, Xiaojun Quan

Large reasoning models (LRMs) excel at solving complex tasks by leveraging long chain-of-thought (CoT) reasoning. However, this often leads to overthinking on simple tasks, resulting in unnecessary computational overhead…

Thinking in Streaming Video

2026-03-13 · Zikang Liu, Longteng Guo, Handong Li, Ru Zhen 외 arxiv

Real-time understanding of continuous video streams is essential for interactive assistants and multimodal agents operating in dynamic environments. However, most existing video reasoning approaches follow a batch paradi…

Reinforcement Learning

The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning

2025-10-08 · Milad Aghajohari, Kamran Chitsaz, Amirhossein Kazemnejad, Sarath Chandar 외 arxiv

Reinforcement learning (RL) has recently become a strong recipe for training reasoning LLMs that produce long chains of thought (LongCoT). Yet the standard RL "thinking environment", where the state is the prompt plus al…

Reinforcement Learning

The Relationship Between Reasoning and Performance in Large Language Models -- o3 (mini) Thinks Harder, Not Longer

2025-02-21 · Marthe Ballon, Andres Algaba, Vincent Ginis

Large language models have demonstrated remarkable progress in mathematical reasoning, leveraging chain-of-thought and test-time compute scaling. However, many open questions remain regarding the interplay between reason…

MathMathematical Reasoning