paper-with-me

홈 › Papers

Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies

2025-12-22 · Yuqiao Tan, Minzheng Wang, Shizhu He, Huanxuan Liao, Chengfeng Zhao, Qiunan Lu, Tian Liang, Jun Zhao, Kang Liu arxiv

Existing reinforcement learning (RL) approaches treat large language models (LLMs) as a unified policy, overlooking their internal mechanisms. In this paper, we decompose the LLM-based policy into Internal Layer Policies and Internal Modular Policies via the Transformer's residual stream. Our entropy analysis of internal policy reveals distinct patterns: (1) universally, internal policies evolve from high-entropy exploration in early layers to deterministic refinement in the top layers; and (2) Qwen exhibits an explicit progressive reasoning structure, contrasting with the abrupt convergence in Llama. Furthermore, we discover that optimizing internal layers induces feature refinement, forcing lower layers to capture high-level reasoning representations early. Motivated by these findings, we propose Bottom-up Policy Optimization (BuPO), a novel RL paradigm that reconstructs the LLM's reasoning foundation from the bottom up by optimizing internal layers in early stages. Extensive experiments on complex reasoning benchmarks demonstrate the effectiveness of BuPO.

📄 PDF Abstract BibTeX arXiv:2512.19673

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Plan Your Target and Learn Your Skills: State-Only Imitation Learning via Decoupled Policy Optimization

2021-09-29 · NeurIPS 2021 12 · Minghuan Liu, Zhengbang Zhu, Yuzheng Zhuang, Weinan Zhang 외

State-only imitation learning (SOIL) enables agents to learn from massive demonstrations without explicit action or reward information. However, previous methods attempt to learn the implicit state-to-action mapping poli…

Imitation LearningReinforcement Learning (RL)

Plan Your Target and Learn Your Skills: Transferable State-Only Imitation Learning via Decoupled Policy Optimization

2022-03-04 · Minghuan Liu, Zhengbang Zhu, Yuzheng Zhuang, Weinan Zhang 외

Recent progress in state-only imitation learning extends the scope of applicability of imitation learning to real-world settings by relieving the need for observing expert actions. However, existing solutions only learn …

Imitation LearningTransfer Learning

Bottom-Up Meta-Policy Search

2019-10-22 · Luckeciano C. Melo, Marcos R. O. A. Maximo, Adilson Marques da Cunha

Despite of the recent progress in agents that learn through interaction, there are several challenges in terms of sample efficiency and generalization across unseen behaviors during training. To mitigate these problems, …

Meta-LearningReinforcement Learning

Is Crowdsourcing Breaking Your Bank? Cost-Effective Fine-Tuning of Pre-trained Language Models with Proximal Policy Optimization

2024-02-28 · Shuo Yang, Gjergji Kasneci

Wide usage of ChatGPT has highlighted the potential of reinforcement learning from human feedback. However, its training pipeline relies on manual ranking, a resource-intensive process. To reduce labor costs, we propose …

Language ModelingLanguage Modelling

Control Your Robot: A Unified System for Robot Control and Policy Deployment

2025-09-28 · Tian Nian, Weijie Ke, Shaolong Zhu, Bingshan Hu arxiv

Cross-platform robot control remains difficult because hardware interfaces, data formats, and control paradigms vary widely, which fragments toolchains and slows deployment. To address this, we present Control Your Robot…