paper-with-me

홈 › Papers

fg-expo: Frontier-guided exploration-prioritized policy optimization via adaptive kl and gaussian curriculum

2026-05-12 · Mingxiong Lin, Zhangquan Gong, Maowen Tang, Qian Li, Chuangchuang Wang, Jian Ma, Sutian Huang, Kai Tang, Haonan Lu arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, with Group Relative Policy Optimization (GRPO) serving as the dominant algorithm. We identify two overlooked inefficiencies inherent in GRPO. First, a fixed KL coefficient overly restricts policy exploration at moments when the model needs to diverge significantly from the reference policy. Second, uniform question sampling overlooks that moderately difficult problems produce the most informative gradient signals. We propose FG-ExPO, short for Frontier-Guided Exploration-Prioritized Policy Optimization, which integrates two lightweight components. Accuracy-Conditioned KL Scaling (AKL) adjusts the KL penalty strength through a smooth nonlinear function of batch average accuracy, loosening the constraint when the model performs poorly and strengthening it when the model achieves satisfactory results. Gaussian Curriculum Sampling (GCS) assigns sampling weights to questions following a Gaussian distribution centered at a moderate accuracy level around 0.5, focusing model training on its learning frontier. We conduct evaluations on DeepSeek-R1-Distill-Qwen-1.5B and Qwen3-8B-Base across six mainstream mathematical reasoning benchmarks. Experimental results demonstrate that FG-ExPO consistently outperforms vanilla GRPO. It delivers an absolute improvement of 13.34 on the AIME 2025 pass@32 metric, rising from 63.33 percent to 76.67 percent, and obtains an average pass@32 gain of 2.66 on the 8B model. The substantially larger performance gains observed on pass@32 compared to pass@1 verify that FG-ExPO enlarges the model's effective exploration space under a fixed inference budget.

📄 PDF Abstract BibTeX arXiv:2605.11403

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMathematical Reasoning

Similar Papers 제목 키워드 기반

expo: Exploration-prioritized policy optimization via adaptive kl regulation and gaussian curriculum sampling

2026-05-11 · Mingxiong Lin, Zhangquan Gong, Maowen Tang, Qian Li 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, where Group Relative Policy Optimization (GRPO) serves as the mainstream algorithm. We point out two …

Reinforcement LearningMathematical Reasoning

Semantic-Aware Guided Drone Exploration for Language-Conditioned 3D Indoor Mapping

2026-05-22 · Nitin Vegesna, Avideh Zakhor arxiv

We present Semantic-Aware Guided Exploration, SAGE, a system for open-vocabulary exploration in unknown 3D indoor environments that preserves coverage-oriented behavior while allowing semantic cues to reprioritize fronti…

Learning-Guided Sparsification of Dynamic Graphs in Robotic Exploration

2026-04-15 · Adithya V. Sastry, Bibek Poudel, Weizi Li arxiv

Many robotic exploration algorithms rely on graph structures for frontier-based exploration and dynamic path planning. However, these graphs grow rapidly, accumulating redundant information and impacting performance. We …

Search Inspired Exploration in Reinforcement Learning

2026-01-31 · Georgios Sotirchos, Zlatan Ajanović, Jens Kober arxiv

Exploration in environments with sparse rewards remains a fundamental challenge in reinforcement learning (RL). Existing approaches such as curriculum learning and Go-Explore often rely on hand-crafted heuristics, while …

Reinforcement Learning

FH-DRL: Exponential-Hyperbolic Frontier Heuristics with DRL for accelerated Exploration in Unknown Environments

2024-07-26 · Seunghyeop Nam, Tuan Anh Nguyen, Eunmi Choi, Dugki Min

Autonomous robot exploration in large-scale or cluttered environments remains a central challenge in intelligent vehicle applications, where partial or absent prior maps constrain reliable navigation. This paper introduc…

Autonomous DrivingAutonomous NavigationAutonomous VehiclesDeep Reinforcement Learning