paper-with-me

Papers

Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning

2026-04-11 · Yiming Huang, Zhenbo Shi, Shuzheng Gao, Cuiyun Gao, Peiyi Han, Chuanyi Liu arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) is an essential paradigm that enhances the reasoning capabilities of Large Language Models (LLMs). However, existing methods typically rely on static policy optimization schemes that misalign with the model's evolving reasoning capabilities. To address this issue, we propose Adaptive Power-Mean Policy Optimization (APMPO), which comprises two main innovations: Power-Mean Policy Optimization (PMPO) and Feedback-Adaptive Clipping (FAC). Specifically, PMPO introduces a generalized power-mean objective. This enables the model to adaptively transition from the signal-amplifying behavior of the arithmetic mean to the consistency-enforcing behavior of the geometric mean. FAC adaptively adjusts clipping bounds based on real-time reward statistics to overcome the limitations of static mechanisms. Capitalizing on these innovations, APMPO improves learning dynamics and reasoning performance. Extensive experiments on nine datasets across three reasoning tasks showcase the superiority of APMPO over state-of-the-art RLVR-based baselines. For instance, APMPO boosts the average Pass@1 score on mathematical reasoning benchmarks by 3.0 points compared to GRPO when using Qwen2.5-3B-Instruct.

📄 PDF Abstract BibTeX arXiv:2605.04066

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMathematical Reasoning

Similar Papers 제목 키워드 기반

THRIVE: Therapeutic Humanoid Robot In Virtual Environment

2026-08-14 · Jin Xu, Yu-Ping Chen, Ayanna Howard arxiv

This paper presents THRIVE (Therapeutic Humanoid Robot In Virtual Environment), an at-home rehabilitation platform that integrates a suite of virtual-reality upper-body rehabilitation games, a real-time camera-based moti…

New And Surprising Ways to Be Mean. Adversarial NPCs with Coupled Empowerment Minimisation

2018-06-04 · Christian Guckelsberger, Christoph Salge, Julian Togelius

Creating Non-Player Characters (NPCs) that can react robustly to unforeseen player behaviour or novel game content is difficult and time-consuming. This hinders the design of believable characters, and the inclusion of N…

Partially-Observable Sequential Change-Point Detection for Autocorrelated Data via Upper Confidence Region

2024-03-30 · Haijie Xu, Xiaochen Xian, Chen Zhang, Kaibo Liu

Sequential change point detection for multivariate autocorrelated data is a very common problem in practice. However, when the sensing resources are limited, only a subset of variables from the multivariate system can be…

Change Point Detection

Direct Adaptive Control of Grid-Connected Power Converters via Output-Feedback Data-Enabled Policy Optimization

2024-11-06 · Feiran Zhao, Ruohan Leng, Linbin Huang, Huanhai Xin 외

Power electronic converters are becoming the main components of modern power systems due to the increasing integration of renewable energy sources. However, power converters may become unstable when interacting with the …

Adaptivity in Adaptive Submodularity

2019-11-09 · Hossein Esfandiari, Amin Karbasi, Vahab Mirrokni

Adaptive sequential decision making is one of the central challenges in machine learning and artificial intelligence. In such problems, the goal is to design an interactive policy that plans for an action to take, from a…

Active LearningDecision MakingExperimental DesignSequential Decision Making