paper-with-me

홈 › Papers

KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical Reasoning

2025-05-22 · Wei Sun, Wen Yang, Pu Jian, Qianlong Du, Fuwei Cui, Shuo Ren, Jiajun Zhang

Recent advances have demonstrated that integrating reinforcement learning with rule-based rewards can significantly enhance the reasoning capabilities of large language models, even without supervised fine-tuning. However, prevalent reinforcement learning algorithms such as GRPO and its variants like DAPO, suffer from a coarse granularity issue when computing the advantage. Specifically, they compute rollout-level advantages that assign identical values to every token within a sequence, failing to capture token-specific contributions and hindering effective learning. To address this limitation, we propose Key-token Advantage Estimation (KTAE) - a novel algorithm that estimates fine-grained, token-level advantages without introducing additional models. KTAE leverages the correctness of sampled rollouts and applies statistical analysis to quantify the importance of individual tokens within a sequence to the final outcome. This quantified token-level importance is then combined with the rollout-level advantage to obtain a more fine-grained token-level advantage estimation. Empirical results show that models trained with GRPO+KTAE and DAPO+KTAE outperform baseline methods across five mathematical reasoning benchmarks. Notably, they achieve higher accuracy with shorter responses and even surpass R1-Distill-Qwen-1.5B using the same base model.

📄 PDF Abstract BibTeX arXiv:2505.16826

Code (1)

xiaolizh1/ktae 공식 구현 pytorch

Tasks

Mathematical Reasoningreinforcement-learningReinforcement Learning

Methods 이 논문이 사용한 방법론

DAPO Dialogue-Adaptive Pre-training Objective (DAPO) is a pre-training objective for dialogue adaptation, which is designed to measure qualities of dialogues from multiple…
BASE 설명 없음

Similar Papers 제목 키워드 기반

InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models

2026-06-01 · Xinxin Liu, Shiwei Gan, Xiao Liu, Yafeng Yin 외 arxiv

Video Large Language Models (Video-LLMs) achieve strong performance in video understanding, but their excessive visual tokens bring substantial computational overhead. Existing training-free compression methods improve i…

TokenSplat: Token-aligned 3D Gaussian Splatting for Feed-forward Pose-free Reconstruction

2026-02-28 · Yihui Li, Chengxin Lv, Zichen Tang, Hongyu Yang 외 arxiv

We present TokenSplat, a feed-forward framework for joint 3D Gaussian reconstruction and camera pose estimation from unposed multi-view images. At its core, TokenSplat introduces a Token-aligned Gaussian Prediction modul…

Camera Pose Estimation

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits

2026-05-09 · Tianhao Cheng, Zeyu Huang, Zihan Qiu, Yu Cheng 외 arxiv

A commonly accepted explanation of critic-free RL for LLMs, based on sequence-level rewards, is that it reinforces successful rollouts with a positive advantage while penalizing failed ones. In contrast, we study critic-…

Segmental Advantage Estimation: Enhancing PPO for Long-Context LLM Training

2026-01-12 · Xue Gong, Qi Yi, Ziyuan Nan, Guanhua Huang 외 arxiv

Training Large Language Models (LLMs) for reasoning tasks is increasingly driven by Reinforcement Learning with Verifiable Rewards (RLVR), where Proximal Policy Optimization (PPO) provides a principled framework for stab…

Reinforcement Learning

SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

2026-09-24 · Yan Zhan, Shaobo Liu, Qiunan Liu, Yuanjun Shi 외 hf

Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with user-facing natural language summaries. This output heterogeneity presents a structural failure mode in standard on-policy …

Reinforcement Learning