paper-with-me

홈 › Papers

Rethinking GSPO: The Perplexity-Entropy Equivalence

2025-10-27 · Chi Liu arxiv

We provide a new perspective on GSPO's length-normalized importance ratios by establishing their connection to information-theoretic quantities. We show that GSPO's sequence-level weight $s(θ) = (π_θ/π_{θ_{\text{old}}})^{1/|y|}$ can be equivalently expressed as the inverse perplexity ratio $\text{PPL}_{θ_{\text{old}}}/\text{PPL}_θ$ and as the exponential cross-entropy change $\exp(ΔH)$. While the perplexity-entropy relationship follows from standard definitions, this observation provides a useful lens for understanding GSPO: the algorithm weights policy gradient updates by perplexity ratios, offering an information-theoretic interpretation of the importance weights. This perspective helps explain GSPO's empirical properties, including log-domain variance reduction through geometric averaging and stability in training mixture-of-experts models. We validate the mathematical equivalences and variance predictions through controlled experiments on mathematical reasoning tasks.

📄 PDF Abstract BibTeX arXiv:2510.23142

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

Entropy-Gated Selective Policy Optimization:Token-Level Gradient Allocation for Hybrid Training of Large Language Models

2026-02-03 · Yuelin Hu, Zhengxue Cheng, Wei Liu, Li Song arxiv

Hybrid training methods for large language models combine supervised fine tuning (SFT) on expert demonstrations with reinforcement learning (RL) on model rollouts, typically at the sample level. We propose Entropy Gated …

Reinforcement LearningMathematical Reasoning

SSPO: Subsentence-level Policy Optimization

2025-11-06 · Kun Yang, Zikang chen, Yanmeng Wang, Zhigen Li 외 arxiv

As a key component of large language model (LLM) post-training, Reinforcement Learning from Verifiable Rewards (RLVR) has substantially improved reasoning performance. However, existing RLVR algorithms exhibit distinct s…

Reinforcement Learning

Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance

2024-12-13 · Jiahao Lyu, Wei Wang, Dongbao Yang, Jinwen Zhong 외

Scene text spotting has attracted the enthusiasm of relative researchers in recent years. Most existing scene text spotters follow the detection-then-recognition paradigm, where the vanilla detection module hardly determ…

Scene Text RecognitionText Spotting

Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR

2026-01-09 · Zijun Min, Bingshuai Liu, Ante Wang, Long Zhang 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising framework for optimizing large language models in reasoning tasks. However, existing RLVR algorithms focus on different granularities, and each has…

Reinforcement LearningMathematical Reasoning

When RL Suppresses Its Own Vocabulary: Recovering Reasoning Diversity in Puzzle-to-Math Transfer

2026-05-28 · Mayug Maniparambil, Arjun Karuvally, Terrence Sejnowski, Fergal Reid arxiv

Reinforcement learning using verifiable rewards (RLVR) improves LLM reasoning, but the conditions under which it transfers across domains -- and why it does so -- remain under-explored. We study cross-domain transfer in …

Reinforcement Learning