paper-with-me

홈 › Papers

Cyclical Entropy Eruption: Entropy Dynamics in Agent Reinforcement Learning

2026-05-27 · Wendi Li, Shawn Im, Sharon Li arxiv

Agentic large language models are increasingly used to solve real-world tasks by reasoning over goals, invoking tools, and interacting with external environments. Reinforcement learning provides a natural framework for improving these behaviors, and recent agent RL methods have achieved strong results across domains. However, the training dynamics of agent RL remain poorly understood, limiting our ability to diagnose instabilities and design more effective training algorithms. In this work, we identify a previously underexplored phenomenon in agent RL, which we term cyclical entropy eruption. Unlike single-turn reasoning RL, where entropy typically collapses and stays low, agent RL training exhibits unique recurring cycles of sharp entropy eruption and gradual subsidence. We decompose this dynamic into three phases and provide theoretical and empirical analyses of each, explaining the mechanisms underlying its cyclical oscillation. We further show that degenerate patterns such as sentence duplication and hallucination, once acquired during eruption, can persist and accumulate across cycles. Motivated by these findings, we propose SEAL (Separation-Enhanced Agent Learning), a lightweight auxiliary loss that separates correct and incorrect trajectories in representation space, directly targeting the root cause of entropy eruption. Experiments across multiple benchmarks, models, and RL algorithms demonstrate that SEAL stabilizes training and yields stronger downstream agent performance.

📄 PDF Abstract BibTeX arXiv:2605.27954

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Generalized Derangetropy Functionals for Modeling Cyclical Information Flow

2025-04-20 · Masoud Ataei, Xiaogang Wang

This paper introduces a framework for modeling cyclical and feedback-driven information flow through a generalized family of entropy-modulated transformations called derangetropy functionals. Unlike scalar and static ent…

Cyclical Focal Loss

2022-02-16 · Leslie N. Smith

The cross-entropy softmax loss is the primary loss function used to train deep neural networks. On the other hand, the focal loss function has been demonstrated to provide improved performance when there is an imbalance …

When Does Multi-Agent Collaboration Help? An Entropy Perspective

2026-02-04 · Yuxuan Zhao, Sijia Chen, Ningxin Su arxiv

Multi-agent systems (MAS) have emerged as a prominent paradigm for leveraging large language models (LLMs) to tackle complex tasks. However, the mechanisms governing the effectiveness of MAS built upon publicly available…

Trajectory Entropy Reinforcement Learning for Predictable and Robust Control

2025-05-07 · Bang You, Chenxu Wang, Huaping Liu

Simplicity is a critical inductive bias for designing data-driven controllers, especially when robustness is important. Despite the impressive results of deep reinforcement learning in complex control tasks, it is prone …

Deep Reinforcement LearningInductive Biasreinforcement-learningReinforcement Learning

Accurate Evaluation of Asset Pricing Under Uncertainty and Ambiguity of Information

2018-03-26

Since exchange economy considerably varies in the market assets, asset prices have become an attractive research area for investigating and modeling ambiguous and uncertain information in today markets. This paper propos…

Bayesian Inference