paper-with-me

홈 › Papers

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity

2026-02-11 · Haihui Pan, Yuzhong Hong, Kaichen Zhang, Shaoke Lv, Junwei Bao, Hongfei Jiang, Yang Song arxiv

In many large language model (LLM) alignment applications, users expect not only high-quality outputs but also substantial diversity. However, existing methods often face a fundamental trade-off between these objectives: approaches that improve output quality tend to reduce diversity, while methods that increase diversity often do so at the expense of quality. In this work, we propose Quality-constrained Entropy Maximization Policy Optimization (QEMPO), a novel framework that enhances the diversity of LLM outputs while explicitly preserving output quality. QEMPO is grounded in a strong theoretical foundation: we derive a closed-form analytical solution that provably maximizes entropy-a principled measure of diversity-subject to a quality constraint, with guarantees on optimality under the defined objective. Leveraging this solution, QEMPO naturally supports both online and offline training settings. Empirical results demonstrate that QEMPO consistently improves output diversity without sacrificing quality, and in many cases yields gains in both dimensions compared to existing baselines, aligning with our theoretical guarantees.

📄 PDF Abstract BibTeX arXiv:2602.15894

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Manifold Sampling via Entropy Maximization

2026-05-12 · Cornelius V. Braun, Tilman Burghoff, Marc Toussaint arxiv

Sampling from constrained distributions has a wide range of applications, including in Bayesian optimization and robotics. Prior work establishes convergence and feasibility guarantees for constrained sampling, but assum…

Density Estimation

How to Explore with Belief: State Entropy Maximization in POMDPs

2024-06-04 · Riccardo Zamboni, Duilio Cirino, Marcello Restelli, Mirco Mutti

Recent works have studied *state entropy maximization* in reinforcement learning, in which the agent's objective is to learn a policy inducing high entropy over states visitation (Hazan et al., 2019). They typically assu…

Hallucination

A general Markov decision process formalism for action-state entropy-regularized reward maximization

2023-02-02 · Dmytro Grytskyy, Jorge Ramírez-Ruiz, Rubén Moreno-Bote

Previous work has separately addressed different forms of action, state and action-state entropy regularization, pure exploration and space occupation. These problems have become extremely relevant for regularization, ge…

Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization

2026-05-11 · Mengqi He, Xinyu Tian, Xin Shen, Shu Zou 외 arxiv

Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transferability, casting doubt on the feasibility of transferable multimodal jailbreaks…

Global Optimality for Constrained Exploration via Penalty Regularization

2026-04-30 · Florian Wolf, Ilyas Fatkhullin, Niao He arxiv

Efficient exploration is a central problem in reinforcement learning and is often formalized as maximizing the entropy of the state-action occupancy measure. While unconstrained maximum-entropy exploration is relatively …

Reinforcement Learning