paper-with-me

Papers

Contextual Bilevel Reinforcement Learning for Incentive Alignment

2024-06-03 · Vinzenz Thoma, Barna Pasztor, Andreas Krause, Giorgia Ramponi, Yifan Hu

The optimal policy in various real-world strategic decision-making problems depends both on the environmental configuration and exogenous events. For these settings, we introduce Contextual Bilevel Reinforcement Learning (CB-RL), a stochastic bilevel decision-making model, where the lower level consists of solving a contextual Markov Decision Process (CMDP). CB-RL can be viewed as a Stackelberg Game where the leader and a random context beyond the leader's control together decide the setup of many MDPs that potentially multiple followers best respond to. This framework extends beyond traditional bilevel optimization and finds relevance in diverse fields such as RLHF, tax design, reward shaping, contract theory and mechanism design. We propose a stochastic Hyper Policy Gradient Descent (HPGD) algorithm to solve CB-RL, and demonstrate its convergence. Notably, HPGD uses stochastic hypergradient estimates, based on observations of the followers' trajectories. Therefore, it allows followers to use any training procedure and the leader to be agnostic of the specific algorithm, which aligns with various real-world scenarios. We further consider the setting when the leader can influence the training of followers and propose an accelerated algorithm. We empirically demonstrate the performance of our algorithm for reward shaping and tax design.

📄 PDF Abstract BibTeX arXiv:2406.01575

Code (1)

lasgroup/hpgd 공식 구현 jax

Tasks

Bilevel OptimizationDecision Makingreinforcement-learningReinforcement Learning

Similar Papers 제목 키워드 기반

Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF

2024-02-10 · Han Shen, Zhuoran Yang, Tianyi Chen

Bilevel optimization has been recently applied to many machine learning tasks. However, their applications have been restricted to the supervised learning setting, where static objective functions with benign structures …

Bilevel Optimizationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

On The Sample Complexity Bounds In Bilevel Reinforcement Learning

2025-03-22 · Mudit Gaur, Amrit Singh Bedi, Raghu Pasupathu, Vaneet Aggarwal

Bilevel reinforcement learning (BRL) has emerged as a powerful mathematical framework for studying generative AI alignment and related problems. While several principled algorithmic frameworks have been proposed, key the…

Bilevel Optimizationreinforcement-learningReinforcement Learning

Bilevel Optimization over Saddle Points of Zero-Sum Markov Games

2026-05-26 · Zihao Zheng, Irwin King, Songtao Lu arxiv

Reinforcement learning (RL) often has a hierarchical structure, where an upper-level (UL) learner selects model parameters and a lower-level (LL) decision-making process responds, naturally leading to a bilevel optimizat…

Reinforcement LearningBilevel Optimization

Inducing Equilibria via Incentives: Simultaneous Design-and-Play Ensures Global Convergence

2021-10-04 · Boyi Liu, Jiayang Li, Zhuoran Yang, Hoi-To Wai 외

To regulate a social system comprised of self-interested agents, economic incentives are often required to induce a desirable outcome. This incentive design problem naturally possesses a bilevel structure, in which a des…

Bilevel Optimization

AI Alignment via Incentives and Correction

2026-05-02 · Rohit Agarwal, Joshua Lin, Mark Braverman, Elad Hazan arxiv

We study AI alignment through the lens of law-and-economics models of deterrence and enforcement. In these models, misconduct is not treated as an external failure, but as a strategic response to incentives: an actor wei…

Bilevel Optimization