paper-with-me

Papers

Burning RED: Unlocking Subtask-Driven Reinforcement Learning and Risk-Awareness in Average-Reward Markov Decision Processes

2024-10-14 · Juan Sebastian Rojas, Chi-Guhn Lee

Average-reward Markov decision processes (MDPs) provide a foundational framework for sequential decision-making under uncertainty. However, average-reward MDPs have remained largely unexplored in reinforcement learning (RL) settings, with the majority of RL-based efforts having been allocated to episodic and discounted MDPs. In this work, we study a unique structural property of average-reward MDPs and utilize it to introduce Reward-Extended Differential (or RED) reinforcement learning: a novel RL framework that can be used to effectively and efficiently solve various learning objectives, or subtasks, simultaneously in the average-reward setting. We introduce a family of RED learning algorithms for prediction and control, including proven-convergent algorithms for the tabular case. We then showcase the power of these algorithms by demonstrating how they can be used to learn a policy that optimizes, for the first time, the well-known conditional value-at-risk (CVaR) risk measure in a fully-online manner, without the use of an explicit bi-level optimization scheme or an augmented state-space.

📄 PDF Abstract BibTeX arXiv:2410.10578

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingDecision Making Under Uncertaintyreinforcement-learningReinforcement LearningReinforcement Learning (RL)Sequential Decision Making

Similar Papers 제목 키워드 기반

Active Disruption Avoidance and Trajectory Design for Tokamak Ramp-downs with Neural Differential Equations and Reinforcement Learning

2024-02-14 · Allen M. Wang, Oswin So, Charles Dawson, Darren T. Garnier 외

The tokamak offers a promising path to fusion energy, but plasma disruptions pose a major economic risk, motivating considerable advances in disruption avoidance. This work develops a reinforcement learning approach to t…

GPU

Detecting Crop Burning in India using Satellite Data

2022-09-21 · Kendra Walker, Ben Moscona, Kelsey Jack, Seema Jayachandran 외

Crop residue burning is a major source of air pollution in many parts of the world, notably South Asia. Policymakers, practitioners and researchers have invested in both measuring impacts and developing interventions to …

SDRL: Interpretable and Data-efficient Deep Reinforcement Learning Leveraging Symbolic Planning

2018-10-31 · Daoming Lyu, Fangkai Yang, Bo Liu, Steven Gustafson

Deep reinforcement learning (DRL) has gained great success by learning directly from high-dimensional sensory inputs, yet is notorious for the lack of interpretability. Interpretability of the subtasks is critical in hie…

Decision MakingDeep Reinforcement Learningreinforcement-learningReinforcement Learning+2

Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation

2025-11-04 · Zhiwei Zhang, Xiaomin Li, Yudi Lin, Hui Liu 외 arxiv

Large Language Models (LLMs) trained with reinforcement learning and verifiable rewards have achieved strong results on complex reasoning tasks. Recent work extends this paradigm to a multi-agent setting, where a meta-th…

Reinforcement Learning

CBAG: An Efficient Genetic Algorithm for the Graph Burning Problem

2022-08-01 · Mahdi Nazeri, Ali Mollahosseini, Iman Izadi

Information spread is an intriguing topic to study in network science, which investigates how information, influence, or contagion propagate through networks. Graph burning is a simplified deterministic model for how inf…