paper-with-me

홈 › Papers

Breaking the Grid: Distance-Guided Reinforcement Learning in Large Discrete Action Spaces

2026-02-09 · Heiko Hoppe, Fabian Akkerman, Wouter van Heeswijk, Maximilian Schiffer arxiv

Reinforcement Learning (RL) is increasingly applied to large-scale decision-making problems like logistics, scheduling, and recommender systems, but existing algorithms struggle with the curse of dimensionality in such large discrete action spaces. We propose Distance-Guided Reinforcement Learning (DGRL), combining Sampled Dynamic Neighborhoods and Distance-Based Updates to enable efficient RL in problems with up to $10^{20}$ actions. Unlike prior methods, DGRL performs stochastic volumetric exploration and transforms policy optimization into a stable regression task, decoupling gradient variance from action space cardinality. On structured tasks, DGRL provably guarantees local value improvement. DGRL naturally generalizes to hybrid continuous-discrete action spaces. We demonstrate performance improvements of up to 66% against state-of-the-art benchmarks across regularly and irregularly structured environments, while simultaneously improving convergence speed and computational complexity.

📄 PDF Abstract BibTeX arXiv:2602.08616

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking

2026-02-06 · Sung-Hoon Yoon, Ruizhi Qian, Minda Zhao, Weiyue Li 외 arxiv

Large Language Models (LLMs) have become integral to many domains, making their safety a critical priority. Prior jailbreaking research has explored diverse approaches, including prompt optimization, automated red teamin…

Reinforcement LearningRed Teaming

Breaking the Performance Ceiling in Complex Reinforcement Learning requires Inference Strategies

2025-05-27 · Felix Chalumeau, Daniel Rajaonarivonivelomanantsoa, Ruan de Kock, Claude Formanek 외

Reinforcement learning (RL) systems have countless applications, from energy-grid management to protein design. However, such real-world scenarios are often extremely difficult, combinatorial in nature, and require compl…

Protein DesignReinforcement Learning (RL)

Partially Equivariant Reinforcement Learning in Symmetry-Breaking Environments

2025-11-30 · Junwoo Chang, Minwoo Park, Joohwan Seo, Roberto Horowitz 외 arxiv

Group symmetries provide a powerful inductive bias for reinforcement learning (RL), enabling efficient generalization across symmetric states and actions via group-invariant Markov Decision Processes (MDPs). However, rea…

Reinforcement LearningContinuous Control

High-Fidelity Simulation and Novel Data Analysis of the Bubble Creation and Sound Generation Processes in Breaking Waves

2022-11-06 · Qiang Gao, Grant B. Deane, Saswata Basak, Umberto Bitencourt 외

Recent increases in computing power have enabled the numerical simulation of many complex flow problems that are of practical and strategic interest for naval applications. A noticeable area of advancement is the computa…

SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models

2025-08-21 · Peng Ding, Wen Sun, Dailin Li, Wei Zou 외 arxiv

Large Language Models (LLMs) excel at various natural language processing tasks but remain vulnerable to jailbreaking attacks that induce harmful content generation. In this paper, we reveal a critical safety inconsisten…

Reinforcement Learning