paper-with-me

홈 › Papers

Counteractive RL: Rethinking Core Principles for Efficient and Scalable Deep Reinforcement Learning

2026-03-16 · Ezgi Korkmaz arxiv

Following the pivotal success of learning strategies to win at tasks, solely by interacting with an environment without any supervision, agents have gained the ability to make sequential decisions in complex MDPs. Yet, reinforcement learning policies face exponentially growing state spaces in high dimensional MDPs resulting in a dichotomy between computational complexity and policy success. In our paper we focus on the agent's interaction with the environment in a high-dimensional MDP during the learning phase and we introduce a theoretically-founded novel paradigm based on experiences obtained through counteractive actions. Our analysis and method provide a theoretical basis for efficient, effective, scalable and accelerated learning, and further comes with zero additional computational complexity while leading to significant acceleration in training. We conduct extensive experiments in the Arcade Learning Environment with high-dimensional state representation MDPs. The experimental results further verify our theoretical analysis, and our method achieves significant performance increase with substantial sample-efficiency in high-dimensional environments.

📄 PDF Abstract BibTeX arXiv:2603.15871

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Position Paper: Rethinking Privacy in RL for Sequential Decision-making in the Age of LLMs

2025-04-15 · Flint Xiaofeng Fan, Cheston Tan, Roger Wattenhofer, Yew-Soon Ong

The rise of reinforcement learning (RL) in critical real-world applications demands a fundamental rethinking of privacy in AI systems. Traditional privacy frameworks, designed to protect isolated data points, fall short …

Autonomous VehiclesDecision MakingPositionReinforcement Learning (RL)+1

Reasoning Core: A Scalable RL Environment for LLM Symbolic Reasoning

2025-09-22 · Valentin Lacombe, Valentin Quesnel, Damien Sileo arxiv

We introduce Reasoning Core, a new scalable environment for Reinforcement Learning with Verifiable Rewards (RLVR), designed to advance foundational symbolic reasoning in Large Language Models (LLMs). Unlike existing benc…

Reinforcement Learning

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?

2025-10-12 · Zihan Chen, Yiming Zhang, Hengguang Zhou, Zenghui Ding 외 arxiv

Current benchmarks are inadequate for evaluating progress in reinforcement learning (RL) for large language models (LLMs).Despite recent benchmark gains reported for RL, we find that training on these benchmarks' trainin…

Reinforcement Learning

ElegantRL-Podracer: Scalable and Elastic Library for Cloud-Native Deep Reinforcement Learning

2021-12-11 · Xiao-Yang Liu, Zechu Li, Zhuoran Yang, Jiahao Zheng 외

Deep reinforcement learning (DRL) has revolutionized learning and actuation in applications such as game playing and robotic control. The cost of data collection, i.e., generating transitions from agent-environment inter…

Deep Reinforcement LearningGPUreinforcement-learningReinforcement Learning+3

Collaborative Disagreement Resolution for Scalable Oversight

2026-06-02 · Yuyang Jiang, Chacha Chen, Teng Wu, Liwen Sun 외 arxiv

Debate, where AI agents argue opposing positions, has emerged as a key approach to scalable oversight. However, debate faces a fundamental tension: models are incentivized to be persuasive to the judge, which may not alw…