General Reinforcement Learning
6개 벤치마크 · 논문 94편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
OpenSpiel: A Framework for Reinforcement Learning in Games
Stabilizing Transformers for Reinforcement Learning
Gibson Env: Real-World Perception for Embodied Agents
Action Branching Architectures for Deep Reinforcement Learning
Adaptive Rational Activations to Boost Deep Reinforcement Learning
Papers
Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction
Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge. As embodied agents become increasingly capable, there is a growing demand for compact mode…
General Reinforcement LearningA Unified Pair-GRPO Family: From Implicit to Explicit Preference Constraints for Stable and General RL Alignment
Large language model (LLM) alignment via reinforcement learning from human preferences (RLHF) suffers from unstable policy updates, ambiguous gradient directions, poor interpretability, and high gradient variance in main…
General Reinforcement LearningContinuous ControlReasoning as Compression: Unifying Budget Forcing via the Conditional Information Bottleneck
\ac{CoT} prompting improves LLM accuracy on complex tasks but often increases token usage and inference cost. Existing ``Budget Forcing'' methods reduce cost via fine-tuning with heuristic length penalties, suppressing b…
General Reinforcement LearningA Model-Free Universal AI
In general reinforcement learning, all established optimal agents, including AIXI, are model-based, explicitly maintaining and using environment models. This paper introduces Universal AI with Q-Induction (AIQI), the fir…
General Reinforcement LearningFaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization
Reinforcement Learning with Verifiable Rewards (RLVR) has markedly improved the performance of Large Language Models (LLMs) on tasks requiring multi-step reasoning. However, most RLVR pipelines rely on sparse outcome-bas…
General Reinforcement LearningDARL: Encouraging Diverse Answers for General Reasoning without Verifiers
Reinforcement Learning with Verifiable Rewards (RLVR) has demonstrated promising gains in enhancing the reasoning capabilities of large language models. However, its dependence on domain-specific verifiers significantly …
General Reinforcement Learning