Reinforcement Learning (RL)
2개 벤치마크 · 논문 15,112편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Continuous control with deep reinforcement learning
Playing Atari with Deep Reinforcement Learning
Deep Reinforcement Learning with Double Q-learning
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
Papers
Reinforcement Learning Enabled Adaptive Multi-Task Control for Bipedal Soccer Robots
Developing bipedal football robots in dynamiccombat environments presents challenges related to motionstability and deep coupling of multiple tasks, as well ascontrol switching issues between different states such as up-…
Reinforcement Learning (RL)Aligning Humans and Robots via Reinforcement Learning from Implicit Human Feedback
Conventional reinforcement learning (RL) ap proaches often struggle to learn effective policies under sparse reward conditions, necessitating the manual design of complex, task-specific reward functions. To address this …
EEGMuJoCoreinforcement-learningReinforcement Learning+1VAR-MATH: Probing True Mathematical Reasoning in Large Language Models via Symbolic Multi-Instance Benchmarks
Recent advances in reinforcement learning (RL) have led to substantial improvements in the mathematical reasoning abilities of large language models (LLMs), as measured by standard benchmarks. However, these gains often …
MathMathematical ReasoningReinforcement Learning (RL)QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation
Reinforcement learning (RL) has become a key component in training large language reasoning models (LLMs). However, recent studies questions its effectiveness in improving multi-step reasoning-particularly on hard proble…
MathReinforcement Learning (RL)Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
In the era of Large Language Models (LLMs), alignment has emerged as a fundamental yet challenging problem in the pursuit of more reliable, controllable, and capable machine intelligence. The recent success of reasoning …
Language ModelingLanguage ModellingLarge Language ModelReinforcement Learning (RL)Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
Behavior Cloning (BC) on curated (or filtered) data is the predominant paradigm for supervised fine-tuning (SFT) of large language models; as well as for imitation learning of control policies. Here, we draw on a connect…
continuous-controlContinuous ControlImitation LearningReinforcement Learning (RL)