paper-with-me

Reinforcement Learning (RL)

2개 벤치마크 · 논문 15,112편 · 이 태스크의 논문 보기 →

Benchmarks

ProcGen

결과 2개

.

결과 1개

Most implemented

Papers

Reinforcement Learning Enabled Adaptive Multi-Task Control for Bipedal Soccer Robots

2026-04-21 · Yulai Zhang, Yinrong Zhang, Ting Wu, Linqi Ye arxiv

Developing bipedal football robots in dynamiccombat environments presents challenges related to motionstability and deep coupling of multiple tasks, as well ascontrol switching issues between different states such as up-…

Reinforcement Learning (RL)

Aligning Humans and Robots via Reinforcement Learning from Implicit Human Feedback

2025-07-17 · Suzie Kim, Hye-Bin Shin, Seong-Whan Lee

Conventional reinforcement learning (RL) ap proaches often struggle to learn effective policies under sparse reward conditions, necessitating the manual design of complex, task-specific reward functions. To address this …

EEGMuJoCoreinforcement-learningReinforcement Learning+1

VAR-MATH: Probing True Mathematical Reasoning in Large Language Models via Symbolic Multi-Instance Benchmarks

2025-07-17 · Jian Yao, Ran Cheng, Kay Chen Tan

Recent advances in reinforcement learning (RL) have led to substantial improvements in the mathematical reasoning abilities of large language models (LLMs), as measured by standard benchmarks. However, these gains often …

MathMathematical ReasoningReinforcement Learning (RL)

QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation

2025-07-17 · Jiazheng Li, Hong Lu, Kaiyue Wen, Zaiwen Yang 외

Reinforcement learning (RL) has become a key component in training large language reasoning models (LLMs). However, recent studies questions its effectiveness in improving multi-step reasoning-particularly on hard proble…

MathReinforcement Learning (RL)

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities

2025-07-17 · Hao Sun, Mihaela van der Schaar

In the era of Large Language Models (LLMs), alignment has emerged as a fundamental yet challenging problem in the pursuit of more reliable, controllable, and capable machine intelligence. The recent success of reasoning …

Language ModelingLanguage ModellingLarge Language ModelReinforcement Learning (RL)

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)

2025-07-17 · Chongli Qin, Jost Tobias Springenberg

Behavior Cloning (BC) on curated (or filtered) data is the predominant paradigm for supervised fine-tuning (SFT) of large language models; as well as for imitation learning of control policies. Here, we draw on a connect…

continuous-controlContinuous ControlImitation LearningReinforcement Learning (RL)

전체 15,112편 보기 →