paper-with-me

General Reinforcement Learning

6개 벤치마크 · 논문 94편 · 이 태스크의 논문 보기 →

Benchmarks

Most implemented

Papers

Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction

2026-07-21 · Jialian Li, Junhong Liu, Yuchen Cao, Weiran Guo 외 arxiv

Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge. As embodied agents become increasingly capable, there is a growing demand for compact mode…

General Reinforcement Learning

A Unified Pair-GRPO Family: From Implicit to Explicit Preference Constraints for Stable and General RL Alignment

2026-05-07 · Hao Yu arxiv

Large language model (LLM) alignment via reinforcement learning from human preferences (RLHF) suffers from unstable policy updates, ambiguous gradient directions, poor interpretability, and high gradient variance in main…

General Reinforcement LearningContinuous Control

Reasoning as Compression: Unifying Budget Forcing via the Conditional Information Bottleneck

2026-03-09 · Fabio Valerio Massoli, Andrey Kuzmin, Arash Behboodi arxiv

\ac{CoT} prompting improves LLM accuracy on complex tasks but often increases token usage and inference cost. Existing ``Budget Forcing'' methods reduce cost via fine-tuning with heuristic length penalties, suppressing b…

General Reinforcement Learning

A Model-Free Universal AI

2026-02-26 · Yegon Kim, Juho Lee arxiv

In general reinforcement learning, all established optimal agents, including AIXI, are model-based, explicitly maintaining and using environment models. This paper introduces Universal AI with Q-Induction (AIQI), the fir…

General Reinforcement Learning

FaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization

2026-02-03 · Runquan Gui, Yafu Li, Xiaoye Qu, Ziyan Liu 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has markedly improved the performance of Large Language Models (LLMs) on tasks requiring multi-step reasoning. However, most RLVR pipelines rely on sparse outcome-bas…

General Reinforcement Learning

DARL: Encouraging Diverse Answers for General Reasoning without Verifiers

2026-01-21 · Chongxuan Huang, Lei Lin, Xiaodong Shi, Wenping Hu 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has demonstrated promising gains in enhancing the reasoning capabilities of large language models. However, its dependence on domain-specific verifiers significantly …

General Reinforcement Learning

전체 94편 보기 →