paper-with-me

Papers General Reinforcement Learning

“General Reinforcement Learning” 태그가 달린 논문 94편 · 필터 해제

Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction

2026-07-21 · Jialian Li, Junhong Liu, Yuchen Cao, Weiran Guo 외 arxiv

Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge. As embodied agents become increasingly capable, there is a growing demand for compact mode…

General Reinforcement Learning

A Unified Pair-GRPO Family: From Implicit to Explicit Preference Constraints for Stable and General RL Alignment

2026-05-07 · Hao Yu arxiv

Large language model (LLM) alignment via reinforcement learning from human preferences (RLHF) suffers from unstable policy updates, ambiguous gradient directions, poor interpretability, and high gradient variance in main…

General Reinforcement LearningContinuous Control

Reasoning as Compression: Unifying Budget Forcing via the Conditional Information Bottleneck

2026-03-09 · Fabio Valerio Massoli, Andrey Kuzmin, Arash Behboodi arxiv

\ac{CoT} prompting improves LLM accuracy on complex tasks but often increases token usage and inference cost. Existing ``Budget Forcing'' methods reduce cost via fine-tuning with heuristic length penalties, suppressing b…

General Reinforcement Learning

A Model-Free Universal AI

2026-02-26 · Yegon Kim, Juho Lee arxiv

In general reinforcement learning, all established optimal agents, including AIXI, are model-based, explicitly maintaining and using environment models. This paper introduces Universal AI with Q-Induction (AIQI), the fir…

General Reinforcement Learning

FaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization

2026-02-03 · Runquan Gui, Yafu Li, Xiaoye Qu, Ziyan Liu 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has markedly improved the performance of Large Language Models (LLMs) on tasks requiring multi-step reasoning. However, most RLVR pipelines rely on sparse outcome-bas…

General Reinforcement Learning

DARL: Encouraging Diverse Answers for General Reasoning without Verifiers

2026-01-21 · Chongxuan Huang, Lei Lin, Xiaodong Shi, Wenping Hu 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has demonstrated promising gains in enhancing the reasoning capabilities of large language models. However, its dependence on domain-specific verifiers significantly …

General Reinforcement Learning

TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning

2025-09-30 · Zhepei Wei, Xiao Yang, Kai Sun, Jiaqi Wang 외 arxiv

While large language models (LLMs) have demonstrated strong performance on factoid question answering, they are still prone to hallucination and untruthful responses, particularly when tasks demand information outside th…

General Reinforcement LearningQuestion Answering

Deterministic Policy Gradient for Reinforcement Learning with Continuous Time and State

2025-09-28 · Ziheng Cheng, Xin Guo, Yufei Zhang arxiv

The theory of continuous-time reinforcement learning (RL) has progressed rapidly in recent years. While the ultimate objective of RL is typically to learn deterministic control policies, most existing continuous-time RL …

General Reinforcement Learning

UR$^2$: Unify RAG and Reasoning through Reinforcement Learning

2025-08-08 · Weitao Li, Boran Xiang, Xiaolong Wang, Zhinan Gou 외 arxiv

Large Language Models (LLMs) have shown strong capabilities through two complementary paradigms: Retrieval-Augmented Generation (RAG) for knowledge grounding and Reinforcement Learning from Verifiable Rewards (RLVR) for …

General Reinforcement LearningMathematical Reasoning

Benchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains

2025-07-31 · Ruo Yu Tao, Kaicheng Guo, Cameron Allen, George Konidaris arxiv

Mitigating partial observability is a necessary but challenging task for general reinforcement learning algorithms. To improve an algorithm's ability to mitigate partial observability, researchers need comprehensive benc…

General Reinforcement Learning

PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning

2025-06-17 · Yizhen Zhang, Yang Ding, Shuoshuo Zhang, Xinchen Zhang 외

Inspired by the impressive reasoning capabilities demonstrated by reinforcement learning approaches like DeepSeek-R1, recent emerging research has begun exploring the use of reinforcement learning (RL) to enhance vision-…

General Reinforcement LearningMultimodal Reasoningreinforcement-learningReinforcement Learning+2

NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning

2025-05-21 · Wei Liu, Siya Qi, Xinyu Wang, Chen Qian 외

Recent advances such as DeepSeek R1-Zero highlight the effectiveness of incentive training, a reinforcement learning paradigm that computes rewards solely based on the final answer part of a language model's output, ther…

General Reinforcement LearningLogical ReasoningOpen-Domain Question Answeringreinforcement-learning+1

High-order Regularization for Machine Learning and Learning-based Control

2025-05-13 · Xinghua Liu, Ming Cao

The paper proposes a novel regularization procedure for machine learning. The proposed high-order regularization (HR) provides new insight into regularization, which is widely used to train a neural network that can be u…

General Reinforcement Learning

Towards More Efficient, Robust, Instance-adaptive, and Generalizable Sequential Decision making

2025-04-12 · Zhiyong Wang

The primary goal of my Ph.D. study is to develop provably efficient and practical algorithms for data-driven sequential decision-making under uncertainty. My work focuses on reinforcement learning (RL), multi-armed bandi…

Decision MakingDecision Making Under UncertaintyGeneral Reinforcement LearningMulti-Armed Bandits+5

Rec-R1: Bridging Generative Large Language Models and User-Centric Recommendation Systems via Reinforcement Learning

2025-03-31 · Jiacheng Lin, Tian Wang, Kun Qian

We propose Rec-R1, a general reinforcement learning framework that bridges large language models (LLMs) with recommendation systems through closed-loop optimization. Unlike prompting and supervised fine-tuning (SFT), Rec…

General Reinforcement LearningInstruction FollowingRecommendation SystemsSequential Recommendation

The Problem of Social Cost in Multi-Agent General Reinforcement Learning: Survey and Synthesis

2024-12-03 · Kee Siong Ng, Samuel Yang-Zhao, Timothy Cadogan-Cowper

The AI safety literature is full of examples of powerful AI agents that, in blindly pursuing a specific and usually narrow objective, ends up with unacceptable and even catastrophic collateral damage to others. In this p…

General Reinforcement LearningMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning

Hypercube Policy Regularization Framework for Offline Reinforcement Learning

2024-11-07 · Yi Shen, Hanyan Huang

Offline reinforcement learning has received extensive attention from scholars because it avoids the interaction between the agent and the environment by learning a policy through a static dataset. However, general reinfo…

D4RLGeneral Reinforcement Learningreinforcement-learningReinforcement Learning

Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks

2024-10-30 · Michael Matthews, Michael Beukman, Chris Lu, Jakob Foerster

While large models trained with self-supervised learning on offline datasets have shown remarkable capabilities in text and image domains, achieving the same generalisation for agents that act in sequential decision prob…

General Reinforcement LearningReinforcement Learning (RL)Self-Supervised Learning

Reinforcement Learning: Tutorial and Survey

2024-07-18 · OSF Preprints 2024 7 · Benyamin Ghojogh, Ali Ghodsi

This is a tutorial and survey paper on reinforcement learning, from fundamental reinforcement learning to deep reinforcement learning. It starts with introducing the elements of reinforcement learning. Then, Markov decis…

Deep Reinforcement LearningGeneral Reinforcement LearningQ-Learningreinforcement-learning+3

Dynamic Knowledge Injection for AIXI Agents

2023-12-18 · Samuel Yang-Zhao, Kee Siong Ng, Marcus Hutter

Prior approximations of AIXI, a Bayesian optimality notion for general reinforcement learning, can only approximate AIXI's Bayesian environment model using an a-priori defined set of models. This is a fundamental source …

General Reinforcement Learning
1–20 / 94 다음 →