Papers General Reinforcement Learning
“General Reinforcement Learning” 태그가 달린 논문 94편 · 필터 해제
Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction
Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge. As embodied agents become increasingly capable, there is a growing demand for compact mode…
General Reinforcement LearningA Unified Pair-GRPO Family: From Implicit to Explicit Preference Constraints for Stable and General RL Alignment
Large language model (LLM) alignment via reinforcement learning from human preferences (RLHF) suffers from unstable policy updates, ambiguous gradient directions, poor interpretability, and high gradient variance in main…
General Reinforcement LearningContinuous ControlReasoning as Compression: Unifying Budget Forcing via the Conditional Information Bottleneck
\ac{CoT} prompting improves LLM accuracy on complex tasks but often increases token usage and inference cost. Existing ``Budget Forcing'' methods reduce cost via fine-tuning with heuristic length penalties, suppressing b…
General Reinforcement LearningA Model-Free Universal AI
In general reinforcement learning, all established optimal agents, including AIXI, are model-based, explicitly maintaining and using environment models. This paper introduces Universal AI with Q-Induction (AIQI), the fir…
General Reinforcement LearningFaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization
Reinforcement Learning with Verifiable Rewards (RLVR) has markedly improved the performance of Large Language Models (LLMs) on tasks requiring multi-step reasoning. However, most RLVR pipelines rely on sparse outcome-bas…
General Reinforcement LearningDARL: Encouraging Diverse Answers for General Reasoning without Verifiers
Reinforcement Learning with Verifiable Rewards (RLVR) has demonstrated promising gains in enhancing the reasoning capabilities of large language models. However, its dependence on domain-specific verifiers significantly …
General Reinforcement LearningTruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
While large language models (LLMs) have demonstrated strong performance on factoid question answering, they are still prone to hallucination and untruthful responses, particularly when tasks demand information outside th…
General Reinforcement LearningQuestion AnsweringDeterministic Policy Gradient for Reinforcement Learning with Continuous Time and State
The theory of continuous-time reinforcement learning (RL) has progressed rapidly in recent years. While the ultimate objective of RL is typically to learn deterministic control policies, most existing continuous-time RL …
General Reinforcement LearningUR$^2$: Unify RAG and Reasoning through Reinforcement Learning
Large Language Models (LLMs) have shown strong capabilities through two complementary paradigms: Retrieval-Augmented Generation (RAG) for knowledge grounding and Reinforcement Learning from Verifiable Rewards (RLVR) for …
General Reinforcement LearningMathematical ReasoningBenchmarking Partial Observability in Reinforcement Learning with a Suite of Memory-Improvable Domains
Mitigating partial observability is a necessary but challenging task for general reinforcement learning algorithms. To improve an algorithm's ability to mitigate partial observability, researchers need comprehensive benc…
General Reinforcement LearningPeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
Inspired by the impressive reasoning capabilities demonstrated by reinforcement learning approaches like DeepSeek-R1, recent emerging research has begun exploring the use of reinforcement learning (RL) to enhance vision-…
General Reinforcement LearningMultimodal Reasoningreinforcement-learningReinforcement Learning+2NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning
Recent advances such as DeepSeek R1-Zero highlight the effectiveness of incentive training, a reinforcement learning paradigm that computes rewards solely based on the final answer part of a language model's output, ther…
General Reinforcement LearningLogical ReasoningOpen-Domain Question Answeringreinforcement-learning+1High-order Regularization for Machine Learning and Learning-based Control
The paper proposes a novel regularization procedure for machine learning. The proposed high-order regularization (HR) provides new insight into regularization, which is widely used to train a neural network that can be u…
General Reinforcement LearningTowards More Efficient, Robust, Instance-adaptive, and Generalizable Sequential Decision making
The primary goal of my Ph.D. study is to develop provably efficient and practical algorithms for data-driven sequential decision-making under uncertainty. My work focuses on reinforcement learning (RL), multi-armed bandi…
Decision MakingDecision Making Under UncertaintyGeneral Reinforcement LearningMulti-Armed Bandits+5Rec-R1: Bridging Generative Large Language Models and User-Centric Recommendation Systems via Reinforcement Learning
We propose Rec-R1, a general reinforcement learning framework that bridges large language models (LLMs) with recommendation systems through closed-loop optimization. Unlike prompting and supervised fine-tuning (SFT), Rec…
General Reinforcement LearningInstruction FollowingRecommendation SystemsSequential RecommendationThe Problem of Social Cost in Multi-Agent General Reinforcement Learning: Survey and Synthesis
The AI safety literature is full of examples of powerful AI agents that, in blindly pursuing a specific and usually narrow objective, ends up with unacceptable and even catastrophic collateral damage to others. In this p…
General Reinforcement LearningMulti-agent Reinforcement Learningreinforcement-learningReinforcement LearningHypercube Policy Regularization Framework for Offline Reinforcement Learning
Offline reinforcement learning has received extensive attention from scholars because it avoids the interaction between the agent and the environment by learning a policy through a static dataset. However, general reinfo…
D4RLGeneral Reinforcement Learningreinforcement-learningReinforcement LearningKinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks
While large models trained with self-supervised learning on offline datasets have shown remarkable capabilities in text and image domains, achieving the same generalisation for agents that act in sequential decision prob…
General Reinforcement LearningReinforcement Learning (RL)Self-Supervised LearningReinforcement Learning: Tutorial and Survey
This is a tutorial and survey paper on reinforcement learning, from fundamental reinforcement learning to deep reinforcement learning. It starts with introducing the elements of reinforcement learning. Then, Markov decis…
Deep Reinforcement LearningGeneral Reinforcement LearningQ-Learningreinforcement-learning+3Dynamic Knowledge Injection for AIXI Agents
Prior approximations of AIXI, a Bayesian optimality notion for general reinforcement learning, can only approximate AIXI's Bayesian environment model using an a-priori defined set of models. This is a fundamental source …
General Reinforcement Learning