Papers General Reinforcement Learning
“General Reinforcement Learning” 태그가 달린 논문 94편 · 필터 해제
Dropout Strategy in Reinforcement Learning: Limiting the Surrogate Objective Variance in Policy Optimization Methods
Policy-based reinforcement learning algorithms are widely used in various fields. Among them, mainstream policy optimization algorithms such as TRPO and PPO introduce importance sampling into policy iteration, which allo…
General Reinforcement Learningreinforcement-learningReinforcement LearningReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reinforcement Learning from Human Feedback (RLHF) is key to aligning Large Language Models (LLMs), typically paired with the Proximal Policy Optimization (PPO) algorithm. While PPO is a powerful method designed for gener…
General Reinforcement LearningGPUreinforcement-learningDiscovering General Reinforcement Learning Algorithms with Adversarial Environment Design
The past decade has seen vast progress in deep reinforcement learning (RL) on the back of algorithms manually designed by human researchers. Recently, it has been shown that it is possible to meta-learn update rules, wit…
Deep Reinforcement LearningGeneral Reinforcement Learningreinforcement-learningReinforcement Learning+1Image Transformation Sequence Retrieval with General Reinforcement Learning
In this work, the novel Image Transformation Sequence Retrieval (ITSR) task is presented, in which a model must retrieve the sequence of transformations between two given images that act as source and target, respectivel…
General Reinforcement LearningModel-based Reinforcement Learningreinforcement-learningReinforcement Learning+1L-SA: Learning Under-Explored Targets in Multi-Target Reinforcement Learning
Tasks that involve interaction with various targets are called multi-target tasks. When applying general reinforcement learning approaches for such tasks, certain targets that are difficult to access or interact with may…
General Reinforcement Learningreinforcement-learningVisual NavigationComputably Continuous Reinforcement-Learning Objectives are PAC-learnable
In reinforcement learning, the classic objectives of maximizing discounted and finite-horizon cumulative rewards are PAC-learnable: There are algorithms that learn a near-optimal policy with high probability using a fini…
General Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Policy Mirror Descent Inherently Explores Action Space
Explicit exploration in the action space was assumed to be indispensable for online policy gradient methods to avoid a drastic degradation in sample complexity, for solving general reinforcement learning problems over fi…
Efficient ExplorationGeneral Reinforcement LearningPolicy Gradient MethodsLearning to Backdoor Federated Learning
In a federated learning (FL) system, malicious participants can easily embed backdoors into the aggregated model while maintaining the model's performance on the main task. To this end, various defenses, including traini…
Backdoor AttackFederated LearningGeneral Reinforcement LearningComputational Dualism and Objective Superintelligence
The concept of intelligent software is flawed. The behaviour of software is determined by the hardware that "interprets" it. This undermines claims regarding the behaviour of theorised, software superintelligence. Here w…
General Reinforcement LearningAccuracy-Guaranteed Collaborative DNN Inference in Industrial IoT via Deep Reinforcement Learning
Collaboration among industrial Internet of Things (IoT) devices and edge networks is essential to support computation-intensive deep neural network (DNN) inference services which require low delay and high accuracy. Samp…
Deep Reinforcement LearningEdge-computingGeneral Reinforcement Learningreinforcement-learning+1AcceRL: Policy Acceleration Framework for Deep Reinforcement Learning
Deep reinforcement learning has achieved great success in various fields with its super decision-making ability. However, the policy learning process requires a large amount of training time, causing energy consumption. …
Decision MakingDeep Reinforcement LearningGeneral Reinforcement LearningNeural Network Compression+3DeFIX: Detecting and Fixing Failure Scenarios with Reinforcement Learning in Imitation Learning Based Autonomous Driving
Safely navigating through an urban environment without violating any traffic rules is a crucial performance target for reliable autonomous driving. In this paper, we present a Reinforcement Learning (RL) based methodolog…
Autonomous DrivingCARLA MAP LeaderboardGeneral Reinforcement LearningImitation Learning+1Intelligent Resource Allocation in Joint Radar-Communication With Graph Neural Networks
Autonomous vehicles produce high data rates of sensory information from sensing systems. To achieve the advantages of sensor fusion among different vehicles in a cooperative driving scenario, high data-rate communication…
Autonomous DrivingAutonomous VehiclesDeep Reinforcement LearningDistributional Reinforcement Learning+4Learning Deformable Object Manipulation from Expert Demonstrations
We present a novel Learning from Demonstration (LfD) method, Deformable Manipulation from Demonstrations (DMfD), to solve deformable manipulation tasks using states or images as inputs, given expert demonstrations. Our m…
Deformable Object ManipulationGeneral Reinforcement LearningObjectComputable Artificial General Intelligence
Artificial general intelligence (AGI) may herald our extinction, according to AI safety research. Yet claims regarding AGI must rely upon mathematical formalisms -- theoretical agents we may analyse or attempt to build. …
General Reinforcement LearningPhilosophyD3PG: Dirichlet DDPG for Task Partitioning and Offloading With Constrained Hybrid Action Space in Mobile-Edge Computing
Mobile-edge computing (MEC) has been regarded as a promising paradigm to reduce service latency for data processing in the Internet of Things (IoT) by provisioning computing resources at the network edges. In this work, …
Deep Reinforcement LearningEdge-computingGeneral Reinforcement LearningMultiobjective Optimization+2Doubly-Robust Estimation for Correcting Position-Bias in Click Feedback for Unbiased Learning to Rank
Clicks on rankings suffer from position-bias: generally items on lower ranks are less likely to be examined - and thus clicked - by users, in spite of their actual preferences between items. The prevalent approach to unb…
counterfactualGeneral Reinforcement LearningLearning-To-RankPositionAbstractions of General Reinforcement Learning
The field of artificial intelligence (AI) is devoted to the creation of artificial decision-makers that can perform (at least) on par with the human counterparts on a domain of interest. Unlike the agents in traditional …
General Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Reducing Planning Complexity of General Reinforcement Learning with Non-Markovian Abstractions
The field of General Reinforcement Learning (GRL) formulates the problem of sequential decision-making from ground up. The history of interaction constitutes a "ground" state of the system, which never repeats. On the on…
Decision MakingGeneral Reinforcement Learningreinforcement-learningReinforcement Learning (RL)+1Intelligent Trading Systems: A Sentiment-Aware Reinforcement Learning Approach
The feasibility of making profitable trades on a single asset on stock exchanges based on patterns identification has long attracted researchers. Reinforcement Learning (RL) and Natural Language Processing have gained no…
Algorithmic TradingGeneral Reinforcement Learningreinforcement-learningReinforcement Learning+4