paper-with-me

Papers

Dynamic Preference Multi-Objective Reinforcement Learning for Internet Network Management

2025-06-16 · DongNyeong Heo, Daniela Noemi Rim, Heeyoul Choi

An internet network service provider manages its network with multiple objectives, such as high quality of service (QoS) and minimum computing resource usage. To achieve these objectives, a reinforcement learning-based (RL) algorithm has been proposed to train its network management agent. Usually, their algorithms optimize their agents with respect to a single static reward formulation consisting of multiple objectives with fixed importance factors, which we call preferences. However, in practice, the preference could vary according to network status, external concerns and so on. For example, when a server shuts down and it can cause other servers' traffic overloads leading to additional shutdowns, it is plausible to reduce the preference of QoS while increasing the preference of minimum computing resource usages. In this paper, we propose new RL-based network management agents that can select actions based on both states and preferences. With our proposed approach, we expect a single agent to generalize on various states and preferences. Furthermore, we propose a numerical method that can estimate the distribution of preference that is advantageous for unbiased training. Our experiment results show that the RL agents trained based on our proposed approach significantly generalize better with various preferences than the previous RL approaches, which assume static preference during training. Moreover, we demonstrate several analyses that show the advantages of our numerical estimation method.

📄 PDF Abstract BibTeX arXiv:2506.13153

Code (0)

등록된 구현이 없습니다.

Tasks

ManagementMulti-Objective Reinforcement Learning

Similar Papers 제목 키워드 기반

Inferring Preferences from Demonstrations in Multi-objective Reinforcement Learning: A Dynamic Weight-based Approach

2023-04-27 · Junlin Lu, Patrick Mannion, Karl Mason

Many decision-making problems feature multiple objectives. In such problems, it is not always possible to know the preferences of a decision-maker for different objectives. However, it is often possible to observe the be…

Decision MakingMulti-Objective Reinforcement Learning

PD-MORL: Preference-Driven Multi-Objective Reinforcement Learning Algorithm

2022-08-16 · Toygun Basaklar, Suat Gumussoy, Umit Y. Ogras

Multi-objective reinforcement learning (MORL) approaches have emerged to tackle many real-world problems with multiple conflicting objectives by maximizing a joint objective function weighted by a preference vector. Thes…

continuous-controlContinuous ControlMulti-Objective Reinforcement Learningreinforcement-learning+2

Multi-Preference Lambda-weighted Listwise DPO for Dynamic Preference Alignment

2025-06-24 · Yuhui Sun, Xiyao Wang, Zixi Li, Jinman Zhao

While large-scale unsupervised language models (LMs) capture broad world knowledge and reasoning capabilities, steering their behavior toward desired objectives remains challenging due to the lack of explicit supervision…

Informativenessreinforcement-learningReinforcement LearningWorld Knowledge

Multi-Objective Reinforcement Learning for Adaptive Personalized Autonomous Driving

2025-05-08 · Hendrik Surmann, Jorge de Heuvel, Maren Bennewitz

Human drivers exhibit individual preferences regarding driving style. Adapting autonomous vehicles to these preferences is essential for user trust and satisfaction. However, existing end-to-end driving approaches often …

Autonomous DrivingAutonomous VehiclesCollision AvoidanceMulti-Objective Reinforcement Learning+2

gTLO: A Generalized and Non-linear Multi-Objective Deep Reinforcement Learning Approach

2022-04-11 · Johannes Dornheim

In real-world decision optimization, often multiple competing objectives must be taken into account. Following classical reinforcement learning, these objectives have to be combined into a single reward function. In cont…

Deep Reinforcement LearningDeep-Sea Treasure, Image versionMulti-Objective Reinforcement Learningreinforcement-learning+2