paper-with-me

Papers

Preference-Agile Multi-Objective Optimization for Real-time Vehicle Dispatching

2026-04-12 · Jiahuan Jin, Wenhao Zhao, Rong Qu, Jianfeng Ren, Xinan Chen, Qingfu Zhang, Ruibin Bai arxiv

Multi-objective optimization (MOO) has been widely studied in literature because of its versatility in human-centered decision making in real-life applications. Recently, demand for dynamic MOO is fast-emerging due to tough market dynamics that require real-time re-adjustments of priorities for different objectives. However, most existing studies focus either on deterministic MOO problems which are not practical, or non-sequential dynamic MOO decision problems that cannot deal with some real-life complexities. To address these challenges, a preference-agile multi-objective optimization (PAMOO) is proposed in this paper to permit users to dynamically adjust and interactively assign the preferences on the fly. To achieve this, a novel uniform model within a deep reinforcement learning (DRL) framework is proposed that can take as inputs users' dynamic preference vectors explicitly. Additionally, a calibration function is fitted to ensure high quality alignment between the preference vector inputs and the output DRL decision policy. Extensive experiments on challenging real-life vehicle dispatching problems at a container terminal showed that PAMOO obtains superior performance and generalization ability when compared with two most popular MOO methods. Our method presents the first dynamic MOO method for challenging \rev{dynamic sequential MOO decision problems

📄 PDF Abstract BibTeX arXiv:2604.10664

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningDecision Making

Similar Papers 제목 키워드 기반

TUR-DPO: Topology- and Uncertainty-Aware Direct Preference Optimization

2026-04-30 · Abdulhady Abas Abdullah, Fatemeh Daneshfar, Seyedali Mirjalili, Mourad Oussalah arxiv

Aligning large language models (LLMs) with human preferences is commonly done via reinforcement learning from human feedback (RLHF) with Proximal Policy Optimization (PPO) or, more simply, via Direct Preference Optimizat…

Reinforcement LearningMathematical ReasoningQuestion Answering

Automatic Preference Based Multi-objective Evolutionary Algorithm on Vehicle Fleet Maintenance Scheduling Optimization

2021-01-23 · Yali Wang, Steffen Limmer, Markus Olhofer, Michael Emmerich 외

A preference based multi-objective evolutionary algorithm is proposed for generating solutions in an automatically detected knee point region. It is named Automatic Preference based DI-MOEA (AP-DI-MOEA) where DI-MOEA sta…

DiversityScheduling

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models

2025-05-16 · Akhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng Wen

Post-training of LLMs with RLHF, and subsequently preference optimization algorithms such as DPO, IPO, etc., made a big difference in improving human alignment. However, all such techniques can only work with a single (h…

Active Preference Learning over Latent Preference Archetypes for Many-Objective Bayesian Optimization

2026-03-30 · Manisha Dubey, Sebastiaan De Peuter, Wanrong Wang, Samuel Kaski arxiv

Preference-based many-objective Bayesian optimization typically assumes that all pairwise comparisons arise from a single latent utility function, despite real decision makers often exhibiting multiple latent preference …

Efficient First-Order Optimization on the Pareto Set for Multi-Objective Learning under Preference Guidance

2025-03-26 · Lisha Chen, Quan Xiao, Ellen Hidemi Fukuda, Xinyi Chen 외

Multi-objective learning under user-specified preference is common in real-world problems such as multi-lingual speech recognition under fairness. In this work, we frame such a problem as a semivectorial bilevel optimiza…

Bilevel OptimizationFairnessspeech-recognitionSpeech Recognition