paper-with-me

홈 › Papers

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach

2025-07-29 · Fnu Hairi, Jiao Yang, Tianchen Zhou, Haibo Yang, Chaosheng Dong, Fan Yang, Michinari Momma, Yan Gao, Jia Liu arxiv

In many multi-objective reinforcement learning (MORL) applications, being able to systematically explore the Pareto-stationary solutions under multiple non-convex reward objectives with theoretical finite-time sample complexity guarantee is an important and yet under-explored problem. This motivates us to take the first step and fill the important gap in MORL. Specifically, in this paper, we propose a \uline{M}ulti-\uline{O}bjective weighted-\uline{CH}ebyshev \uline{A}ctor-critic (MOCHA) algorithm for MORL, which judiciously integrates the weighted-Chebychev (WC) and actor-critic framework to enable Pareto-stationarity exploration systematically with finite-time sample complexity guarantee. Sample complexity result of MOCHA algorithm reveals an interesting dependency on $p_{\min}$ in finding an $ε$-Pareto-stationary solution, where $p_{\min}$ denotes the minimum entry of a given weight vector $\mathbf{p}$ in WC-scarlarization. By carefully choosing learning rates, the sample complexity for each exploration can be $\tilde{\mathcal{O}}(ε^{-2})$. Furthermore, simulation studies on a large KuaiRand offline dataset, show that the performance of MOCHA algorithm significantly outperforms other baseline MORL approaches.

📄 PDF Abstract BibTeX arXiv:2507.21397

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Accelerating Multi-Objective Bayesian Optimisation via Predictive-Gradient Catalysts

2026-06-05 · Alma Rahat, Tinkle Chugh, Jonathan Fieldsend, Richard Allmendinger arxiv

This paper presents a general acceleration mechanism for multi-objective Bayesian optimisation (MOBO) that leverages Gaussian process predictive gradients as auxiliary signals. Rather than replacing existing Pareto-compl…

Multi-Objective Bilevel Learning

2025-11-11 · Zhiyao Zhang, Zhuqing Liu, Xin Zhang, Wen-Yen Chen 외 arxiv

As machine learning (ML) applications grow increasingly complex in recent years, modern ML frameworks often need to address multiple potentially conflicting objectives with coupled decision variables across different lay…

Optimization on Pareto sets: On a theory of multi-objective optimization

2023-08-04 · Abhishek Roy, Geelon So, Yi-An Ma

In multi-objective optimization, a single decision vector must balance the trade-offs between many objectives. Solutions achieving an optimal trade-off are said to be Pareto optimal: these are decision vectors for which …

Multi-objective Neural Architecture Search via Non-stationary Policy Gradient

2020-01-23 · Zewei Chen, Fengwei Zhou, George Trimponias, Zhenguo Li

Multi-objective Neural Architecture Search (NAS) aims to discover novel architectures in the presence of multiple conflicting objectives. Despite recent progress, the problem of approximating the full Pareto front accura…

Neural Architecture SearchReinforcement LearningReinforcement Learning (RL)

An Adaptive KKT-Based Indicator for Convergence Assessment in Multi-Objective Optimization

2026-03-04 · Thiago Santos, Sebastiao Xavier arxiv

Performance indicators are essential tools for assessing the convergence behavior of multi-objective optimization algorithms, particularly when the true Pareto front is unknown or difficult to approximate. Classical refe…