paper-with-me

Papers

Controllability in preference-conditioned multi-objective reinforcement learning

2026-05-11 · Pau de las Heras Molins, Beyazit Yalcinkaya, Lasse Peters, David Fridovich-Keil, Georgios Bakirtzis arxiv

Multi-objective reinforcement learning (MORL) allows a user to express preference over outcomes in terms of the relative importance of the objectives, but standard metrics cannot capture whether changes in preference reliably change the agent's behavior in the intended way, a property termed controllability. As a result, preference-conditioned agents can score well on standard MORL metrics while being insensitive to the preference input. If the ability to control agents cannot be reliably assessed, the symbolic interface that MORL provides between user intent and agent behavior is broken. Mainstream MORL metrics alone fail to measure the controllability of preference-conditioned agents, motivating a complementary metric specifically designed to that end. We hope the results spur discussion in the community on existing evaluation protocols to consolidate advances in preference adaptation in MORL to larger and more complex problems.

📄 PDF Abstract BibTeX arXiv:2605.10585

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

One Model for All: Multi-Objective Controllable Language Models

2026-04-06 · Qiang He, Yucheng Yang, Tianyi Zhou, Meng Fang 외 arxiv

Aligning large language models (LLMs) with human preferences is critical for enhancing LLMs' safety, helpfulness, humor, faithfulness, etc. Current reinforcement learning from human feedback (RLHF) mainly focuses on a fi…

Computational EfficiencyReinforcement Learning

PCHC: Enabling Preference Conditioned Humanoid Control via Multi-Objective Reinforcement Learning

2026-03-25 · Huanyu Li, Dewei Wang, Xinmiao Wang, Xinzhe Liu 외 arxiv

Humanoid robots often need to balance competing objectives, such as maximizing speed while minimizing energy consumption. While current reinforcement learning (RL) methods can master complex skills like fall recovery and…

Reinforcement Learning

Regularized Conditional Diffusion Model for Multi-Task Preference Alignment

2024-04-07 · Xudong Yu, Chenjia Bai, Haoran He, Changhong Wang 외

Sequential decision-making is desired to align with human intents and exhibit versatility across various tasks. Previous methods formulate it as a conditional generation process, utilizing return-conditioned diffusion mo…

D4RLDecision MakingSequential Decision Making

Emotional Preferences as Goal-Priority Regulation

2026-08-27 · Shiqi Liu, Yihua Tan, Hu Fu, Guanyu Qi arxiv

A core question in decision-making for agents is whether the relative priorities of competing lower-level objectives can be determined by emotional preferences autonomously generated by higher-level goals, rather than be…

Reinforcement Learning

Hindsight Preference Replay Improves Preference-Conditioned Multi-Objective Reinforcement Learning

2026-01-08 · Jonaid Shianifar, Michael Schukat, Karl Mason arxiv

Multi-objective reinforcement learning (MORL) enables agents to optimize vector-valued rewards while respecting user preferences. CAPQL, a preference-conditioned actor-critic method, achieves this by conditioning on weig…

Reinforcement Learning