paper-with-me

Papers

Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts

2026-03-24 · Xianwei Cao, Dou Quan, Zhenliang Zhang, Shuang Wang arxiv

Humans often juggle multiple, sometimes conflicting objectives and shift their priorities as circumstances change, rather than following a fixed objective function. In contrast, most computational decision-making and multi-objective RL methods assume static preference weights or a known scalar reward. In this work, we study sequential decision-making problem when these preference weights are unobserved latent variables that drift with context. Specifically, we propose Dynamic Preference Inference (DPI), a cognitively inspired framework in which an agent maintains a probabilistic belief over preference weights, updates this belief from recent interaction, and conditions its policy on inferred preferences. We instantiate DPI as a variational preference inference module trained jointly with a preference-conditioned actor-critic, using vector-valued returns as evidence about latent trade-offs. In queueing, maze, and multi-objective continuous-control environments with event-driven changes in objectives, DPI adapts its inferred preferences to new regimes and achieves higher post-shift performance than fixed-weight and heuristic envelope baselines.

📄 PDF Abstract BibTeX arXiv:2603.22813

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

What Matters to You? Towards Visual Representation Alignment for Robot Learning

2023-10-11 · Ran Tian, Chenfeng Xu, Masayoshi Tomizuka, Jitendra Malik 외

When operating in service of people, robots need to optimize rewards aligned with end-user preferences. Since robots will rely on raw perceptual inputs like RGB images, their rewards will inevitably use visual representa…

Zero-shot Generalization

Silence Routing: When Not Speaking Improves Collective Judgment

2026-02-09 · Itsuki Fujisaki, Kunhao Yang arxiv

The wisdom of crowds has been shown to operate not only for factual judgments but also in matters of taste, where accuracy is defined relative to an individual's preferences. However, it remains unclear how different typ…

Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward Modeling

2026-04-07 · Qiyuan Chen, Hongsen Huang, Jiahe Chen, Qian Shao 외 arxiv

Vision-language reward modeling faces a dilemma: generative approaches are interpretable but slow, while discriminative ones are efficient but act as opaque "black boxes." To bridge this gap, we propose VL-MDR (Vision-La…

Doing the right thing (or not) in a lemons-like situation: on the role of social preferences and Kantian moral concerns

2024-05-21 · Ingela Alger, José Ignacio Rivero-Wildemauwe

We conduct a laboratory experiment using framing to assess the willingness to ``sell a lemon'', i.e., to undertake an action that benefits self but hurts the other (the ``buyer''). We seek to disentangle the role of othe…

What Matters in Data for DPO?

2025-08-23 · Yu Pan, Zhongze Cai, Guanting Chen, Huaiyang Zhong 외 arxiv

Direct Preference Optimization (DPO) has emerged as a simple and effective approach for aligning large language models (LLMs) with human preferences, bypassing the need for a learned reward model. Despite its growing ado…