paper-with-me

홈 › Papers

Contextual Preference Distribution Learning

2026-03-17 · Benjamin Hudson, Laurent Charlin, Emma Frejinger arxiv

Decision-making problems often feature uncertainty stemming from heterogeneous and context-dependent human preferences. To address this, we propose a sequential learning-and-optimization pipeline to learn preference distributions and leverage them to solve downstream problems, for example risk-averse formulations. We focus on human choice settings that can be formulated as (integer) linear programs. In such settings, existing inverse optimization and choice modelling methods infer preferences from observed choices but typically produce point estimates or fail to capture contextual shifts, making them unsuitable for risk-averse decision-making. Using a bounded-variance score function gradient estimator, we train a predictive model mapping contextual features to a rich class of parameterizable distributions. This approach yields a maximum likelihood estimate. The model generates scenarios for unseen contexts in the subsequent optimization phase. In a synthetic ridesharing environment, our approach reduces average post-decision surprise by up to 114$\times$ compared to a risk-neutral approach with perfect predictions and up to 25$\times$ compared to leading risk-averse baselines.

📄 PDF Abstract BibTeX arXiv:2603.17139

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Vector preference-based contextual bandits under distributional shifts

2025-08-21 · Apurv Shukla, P. R. Kumar arxiv

We consider contextual bandit learning under distribution shift when reward vectors are ordered according to a given preference cone. We propose an adaptive-discretization and optimistic elimination based policy that sel…

Low-Rank Contextual Reinforcement Learning from Heterogeneous Human Feedback

2024-12-27 · Seong Jin Lee, Will Wei Sun, Yufeng Liu

Reinforcement learning from human feedback (RLHF) has become a cornerstone for aligning large language models with human preferences. However, the heterogeneity of human feedback, driven by diverse individual contexts an…

Computational Efficiencyreinforcement-learningReinforcement Learning

Aligning Crowd Feedback via Distributional Preference Reward Modeling

2024-02-15 · Dexun Li, Cong Zhang, Kuicai Dong, Derrick Goh Xin Deik 외

Deep Reinforcement Learning is widely used for aligning Large Language Models (LLM) with human preference. However, the conventional reward modelling is predominantly dependent on human annotations provided by a select c…

Deep Reinforcement Learning

Dynamic Incentive-aware Learning: Robust Pricing in Contextual Auctions

2020-02-25 · NeurIPS 2019 12 · Negin Golrezaei, Adel Javanmard, Vahab Mirrokni

Motivated by pricing in ad exchange markets, we consider the problem of robust learning of reserve prices against strategic buyers in repeated contextual second-price auctions. Buyers' valuations for an item depend on th…

Learning to Route LLMs from Implicit Cost-Performance Preferences via Meta-Learning

2026-06-04 · Jiahao Zeng, Ming Tang, Ningning Ding arxiv

Large language models (LLMs) present a trade-off between performance and cost, where more powerful models incur greater expense. LLM routing aims to mitigate expenses while maintaining performance by sending queries to t…