paper-with-me

Papers

Sequential Preference-Based Optimization

2018-01-09 · Ian Dewancker, Jakob Bauer, Michael McCourt

Many real-world engineering problems rely on human preferences to guide their design and optimization. We present PrefOpt, an open source package to simplify sequential optimization tasks that incorporate human preference feedback. Our approach extends an existing latent variable model for binary preferences to allow for observations of equivalent preference from users.

📄 PDF Abstract BibTeX arXiv:1801.02788

Code (1)

prefopt/prefopt 공식 구현

Similar Papers 제목 키워드 기반

Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game

2025-12-18 · Barna Pásztor, Thomas Kleine Buening, Andreas Krause arxiv

We introduce Stackelberg Learning from Human Feedback (SLHF), a new framework for preference optimization. SLHF frames the alignment problem as a sequential-move game between two policies: a Leader, which commits to an a…

Reinforcement Learning

Shapley-based Data Valuation for LLM Alignment via Sequential Preference Optimization

2025-12-12 · Mélissa Tamine, Otmane Sakhi, Benjamin Heymann, Maxime Vono 외 arxiv

Data valuation is a natural framework for understanding which preference datasets matter most when aligning a Large Language Model (LLM) using multiple sources. The standard game-theoretic approach assigns each dataset a…

SPO: Multi-Dimensional Preference Sequential Alignment With Implicit Reward Modeling

2024-05-21 · Xingzhou Lou, Junge Zhang, Jian Xie, Lifeng Liu 외

Human preference alignment is critical in building powerful and reliable large language models (LLMs). However, current methods either ignore the multi-dimensionality of human preferences (e.g. helpfulness and harmlessne…

Sequential Learning-based IaaS Composition

2021-02-24 · Sajib Mistry, Sheik Mohammad Mostakim Fattah, Athman Bouguettaya

We propose a novel IaaS composition framework that selects an optimal set of consumer requests according to the provider's qualitative preferences on long-term service provisions. Decision variables are included in the t…

ClusteringQ-Learning

Beyond Uniform Forgetting: A Study of Sequential Direct Preference Optimization Across Preference Settings

2026-06-18 · Pranav Bhandari, Nicolas Fay, Amitava Datta, Usman Naseem 외 arxiv

Aligning language models with human preferences often requires optimising multiple behavioural objectives. A practical approach is to apply these objectives sequentially using preference optimisation methods such as Dire…