paper-with-me

Papers

ADAPT: Actively Discovering and Adapting to Preferences for any Task

2025-04-05 · Maithili Patel, Xavier Puig, Ruta Desai, Roozbeh Mottaghi, Sonia Chernova, Joanne Truong, Akshara Rai

Assistive agents should be able to perform under-specified long-horizon tasks while respecting user preferences. We introduce Actively Discovering and Adapting to Preferences for any Task (ADAPT) -- a benchmark designed to evaluate agents' ability to adhere to user preferences across various household tasks through active questioning. Next, we propose Reflection-DPO, a novel training approach for adapting large language models (LLMs) to the task of active questioning. Reflection-DPO finetunes a 'student' LLM to follow the actions of a privileged 'teacher' LLM, and optionally ask a question to gather necessary information to better predict the teacher action. We find that prior approaches that use state-of-the-art LLMs fail to sufficiently follow user preferences in ADAPT due to insufficient questioning and poor adherence to elicited preferences. In contrast, Reflection-DPO achieves a higher rate of satisfying user preferences, outperforming a zero-shot chain-of-thought baseline by 6.1% on unseen users.

📄 PDF Abstract BibTeX arXiv:2504.04040

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward

2025-04-04 · Yanming Wan, Jiaxing Wu, Marwa Abdulhai, Lior Shani 외

Effective conversational agents must be able to personalize their behavior to suit a user's preferences, personality, and attributes, whether they are assisting with writing tasks or operating in domains like education o…

A Personalized Reinforcement Learning Summarization Service for Learning Structure from Unstructured Data

2023-07-09 · Samira Ghodratnama, Amin Beheshti, Mehrdad Zakershahrak

The exponential growth of textual data has created a crucial need for tools that assist users in extracting meaningful insights. Traditional document summarization approaches often fail to meet individual user requiremen…

Document Summarizationreinforcement-learning

FLoRA: Sample-Efficient Preference-based RL via Low-Rank Style Adaptation of Reward Functions

2025-04-14 · Daniel Marta, Simon Holk, Miguel Vasco, Jens Lundell 외

Preference-based reinforcement learning (PbRL) is a suitable approach for style adaptation of pre-trained robotic behavior: adapting the robot's policy to follow human user preferences while still being able to perform t…

PROPER: A Progressive Learning Framework for Personalized Large Language Models with Group-Level Adaptation

2025-03-03 · Linhai Zhang, Jialong Wu, Deyu Zhou, Yulan He

Personalized large language models (LLMs) aim to tailor their outputs to user preferences. Recent advances in parameter-efficient fine-tuning (PEFT) methods have highlighted the effectiveness of adapting population-level…

Mixture-of-Expertsparameter-efficient fine-tuning

Fast Adaptation with Bradley-Terry Preference Models in Text-To-Image Classification and Generation

2023-07-15 · Victor Gallego

Recently, large multimodal models, such as CLIP and Stable Diffusion have experimented tremendous successes in both foundations and applications. However, as these models increase in parameter size and computational requ…

image-classificationImage Classification