paper-with-me

Papers

Diversity from Human Feedback

2023-10-10 · Ren-Jian Wang, Ke Xue, Yutong Wang, Peng Yang, Haobo Fu, Qiang Fu, Chao Qian

Diversity plays a significant role in many problems, such as ensemble learning, reinforcement learning, and combinatorial optimization. How to define the diversity measure is a longstanding problem. Many methods rely on expert experience to define a proper behavior space and then obtain the diversity measure, which is, however, challenging in many scenarios. In this paper, we propose the problem of learning a behavior space from human feedback and present a general method called Diversity from Human Feedback (DivHF) to solve it. DivHF learns a behavior descriptor consistent with human preference by querying human feedback. The learned behavior descriptor can be combined with any distance measure to define a diversity measure. We demonstrate the effectiveness of DivHF by integrating it with the Quality-Diversity optimization algorithm MAP-Elites and conducting experiments on the QDax suite. The results show that DivHF learns a behavior space that aligns better with human requirements compared to direct data-driven approaches and leads to more diverse solutions under human preference. Our contributions include formulating the problem, proposing the DivHF method, and demonstrating its effectiveness through experiments.

📄 PDF Abstract BibTeX arXiv:2310.06648

Code (0)

등록된 구현이 없습니다.

Tasks

Combinatorial OptimizationDiversityEnsemble Learning

Similar Papers 제목 키워드 기반

Quality Diversity through Human Feedback: Towards Open-Ended Diversity-Driven Optimization

2023-10-18 · Li Ding, Jenny Zhang, Jeff Clune, Lee Spector 외

Reinforcement Learning from Human Feedback (RLHF) has shown potential in qualitative tasks where easily defined performance measures are lacking. However, there are drawbacks when RLHF is commonly used to optimize for av…

DiversityImage Generationreinforcement-learningReinforcement Learning+4

Interpretable Directed Diversity: Leveraging Model Explanations for Iterative Crowd Ideation

2021-09-21 · Yunlong Wang, Priyadarshini Venkatesh, Brian Y. Lim

Feedback in creativity support tools can help crowdworkers to improve their ideations. However, current feedback methods require human assessment from facilitators or peers. This is not scalable to large crowds. We propo…

counterfactualDiversity

Curiosity-Driven Reinforcement Learning from Human Feedback

2025-01-20 · Haoran Sun, Yekun Chai, Shuohuan Wang, Yu Sun 외

Reinforcement learning from human feedback (RLHF) has proven effective in aligning large language models (LLMs) with human preferences, but often at the cost of reduced output diversity. This trade-off between diversity …

DiversityInstruction Followingreinforcement-learningReinforcement Learning+1

Quality-Diversity through AI Feedback

2023-10-19 · Herbie Bradley, Andrew Dai, Hannah Teufel, Jenny Zhang 외

In many text-generation problems, users may prefer not only a single response, but a diverse range of high-quality outputs from which to choose. Quality-diversity (QD) search algorithms aim at such outcomes, by continual…

DiversityText Generation

Effects of AI Feedback on Learning, the Skill Gap, and Intellectual Diversity

2024-09-27 · Christoph Riedl, Eric Bogert

Can human decision-makers learn from AI feedback? Using data on 52,000 decision-makers from a large online chess platform, we investigate how their AI use affects three interrelated long-term outcomes: Learning, skill ga…

Diversity