paper-with-me

Papers

Quality-Diversity through AI Feedback

2023-10-19 · Herbie Bradley, Andrew Dai, Hannah Teufel, Jenny Zhang, Koen Oostermeijer, Marco Bellagente, Jeff Clune, Kenneth Stanley, Grégory Schott, Joel Lehman

In many text-generation problems, users may prefer not only a single response, but a diverse range of high-quality outputs from which to choose. Quality-diversity (QD) search algorithms aim at such outcomes, by continually improving and diversifying a population of candidates. However, the applicability of QD to qualitative domains, like creative writing, has been limited by the difficulty of algorithmically specifying measures of quality and diversity. Interestingly, recent developments in language models (LMs) have enabled guiding search through AI feedback, wherein LMs are prompted in natural language to evaluate qualitative aspects of text. Leveraging this development, we introduce Quality-Diversity through AI Feedback (QDAIF), wherein an evolutionary algorithm applies LMs to both generate variation and evaluate the quality and diversity of candidate text. When assessed on creative writing domains, QDAIF covers more of a specified search space with high-quality samples than do non-QD controls. Further, human evaluation of QDAIF-generated creative texts validates reasonable agreement between AI and human evaluation. Our results thus highlight the potential of AI feedback to guide open-ended search for creative and original solutions, providing a recipe that seemingly generalizes to many domains and modalities. In this way, QDAIF is a step towards AI systems that can independently search, diversify, evaluate, and improve, which are among the core skills underlying human society's capacity for innovation.

📄 PDF Abstract BibTeX arXiv:2310.13032

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityText Generation

Similar Papers 제목 키워드 기반

Quality Diversity through Human Feedback: Towards Open-Ended Diversity-Driven Optimization

2023-10-18 · Li Ding, Jenny Zhang, Jeff Clune, Lee Spector 외

Reinforcement Learning from Human Feedback (RLHF) has shown potential in qualitative tasks where easily defined performance measures are lacking. However, there are drawbacks when RLHF is commonly used to optimize for av…

DiversityImage Generationreinforcement-learningReinforcement Learning+4

Diversity from Human Feedback

2023-10-10 · Ren-Jian Wang, Ke Xue, Yutong Wang, Peng Yang 외

Diversity plays a significant role in many problems, such as ensemble learning, reinforcement learning, and combinatorial optimization. How to define the diversity measure is a longstanding problem. Many methods rely on …

Combinatorial OptimizationDiversityEnsemble Learning

Curiosity-Driven Reinforcement Learning from Human Feedback

2025-01-20 · Haoran Sun, Yekun Chai, Shuohuan Wang, Yu Sun 외

Reinforcement learning from human feedback (RLHF) has proven effective in aligning large language models (LLMs) with human preferences, but often at the cost of reduced output diversity. This trade-off between diversity …

DiversityInstruction Followingreinforcement-learningReinforcement Learning+1

Interpretable Directed Diversity: Leveraging Model Explanations for Iterative Crowd Ideation

2021-09-21 · Yunlong Wang, Priyadarshini Venkatesh, Brian Y. Lim

Feedback in creativity support tools can help crowdworkers to improve their ideations. However, current feedback methods require human assessment from facilitators or peers. This is not scalable to large crowds. We propo…

counterfactualDiversity

Adaptive Quality-Diversity Trade-offs for Large-Scale Batch Recommendation

2026-02-02 · Clémence Réda, Tomas Rigaux, Hiba Bederina, Koh Takeuchi 외 arxiv

A core research question in recommender systems is to propose batches of highly relevant and diverse items, that is, items personalized to the user's preferences, but which also might get the user out of their comfort zo…

Movie RecommendationPoint Processes