paper-with-me

Papers

Towards Solving Fuzzy Tasks with Human Feedback: A Retrospective of the MineRL BASALT 2022 Competition

2023-03-23 · Stephanie Milani, Anssi Kanervisto, Karolis Ramanauskas, Sander Schulhoff, Brandon Houghton, Sharada Mohanty, Byron Galbraith, Ke Chen, Yan Song, Tianze Zhou, Bingquan Yu, He Liu, Kai Guan, Yujing Hu, Tangjie Lv, Federico Malato, Florian Leopold, Amogh Raut, Ville Hautamäki, Andrew Melnik, Shu Ishida, João F. Henriques, Robert Klassert, Walter Laurito, Ellen Novoseller, Vinicius G. Goecks, Nicholas Waytowich, David Watkins, Josh Miller, Rohin Shah

To facilitate research in the direction of fine-tuning foundation models from human feedback, we held the MineRL BASALT Competition on Fine-Tuning from Human Feedback at NeurIPS 2022. The BASALT challenge asks teams to compete to develop algorithms to solve tasks with hard-to-specify reward functions in Minecraft. Through this competition, we aimed to promote the development of algorithms that use human feedback as channels to learn the desired behavior. We describe the competition and provide an overview of the top solutions. We conclude by discussing the impact of the competition and future directions for improvement.

📄 PDF Abstract BibTeX arXiv:2303.13512

Code (0)

등록된 구현이 없습니다.

Tasks

Minecraft

Similar Papers 제목 키워드 기반

RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback

2026-03-09 · Xiaoying Zhang, Zichen Liu, Yipeng Zhang, Xia Hu 외 arxiv

Standard reinforcement learning (RL) for large language model (LLM) agents primarily optimizes extrinsic task rewards, often favoring isolated task completion over continual adaptation. This paradigm can cause premature …

Reinforcement LearningTest-time Adaptation

A knowledge-based intelligent system for control of dirt recognition process in the smart washing machines

2019-05-02 · Mohsen Annabestani, Alireza Rowhanimanesh, Akram Rezaei, Ladan Avazpour 외

In this paper, we propose an intelligence approach based on fuzzy logic to modeling human intelligence in washing clothes. At first, an intelligent feedback loop is designed for perception-based sensing of dirt inspired …

Decision Making

Retrospective on the 2021 BASALT Competition on Learning from Human Feedback

2022-04-14 · Rohin Shah, Steven H. Wang, Cody Wild, Stephanie Milani 외

We held the first-ever MineRL Benchmark for Agents that Solve Almost-Lifelike Tasks (MineRL BASALT) Competition at the Thirty-fifth Conference on Neural Information Processing Systems (NeurIPS 2021). The goal of the comp…

Minecraft

Model Predictive Control for T-S Fuzzy Markovian Jump Systems Using Dynamic Prediction Optimization

2024-08-27 · Bin Zhang

In this paper, the model predictive control (MPC) problem is investigated for the constrained discrete-time Takagi-Sugeno fuzzy Markovian jump systems (FMJSs) under imperfect premise matching rules. To strike a balance b…

Model Predictive Control

A Novel Robust Extended Dissipativity State Feedback Control system design for Interval Type-2 Fuzzy Takagi-Sugeno Large-Scale Systems

2021-08-31 · Mojtaba Asadi Jokar, Iman Zamani, Mohamad Manthouri, Mohammad Sarbaz

In this paper, we use the advantage of large-scale systems modeling based on the type-2 fuzzy Takagi-Sugeno model to cover the uncertainties caused by large-scale systems modeling. The advantage of using membership funct…