Towards Solving Fuzzy Tasks with Human Feedback: A Retrospective of the MineRL BASALT 2022 Competition
To facilitate research in the direction of fine-tuning foundation models from human feedback, we held the MineRL BASALT Competition on Fine-Tuning from Human Feedback at NeurIPS 2022. The BASALT challenge asks teams to compete to develop algorithms to solve tasks with hard-to-specify reward functions in Minecraft. Through this competition, we aimed to promote the development of algorithms that use human feedback as channels to learn the desired behavior. We describe the competition and provide an overview of the top solutions. We conclude by discussing the impact of the competition and future directions for improvement.
Code (0)
등록된 구현이 없습니다.
Tasks
MinecraftSimilar Papers 제목 키워드 기반
RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback
Standard reinforcement learning (RL) for large language model (LLM) agents primarily optimizes extrinsic task rewards, often favoring isolated task completion over continual adaptation. This paradigm can cause premature …
Reinforcement LearningTest-time AdaptationA knowledge-based intelligent system for control of dirt recognition process in the smart washing machines
In this paper, we propose an intelligence approach based on fuzzy logic to modeling human intelligence in washing clothes. At first, an intelligent feedback loop is designed for perception-based sensing of dirt inspired …
Decision MakingRetrospective on the 2021 BASALT Competition on Learning from Human Feedback
We held the first-ever MineRL Benchmark for Agents that Solve Almost-Lifelike Tasks (MineRL BASALT) Competition at the Thirty-fifth Conference on Neural Information Processing Systems (NeurIPS 2021). The goal of the comp…
MinecraftModel Predictive Control for T-S Fuzzy Markovian Jump Systems Using Dynamic Prediction Optimization
In this paper, the model predictive control (MPC) problem is investigated for the constrained discrete-time Takagi-Sugeno fuzzy Markovian jump systems (FMJSs) under imperfect premise matching rules. To strike a balance b…
Model Predictive ControlA Novel Robust Extended Dissipativity State Feedback Control system design for Interval Type-2 Fuzzy Takagi-Sugeno Large-Scale Systems
In this paper, we use the advantage of large-scale systems modeling based on the type-2 fuzzy Takagi-Sugeno model to cover the uncertainties caused by large-scale systems modeling. The advantage of using membership funct…