The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values
Human feedback is increasingly used to steer the behaviours of Large Language Models (LLMs). However, it is unclear how to collect and incorporate feedback in a way that is efficient, effective and unbiased, especially for highly subjective human preferences and values. In this paper, we survey existing approaches for learning from human feedback, drawing on 95 papers primarily from the ACL and arXiv repositories.First, we summarise the past, pre-LLM trends for integrating human feedback into language models. Second, we give an overview of present techniques and practices, as well as the motivations for using feedback; conceptual frameworks for defining values and preferences; and how feedback is collected and from whom. Finally, we encourage a better future of feedback learning in LLMs by raising five unresolved conceptual and practical challenges.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
You're (Not) My Type -- Can LLMs Generate Feedback of Specific Types for Introductory Programming Tasks?
Background: Feedback as one of the most influential factors for learning has been subject to a great body of research. It plays a key role in the development of educational technology systems and is traditionally rooted …
Dynamic Biases of Static Panel Data Estimators
This paper identifies an important bias - termed dynamic bias - in fixed effects panel estimators that arises when dynamic feedback is ignored in the estimating equation. Dynamic feedback occurs if past outcomes impact c…
Context-Aware Attentive Knowledge Tracing
Knowledge tracing (KT) refers to the problem of predicting future learner performance given their past performance in educational applications. Recent developments in KT using flexible deep neural network-based models ex…
Knowledge TracingDESIRE: Distant Future Prediction in Dynamic Scenes with Interacting Agents
We introduce a Deep Stochastic IOC RNN Encoderdecoder framework, DESIRE, for the task of future predictions of multiple interacting agents in dynamic scenes. DESIRE effectively predicts future locations of objects in mul…
Future predictionMulti Future Trajectory PredictionPredictionTrajectory PredictionA Global Past-Future Early Exit Method for Accelerating Inference of Pre-trained Language Models
Early exit mechanism aims to accelerate the inference speed of large-scale pre-trained language models. The essential idea is to exit early without passing through all the inference layers at the inference stage. To make…