Continually Improving Extractive QA via Human Feedback
We study continually improving an extractive question answering (QA) system via human user feedback. We design and deploy an iterative approach, where information-seeking users ask questions, receive model-predicted answers, and provide feedback. We conduct experiments involving thousands of user interactions under diverse setups to broaden the understanding of learning from feedback over time. Our experiments show effective improvement from user feedback of extractive QA models over time across different data regimes, including significant potential for domain adaptation.
Code (1)
Tasks
Domain AdaptationExtractive Question-AnsweringQuestion AnsweringSimilar Papers 제목 키워드 기반
Towards Enhancing Coherence in Extractive Summarization: Dataset and Experiments with LLMs
Extractive summarization plays a pivotal role in natural language processing due to its wide-range applications in summarizing diverse content efficiently, while also being faithful to the original content. Despite signi…
Extractive SummarizationSimulating Bandit Learning from User Feedback for Extractive Question Answering
We study learning from user feedback for extractive question answering by simulating feedback using supervised data. We cast the problem as contextual bandit learning, and analyze the characteristics of several learning …
Extractive Question-AnsweringQuestion AnsweringSimulating Bandit Learning from User Feedback for Extractive Question Answering
We study learning from user feedback for extractive question answering by simulating feedback using supervised data. We cast the problem as contextual bandit learning, and analyze the characteristics of several learning …
Extractive Question-AnsweringQuestion AnsweringContinual Learning for Instruction Following from Realtime Feedback
We propose and deploy an approach to continually train an instruction-following agent from feedback provided by users during collaborative interactions. During interaction, human users instruct an agent using natural lan…
Continual LearningInstruction FollowingContinual Learning with Delayed Feedback
Most of the artificial neural networks are using the benefit of labeled datasets whereas in human brain, the learning is often unsupervised. The feedback or a label for a given input or a sensory stimuli is not often ava…
Continual Learning