Do not let the history haunt you: Mitigating Compounding Errors in Conversational Question Answering
The Conversational Question Answering (CoQA) task involves answering a sequence of inter-related conversational questions about a contextual paragraph. Although existing approaches employ human-written ground-truth answers for answering conversational questions at test time, in a realistic scenario, the CoQA model will not have any access to ground-truth answers for the previous questions, compelling the model to rely upon its own previously predicted answers for answering the subsequent questions. In this paper, we find that compounding errors occur when using previously predicted answers at test time, significantly lowering the performance of CoQA systems. To solve this problem, we propose a sampling strategy that dynamically selects between target answers and model predictions during training, thereby closely simulating the situation at test time. Further, we analyse the severity of this phenomena as a function of the question type, conversation length and domain type.
Code (0)
등록된 구현이 없습니다.
Tasks
Conversational Question AnsweringQuestion AnsweringSimilar Papers 제목 키워드 기반
Do not let the history haunt you -- Mitigating Compounding Errors in Conversational Question Answering
The Conversational Question Answering (CoQA) task involves answering a sequence of inter-related conversational questions about a contextual paragraph. Although existing approaches employ human-written ground-truth answe…
Conversational Question AnsweringQuestion AnsweringNon-Adversarial Imitation Learning Provably Free of Compounding Errors: The Value Flow Mechanism
Adversarial imitation learning (AIL) achieves high-quality imitation by mitigating compounding errors inherent to behavioral cloning (BC), yet its adversarial optimization frequently leads to training instability. A clas…
Context-Driven Incremental Compression for Multi-Turn Dialogue Generation
Modern conversational agents condition on an ever-growing dialogue history at each turn, incurring redundant attention and encoding costs that grow with conversation length. Naive truncation or summarization degrades fid…
Dialogue GenerationLong-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism
Popular offline reinforcement learning (RL) methods rely on explicit conservatism, penalizing out-of-dataset actions or restricting rollout horizons. We question the universality of this principle and revisit a complemen…
Reinforcement LearningTest-time AdaptationRevisiting Open-Loop Execution in Robotics: Toward Reactive, Higher-Performing Policies
Action chunking --- the practice of predicting a sequence of actions and executing a prefix open-loop --- has emerged as a key enabler of recent progress in imitation learning for robotic manipulation. However, executing…