Multi-stage Training with Improved Negative Contrast for Neural Passage Retrieval
In the context of neural passage retrieval, we study three promising techniques: synthetic data generation, negative sampling, and fusion. We systematically investigate how these techniques contribute to the performance of the retrieval system and how they complement each other. We propose a multi-stage framework comprising of pre-training with synthetic data, fine-tuning with labeled data, and negative sampling at both stages. We study six negative sampling strategies and apply them to the fine-tuning stage and, as a noteworthy novelty, to the synthetic data that we use for pre-training. Also, we explore fusion methods that combine negatives from different strategies. We evaluate our system using two passage retrieval tasks for open-domain QA and using MS MARCO. Our experiments show that augmenting the negative contrast in both stages is effective to improve passage retrieval accuracy and, importantly, they also show that synthetic data generation and negative sampling have additive benefits. Moreover, using the fusion of different kinds allows us to reach performance that establishes a new state-of-the-art level in two of the tasks we evaluated.
Code (0)
등록된 구현이 없습니다.
Tasks
Passage RetrievalRetrievalSynthetic Data GenerationSimilar Papers 제목 키워드 기반
Robust Data2vec: Noise-robust Speech Representation Learning for ASR by Combining Regression and Improved Contrastive Learning
Self-supervised pre-training methods based on contrastive learning or regression tasks can utilize more unlabeled data to improve the performance of automatic speech recognition (ASR). However, the robustness impact of c…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive Learningregression+4Neural Passage Retrieval with Improved Negative Contrast
In this paper we explore the effects of negative sampling in dual encoder models used to retrieve passages for automatic question answering. We explore four negative sampling strategies that complement the straightforwar…
Open-Domain Question AnsweringPassage RetrievalQuestion AnsweringRetrieval+2Self-supervised pre-training and contrastive representation learning for multiple-choice video QA
Video Question Answering (Video QA) requires fine-grained understanding of both video and language modalities to answer the given questions. In this paper, we propose novel training schemes for multiple-choice video ques…
Auxiliary LearningContrastive LearningMultiple-choiceQuestion Answering+2When hard negative sampling meets supervised contrastive learning
State-of-the-art image models predominantly follow a two-stage strategy: pre-training on large datasets and fine-tuning with cross-entropy loss. Many studies have shown that using cross-entropy can result in sub-optimal …
Contrastive LearningFew-Shot LearningUnderwater Image Enhancement with Cascaded Contrastive Learning
Underwater image enhancement (UIE) is a highly challenging task due to the complexity of underwater environment and the diversity of underwater image degradation. Due to the application of deep learning, current UIE meth…
Contrastive LearningImage EnhancementUIE