paper-with-me

홈 › Papers

Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data

2025-05-25 · Manveer Singh Tamber, Suleman Kazi, Vivek Sourabh, Jimmy Lin

We investigate improving the retrieval effectiveness of embedding models through the lens of corpus-specific fine-tuning. Prior work has shown that fine-tuning with queries generated using a dataset's retrieval corpus can boost retrieval effectiveness for the dataset. However, we find that surprisingly, fine-tuning using the conventional InfoNCE contrastive loss often reduces effectiveness in state-of-the-art models. To overcome this, we revisit cross-encoder listwise distillation and demonstrate that, unlike using contrastive learning alone, listwise distillation can help more consistently improve retrieval effectiveness across multiple datasets. Additionally, we show that synthesizing more training data using diverse query types (such as claims, keywords, and questions) yields greater effectiveness than using any single query type alone, regardless of the query type used in evaluation. Our findings further indicate that synthetic queries offer comparable utility to human-written queries for training. We use our approach to train an embedding model that achieves state-of-the-art effectiveness among BERT embedding models. We release our model and both query generation and training code to facilitate further research.

📄 PDF Abstract BibTeX arXiv:2505.19274

Code (1)

manveertamber/cadet-dense-retrieval 공식 구현 pytorch

Tasks

Contrastive LearningRetrieval

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
WordPiece 설명 없음
Weight Decay 설명 없음
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

FlowTrack: Revisiting Optical Flow for Long-Range Dense Tracking

2024-01-01 · CVPR 2024 1 · Seokju Cho, Jiahui Huang, Seungryong Kim, Joon-Young Lee

In the domain of video tracking existing methods often grapple with a trade-off between spatial density and temporal range. Current approaches in dense optical flow estimators excel in providing spatially dense track…

Optical Flow Estimation

Hierarchical Topology Isomorphism Expertise Embedded Graph Contrastive Learning

2023-12-21 · Jiangmeng Li, Yifan Jin, Hang Gao, Wenwen Qiang 외

Graph contrastive learning (GCL) aims to align the positive features while differentiating the negative features in the latent space by minimizing a pair-wise contrastive loss. As the embodiment of an outstanding discrim…

Contrastive LearningGraph Representation LearningRepresentation LearningTransfer Learning

Should we pre-train a decoder in contrastive learning for dense prediction tasks?

2025-03-21 · Sébastien Quetin, Tapotosh Ghosh, Farhad Maleki

Contrastive learning in self-supervised settings primarily focuses on pre-training encoders, while decoders are typically introduced and trained separately for downstream dense prediction tasks. This conventional approac…

Contrastive LearningDecoderInstance Segmentationobject-detection+3

Multi-Scale DenseNet-Based Electricity Theft Detection

2018-05-24 · Bo Li, Kele Xu, Xiaoyan Cui, Yiheng Wang 외

Electricity theft detection issue has drawn lots of attention during last decades. Timely identification of the electricity theft in the power system is crucial for the safety and availability of the system. Although sus…

Feature Engineering

Building Scalable Video Understanding Benchmarks through Sports

2023-01-17 · Aniket Agarwal, Alex Zhang, Karthik Narasimhan, Igor Gilitschenski 외

Existing benchmarks for evaluating long video understanding falls short on two critical aspects, either lacking in scale or quality of annotations. These limitations arise from the difficulty in collecting dense annotati…

Video Understanding