paper-with-me

홈 › Papers

Contextualization with SPLADE for High Recall Retrieval

2024-05-07 · Eugene Yang

High Recall Retrieval (HRR), such as eDiscovery and medical systematic review, is a search problem that optimizes the cost of retrieving most relevant documents in a given collection. Iterative approaches, such as iterative relevance feedback and uncertainty sampling, are shown to be effective under various operational scenarios. Despite neural models demonstrating success in other text-related tasks, linear models such as logistic regression, in general, are still more effective and efficient in HRR since the model is trained and retrieves documents from the same fixed collection. In this work, we leverage SPLADE, an efficient retrieval model that transforms documents into contextualized sparse vectors, for HRR. Our approach combines the best of both worlds, leveraging both the contextualization from pretrained language models and the efficiency of linear models. It reduces 10% and 18% of the review cost in two HRR evaluation collections under a one-phase review workflow with a target recall of 80%. The experiment is implemented with TARexp and is available at https://github.com/eugene-yang/LSR-for-TAR.

📄 PDF Abstract BibTeX arXiv:2405.03972

Code (1)

eugene-yang/lsr-for-tar 공식 구현 pytorch

Tasks

RetrievalTAR

Similar Papers 제목 키워드 기반

A Comparative Study of Text Retrieval Models on DaReCzech

2024-11-19 · Jakub Stetina, Martin Fajcik, Michal Stefanik, Michal Hradis

This article presents a comprehensive evaluation of 7 off-the-shelf document retrieval models: Splade, Plaid, Plaid-X, SimCSE, Contriever, OpenAI ADA and Gemma2 chosen to determine their performance on the Czech retrieva…

Information RetrievalMachine TranslationRetrievalText Retrieval

Efficiency and Effectiveness of SPLADE Models on Billion-Scale Web Document Title

2025-11-27 · Taeryun Won, Tae Kwan Lee, Hiun Kim, Hyemin Lee arxiv

This paper presents a comprehensive comparison of BM25, SPLADE, and Expanded-SPLADE models in the context of large-scale web document retrieval. We evaluate the effectiveness and efficiency of these models on datasets sp…

The Pre-Training Study of Expanded-SPLADE Models on Web Document Titles

2026-05-02 · Hiun Kim, Tae Kwan Lee, Taeryun Won arxiv

Masked Language Modeling (MLM) pre-training is one of the primary ways to initialize Neural Information Retrieval (IR) models prior to retrieval fine-tuning. However, studies show that MLM pre-trained models have limited…

Information RetrievalTransfer Learning

Mistral-SPLADE: LLMs for better Learned Sparse Retrieval

2024-08-20 · Meet Doshi, Vishwajeet Kumar, Rudra Murthy, Vignesh P 외

Learned Sparse Retrievers (LSR) have evolved into an effective retrieval strategy that can bridge the gap between traditional keyword-based sparse retrievers and embedding-based dense retrievers. At its core, learned spa…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+3

The Role of Vocabularies in Learning Sparse Representations for Ranking

2025-09-20 · Hiun Kim, Tae Kwan Lee, Taeryun Won arxiv

Learned Sparse Retrieval (LSR) such as SPLADE has growing interest for effective semantic 1st stage matching while enjoying the efficiency of inverted indices. A recent work on learning SPLADE models with expanded vocabu…