paper-with-me

홈 › Papers

An Early FIRST Reproduction and Improvements to Single-Token Decoding for Fast Listwise Reranking

2024-11-08 · Zijian Chen, Ronak Pradeep, Jimmy Lin

Recent advances have demonstrated that large language models (LLMs) excel as listwise rerankers, but their high computational demands remain a barrier to widespread adoption. Further, the traditional language modeling (LM) objective is not ideally suited for reranking tasks. FIRST is a novel approach that addresses these challenges by integrating a learning-to-rank objective and leveraging the logits of only the first generated token, thereby significantly reducing inference latency compared to traditional LLM rerankers. In this study, we extend the evaluation of FIRST to the TREC Deep Learning datasets (DL19-22), validating its robustness across diverse domains. We investigate the influence of different first-stage retrievers on FIRST rerankers, observing diminishing returns and patterns consistent with traditional LLM rerankers. Through applying the FIRST objective to a broader range of backbone models, we achieve effectiveness surpassing the original implementation. Our experiments confirm that fast reranking with single-token logits does not compromise out-of-domain reranking quality. To better quantify the computational savings in the original study, we measure and compare latency to find a 21%-42% gain across various models and benchmarks. Moreover, while LM training implicitly improves zero-shot single-token reranking, our experiments also raise questions about whether LM pre-training may hinder subsequent fine-tuning with the FIRST objective. These findings pave the way for more efficient and effective listwise reranking in future applications.

📄 PDF Abstract BibTeX arXiv:2411.05508

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLearning-To-RankReranking

Similar Papers 제목 키워드 기반

EarlyTom: Early Token Compression Completes Fast Video Understanding

2026-05-28 · Hesong Wang, Xin Jin, Lu Lu, Chenhaowen Li 외 arxiv

Video large language models (Video-LLMs) have demonstrated strong capabilities in video understanding tasks. However, their practical deployment is still hindered by the inefficiency introduced by processing massive amou…

Early estimates on monkeypox incubation period, generation time and reproduction number in Italy, May-June 2022

2022-07-27 · Giorgio Guzzetta, Alessia Mammone, Federica Ferraro, Anna Caraglia 외

We analyzed the first 255 PCR-confirmed cases of monkeypox occurred in Italy in 2022. Preliminary estimates indicate: mean incubation period of 9.1 days (95%CI of the mean: 6.5-10.9); mean generation time of 12.5 days (9…

SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

2026-01-20 · Aditya Bharat Soni, Rajat Ghosh, Vaishnavi Bhargava, Valerie Chen 외 arxiv

Software testing is crucial for ensuring the correctness and reliability of software systems. Automated generation of issue reproduction tests from natural language issue descriptions enhances developer productivity by s…

Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets

2026-08-19 · Tate Berenbaum, Muthaiah Venkatachalam arxiv

Modern Intel AI PCs ship capable integrated GPUs and NPUs with 16+ GB of unified memory, and they spend considerable time idle. That is not enough memory to fit a large model such as a 70B-parameter LLM. We show that a h…

The early phase of the COVID-19 outbreak in Lombardy, Italy

2020-03-20

In the night of February 20, 2020, the first case of novel coronavirus disease (COVID-19) was confirmed in the Lombardy Region, Italy. In the week that followed, Lombardy experienced a very rapid increase in the number o…