paper-with-me

홈 › Papers

Efficient and Effective Tail Latency Minimization in Multi-Stage Retrieval Systems

2017-04-20 · Mackenzie Joel, Culpepper J. Shane, Blanco Roi, Crane Matt, Clarke Charles L. A., Lin Jimmy

Scalable web search systems typically employ multi-stage retrieval architectures, where an initial stage generates a set of candidate documents that are then pruned and re-ranked. Since subsequent stages typically exploit a multitude of features of varying costs using machine-learned models, reducing the number of documents that are considered at each stage improves latency. In this work, we propose and validate a unified framework that can be used to predict a wide range of performance-sensitive parameters which minimize effectiveness loss, while simultaneously minimizing query latency, across all stages of a multi-stage search architecture. Furthermore, our framework can be easily applied in large-scale IR systems, can be trained without explicitly requiring relevance judgments, and can target a variety of different efficiency-effectiveness trade-offs, making it well suited to a wide range of search scenarios. Our results show that we can reliably predict a number of different parameters on a per-query basis, while simultaneously detecting and minimizing the likelihood of tail-latency queries that exceed a pre-specified performance budget. As a proof of concept, we use the prediction framework to help alleviate the problem of tail-latency queries in early stage retrieval. On the standard ClueWeb09B collection and 31k queries, we show that our new hybrid system can reliably achieve a maximum query time of 200 ms with a 99.99% response time guarantee without a significant loss in overall effectiveness. The solutions presented are practical, and can easily be used in large-scale distributed search engine deployments with a small amount of additional overhead.

📄 PDF Abstract BibTeX arXiv:1704.03970

Code (0)

등록된 구현이 없습니다.

Tasks

Retrieval

Similar Papers 제목 키워드 기반

Class-Conditional Sharpness-Aware Minimization for Deep Long-Tailed Recognition

2023-01-01 · CVPR 2023 1 · Zhipeng Zhou, Lanqing Li, Peilin Zhao, Pheng-Ann Heng 외

It's widely acknowledged that deep learning models with flatter minima in its loss landscape tend to generalize better. However, such property is under-explored in deep long-tailed recognition (DLTR), a practical pro…

Long-tail Learning

Deep Unfolding Tensor Rank Minimization With Generalized Detail Injection for Pansharpening

2024-04-22 · IEEE Transactions on Geoscience and Remote Sensing 2024 4 · Truong Thanh Nhat Mai, Edmund Y. Lam, Chul Lee

Pansharpening aims to generate a high-resolution multispectral (HRMS) image by merging a low-resolution multispectral (LRMS) image with a high-resolution panchromatic (PAN) image. While traditional model-based pansharpen…

Pansharpening

Lotus: learning-based online thermal and latency variation management for two-stage detectors on edge devices

2024-10-01 · Yifan Gong, Yushu Wu, Zheng Zhan, Pu Zhao 외

Two-stage object detectors exhibit high accuracy and precise localization, especially for identifying small objects that are favorable for various edge applications. However, the high computation costs associated with tw…

CPUDeep Reinforcement LearningGPUManagement+1

Federated Learning with Correlated Data: Taming the Tail for Age-Optimal Industrial IoT

2021-08-17 · Chen-Feng Liu, Mehdi Bennis

While information delivery in industrial Internet of things demands reliability and latency guarantees, the freshness of the controller's available information, measured by the age of information (AoI), is paramount for …

Federated LearningModel Selection

Graph-Based Simplex Method for Pairwise Energy Minimization With Binary Variables

2015-06-01 · CVPR 2015 6 · Daniel Prusa

We show how the simplex algorithm can be tailored to the linear programming relaxation of pairwise energy minimization with binary variables. A special structure formed by basic and nonbasic variables in each stage of th…