paper-with-me

홈 › Papers

LeanVec: Searching vectors faster by making them fit

2023-12-26 · Mariano Tepper, Ishwar Singh Bhati, Cecilia Aguerrebere, Mark Hildebrand, Ted Willke

Modern deep learning models have the ability to generate high-dimensional vectors whose similarity reflects semantic resemblance. Thus, similarity search, i.e., the operation of retrieving those vectors in a large collection that are similar to a given query, has become a critical component of a wide range of applications that demand highly accurate and timely answers. In this setting, the high vector dimensionality puts similarity search systems under compute and memory pressure, leading to subpar performance. Additionally, cross-modal retrieval tasks have become increasingly common, e.g., where a user inputs a text query to find the most relevant images for that query. However, these queries often have different distributions than the database embeddings, making it challenging to achieve high accuracy. In this work, we present LeanVec, a framework that combines linear dimensionality reduction with vector quantization to accelerate similarity search on high-dimensional vectors while maintaining accuracy. We present LeanVec variants for in-distribution (ID) and out-of-distribution (OOD) queries. LeanVec-ID yields accuracies on par with those from recently introduced deep learning alternatives whose computational overhead precludes their usage in practice. LeanVec-OOD uses two novel techniques for dimensionality reduction that consider the query and database distributions to simultaneously boost the accuracy and the performance of the framework even further (even presenting competitive results when the query and database distributions match). All in all, our extensive and varied experimental results show that LeanVec produces state-of-the-art results, with up to 3.7x improvement in search throughput and up to 4.9x faster index build time over the state of the art.

📄 PDF Abstract BibTeX arXiv:2312.16335

Code (2)

IntelLabs/ScalableVectorSearch 공식 구현
intellabs/vectorsearchdatasets 공식 구현

Tasks

Cross-Modal RetrievalDimensionality ReductionQuantization

Similar Papers 제목 키워드 기반

GleanVec: Accelerating vector search with minimalist nonlinear dimensionality reduction

2024-10-14 · Mariano Tepper, Ishwar Singh Bhati, Cecilia Aguerrebere, Ted Willke

Embedding models can generate high-dimensional vectors whose similarity reflects semantic affinities. Thus, accurately and timely retrieving those vectors in a large collection that are similar to a given query has becom…

Cross-Modal RetrievalDimensionality Reduction

4bit-Quantization in Vector-Embedding for RAG

2025-01-17 · Taehee Jeong

Retrieval-augmented generation (RAG) is a promising technique that has shown great potential in addressing some of the limitations of large language models (LLMs). LLMs have two major limitations: they can contain outdat…

QuantizationRAGRetrieval-augmented Generation

SOLAR: Sparse Orthogonal Learned and Random Embeddings

2020-08-30 · ICLR 2021 1 · Tharun Medini, Beidi Chen, Anshumali Shrivastava

Dense embedding models are commonly deployed in commercial search engines, wherein all the document vectors are pre-computed, and near-neighbor search (NNS) is performed with the query vector to find relevant documents. …

Multi-Label ClassificationMUlTI-LABEL-ClASSIFICATION

Faster Local Solvers for Graph Diffusion Equations

2024-10-29 · Jiahe Bai, Baojian Zhou, Deqing Yang, Yanghua Xiao

Efficient computation of graph diffusion equations (GDEs), such as Personalized PageRank, Katz centrality, and the Heat kernel, is crucial for clustering, training neural networks, and many other graph-related problems. …

VDMS: Efficient Big-Visual-Data Access for Machine Learning Workloads

2018-10-28 · Luis Remis, Vishakha Gupta-Cledat, Christina Strong, Ragaad Altarawneh

We introduce the Visual Data Management System (VDMS), which enables faster access to big-visual-data and adds support to visual analytics. This is achieved by searching for relevant visual data via metadata stored as a …

BIG-bench Machine LearningManagement