paper-with-me

홈 › Papers

Fast Processing and Querying of 170TB of Genomics Data via a Repeated And Merged BloOm Filter (RAMBO)

2019-10-10 · Gaurav Gupta, Minghao Yan, Benjamin Coleman, Bryce Kille, R. A. Leo Elworth, Tharun Medini, Todd Treangen, Anshumali Shrivastava

DNA sequencing, especially of microbial genomes and metagenomes, has been at the core of recent research advances in large-scale comparative genomics. The data deluge has resulted in exponential growth in genomic datasets over the past years and has shown no sign of slowing down. Several recent attempts have been made to tame the computational burden of sequence search on these terabyte and petabyte-scale datasets, including raw reads and assembled genomes. However, no known implementation provides both fast query and construction time, keeps the low false-positive requirement, and offers cheap storage of the data structure. We propose a data structure for search called RAMBO (Repeated And Merged BloOm Filter) which is significantly faster in query time than state-of-the-art genome indexing methods- COBS (Compact bit-sliced signature index), Sequence Bloom Trees, HowDeSBT, and SSBT. Furthermore, it supports insertion and query process parallelism, cheap updates for streaming inputs, has a zero false-negative rate, a low false-positive rate, and a small index size. RAMBO converts the search problem into set membership testing among $K$ documents. Interestingly, it is a count-min sketch type arrangement of a membership testing utility (Bloom Filter in our case). The simplicity of the algorithm and embarrassingly parallel architecture allows us to stream and index a 170TB whole-genome sequence dataset in a mere 9 hours on a cluster of 100 nodes while competing methods require weeks.

📄 PDF Abstract BibTeX arXiv:1910.04358

Code (1)

gaurav16gupta/rambo_msmt 공식 구현

Similar Papers 제목 키워드 기반

Scaling GraphLLM with Bilevel-Optimized Sparse Querying

2026-01-30 · Yangzhe Peng, Haiquan Qiu, Quanming Yao, Kun He arxiv

LLMs have recently shown strong potential in enhancing node-level tasks on text-attributed graphs (TAGs) by providing explanation features. However, their practical use is severely limited by the high computational and m…

LRez: C++ API and toolkit for analyzing and managing Linked-Reads data

2021-03-26 · Pierre Morisse, Claire Lemaitre, Fabrice Legeai

Linked-Reads technologies, such as 10x Genomics, combine both the high-quality and low cost of short-reads sequencing and a long-range information, through the use of barcodes able to tag reads which originate from a com…

ManagementTAG

Processing-in-memory for genomics workloads

2025-05-31 · William Andrew Simon, Leonid Yavits, Konstantina Koliogeorgi, Yann Falevoz 외

Low-cost, high-throughput DNA and RNA sequencing (HTS) data is the main workforce for the life sciences. Genome sequencing is now becoming a part of Predictive, Preventive, Personalized, and Participatory (termed 'P4') m…

Large-Scale Video Search with Efficient Temporal Voting Structure

2016-07-25 · Ersin Esen, Savas Ozkan, Ilkay Atil

In this work, we propose a fast content-based video querying system for large-scale video search. The proposed system is distinguished from similar works with two major contributions. First contribution is superiority of…

An AI-powered Knowledge Hub for Potato Functional Genomics

2025-05-30 · Jia Yuxin, Li Jinye, Jia Yudong, Li Futing 외

Potato functional genomics lags due to unsystematic gene information curation, gene identifier inconsistencies across reference genome versions, and the increasing volume of research publications. To address these limita…

AI AgentHallucinationRAGRetrieval-augmented Generation