paper-with-me

홈 › Papers

FPScreen: A Rapid Similarity Search Tool for Massive Molecular Library Based on Molecular Fingerprint Comparison

2019-06-13 · Lijun Wang, Jianbing Gong, Yingxia Zhang, Tianmou Liu, Junhui Gao

We designed a fast similarity search engine for large molecular libraries: FPScreen. We downloaded 100 million molecules' structure files in PubChem with SDF extension, then applied a computational chemistry tool RDKit to convert each structure file into one line of text in MACCS format and stored them in a text file as our molecule library. The similarity search engine compares the similarity while traversing the 166-bit strings in the library file line by line. FPScreen can complete similarity search through 100 million entries in our molecule library within one hour. That is very fast as a biology computation tool. Additionally, we divided our library into several strides for parallel processing. FPScreen was developed in WEB mode.

📄 PDF Abstract BibTeX arXiv:1906.06170

Code (0)

등록된 구현이 없습니다.

Tasks

Computational chemistry

Similar Papers 제목 키워드 기반

TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity

2026-08-08 · Yen-Ku Liu, Hongjie Chen, Ryan A. Rossi, Franck Dernoncourt hf

The rapid advancement of artificial intelligence (AI) has significantly accelerated research in time-series analysis, particularly in forecasting, classification, and generation tasks. Recent models, especially foundatio…

PhosNetVis: a web-based tool for fast kinase-substrate enrichment analysis and interactive 2D/3D network visualizations of phosphoproteomics data

2024-02-07 · Osho Rawal, Berk Turhan, Irene Font Peradejordi, Shreya Chandrasekar 외

Protein phosphorylation involves the reversible modification of a protein (substrate) residue by another protein (kinase). Liquid chromatography-mass spectrometry studies are rapidly generating massive protein phosphoryl…

Mining for Strong Gravitational Lenses with Self-supervised Learning

2021-09-30 · George Stein, Jacqueline Blaum, Peter Harrington, Tomislav Medan 외

We employ self-supervised representation learning to distill information from 76 million galaxy images from the Dark Energy Spectroscopic Instrument Legacy Imaging Surveys' Data Release 9. Targeting the identification of…

CPURepresentation LearningSelf-Supervised Learning

Self-supervised similarity search for large scientific datasets

2021-10-25 · George Stein, Peter Harrington, Jacqueline Blaum, Tomislav Medan 외

We present the use of self-supervised learning to explore and exploit large unlabeled datasets. Focusing on 42 million galaxy images from the latest data release of the Dark Energy Spectroscopic Instrument (DESI) Legacy …

Self-Supervised LearningSemantic SimilaritySemantic Textual Similarity

Mahalanobis Distance Metric Learning Algorithm for Instance-based Data Stream Classification

2016-04-17 · Jorge Luis Rivero Perez, Bernardete Ribeiro, Carlos Morell Perez

With the massive data challenges nowadays and the rapid growing of technology, stream mining has recently received considerable attention. To address the large number of scenarios in which this phenomenon manifests itsel…

ClassificationDrift DetectionGeneral ClassificationMetric Learning