paper-with-me

Papers

BioChemInsight: An Open-Source Toolkit for Automated Identification and Recognition of Optical Chemical Structures and Activity Data in Scientific Publications

2025-04-12 · Zhe Wang, Fangtian Fu, Wei zhang, Lige Yan, Yan Meng, Jianping Wu, Hui Wu, Gang Xu, Si Chen

Automated extraction of chemical structures and their bioactivity data is crucial for accelerating drug discovery and enabling data-driven pharmaceutical research. Existing optical chemical structure recognition (OCSR) tools fail to autonomously associate molecular structures with their bioactivity profiles, creating a critical bottleneck in structure-activity relationship (SAR) analysis. Here, we present BioChemInsight, an open-source pipeline that integrates: (1) DECIMER Segmentation and MolVec for chemical structure recognition, (2) Qwen2.5-VL-32B for compound identifier association, and (3) PaddleOCR with Gemini-2.0-flash for bioactivity extraction and unit normalization. We evaluated the performance of BioChemInsight on 25 patents and 17 articles. BioChemInsight achieved 95% accuracy for tabular patent data (structure/identifier recognition), with lower accuracy in non-tabular patents (~80% structures, ~75% identifiers), plus 92.2 % bioactivity extraction accuracy. For articles, it attained >99% identifiers and 78-80% structure accuracy in non-tabular formats, plus 97.4% bioactivity extraction accuracy. The system generates ready-to-use SAR datasets, reducing data preprocessing time from weeks to hours while enabling applications in high-throughput screening and ML-driven drug design (https://github.com/dahuilangda/BioChemInsight).

📄 PDF Abstract BibTeX arXiv:2504.10525

Code (1)

dahuilangda/biocheminsight 공식 구현 pytorch

Tasks

ArticlesDrug DesignDrug Discovery

Similar Papers 제목 키워드 기반

MAPA Project: Ready-to-Go Open-Source Datasets and Deep Learning Technology to Remove Identifying Information from Text Documents

2022-06-01 · LEGAL (LREC) 2022 6 · Victoria Arranz, Khalid Choukri, Montse Cuadros, Aitor García Pablos 외

This paper presents the outcomes of the MAPA project, a set of annotated corpora for 24 languages of the European Union and an open-source customisable toolkit able to detect and substitute sensitive information in text …

De-identificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

The Multilingual Anonymisation Toolkit for Public Administrations (MAPA) Project

2020-11-01 · EAMT 2020 11 · Ēriks Ajausks, Victoria Arranz, Laurent Bié, Aleix Cerdà-i-Cucó 외

We describe the MAPA project, funded under the Connecting Europe Facility programme, whose goal is the development of an open-source de-identification toolkit for all official European Union languages. It will be develop…

De-identification

Automating Thematic Review of Prevention of Future Deaths Reports: Replicating the ONS Child Suicide Study using Large Language Models

2025-07-28 · Sam Osian, Arpan Dutta, Sahil Bhandari, Iain E. Buchan 외 arxiv

Prevention of Future Deaths (PFD) reports, issued by coroners in England and Wales, flag systemic hazards that may lead to further loss of life. Analysis of these reports has previously been constrained by the manual eff…

WildlifeDatasets: An open-source toolkit for animal re-identification

2023-11-15 · Vojtěch Čermák, Lukas Picek, Lukáš Adam, Kostas Papafitsoros

In this paper, we present WildlifeDatasets (https://github.com/WildlifeDatasets/wildlife-datasets) - an open-source toolkit intended primarily for ecologists and computer-vision / machine-learning researchers. The Wildli…

A Software Toolkit for Pre-processing Sign Language Video Streams

2022-06-01 · SLTAT (LREC) 2022 6 · Fabrizio Nunnari

We present the requirements, design guidelines, and the software architecture of an open-source toolkit dedicated to the pre-processing of sign language video material. The toolkit is a collection of functions and comman…