Easy Semantification of Bioassays
Biological data and knowledge bases increasingly rely on Semantic Web technologies and the use of knowledge graphs for data integration, retrieval and federated queries. We propose a solution for automatically semantifying biological assays. Our solution contrasts the problem of automated semantification as labeling versus clustering where the two methods are on opposite ends of the method complexity spectrum. Characteristically modeling our problem, we find the clustering solution significantly outperforms a deep neural network state-of-the-art labeling approach. This novel contribution is based on two factors: 1) a learning objective closely modeled after the data outperforms an alternative approach with sophisticated semantic modeling; 2) automatically semantifying biological assays achieves a high performance F1 of nearly 83%, which to our knowledge is the first reported standardized evaluation of the task offering a strong benchmark model.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringData IntegrationKnowledge GraphsRetrievalSimilar Papers 제목 키워드 기반
The Digitalization of Bioassays in the Open Research Knowledge Graph
Background: Recent years are seeing a growing impetus in the semantification of scholarly knowledge at the fine-grained level of scientific entities in knowledge graphs. The Open Research Knowledge Graph (ORKG) https://w…
Knowledge GraphsSciBERT-based Semantification of Bioassays in the Open Research Knowledge Graph
As a novel contribution to the problem of semantifying biological assays, in this paper, we propose a neural-network-based approach to automatically semantify, thereby structure, unstructured bioassay text descriptions. …
Clustering Semantic Predicates in the Open Research Knowledge Graph
When semantically describing knowledge graphs (KGs), users have to make a critical choice of a vocabulary (i.e. predicates and resources). The success of KG building is determined by the convergence of shared vocabularie…
Clusteringgraph constructionKnowledge GraphsPredicting Novel Functional Roles of Designed Small Biomolecules: An ML Approach Utilizing PubChem Compound and Substance Identifiers (CID-SID ML model)
Significance and Object: The proposed methodology aims to provide time- and cost-effective approach for the early stage in drug discovery. The machine learning models developed in this study used only the identification …
Drug DiscoveryGraph Memory Networks for Molecular Activity Prediction
Molecular activity prediction is critical in drug design. Machine learning techniques such as kernel methods and random forests have been successful for this task. These models require fixed-size feature vectors as input…
Activity PredictionDrug DesignMulti-Task LearningPrediction