paper-with-me

홈 › Papers

InstructBio: A Large-scale Semi-supervised Learning Paradigm for Biochemical Problems

2023-04-08 · Fang Wu, Huiling Qin, Siyuan Li, Stan Z. Li, Xianyuan Zhan, Jinbo Xu

In the field of artificial intelligence for science, it is consistently an essential challenge to face a limited amount of labeled data for real-world problems. The prevailing approach is to pretrain a powerful task-agnostic model on a large unlabeled corpus but may struggle to transfer knowledge to downstream tasks. In this study, we propose InstructMol, a semi-supervised learning algorithm, to take better advantage of unlabeled examples. It introduces an instructor model to provide the confidence ratios as the measurement of pseudo-labels' reliability. These confidence scores then guide the target model to pay distinct attention to different data points, avoiding the over-reliance on labeled data and the negative influence of incorrect pseudo-annotations. Comprehensive experiments show that InstructBio substantially improves the generalization ability of molecular models, in not only molecular property predictions but also activity cliff estimations, demonstrating the superiority of the proposed method. Furthermore, our evidence indicates that InstructBio can be equipped with cutting-edge pretraining methods and used to establish large-scale and task-specific pseudo-labeled molecular datasets, which reduces the predictive errors and shortens the training process. Our work provides strong evidence that semi-supervised learning can be a promising tool to overcome the data scarcity limitation and advance molecular representation learning.

📄 PDF Abstract BibTeX arXiv:2304.03906

Code (1)

smiles724/instructbio 공식 구현 pytorch

Tasks

molecular representationRepresentation Learning

Similar Papers 제목 키워드 기반

InstructBioMol: Advancing Biomolecule Understanding and Design Following Human Instructions

2024-10-10 · Xiang Zhuang, Keyan Ding, Tianwen Lyu, Yinuo Jiang 외

Understanding and designing biomolecules, such as proteins and small molecules, is central to advancing drug discovery, synthetic biology, and enzyme engineering. Recent breakthroughs in Artificial Intelligence (AI) have…

Data IntegrationDrug Discovery

A Preliminary Study on Environmental Sound Classification Leveraging Large-Scale Pretrained Model and Semi-Supervised Learning

2021-10-01 · ROCLING 2021 10 · You-Sheng Tsao, Tien-Hong Lo, Jiun-Ting Li, Shi-Yan Weng 외

With the widespread commercialization of smart devices, research on environmental sound classification has gained more and more attention in recent years. In this paper, we set out to make effective use of large-scale au…

ClassificationData AugmentationEnvironmental Sound ClassificationSound Classification+1

A Large-scale Evaluation of Pretraining Paradigms for the Detection of Defects in Electroluminescence Solar Cell Images

2024-02-27 · David Torpey, Lawrence Pratt, Richard Klein

Pretraining has been shown to improve performance in many domains, including semantic segmentation, especially in domains with limited labelled data. In this work, we perform a large-scale evaluation and benchmarking of …

BenchmarkingDefect DetectionSegmentationSemantic Segmentation

Collaborative Learning for Semi-Supervised LiDAR Semantic Segmentation

2026-05-16 · Bin Yang, Alexandru Paul Condurache arxiv

Annotating large-scale LiDAR point clouds for 3D semantic segmentation is costly and time-consuming, which motivates the use of semi-supervised learning (SemiSL). Standard LiDAR SemiSL methods typically adopt a two-step …

LIDAR Semantic Segmentation3D Semantic SegmentationPoint Clouds

Billion-scale semi-supervised learning for image classification

2019-05-02 · I. Zeki Yalniz, Hervé Jégou, Kan Chen, Manohar Paluri 외

This paper presents a study of semi-supervised learning with large convolutional networks. We propose a pipeline, based on a teacher/student paradigm, that leverages a large collection of unlabelled images (up to 1 billi…

ClassificationGeneral Classificationimage-classificationImage Classification+2