paper-with-me

Papers

Interpretable RNA Foundation Model from Unannotated Data for Highly Accurate RNA Structure and Function Predictions

2022-04-01 · Jiayang Chen, Zhihang Hu, Siqi Sun, Qingxiong Tan, YiXuan Wang, Qinze Yu, Licheng Zong, Liang Hong, Jin Xiao, Tao Shen, Irwin King, Yu Li

Non-coding RNA structure and function are essential to understanding various biological processes, such as cell signaling, gene expression, and post-transcriptional regulations. These are all among the core problems in the RNA field. With the rapid growth of sequencing technology, we have accumulated a massive amount of unannotated RNA sequences. On the other hand, expensive experimental observatory results in only limited numbers of annotated data and 3D structures. Hence, it is still challenging to design computational methods for predicting their structures and functions. The lack of annotated data and systematic study causes inferior performance. To resolve the issue, we propose a novel RNA foundation model (RNA-FM) to take advantage of all the 23 million non-coding RNA sequences through self-supervised learning. Within this approach, we discover that the pre-trained RNA-FM could infer sequential and evolutionary information of non-coding RNAs without using any labels. Furthermore, we demonstrate RNA-FM's effectiveness by applying it to the downstream secondary/3D structure prediction, SARS-CoV-2 genome structure and evolution prediction, protein-RNA binding preference modeling, and gene expression regulation modeling. The comprehensive experiments show that the proposed method improves the RNA structural and functional modelling results significantly and consistently. Despite only being trained with unlabelled data, RNA-FM can serve as the foundational model for the field.

📄 PDF Abstract BibTeX arXiv:2204.00300

Code (4)

ml4bio/rna-fm 공식 구현 pytorch
cgoliver/rnaglib pytorch
lulab/OligoFormer
sinc-lab/rna-llm-folding pytorch

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

MIS-FM: 3D Medical Image Segmentation using Foundation Models Pretrained on a Large-Scale Unannotated Dataset

2023-06-29 · Guotai Wang, Jianghao Wu, Xiangde Luo, Xinglong Liu 외

Pretraining with large-scale 3D volumes has a potential for improving the segmentation performance on a target medical image dataset where the training images and annotations are limited. Due to the high cost of acquirin…

Image SegmentationMedical Image SegmentationSegmentationSelf-Supervised Learning+1

Machine Learning for Medicine Must Be Interpretable, Shareable, Reproducible and Accountable by Design

2025-08-22 · Ayyüce Begüm Bektaş, Mithat Gönen arxiv

This paper claims that machine learning models deployed in high stakes domains such as medicine must be interpretable, shareable, reproducible and accountable. We argue that these principles should form the foundational …

Federated Learning

AdaFusion: Prompt-Guided Inference with Adaptive Fusion of Pathology Foundation Models

2025-08-07 · Yuxiang Xiao, Yang Hu, Bin Li, Tianyang Zhang 외 arxiv

Pathology foundation models (PFMs) have demonstrated strong representational capabilities through self-supervised pre-training on large-scale, unannotated histopathology image datasets. However, their diverse yet opaque …

Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models

2024-02-14 · Goutham Rajendran, Simon Buchholz, Bryon Aragam, Bernhard Schölkopf 외

To build intelligent machine learning systems, there are two broad approaches. One approach is to build inherently interpretable models, as endeavored by the growing field of causal representation learning. The other app…

Representation Learning

Knowledge-Augmented Contrastive Learning for Abnormality Classification and Localization in Chest X-rays with Radiomics using a Feedback Loop

2021-04-11 · Yan Han, Chongyan Chen, Ahmed Tewfik, Benjamin Glicksberg 외

Building a highly accurate predictive model for classification and localization of abnormalities in chest X-rays usually requires a large number of manually annotated labels and pixel regions (bounding boxes) of abnormal…

Contrastive Learning