paper-with-me

Papers

Integrating Protein Sequence and Expression Level to Analysis Molecular Characterization of Breast Cancer Subtypes

2024-10-02 · Hossein Sholehrasa

Breast cancer's complexity and variability pose significant challenges in understanding its progression and guiding effective treatment. This study aims to integrate protein sequence data with expression levels to improve the molecular characterization of breast cancer subtypes and predict clinical outcomes. Using ProtGPT2, a language model designed for protein sequences, we generated embeddings that capture the functional and structural properties of proteins sequence. These embeddings were integrated with protein expression level to form enriched biological representations, which were analyzed using machine learning methods like ensemble K-means for clustering and XGBoost for classification. Our approach enabled successful clustering of patients into biologically distinct groups and accurately predicted clinical outcomes such as survival and biomarkers status, achieving high performance metrics, notably an F1 score of 0.88 for survival and 0.87 for biomarkers status prediction. Feature importance analysis identified KMT2C, CLASP2, and MYO1B as key proteins involved in hormone signaling, cytoskeletal remodeling, and therapy resistance in hormone receptor-positive and triple-negative breast cancer, with potential influence on breast cancer subtype behavior and progression. Furthermore, protein-protein interaction networks and correlation analyses revealed functional interdependencies among proteins that may influence breast cancer subtype behavior and progression. These findings suggest that integrating protein sequence and expression data provides valuable insights into tumor biology and has significant potential to enhance personalized treatment strategies in breast cancer care.

📄 PDF Abstract BibTeX arXiv:2410.01755

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringFeature ImportanceLanguage Modelling

Similar Papers 제목 키워드 기반

STELLA: Towards Protein Function Prediction with Multimodal LLMs Integrating Sequence-Structure Representations

2025-06-04 · Hongwang Xiao, Wenjun Lin, Xi Chen, Hui Wang 외

Protein biology focuses on the intricate relationships among sequences, structures, and functions. Deciphering protein functions is crucial for understanding biological processes, advancing drug discovery, and enabling s…

Drug DiscoveryGeneral KnowledgePredictionProtein Function Prediction

CodonMPNN for Organism Specific and Codon Optimal Inverse Folding

2024-09-25 · Hannes Stark, Umesh Padia, Julia Balla, Cameron Diao 외

Generating protein sequences conditioned on protein structures is an impactful technique for protein engineering. When synthesizing engineered proteins, they are commonly translated into DNA and expressed in an organism …

Integrating Heterogeneous Gene Expression Data through Knowledge Graphs for Improving Diabetes Prediction

2024-04-23 · Rita T. Sousa, Heiko Paulheim

Diabetes is a worldwide health issue affecting millions of people. Machine learning methods have shown promising results in improving diabetes prediction, particularly through the analysis of diverse data types, namely g…

Data IntegrationDiabetes PredictionKnowledge Graphs

ProtDAT: A Unified Framework for Protein Sequence Design from Any Protein Text Description

2024-12-05 · Xiao-Yu Guo, Yi-Fan Li, YuAn Liu, Xiaoyong Pan 외

Protein design has become a critical method in advancing significant potential for various applications such as drug development and enzyme engineering. However, protein design methods utilizing large language models wit…

DescriptiveProtein Design

Deciphering Cell Systems: Machine Learning Perspectives And Approaches For The Analysis Of Single-Cell Data

2024-09-28 · Yongjian Yang

This dissertation explores the application of machine learning in molecular biology, focusing on gene expression regulation and cellular behavior at the single-cell level. Using modern neural networks, the research addre…