Protein Word Detection using Text Segmentation Techniques
Literature in Molecular Biology is abundant with linguistic metaphors. There have been works in the past that attempt to draw parallels between linguistics and biology, driven by the fundamental premise that proteins have a language of their own. Since word detection is crucial to the decipherment of any unknown language, we attempt to establish a problem mapping from natural language text to protein sequences at the level of words. Towards this end, we explore the use of an unsupervised text segmentation algorithm to the task of extracting {``}biological words{''} from protein sequences. In particular, we demonstrate the effectiveness of using domain knowledge to complement data driven approaches in the text segmentation task, as well as in its biological counterpart. We also propose a novel extrinsic evaluation measure for protein words through protein family classification.
Code (0)
등록된 구현이 없습니다.
Tasks
DeciphermentGeneral ClassificationSegmentationText SegmentationSimilar Papers 제목 키워드 기반
Linguistic Laws Meet Protein Sequences: A Comparative Analysis of Subword Tokenization Methods
Tokenization is a crucial step in processing protein sequences for machine learning models, as proteins are complex sequences of amino acids that require meaningful segmentation to capture their functional and structural…
Exploring Data-Driven Chemical SMILES Tokenization Approaches to Identify Key Protein-Ligand Binding Moieties
Machine learning models have found numerous successful applications in computational drug discovery. A large body of these models represents molecules as sequences since molecular sequences are easily available, simple, …
Drug DesignDrug DiscoveryProperty PredictionImproving detection of protein-ligand binding sites with 3D segmentation
In recent years machine learning (ML) took bio- and cheminformatics fields by storm, providing new solutions for a vast repertoire of problems related to protein sequence, structure, and interactions analysis. ML techniq…
Drug DiscoveryArabic Handwritten Document OCR Solution with Binarization and Adaptive Scale Fusion Detection
The problem of converting images of text into plain text is a widely researched topic in both academia and industry. Arabic handwritten Text Recognation (AHTR) poses additional challenges due to diverse handwriting style…
BinarizationOptical Character Recognition (OCR)Prot2Chat: Protein LLM with Early-Fusion of Text, Sequence and Structure
Motivation: Proteins are of great significance in living organisms. However, understanding their functions encounters numerous challenges, such as insufficient integration of multimodal information, a large number of tra…
Answer GenerationDecoderLanguage ModelingLanguage Modelling+2