paper-with-me

Papers

Protein identification with deep learning: from abc to xyz

2017-10-08 · Ngoc Hieu Tran, Zachariah Levine, Lei Xin, Baozhen Shan, Ming Li

Proteins are the main workhorses of biological functions in a cell, a tissue, or an organism. Identification and quantification of proteins in a given sample, e.g. a cell type under normal/disease conditions, are fundamental tasks for the understanding of human health and disease. In this paper, we present DeepNovo, a deep learning-based tool to address the problem of protein identification from tandem mass spectrometry data. The idea was first proposed in the context of de novo peptide sequencing [1] in which convolutional neural networks and recurrent neural networks were applied to predict the amino acid sequence of a peptide from its spectrum, a similar task to generating a caption from an image. We further develop DeepNovo to perform sequence database search, the main technique for peptide identification that greatly benefits from numerous existing protein databases. We combine two modules de novo sequencing and database search into a single deep learning framework for peptide identification, and integrate de Bruijn graph assembly technique to offer a complete solution to reconstruct protein sequences from tandem mass spectrometry data. This paper describes a comprehensive protocol of DeepNovo for protein identification, including training neural network models, dynamic programming search, database querying, estimation of false discovery rate, and de Bruijn graph assembly. Training and testing data, model implementations, and comprehensive tutorials in form of IPython notebooks are available in our GitHub repository (https://github.com/nh2tran/DeepNovo).

📄 PDF Abstract BibTeX arXiv:1710.02765

Code (1)

nh2tran/DeepNovo 공식 구현 tf

Tasks

Deep Learningde novo peptide sequencing

Similar Papers 제목 키워드 기반

Single-Molecule Protein Identification by Sub-Nanopore Sensors

2016-04-08 · Mikhail Kolmogorov, Eamonn Kennedy, Zhuxin Dong, Gregory Timp 외

Recent advances in top-down mass spectrometry enabled identification of intact proteins, but this technology still faces challenges. For example, top-down mass spectrometry suffers from a lack of sensitivity since the io…

idMotif: An Interactive Motif Identification in Protein Sequences

2024-02-04 · Ji Hwan Park, Vikash Prasad, Sydney Newsom, Fares Najar 외

This article introduces idMotif, a visual analytics framework designed to aid domain experts in the identification of motifs within protein sequences. Motifs, short sequences of amino acids, are critical for understandin…

Deep Learning

HMACA: Towards Proposing a Cellular Automata Based Tool for Protein Coding, Promoter Region Identification and Protein Structure Prediction

2014-01-21 · Pokkuluri Kiran Sree, Inampudi Ramesh Babu, SSSN Usha Devi N

Human body consists of lot of cells, each cell consist of DeOxaRibo Nucleic Acid (DNA). Identifying the genes from the DNA sequences is a very difficult task. But identifying the coding regions is more complex task compa…

Protein Structure Prediction

A multi-layer refined network model for the identification of essential proteins

2023-12-06 · Haoyue Wang, Li Pan, Bo Yang, Junqiang Jiang 외

The identification of essential proteins in protein-protein interaction networks (PINs) can help to discover drug targets and prevent disease. In order to improve the accuracy of the identification of essential proteins,…

Specificity

Characterization of protein complexes using chemical cross-linking coupled electrospray mass spectrometry

2016-06-14

Identification and characterization of large protein complexes is a mainstay of biochemical toolboxes. Utilization of cross-linking chemicals can facilitate the capture and identification of transient or weak interaction…

Cultural Vocal Bursts Intensity Prediction