paper-with-me

Papers

Biological Sequence Kernels with Guaranteed Flexibility

2023-04-06 · Alan Nawzad Amin, Eli Nathan Weinstein, Debora Susan Marks

Applying machine learning to biological sequences - DNA, RNA and protein - has enormous potential to advance human health, environmental sustainability, and fundamental biological understanding. However, many existing machine learning methods are ineffective or unreliable in this problem domain. We study these challenges theoretically, through the lens of kernels. Methods based on kernels are ubiquitous: they are used to predict molecular phenotypes, design novel proteins, compare sequence distributions, and more. Many methods that do not use kernels explicitly still rely on them implicitly, including a wide variety of both deep learning and physics-based techniques. While kernels for other types of data are well-studied theoretically, the structure of biological sequence space (discrete, variable length sequences), as well as biological notions of sequence similarity, present unique mathematical challenges. We formally analyze how well kernels for biological sequences can approximate arbitrary functions on sequence space and how well they can distinguish different sequence distributions. In particular, we establish conditions under which biological sequence kernels are universal, characteristic and metrize the space of distributions. We show that a large number of existing kernel-based machine learning methods for biological sequences fail to meet our conditions and can as a consequence fail severely. We develop straightforward and computationally tractable ways of modifying existing kernels to satisfy our conditions, imbuing them with strong guarantees on accuracy and reliability. Our proof techniques build on and extend the theory of kernels with discrete masses. We illustrate our theoretical results in simulation and on real biological data sets.

📄 PDF Abstract BibTeX arXiv:2304.03775

Code (1)

alannawzadamin/kernels-with-guarantees 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

Recurrent Kernel Networks

2019-06-07 · NeurIPS 2019 12 · Dexiong Chen, Laurent Jacob, Julien Mairal

Substring kernels are classical tools for representing biological sequences or text. However, when large amounts of annotated data are available, models that allow end-to-end training such as neural networks are often pr…

Functional Synthetic Biology

2022-07-01 · Ibrahim Aldulijan, Jacob Beal, Sonja Billerbeck, Jeff Bouffard 외

Synthetic biologists have made great progress over the past decade in developing methods for modular assembly of genetic sequences and in engineering biological systems with a wide variety of functions in various context…

Boosting t-SNE Efficiency for Sequencing Data: Insights from Kernel Selection

2025-12-17 · Avais Jan, Prakash Chourasia, Sarwan Ali, Murray Patterson arxiv

Dimensionality reduction techniques are essential for visualizing and analyzing high-dimensional biological sequencing data. t-distributed Stochastic Neighbor Embedding (t-SNE) is widely used for this purpose, traditiona…

Dimensionality Reduction

Flexible Flows for Biological Sequence Design

2026-06-09 · Yogesh Verma, Dani Korpela, Harri Lähdesmäki, Vikas Garg arxiv

Designing functional biological sequences requires navigating vast discrete spaces under strict evolutionary and biophysical constraints. Discrete Flow Matching (DFM) offers a generative framework over such spaces, but e…

Density Estimation

Control in Boolean networks with model checking

2021-12-20 · Laura Cifuentes-Fontanals, Elisa Tonello, Heike Siebert

Understanding control mechanisms in biological systems plays a crucial role in important applications, for instance in cell reprogramming. Boolean modeling allows the identification of possible efficient strategies, help…

model