paper-with-me

홈 › Papers

Long-range gene expression prediction with token alignment of large language model

2024-10-02 · Edouardo Honig, Huixin Zhan, Ying Nian Wu, Zijun Frank Zhang

Gene expression is a cellular process that plays a fundamental role in human phenotypical variations and diseases. Despite advances of deep learning models for gene expression prediction, recent benchmarks have revealed their inability to learn distal regulatory grammar. Here, we address this challenge by leveraging a pretrained large language model to enhance gene expression prediction. We introduce Genetic sequence Token Alignment (GTA), which aligns genetic sequence features with natural language tokens, allowing for symbolic reasoning of genomic sequence features via the frozen language model. This cross-modal adaptation learns the regulatory grammar and allows us to further incorporate gene-specific human annotations as prompts, enabling in-context learning that is not possible with existing models. Trained on lymphoblastoid cells, GTA was evaluated on cells from the Geuvadis consortium and outperforms state-of-the-art models such as Enformer, achieving a Spearman correlation of 0.65, a 10\% improvement. Additionally, GTA offers improved interpretation of long-range interactions through the identification of the most meaningful sections of the input genetic context. GTA represents a powerful and novel cross-modal approach to gene expression prediction by utilizing a pretrained language model, in a paradigm shift from conventional gene expression models trained only on sequence data.

📄 PDF Abstract BibTeX arXiv:2410.01858

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningLanguage ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Provable Long-Range Benefits of Next-Token Prediction

2025-12-08 · Xinyuan Cao, Santosh S. Vempala arxiv

Why do modern language models, trained to do well on next-word prediction, appear to generate coherent documents and capture long-range structure? Here we show that next-token prediction is provably powerful for learning…

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation

2025-05-30 · Wenrui Liu, Qian Chen, Wen Wang, Yafeng Chen 외

Neural audio codecs, used as speech tokenizers, have demonstrated remarkable potential in the field of speech generation. However, to ensure high-fidelity audio reconstruction, neural audio codecs typically encode audio …

Language ModelingLanguage Modellingtext-to-speechText to Speech

Do Long-Range Language Models Actually Use Long-Range Context?

2021-09-19 · EMNLP 2021 11 · Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, Mohit Iyyer

Language models are generally trained on short, truncated input sequences, which limits their ability to use discourse-level information present in long-range context to improve their predictions. Recent efforts to impro…

2k8kSentence

LOGO-Former: Local-Global Spatio-Temporal Transformer for Dynamic Facial Expression Recognition

2023-05-05 · Fuyan Ma, Bin Sun, Shutao Li

Previous methods for dynamic facial expression recognition (DFER) in the wild are mainly based on Convolutional Neural Networks (CNNs), whose local operations ignore the long-range dependencies in videos. Transformer-bas…

Dynamic Facial Expression RecognitionFacial Expression Recognition

Learning Python Code Suggestion with a Sparse Pointer Network

2016-11-24 · Avishkar Bhoopchand, Tim Rocktäschel, Earl Barr, Sebastian Riedel

To enhance developer productivity, all modern integrated development environments (IDEs) include code suggestion functionality that proposes likely next tokens at the cursor. While current IDEs work well for statically-t…

Language ModelingLanguage Modelling