paper-with-me

홈 › Papers

DNABERT-2: Fine-Tuning a Genomic Language Model for Colorectal Gene Enhancer Classification

2025-09-28 · Darren King, Yaser Atlasi, Gholamreza Rafiee arxiv

Gene enhancers control when and where genes switch on, yet their sequence diversity and tissue specificity make them hard to pinpoint in colorectal cancer. We take a sequence-only route and fine-tune DNABERT-2, a transformer genomic language model that uses byte-pair encoding to learn variable-length tokens from DNA. Using assays curated via the Johnston Cancer Research Centre at Queen's University Belfast, we assembled a balanced corpus of 2.34 million 1 kb enhancer sequences, applied summit-centered extraction and rigorous de-duplication including reverse-complement collapse, and split the data stratified by class. With a 4096-term vocabulary and a 232-token context chosen empirically, the DNABERT-2-117M classifier was trained with Optuna-tuned hyperparameters and evaluated on 350742 held-out sequences. The model reached PR-AUC 0.759, ROC-AUC 0.743, and best F1 0.704 at an optimized threshold (0.359), with recall 0.835 and precision 0.609. Against a CNN-based EnhancerNet trained on the same data, DNABERT-2 delivered stronger threshold-independent ranking and higher recall, although point accuracy was lower. To our knowledge, this is the first study to apply a second-generation genomic language model with BPE tokenization to enhancer classification in colorectal cancer, demonstrating the feasibility of capturing tumor-associated regulatory signals directly from DNA sequence alone. Overall, our results show that transformer-based genomic models can move beyond motif-level encodings toward holistic classification of regulatory elements, offering a novel path for cancer genomics. Next steps will focus on improving precision, exploring hybrid CNN-transformer designs, and validating across independent datasets to strengthen real-world utility.

📄 PDF Abstract BibTeX arXiv:2509.25274

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasks

2026-06-29 · Romain Karpinsky, Julien Mozziconacci, Mickaël Delcey arxiv

Recent breakthroughs in foundation models and Large Language Models (LLMs) have introduced new opportunities for studying and decoding genomic sequences. Several state-of-the-art approaches, such as DNABERT2, rely on tra…

Embedding Is (Almost) All You Need: Retrieval-Augmented Inference for Generalizable Genomic Prediction Tasks

2025-08-06 · Nirjhor Datta, Swakkhar Shatabda, M Sohel Rahman arxiv

Large pre-trained DNA language models such as DNABERT-2, Nucleotide Transformer, and HyenaDNA have demonstrated strong performance on various genomic benchmarks. However, most applications rely on expensive fine-tuning, …

Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation Models

2025-09-13 · Weimin Wu, Xuefeng Song, Yibo Wen, Qinjie Lin 외 arxiv

We introduce Genome-Factory, the first integrated Python library for tuning, deploying, and interpreting genomic foundation models. Our core contribution is to simplify and unify the workflow for genomic model developmen…

parameter-efficient fine-tuning

DNABERT-S: Pioneering Species Differentiation with Species-Aware DNA Embeddings

2024-02-13 · Zhihan Zhou, Weimin Wu, Harrison Ho, Jiayi Wang 외

We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic se…

Contrastive Learning

GeneMask: Fast Pretraining of Gene Sequences to Enable Few-Shot Learning

2023-07-29 · Soumyadeep Roy, Jonas Wallat, Sowmya S Sundaram, Wolfgang Nejdl 외

Large-scale language models such as DNABert and LOGO aim to learn optimal gene representations and are trained on the entire Human Reference Genome. However, standard tokenization schemes involve a simple sliding window …

Few-Shot LearningLanguage ModelingLanguage ModellingMasked Language Modeling