paper-with-me

홈 › Papers

HybriDNA: A Hybrid Transformer-Mamba2 Long-Range DNA Language Model

2025-02-15 · Mingqian Ma, Guoqing Liu, Chuan Cao, Pan Deng, Tri Dao, Albert Gu, Peiran Jin, Zhao Yang, Yingce Xia, Renqian Luo, Pipi Hu, Zun Wang, Yuan-Jyue Chen, Haiguang Liu, Tao Qin

Advances in natural language processing and large language models have sparked growing interest in modeling DNA, often referred to as the "language of life". However, DNA modeling poses unique challenges. First, it requires the ability to process ultra-long DNA sequences while preserving single-nucleotide resolution, as individual nucleotides play a critical role in DNA function. Second, success in this domain requires excelling at both generative and understanding tasks: generative tasks hold potential for therapeutic and industrial applications, while understanding tasks provide crucial insights into biological mechanisms and diseases. To address these challenges, we propose HybriDNA, a decoder-only DNA language model that incorporates a hybrid Transformer-Mamba2 architecture, seamlessly integrating the strengths of attention mechanisms with selective state-space models. This hybrid design enables HybriDNA to efficiently process DNA sequences up to 131kb in length with single-nucleotide resolution. HybriDNA achieves state-of-the-art performance across 33 DNA understanding datasets curated from the BEND, GUE, and LRB benchmarks, and demonstrates exceptional capability in generating synthetic cis-regulatory elements (CREs) with desired properties. Furthermore, we show that HybriDNA adheres to expected scaling laws, with performance improving consistently as the model scales from 300M to 3B and 7B parameters. These findings underscore HybriDNA's versatility and its potential to advance DNA research and applications, paving the way for innovations in understanding and engineering the "language of life".

📄 PDF Abstract BibTeX arXiv:2502.10807

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingState Space Models

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

HybridTM: Combining Transformer and Mamba for 3D Semantic Segmentation

2025-07-24 · Xinyu Wang, Jinghua Hou, Zhe Liu, Yingying Zhu arxiv

Transformer-based methods have demonstrated remarkable capabilities in 3D semantic segmentation through their powerful attention mechanisms, but the quadratic complexity limits their modeling of long-range dependencies i…

3D Semantic SegmentationPoint Clouds

MatIR: A Hybrid Mamba-Transformer Image Restoration Model

2025-01-30 · Juan Wen, Weiyan Hou, Luc van Gool, Radu Timofte

In recent years, Transformers-based models have made significant progress in the field of image restoration by leveraging their inherent ability to capture complex contextual features. Recently, Mamba models have made a …

Computational EfficiencyImage InpaintingImage RestorationMamba+1

SST: Multi-Scale Hybrid Mamba-Transformer Experts for Long-Short Range Time Series Forecasting

2024-04-23 · Xiongxiao Xu, Canyu Chen, Yueqing Liang, Baixiang Huang 외

Despite significant progress in time series forecasting, existing forecasters often overlook the heterogeneity between long-range and short-range time series, leading to performance degradation in practical applications.…

MambaTime SeriesTime Series ForecastingWeather Forecasting

HMT-UNet: A hybird Mamba-Transformer Vision UNet for Medical Image Segmentation

2024-08-21 · Mingya Zhang, Zhihao Chen, Yiyuan Ge, Xianping Tao

In the field of medical image segmentation, models based on both CNN and Transformer have been thoroughly investigated. However, CNNs have limited modeling capabilities for long-range dependencies, making it challenging …

Image SegmentationMambaMedical Image SegmentationSegmentation+2

MambaVesselNet++: A Hybrid CNN-Mamba Architecture for Medical Image Segmentation

2025-07-26 · Qing Xu, Yanming Chen, Yue Li, Ziyu Liu 외 arxiv

Medical image segmentation plays an important role in computer-aided diagnosis. Traditional convolution-based U-shape segmentation architectures are usually limited by the local receptive field. Existing vision transform…

Medical Image SegmentationInstance Segmentation