paper-with-me

홈 › Papers

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction

2025-11-04 · Zhongmin Li, Runze Ma, Jiahao Tan, Chengzi Tan, Shuangjia Zheng arxiv

Nucleotide sequence variation can induce significant shifts in functional fitness. Recent nucleotide foundation models promise to predict such fitness effects directly from sequence, yet heterogeneous datasets and inconsistent preprocessing make it difficult to compare methods fairly across DNA and RNA families. Here we introduce NABench, a large-scale, systematic benchmark for nucleic acid fitness prediction. NABench aggregates 162 high-throughput assays and curates 2.6 million mutated sequences spanning diverse DNA and RNA families, with standardized splits and rich metadata. We show that NABench surpasses prior nucleotide fitness benchmarks in scale, diversity, and data quality. Under a unified evaluation suite, we rigorously assess 29 representative foundation models across zero-shot, few-shot prediction, transfer learning, and supervised settings. The results quantify performance heterogeneity across tasks and nucleic-acid types, demonstrating clear strengths and failure modes for different modeling choices and establishing strong, reproducible baselines. We release NABench to advance nucleic acid modeling, supporting downstream applications in RNA/DNA design, synthetic biology, and biochemistry. Our code is available at https://github.com/mrzzmrzz/NABench.

📄 PDF Abstract BibTeX arXiv:2511.02888

Code (0)

등록된 구현이 없습니다.

Tasks

Transfer Learning

Similar Papers 제목 키워드 기반

ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation

2025-10-09 · Qin Liu, Jacob Dineen, Yuxi Huang, Sheng Zhang 외 arxiv

Benchmarks are central to measuring the capabilities of large language models and guiding model development, yet widespread data leakage from pretraining corpora undermines their validity. Models can match memorized cont…

Dynabench: Rethinking Benchmarking in NLP

2021-04-07 · NAACL 2021 4 · Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik 외

We introduce Dynabench, an open-source platform for dynamic dataset creation and model benchmarking. Dynabench runs in a web browser and supports human-and-model-in-the-loop dataset creation: annotators seek to create ex…

Benchmarking

HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution

2023-06-27 · NeurIPS 2023 11 · Eric Nguyen, Michael Poli, Marjan Faizi, Armin Thomas 외

Genomic (DNA) sequences encode an enormous amount of information for gene regulation and protein synthesis. Similar to natural language models, researchers have proposed foundation models in genomics to learn generalizab…

4kIn-Context LearningLanguage ModellingLarge Language Model

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models

2024-06-14 · Yuchen Ren, ZhiYuan Chen, Lifeng Qiao, Hongtai Jing 외

RNA plays a pivotal role in translating genetic instructions into functional outcomes, underscoring its importance in biological processes and disease mechanisms. Despite the emergence of numerous deep learning approache…

Language Modelling

ViroBench: Benchmarking Nucleotide Foundation Models on Viral Genomics Tasks

2026-05-25 · Dongxin Ye, Fang Hu, Han Hu, Shu Hu 외 arxiv

Nucleotide sequences constitute the fundamental genetic basis of biological systems, rendering viral genomic analysis critical for biomedical advancement. Despite progress in biological foundation models, specifically nu…