paper-with-me

홈 › Papers

scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis

2026-02-09 · Kenny Workman, Zhen Yang, Harihara Muralidharan, Aidan Abdulali, Hannah Le arxiv

As single-cell RNA sequencing datasets grow in adoption, scale, and complexity, data analysis remains a bottleneck for many research groups. Although frontier AI agents have improved dramatically at software engineering and general data analysis, it remains unclear whether they can extract biological insight from messy, real-world single-cell datasets. We introduce scBench, a benchmark of 394 verifiable problems derived from practical scRNA-seq workflows spanning six sequencing platforms and seven task categories. Each problem provides a snapshot of experimental data immediately prior to an analysis step and a deterministic grader that evaluates recovery of a key biological result. Benchmark data on eight frontier models shows that accuracy ranges from 29-53%, with strong model-task and model-platform interactions. Platform choice affects accuracy as much as model choice, with 40+ percentage point drops on less-documented technologies. scBench complements SpatialBench to cover the two dominant single-cell modalities, serving both as a measurement tool and a diagnostic lens for developing agents that can analyze real scRNA-seq datasets faithfully and reproducibly.

📄 PDF Abstract BibTeX arXiv:2602.09063

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology

2026-06-25 · Ian Diks, Zhen Yang, Arjun Banerjee, Tim Proctor 외 arxiv

Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay context, and auxiliary evidence. Existing AI-biology benchm…

CSCBench: A PVC Diagnostic Benchmark for Commodity Supply Chain Reasoning

2026-01-05 · Yaxin Cui, Yuanqiang Zeng, Jiapeng Yan, Keling Lin 외 arxiv

Large Language Models (LLMs) have achieved remarkable success in general benchmarks, yet their competence in commodity supply chains (CSCs) -- a domain governed by institutional rule systems and feasibility constraints -…

SCBench: A KV Cache-Centric Analysis of Long-Context Methods

2024-12-13 · Yucheng Li, Huiqiang Jiang, Qianhui Wu, Xufang Luo 외

Long-context LLMs have enabled numerous downstream applications but also introduced significant challenges related to computational and memory efficiency. To address these challenges, optimizations for long-context infer…

MambaQuantizationRetrievalSemantic Retrieval

scPilot: Large Language Model Reasoning Toward Automated Single-Cell Analysis and Discovery

2026-02-12 · Yiming Gao, Zhen Wang, Jefferson Chen, Mark Antkowiak 외 arxiv

We present scPilot, the first systematic framework to practice omics-native reasoning: a large language model (LLM) converses in natural language while directly inspecting single-cell RNA-seq data and on-demand bioinform…

VSCBench: Bridging the Gap in Vision-Language Model Safety Calibration

2025-05-26 · Jiahui Geng, Qing Li, Zongxiong Chen, Yuxia Wang 외

The rapid advancement of vision-language models (VLMs) has brought a lot of attention to their safety alignment. However, existing methods have primarily focused on model undersafety, where the model responds to hazardou…

Language ModelingLanguage ModellingSafety Alignment