paper-with-me

홈 › Papers

Verifiable Benchmarking of Long-Horizon Spatial Biology

2026-05-27 · Ian Diks, Harihara Muralidharan, Tim Proctor, Kenny Workman arxiv

AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or localized analysis steps rather than end-to-end scientific reasoning over spatial measurements. We introduce SpatialBench-Long, a benchmark for long-horizon spatial biology in which agents must recover biological claims from raw or near-raw data and calibrated experimental context without prescribed methods. SpatialBench-Long contains 24 evaluations across primary pancreatic ductal adenocarcinoma (PDAC), engineered glioblastoma organoids and in vivo tumors, Cas9 lineage-traced lung adenocarcinoma, and mouse optic nerve aging/intervention systems, spanning CosMx, Visium, Xenium, multiplexed error-robust fluorescence in situ hybridization (MERFISH), single-cell RNA sequencing (scRNA-seq), Slide-seq, Slide-tags, histology, and lineage-recording data. Candidate claims are hardened through reproduction, independent scientist review, and trajectory inspection. Final answers are graded deterministically over controlled vocabularies and symbols with companion rubrics capturing progress through key analysis chokepoints. Across the SpatialBench-Long benchmark, three model-harness pairs tie at 8/72 runs (11.1\%): Gemini 3.5 Flash / Pi terminal coding harness, GPT-5.5 / Pi, and GPT-5.5 / OpenAI Codex. SpatialBench-Long tests whether agents can move beyond executing procedural analysis to deriving accurate scientific conclusions from complex spatial measurements.

📄 PDF Abstract BibTeX arXiv:2605.28065

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology

2026-06-25 · Ian Diks, Zhen Yang, Arjun Banerjee, Tim Proctor 외 arxiv

Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay context, and auxiliary evidence. Existing AI-biology benchm…

LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning

2026-04-15 · Sumeet Ramesh Motwani, Daniel Nichols, Charles London, Peggy Li 외 arxiv

As language models are increasingly deployed for complex autonomous tasks, their ability to reason accurately over longer horizons becomes critical. An essential component of this ability is planning and managing a long,…

DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints

2026-01-26 · Yinger Zhang, Shutong Jiang, Renhao Li, Jianhong Tu 외 arxiv

While agent evaluation has shifted toward long-horizon tasks, most benchmarks still emphasize local, step-level reasoning rather than the global constrained optimization (e.g., time and financial budgets) that demands ge…

SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception-Memory Integration in Embodied Environments

2026-04-24 · Chih-Ting Liao, Xi Xiao, Chunlei Meng, Zhangquan Chen 외 arxiv

Multimodal large language models (MLLMs) have advanced static visual--spatial reasoning, yet they often fail to preserve long-horizon spatial coherence in embodied settings where beliefs must be continuously revised from…

Semantic SegmentationSpatial Reasoning

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

2026-08-05 · Bohai Gu, Yueyang Yuan, Taiyi Wu, Dazhao Du 외 hf

Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods such as reinforcement learning (RL) can improve these models, but the…

Reinforcement Learning