paper-with-me

홈 › Papers

$A^3$-Bench: Benchmarking Memory-Driven Scientific Reasoning via Anchor and Attractor Activation

2026-01-14 · Jian Zhang, Yu He, Zhiyuan Wang, Zhangqi Wang, Kai He, Fangzhi Xu, Qika Lin, Jun Liu arxiv

Scientific reasoning relies not only on logical inference but also on activating prior knowledge and experiential structures. Memory can efficiently reuse knowledge and enhance reasoning consistency and stability. However, existing benchmarks mainly evaluate final answers or step-by-step coherence, overlooking the \textit{memory-driven} mechanisms that underlie human reasoning, which involves activating anchors and attractors, then integrating them into multi-step inference. To address this gap, we propose $A^3$-Bench~ https://a3-bench.github.io, a benchmark designed to evaluate scientific reasoning through dual-scale memory-driven activation, grounded in Anchor and Attractor Activation. First, we annotate 2,198 science reasoning problems across domains using the SAPM process(subject, anchor & attractor, problem, and memory developing). Second, we introduce a dual-scale memory evaluation framework utilizing anchors and attractors, along with the AAUI(Anchor--Attractor Utilization Index) metric to measure memory activation rates. Finally, through experiments with various base models and paradigms, we validate $A^3$-Bench and analyze how memory activation impacts reasoning performance, providing insights into memory-driven scientific reasoning.

📄 PDF Abstract BibTeX arXiv:2601.09274

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility

2026-01-17 · Honglin Lin, Chonghan Qin, Zheng Liu, Qizhi Pei 외 arxiv

While synthetic data has proven effective for improving scientific reasoning in the text domain, multimodal reasoning remains constrained by the difficulty of synthesizing scientifically rigorous images. Existing Text-to…

Multimodal Reasoning

Benchmarking AI scientists in omics data-driven biological research

2025-05-13 · Erpai Luo, Jinmeng Jia, Yifan Xiong, Xiangyu Li 외

The rise of large language models and multi-agent systems has sparked growing interest in AI scientists capable of autonomous biological research. However, existing benchmarks either focus on reasoning without data or on…

BenchmarkingMultiple-choicescientific discovery

SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence

2025-12-26 · Yiheng Wang, Yixin Chen, Shuo Li, Yifan Zhou 외 arxiv

We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike general-purpose evaluation platforms, SciEva…

Multimodal ReasoningCode Generation

AInsteinBench: Benchmarking Coding Agents on Scientific Repositories

2025-12-24 · Titouan Duston, Shuo Xin, Yang Sun, Daoguang Zan 외 arxiv

We introduce AInsteinBench, a large-scale benchmark for evaluating whether large language model (LLM) agents can operate as scientific computing development agents within real research software ecosystems. Unlike existin…

Code Generation

SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models

2024-06-13 · Kehua Feng, Keyan Ding, Weijie Wang, Xiang Zhuang 외

Large language models (LLMs) have gained increasing prominence in scientific research, but there is a lack of comprehensive benchmarks to fully evaluate their proficiency in understanding and mastering scientific knowled…

Benchmarking