paper-with-me

Papers

SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration

2026-07-23 · Yinhao Tang, Youqing Fang, Yanan Sun, Wenran Liu, Weiming Zhang, Bin Liu, Kuikun Liu, Wenwei Zhang, Kai Chen arxiv

Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retrieval or static scientific question answering, and therefore fail to assess key capabilities required in realistic scientific research workflows. We introduce SciExplore, a benchmark designed to evaluate scientific information-seeking and reasoning capabilities of LLMs and agents. SciExplore comprises four task types covering 103 expert-curated tasks across more than ten scientific disciplines: scientific database navigation, ambiguous literature retrieval, missing reference completion, and cross-source structured knowledge synthesis, which probe progressively higher-level abilities from entity-level reasoning and document-level identification to evidence-level grounding and domain-level synthesis. We evaluate over ten state-of-the-art LLMs and autonomous agents on SciExplore, revealing substantial performance gaps with performance degrading sharply as task complexity increases and extremely low accuracy on the most challenging structured synthesis tasks. These results highlight significant limitations of current models and agents in realistic scientific information-seeking scenarios.

📄 PDF Abstract BibTeX arXiv:2607.20926

Code (3)

Aaron617/agent-arXiv-daily ★ 10
Tavish9/awesome-daily-AI-arxiv ★ 112
arxivsub/arXivSub_daily_arxiv ★ 3

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Agentic Exploration of Physics Models

2025-09-29 · Maximilian Nägele, Florian Marquardt arxiv

The process of scientific discovery relies on an interplay of observations, analysis, and hypothesis generation. Machine learning is increasingly being adopted to address individual aspects of this process. However, it r…

Benchmarking Visual Localization for Autonomous Navigation

2022-03-24 · Lauri Suomela, Jussi Kalliola, Atakan Dag, Harry Edelman 외

This work introduces a simulator-based benchmark for visual localization in the autonomous navigation context. The dynamic benchmark enables investigation of how variables such as the time of day, weather, and camera per…

Autonomous NavigationBenchmarkingMotion PlanningVisual Localization+1

FIRE-Bench: Evaluating Agents on the Rediscovery of Scientific Insights

2026-02-02 · Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su 외 arxiv

Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge. Existing ben…

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

2026-05-28 · Wanghan Xu, Shuo Li, Tianlin Ye, Qinglong Cao 외 arxiv

AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify. We present ResearchClawBench, a benchmark for evaluating autonomous scientific …

Recent Advancements in Deep Learning Applications and Methods for Autonomous Navigation: A Comprehensive Review

2023-02-22 · Arman Asgharpoor Golroudbari, Mohammad Hossein Sabour

This review article is an attempt to survey all recent AI based techniques used to deal with major functions in This review paper presents a comprehensive overview of end-to-end deep learning frameworks used in the conte…

Autonomous NavigationAutonomous VehiclesDeep Learning