paper-with-me

홈 › Papers

SciAgent: Tool-augmented Language Models for Scientific Reasoning

2024-02-18 · Yubo Ma, Zhibin Gou, Junheng Hao, Ruochen Xu, Shuohang Wang, Liangming Pan, Yujiu Yang, Yixin Cao, Aixin Sun, Hany Awadalla, Weizhu Chen

Scientific reasoning poses an excessive challenge for even the most advanced Large Language Models (LLMs). To make this task more practical and solvable for LLMs, we introduce a new task setting named tool-augmented scientific reasoning. This setting supplements LLMs with scalable toolsets, and shifts the focus from pursuing an omniscient problem solver to a proficient tool-user. To facilitate the research of such setting, we construct a tool-augmented training corpus named MathFunc which encompasses over 30,000 samples and roughly 6,000 tools. Building on MathFunc, we develop SciAgent to retrieve, understand and, if necessary, use tools for scientific problem solving. Additionally, we craft a benchmark, SciToolBench, spanning five scientific domains to evaluate LLMs' abilities with tool assistance. Extensive experiments on SciToolBench confirm the effectiveness of SciAgent. Notably, SciAgent-Mistral-7B surpasses other LLMs with the same size by more than 13% in absolute accuracy. Furthermore, SciAgent-DeepMath-7B shows much superior performance than ChatGPT.

📄 PDF Abstract BibTeX arXiv:2402.11451

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning

2025-11-11 · Xuchen Li, Ruitao Wu, Xuanbo Liu, Xukai Wang 외 arxiv

Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified …

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents

2026-02-13 · Yujiong Shen, Yajie Yang, Zhiheng Xi, Binze Hu 외 arxiv

Scientific reasoning inherently demands integrating sophisticated toolkits to navigate domain-specific knowledge. Yet, current benchmarks largely overlook agents' ability to orchestrate tools for such rigorous workflows.…

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales

2026-06-10 · Tianyu Liu, Allen Xin Wang, Antonia Panescu, Lisa Xinyi Chen 외 arxiv

AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchmarks for AI agents rarely capture the com…

SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning

2024-09-09 · Alireza Ghafarollahi, Markus J. Buehler

A key challenge in artificial intelligence is the creation of systems capable of autonomously advancing scientific understanding by exploring novel domains, identifying complex patterns, and uncovering previously unseen …

AI AgentKnowledge Graphsscientific discovery

SciTrace: Trajectory-Aware Safety Reasoning for Scientific Discovery Agents

2026-06-06 · Tanush Swaminathan, Runmin Jiang, Letian Zhang, Min Xu arxiv

LLM-based scientific agents have shown strong capacity for autonomous research, yet their safety layers remain structurally divorced from core reasoning: they inspect pipeline outputs rather than shaping the deliberation…

Adversarial Robustness