paper-with-me

Papers

Toward Evaluation Frameworks for Multi-Agent Scientific AI Systems

2026-03-18 · Marcin Abram arxiv

We analyze the challenges of benchmarking scientific (multi)-agentic systems, including the difficulty of distinguishing reasoning from retrieval, the risks of data/model contamination, the lack of reliable ground truth for novel research problems, the complications introduced by tool use, and the replication challenges due to the continuously changing/updating knowledge base. We discuss strategies for constructing contamination-resistant problems, generating scalable families of tasks, and the need for evaluating systems through multi-turn interactions that better reflect real scientific practice. As an early feasibility test, we demonstrate how to construct a dataset of novel research ideas to test the out-of-sample performance of our system. We also discuss the results of interviews with several researchers and engineers working in quantum science. Through those interviews, we examine how scientists expect to interact with AI systems and how these expectations should shape evaluation methods.

📄 PDF Abstract BibTeX arXiv:2603.26718

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios

2025-08-25 · Bingxi Zhao, Lin Geng Foo, Ping Hu, Christian Theobalt 외 arxiv

Recent advances in the intrinsic reasoning capabilities of large language models (LLMs) have given rise to LLM-based agent systems that exhibit near-human performance on a variety of automated tasks. However, although th…

INDIBATOR: Diverse and Fact-Grounded Individuality for Multi-Agent Debate in Molecular Discovery

2026-02-02 · Yunhui Jang, Seonghyun Park, Jaehyung Kim, Sungsoo Ahn arxiv

Multi-agent systems have emerged as a powerful paradigm for automating scientific discovery. To differentiate agent behavior in the multi-agent system, current frameworks typically assign generic role-based personas such…

Towards Scientific Discovery with Generative AI: Progress, Opportunities, and Challenges

2024-12-16 · Chandan K Reddy, Parshin Shojaee

Scientific discovery is a complex cognitive process that has driven human knowledge and technological progress for centuries. While artificial intelligence (AI) has made significant advances in automating aspects of scie…

Automated Theorem Provingscientific discovery

Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions

2025-03-12 · Mourad Gridach, Jay Nanavati, Khaldoun Zine El Abidine, Lenon Mendes 외

The integration of Agentic AI into scientific discovery marks a new frontier in research automation. These AI systems, capable of reasoning, planning, and autonomous decision-making, are transforming how scientists perfo…

Decision Makingscientific discovery

Survey on Evaluation of LLM-based Agents

2025-03-20 · Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel 외

The emergence of LLM-based agents represents a paradigm shift in AI, enabling autonomous systems to plan, reason, use tools, and maintain memory while interacting with dynamic environments. This paper provides the first …

Survey