paper-with-me

홈 › Papers

CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation

2025-03-20 · Peter Jansen, Oyvind Tafjord, Marissa Radensky, Pao Siangliulue, Tom Hope, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Daniel S. Weld, Peter Clark

Despite the surge of interest in autonomous scientific discovery (ASD) of software artifacts (e.g., improved ML algorithms), current ASD systems face two key limitations: (1) they largely explore variants of existing codebases or similarly constrained design spaces, and (2) they produce large volumes of research artifacts (such as automatically generated papers and code) that are typically evaluated using conference-style paper review with limited evaluation of code. In this work we introduce CodeScientist, a novel ASD system that frames ideation and experiment construction as a form of genetic search jointly over combinations of research articles and codeblocks defining common actions in a domain (like prompting a language model). We use this paradigm to conduct hundreds of automated experiments on machine-generated ideas broadly in the domain of agents and virtual environments, with the system returning 19 discoveries, 6 of which were judged as being both at least minimally sound and incrementally novel after a multi-faceted evaluation beyond that typically conducted in prior work, including external (conference-style) review, code review, and replication attempts. Moreover, the discoveries span new tasks, agents, metrics, and data, suggesting a qualitative shift from benchmark optimization to broader discoveries.

📄 PDF Abstract BibTeX arXiv:2503.22708

Code (0)

등록된 구현이 없습니다.

Tasks

Articlesscientific discovery

Similar Papers 제목 키워드 기반

HARPA: A Testability-Driven, Literature-Grounded Framework for Research Ideation

2025-10-01 · Rosni Vasu, Peter Jansen, Pao Siangliulue, Cristina Sarasua 외 arxiv

While there has been a surge of interest in automated scientific discovery (ASD), especially with the emergence of LLMs, it remains challenging for tools to generate hypotheses that are both testable and grounded in the …

Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials Science

2025-06-04 · Peter Jansen, Samiah Hassan, Ruoyao Wang

Contemporary approaches to assisted scientific discovery use language models to automatically generate large numbers of potential hypothesis to test, while also automatically generating code-based experiments to test tho…

ArticlesCode GenerationRetrieval-augmented Generationscientific discovery

DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents

2024-06-10 · Peter Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom 외

Automated scientific discovery promises to accelerate progress across scientific domains. However, developing and evaluating an AI agent's capacity for end-to-end scientific reasoning is challenging as running real-world…

Benchmarkingscientific discovery

PiFlow: Principle-aware Scientific Discovery with Multi-Agent Collaboration

2025-05-21 · Yingming Pu, Tao Lin, Hongyu Chen

Large Language Model (LLM)-based multi-agent systems (MAS) demonstrate remarkable potential for scientific discovery. Existing approaches, however, often automate scientific discovery using predefined workflows that lack…

Large Language Modelscientific discovery

Integration of Scanning Probe Microscope with High-Performance Computing: fixed-policy and reward-driven workflows implementation

2024-05-20 · Yu Liu, Utkarsh Pratiush, Jason Bemis, Roger Proksch 외

The rapid development of computation power and machine learning algorithms has paved the way for automating scientific discovery with a scanning probe microscope (SPM). The key elements towards operationalization of auto…

scientific discovery