paper-with-me

홈 › Papers

Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents

2025-02-22 · Patrick Tser Jern Kon, Jiachen Liu, Qiuyi Ding, Yiming Qiu, Zhenning Yang, Yibo Huang, Jayanth Srinivasa, Myungjin Lee, Mosharaf Chowdhury, Ang Chen

Scientific experimentation, a cornerstone of human progress, demands rigor in reliability, methodical control, and interpretability to yield meaningful results. Despite the growing capabilities of large language models (LLMs) in automating different aspects of the scientific process, automating rigorous experimentation remains a significant challenge. To address this gap, we propose Curie, an AI agent framework designed to embed rigor into the experimentation process through three key components: an intra-agent rigor module to enhance reliability, an inter-agent rigor module to maintain methodical control, and an experiment knowledge module to enhance interpretability. To evaluate Curie, we design a novel experimental benchmark composed of 46 questions across four computer science domains, derived from influential research papers, and widely adopted open-source projects. Compared to the strongest baseline tested, we achieve a 3.4$\times$ improvement in correctly answering experimental questions.Curie is open-sourced at https://github.com/Just-Curieous/Curie.

📄 PDF Abstract BibTeX arXiv:2502.16069

Code (1)

just-curieous/curie 공식 구현

Tasks

AI Agent

Similar Papers 제목 키워드 기반

EXP-Bench: Can AI Conduct AI Research Experiments?

2025-05-30 · Patrick Tser Jern Kon, Jiachen Liu, Xinyi Zhu, Qiuyi Ding 외

Automating AI research holds immense potential for accelerating scientific progress, yet current AI agents struggle with the complexities of rigorous, end-to-end experimentation. We introduce EXP-Bench, a novel benchmark…

CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

2025-03-14 · HAO CUI, Zahra Shamsi, Gowoon Cheon, Xuejian Ma 외

Scientific problem-solving involves synthesizing information while applying expert knowledge. We introduce CURIE, a scientific long-Context Understanding,Reasoning and Information Extraction benchmark to measure the pote…

Long-Context Understanding

ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System

2026-07-15 · Yutong He, Daibo Li, Guohong Li, Jiahe Geng 외 arxiv

Recent advances in Large Language Models have fueled autonomous AI agents capable of tackling complex scientific tasks, yet existing automated research systems remain predominantly focused on empirically driven domains w…

Explainable AI for Curie Temperature Prediction in Magnetic Materials

2025-08-09 · M. Adeel Ajaib, Fariha Nasir, Abdul Rehman arxiv

We explore machine learning techniques for predicting Curie temperatures of magnetic materials using the NEMAD database. By augmenting the dataset with composition-based and domain-aware descriptors, we evaluate the perf…

Finding Symmetry Breaking Order Parameters with Euclidean Neural Networks

2020-07-04 · Tess E. Smidt, Mario Geiger, Benjamin Kurt Miller

Curie's principle states that "when effects show certain asymmetry, this asymmetry must be found in the causes that gave rise to them". We demonstrate that symmetry equivariant neural networks uphold Curie's principle an…