paper-with-me

홈 › Papers

SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs

2026-03-11 · Yadi Cao, Sicheng Lai, Jiahe Huang, Yang Zhang, Zach Lawrence, Rohan Bhakta, Izzy F. Thomas, Mingyun Cao, Chung-Hao Tsai, Zihao Zhou, Yidong Zhao, Hao Liu, Alessandro Marinoni, Alexey Arefiev, Rose Yu arxiv

Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@k become impractical under realistic budget constraints. To address this gap, we introduce SimulCost, the first benchmark targeting cost-sensitive parameter tuning in physics simulations. SimulCost compares LLM tuning cost-sensitive parameters against traditional scanning approach in both accuracy and computational cost, spanning 2,643 single-round (initial guess) and 2,304 multi-round (adjustment by trial-and-error) tasks across 11 simulators from fluid dynamics, solid mechanics, and plasma physics, whose costs are analytically defined and platform-independent. A twelfth simulator, a production plasma code measurable only by wall clock, is reported separately. Frontier LLMs achieve 45-62% success rates in single-round mode, dropping to 34-50% under high accuracy requirements, rendering their initial guesses unreliable especially for high accuracy tasks. Multi-round mode improves rates to 66-81%, but LLMs are 1.5-2.7x slower than traditional scanning, making them uneconomical choices. We also investigate parameter group correlations for knowledge transfer potential, and the impact of in-context examples and reasoning effort, providing practical implications for deployment and fine-tuning. We open-source SimulCost as a static benchmark and extensible toolkit to facilitate research on improving cost-aware agentic designs for physics simulations, and for expanding new simulation environments. Code and data are available at https://github.com/Rose-STL-Lab/SimulCost-Bench

📄 PDF Abstract BibTeX arXiv:2603.20253

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AutoNMT: A Framework to Streamline the Research of Seq2Seq Models

2023-02-09 · Salvador Carrión, Francisco Casacuberta

We present AutoNMT, a framework to streamline the research of seq-to-seq models by automating the data pipeline (i.e., file management, data preprocessing, and exploratory analysis), automating experimentation in a toolk…

Management

MolMole: Molecule Mining from Scientific Literature

2025-04-30 · LG AI Research, Sehyun Chun, Jiye Kim, Ahra Jo 외

The extraction of molecular structures and reaction data from scientific documents is challenging due to their varied, unstructured chemical formats and complex document layouts. To address this, we introduce MolMole, a …

CEBench: A Benchmarking Toolkit for the Cost-Effectiveness of LLM Pipelines

2024-06-20 · Wenbo Sun, Jiaqi Wang, Qiming Guo, Ziyu Li 외

Online Large Language Model (LLM) services such as ChatGPT and Claude 3 have transformed business operations and academic research by effortlessly enabling new opportunities. However, due to data-sharing restrictions, se…

BenchmarkingDecision MakingLanguage ModelingLanguage Modelling+1

NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning

2026-01-07 · Zhongtao Miao, Kaiyan Zhao, Masaaki Nagata, Yoshimasa Tsuruoka arxiv

Neologism-aware machine translation aims to translate source sentences containing neologisms into target languages. This field remains underexplored compared with general machine translation (MT). In this paper, we propo…

Reinforcement LearningMachine Translation

FairDiverse: A Comprehensive Toolkit for Fair and Diverse Information Retrieval Algorithms

2025-02-17 · Chen Xu, Zhirui Deng, Clara Rus, Xiaopeng Ye 외

In modern information retrieval (IR). achieving more than just accuracy is essential to sustaining a healthy ecosystem, especially when addressing fairness and diversity considerations. To meet these needs, various datas…

DiversityFairnessInformation RetrievalRetrieval