paper-with-me

홈 › Papers

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics

2025-09-19 · Nithin Somasekharan, Ling Yue, Yadi Cao, Weichao Li, Patrick Emami, Pochinapeddi Sai Bhargav, Anurag Acharya, Xingyu Xie, Shaowu Pan arxiv

Large Language Models (LLMs) have demonstrated strong performance across general NLP tasks, but their utility in automating numerical experiments of complex physical system -- a critical and labor-intensive component -- remains underexplored. As the major workhorse of computational science over the past decades, Computational Fluid Dynamics (CFD) offers a uniquely challenging testbed for evaluating the scientific capabilities of LLMs. We introduce CFDLLMBench, a benchmark suite comprising three complementary components -- CFDQuery, CFDCodeBench, and FoamBench -- designed to holistically evaluate LLM performance across three key competencies: graduate-level CFD knowledge, numerical and physical reasoning of CFD, and context-dependent implementation of CFD workflows. Grounded in real-world CFD practices, our benchmark combines a detailed task taxonomy with a rigorous evaluation framework to deliver reproducible results and quantify LLM performance across code executability, solution accuracy, and numerical convergence behavior. CFDLLMBench establishes a solid foundation for the development and evaluation of LLM-driven automation of numerical experiments for complex physical systems. Code and data are available at https://github.com/NREL-Theseus/cfdllmbench/.

📄 PDF Abstract BibTeX arXiv:2509.20374

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SWE-Mutation: Can LLMs Generate Reliable Test Suites in Software Engineering?

2026-05-21 · Yuxuan Sun, Yuze Zhao, Yufeng Wang, Yao Du 외 arxiv

Evaluating software engineering capabilities has become a core component of modern large language models (LLMs); however, the key bottleneck hindering further scaling lies not in the scarcity of high-quality solutions, b…

Reinforcement LearningProgram Repair

Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis

2025-08-27 · Anusha Kamath, Kanishk Singla, Rakesh Paul, Raviraj Joshi 외 arxiv

Evaluating instruction-tuned Large Language Models (LLMs) in Hindi is challenging due to a lack of high-quality benchmarks, as direct translation of English datasets fails to capture crucial linguistic and cultural nuanc…

MTG: A Benchmarking Suite for Multilingual Text Generation

2021-10-16 · ACL ARR October 2021 10 · Anonymous

We introduce MTG, a new benchmark suite for training and evaluating multilingual text generation. It is the first and largest multilingual multiway text generation benchmark with 400k human-annotated data for four tasks …

BenchmarkingQuestion GenerationQuestion-GenerationStory Generation+3

MTG: A Benchmark Suite for Multilingual Text Generation

2021-08-13 · Findings (NAACL) 2022 7 · Yiran Chen, Zhenqiao Song, Xianze Wu, Danqing Wang 외

We introduce MTG, a new benchmark suite for training and evaluating multilingual text generation. It is the first-proposed multilingual multiway text generation dataset with the largest human-annotated data (400k). It in…

Question GenerationQuestion-GenerationStory GenerationText Generation+2

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

2026-07-29 · Jingbo Zhou, Yusai Zhao, Qi Bao, Jingjia Cao 외 hf

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at …