paper-with-me

Papers

Towards a Large Physics Benchmark

2025-07-29 · Kristian G. Barman, Sascha Caron, Faegheh Hasibi, Eugene Shalugin, Yoris Marcet, Johannes Otte, Henk W. de Regt, Merijn Moody arxiv

We introduce a benchmark framework developed by and for the scientific community to evaluate, monitor and steer large language model development in fundamental physics. Building on philosophical concepts of scientific understanding and creativity, we develop a scoring system in which each question is scored by an expert for its correctness, difficulty, and surprise. The questions are of three forms: (i) multiple-choice questions for conceptual understanding, (ii) analytical problems requiring mathematical derivation, and (iii) openended tasks requiring complex problem solving. Our current dataset contains diverse set of examples, including a machine learning challenge to classify high-energy physics events, such as the four top quark signal. To ensure continued relevance, we propose a living benchmark, where physicists contribute questions, for instance alongside new publications. We invite contributions via: http://www.physicsbenchmarks.org/. We hope that this benchmark will enable a targeted AI development that can make a meaningful contribution to fundamental physics research.

📄 PDF Abstract BibTeX arXiv:2507.21695

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models

2025-02-01 · Xin Xu, Qiyun Xu, Tong Xiao, Tianhao Chen 외

Large language models (LLMs) have demonstrated remarkable capabilities in solving complex reasoning tasks, particularly in mathematics. However, the domain of physics reasoning presents unique challenges that have receiv…

Math

PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions

2025-05-21 · Song Dai, Yibo Yan, Jiamin Su, Dongfang Zihao 외

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in diverse reasoning tasks, yet their application to complex physics reasoning remains underexplored. Physics reasoning presents unique c…

PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning

2025-02-17 · Xinyu Zhang, Yuxuan Dong, Yanrui Wu, Jiaxing Huang 외

Large language models demonstrate remarkable capabilities across various domains, especially mathematics and logic reasoning. However, current evaluations overlook physics-based reasoning - a complex task requiring physi…

OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora

2026-08-26 · Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks. However, their development in the physics domain is significantly hindered by the lack of…

PhysUniBench: An Undergraduate-Level Physics Reasoning Benchmark for Multimodal Models

2025-06-21 · Lintao Wang, Encheng Su, Jiaqi Liu, Pengze Li 외

Physics problem-solving is a challenging domain for large AI models, requiring integration of conceptual understanding, mathematical reasoning, and interpretation of physical diagrams. Current evaluation methodologies sh…

Mathematical ReasoningMultiple-choice