paper-with-me

Papers

Lean4Physics: Comprehensive Reasoning Framework for College-level Physics in Lean4

2025-10-30 · Yuxin Li, Minghao Liu, Ruida Wang, Wenzhao Ji, Zhitao He, Rui Pan, Junming Huang, Tong Zhang, Yi R. Fung arxiv

We present Lean4PHYS, a comprehensive reasoning framework for college-level physics problems in Lean4. Lean4PHYS includes *LeanPhysBench*, a college-level benchmark for formal physics reasoning in Lean4, which contains 200 hand-crafted and peer-reviewed statements derived from university textbooks and physics competition problems. To establish a solid foundation for formal reasoning in physics, we also introduce *PhysLib*, a community-driven repository containing fundamental unit systems and theorems essential for formal physics reasoning. Based on the benchmark and Lean4 repository we composed in Lean4PHYS, we report baseline results using major expert Math Lean4 provers and state-of-the-art closed-source models, with the best performance of DeepSeek-Prover-V2-7B achieving only 16% and Claude-Sonnet-4 achieving 35%. We also conduct a detailed analysis showing that our *PhysLib* can achieve an average improvement of 11.75% in model performance. This demonstrates the challenging nature of our *LeanPhysBench* and the effectiveness of *PhysLib*. To the best of our knowledge, this is the first study to provide a physics benchmark in Lean4.

📄 PDF Abstract BibTeX arXiv:2510.26094

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark

2024-02-06 · Zhenwen Liang, Kehan Guo, Gang Liu, Taicheng Guo 외

The paper introduces SceMQA, a novel benchmark for scientific multimodal question answering at the college entrance level. It addresses a critical educational phase often overlooked in existing benchmarks, spanning high …

Multiple-choiceQuestion Answering

CLR-Bench: Evaluating Large Language Models in College-level Reasoning

2024-10-23 · Junnan Dong, Zijin Hong, Yuanchen Bei, Feiran Huang 외

Large language models (LLMs) have demonstrated their remarkable performance across various language understanding tasks. While emerging benchmarks have been proposed to evaluate LLMs in various domains such as mathematic…

PhysProver: Advancing Automatic Theorem Proving for Physics

2026-01-22 · Hanning Zhang, Ruida Wang, Rui Pan, Wenyuan Wang 외 arxiv

The combination of verifiable languages and LLMs has significantly influenced both the mathematical and computer science communities because it provides a rigorous foundation for theorem proving. Recent advancements in t…

Reinforcement LearningMathematical Reasoning

OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

2024-02-21 · Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu 외

Recent advancements have seen Large Language Models (LLMs) and Large Multimodal Models (LMMs) surpassing general human capabilities in various tasks, approaching the proficiency level of human experts across multiple dom…

Logical Fallacies

PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning

2025-02-17 · Xinyu Zhang, Yuxuan Dong, Yanrui Wu, Jiaxing Huang 외

Large language models demonstrate remarkable capabilities across various domains, especially mathematics and logic reasoning. However, current evaluations overlook physics-based reasoning - a complex task requiring physi…