paper-with-me

홈 › Papers

Evaluating GPT- and Reasoning-based Large Language Models on Physics Olympiad Problems: Surpassing Human Performance and Implications for Educational Assessment

2025-05-14 · Paul Tschisgale, Holger Maus, Fabian Kieser, Ben Kroehs, Stefan Petersen, Peter Wulff

Large language models (LLMs) are now widely accessible, reaching learners at all educational levels. This development has raised concerns that their use may circumvent essential learning processes and compromise the integrity of established assessment formats. In physics education, where problem solving plays a central role in instruction and assessment, it is therefore essential to understand the physics-specific problem-solving capabilities of LLMs. Such understanding is key to informing responsible and pedagogically sound approaches to integrating LLMs into instruction and assessment. This study therefore compares the problem-solving performance of a general-purpose LLM (GPT-4o, using varying prompting techniques) and a reasoning-optimized model (o1-preview) with that of participants of the German Physics Olympiad, based on a set of well-defined Olympiad problems. In addition to evaluating the correctness of the generated solutions, the study analyzes characteristic strengths and limitations of LLM-generated solutions. The findings of this study indicate that both tested LLMs (GPT-4o and o1-preview) demonstrate advanced problem-solving capabilities on Olympiad-type physics problems, on average outperforming the human participants. Prompting techniques had little effect on GPT-4o's performance, while o1-preview almost consistently outperformed both GPT-4o and the human benchmark. Based on these findings, the study discusses implications for the design of summative and formative assessment in physics education, including how to uphold assessment integrity and support students in critically engaging with LLMs.

📄 PDF Abstract BibTeX arXiv:2505.09438

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Uphold 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

2024-02-21 · Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu 외

Recent advancements have seen Large Language Models (LLMs) and Large Multimodal Models (LMMs) surpassing general human capabilities in various tasks, approaching the proficiency level of human experts across multiple dom…

Logical Fallacies

P1: Mastering Physics Olympiads with Reinforcement Learning

2025-11-17 · Jiacheng Chen, Qianjia Cheng, Fangchen Yu, Haiyuan Wan 외 arxiv

Recent progress in large language models (LLMs) has moved the frontier from puzzle-solving to science-grade reasoning-the kind needed to tackle problems whose answers must stand against nature, not merely fit a rubric. P…

Reinforcement Learning

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model

2026-04-22 · Qiguang Chen, Chengyu Luan, Jiajun Wu, Qiming Yu 외 arxiv

Large vision-language models (LVLMs) have made substantial advances in reasoning tasks at the Olympiad level. Nevertheless, current Olympiad-level multimodal reasoning benchmarks for these models often emphasize single-i…

Multimodal Reasoning

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

2025-04-22 · Shi Qiu, Shaoyang Guo, Zhuo-Yang Song, Yunbo Sun 외

Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed evaluation items. These deficiencies nece…

Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning

2026-05-13 · Shan Yang arxiv

We audit the multimodal-physics evaluation pipeline end-to-end and document three undetected construction practices that distort how the field measures vision-language reasoning: train-eval contamination, translation dri…