paper-with-me

홈 › Papers

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

2025-04-22 · Shi Qiu, Shaoyang Guo, Zhuo-Yang Song, Yunbo Sun, Zeyu Cai, Jiashen Wei, Tianyu Luo, Yixuan Yin, Haoxu Zhang, Yi Hu, Chenyang Wang, Chencheng Tang, Haoling Chang, Qi Liu, Ziheng Zhou, Tianyu Zhang, Jingtian Zhang, Zhangyi Liu, Minghao Li, Yuku Zhang, Boxuan Jing, Xianqi Yin, Yutong Ren, Zizhuo Fu, Jiaming Ji, Weike Wang, Xudong Tian, Anqi Lv, Laifu Man, Jianxiang Li, Feiyu Tao, Qihua Sun, Zhou Liang, Yushu Mu, Zhongxuan Li, Jing-Jun Zhang, Shutao Zhang, Xiaotian Li, Xingqi Xia, Jiawei Lin, Zheyu Shen, Jiahang Chen, Qiuhao Xiong, Binran Wang, Fengyuan Wang, Ziyang Ni, Bohan Zhang, Fan Cui, Changkun Shao, Qing-Hong Cao, Ming-Xing Luo, Yaodong Yang, Muhan Zhang, Hua Xing Zhu

Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed evaluation items. These deficiencies necessitate more rigorous assessment methods. To address these limitations, we introduce PHYBench, a benchmark of 500 original physics problems ranging from high school to Physics Olympiad difficulty. PHYBench addresses data contamination through original content and employs a systematic curation pipeline to eliminate flawed items. Evaluations show that PHYBench activates more tokens and provides stronger differentiation between reasoning models compared to other baselines like AIME 2024, OlympiadBench and GPQA. Even the best-performing model, Gemini 2.5 Pro, achieves only 36.9% accuracy compared to human experts' 61.9%. To further enhance evaluation precision, we introduce the Expression Edit Distance (EED) Score for mathematical expression assessment, which improves sample efficiency by 204% over binary scoring. Moreover, PHYBench effectively elicits multi-step and multi-condition reasoning, providing a platform for examining models' reasoning robustness, preferences, and deficiencies. The benchmark results and dataset are publicly available at https://www.phybench.cn/.

📄 PDF Abstract BibTeX arXiv:2504.16074

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors?

2026-06-06 · Bin Zhu, Yanhao Jia, Kexin Zhao, Jie Wang 외 arxiv

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in open-world reasoning and understanding. However, a critical ambiguity persists: it remains unclear whether these…

Visual Question AnsweringMultimodal Reasoning

Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Models

2025-10-30 · Shiho Matta, Lis Kanashiro Pereira, Peitao Han, Fei Cheng 외 arxiv

Modern vision-language models (VLMs) excel at many multimodal tasks, yet their grasp of temporal information in video remains weak and has not been adequately evaluated. We probe this gap with a deceptively simple but re…

PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models

2024-06-17 · Fanqing Meng, Wenqi Shao, Lixin Luo, Yahong Wang 외

Text-to-image (T2I) models have made substantial progress in generating images from textual prompts. However, they frequently fail to produce images consistent with physical commonsense, a vital capability for applicatio…

Image Generation

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction

2026-02-09 · Jiarong Liang, Max Ku, Ka-Hei Hui, Ping Nie 외 arxiv

Evaluating whether Multimodal Large Language Models (MLLMs) genuinely reason about physical dynamics remains challenging. Most existing benchmarks rely on recognition-style protocols such as Visual Question Answering (VQ…

Visual Question AnsweringVideo ReconstructionScene Understanding

STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence

2025-10-28 · Zihan Liu, Zhikang Niu, Qiuyang Xiao, Zhisheng Zheng 외 arxiv

Despite rapid progress in Multi-modal Large Language Models and Large Audio-Language Models, existing audio benchmarks largely test semantics that can be recovered from text captions, masking deficits in fine-grained per…