paper-with-me

홈 › Papers

PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding

2025-01-27 · Wei Chow, Jiageng Mao, Boyi Li, Daniel Seita, Vitor Guizilini, Yue Wang

Understanding the physical world is a fundamental challenge in embodied AI, critical for enabling agents to perform complex tasks and operate safely in real-world environments. While Vision-Language Models (VLMs) have shown great promise in reasoning and task planning for embodied agents, their ability to comprehend physical phenomena remains extremely limited. To close this gap, we introduce PhysBench, a comprehensive benchmark designed to evaluate VLMs' physical world understanding capability across a diverse set of tasks. PhysBench contains 10,002 entries of interleaved video-image-text data, categorized into four major domains: physical object properties, physical object relationships, physical scene understanding, and physics-based dynamics, further divided into 19 subclasses and 8 distinct capability dimensions. Our extensive experiments, conducted on 75 representative VLMs, reveal that while these models excel in common-sense reasoning, they struggle with understanding the physical world -- likely due to the absence of physical knowledge in their training data and the lack of embedded physical priors. To tackle the shortfall, we introduce PhysAgent, a novel framework that combines the generalization strengths of VLMs with the specialized expertise of vision models, significantly enhancing VLMs' physical understanding across a variety of tasks, including an 18.4\% improvement on GPT-4o. Furthermore, our results demonstrate that enhancing VLMs' physical world understanding capabilities can help embodied agents such as MOKA. We believe that PhysBench and PhysAgent offer valuable insights and contribute to bridging the gap between VLMs and physical world understanding.

📄 PDF Abstract BibTeX arXiv:2501.16411

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingCommon Sense ReasoningScene UnderstandingTask Planning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

PhysBrain 1.0 Technical Report

2026-05-14 · Shijie Lian, Bin Yu, Xiaopeng Lin, Changti Wu 외 arxiv

Vision-language-action models have advanced rapidly, but robot trajectories alone provide limited coverage for learning broad physical understanding. PhysBrain 1.0 studies a complementary route: converting large-scale hu…

CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics

2025-08-25 · Weida Wang, Dongchen Huang, Jiatong Li, Tengchao Yang 외 arxiv

We introduce CMPhysBench, designed to assess the proficiency of Large Language Models (LLMs) in Condensed Matter Physics, as a novel Benchmark. CMPhysBench is composed of more than 520 graduate-level meticulously curated…

PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model

2026-04-27 · Sinin Zhang, Yunfei Xie, Yuxuan Cheng, Haoyu Zhang 외 arxiv

Vision-Language Models (VLMs) have demonstrated strong performance on textbook-style physics problems, yet they frequently fail when confronted with dynamic real-world scenarios that require temporal consistency and caus…

T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation

2025-05-01 · Xuyang Guo, Jiayan Huo, Zhenmei Shi, Zhao Song 외

Text-to-video generative models have made significant strides in recent years, producing high-quality videos that excel in both aesthetic appeal and accurate instruction following, and have become central to digital art …

counterfactualInstruction FollowingText-to-Video GenerationVideo Generation

Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos

2026-01-23 · Meng Cao, Haoran Tang, Haoze Zhao, Mingfei Han 외 arxiv

Understanding the physical world, including object dynamics, material properties, and causal interactions, remains a core challenge in artificial intelligence. Although recent multi-modal large language models (MLLMs) ha…