paper-with-me

홈 › Papers

PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models

2024-06-17 · Fanqing Meng, Wenqi Shao, Lixin Luo, Yahong Wang, Yiran Chen, Quanfeng Lu, Yue Yang, Tianshuo Yang, Kaipeng Zhang, Yu Qiao, Ping Luo

Text-to-image (T2I) models have made substantial progress in generating images from textual prompts. However, they frequently fail to produce images consistent with physical commonsense, a vital capability for applications in world simulation and everyday tasks. Current T2I evaluation benchmarks focus on metrics such as accuracy, bias, and safety, neglecting the evaluation of models' internal knowledge, particularly physical commonsense. To address this issue, we introduce PhyBench, a comprehensive T2I evaluation dataset comprising 700 prompts across 4 primary categories: mechanics, optics, thermodynamics, and material properties, encompassing 31 distinct physical scenarios. We assess 6 prominent T2I models, including proprietary models DALLE3 and Gemini, and demonstrate that incorporating physical principles into prompts enhances the models' ability to generate physically accurate images. Our findings reveal that: (1) even advanced models frequently err in various physical scenarios, except for optics; (2) GPT-4o, with item-specific scoring instructions, effectively evaluates the models' understanding of physical commonsense, closely aligning with human assessments; and (3) current T2I models are primarily focused on text-to-image translation, lacking profound reasoning regarding physical commonsense. We advocate for increased attention to the inherent knowledge within T2I models, beyond their utility as mere image generation tools. The data will be available soon.

📄 PDF Abstract BibTeX arXiv:2406.11802

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

2025-04-22 · Shi Qiu, Shaoyang Guo, Zhuo-Yang Song, Yunbo Sun 외

Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed evaluation items. These deficiencies nece…

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors?

2026-06-06 · Bin Zhu, Yanhao Jia, Kexin Zhao, Jie Wang 외 arxiv

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in open-world reasoning and understanding. However, a critical ambiguity persists: it remains unclear whether these…

Visual Question AnsweringMultimodal Reasoning

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction

2026-02-09 · Jiarong Liang, Max Ku, Ka-Hei Hui, Ping Nie 외 arxiv

Evaluating whether Multimodal Large Language Models (MLLMs) genuinely reason about physical dynamics remains challenging. Most existing benchmarks rely on recognition-style protocols such as Visual Question Answering (VQ…

Visual Question AnsweringVideo ReconstructionScene Understanding

Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Models

2025-10-30 · Shiho Matta, Lis Kanashiro Pereira, Peitao Han, Fei Cheng 외 arxiv

Modern vision-language models (VLMs) excel at many multimodal tasks, yet their grasp of temporal information in video remains weak and has not been adequately evaluated. We probe this gap with a deceptively simple but re…

PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos

2024-12-02 · Meng Cao, Haoran Tang, Haoze Zhao, Hangyu Guo 외

Recent advancements in video-based large language models (Video LLMs) have witnessed the emergence of diverse capabilities to reason and interpret dynamic visual content. Among them, gameplay videos stand out as a distin…

Question AnsweringVideo Understanding