paper-with-me

홈 › Papers

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors?

2026-06-06 · Bin Zhu, Yanhao Jia, Kexin Zhao, Jie Wang, Munan Ning, Hao Li, Yuwei Niu, Tanqing Sun, Huangchong Yan, Mingjun Pan, Xinyi Wu, Qishen Yin, Yunyang Ge, Shuai Zhao, Li Yuan arxiv

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in open-world reasoning and understanding. However, a critical ambiguity persists: it remains unclear whether these models genuinely synthesize cross-modal information to construct physically grounded reasoning chains, or if they merely exploit strong language priors to mask single-modality reliance, thereby hallucinating advanced multimodal capabilities. Motivated by this, and to rigorously mitigate language modality bias and shortcuts, we propose a novel multimodal Chrono}logical Physical Dynamics Reasoning Benchmark ChronoPhyBench, which unifies next state prediction with Visual Question Answering (VQA) paradigms by conditioning on historical video context and textual captions to enforce models to deduce subsequent physical states through both single image selection and the inherently more complex task of multiple frame chronological sorting. Concurrently, we construct a large-scale multimodal reasoning dataset curated using the ChronoPhyBench criteria, comprising over 10,000 long-form videos paired with meticulously annotated captions, totaling 5M tokens. Our experimental evaluations reveal a stark contrast to conclusions drawn by previous benchmarks. The capacity of current open-source models to perform physically grounded multimodal reasoning remains in its infancy. Ultimately, this work seeks to systematically stress-test the reasoning capabilities of multimodal models, quantify hallucination rates, and advance the development of Physical AI, thereby providing the community with a robust and transparent evaluation framework toward Artificial General Intelligence (AGI).

📄 PDF Abstract BibTeX arXiv:2606.07962

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringMultimodal Reasoning

Similar Papers 제목 키워드 기반

GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation

2025-12-19 · Rang Li, Lei Li, Shuhuai Ren, Hao Tian 외 arxiv

Visual grounding, localizing objects from natural language descriptions, represents a critical bridge between language and vision understanding. While multimodal large language models (MLLMs) achieve impressive scores on…

Visual Grounding

Understanding the Role of LLMs in Multimodal Evaluation Benchmarks

2024-10-16 · Botian Jiang, Lei LI, Xiaonan Li, Zhaowei Li 외

The rapid advancement of Multimodal Large Language Models (MLLMs) has been accompanied by the development of various benchmarks to evaluate their capabilities. However, the true nature of these evaluations and the extent…

BenchmarkingLarge Language ModelMultimodal ReasoningWorld Knowledge

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

2025-11-25 · Tianxiang Jiang, Sheng Xia, Yicheng Xu, Linquan Wu 외 arxiv

While Multimodal Large Language Models (MLLMs) have become adept at recognizing objects, they often lack the intuitive, human-like understanding of the world's underlying physical and social principles. This high-level v…

Reinforcement Learning

Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?

2026-05-21 · Caixin Kang, Tianyu Yan, Sitong Gong, Mingfang Zhang 외 arxiv

Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchmarks evaluate this capability solely on numerical Big Five score predi…

EgoSound: Benchmarking Sound Understanding in Egocentric Videos

2026-02-15 · Bingwen Zhu, Yuqian Fu, Qiaole Dong, Guolei Sun 외 arxiv

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about…

Causal Inference