paper-with-me

홈 › Papers

NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics

2026-05-08 · Jian Lan, Zhicheng Liu, Xinpeng Wang, Yuhao Zhou, Haokun Chen, Jiancheng Lv, Barbara Plank, Thomas Seidl arxiv

The ability to derive precise spatial and physical insights is a cornerstone of vision-language models (VLMs), yet their poor performances in related spatial intelligence tasks such as physical reasoning remain a fundamental barrier. The community critically lacks a scientific analysis revealing whether VLMs faithfully reach answers or plausibly make guesses. This work aims to provide a fundamental understanding of how VLMs perceive the physical world, and utilize physical laws, while assessing the reliability of model confidence. We propose NICE and FACT, a dual-diagnostic paradigm that explicitly decomposes quantitative reasoning for kinematic physics: FACT diagnoses visual fidelity, physical law comprehension, and temporal grounding. NICE studies our novel neighborhood-informed calibration method and novel metrics to evaluate and calibrate confidence reliability. Evaluated across 6 latest state-of-the-art VLMs, we uncover that models fail to identify visual preconditions or utilize necessary physical laws to reach answers. This work highlights and establishes a standardized diagnostic paradigm to guide the development of faithful, physically-grounded VLMs.

📄 PDF Abstract BibTeX arXiv:2605.08452

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HarmonicEval: Multi-modal, Multi-task, Multi-criteria Automatic Evaluation Using a Vision Language Model

2024-12-19 · Masanari Ohi, Masahiro Kaneko, Naoaki Okazaki, Nakamasa Inoue

Vision-language models (VLMs) have shown impressive abilities in text and image understanding. However, existing metrics for evaluating the text generated by VLMs focus exclusively on overall quality, leading to two limi…

Language ModelingLanguage Modelling

O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language Models

2025-03-15 · CVPR 2025 1 · Ashshak Sharifdeen, Muhammad Akhtar Munir, Sanoojan Baliah, Salman Khan 외

Test-time prompt tuning for vision-language models (VLMs) is getting attention because of their ability to learn with unlabeled data without fine-tuning. Although test-time prompt tuning methods for VLMs can boost accura…

MemLeak: Diagnosing Information Leaks in Multimodal Agent Memory

2026-06-29 · Kuan Wang, Chao Zhang arxiv

When a multimodal AI agent is asked to forget a fact, current memory systems usually delete the text entry and report success. We find that the fact can remain recoverable from retained user images, including images tagg…

EvenNICER-SLAM: Event-based Neural Implicit Encoding SLAM

2024-10-04 · Shi Chen, Danda Pani Paudel, Luc van Gool

The advancement of dense visual simultaneous localization and mapping (SLAM) has been greatly facilitated by the emergence of neural implicit representations. Neural implicit encoding SLAM, a typical example of which is …

Simultaneous Localization and Mapping

SalArt-VQA: Diagnosing Whether VLMs Understand Salient Artifacts in Generated Images

2026-06-10 · Xiaoxiao Sun, Ruotian Zhang, Junzhe Huang, James Burgess 외 arxiv

Vision-language models (VLMs) are increasingly used to detect whether AI-generated images contain visible artifacts, yet their ability to analyze such artifacts remains poorly understood. A correct image-level decision c…

Artifact Detection