paper-with-me

Papers

Multimodal Language Models Cannot Spot Spatial Inconsistencies

2026-04-01 · Om Khangaonkar, Hadi J. Rad, Hamed Pirsiavash arxiv

Spatial consistency is a fundamental property of the visual world and a key requirement for models that aim to understand physical reality. Despite recent advances, multimodal large language models (MLLMs) often struggle to reason about 3D geometry across multiple views. Rather than asking models to describe scene attributes, we introduce a more challenging task: given two views of the same scene, identify the object that violates 3D motion consistency. We propose a simple and scalable method for generating realistic, spatially inconsistent image pairs from multi-view scenes, enabling systematic evaluation of this capability. Our results show that state-of-the-art MLLMs significantly underperform human observers and exhibit substantial variability across different scene attributes, revealing a fragile and incomplete understanding of 3D structure. We hope our findings underscore the need for approaches that develop a more deeply grounded understanding of the physical world.

📄 PDF Abstract BibTeX arXiv:2604.00799

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SPOT: Bridging Natural Language and Geospatial Search for Investigative Journalists

2025-06-16 · Lynn Khellaf, Ipek Baris Schlicht, Tilman Mirass, Julia Bayer 외

OpenStreetMap (OSM) is a vital resource for investigative journalists doing geolocation verification. However, existing tools to query OSM data such as Overpass Turbo require familiarity with complex query languages, cre…

Fact CheckingTAG

LATTICE: Graph Self-Supervised Learning for Multimodal Spatial Omics Integration

2026-07-15 · Jagan Mohan Reddy Dwarampudi, Veena Kochat, Suresh Satpati, Kunal Rai 외 arxiv

Spatially resolved omics studies increasingly combine transcriptomic and epigenomic assays, yet downstream analysis is often still performed using single-modality pipelines. We present LATTICE (Latent Alignment of Tissue…

Self-Supervised Learning

ST-Align: A Multimodal Foundation Model for Image-Gene Alignment in Spatial Transcriptomics

2024-11-25 · Yuxiang Lin, Ling Luo, Ying Chen, Xushi Zhang 외

Spatial transcriptomics (ST) provides high-resolution pathological images and whole-transcriptomic expression profiles at individual spots across whole-slide scales. This setting makes it an ideal data source to develop …

MedSPOT: A Workflow-Aware Sequential Grounding Benchmark for Clinical GUI

2026-03-20 · Rozain Shakeel, Abdul Rahman Mohammad Ali, Muneeb Mushtaq, Tausifa Jan Saleem 외 arxiv

Despite the rapid progress of Multimodal Large Language Models (MLLMs), their ability to perform reliable visual grounding in high-stakes clinical software environments remains underexplored. Existing GUI benchmarks larg…

Visual Grounding

Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models

2025-02-22 · Qianqi Yan, Yue Fan, Hongquan Li, Shan Jiang 외

Existing Multimodal Large Language Models (MLLMs) are predominantly trained and tested on consistent visual-textual inputs, leaving open the question of whether they can handle inconsistencies in real-world, layout-rich …

Multimodal Reasoning