paper-with-me

Papers

Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models

2023-08-18 · Navid Rajabi, Jana Kosecka

Large vision-and-language models (VLMs) trained to match images with text on large-scale datasets of image-text pairs have shown impressive generalization ability on several vision and language tasks. Several recent works, however, showed that these models lack fine-grained understanding, such as the ability to count and recognize verbs, attributes, or relationships. The focus of this work is to study the understanding of spatial relations. This has been tackled previously using image-text matching (e.g., Visual Spatial Reasoning benchmark) or visual question answering (e.g., GQA or VQAv2), both showing poor performance and a large gap compared to human performance. In this work, we show qualitatively (using explainability tools) and quantitatively (using object detectors) that the poor object localization "grounding" ability of the models is a contributing factor to the poor image-text matching performance. We propose an alternative fine-grained, compositional approach for recognizing and ranking spatial clauses that combines the evidence from grounding noun phrases corresponding to objects and their locations to compute the final rank of the spatial clause. We demonstrate the approach on representative VLMs (such as LXMERT, GPV, and MDETR) and compare and highlight their abilities to reason about spatial relationships.

📄 PDF Abstract BibTeX arXiv:2308.09778

Code (0)

등록된 구현이 없습니다.

Tasks

Image-text matchingObject LocalizationQuestion AnsweringSpatial ReasoningText MatchingVisual Question AnsweringVisual Reasoning

Methods 이 논문이 사용한 방법론

LXMERT LXMERT is a model for learning vision-and-language cross-modality representations. It consists of a Transformer model that consists three encoders: object relationship encoder, a…
Focus 설명 없음

Similar Papers 제목 키워드 기반

TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation

2026-03-19 · Yan Shu, Bin Ren, Zhitong Xiong, Xiao Xiang Zhu 외 arxiv

Vision-language models (VLMs) have shown promise in earth observation (EO), yet they struggle with tasks that require grounding complex spatial reasoning in precise pixel-level visual representations. To address this pro…

Temporal SequencesSpatial ReasoningVisual Reasoning

Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning

2024-12-16 · Qi Sun, Pengfei Hong, Tej Deep Pala, Vernon Toh 외

Traditional reinforcement learning-based robotic control methods are often task-specific and fail to generalize across diverse environments or unseen objects and instructions. Visual Language Models (VLMs) demonstrate st…

HallucinationRobot ManipulationScene UnderstandingSpatial Reasoning+1

CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning

2026-08-03 · Xuehang Guo, Pingyue Zhang, Ruiyi Zhang, Zhenhailong Wang 외 hf

Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning…

Chart Question AnsweringMultimodal ReasoningLogical ReasoningVisual Reasoning

SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving

2026-01-24 · Ashutosh Bajpai, Akshat Bhandari, Akshay Nambi, Tanmoy Chakraborty arxiv

Multimodal Small-to-Medium sized Language Models (MSLMs) have demonstrated strong capabilities in integrating visual and textual information but still face significant limitations in visual comprehension and mathematical…

Mathematical ReasoningData Augmentation

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

2026-05-05 · Lin Song, Wenbo Li, Guoqing Ma, Wei Tang 외 arxiv

We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text-to-image generation, and instruction-guided image editing. JoyAI-Image couples a spatially enhanced Multimodal Large Language M…

Text-to-Image GenerationImage Editing