paper-with-me

홈 › Papers

SIFThinker: Spatially-Aware Image Focus for Visual Reasoning

2025-08-08 · Zhangquan Chen, Ruihui Zhao, Chuwei Luo, Mingze Sun, Xinlei Yu, Yangyang Kang, Ruqi Huang arxiv

Current multimodal large language models (MLLMs) still face significant challenges in complex visual tasks (e.g., spatial understanding, fine-grained perception). Prior methods have tried to incorporate visual reasoning, however, they fail to leverage attention correction with spatial cues to iteratively refine their focus on prompt-relevant regions. In this paper, we introduce SIFThinker, a spatially-aware "think-with-images" framework that mimics human visual perception. Specifically, SIFThinker enables attention correcting and image region focusing by interleaving depth-enhanced bounding boxes and natural language. Our contributions are twofold: First, we introduce a reverse-expansion-forward-inference strategy that facilitates the generation of interleaved image-text chains of thought for process-level supervision, which in turn leads to the construction of the SIF-50K dataset. Besides, we propose GRPO-SIF, a reinforced training paradigm that integrates depth-informed visual grounding into a unified reasoning pipeline, teaching the model to dynamically correct and focus on prompt-relevant regions. Extensive experiments demonstrate that SIFThinker outperforms state-of-the-art methods in spatial understanding and fine-grained visual perception, while maintaining strong general capabilities, highlighting the effectiveness of our method. Code: https://github.com/zhangquanchen/SIFThinker.

📄 PDF Abstract BibTeX arXiv:2508.06259

Code (0)

등록된 구현이 없습니다.

Tasks

Visual GroundingVisual Reasoning

Similar Papers 제목 키워드 기반

Spatially Aware Multimodal Transformers for TextVQA

2020-07-23 · ECCV 2020 8 · Yash Kant, Dhruv Batra, Peter Anderson, Alex Schwing 외

Textual cues are essential for everyday tasks like buying groceries and using public transport. To develop this assistive technology, we study the TextVQA task, i.e., reasoning about text in images to answer a question. …

Optical Character Recognition (OCR)Spatial ReasoningTextVQAVisual Grounding+1

Automatic Spatially-aware Fashion Concept Discovery

2017-08-03 · ICCV 2017 10 · Xintong Han, Zuxuan Wu, Phoenix X. Huang, Xiao Zhang 외

This paper proposes an automatic spatially-aware concept discovery approach using weakly labeled image-text data from shopping websites. We first fine-tune GoogleNet by jointly modeling clothing images and their correspo…

AttributeClusteringImage Retrieval with Multi-Modal QueryRetrieval

Self-Supervised Spatially Variant PSF Estimation for Aberration-Aware Depth-from-Defocus

2024-02-28 · Zhuofeng Wu, Yusuke Monno, Masatoshi Okutomi

In this paper, we address the task of aberration-aware depth-from-defocus (DfD), which takes account of spatially variant point spread functions (PSFs) of a real camera. To effectively obtain the spatially variant PSFs o…

Depth EstimationSelf-Supervised Learning

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models

2026-05-13 · Junha Song, Byeongho Heo, Geonmo Gu, Jaegul Choo 외 arxiv

When humans describe a visual scene, they do not process the entire image uniformly; instead, they selectively fixate on regions relevant to their intended description. In contrast, current multimodal large language mode…

Show and Tell: Visually Explainable Deep Neural Nets via Spatially-Aware Concept Bottleneck Models

2025-02-27 · CVPR 2025 1 · Itay Benou, Tammy Riklin-Raviv

Modern deep neural networks have now reached human-level performance across a variety of tasks. However, unlike humans they lack the ability to explain their decisions by showing where and telling what concepts guided th…

Zero Shot Segmentation