paper-with-me

Papers

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models

2026-05-11 · Tingshu Mou, Jiabo He, Renying Wang, Ce Liu, Hao Yang, Tiehua Zhang, Jingjing Chen, Xingjun Ma arxiv

Recent advances in Multi-modal Large Language Models (MLLMs) target 3D spatial intelligence, yet the progress has been largely driven by post-training on curated benchmarks, leaving the inference-time approach relatively underexplored. In this paper, we take a training-free perspective and introduce ViSRA, a human-aligned Video-based Spatial Reasoning Agent, as a framework to probe the spatial reasoning mechanism of MLLMs. ViSRA elicits spatial reasoning in a modular and extensible manner by leveraging explicit spatial information from expert models, enabling a plug-and-play flexible paradigm. ViSRA offers two key advantages: (1) human-aligned and transferable 3D understanding rather than task-specific overfitting; and (2) no post-training computational cost along with heavy manual curation of spatial reasoning datasets. Experimental results demonstrate consistent improvement across a set of MLLMs on both existing benchmarks and unseen 3D spatial reasoning tasks, with ViSRA outperforming baselines by up to a 15.6% and 28.9% absolute margin respectively.

📄 PDF Abstract BibTeX arXiv:2605.10106

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation

2025-10-10 · Yubo Sun, Chunyi Peng, Yukun Yan, Shi Yu 외 arxiv

Visual Retrieval-Augmented Generation (VRAG) has emerged as a promising paradigm for equipping Vision-Language Models (VLMs) with external visual evidence, enabling them to go beyond parametric knowledge when answering v…

Visual Question Answering

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

2024-10-14 · Shi Yu, Chaoyue Tang, Bokai Xu, Junbo Cui 외

Retrieval-augmented generation (RAG) is an effective technique that enables large language models (LLMs) to utilize external knowledge sources for generation. However, current RAG systems are solely based on text, render…

RAGRetrievalRetrieval-augmented Generation

RobustVisRAG: Causality-Aware Vision-Based Retrieval-Augmented Generation under Visual Degradations

2026-02-25 · I-Hsiang Chen, Yu-Wei Liu, Tse-Yu Wu, Yu-Chien Chiang 외 arxiv

Vision-based Retrieval-Augmented Generation (VisRAG) leverages vision-language models (VLMs) to jointly retrieve relevant visual documents and generate grounded answers based on multimodal evidence. However, existing Vis…

Zero-shot Generalization

AiSciVision: A Framework for Specializing Large Multimodal Models in Scientific Image Classification

2024-10-28 · Brendan Hogan, Anmol Kabra, Felipe Siqueira Pacheco, Laura Greenstreet 외

Trust and interpretability are crucial for the use of Artificial Intelligence (AI) in scientific research, but current models often operate as black boxes offering limited transparency and justifications for their output…

image-classificationImage ClassificationRetrieval-augmented Generationscientific discovery

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence

2026-06-18 · Yalun Dai, Hao Li, Shulin Tian, Runmao Yao 외 arxiv

Real-world spatial intelligence requires reasoning over a continuous and evolving 3D world, yet existing VLMs and tool-augmented agents largely remain tied to static, stateless inference from isolated visual observations…

Spatial Reasoning