paper-with-me

홈 › Papers

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models

2026-02-09 · Masanari Oi, Koki Maeda, Ryuto Koike, Daisuke Oba, Nakamasa Inoue, Naoaki Okazaki arxiv

While multimodal large language models (MLLMs) have made substantial progress in single-image spatial reasoning, multi-image spatial reasoning, which requires integration of information from multiple viewpoints, remains challenging. Cognitive studies suggest that humans address such tasks through two mechanisms: cross-view correspondence, which identifies regions across different views that correspond to the same physical locations, and stepwise viewpoint transformation, which composes relative viewpoint changes sequentially. However, existing studies incorporate these mechanisms only partially and often implicitly, without explicit supervision for both. We propose Human-Aware Training for Cross-view correspondence and viewpoint cHange (HATCH), a training framework with two complementary objectives: (1) Patch-Level Spatial Alignment, which encourages patch representations to align across views for spatially corresponding regions, and (2) Action-then-Answer Reasoning, which requires the model to generate explicit viewpoint transition actions before predicting the final answer. Experiments on three benchmarks demonstrate that HATCH consistently outperforms baselines of comparable size by a clear margin and achieves competitive results against much larger models, while preserving single-image reasoning capabilities.

📄 PDF Abstract BibTeX arXiv:2602.08735

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Learning to Animate Images from A Few Videos to Portray Delicate Human Actions

2025-03-01 · Haoxin Li, Yingchen Yu, Qilong Wu, Hanwang Zhang 외

Despite recent progress, video generative models still struggle to animate human actions from static images, particularly when handling uncommon actions whose training data are limited. In this paper, we investigate the …

DecoderFew-Shot Learning

MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion

2023-07-03 · NeurIPS 2023 11 · Shitao Tang, Fuyang Zhang, Jiacheng Chen, Peng Wang 외

This paper introduces MVDiffusion, a simple yet effective method for generating consistent multi-view images from text prompts given pixel-to-pixel correspondences (e.g., perspective crops from a panorama or multi-view i…

Image Generation

Virtual Correspondence: Humans as a Cue for Extreme-View Geometry

2022-06-16 · CVPR 2022 1 · Wei-Chiu Ma, Anqi Joyce Yang, Shenlong Wang, Raquel Urtasun 외

Recovering the spatial layout of the cameras and the geometry of the scene from extreme-view images is a longstanding challenge in computer vision. Prevailing 3D reconstruction algorithms often adopt the image matching p…

3D ReconstructionCamera Pose EstimationNovel View SynthesisPose Estimation

CrossVideoMAE: Contrastive Spatiotemporal and Semantic Representation Learning from Videos and Images with Masked Autoencoders

2025-02-08 · Shihab Aaqil Ahamed, Malitha Gunawardhana, Liel David, Michael Sidorov 외

Current video-based Masked Autoencoders (MAEs) primarily learn general spatial-temporal patterns from a visual perspective but often overlook nuanced semantic attributes like specific interactions or sequences that defin…

Representation Learning

TransforMatcher: Match-to-Match Attention for Semantic Correspondence

2022-05-23 · CVPR 2022 1 · SeungWook Kim, Juhong Min, Minsu Cho

Establishing correspondences between images remains a challenging task, especially under large appearance changes due to different viewpoints or intra-class variations. In this work, we introduce a strong semantic image …

Semantic correspondence