paper-with-me

홈 › Papers

ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining

2025-12-22 · Zhenyang Huang, Xiao Yu, Yi Zhang, Decheng Wang, Hang Ruan arxiv

Remote sensing image change detection is one of the fundamental tasks in remote sensing intelligent interpretation. Its core objective is to identify changes within change regions of interest (CRoI). Current multimodal large models encode rich human semantic knowledge, which is utilized for guidance in tasks such as remote sensing change detection. However, existing methods that use semantic guidance for detecting users' CRoI overly rely on explicit textual descriptions of CRoI, leading to the problem of near-complete performance failure when presented with implicit CRoI textual descriptions. This paper proposes a multimodal reasoning change detection model named ReasonCD, capable of mining users' implicit task intent. The model leverages the powerful reasoning capabilities of pre-trained large language models to mine users' implicit task intents and subsequently obtains different change detection results based on these intents. Experiments on public datasets demonstrate that the model achieves excellent change detection performance, with an F1 score of 92.1\% on the BCDD dataset. Furthermore, to validate its superior reasoning functionality, this paper annotates a subset of reasoning data based on the SECOND dataset. Experimental results show that the model not only excels at basic reasoning-based change detection tasks but can also explain the reasoning process to aid human decision-making.

📄 PDF Abstract BibTeX arXiv:2512.19354

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningChange Detection

Similar Papers 제목 키워드 기반

Safe Semantics, Unsafe Interpretations: Tackling Implicit Reasoning Safety in Large Vision-Language Models

2025-08-12 · Wei Cai, Jian Zhao, Yuchu Jiang, Tianle Zhang 외 arxiv

Large Vision-Language Models face growing safety challenges with multimodal inputs. This paper introduces the concept of Implicit Reasoning Safety, a vulnerability in LVLMs. Benign combined inputs trigger unsafe LVLM out…

CoVR-R:Reason-Aware Composed Video Retrieval

2026-03-20 · Omkar Thawakar, Dmitry Demidov, Vaishnav Potlapalli, Sai Prasanna Teja Reddy Bogireddy 외 arxiv

Composed Video Retrieval (CoVR) aims to find a target video given a reference video and a textual modification. Prior work assumes the modification text fully specifies the visual changes, overlooking after-effects and i…

Video Retrieval

DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models

2026-07-09 · Pengjie Wang, Linger Deng, Zujia Zhang, Shaojie Zhang 외 arxiv

Current Unified Large Multimodal Models (ULMMs) support interleaved multimodal reasoning through textual reasoning and intermediate visual states, but typically generate each visual state as a full image. This full-image…

Multimodal ReasoningImage Generation

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models

2026-02-09 · Masanari Oi, Koki Maeda, Ryuto Koike, Daisuke Oba 외 arxiv

While multimodal large language models (MLLMs) have made substantial progress in single-image spatial reasoning, multi-image spatial reasoning, which requires integration of information from multiple viewpoints, remains …

Spatial Reasoning

Towards Explainable Harmful Meme Detection through Multimodal Debate between Large Language Models

2024-01-24 · Hongzhan Lin, Ziyang Luo, Wei Gao, Jing Ma 외

The age of social media is flooded with Internet memes, necessitating a clear grasp and effective identification of harmful ones. This task presents a significant challenge due to the implicit meaning embedded in memes, …

Hateful Meme ClassificationLanguage ModellingSmall Language ModelText Generation