paper-with-me

Papers

Interaction-Consistent Object Removal via MLLM-Based Reasoning

2026-02-01 · Ching-Kai Huang, Wen-Chieh Lin, Yan-Cen Lee arxiv

Image-based object removal often erases only the named target, leaving behind interaction evidence that renders the result semantically inconsistent. We formalize this problem as Interaction-Consistent Object Removal (ICOR), which requires removing not only the target object but also associated interaction elements, such as lighting-dependent effects, physically connected objects, targetproduced elements, and contextually linked objects. To address this task, we propose Reasoning-Enhanced Object Removal with MLLM (REORM), a reasoningenhanced object removal framework that leverages multimodal large language models to infer which elements must be jointly removed. REORM features a modular design that integrates MLLM-driven analysis, mask-guided removal, and a self-correction mechanism, along with a local-deployment variant that supports accurate editing under limited resources. To support evaluation, we introduce ICOREval, a benchmark consisting of instruction-driven removals with rich interaction dependencies. On ICOREval, REORM outperforms state-of-the-art image editing systems, demonstrating its effectiveness in producing interactionconsistent results.

📄 PDF Abstract BibTeX arXiv:2602.01298

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting

2025-10-18 · Changyue Shi, Minghao Chen, Yiping Mao, Chuxiao Yang 외 arxiv

Bridging the gap between complex human instructions and precise 3D object grounding remains a significant challenge in vision and robotics. Existing 3D segmentation methods often struggle to interpret ambiguous, reasonin…

Style Transfer

VOID: Video Object and Interaction Deletion

2026-04-02 · Saman Motamed, William Harvey, Benjamin Klein, Luc Van Gool 외 arxiv

Existing video object removal methods excel at inpainting content "behind" the object and correcting appearance-level artifacts such as shadows and reflections. However, when the removed object has more significant inter…

ObjChangeVR: Object State Change Reasoning from Continuous Egocentric Views in VR Environments

2026-02-27 · Shiyi Ding, Shaoen Wu, Ying Chen arxiv

Recent advances in multimodal large language models (MLLMs) offer a promising approach for natural language-based scene change queries in virtual reality (VR). Prior work on applying MLLMs for object state understanding …

HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection

2025-10-07 · Junwen Chen, Peilin Xiong, Keiji Yanai arxiv

Recent human-object interaction detection (HOID) methods highly require prior knowledge from vision-language models (VLMs) to enhance the interaction recognition capabilities. The training strategies and model architectu…

Human-Object Interaction DetectionReinforcement Learning

Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion

2026-03-06 · Bohai Gu, Taiyi Wu, Dazhao Du, Jian Liu 외 arxiv

Modern video editing techniques have achieved high visual fidelity when inserting video objects. However, they focus on optimizing visual fidelity rather than physical causality, leading to edits that are physically inco…

Scene Understanding