paper-with-me

홈 › Papers

Understanding the Implicit User Intention via Reasoning with Large Language Model for Image Editing

2025-10-31 · Yijia Wang, Yiqing Shen, Weiming Chen, Zhihai He arxiv

Existing image editing methods can handle simple editing instructions very well. To deal with complex editing instructions, they often need to jointly fine-tune the large language models (LLMs) and diffusion models (DMs), which involves very high computational complexity and training cost. To address this issue, we propose a new method, called \textbf{C}omplex \textbf{I}mage \textbf{E}diting via \textbf{L}LM \textbf{R}easoning (CIELR), which converts a complex user instruction into a set of simple and explicit editing actions, eliminating the need for jointly fine-tuning the large language models and diffusion models. Specifically, we first construct a structured semantic representation of the input image using foundation models. Then, we introduce an iterative update mechanism that can progressively refine this representation, obtaining a fine-grained visual representation of the image scene. This allows us to perform complex and flexible image editing tasks. Extensive experiments on the SmartEdit Reasoning Scenario Set show that our method surpasses the previous state-of-the-art by 9.955 dB in PSNR, indicating its superior preservation of regions that should remain consistent. Due to the limited number of samples of public datasets of complex image editing with reasoning, we construct a benchmark named CIEBench, containing 86 image samples, together with a metric specifically for reasoning-based image editing. CIELR also outperforms previous methods on this benchmark. The code and dataset are available at \href{https://github.com/Jia-shao/Reasoning-Editing}{https://github.com/Jia-shao/Reasoning-Editing}.

📄 PDF Abstract BibTeX arXiv:2510.27335

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

A Data Synthesis Method Driven by Large Language Models for Proactive Mining of Implicit User Intentions in Tourism

2025-05-14 · Jinqiang Wang, Huansheng Ning, Tao Zhu, Jianguo Ding

In the tourism domain, Large Language Models (LLMs) often struggle to mine implicit user intentions from tourists' ambiguous inquiries and lack the capacity to proactively guide users toward clarifying their needs. A cri…

Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents

2024-02-14 · Cheng Qian, Bingxiang He, Zhong Zhuang, Jia Deng 외

Current language model-driven agents often lack mechanisms for effective user participation, which is crucial given the vagueness commonly found in user instructions. Although adept at devising strategies and performing …

Language ModelingLanguage Modelling

Uncertain Multimodal Intention and Emotion Understanding in the Wild

2025-01-01 · CVPR 2025 1 · Qu Yang, Qinghongya Shi, Tongxin Wang, Mang Ye

Understanding intention and emotion from social media poses unique challenges due to the inherent uncertainty in multimodal data, where posts often contain incomplete or missing modalities. While this uncertainty ref…

A scalable framework for learning from implicit user feedback to improve natural language understanding in large-scale conversational AI systems

2020-10-23 · EMNLP 2021 11 · Sunghyun Park, Han Li, Ameen Patel, Sidharth Mudgal 외

Natural Language Understanding (NLU) is an established component within a conversational AI or digital assistant system, and it is responsible for producing semantic understanding of a user request. We propose a scalable…

Natural Language Understanding

Comparing Apples to Oranges: LLM-powered Multimodal Intention Prediction in an Object Categorization Task

2024-04-12 · Hassan Ali, Philipp Allgeuer, Stefan Wermter

Human intention-based systems enable robots to perceive and interpret user actions to interact with humans and adapt to their behavior proactively. Therefore, intention prediction is pivotal in creating a natural interac…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Object Categorizationspeech-recognition+2