paper-with-me

Papers

MoEdit: On Learning Quantity Perception for Multi-object Image Editing

2025-03-13 · CVPR 2025 1 · Yanfeng Li, Kahou Chan, Yue Sun, ChanTong Lam, Tong Tong, Zitong Yu, Keren Fu, Xiaohong Liu, Tao Tan

Multi-object images are prevalent in various real-world scenarios, including augmented reality, advertisement design, and medical imaging. Efficient and precise editing of these images is critical for these applications. With the advent of Stable Diffusion (SD), high-quality image generation and editing have entered a new era. However, existing methods often struggle to consider each object both individually and part of the whole image editing, both of which are crucial for ensuring consistent quantity perception, resulting in suboptimal perceptual performance. To address these challenges, we propose MoEdit, an auxiliary-free multi-object image editing framework. MoEdit facilitates high-quality multi-object image editing in terms of style transfer, object reinvention, and background regeneration, while ensuring consistent quantity perception between inputs and outputs, even with a large number of objects. To achieve this, we introduce the Feature Compensation (FeCom) module, which ensures the distinction and separability of each object attribute by minimizing the in-between interlacing. Additionally, we present the Quantity Attention (QTTN) module, which perceives and preserves quantity consistency by effective control in editing, without relying on auxiliary tools. By leveraging the SD model, MoEdit enables customized preservation and modification of specific concepts in inputs with high quality. Experimental results demonstrate that our MoEdit achieves State-Of-The-Art (SOTA) performance in multi-object image editing. Data and codes will be available at https://github.com/Tear-kitty/MoEdit.

📄 PDF Abstract BibTeX arXiv:2503.10112

Code (1)

tear-kitty/moedit 공식 구현

Tasks

AttributeImage GenerationObjectStyle Transfer

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

EmoEdit: Evoking Emotions through Image Manipulation

2024-05-21 · CVPR 2025 1 · Jingyuan Yang, Jiawei Feng, Weibin Luo, Dani Lischinski 외

Affective Image Manipulation (AIM) seeks to modify user-provided images to evoke specific emotional responses. This task is inherently complex due to its twofold objective: significantly evoking the intended emotion, whi…

Image ManipulationLanguage Modelling

Visual Perception, Quantity of Information Function and the Concept of the Quantity of Information Continuous Splines

2025-02-03 · Rushan Ziatdinov

The geometric shapes of the outside world objects hide an undisclosed emotional, psychological, artistic, aesthetic and shape-generating potential; they may attract or cause fear as well as a variety of other emotions. T…

Situational Perception Guided Image Matting

2022-04-20 · Bo Xu, Jiake Xie, Han Huang, Ziwen Li 외

Most automatic matting methods try to separate the salient foreground from the background. However, the insufficient quantity and subjective bias of the current existing matting datasets make it difficult to fully explor…

Image MattingObject

A Number Sense as an Emergent Property of the Manipulating Brain

2020-12-08 · Neehar Kondapaneni, Pietro Perona

The ability to understand and manipulate numbers and quantities emerges during childhood, but the mechanism through which humans acquire and develop this ability is still poorly understood. We explore this question throu…

CountDiffusion: Text-to-Image Synthesis with Training-Free Counting-Guidance Diffusion

2025-05-07 · Yanyu Li, Pencheng Wan, Liang Han, YaoWei Wang 외

Stable Diffusion has advanced text-to-image synthesis, but training models to generate images with accurate object quantity is still difficult due to the high computational cost and the challenge of teaching models the a…

DenoisingImage GenerationObject