paper-with-me

Papers

VENUS: Visual Editing with Noise Inversion Using Scene Graphs

2026-01-12 · Thanh-Nhan Vo, Trong-Thuan Nguyen, Tam V. Nguyen, Minh-Triet Tran arxiv

State-of-the-art text-based image editing models often struggle to balance background preservation with semantic consistency, frequently resulting either in the synthesis of entirely new images or in outputs that fail to realize the intended edits. In contrast, scene graph-based image editing addresses this limitation by providing a structured representation of semantic entities and their relations, thereby offering improved controllability. However, existing scene graph editing methods typically depend on model fine-tuning, which incurs high computational cost and limits scalability. To this end, we introduce VENUS (Visual Editing with Noise inversion Using Scene graphs), a training-free framework for scene graph-guided image editing. Specifically, VENUS employs a split prompt conditioning strategy that disentangles the target object of the edit from its background context, while simultaneously leveraging noise inversion to preserve fidelity in unedited regions. Moreover, our proposed approach integrates scene graphs extracted from multimodal large language models with diffusion backbones, without requiring any additional training. Empirically, VENUS substantially improves both background preservation and semantic alignment on PIE-Bench, increasing PSNR from 22.45 to 24.80, SSIM from 0.79 to 0.84, and reducing LPIPS from 0.100 to 0.070 relative to the state-of-the-art scene graph editing model (SGEdit). In addition, VENUS enhances semantic consistency as measured by CLIP similarity (24.97 vs. 24.19). On EditVal, VENUS achieves the highest fidelity with a 0.87 DINO score and, crucially, reduces per-image runtime from 6-10 minutes to only 20-30 seconds. Beyond scene graph-based editing, VENUS also surpasses strong text-based editing baselines such as LEDIT++ and P2P+DirInv, thereby demonstrating consistent improvements across both paradigms.

📄 PDF Abstract BibTeX arXiv:2601.07219

Code (0)

등록된 구현이 없습니다.

Tasks

Text-based Image Editing

Similar Papers 제목 키워드 기반

Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing

2025-09-02 · Quan Dao, Xiaoxiao He, Ligong Han, Ngan Hoai Nguyen 외 arxiv

Visual autoregressive models (VAR) have recently emerged as a promising class of generative models, achieving performance comparable to diffusion models in text-to-image generation tasks. While conditional generation has…

Text-based Image EditingText-to-Image Generation

Make It So: Steering StyleGAN for Any Image Inversion and Editing

2023-04-27 · Anand Bhattad, Viraj Shah, Derek Hoiem, D. A. Forsyth

StyleGAN's disentangled style representation enables powerful image editing by manipulating the latent variables, but accurately mapping real-world images to their latent variables (GAN inversion) remains a challenge. Ex…

Dual-Schedule Inversion: Training- and Tuning-Free Inversion for Real Image Editing

2024-12-15 · Jiancheng Huang, Yi Huang, Jianzhuang Liu, Donghao Zhou 외

Text-conditional image editing is a practical AIGC task that has recently emerged with great commercial and academic value. For real image editing, most diffusion model-based methods use DDIM Inversion as the first stage…

Consistent-Inversion: Reverse Consistency Guidance for Structure-Preserving Visual Editing

2026-06-05 · Xiaocheng Lu, Jingcai Guo, Song Guo arxiv

Text-guided diffusion models have become effective tools for real-image visual editing, where the edited image must follow a target instruction while preserving editing-irrelevant structure. Most training-free editors re…

Wavelet-Guided Semantic Signal Compensation for Inversion-Free Image Editing

2026-07-02 · Anqi Tang, Wenhao Sun, Zhaoqiang Liu arxiv

Text-guided image editing aims to modify visual content according to a target prompt while preserving the background. Recent inversion-free image editing frameworks such as FlowEdit have demonstrated strong editing capab…

Image Editing