paper-with-me

Papers

Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceivers

2025-11-24 · Yiqing Shi, Yiren Song, Mike Zheng Shou arxiv

Recent advances in diffusion transformers have shown remarkable generalization in visual synthesis, yet most dense perception methods still rely on text-to-image (T2I) generators designed for stochastic generation. We revisit this paradigm and show that image editing diffusion models are inherently image-to-image consistent, providing a more suitable foundation for dense perception task. We introduce Edit2Perceive, a unified diffusion framework that adapts editing models for depth, normal, and matting. Built upon the FLUX.1 Kontext architecture, our approach employs full-parameter fine-tuning and a pixel-space consistency loss to enforce structure-preserving refinement across intermediate denoising states. Moreover, our single-step deterministic inference yields up to faster runtime while training on relatively small datasets. Extensive experiments demonstrate comprehensive state-of-the-art results across all three tasks, revealing the strong potential of editing-oriented diffusion transformers for geometry-aware perception.

📄 PDF Abstract BibTeX arXiv:2511.18673

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

DiT4Edit: Diffusion Transformer for Image Editing

2024-11-05 · Kunyu Feng, Yue Ma, Bingyuan Wang, Chenyang Qi 외

Despite recent advances in UNet-based image editing, methods for shape-aware object editing in high-resolution images are still lacking. Compared to UNet, Diffusion Transformers (DiT) demonstrate superior capabilities to…

Image Generation

FeedEdit: Text-Based Image Editing with Dynamic Feedback Regulation

2025-01-01 · CVPR 2025 1 · Fengyi Fu, Lei Zhang, Mengqi Huang, Zhendong Mao

Text-based image editing which aims at generating rigid or non-rigid changes to images conditioned on the given text, has recently attracted considerable interest. Previous works mainly follow the multi-step denoisin…

DenoisingText-based Image Editing

SeedEdit: Align Image Re-Generation to Image Editing

2024-11-11 · Yichun Shi, Peng Wang, Weilin Huang

We introduce SeedEdit, a diffusion model that is able to revise a given image with any text prompt. In our perspective, the key to such a task is to obtain an optimal balance between maintaining the original image, i.e. …

Image Reconstruction

InstructX: Towards Unified Visual Editing with MLLM Guidance

2025-10-09 · Chong Mou, Qichao Sun, Yanze Wu, Pengze Zhang 외 arxiv

With recent advances in Multimodal Large Language Models (MLLMs) showing strong visual understanding and reasoning, interest is growing in using them to improve the editing performance of diffusion models. Despite rapid …

StableVideo: Text-driven Consistency-aware Diffusion Video Editing

2023-08-18 · ICCV 2023 1 · Wenhao Chai, Xun Guo, Gaoang Wang, Yan Lu

Diffusion-based methods can generate realistic images and videos, but they struggle to edit existing objects in a video while preserving their appearance over time. This prevents diffusion models from being applied to na…

Video Editing