paper-with-me

홈 › Papers

Step1X-Edit: A Practical Framework for General Image Editing

2025-04-24 · Shiyu Liu, Yucheng Han, Peng Xing, Fukun Yin, Rui Wang, Wei Cheng, Jiaqi Liao, Yingming Wang, Honghao Fu, Chunrui Han, Guopeng Li, Yuang Peng, Quan Sun, Jingwei Wu, Yan Cai, Zheng Ge, Ranchen Ming, Lei Xia, Xianfang Zeng, Yibo Zhu, Binxing Jiao, Xiangyu Zhang, Gang Yu, Daxin Jiang

In recent years, image editing models have witnessed remarkable and rapid development. The recent unveiling of cutting-edge multimodal models such as GPT-4o and Gemini2 Flash has introduced highly promising image editing capabilities. These models demonstrate an impressive aptitude for fulfilling a vast majority of user-driven editing requirements, marking a significant advancement in the field of image manipulation. However, there is still a large gap between the open-source algorithm with these closed-source models. Thus, in this paper, we aim to release a state-of-the-art image editing model, called Step1X-Edit, which can provide comparable performance against the closed-source models like GPT-4o and Gemini2 Flash. More specifically, we adopt the Multimodal LLM to process the reference image and the user's editing instruction. A latent embedding has been extracted and integrated with a diffusion image decoder to obtain the target image. To train the model, we build a data generation pipeline to produce a high-quality dataset. For evaluation, we develop the GEdit-Bench, a novel benchmark rooted in real-world user instructions. Experimental results on GEdit-Bench demonstrate that Step1X-Edit outperforms existing open-source baselines by a substantial margin and approaches the performance of leading proprietary models, thereby making significant contributions to the field of image editing.

📄 PDF Abstract BibTeX arXiv:2504.17761

Code (1)

stepfun-ai/step1x-edit 공식 구현 pytorch

Tasks

Image EditingImage Manipulation

Methods 이 논문이 사용한 방법론

ADOPT Please enter a description about the method here
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

OSVE: One Step Video Editing with One Step Diffusion Models

2026-07-22 · Habin Lim, Gyeong-Moon Park arxiv

Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling and inversion. We present OSVE, the first framework to successfully adapt one-step Text-to-Image (T2I) models …

Not All Steps are Created Equal: Selective Diffusion Distillation for Image Manipulation

2023-07-17 · ICCV 2023 1 · Luozhou Wang, Shuai Yang, Shu Liu, Ying-Cong Chen

Conditional diffusion models have demonstrated impressive performance in image manipulation tasks. The general pipeline involves adding noise to the image and then denoising it. However, this method faces a trade-off pro…

AllDenoisingImage Manipulation

Rethinking One-Step Image Editing through ChordEdit: Reproduction, Simplification, and New Insights

2026-06-12 · Minghan Li, Jeremy Moebel, Mengyu Wang arxiv

One-step image editing is important for making text-guided editing fast, practical, and easy to deploy, but its underlying mechanism is still not fully understood. We revisit ChordEdit through reproduction, ablation, and…

Image Editing

AutoEdit: Automatic Hyperparameter Tuning for Image Editing

2025-09-18 · Chau Pham, Quan Dao, Mahesh Bhosale, Yunjie Tian 외 arxiv

Recent advances in diffusion models have revolutionized text-guided image editing, yet existing editing methods face critical challenges in hyperparameter identification. To get the reasonable editing performance, these …

Reinforcement LearningImage Editing

High-Fidelity Diffusion-based Image Editing

2023-12-25 · Chen Hou, Guoqiang Wei, Zhibo Chen

Diffusion models have attained remarkable success in the domains of image generation and editing. It is widely recognized that employing larger inversion and denoising steps in diffusion model leads to improved image rec…

DenoisingImage GenerationImage ReconstructionImage-to-Image Translation