paper-with-me

홈 › Papers

OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision

2024-11-11 · Cong Wei, Zheyang Xiong, Weiming Ren, Xinrun Du, Ge Zhang, Wenhu Chen

Instruction-guided image editing methods have demonstrated significant potential by training diffusion models on automatically synthesized or manually annotated image editing pairs. However, these methods remain far from practical, real-life applications. We identify three primary challenges contributing to this gap. Firstly, existing models have limited editing skills due to the biased synthesis process. Secondly, these methods are trained with datasets with a high volume of noise and artifacts. This is due to the application of simple filtering methods like CLIP-score. Thirdly, all these datasets are restricted to a single low resolution and fixed aspect ratio, limiting the versatility to handle real-world use cases. In this paper, we present \omniedit, which is an omnipotent editor to handle seven different image editing tasks with any aspect ratio seamlessly. Our contribution is in four folds: (1) \omniedit is trained by utilizing the supervision from seven different specialist models to ensure task coverage. (2) we utilize importance sampling based on the scores provided by large multimodal models (like GPT-4o) instead of CLIP-score to improve the data quality. (3) we propose a new editing architecture called EditNet to greatly boost the editing success rate, (4) we provide images with different aspect ratios to ensure that our model can handle any image in the wild. We have curated a test set containing images of different aspect ratios, accompanied by diverse instructions to cover different tasks. Both automatic evaluation and human evaluations demonstrate that \omniedit can significantly outperform all the existing models. Our code, dataset and model will be available at \url{https://tiger-ai-lab.github.io/OmniEdit/}

📄 PDF Abstract BibTeX arXiv:2411.07199

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

OmniEdit: A Training-free framework for Lip Synchronization and Audio-Visual Editing

2026-03-10 · Lixiang Lin, Siyuan Jin, Jinshan Zhang arxiv

Lip synchronization and audio-visual editing have emerged as fundamental challenges in multimodal learning, underpinning a wide range of applications, including film production, virtual avatars, and telepresence. Despite…

RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward

2026-02-19 · Qiucheng Wu, Jing Shi, Simon Jenni, Kushal Kafle 외 arxiv

Recent advances in multimodal large language models (MLLMs) have shown great potential for extending vision-language reasoning to professional tool-based image editing, enabling intuitive and creative editing. A promisin…

Reinforcement LearningMultimodal ReasoningImage Editing

UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing

2025-04-22 · Tsu-Jui Fu, Yusu Qian, Chen Chen, Wenze Hu 외

Text-to-Image (T2I) diffusion models have shown impressive results in generating visually compelling images following user prompts. Building on this, various methods further fine-tune the pre-trained T2I model for specif…

Depth EstimationImage Generation

PhotoDoodle: Learning Artistic Image Editing from Few-Shot Pairwise Data

2025-02-20 · Shijie Huang, Yiren Song, Yuxuan Zhang, Hailong Guo 외

We introduce PhotoDoodle, a novel image editing framework designed to facilitate photo doodling by enabling artists to overlay decorative elements onto photographs. Photo doodling is challenging because the inserted elem…

Style Transfer

GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

2025-07-28 · Yuhan Wang, Siwei Yang, Bingchen Zhao, Letian Zhang 외 arxiv

Recent advancements in large multimodal models like GPT-4o have set a new standard for high-fidelity, instruction-guided image editing. However, the proprietary nature of these models and their training data creates a si…

Instruction FollowingImage Editing