paper-with-me

홈 › Papers

Magic Fixup: Streamlining Photo Editing by Watching Dynamic Videos

2024-03-19 · Hadi AlZayer, Zhihao Xia, Xuaner Zhang, Eli Shechtman, Jia-Bin Huang, Michael Gharbi

We propose a generative model that, given a coarsely edited image, synthesizes a photorealistic output that follows the prescribed layout. Our method transfers fine details from the original image and preserves the identity of its parts. Yet, it adapts it to the lighting and context defined by the new layout. Our key insight is that videos are a powerful source of supervision for this task: objects and camera motions provide many observations of how the world changes with viewpoint, lighting, and physical interactions. We construct an image dataset in which each sample is a pair of source and target frames extracted from the same video at randomly chosen time intervals. We warp the source frame toward the target using two motion models that mimic the expected test-time user edits. We supervise our model to translate the warped image into the ground truth, starting from a pretrained diffusion model. Our model design explicitly enables fine detail transfer from the source frame to the generated image, while closely following the user-specified layout. We show that by using simple segmentations and coarse 2D manipulations, we can synthesize a photorealistic edit faithful to the user's input while addressing second-order effects like harmonizing the lighting and physical interactions between edited objects.

📄 PDF Abstract BibTeX arXiv:2403.13044

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

3D-Fixup: Advancing Photo Editing with 3D Priors

2025-05-15 · Yen-Chi Cheng, Krishna Kumar Singh, Jae Shin Yoon, Alex Schwing 외

Despite significant advances in modeling image priors via diffusion models, 3D-aware image editing remains challenging, in part because the object is only specified via a single image. To tackle this challenge, we propos…

Image ManipulationImage to 3D

MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing

2023-06-16 · NeurIPS 2023 11 · Kai Zhang, Lingbo Mo, Wenhu Chen, Huan Sun 외

Text-guided image editing is widely needed in daily life, ranging from personal use to professional applications such as Photoshop. However, existing methods are either zero-shot or trained on an automatically synthesize…

Image Editingtext-guided-image-editing

UltraEdit: Instruction-based Fine-Grained Image Editing at Scale

2024-07-07 · Haozhe Zhao, Xiaojian Ma, Liang Chen, Shuzheng Si 외

This paper presents UltraEdit, a large-scale (approximately 4 million editing samples), automatically generated dataset for instruction-based image editing. Our key idea is to address the drawbacks in existing image edit…

DiversityImage Editing

MagicProp: Diffusion-based Video Editing via Motion-aware Appearance Propagation

2023-09-02 · Hanshu Yan, Jun Hao Liew, Long Mai, Shanchuan Lin 외

This paper addresses the issue of modifying the visual appearance of videos while preserving their motion. A novel framework, named MagicProp, is proposed, which disentangles the video editing process into two stages: ap…

Video Editing

MagicEdit: High-Fidelity and Temporally Coherent Video Editing

2023-08-28 · Jun Hao Liew, Hanshu Yan, Jianfeng Zhang, Zhongcong Xu 외

In this report, we present MagicEdit, a surprisingly simple yet effective solution to the text-guided video editing task. We found that high-fidelity and temporally coherent video-to-video translation can be achieved by …

TranslationVideo Editing