paper-with-me

홈 › Papers

Language-Guided Global Image Editing via Cross-Modal Cyclic Mechanism

2021-01-01 · ICCV 2021 10 · Wentao Jiang, Ning Xu, Jiayun Wang, Chen Gao, Jing Shi, Zhe Lin, Si Liu

Editing an image automatically via a linguistic request can significantly save laborious manual work and is friendly to photography novice. In this paper, we focus on the task of language-guided global image editing. Existing works suffer from imbalanced data distribution of real-world datasets and thus fail to understand language requests well. To handle this issue, we propose to create a cycle with our image generator by creating another model called Editing Description Network (EDNet) which predicts an editing embedding given a pair of images. Given the cycle, we propose several free augmentation strategies to help our model understand various editing requests given the imbalanced dataset. In addition, two other novel ideas are proposed: an Image-Request Attention (IRA) module which allows our method to edit an image spatial-adaptively when the image requires different editing degree at different regions, as well as a new evaluation metric for this task which is more semantic and reasonable than conventional pixel losses (eg L1). Extensive experiments on two benchmark datasets demonstrate the effectiveness of our method over existing approaches.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning by Planning: Language-Guided Global Image Editing

2021-06-24 · CVPR 2021 1 · Jing Shi, Ning Xu, Yihang Xu, Trung Bui 외

Recently, language-guided global image editing draws increasing attention with growing application potentials. However, previous GAN-based methods are not only confined to domain-specific, low-resolution data but also la…

3DitScene: Editing Any Scene via Language-guided Disentangled Gaussian Splatting

2024-05-28 · Qihang Zhang, Yinghao Xu, Chaoyang Wang, Hsin-Ying Lee 외

Scene image editing is crucial for entertainment, photography, and advertising design. Existing methods solely focus on either 2D individual object or 3D global scene editing. This results in a lack of a unified approach…

3D geometryDisentanglement

Guiding Instruction-based Image Editing via Multimodal Large Language Models

2023-09-29 · Tsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang 외

Instruction-based image editing improves the controllability and flexibility of image manipulation via natural commands without elaborate descriptions or regional masks. However, human instructions are sometimes too brie…

Image ManipulationResponse Generation

MiVE: Multiscale Vision-language features for reference-guided video Editing

2026-05-14 · Tong Wang, Meng Zou, Chengjing Wu, Xiaochao Qu 외 arxiv

Reference-guided video editing takes a source video, a text instruction, and a reference image as inputs, requiring the model to faithfully apply the instructed edits while preserving original motion and unedited content…

Structure-Preserving Zero-Shot Image Editing via Stage-Wise Latent Injection in Diffusion Models

2025-04-22 · Dasol Jeong, Donggoo Kang, Jiwon Park, Hyebean Lee 외

We propose a diffusion-based framework for zero-shot image editing that unifies text-guided and reference-guided approaches without requiring fine-tuning. Our method leverages diffusion inversion and timestep-specific nu…

Attribute