paper-with-me

Papers

Target-Free Text-guided Image Manipulation

2022-11-26 · Wan-Cyuan Fan, Cheng-Fu Yang, Chiao-An Yang, Yu-Chiang Frank Wang

We tackle the problem of target-free text-guided image manipulation, which requires one to modify the input reference image based on the given text instruction, while no ground truth target image is observed during training. To address this challenging task, we propose a Cyclic-Manipulation GAN (cManiGAN) in this paper, which is able to realize where and how to edit the image regions of interest. Specifically, the image editor in cManiGAN learns to identify and complete the input image, while cross-modal interpreter and reasoner are deployed to verify the semantic correctness of the output image based on the input instruction. While the former utilizes factual/counterfactual description learning for authenticating the image semantics, the latter predicts the "undo" instruction and provides pixel-level supervision for the training of cManiGAN. With such operational cycle-consistency, our cManiGAN can be trained in the above weakly supervised setting. We conduct extensive experiments on the datasets of CLEVR and COCO, and the effectiveness and generalizability of our proposed method can be successfully verified. Project page: https://sites.google.com/view/wancyuanfan/projects/cmanigan.

📄 PDF Abstract BibTeX arXiv:2211.14544

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualImage Manipulation

Similar Papers 제목 키워드 기반

LDEdit: Towards Generalized Text Guided Image Manipulation via Latent Diffusion Models

2022-10-05 · Paramanand Chandramouli, Kanchana Vaishnavi Gandikota

Research in vision-language models has seen rapid developments off-late, enabling natural language-based interfaces for image generation and manipulation. Many existing text guided manipulation techniques are restricted …

Image GenerationImage ManipulationStyle TransferText to Image Generation+1

Towards Generalized and Training-Free Text-Guided Semantic Manipulation

2025-04-24 · Yu Hong, Xiao Cai, Pengpeng Zeng, Shuai Zhang 외

Text-guided semantic manipulation refers to semantically editing an image generated from a source prompt to match a target prompt, enabling the desired semantic changes (e.g., addition, removal, and style transfer) while…

Style Transfer

WristMimic: Full-Body Humanoid Control with Wrist-Guided Manipulation

2026-07-07 · Wongyun Yu, Youngwoon Kim, Minsu Cho arxiv

Retargeting human object interaction demonstrations to physics based simulation requires reproducing not only body motion but also the object motion and contacts that make manipulation succeed. However, position only han…

Object Tracking

CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable, and Controllable Text-Guided Face Manipulation

2022-10-08 · Chenliang Zhou, Fangcheng Zhong, Cengiz Oztireli

Recently introduced Contrastive Language-Image Pre-Training (CLIP) bridges images and text by embedding them into a joint latent space. This opens the door to ample literature that aims to manipulate an input image by pr…

DisentanglementImage Manipulation

Reversible Inversion for Training-Free Exemplar-guided Image Editing

2025-12-01 · Yuke Li, Lianli Gao, Ji Zhang, Pengpeng Zeng 외 arxiv

Exemplar-guided Image Editing (EIE) aims to modify a source image according to a visual reference. Existing approaches often require large-scale pre-training to learn relationships between the source and reference images…

Image Editing