paper-with-me

홈 › Papers

Combing Text-based and Drag-based Editing for Precise and Flexible Image Editing

2024-10-04 · Ziqi Jiang, Zhen Wang, Long Chen

Precise and flexible image editing remains a fundamental challenge in computer vision. Based on the modified areas, most editing methods can be divided into two main types: global editing and local editing. In this paper, we choose the two most common editing approaches (ie text-based editing and drag-based editing) and analyze their drawbacks. Specifically, text-based methods often fail to describe the desired modifications precisely, while drag-based methods suffer from ambiguity. To address these issues, we proposed \textbf{CLIPDrag}, a novel image editing method that is the first to combine text and drag signals for precise and ambiguity-free manipulations on diffusion models. To fully leverage these two signals, we treat text signals as global guidance and drag points as local information. Then we introduce a novel global-local motion supervision method to integrate text signals into existing drag-based methods by adapting a pre-trained language-vision model like CLIP. Furthermore, we also address the problem of slow convergence in CLIPDrag by presenting a fast point-tracking method that enforces drag points moving toward correct directions. Extensive experiments demonstrate that CLIPDrag outperforms existing single drag-based methods or text-based methods.

📄 PDF Abstract BibTeX arXiv:2410.03097

Code (0)

등록된 구현이 없습니다.

Tasks

Point Tracking

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Drag Your Gaussian: Effective Drag-Based Editing with Score Distillation for 3D Gaussian Splatting

2025-01-30 · Yansong Qu, Dian Chen, Xinyang Li, Xiaofan Li 외

Recent advancements in 3D scene editing have been propelled by the rapid development of generative models. Existing methods typically utilize generative models to perform text-guided editing on 3D representations, such a…

3DGS3D scene Editing

MvDrag3D: Drag-based Creative 3D Editing via Multi-view Generation-Reconstruction Priors

2024-10-21 · Honghua Chen, Yushi Lan, Yongwei Chen, Yifan Zhou 외

Drag-based editing has become popular in 2D content creation, driven by the capabilities of image generative models. However, extending this technique to 3D remains a challenge. Existing 3D drag-based editing methods, wh…

InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing

2025-10-09 · Haoran Yu, Yi Shi arxiv

Text-to-image diffusion models have shown great potential for image editing, with techniques such as text-based and object-dragging methods emerging as key approaches. However, each of these methods has inherent limitati…

Text-based Image EditingImage Reconstruction

TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation

2025-09-26 · Qihang Wang, Yaxiong Wang, Lechao Cheng, Zhun Zhong arxiv

This paper explores image editing under the joint control of text and drag interactions. While recent advances in text-driven and drag-driven editing have achieved remarkable progress, they suffer from complementary limi…

Image ManipulationImage Editing

DragScene: Interactive 3D Scene Editing with Single-view Drag Instructions

2024-12-18 · Chenghao Gu, Zhenzhe Li, Zhengqi Zhang, Yunpeng Bai 외

3D editing has shown remarkable capability in editing scenes based on various instructions. However, existing methods struggle with achieving intuitive, localized editing, such as selectively making flowers blossom. Drag…

3D scene Editing