paper-with-me

Papers

CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models

2024-06-13 · Yigit Ekin, Ahmet Burak Yildirim, Erdem Eren Caglar, Aykut Erdem, Erkut Erdem, Aysegul Dundar

Advanced image editing techniques, particularly inpainting, are essential for seamlessly removing unwanted elements while preserving visual integrity. Traditional GAN-based methods have achieved notable success, but recent advancements in diffusion models have produced superior results due to their training on large-scale datasets, enabling the generation of remarkably realistic inpainted images. Despite their strengths, diffusion models often struggle with object removal tasks without explicit guidance, leading to unintended hallucinations of the removed object. To address this issue, we introduce CLIPAway, a novel approach leveraging CLIP embeddings to focus on background regions while excluding foreground elements. CLIPAway enhances inpainting accuracy and quality by identifying embeddings that prioritize the background, thus achieving seamless object removal. Unlike other methods that rely on specialized training datasets or costly manual annotations, CLIPAway provides a flexible, plug-and-play solution compatible with various diffusion-based inpainting techniques.

📄 PDF Abstract BibTeX arXiv:2406.09368

Code (1)

YigitEkin/CLIPAway 공식 구현 pytorch

Tasks

Object

Methods 이 논문이 사용한 방법론

Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.
Focus 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Harmonizing Lexical Data for their Linking to Knowledge Objects in the Linked Data Framework

2014-08-01 · WS 2014 8 · Thierry Declerck

Harmonizing Diverse Models: A Layer-wise Merging Strategy for Consistent Generation

2025-10-16 · Xujun Peng, Anoop Kumar, Jingyu Wu, Parker Glenn 외 arxiv

Retrieval-Augmented Generation (RAG) systems leverage Large Language Models (LLMs) to generate accurate and reliable responses that are grounded in retrieved context. However, LLMs often generate inconsistent outputs for…

Synthetic Data Generation

Removing Dynamic Objects for Static Scene Reconstruction using Light Fields

2020-03-24 · Pushyami Kaveti, Sammie Katt, Hanumant Singh

There is a general expectation that robots should operate in environments that consist of static and dynamic entities including people, furniture and automobiles. These dynamic environments pose challenges to visual simu…

GPUSemantic SegmentationSimultaneous Localization and Mapping

OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward

2025-08-26 · Chunlin Zhong, Qiuxia Hou, Zhangjun Zhou, Shuang Hao 외 arxiv

Video captioning aims to generate comprehensive and coherent descriptions of the video content, contributing to the advancement of both video understanding and generation. However, existing methods often suffer from moti…

Video Captioning

Using Word Embeddings to Analyze Protests News

2022-03-11 · Maria Alejandra Cardoza Ceron

The first two tasks of the CLEF 2019 ProtestNews events focused on distinguishing between protest and non-protest related news articles and sentences in a binary classification task. Among the submissions, two well perfo…

ArticlesBinary ClassificationWord Embeddings