paper-with-me

홈 › Papers

Tune-An-Ellipse: CLIP Has Potential to Find What You Want

2024-01-01 · CVPR 2024 1 · Jinheng Xie, Songhe Deng, Bing Li, Haozhe Liu, Yawen Huang, Yefeng Zheng, Jurgen Schmidhuber, Bernard Ghanem, Linlin Shen, Mike Zheng Shou

Visual prompting of large vision language models such as CLIP exhibits intriguing zero-shot capabilities. A manually drawn red circle commonly used for highlighting can guide CLIP's attention to the surrounding region to identify specific objects within an image. Without precise object proposals however it is insufficient for localization. Our novel simple yet effective approach i.e. Differentiable Visual Prompting enables CLIP to zero-shot localize: given an image and a text prompt describing an object we first pick a rendered ellipse from uniformly distributed anchor ellipses on the image grid via visual prompting then use three loss functions to tune the ellipse coefficients to encapsulate the target region gradually. This yields promising experimental results for referring expression comprehension without precisely specified object proposals. In addition we systematically present the limitations of visual prompting inherent in CLIP and discuss potential solutions.

📄 PDF Abstract BibTeX

Code (1)

showlab/tune-an-ellipse 공식 구현 pytorch

Tasks

ObjectReferring ExpressionReferring Expression ComprehensionVisual Prompting

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

EllipseLIO: Adaptive LiDAR Inertial Odometry with an Ellipsoid Representation

2026-05-20 · Rowan Border, Margarita Chli arxiv

LiDAR Inertial Odometry (LIO) is a critical component for many mobile robots that need to navigate without relying on external positioning (e.g., GPS). Platforms that operate autonomously in different environments and wi…

Are Ellipses Important for Machine Translation?

2021-12-01 · CL (ACL) 2021 12 · Payal Khullar

Abstract This article describes an experiment to evaluate the impact of different types of ellipses discussed in theoretical linguistics on Neural Machine Translation (NMT), using English to Hindi/Telugu as source and ta…

Machine TranslationNMTTranslation

Multi-ellipses detection on images inspired by collective animal behavior

2014-05-20 · Erik Cuevas, Maurici Gonzalez, Daniel Zaldivar, Marco Perez

This paper presents a novel and effective technique for extracting multiple ellipses from an image. The approach employs an evolutionary algorithm to mimic the way animals behave collectively assuming the overall detecti…

Homography Estimation From the Common Self-Polar Triangle of Separate Ellipses

2016-06-01 · CVPR 2016 6 · Haifei Huang, HUI ZHANG, Yiu-ming Cheung

How to avoid ambiguity is a challenging problem for conic-based homography estimation. In this paper, we address the problem of homography estimation from two separate ellipses. We find that any two ellipses have a uniqu…

Homography Estimation

Detecting AI-Generated Images via CLIP

2024-04-12 · A. G. Moskowitz, T. Gaona, J. Peterson

As AI-generated image (AIGI) methods become more powerful and accessible, it has become a critical task to determine if an image is real or AI-generated. Because AIGI lack the signatures of photographs and have their own…

GPU