Regression-Based Image Alignment for General Object Categories
Gradient-descent methods have exhibited fast and reliable performance for image alignment in the facial domain, but have largely been ignored by the broader vision community. They require the image function be smooth and (numerically) differentiable -- properties that hold for pixel-based representations obeying natural image statistics, but not for more general classes of non-linear feature transforms. We show that transforms such as Dense SIFT can be incorporated into a Lucas Kanade alignment framework by predicting descent directions via regression. This enables robust matching of instances from general object categories whilst maintaining desirable properties of Lucas Kanade such as the capacity to handle high-dimensional warp parametrizations and a fast rate of convergence. We present alignment results on a number of objects from ImageNet, and an extension of the method to unsupervised joint alignment of objects from a corpus of images.
Code (0)
등록된 구현이 없습니다.
Tasks
ObjectregressionSimilar Papers 제목 키워드 기반
EdaDet: Open-Vocabulary Object Detection Using Early Dense Alignment
Vision-language models such as CLIP have boosted the performance of open-vocabulary object detection, where the detector is trained on base categories but required to detect novel categories. Existing methods leverage CL…
Objectobject-detectionObject DetectionOpen-vocabulary object detection+2Zero-shot Inexact CAD Model Alignment from a Single Image
One practical approach to infer 3D scene structure from a single image is to retrieve a closely matching 3D model from a database and align it with the object in the image. Existing methods rely on supervised training wi…
Expanding Zero-Shot Object Counting with Rich Prompts
Expanding pre-trained zero-shot counting models to handle unseen categories requires more than simply adding new prompts, as this approach does not achieve the necessary alignment between text and visual features for acc…
ObjectObject CountingZero-Shot CountingOpen-Vocabulary Object Detection With an Open Corpus
Existing open vocabulary object detection (OVD) works expand the object detector toward open categories by replacing the classifier with the category text embeddings and optimizing the region-text alignment on data o…
Objectobject-detectionObject DetectionOpen-vocabulary object detection+1GridCLIP: One-Stage Object Detection by Grid-Level CLIP Representation Learning
A vision-language foundation model pretrained on very large-scale image-text paired data has the potential to provide generalizable knowledge representation for downstream visual recognition and detection tasks, especial…
object-detectionObject DetectionRepresentation Learning