paper-with-me

Papers

A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic Correspondence

2023-05-24 · NeurIPS 2023 11 · Junyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Polania Cabrera, Varun Jampani, Deqing Sun, Ming-Hsuan Yang

Text-to-image diffusion models have made significant advances in generating and editing high-quality images. As a result, numerous approaches have explored the ability of diffusion model features to understand and process single images for downstream tasks, e.g., classification, semantic segmentation, and stylization. However, significantly less is known about what these features reveal across multiple, different images and objects. In this work, we exploit Stable Diffusion (SD) features for semantic and dense correspondence and discover that with simple post-processing, SD features can perform quantitatively similar to SOTA representations. Interestingly, the qualitative analysis reveals that SD features have very different properties compared to existing representation learning features, such as the recently released DINOv2: while DINOv2 provides sparse but accurate matches, SD features provide high-quality spatial information but sometimes inaccurate semantic matches. We demonstrate that a simple fusion of these two features works surprisingly well, and a zero-shot evaluation using nearest neighbors on these fused features provides a significant performance gain over state-of-the-art methods on benchmark datasets, e.g., SPair-71k, PF-Pascal, and TSS. We also show that these correspondences can enable interesting applications such as instance swapping in two images.

📄 PDF Abstract BibTeX arXiv:2305.15347

Code (1)

Junyi42/sd-dino 공식 구현 pytorch

Tasks

Dense Pixel Correspondence EstimationRepresentation LearningSemantic correspondenceSemantic Segmentation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

A General Protocol to Probe Large Vision Models for 3D Physical Understanding

2023-10-10 · Guanqi Zhan, Chuanxia Zheng, Weidi Xie, Andrew Zisserman

Our objective in this paper is to probe large vision models to determine to what extent they 'understand' different physical properties of the 3D scene depicted in an image. To this end, we make the following contributio…

Stale Diffusion: Hyper-realistic 5D Movie Generation Using Old-school Methods

2024-04-01 · Joao F. Henriques, Dylan Campbell, Tengda Han

Two years ago, Stable Diffusion achieved super-human performance at generating images with super-human numbers of fingers. Following the steady decline of its technical novelty, we propose Stale Diffusion, a method that …

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

2026-07-31 · Mingxin Wang, Bin Hu, Bin Qian, Kaitao Jiang 외 hf

World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant st…

Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence

2026-05-28 · Artur Jesslen, Olaf Dünkel, Adam Kortylewski arxiv

Foundation features from self-supervised vision models and text-to-image diffusion models have proven effective for semantic correspondence estimation. However, because these features are learned primarily from 2D image …

Semantic correspondence

Gromov Wasserstein Optimal Transport for Semantic Correspondences

2026-02-03 · Francis Snelgar, Stephen Gould, Ming Xu, Liang Zheng 외 arxiv

Establishing correspondences between image pairs is a long studied problem in computer vision. With recent large-scale foundation models showing strong zero-shot performance on downstream tasks including classification a…

Semantic correspondence