paper-with-me

Papers

Diffusion Representations for Fine-Grained Image Classification: A Marine Plankton Case Study

2026-01-19 · A. Nieto Juscafresa, Á. Mazcuñán Herreros, J. Sullivan arxiv

Diffusion models have emerged as state-of-the-art generative methods for image synthesis, yet their potential as general-purpose feature encoders remains underexplored. Trained for denoising and generation without labels, they can be interpreted as self-supervised learners that capture both low- and high-level structure. We show that a frozen diffusion backbone enables strong fine-grained recognition by probing intermediate denoising features across layers and timesteps and training a linear classifier for each pair. We evaluate this in a real-world plankton-monitoring setting with practical impact, using controlled and comparable training setups against established supervised and self-supervised baselines. Frozen diffusion features are competitive with supervised baselines and outperform other self-supervised methods in both balanced and naturally long-tailed settings. Out-of-distribution evaluations on temporally and geographically shifted plankton datasets further show that frozen diffusion features maintain strong accuracy and Macro F1 under substantial distribution shift.

📄 PDF Abstract BibTeX arXiv:2601.13416

Code (0)

등록된 구현이 없습니다.

Tasks

Fine-Grained Image Classification

Similar Papers 제목 키워드 기반

TaxaDiffusion: Progressively Trained Diffusion Model for Fine-Grained Species Generation

2025-06-02 · Amin Karimi Monsefi, Mridul Khurana, Rajiv Ramnath, Anuj Karpatne 외

We propose TaxaDiffusion, a taxonomy-informed training framework for diffusion models to generate fine-grained animal images with high morphological and identity accuracy. Unlike standard approaches that treat each speci…

Image GenerationTransfer Learning

Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control

2024-05-09 · Gunshi Gupta, Karmesh Yadav, Yarin Gal, Dhruv Batra 외

Embodied AI agents require a fine-grained understanding of the physical world mediated through visual and language inputs. Such capabilities are difficult to learn solely from task-specific data. This has led to the emer…

Representation LearningScene Understanding

Text-to-Image Diffusion Models are Zero-Shot Classifiers

2023-03-27 · Kevin Clark, Priyank Jaini

The excellent generative capabilities of text-to-image diffusion models suggest they learn informative representations of image-text data. However, what knowledge their representations capture is not fully understood, an…

AttributeContrastive Learningimage-classificationImage Classification+1

Text-to-Image Diffusion Models are Zero Shot Classifiers

2023-09-21 · NeurIPS 2023 11

The excellent generative capabilities of text-to-image diffusion models suggest they learn informative representations of image-text data. However, what knowledge their representations capture is not fully understood, an…

Do text-free diffusion models learn discriminative visual representations?

2023-11-29 · Soumik Mukhopadhyay, Matthew Gwilliam, Yosuke Yamaguchi, Vatsal Agarwal 외

While many unsupervised learning models focus on one family of tasks, either generative or discriminative, we explore the possibility of a unified representation learner: a model which addresses both families of tasks si…

image-classificationImage Classificationobject-detectionObject Detection+2