paper-with-me

Papers

DINO-QPM: Adapting Visual Foundation Models for Globally Interpretable Image Classification

2026-04-08 · Robert Zimmermann, Thomas Norrenbrock, Bodo Rosenhahn arxiv

Although visual foundation models like DINOv2 provide state-of-the-art performance as feature extractors, their complex, high-dimensional representations create substantial hurdles for interpretability. This work proposes DINO-QPM, which converts these powerful but entangled features into contrastive, class-independent representations that are interpretable by humans. DINO-QPM is a lightweight interpretability adapter that pursues globally interpretable image classification, adapting the Quadratic Programming Enhanced Model (QPM) to operate on strictly frozen DINO backbones. While classification with visual foundation models typically relies on the \texttt{CLS} token, we deliberately diverge from this standard. By leveraging average-pooling, we directly connect the patch embeddings to the model's features and therefore enable spatial localisation of DINO-QPM's globally interpretable features within the input space. Furthermore, we apply a sparsity loss to minimise spatial scatter and background noise, ensuring that explanations are grounded in relevant object parts. With DINO-QPM we make the level of interpretability of QPM available as an adapter while exceeding the accuracy of DINOv2 linear probe. Evaluated through an introduced Plausibility metric and other interpretability metrics, extensive experiments demonstrate that DINO-QPM is superior to other applicable methods for frozen visual foundation models in both classification accuracy and explanation quality.

📄 PDF Abstract BibTeX arXiv:2604.07166

Code (0)

등록된 구현이 없습니다.

Tasks

Image Classification

Similar Papers 제목 키워드 기반

DINO-Mix: Enhancing Visual Place Recognition with Foundational Vision Model and Feature Mixing

2023-11-01 · Gaoshuang Huang, Yang Zhou, Xiaofei Hu, Chenglong Zhang 외

Utilizing visual place recognition (VPR) technology to ascertain the geographical location of publicly available images is a pressing issue for real-world VPR applications. Although most current VPR methods achieve favor…

Visual Place Recognition

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

2025-01-24 · Tianming Liang, Kun-Yu Lin, Chaolei Tan, JianGuo Zhang 외

Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. Despite notable progress in recent years, current RVOS models remain struggle to handle complicate…

DecoderObjectReferring Expression SegmentationReferring Video Object Segmentation+4

NeuroSeg Meets DINOv3: Transferring 2D Self-Supervised Visual Priors to 3D Neuron Segmentation via DINOv3 Initialization

2026-03-24 · Yik San Cheng, Runkai Zhao, Weidong Cai arxiv

2D visual foundation models, such as DINOv3, a self-supervised model trained on large-scale natural images, have demonstrated strong zero-shot generalization, capturing both rich global context and fine-grained structura…

Zero-shot Generalization

MedDINOv3: How to adapt vision foundation models for medical image segmentation?

2025-09-02 · Yuheng Li, Yizhou Wu, Yuxiang Lai, Mingzhe Hu 외 arxiv

Accurate segmentation of organs and tumors in CT and MRI scans is essential for diagnosis, treatment planning, and disease monitoring. While deep learning has advanced automated segmentation, most models remain task-spec…

Medical Image Segmentation

Block Expanded DINORET: Adapting Natural Domain Foundation Models for Retinal Imaging Without Catastrophic Forgetting

2024-09-25 · Jay Zoellin, Colin Merk, Mischa Buob, Amr Saad 외

Integrating deep learning into medical imaging is poised to greatly advance diagnostic methods but it faces challenges with generalizability. Foundation models, based on self-supervised learning, address these issues and…

DiagnosticDomain AdaptationFew-Shot Learningparameter-efficient fine-tuning+1