paper-with-me

홈 › Papers

Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs

2025-05-02 · Hari Chandana Kuchibhotla, Sai Srinivas Kancheti, Abbavaram Gowtham Reddy, Vineeth N Balasubramanian

Fine-grained Visual Recognition (FGVR) involves distinguishing between visually similar categories, which is inherently challenging due to subtle inter-class differences and the need for large, expert-annotated datasets. In domains like medical imaging, such curated datasets are unavailable due to issues like privacy concerns and high annotation costs. In such scenarios lacking labeled data, an FGVR model cannot rely on a predefined set of training labels, and hence has an unconstrained output space for predictions. We refer to this task as Vocabulary-Free FGVR (VF-FGVR), where a model must predict labels from an unconstrained output space without prior label information. While recent Multimodal Large Language Models (MLLMs) show potential for VF-FGVR, querying these models for each test input is impractical because of high costs and prohibitive inference times. To address these limitations, we introduce \textbf{Nea}rest-Neighbor Label \textbf{R}efinement (NeaR), a novel approach that fine-tunes a downstream CLIP model using labels generated by an MLLM. Our approach constructs a weakly supervised dataset from a small, unlabeled training set, leveraging MLLMs for label generation. NeaR is designed to handle the noise, stochasticity, and open-endedness inherent in labels generated by MLLMs, and establishes a new benchmark for efficient VF-FGVR.

📄 PDF Abstract BibTeX arXiv:2505.01064

Code (0)

등록된 구현이 없습니다.

Tasks

Fine-Grained Visual Recognition

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs

2025-12-21 · Dmitry Demidov, Zaigham Zaheer, Zongyan Han, Omkar Thawakar 외 arxiv

Vocabulary-free fine-grained image recognition aims to distinguish visually similar categories within a meta-class without a fixed, human-defined label set. Existing solutions for this problem are limited by either the u…

Fine-Grained Visual RecognitionFine-Grained Image Recognition

Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model

2025-07-30 · Dmitry Demidov, Zaigham Zaheer, Omkar Thawakar, Salman Khan 외 arxiv

Fine-grained image classification, the task of distinguishing between visually similar subcategories within a broader category (e.g., bird species, car models, flower types), is a challenging computer vision problem. Tra…

Fine-Grained Image ClassificationFine-Grained Visual Recognition

ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction

2025-06-10 · Juan Yeo, Soonwoo Cha, Jiwoo Song, Hyunbin Jin 외

Vision-language models such as CLIP have recently propelled open-vocabulary dense prediction tasks by enabling recognition of a broad range of visual concepts. However, CLIP still struggles with fine-grained, region-leve…

object-detectionObject DetectionOpen-vocabulary object detectionOpen Vocabulary Object Detection+1

The devil is in the fine-grained details: Evaluating open-vocabulary object detectors for fine-grained understanding

2023-11-29 · CVPR 2024 1 · Lorenzo Bianchi, Fabio Carrara, Nicola Messina, Claudio Gennaro 외

Recent advancements in large vision-language models enabled visual object detection in open-vocabulary scenarios, where object classes are defined in free-text formats during inference. In this paper, we aim to probe the…

Objectobject-detectionObject DetectionOpen-vocabulary object detection+1

Is CLIP the main roadblock for fine-grained open-world perception?

2024-04-04 · Lorenzo Bianchi, Fabio Carrara, Nicola Messina, Fabrizio Falchi

Modern applications increasingly demand flexible computer vision models that adapt to novel concepts not encountered during training. This necessity is pivotal in emerging domains like extended reality, robotics, and aut…

Autonomous DrivingNovel ConceptsObjectobject-detection+4