paper-with-me

Papers

Active Learning for Vision-Language Models

2024-10-29 · Bardia Safaei, Vishal M. Patel

Pre-trained vision-language models (VLMs) like CLIP have demonstrated impressive zero-shot performance on a wide range of downstream computer vision tasks. However, there still exists a considerable performance gap between these models and a supervised deep model trained on a downstream dataset. To bridge this gap, we propose a novel active learning (AL) framework that enhances the zero-shot classification performance of VLMs by selecting only a few informative samples from the unlabeled data for annotation during training. To achieve this, our approach first calibrates the predicted entropy of VLMs and then utilizes a combination of self-uncertainty and neighbor-aware uncertainty to calculate a reliable uncertainty measure for active sample selection. Our extensive experiments show that the proposed approach outperforms existing AL approaches on several image classification datasets, and significantly enhances the zero-shot performance of VLMs.

📄 PDF Abstract BibTeX arXiv:2410.22187

Code (0)

등록된 구현이 없습니다.

Tasks

Active Learningimage-classificationImage Classificationzero-shot-classificationZero-Shot Learning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation

2026-01-13 · Zhenyang Liu, Yongchong Gu, Yikai Wang, Xiangyang Xue 외 arxiv

Recent advances in robot manipulation have leveraged pre-trained vision-language models (VLMs) and explored integrating 3D spatial signals into these models for effective action prediction, giving rise to the promising v…

Robot Manipulation

An Exam for Active Observers

2026-07-17 · Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu 외 arxiv

Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essentia…

ViThinker: Active Vision-Language Reasoning via Dynamic Perceptual Querying

2026-02-02 · Weihang You, Qingchan Zhu, David Liu, Yi Pan 외 arxiv

Chain-of-Thought (CoT) reasoning excels in language models but struggles in vision-language models due to premature visual-to-text conversion that discards continuous information such as geometry and spatial layout. Whil…

Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration

2026-02-11 · Jinghan He, Junfeng Fang, Feng Xiong, Zijun Yao 외 arxiv

Self-play has enabled large language models to autonomously improve through self-generated challenges. However, existing self-play methods for vision-language models rely on passive interaction with static image collecti…

An interactive enhanced driving dataset for autonomous driving

2026-02-24 · Haojie Feng, Peizhi Zhang, Mengjie Tian, Xinrui Zhang 외 arxiv

The evolution of autonomous driving towards full automation demands robust interactive capabilities; however, the development of Vision-Language-Action (VLA) models is constrained by the sparsity of interactive scenarios…

Autonomous Driving