paper-with-me

홈 › Papers

Enabling Calibration In The Zero-Shot Inference of Large Vision-Language Models

2023-03-11 · Will LeVine, Benjamin Pikus, Pranav Raja, Fernando Amat Gil

Calibration of deep learning models is crucial to their trustworthiness and safe usage, and as such, has been extensively studied in supervised classification models, with methods crafted to decrease miscalibration. However, there has yet to be a comprehensive study of the calibration of vision-language models that are used for zero-shot inference, like CLIP. We measure calibration across relevant variables like prompt, dataset, and architecture, and find that zero-shot inference with CLIP is miscalibrated. Furthermore, we propose a modified version of temperature scaling that is aligned with the common use cases of CLIP as a zero-shot inference model, and show that a single learned temperature generalizes for each specific CLIP model (defined by a chosen pre-training dataset and architecture) across inference dataset and prompt choice.

📄 PDF Abstract BibTeX arXiv:2303.12748

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Task Calibration: Calibrating Large Language Models on Inference Tasks

2024-10-24 · Yingjie Li, Yun Luo, Xiaotian Xie, Yue Zhang

Large language models (LLMs) have exhibited impressive zero-shot performance on inference tasks. However, LLMs may suffer from spurious correlations between input texts and output labels, which limits LLMs' ability to re…

Natural Language Understanding

Calibrate to Discriminate: Improve In-Context Learning with Label-Free Comparative Inference

2024-10-03 · Wei Cheng, Tianlu Wang, Yanmin Ji, Fan Yang 외

While in-context learning with large language models (LLMs) has shown impressive performance, we have discovered a unique miscalibration behavior where both correct and incorrect predictions are assigned the same level o…

In-Context Learning

Does Bigger Mean Better? Comparitive Analysis of CNNs and Biomedical Vision Language Modles in Medical Diagnosis

2025-10-01 · Ran Tong, Jiaqi Liu, Tong Wang, Xin Hu 외 arxiv

The accurate interpretation of chest radiographs using automated methods is a critical task in medical imaging. This paper presents a comparative analysis between a supervised lightweight Convolutional Neural Network (CN…

Pneumonia DetectionMedical Diagnosis

TinyVLM: Zero-Shot Object Detection on Microcontrollers via Vision-Language Distillation with Matryoshka Embeddings

2026-02-24 · Bibin Wilson arxiv

Zero-shot object detection enables recognising novel objects without task-specific training, but current approaches rely on large vision language models (VLMs) like CLIP that require hundreds of megabytes of memory - far…

Zero-Shot Object Detection

Estimating Uncertainty in Multimodal Foundation Models using Public Internet Data

2023-10-15 · Shiladitya Dutta, Hongbo Wei, Lars van der Laan, Ahmed M. Alaa

Foundation models are trained on vast amounts of data at scale using self-supervised learning, enabling adaptation to a wide range of downstream tasks. At test time, these models exhibit zero-shot capabilities through wh…

Conformal PredictionPredictionSelf-Supervised Learningzero-shot-classification+1