paper-with-me

홈 › Papers

VL-Taboo: An Analysis of Attribute-based Zero-shot Capabilities of Vision-Language Models

2022-09-12 · Felix Vogel, Nina Shvetsova, Leonid Karlinsky, Hilde Kuehne

Vision-language models trained on large, randomly collected data had significant impact in many areas since they appeared. But as they show great performance in various fields, such as image-text-retrieval, their inner workings are still not fully understood. The current work analyses the true zero-shot capabilities of those models. We start from the analysis of the training corpus assessing to what extent (and which of) the test classes are really zero-shot and how this correlates with individual classes performance. We follow up with the analysis of the attribute-based zero-shot learning capabilities of these models, evaluating how well this classical zero-shot notion emerges from large-scale webly supervision. We leverage the recently released LAION400M data corpus as well as the publicly available pretrained models of CLIP, OpenCLIP, and FLAVA, evaluating the attribute-based zero-shot capabilities on CUB and AWA2 benchmarks. Our analysis shows that: (i) most of the classes in popular zero-shot benchmarks are observed (a lot) during pre-training; (ii) zero-shot performance mainly comes out of models' capability of recognizing class labels, whenever they are present in the text, and a significantly lower performing capability of attribute-based zeroshot learning is only observed when class labels are not used; (iii) the number of the attributes used can have a significant effect on performance, and can easily cause a significant performance decrease.

📄 PDF Abstract BibTeX arXiv:2209.06103

Code (1)

felixvogel02/vl-taboo 공식 구현 pytorch

Tasks

AttributeImage-text RetrievalRetrievalText RetrievalZero-Shot Learning

Methods 이 논문이 사용한 방법론

Test 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
FLAVA FLAVA aims at building a single holistic universal model that targets all modalities at once. FLAVA is a language vision alignment model that learns strong representations from…

Similar Papers 제목 키워드 기반

A Zero-Shot Classification Approach for a Word-Guessing Challenge

2022-06-27 · Nicos Isaak

The Taboo Challenge competition, a task based on the well-known Taboo game, has been proposed to stimulate research in the AI field. The challenge requires building systems able to comprehend the implied inferences betwe…

ClassificationLanguage ModelingLanguage Modellingzero-shot-classification+1

What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content

2025-07-31 · Alfio Ferrara, Sergio Picascia, Laura Pinnavaia, Vojimir Ranitovic 외 arxiv

Proprietary Large Language Models (LLMs) have shown tendencies toward politeness, formality, and implicit content moderation. While previous research has primarily focused on explicitly training models to moderate and de…

The Taboo Trap: Behavioural Detection of Adversarial Samples

2018-11-18 · Ilia Shumailov, Yiren Zhao, Robert Mullins, Ross Anderson

Deep Neural Networks (DNNs) have become a powerful toolfor a wide range of problems. Yet recent work has found an increasing variety of adversarial samplesthat can fool them. Most existing detection mechanisms against ad…

Exposing the limits of Zero-shot Cross-lingual Hate Speech Detection

2021-08-01 · ACL 2021 5 · Debora Nozza

Reducing and counter-acting hate speech on Social Media is a significant concern. Most of the proposed automatic methods are conducted exclusively on English and very few consistently labeled, non-English resources have …

Cross-Lingual TransferHate Speech DetectionTransfer LearningZero-Shot Cross-Lingual Transfer

Tight Lower Bounds on Worst-Case Guarantees for Zero-Shot Learning with Attributes

2022-05-25 · Alessio Mazzetto, Cristina Menghini, Andrew Yuan, Eli Upfal 외

We develop a rigorous mathematical analysis of zero-shot learning with attributes. In this setting, the goal is to label novel classes with no training data, only detectors for attributes and a description of how those a…

AttributeZero-Shot Learning