paper-with-me

홈 › Papers

Parts of Speech–Grounded Subspaces in Vision-Language Models

2023-09-21 · NeurIPS 2023 11

Latent image representations arising from vision-language models have proved immensely useful for a variety of downstream tasks. However, their utility is limited by their entanglement with respect to different visual attributes. For instance, recent work has shown that CLIP image representations are often biased toward specific visual properties (such as objects or actions) in an unpredictable manner. In this paper, we propose to separate representations of the different visual modalities in CLIP’s joint vision-language space by leveraging the association between parts of speech and specific visual modes of variation (e.g. nouns relate to objects, adjectives describe appearance). This is achieved by formulating an appropriate component analysis model that learns subspaces capturing variability corresponding to a specific part of speech, while jointly minimising variability to the rest. Such a subspace yields disentangled representations of the different visual properties of an image or text in closed form while respecting the underlying geometry of the manifold on which the representations lie. What’s more, we show the proposed model additionally facilitates learning subspaces corresponding to specific visual appearances (e.g. artists’ painting styles), which enables the selective removal of entire visual themes from CLIP-based text-to-image synthesis. We validate the model both qualitatively, by visualising the subspace projections with a text-to-image model and by preventing the imitation of artists’ styles, and quantitatively, through class invariance metrics and improvements to baseline zero-shot classification.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Parts of Speech-Grounded Subspaces in Vision-Language Models

2023-05-23 · James Oldfield, Christos Tzelepis, Yannis Panagakis, Mihalis A. Nicolaou 외

Latent image representations arising from vision-language models have proved immensely useful for a variety of downstream tasks. However, their utility is limited by their entanglement with respect to different visual at…

Image GenerationPOSzero-shot-classificationZero-Shot Learning

Hindi as a Second Language: Improving Visually Grounded Speech with Semantically Similar Samples

2023-03-30 · Hyeonggon Ryu, Arda Senocak, In So Kweon, Joon Son Chung

The objective of this work is to explore the learning of visually grounded speech models (VGS) from multilingual perspective. Bilingual VGS models are generally trained with an equal number of spoken captions from both l…

Cross-Modal RetrievalRetrieval

DisCGen: A Framework for Discourse-Informed Counterspeech Generation

2023-11-29 · Sabit Hassan, Malihe Alikhani

Counterspeech can be an effective method for battling hateful content on social media. Automated counterspeech generation can aid in this process. Generated counterspeech, however, can be viable only when grounded in the…

Connecting Speech to Words through Images

2026-06-15 · Gabriel Pirlogeanu, Dan Oneata, Horia Cucu, Herman Kamper arxiv

How can we learn the mapping between written words and their spoken counterparts in the absence of explicit textual supervision? We present a visually grounded method for building a vocabulary of spoken words using only …

Image CaptioningKeyword Spotting

XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models

2026-06-15 · Yupei Li, Qiyang Sun, Xiaoliang Wu, Chenxi Wang 외 arxiv

Speech deepfake detection (SDD) systems require trustworthy explanations for reliable decision-making. Existing explanation ways mainly fall into two categories. Traditional explainable AI (XAI), such as gradient-based a…

Explanation GenerationDeepFake Detection