Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
Despite the impressive advancements achieved through vision-and-language pretraining, it remains unclear whether this joint learning paradigm can help understand each individual modality. In this work, we conduct a comparative analysis of the visual representations in existing vision-and-language models and vision-only models by probing a broad range of tasks, aiming to assess the quality of the learned representations in a nuanced manner. Interestingly, our empirical observations suggest that vision-and-language models are better at label prediction tasks like object and attribute prediction, while vision-only models are stronger at dense prediction tasks that require more localized information. We hope our study sheds light on the role of language in visual learning, and serves as an empirical guide for various pretrained models. Code will be released at https://github.com/Lizw14/visual_probing
Code (0)
등록된 구현이 없습니다.
Tasks
AttributePredictionRepresentation LearningSimilar Papers 제목 키워드 기반
TCT: A Cross-supervised Learning Method for Multimodal Sequence Representation
Multimodalities provide promising performance than unimodality in most tasks. However, learning the semantic of the representations from multimodalities efficiently is extremely challenging. To tackle this, we propose th…
Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization
Dense-localization Audio-Visual Events (DAVE) aims to identify time boundaries and corresponding categories for events that are both audible and visible in a long video, where events may co-occur and exhibit varying dura…
cross-modal alignmentBrain-Inspired Multimodal Spiking Neural Network for Image-Text Retrieval
Spiking neural networks (SNNs) have recently shown strong potential in unimodal visual and textual tasks, yet building a directly trained, low-energy, and high-performance SNN for multimodal applications such as image-te…
Text RetrievalAdversarial Multimodal Domain Transfer for Video-Level Sentiment Analysis
Video-level sentiment analysis is a challenging task and requires systems to obtain discriminative multimodal representations that can capture difference in sentiments across various modalities. However, due to diverse …
Multimodal Sentiment AnalysisSentiment AnalysisSee & Sniff: Learning Visuo-Olfactory Representations
While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olfactory data. We introduce SmellNet-V, a scalable visuo-olfactory dat…
Cross-Modal Retrieval