paper-with-me

Papers

Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models

2022-12-01 · Zhuowan Li, Cihang Xie, Benjamin Van Durme, Alan Yuille

Despite the impressive advancements achieved through vision-and-language pretraining, it remains unclear whether this joint learning paradigm can help understand each individual modality. In this work, we conduct a comparative analysis of the visual representations in existing vision-and-language models and vision-only models by probing a broad range of tasks, aiming to assess the quality of the learned representations in a nuanced manner. Interestingly, our empirical observations suggest that vision-and-language models are better at label prediction tasks like object and attribute prediction, while vision-only models are stronger at dense prediction tasks that require more localized information. We hope our study sheds light on the role of language in visual learning, and serves as an empirical guide for various pretrained models. Code will be released at https://github.com/Lizw14/visual_probing

📄 PDF Abstract BibTeX arXiv:2212.00281

Code (0)

등록된 구현이 없습니다.

Tasks

AttributePredictionRepresentation Learning

Similar Papers 제목 키워드 기반

TCT: A Cross-supervised Learning Method for Multimodal Sequence Representation

2019-10-23 · Wubo Li, Wei Zou, Xiangang Li

Multimodalities provide promising performance than unimodality in most tasks. However, learning the semantic of the representations from multimodalities efficiently is extremely challenging. To tackle this, we propose th…

Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization

2024-09-12 · Ling Xing, Hongyu Qu, Rui Yan, Xiangbo Shu 외

Dense-localization Audio-Visual Events (DAVE) aims to identify time boundaries and corresponding categories for events that are both audible and visible in a long video, where events may co-occur and exhibit varying dura…

cross-modal alignment

Brain-Inspired Multimodal Spiking Neural Network for Image-Text Retrieval

2026-03-25 · Xintao Zong, Xian Zhong, Wenxuan Liu, Jianhao Ding 외 arxiv

Spiking neural networks (SNNs) have recently shown strong potential in unimodal visual and textual tasks, yet building a directly trained, low-energy, and high-performance SNN for multimodal applications such as image-te…

Text Retrieval

Adversarial Multimodal Domain Transfer for Video-Level Sentiment Analysis

2022-05-11 · IEEE Access 2022 5 · Wang Yanan; Wu Jianming; Furumai Kazuaki; Wada Shinya; Kurihara Satoshi

Video-level sentiment analysis is a challenging task and requires systems to obtain discriminative multimodal representations that can capture difference in sentiments across various modalities. However, due to diverse …

Multimodal Sentiment AnalysisSentiment Analysis

See & Sniff: Learning Visuo-Olfactory Representations

2026-06-25 · Seongyu Kim, Seungwoo Lee, Hyeonggon Ryu, Joon Son Chung 외 arxiv

While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olfactory data. We introduce SmellNet-V, a scalable visuo-olfactory dat…

Cross-Modal Retrieval