paper-with-me

홈 › Papers

Probing Cross-modal Semantics Alignment Capability from the Textual Perspective

2022-10-18 · Zheng Ma, Shi Zong, Mianzhi Pan, Jianbing Zhang, ShuJian Huang, Xinyu Dai, Jiajun Chen

In recent years, vision and language pre-training (VLP) models have advanced the state-of-the-art results in a variety of cross-modal downstream tasks. Aligning cross-modal semantics is claimed to be one of the essential capabilities of VLP models. However, it still remains unclear about the inner working mechanism of alignment in VLP models. In this paper, we propose a new probing method that is based on image captioning to first empirically study the cross-modal semantics alignment of VLP models. Our probing method is built upon the fact that given an image-caption pair, the VLP models will give a score, indicating how well two modalities are aligned; maximizing such scores will generate sentences that VLP models believe are of good alignment. Analyzing these sentences thus will reveal in what way different modalities are aligned and how well these alignments are in VLP models. We apply our probing method to five popular VLP models, including UNITER, ROSITA, ViLBERT, CLIP, and LXMERT, and provide a comprehensive analysis of the generated captions guided by these models. Our results show that VLP models (1) focus more on just aligning objects with visual words, while neglecting global semantics; (2) prefer fixed sentence patterns, thus ignoring more important textual information including fluency and grammar; and (3) deem the captions with more visual words are better aligned with images. These findings indicate that VLP models still have weaknesses in cross-modal semantics alignment and we hope this work will draw researchers' attention to such problems when designing a new VLP model.

📄 PDF Abstract BibTeX arXiv:2210.09550

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningSentence

Methods 이 논문이 사용한 방법론

UNITER UNITER or UNiversal Image-TExt Representation model is a large-scale pre-trained model for joint multimodal embedding. It is pre-trained using four image-text datasets COCO,…
ViLBERT Vision-and-Language BERT (ViLBERT) is a BERT-based model for learning task-agnostic joint representations of image content and…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
LXMERT LXMERT is a model for learning vision-and-language cross-modality representations. It consists of a Transformer model that consists three encoders: object relationship encoder, a…

Similar Papers 제목 키워드 기반

ModalChorus: Visual Probing and Alignment of Multi-modal Embeddings via Modal Fusion Map

2024-07-17 · Yilin Ye, Shishi Xiao, Xingchen Zeng, Wei Zeng

Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal f…

Cross-Modal RetrievalDimensionality Reductionzero-shot-classificationZero-Shot Learning

Universal Scene Graph Generation

2025-03-19 · CVPR 2025 1 · Shengqiong Wu, Hao Fei, Tat-Seng Chua

Scene graph (SG) representations can neatly and efficiently describe scene semantics, which has driven sustained intensive research in SG generation. In the real world, multiple modalities often coexist, with different t…

Graph GenerationScene Graph Generation

Can Linguistic Knowledge Improve Multimodal Alignment in Vision-Language Pretraining?

2023-08-24 · Fei Wang, Liang Ding, Jun Rao, Ye Liu 외

The multimedia community has shown a significant interest in perceiving and representing the physical world with multimodal pretrained neural network models, and among them, the visual-language pertaining (VLP) is, curre…

AttributeNegationSentence

Align3D-AD: Cross-Modal Feature Alignment and Dual-Prompt Learning for Zero-shot 3D Anomaly Detection

2026-05-07 · Letian Bai, Xuanming Cao, Juan Du, Chengyu Tao arxiv

Zero-shot 3D anomaly detection aims to identify anomalies without access to training data from target categories. However, existing methods mainly rely on projecting 3D observations into multi-view representations that p…

3D Anomaly Detection

The Attention Triangle in Audio-Video Models

2026-09-03 · Sagi Polaczek, Noa Kraicer, Gal Metzer, Zhuo Ning 외 hf

Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and…

Audio GenerationVideo Generation