paper-with-me

홈 › Papers

Design of the topology for contrastive visual-textual alignment

2022-09-05 · Zhun Sun

Cosine similarity is the common choice for measuring the distance between the feature representations in contrastive visual-textual alignment learning. However, empirically a learnable softmax temperature parameter is required when learning on large-scale noisy training data. In this work, we first discuss the role of softmax temperature from the embedding space's topological properties. We argue that the softmax temperature is the key mechanism for contrastive learning on noisy training data. It acts as a scaling factor of the distance range (e.g. [-1, 1] for the cosine similarity), and its learned value indicates the level of noise in the training data. Then, we propose an alternative design of the topology for the embedding alignment. We make use of multiple class tokens in the transformer architecture; then map the feature representations onto an oblique manifold endowed with the negative inner product as the distance function. With this configuration, we largely improve the zero-shot classification performance of baseline CLIP models pre-trained on large-scale datasets by an average of 6.1\%.

📄 PDF Abstract BibTeX arXiv:2209.02127

Code (1)

minogame/clip-mtob 공식 구현 pytorch

Tasks

Contrastive LearningImage-to-Text RetrievalRetrievalText Retrievalzero-shot-classificationZero-Shot Learning

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

CVT-SLR: Contrastive Visual-Textual Transformation for Sign Language Recognition with Variational Alignment

2023-03-10 · CVPR 2023 1 · Jiangbin Zheng, Yile Wang, Cheng Tan, Siyuan Li 외

Sign language recognition (SLR) is a weakly supervised task that annotates sign videos as textual glosses. Recent studies show that insufficient training caused by the lack of large-scale available sign datasets becomes …

cross-modal alignmentSign Language Recognition

$β$-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment

2025-12-14 · Fatimah Zohra, Chen Zhao, Hani Itani, Bernard Ghanem arxiv

CLIP achieves strong zero-shot image-text retrieval by aligning global vision and text representations, yet it falls behind on fine-grained tasks even when fine-tuned on long, detailed captions. In this work, we propose …

Contrastive LearningText Retrieval

Object Attribute Matters in Visual Question Answering

2023-12-20 · Peize Li, Qingyi Si, Peng Fu, Zheng Lin 외

Visual question answering is a multimodal task that requires the joint comprehension of visual and textual information. However, integrating visual and textual semantics solely through attention layers is insufficient to…

AttributeGraph Neural NetworkKnowledge DistillationObject+5

Text-driven 3D Human Generation via Contrastive Preference Optimization

2025-02-13 · Pengfei Zhou, Xukun Shen, Yong Hu

Recent advances in Score Distillation Sampling (SDS) have improved 3D human generation from textual descriptions. However, existing methods still face challenges in accurately aligning 3D models with long and complex tex…

Negation

Visual-Semantic Contrastive Alignment for Few-Shot Image Classification

2022-10-20 · Mohamed Afham, Ranga Rodrigo

Few-Shot learning aims to train and optimize a model that can adapt to unseen visual classes with only a few labeled examples. The existing few-shot learning (FSL) methods, heavily rely only on visual data, thus fail to …

ClassificationContrastive LearningFew-Shot Image ClassificationFew-Shot Learning+2