paper-with-me

홈 › Papers

PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining

2022-04-29 · Yuting Gao, Jinfeng Liu, Zihan Xu, Jun Zhang, Ke Li, Rongrong Ji, Chunhua Shen

Large-scale vision-language pre-training has achieved promising results on downstream tasks. Existing methods highly rely on the assumption that the image-text pairs crawled from the Internet are in perfect one-to-one correspondence. However, in real scenarios, this assumption can be difficult to hold: the text description, obtained by crawling the affiliated metadata of the image, often suffers from the semantic mismatch and the mutual compatibility. To address these issues, we introduce PyramidCLIP, which constructs an input pyramid with different semantic levels for each modality, and aligns visual elements and linguistic elements in the form of hierarchy via peer-level semantics alignment and cross-level relation alignment. Furthermore, we soften the loss of negative samples (unpaired samples) so as to weaken the strict constraint during the pre-training stage, thus mitigating the risk of forcing the model to distinguish compatible negative pairs. Experiments on five downstream tasks demonstrate the effectiveness of the proposed PyramidCLIP. In particular, with the same amount of 15 million pre-training image-text pairs, PyramidCLIP exceeds CLIP on ImageNet zero-shot classification top-1 accuracy by 10.6%/13.2%/10.0% with ResNet50/ViT-B32/ViT-B16 based image encoder respectively. When scaling to larger datasets, PyramidCLIP achieves the state-of-the-art results on several downstream tasks. In particular, the results of PyramidCLIP-ResNet50 trained on 143M image-text pairs surpass that of CLIP using 400M data on ImageNet zero-shot classification task, significantly improving the data efficiency of CLIP.

📄 PDF Abstract BibTeX arXiv:2204.14095

Code (0)

등록된 구현이 없습니다.

Tasks

Image ClassificationLanguage ModelingLanguage ModellingObject Detectionzero-shot-classificationZero-Shot Image ClassificationZero-Shot Learning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Hierarchical Pre-Training of Vision Encoders with Large Language Models

2026-03-31 · Eugene Lee, Ting-Yu Chang, Jui-Huang Tsai, Jiajie Diao 외 arxiv

The field of computer vision has experienced significant advancements through scalable vision encoders and multimodal pre-training frameworks. However, existing approaches often treat vision encoders and large language m…

Representation LearningImage Classification

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models

2026-01-28 · Kun Wang, Xiao Feng, Mingcheng Qu, Tonghua Su arxiv

Vision Language Action (VLA) models have recently shown great potential in bridging multimodal perception with robotic control. However, existing methods often rely on direct fine-tuning of pre-trained Vision-Language Mo…

Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models

2025-01-01 · Emily Johnson, Noah Wilson

Text-to-image generation has witnessed significant advancements with the integration of Large Vision-Language Models (LVLMs), yet challenges remain in aligning complex textual descriptions with high-quality, visually coh…

Image GenerationText to Image GenerationText-to-Image Generation

HiLa: Hierarchical Vision-Language Collaboration for Cancer Survival Prediction

2025-07-07 · Jiaqi Cui, Lu Wen, Yuchen Fei, Bo Liu 외 arxiv

Survival prediction using whole-slide images (WSIs) is crucial in cancer re-search. Despite notable success, existing approaches are limited by their reliance on sparse slide-level labels, which hinders the learning of d…

Contrastive Learning

IMITATE: Clinical Prior Guided Hierarchical Vision-Language Pre-training

2023-10-11 · Che Liu, Sibo Cheng, Miaojing Shi, Anand Shah 외

In the field of medical Vision-Language Pre-training (VLP), significant efforts have been devoted to deriving text and image features from both clinical reports and associated medical images. However, most existing metho…

Contrastive LearningDescriptive