paper-with-me

홈 › Papers

UniCLIP: Unified Framework for Contrastive Language-Image Pre-training

2022-09-27 · Janghyeon Lee, Jongsuk Kim, Hyounguk Shon, Bumsoo Kim, Seung Hwan Kim, Honglak Lee, Junmo Kim

Pre-training vision-language models with contrastive objectives has shown promising results that are both scalable to large uncurated datasets and transferable to many downstream applications. Some following works have targeted to improve data efficiency by adding self-supervision terms, but inter-domain (image-text) contrastive loss and intra-domain (image-image) contrastive loss are defined on individual spaces in those works, so many feasible combinations of supervision are overlooked. To overcome this issue, we propose UniCLIP, a Unified framework for Contrastive Language-Image Pre-training. UniCLIP integrates the contrastive loss of both inter-domain pairs and intra-domain pairs into a single universal space. The discrepancies that occur when integrating contrastive loss between different domains are resolved by the three key components of UniCLIP: (1) augmentation-aware feature embedding, (2) MP-NCE loss, and (3) domain dependent similarity measure. UniCLIP outperforms previous vision-language pre-training methods on various single- and multi-modality downstream tasks. In our experiments, we show that each component that comprises UniCLIP contributes well to the final performance.

📄 PDF Abstract BibTeX arXiv:2209.13430

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios

2024-01-19 · Xiangshuo Qiao, Xianxin Li, Xiaozhe Qu, Jie Zhang 외

Vision-Language Models pre-trained on large-scale image-text datasets have shown superior performance in downstream tasks such as image retrieval. Most of the images for pre-training are presented in the form of open dom…

Common Sense ReasoningImage Retrieval

Unified Medical Image-Text-Label Contrastive Learning With Continuous Prompt

2023-07-12 · Yuhao Wang

Contrastive language-image Pre-training (CLIP) [13] can leverage large datasets of unlabeled Image-Text pairs, which have demonstrated impressive performance in various downstream tasks. Given that annotating medical dat…

Contrastive Learning

Toward Unified Multimodal Representation Learning for Autonomous Driving

2026-03-09 · Ximeng Tao, Dimitar Filev, Gaurav Pandey arxiv

Contrastive Language-Image Pre-training (CLIP) has shown impressive performance in aligning visual and textual representations. Recent studies have extended this paradigm to 3D vision to improve scene understanding for a…

Representation LearningContrastive LearningScene UnderstandingAutonomous Driving

Exploring a Unified Vision-Centric Contrastive Alternatives on Multi-Modal Web Documents

2025-10-21 · Yiqi Lin, Alex Jinpeng Wang, Linjie Li, Zhengyuan Yang 외 arxiv

Contrastive vision-language models such as CLIP have demonstrated strong performance across a wide range of multimodal tasks by learning from aligned image-text pairs. However, their ability to handle complex, real-world…

Representation LearningCross-Modal RetrievalContrastive Learning

CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment

2025-08-08 · Shengzhu Yang, Jiawei Du, Shuai Lu, Weihang Zhang 외 arxiv

Large-scale natural image-text datasets, especially those automatically collected from the web, often suffer from loose semantic alignment due to weak supervision, while medical datasets tend to have high cross-modal cor…

Contrastive Learning