paper-with-me

Papers

ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

2020-01-22 · Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, Arun Sacheti

In this paper, we introduce a new vision-language pre-trained model -- ImageBERT -- for image-text joint embedding. Our model is a Transformer-based model, which takes different modalities as input and models the relationship between them. The model is pre-trained on four tasks simultaneously: Masked Language Modeling (MLM), Masked Object Classification (MOC), Masked Region Feature Regression (MRFR), and Image Text Matching (ITM). To further enhance the pre-training quality, we have collected a Large-scale weAk-supervised Image-Text (LAIT) dataset from Web. We first pre-train the model on this dataset, then conduct a second stage pre-training on Conceptual Captions and SBU Captions. Our experiments show that multi-stage pre-training strategy outperforms single-stage pre-training. We also fine-tune and evaluate our pre-trained ImageBERT model on image retrieval and text retrieval tasks, and achieve new state-of-the-art results on both MSCOCO and Flickr30k datasets.

📄 PDF Abstract BibTeX arXiv:2001.07966

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalImage-text matchingLanguage ModelingLanguage ModellingMasked Language ModelingRetrievalText MatchingText RetrievalZero-Shot Cross-Modal Retrieval

Similar Papers 제목 키워드 기반

MHTN: Modal-adversarial Hybrid Transfer Network for Cross-modal Retrieval

2017-08-08 · Xin Huang, Yuxin Peng, Mingkuan Yuan

Cross-modal retrieval has drawn wide interest for retrieval across different modalities of data. However, existing methods based on DNN face the challenge of insufficient cross-modal training data, which limits the train…

Cross-Modal RetrievalRepresentation LearningRetrievalTransfer Learning

Empowering Time Series Analysis with Large-Scale Multimodal Pretraining

2026-02-05 · Peng Chen, Siyuan Wang, Shiyan Hu, Xingjian Wu 외 arxiv

While existing time series foundation models primarily rely on large-scale unimodal pretraining, they lack complementary modalities to enhance time series understanding. Building multimodal foundation models is a natural…

Time Series ForecastingTime Series AnalysisAnomaly Detection

WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training

2021-03-11 · Yuqi Huo, Manli Zhang, Guangzhen Liu, Haoyu Lu 외

Multi-modal pre-training models have been intensively explored to bridge vision and language in recent years. However, most of them explicitly model the cross-modal interaction between image-text pairs, by assuming that …

Contrastive LearningGPUImage CaptioningImage Retrieval+1

CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models

2025-03-09 · Wei Dai, Peilin Chen, Malinda Lu, Daniel Li 외

Recent advances in clinical AI have enabled remarkable progress across many clinical domains. However, existing benchmarks and models are primarily limited to a small set of modalities and tasks, which hinders the develo…

DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset

2026-01-15 · Hengyu Shen, Tiancheng Gu, Bin Qin, Lan Wu 외 arxiv

Vision-Language Pre-training (VLP) models have achieved remarkable success by leveraging large-scale image-text pairs. While English-centric models like CLIP and SigLIP benefit from massive datasets (e.g., LAION-400M), t…

Cross-Modal Retrieval