paper-with-me

홈 › Papers

Multimodal Pretraining Unmasked: A Meta-Analysis and a Unified Framework of Vision-and-Language BERTs

2020-11-30 · Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, Desmond Elliott

Large-scale pretraining and task-specific fine-tuning is now the standard methodology for many tasks in computer vision and natural language processing. Recently, a multitude of methods have been proposed for pretraining vision and language BERTs to tackle challenges at the intersection of these two key areas of AI. These models can be categorised into either single-stream or dual-stream encoders. We study the differences between these two categories, and show how they can be unified under a single theoretical framework. We then conduct controlled experiments to discern the empirical differences between five V&L BERTs. Our experiments show that training data and hyperparameters are responsible for most of the differences between the reported results, but they also reveal that the embedding layer plays a crucial role in these massive models.

📄 PDF Abstract BibTeX arXiv:2011.15124

Code (3)

e-bug/mpre-unmasked 공식 구현 pytorch
e-bug/volta 공식 구현 pytorch
ivana-13/guided_masking pytorch

Similar Papers 제목 키워드 기반

Beyond Language Modeling: An Exploration of Multimodal Pretraining

2026-03-03 · Shengbang Tong, David Fan, John Nguyen, Ellis Brown 외 arxiv

The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space for native multimodal models remains opaque. We provide empirical clar…

Empowering Time Series Analysis with Large-Scale Multimodal Pretraining

2026-02-05 · Peng Chen, Siyuan Wang, Shiyan Hu, Xingjian Wu 외 arxiv

While existing time series foundation models primarily rely on large-scale unimodal pretraining, they lack complementary modalities to enhance time series understanding. Building multimodal foundation models is a natural…

Time Series ForecastingTime Series AnalysisAnomaly Detection

Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis

2022-04-17 · ACL 2022 5 · Yan Ling, Jianfei Yu, Rui Xia

As an important task in sentiment analysis, Multimodal Aspect-Based Sentiment Analysis (MABSA) has attracted increasing attention in recent years. However, previous approaches either (i) use separately pre-trained visual…

Aspect-Based Sentiment AnalysisAspect-Based Sentiment Analysis (ABSA)DecoderSentiment Analysis

Emerging Properties in Unified Multimodal Pretraining

2025-05-20 · Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou 외

Unifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems. In this work, we introduce BAGEL, an open0source foundational model that natively supports multimoda…

Image EditingImage GenerationImage Manipulation+2

Bootstrapped Masked Autoencoders for Vision BERT Pretraining

2022-07-14 · Xiaoyi Dong, Jianmin Bao, Ting Zhang, Dongdong Chen 외

We propose bootstrapped masked autoencoders (BootMAE), a new approach for vision BERT pretraining. BootMAE improves the original masked autoencoders (MAE) with two core designs: 1) momentum encoder that provides online f…

DecoderObject DetectionPredictionSelf-Supervised Image Classification+1