paper-with-me

홈 › Papers

BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

2022-08-12 · Zhiliang Peng, Li Dong, Hangbo Bao, Qixiang Ye, Furu Wei

Masked image modeling (MIM) has demonstrated impressive results in self-supervised representation learning by recovering corrupted image patches. However, most existing studies operate on low-level image pixels, which hinders the exploitation of high-level semantics for representation models. In this work, we propose to use a semantic-rich visual tokenizer as the reconstruction target for masked prediction, providing a systematic way to promote MIM from pixel-level to semantic-level. Specifically, we propose vector-quantized knowledge distillation to train the tokenizer, which discretizes a continuous semantic space to compact codes. We then pretrain vision Transformers by predicting the original visual tokens for the masked image patches. Furthermore, we introduce a patch aggregation strategy which associates discrete image patches to enhance global semantic representation. Experiments on image classification and semantic segmentation show that BEiT v2 outperforms all compared MIM methods. On ImageNet-1K (224 size), the base-size BEiT v2 achieves 85.5% top-1 accuracy for fine-tuning and 80.1% top-1 accuracy for linear probing. The large-size BEiT v2 obtains 87.3% top-1 accuracy for ImageNet-1K (224 size) fine-tuning, and 56.7% mIoU on ADE20K for semantic segmentation. The code and pretrained models are available at https://aka.ms/beitv2.

📄 PDF Abstract BibTeX arXiv:2208.06366

Code (3)

microsoft/unilm/tree/master/beit2 공식 구현 pytorch
MindSpore-scientific/code-7/tree/main/Beit
leondgarse/keras_cv_attention_models/tree/main/keras_cv_attention_models/beit tf

Tasks

image-classificationImage ClassificationKnowledge DistillationRepresentation LearningSelf-Supervised Image ClassificationSemantic Segmentation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Residual Connection 설명 없음
Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Multi-Head Attention 설명 없음
Vision Transformer The Vision Transformer, or ViT, is a model for image classification that employs a Transformer-like architecture over…

Similar Papers 제목 키워드 기반

VL-BEiT: Generative Vision-Language Pretraining

2022-06-02 · Hangbo Bao, Wenhui Wang, Li Dong, Furu Wei

We introduce a vision-language foundation model called VL-BEiT, which is a bidirectional multimodal Transformer learned by generative pretraining. Our minimalist solution conducts masked prediction on both monomodal and …

image-classificationImage ClassificationImage-text RetrievalLanguage Modeling+9

mc-BEiT: Multi-choice Discretization for Image BERT Pre-training

2022-03-29 · Xiaotong Li, Yixiao Ge, Kun Yi, Zixuan Hu 외

Image BERT pre-training with masked image modeling (MIM) becomes a popular practice to cope with self-supervised representation learning. A seminal work, BEiT, casts MIM as a classification task with a visual vocabulary,…

Instance Segmentationobject-detectionObject DetectionRepresentation Learning+3

E-ViLM: Efficient Video-Language Model via Masked Video Modeling with Semantic Vector-Quantized Tokenizer

2023-11-28 · Jacob Zhiyuan Fang, Skyler Zheng, Vasu Sharma, Robinson Piramuthu

To build scalable models for challenging real-world tasks, it is important to learn from diverse, multi-modal data in various forms (e.g., videos, text, and images). Among the existing works, a plethora of them have focu…

Language ModelingLanguage ModellingQuestion AnsweringText to Video Retrieval+2

BandVQ: Band-Wise Vector-Quantized EEG Foundation Model

2026-05-24 · Jamiyan Sukhbaatar, Satoshi Imamura, Toshihisa Tanaka arxiv

A central challenge in electroencephalography (EEG) foundation modeling is learning transferable representations across recordings with diverse tasks, montages, references, and spectral characteristics. Existing masked m…

Image Compression with Product Quantized Masked Image Modeling

2022-12-14 · Alaaeldin El-Nouby, Matthew J. Muckley, Karen Ullrich, Ivan Laptev 외

Recent neural compression methods have been based on the popular hyperprior framework. It relies on Scalar Quantization and offers a very strong compression performance. This contrasts from recent advances in image gener…

Image CompressionImage GenerationQuantizationRepresentation Learning+1