paper-with-me

EsViT

2000년 도입 · 논문 1편에서 사용

EsViT proposes two techniques for developing efficient self-supervised vision transformers for visual representation leaning: a multi-stage architecture with sparse self-attention and a new pre-training task of region matching. The multi-stage architecture reduces modeling complexity but with a cost of losing the ability to capture fine-grained correspondences between image regions. The new pretraining task allows the model to capture fine-grained region dependencies and as a result significantly improves the quality of the learned vision representations.

출처: Efficient Self-supervised Vision Transformers for Representation Learning

소개 논문: Efficient Self-supervised Vision Transformers for Representation Learning

Vision Transformers · Computer Vision