VirTex
2000년 도입 · 논문 6편에서 사용
VirText, or Visual representations from Textual annotations is a pretraining approach using semantically dense captions to learn visual representations. First a ConvNet and Transformer are jointly trained from scratch to generate natural language captions for images. Then, the learned features are transferred to downstream visual recognition tasks.
출처: VirTex: Learning Visual Representations from Textual Annotations
소개 논문: VirTex: Learning Visual Representations from Textual Annotations
Image Representations · Computer Vision