ViLBERT
Vision-and-Language BERT
2000년 도입 · 논문 30편에서 사용
Vision-and-Language BERT (ViLBERT) is a BERT-based model for learning task-agnostic joint representations of image content and natural language. ViLBERT extend the popular BERT architecture to a multi-modal two-stream model, processing both visual and textual inputs in separate streams that interact through co-attentional transformer layers.
출처: ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
소개 논문: ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Vision and Language Pre-Trained Models · Computer VisionTransformers · Natural Language ProcessingRepresentation Learning · General