UNIMO
2000년 도입 · 논문 4편에서 사용
UNIMO is a multi-modal pre-training architecture that can effectively adapt to both single modal and multimodal understanding and generation tasks. UNIMO learns visual representations and textual representations simultaneously, and unifies them into the same semantic space via cross-modal contrastive learning (CMCL) based on a large-scale corpus of image collections, text corpus and image-text pairs. The CMCL aligns the visual representation and textual representation, and unifies them into the same semantic space based on image-text pairs.
출처: UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning
소개 논문: UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning
Vision and Language Pre-Trained Models · Computer VisionMulti-Modal Methods · Computer Vision