paper-with-me

OneR

One Representation

2000년 도입 · 논문 2편에서 사용

In the OneR method, model input can be one of image, text or image+text, and CMC objective is combined with the traditional image-text contrastive (ITC) loss. Masked modeling is also carried out for all three input types (i.e., image, text and multi-modal). This framework employs no modality-specific architectural component except for the initial token embedding layer, making our model generic and modality-agnostic with minimal inductive bias.

출처: Unifying Vision-Language Representation Space with Single-tower Transformer

소개 논문: Unifying Vision-Language Representation Space with Single-tower Transformer

Vision and Language Pre-Trained Models · Computer Vision