VL-T5
2000년 도입 · 논문 5편에서 사용
VL-T5 is a unified framework that learns different tasks in a single architecture with the same language modeling objective, i.e., multimodal conditional text generation. The model learns to generate labels in text based on the visual and textual inputs. In contrast to other existing methods, the framework unifies tasks as generating text labels conditioned on multimodal inputs. This allows the model to tackle vision-and-language tasks with unified text generation objective. The models use text prefixes to adapt to different tasks.
출처: Unifying Vision-and-Language Tasks via Text Generation
소개 논문: Unifying Vision-and-Language Tasks via Text Generation
Vision and Language Pre-Trained Models · Computer Vision