DALL·E 2
2000년 도입 · 논문 2편에서 사용
DALL·E 2 is a generative text-to-image model made up of two main components: a prior that generates a CLIP image embedding given a text caption, and a decoder that generates an image conditioned on the image embedding.
출처: Hierarchical Text-Conditional Image Generation with CLIP Latents
소개 논문: Hierarchical Text-Conditional Image Generation with CLIP Latents
Image Generation Models · Computer Vision