Text-to-Image Generation
17개 벤치마크 · 논문 1,553편 · 이 태스크의 논문 보기 →
Benchmarks
COCO (Common Objects in Context)
GenEval
CUB
Multi-Modal-CelebA-HQ
DrawBench
Oxford 102 Flowers
Conceptual Captions
COCO
LHQC
MS-COCO
DPG
GeNeVA (CoDraw)
GeNeVA (i-CLEVR)
LAION COCO
T2I-CompBench
Colors
Flickr-8k
Most implemented
Show and Tell: A Neural Image Caption Generator
High-Resolution Image Synthesis with Latent Diffusion Models
Generative Adversarial Text to Image Synthesis
StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks
Papers
Importance-Aware Low-Rank Distillation of Diffusion Transformers
Diffusion Transformers (DiTs) have emerged as a dominant architecture for high-quality text-to-image generation, yet their scale poses challenges for efficient deployment. While truncated singular value decomposition (SV…
Text-to-Image GenerationKnowledge DistillationImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation
We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question answering and image-grounded hallucination …
Visual Question AnsweringText-to-Image GenerationGenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling
Latent generative models typically follow a two-stage pipeline, training a variational autoencoder for reconstruction and then a generative model on the frozen latent space. Since reconstruction-optimized latents are not…
Text-to-Image GenerationRepresentation LearningAbstract4D: A Large-Scale Dataset and Framework for Understanding the Visual Language of Abstract Art
Artificial intelligence can classify artistic styles and synthesize images, but it still lacks a model of the visual language that gives art meaning. Abstract painting minimizes object semantics and foregrounds structura…
Text-to-Image GenerationCross-Modal RetrievalAttribute Token Arithmetic: Disentangled and Continuous Semantic Control for Visual Autoregressive Models
Autoregressive text-to-image generation has recently achieved remarkable progress, offering high-fidelity synthesis via a unified generative framework. However, fine-grained semantic control remains challenging due to th…
Computational EfficiencyText-to-Image GenerationRubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing
Reward models play an essential role in aligning visual generative models, yet most existing visual reward models use a single scalar score or rely on fixed criteria that cannot adapt to different instructions. This limi…
Text-to-Image GenerationImage Editing