Image Generation
93개 벤치마크 · 논문 7,946편 · 이 태스크의 논문 보기 →
Benchmarks
ImageNet 256x256
CIFAR-10
ImageNet 64x64
ImageNet 512x512
FFHQ 256 x 256
CelebA 64x64
ImageNet 32x32
LSUN Bedroom 256 x 256
STL-10
LSUN Churches 256 x 256
ImageNet 128x128
FFHQ 1024 x 1024
CelebA-HQ 256x256
CelebA 256x256
MNIST
WISE
FFHQ
FFHQ-U
Binarized MNIST
CelebA-HQ 1024x1024
CIFAR-100
AFHQ Cat
LSUN Cat 256 x 256
AFHQV2
CelebA-HQ 128x128
Fashion-MNIST
TextAtlasEval
AFHQ Dog
CLEVR
Cityscapes
LSUN Horse 256 x 256
AFHQ Wild
CelebA 128x128
Places50
ARKitScenes
CUB 128 x 128
Pokemon 256x256
Replica
Stanford Cars
Stanford Dogs
VLN-CE
VizDoom
ADE-Indoor
CAT 256x256
CIFAR-10 (10% data)
CIFAR-10 (20% data)
CelebA-HQ 64x64
FFHQ 128 x 128
FFHQ 512 x 512
LSUN Bedroom
LSUN Bedroom 64 x 64
MetFaces
MetFaces-U
ObjectsRoom
Pokemon 1024x1024
ShapeStacks
Stacked MNIST
TiO_2 nanoparticle
AFHQ-v2 64x64
FFHQ 64x64
LSUN Bedroom 128 x 128
LSUN Car 512 x 384
RC-49
iNaturalist 2019
25% ImageNet 128x128
CelebA
CelebA-HQ
CelebA-HQ 512x512
Cityscapes-25K 256x512
Cityscapes-5K 256x512
EMNIST-Letters
GQN
Indian Celebs 256 x 256
KMNIST
LLVIP
LSUN
LSUN Car 256 x 256
LSUN tower 64x64
Landscapes 256 x 256
Multi-dSprites
NASA Perseverance
SDSS Galaxies
Most implemented
Analyzing and Improving the Image Quality of StyleGAN
Wasserstein GAN
Progressive Growing of GANs for Improved Quality, Stability, and Variation
Improved Training of Wasserstein GANs
A Style-Based Generator Architecture for Generative Adversarial Networks
Papers
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's matur…
Surface Normals EstimationMonocular Depth EstimationImage GenerationWeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing
Image generation and editing models have advanced rapidly, yet remain unreliable when prompts require external world knowledge. Bounded and long-tail parametric knowledge prevents direct or reason-then-generate approache…
Image GenerationImage EditingRefDiT: Local Attribute Guidance in Reference-Based Image Generation
Personalization models generate new images guided by a few subject references, while style transfer methods aim to produce images aligned with a global style derived from a reference image. Recent approaches perform well…
Image GenerationStyle TransferEditable Visual Design
While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-…
Image GenerationWhen Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images
*Chulin Zhao and Ruoqi Hu contributed equally to this work. State-of-the-art text-to-image (T2I) models exhibit pronounced and systematic defects when prompts involve intricate compositional factors such as multiple enti…
Image GenerationEfficient Training with Foresight: Multi-Token Auxiliary Supervision for Autoregressive Image Generation
Autoregressive (AR) image generation has shown strong potential for scalable high-fidelity synthesis by modeling images as discrete token sequences. However, traditional next token prediction (NTP) continues to suffer fr…
Image Generation