paper-with-me

Image Generation

93개 벤치마크 · 논문 7,946편 · 이 태스크의 논문 보기 →

Benchmarks

ImageNet 256x256

결과 102개

CIFAR-10

결과 79개

ImageNet 64x64

결과 65개

ImageNet 512x512

결과 52개

FFHQ 256 x 256

결과 51개

CelebA 64x64

결과 39개

ImageNet 32x32

결과 35개

LSUN Bedroom 256 x 256

결과 32개

STL-10

결과 31개

ImageNet 128x128

결과 24개

FFHQ 1024 x 1024

결과 20개

CelebA-HQ 256x256

결과 19개

CelebA 256x256

결과 17개

MNIST

결과 15개

WISE

결과 15개

FFHQ

결과 13개

FFHQ-U

결과 13개

Binarized MNIST

결과 10개

CelebA-HQ 1024x1024

결과 10개

CIFAR-100

결과 9개

AFHQ Cat

결과 8개

LSUN Cat 256 x 256

결과 8개

AFHQV2

결과 7개

CelebA-HQ 128x128

결과 7개

Fashion-MNIST

결과 7개

TextAtlasEval

결과 7개

AFHQ Dog

결과 6개

CLEVR

결과 6개

Cityscapes

결과 6개

LSUN Horse 256 x 256

결과 6개

AFHQ Wild

결과 5개

CelebA 128x128

결과 5개

Places50

결과 5개

ARKitScenes

결과 4개

CUB 128 x 128

결과 4개

Pokemon 256x256

결과 4개

Replica

결과 4개

Stanford Cars

결과 4개

Stanford Dogs

결과 4개

VLN-CE

결과 4개

VizDoom

결과 4개

ADE-Indoor

결과 3개

CAT 256x256

결과 3개

CIFAR-10 (10% data)

결과 3개

CIFAR-10 (20% data)

결과 3개

CelebA-HQ 64x64

결과 3개

FFHQ 128 x 128

결과 3개

FFHQ 512 x 512

결과 3개

LSUN Bedroom

결과 3개

LSUN Bedroom 64 x 64

결과 3개

MetFaces

결과 3개

MetFaces-U

결과 3개

ObjectsRoom

결과 3개

Pokemon 1024x1024

결과 3개

ShapeStacks

결과 3개

Stacked MNIST

결과 3개

TiO_2 nanoparticle

결과 3개

AFHQ-v2 64x64

결과 2개

FFHQ 64x64

결과 2개

LSUN Car 512 x 384

결과 2개

RC-49

결과 2개

iNaturalist 2019

결과 2개

25% ImageNet 128x128

결과 1개

CelebA

결과 1개

CelebA-HQ

결과 1개

CelebA-HQ 512x512

결과 1개

Cityscapes-5K 256x512

결과 1개

EMNIST-Letters

결과 1개

GQN

결과 1개

KMNIST

결과 1개

LLVIP

결과 1개

LSUN

결과 1개

LSUN Car 256 x 256

결과 1개

LSUN tower 64x64

결과 1개

Landscapes 256 x 256

결과 1개

Multi-dSprites

결과 1개

NASA Perseverance

결과 1개

SDSS Galaxies

결과 1개

Most implemented

Wasserstein GAN

2017-01-26 · 구현 120개

Improved Training of Wasserstein GANs

2017-03-31 · 구현 110개

Papers

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

2026-09-08 · Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei 외 hf

Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's matur…

Surface Normals EstimationMonocular Depth EstimationImage Generation

WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing

2026-09-04 · Hui Zhang, Zongkai Liu, Liqiang Niu, Juntao Liu 외 arxiv

Image generation and editing models have advanced rapidly, yet remain unreliable when prompts require external world knowledge. Bounded and long-tail parametric knowledge prevents direct or reason-then-generate approache…

Image GenerationImage Editing

RefDiT: Local Attribute Guidance in Reference-Based Image Generation

2026-09-04 · Rameshwar Mishra, Srikrishna Karanam, A V Subramanyam arxiv

Personalization models generate new images guided by a few subject references, while style transfer methods aim to produce images aligned with a global style derived from a reference image. Recent approaches perform well…

Image GenerationStyle Transfer

Editable Visual Design

2026-09-03 · Junyan Ye, Wei Liu, Dongzhi Jiang, Zichen Wen 외 hf

While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-…

Image Generation

When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images

2026-08-26 · Ruoqi Hu, Chulin Zhao, Jiashuo Chang, Ramon Ruiz-Dolz 외 arxiv

*Chulin Zhao and Ruoqi Hu contributed equally to this work. State-of-the-art text-to-image (T2I) models exhibit pronounced and systematic defects when prompts involve intricate compositional factors such as multiple enti…

Image Generation

Efficient Training with Foresight: Multi-Token Auxiliary Supervision for Autoregressive Image Generation

2026-08-26 · Guo Niu, Xiongfei Yao, Teng Wang, Nannan Zhu arxiv

Autoregressive (AR) image generation has shown strong potential for scalable high-fidelity synthesis by modeling images as discrete token sequences. However, traditional next token prediction (NTP) continues to suffer fr…

Image Generation

전체 7,946편 보기 →