paper-with-me

Text-to-Image Generation

17개 벤치마크 · 논문 1,553편 · 이 태스크의 논문 보기 →

Benchmarks

GenEval

결과 101개

CUB

결과 100개

Multi-Modal-CelebA-HQ

결과 50개

DrawBench

결과 40개

Oxford 102 Flowers

결과 40개

Conceptual Captions

결과 25개

COCO

결과 15개

LHQC

결과 15개

MS-COCO

결과 11개

DPG

결과 10개

GeNeVA (CoDraw)

결과 10개

GeNeVA (i-CLEVR)

결과 10개

LAION COCO

결과 10개

T2I-CompBench

결과 10개

Colors

결과 5개

Flickr-8k

결과 5개

Most implemented

Papers

Importance-Aware Low-Rank Distillation of Diffusion Transformers

2026-09-04 · Denis Zavadski, Sebastian Heid, Damjan Kalšan, Stefan Roth 외 arxiv

Diffusion Transformers (DiTs) have emerged as a dominant architecture for high-quality text-to-image generation, yet their scale poses challenges for efficient deployment. While truncated singular value decomposition (SV…

Text-to-Image GenerationKnowledge Distillation

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

2026-08-31 · Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti, Farina Amir 외 arxiv

We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question answering and image-grounded hallucination …

Visual Question AnsweringText-to-Image Generation

GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling

2026-08-29 · Guangting Zheng, Yiyuan Zhang, Tao Yang, Yunpeng Chen 외 hf

Latent generative models typically follow a two-stage pipeline, training a variational autoencoder for reconstruction and then a generative model on the frozen latent space. Since reconstruction-optimized latents are not…

Text-to-Image GenerationRepresentation Learning

Abstract4D: A Large-Scale Dataset and Framework for Understanding the Visual Language of Abstract Art

2026-08-28 · Haowei Zhang, Yuanpei Zhao, Ji-Zhe Zhou, Mao Li arxiv

Artificial intelligence can classify artistic styles and synthesize images, but it still lacks a model of the visual language that gives art meaning. Abstract painting minimizes object semantics and foregrounds structura…

Text-to-Image GenerationCross-Modal Retrieval

Attribute Token Arithmetic: Disentangled and Continuous Semantic Control for Visual Autoregressive Models

2026-08-28 · Xindi Yang, Yicheng Wu, Cheng Zhang, Jianfei Cai 외 arxiv

Autoregressive text-to-image generation has recently achieved remarkable progress, offering high-fidelity synthesis via a unified generative framework. However, fine-grained semantic control remains challenging due to th…

Computational EfficiencyText-to-Image Generation

RubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing

2026-08-27 · Zijian Kan, Wei Wang, Long Luo, Bing Zhao 외 arxiv

Reward models play an essential role in aligning visual generative models, yet most existing visual reward models use a single scalar score or rely on fixed criteria that cannot adapt to different instructions. This limi…

Text-to-Image GenerationImage Editing

전체 1,553편 보기 →