paper-with-me

홈 › Papers

Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?

2025-09-03 · Ouxiang Li, Yuan Wang, Xinting Hu, Huijuan Huang, Rui Chen, Jiarong Ou, Xin Tao, Pengfei Wan, Xiaojuan Qi, Fuli Feng arxiv

Text-to-image (T2I) generation aims to synthesize images from textual prompts, which jointly specify what must be shown and imply what can be inferred, which thus correspond to two core capabilities: \textbf{\textit{composition}} and \textbf{\textit{reasoning}}. Despite recent advances of T2I models in both composition and reasoning, existing benchmarks remain limited in evaluation. They not only fail to provide comprehensive coverage across and within both capabilities, but also largely restrict evaluation to low scene density and simple one-to-one reasoning. To address these limitations, we propose \textbf{\textsc{T2I-CoReBench}}, a comprehensive and complex benchmark that evaluates both composition and reasoning capabilities of T2I models. To ensure comprehensiveness, we structure composition around scene graph elements (\textit{instance}, \textit{attribute}, and \textit{relation}) and reasoning around the philosophical framework of inference (\textit{deductive}, \textit{inductive}, and \textit{abductive}), formulating a 12-dimensional evaluation taxonomy. To increase complexity, driven by the inherent real-world complexities, we curate each prompt with higher compositional density for composition and greater reasoning intensity for reasoning. To facilitate fine-grained and reliable evaluation, we also pair each evaluation prompt with a checklist that specifies individual \textit{yes/no} questions to assess each intended element independently. In statistics, our benchmark comprises 1,080 challenging prompts and around 13,500 checklist questions. Experiments across 38 current T2I models reveal that their composition capability still remains limited in high compositional scenarios, while the reasoning capability lags even further behind as a critical bottleneck, with all models struggling to infer implicit elements from prompts.

📄 PDF Abstract BibTeX arXiv:2509.03516

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Contextual-based Image Inpainting: Infer, Match, and Translate

2017-11-23 · ECCV 2018 9 · Yuhang Song, Chao Yang, Zhe Lin, Xiaofeng Liu 외

We study the task of image inpainting, which is to fill in the missing region of an incomplete image with plausible contents. To this end, we propose a learning-based approach to generate visually coherent completion giv…

Image InpaintingTranslation

Rethinking Fast Fourier Convolution in Image Inpainting

2023-01-01 · ICCV 2023 1 · Tianyi Chu, Jiafu Chen, Jiakai Sun, Shuobin Lian 외

Recently proposed image inpainting method LaMa builds its network upon Fast Fourier Convolution (FFC), which was originally proposed for high-level vision tasks like image classification. FFC empowers the fully convo…

image-classificationImage ClassificationImage Inpainting

PaintFlow: A Unified Framework for Interactive Oil Paintings Editing and Generation

2025-12-09 · Zhangli Hu, Ye Chen, Jiajun Yao, Bingbing Ni arxiv

Oil painting, as a high-level medium that blends human abstract thinking with artistic expression, poses substantial challenges for digital generation and editing due to its intricate brushstroke dynamics and stylized ch…

Style Transfer

Deep Face Video Inpainting via UV Mapping

2021-09-02 · Wenqi Yang, Zhenfang Chen, Chaofeng Chen, GuanYing Chen 외

This paper addresses the problem of face video inpainting. Existing video inpainting methods target primarily at natural scenes with repetitive patterns. They do not make use of any prior knowledge of the face to help re…

Facial InpaintingVideo Inpainting

Space Narrative: Generating Images and 3D Scenes of Chinese Garden from Text using Deep Learning

2023-11-01 · Jiaxi Shi1, Hao Hua1

The consistent mapping from poems to paintings is essential for the research and restoration of traditional Chinese gardens. But the lack of firsthand ma-terial is a great challenge to the reconstruction work. In this pa…

Unity