paper-with-me

홈 › Papers

Human Evaluation of Text-to-Image Models on a Multi-Task Benchmark

2022-11-22 · Vitali Petsiuk, Alexander E. Siemenn, Saisamrit Surbehera, Zad Chin, Keith Tyser, Gregory Hunter, Arvind Raghavan, Yann Hicke, Bryan A. Plummer, Ori Kerret, Tonio Buonassisi, Kate Saenko, Armando Solar-Lezama, Iddo Drori

We provide a new multi-task benchmark for evaluating text-to-image models. We perform a human evaluation comparing the most common open-source (Stable Diffusion) and commercial (DALL-E 2) models. Twenty computer science AI graduate students evaluated the two models, on three tasks, at three difficulty levels, across ten prompts each, providing 3,600 ratings. Text-to-image generation has seen rapid progress to the point that many recent models have demonstrated their ability to create realistic high-resolution images for various prompts. However, current text-to-image methods and the broader body of research in vision-language understanding still struggle with intricate text prompts that contain many objects with multiple attributes and relationships. We introduce a new text-to-image benchmark that contains a suite of thirty-two tasks over multiple applications that capture a model's ability to handle different features of a text prompt. For example, asking a model to generate a varying number of the same object to measure its ability to count or providing a text prompt with several objects that each have a different attribute to identify its ability to match objects and attributes correctly. Rather than subjectively evaluating text-to-image results on a set of prompts, our new multi-task benchmark consists of challenge tasks at three difficulty levels (easy, medium, and hard) and human ratings for each generated image.

📄 PDF Abstract BibTeX arXiv:2211.12112

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeImage GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks

2023-11-02 · Xinlu Zhang, Yujie Lu, Weizhi Wang, An Yan 외

Automatically evaluating vision-language tasks is challenging, especially when it comes to reflecting human judgments due to limitations in accounting for fine-grained details. Although GPT-4V has shown promising results…

Image GenerationImage to text

Towards Scalable Human-aligned Benchmark for Text-guided Image Editing

2025-05-01 · CVPR 2025 1 · Suho Ryu, Kihyun Kim, Eugene Baek, Dongsoo Shin 외

A variety of text-guided image editing models have been proposed recently. However, there is no widely-accepted standard evaluation method mainly due to the subjective nature of the task, letting researchers rely on manu…

text-guided-image-editing

MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation

2025-05-23 · Jihan Yao, Yushi Hu, Yujie Yi, Bin Han 외

Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align reliably with human evaluation, especially for complex tasks that involve multiple modalities.…

Audio GenerationBenchmarkingImage Generationmultimodal generation+1

HarmonicEval: Multi-modal, Multi-task, Multi-criteria Automatic Evaluation Using a Vision Language Model

2024-12-19 · Masanari Ohi, Masahiro Kaneko, Naoaki Okazaki, Nakamasa Inoue

Vision-language models (VLMs) have shown impressive abilities in text and image understanding. However, existing metrics for evaluating the text generated by VLMs focus exclusively on overall quality, leading to two limi…

Language ModelingLanguage Modelling

Evaluating Hallucination in Text-to-Image Diffusion Models with Scene-Graph based Question-Answering Agent

2024-12-07 · Ziyuan Qin, Dongjie Cheng, Haoyu Wang, Huahui Yi 외

Contemporary Text-to-Image (T2I) models frequently depend on qualitative human evaluations to assess the consistency between synthesized images and the text prompts. There is a demand for quantitative and automatic evalu…

HallucinationQuestion Answering