paper-with-me

홈 › Papers

Toward Verifiable and Reproducible Human Evaluation for Text-to-Image Generation

2023-04-04 · CVPR 2023 1 · Mayu Otani, Riku Togashi, Yu Sawai, Ryosuke Ishigami, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, Shin'ichi Satoh

Human evaluation is critical for validating the performance of text-to-image generative models, as this highly cognitive process requires deep comprehension of text and images. However, our survey of 37 recent papers reveals that many works rely solely on automatic measures (e.g., FID) or perform poorly described human evaluations that are not reliable or repeatable. This paper proposes a standardized and well-defined human evaluation protocol to facilitate verifiable and reproducible human evaluation in future works. In our pilot data collection, we experimentally show that the current automatic measures are incompatible with human perception in evaluating the performance of the text-to-image generation results. Furthermore, we provide insights for designing human evaluation experiments reliably and conclusively. Finally, we make several resources publicly available to the community to facilitate easy and fast implementations.

📄 PDF Abstract BibTeX arXiv:2304.01816

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

AEMA: Verifiable Evaluation Framework for Trustworthy and Controlled Agentic LLM Systems

2026-01-17 · YenTing Lee, Keerthi Koneru, Zahra Moslemi, Sheethal Kumar 외 arxiv

Evaluating large language model (LLM)-based multi-agent systems remains a critical challenge, as these systems must exhibit reliable coordination, transparent decision-making, and verifiable performance across evolving t…

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

2026-05-14 · Md Tahmid Rahman Laskar, Xue-Yong Fu, Seyyed Saeed Sarfjoo, Quinten McNamara 외 arxiv

Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling benchmarks remain text-based. We study whether verified text benchmarks can be converted into controlled audio-based tool ca…

Instruction-Following Evaluation for Large Language Models

2023-11-14 · Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma 외

One core capability of Large Language Models (LLMs) is to follow natural language instructions. However, the evaluation of such abilities is not standardized: Human evaluations are expensive, slow, and not objectively re…

Instruction Following

V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models

2025-11-20 · Yang Luo, Xuanlei Zhao, Baijiong Lin, Lingting Zhu 외 arxiv

Recent progress in generative video models, such as Veo-3, has shown surprising zero-shot reasoning abilities, creating a growing need for systematic and reliable evaluation. We introduce V-ReasonBench, a benchmark desig…

Video Generation

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents

2026-04-08 · Mingyu Ouyang, Siyuan Hu, Kevin Qinghong Lin, Hwee Tou Ng 외 arxiv

Towards an embodied generalist for real-world interaction, Multimodal Large Language Model (MLLM) agents still suffer from challenging latency, sparse feedback, and irreversible mistakes. Video games offer an ideal testb…

Action Parsing