paper-with-me

Papers

Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation

2025-09-25 · Seyed Amir Kasaei, Ali Aghayari, Arash Marioriyad, Niki Sepasian, MohammadAmin Fazli, Mahdieh Soleymani Baghshah, Mohammad Hossein Rohban arxiv

Text-image generation has advanced rapidly, but assessing whether outputs truly capture the objects, attributes, and relations described in prompts remains a central challenge. Evaluation in this space relies heavily on automated metrics, yet these are often adopted by convention or popularity rather than validated against human judgment. Because evaluation and reported progress in the field depend directly on these metrics, it is critical to understand how well they reflect human preferences. To address this, we present a broad study of widely used metrics for compositional text-image evaluation. Our analysis goes beyond simple correlation, examining their behavior across diverse compositional challenges and comparing how different metric families align with human judgments. The results show that no single metric performs consistently across tasks: performance varies with the type of compositional problem. Notably, VQA-based metrics, though popular, are not uniformly superior, while certain embedding-based metrics prove stronger in specific cases. Image-only metrics, as expected, contribute little to compositional evaluation, as they are designed for perceptual quality rather than alignment. These findings underscore the importance of careful and transparent metric selection, both for trustworthy evaluation and for their use as reward models in generation. Project page is available at https://amirkasaei.com/eval-the-evals/ .

📄 PDF Abstract BibTeX arXiv:2509.21227

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge

2026-04-20 · Sua Lee, Sanghee Park, Jinbae Im arxiv

Multimodal Large Language Models (MLLMs) have been increasingly used as automatic evaluators-a paradigm known as MLLM-as-a-Judge. However, their reliability and vulnerabilities to biases remain underexplored. We find tha…

T2I-FineEval: Fine-Grained Compositional Metric for Text-to-Image Evaluation

2025-03-14 · Seyed Mohammad Hadi Hosseini, Amir Mohammad Izadi, Ali Abdollahi, Armin Saghafian 외

Although recent text-to-image generative models have achieved impressive performance, they still often struggle with capturing the compositional complexities of prompts including attribute binding, and spatial relationsh…

AttributeQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

A Gamified Evaluation and Recruitment Platform for Low Resource Language Machine Translation Systems

2025-06-13 · Carlos Rafael Catalan

Human evaluators provide necessary contributions in evaluating large language models. In the context of Machine Translation (MT) systems for low-resource languages (LRLs), this is made even more apparent since popular au…

Machine Translation

Evaluating Object-Centric Models beyond Object Discovery

2026-02-07 · Krishnakant Singh, Simone Schaub-Meyer, Stefan Roth arxiv

Object-centric learning (OCL) aims to learn structured scene representations that support compositional generalization and robustness to out-of-distribution (OOD) data. However, OCL models are often not evaluated regardi…

Image Classification

GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

2024-06-19 · Baiqi Li, Zhiqiu Lin, Deepak Pathak, Jiayao Li 외

While text-to-visual models now produce photo-realistic images and videos, they struggle with compositional text prompts involving attributes, relationships, and higher-order reasoning such as logic and comparison. In th…

BenchmarkingImage GenerationVideo Generation