paper-with-me

Papers

GenEval: A Benchmark Suite for Evaluating Generative Models

2018-09-27 · Anton Bakhtin, Arthur Szlam, Marc'Aurelio Ranzato

Generative models are important for several practical applications, from low level image processing tasks, to model-based planning in robotics. More generally, the study of generative models is motivated by the long-standing endeavor to model uncertainty and to discover structure by leveraging unlabeled data. Unfortunately, the lack of an ultimate task of interest has hindered progress in the field, as there is no established way to compare models and, often times, evaluation is based on mere visual inspection of samples drawn from such models. In this work, we aim at addressing this problem by introducing a new benchmark evaluation suite, dubbed \textit{GenEval}. GenEval hosts a large array of distributions capturing many important properties of real datasets, yet in a controlled setting, such as lower intrinsic dimensionality, multi-modality, compositionality, independence and causal structure. Any model can be easily plugged for evaluation, provided it can generate samples. Our extensive evaluation suggests that different models have different strenghts, and that GenEval is a great tool to gain insights about how models and metrics work. We offer GenEval to the community~\footnote{Available at: \it{coming soon}.} and believe that this benchmark will facilitate comparison and development of new generative models.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

2023-10-17 · NeurIPS 2023 11 · Dhruba Ghosh, Hanna Hajishirzi, Ludwig Schmidt

Recent breakthroughs in diffusion models, multimodal pretraining, and efficient finetuning have led to an explosion of text-to-image generative models. Given human evaluation is expensive and difficult to scale, automate…

AttributeObjectobject-detectionObject Detection

MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine Translation

2022-11-02 · Anna Currey, Maria Nădejde, Raghavendra Pappagari, Mia Mayer 외

As generic machine translation (MT) quality has improved, the need for targeted benchmarks that explore fine-grained aspects of quality has increased. In particular, gender accuracy in translation can have implications i…

counterfactualEthicsMachine TranslationSentence+1

BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation

2026-03-26 · Yan Li, Zezi Zeng, Ziwei Zhou, Xin Gao 외 arxiv

Recent advances in image generation models have expanded their applications beyond aesthetic imagery toward practical visual content creation. However, existing benchmarks mainly focus on natural image synthesis and fail…

Image Generation

FRAbench and GenEval: Scaling Fine-Grained Aspect Evaluation across Tasks, Modalities

2025-05-19 · Shibo Hong, Jiahao Ying, Haiyuan Liang, Mengdi Zhang 외

Evaluating the open-ended outputs of large language models (LLMs) has become a bottleneck as model capabilities, task diversity, and modality coverage rapidly expand. Existing "LLM-as-a-Judge" evaluators are typically na…

Image GenerationText Generation

GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation

2025-12-18 · Amita Kamath, Kai-Wei Chang, Ranjay Krishna, Luke Zettlemoyer 외 arxiv

Automating Text-to-Image (T2I) model evaluation is challenging; a judge model must be used to score correctness, and test prompts must be selected to be challenging for current T2I models but not the judge. We argue that…