paper-with-me

Papers

FRAbench and GenEval: Scaling Fine-Grained Aspect Evaluation across Tasks, Modalities

2025-05-19 · Shibo Hong, Jiahao Ying, Haiyuan Liang, Mengdi Zhang, Jun Kuang, Jiazheng Zhang, Yixin Cao

Evaluating the open-ended outputs of large language models (LLMs) has become a bottleneck as model capabilities, task diversity, and modality coverage rapidly expand. Existing "LLM-as-a-Judge" evaluators are typically narrow in a few tasks, aspects, or modalities, and easily suffer from low consistency. In this paper, we argue that explicit, fine-grained aspect specification is the key to both generalizability and objectivity in automated evaluation. To do so, we introduce a hierarchical aspect taxonomy spanning 112 aspects that unifies evaluation across four representative settings - Natural Language Generation, Image Understanding, Image Generation, and Interleaved Text-and-Image Generation. Building on this taxonomy, we create FRAbench, a benchmark comprising 60.4k pairwise samples with 325k aspect-level labels obtained from a combination of human and LLM annotations. FRAbench provides the first large-scale, multi-modal resource for training and meta-evaluating fine-grained LMM judges. Leveraging FRAbench, we develop GenEval, a fine-grained evaluator generalizable across tasks and modalities. Experiments show that GenEval (i) attains high agreement with GPT-4o and expert annotators, (ii) transfers robustly to unseen tasks and modalities, and (iii) reveals systematic weaknesses of current LMMs on evaluation.

📄 PDF Abstract BibTeX arXiv:2505.12795

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText Generation

Similar Papers 제목 키워드 기반

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

2026-07-31 · Yuan Gao, Zeren Yang, Junnan Li, Shawn 외 arxiv

Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, b…

MT-GenEval: A Counterfactual and Contextual Dataset for Evaluating Gender Accuracy in Machine Translation

2022-11-02 · Anna Currey, Maria Nădejde, Raghavendra Pappagari, Mia Mayer 외

As generic machine translation (MT) quality has improved, the need for targeted benchmarks that explore fine-grained aspects of quality has increased. In particular, gender accuracy in translation can have implications i…

counterfactualEthicsMachine TranslationSentence+1

RAISE: Requirement-Adaptive Evolutionary Refinement for Training-Free Text-to-Image Alignment

2026-02-28 · Liyao Jiang, Ruichen Chen, Chao Gao, Di Niu arxiv

Recent text-to-image (T2I) diffusion models achieve remarkable realism, yet faithful prompt-image alignment remains challenging, particularly for complex prompts with multiple objects, relations, and fine-grained attribu…

Image Generation

Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection

2025-03-15 · Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Arsh Koneru 외

The predominant approach to advancing text-to-image generation has been training-time scaling, where larger models are trained on more data using greater computational resources. While effective, this approach is computa…

Image GenerationText to Image GenerationText-to-Image Generation

Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

2024-10-17 · Lijie Fan, Tianhong Li, Siyang Qin, Yuanzhen Li 외

Scaling up autoregressive models in vision has not proven as beneficial as in large language models. In this work, we investigate this scaling problem in the context of text-to-image generation, focusing on two critical …

Image GenerationText to Image GenerationText-to-Image Generation