paper-with-me

홈 › Papers

GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?

2026-02-05 · Ruihang Li, Leigang Qu, Jingxu Zhang, Dongnan Gui, Mengde Xu, Xiaosong Zhang, Han Hu, Wenjie Wang, Jiaqi Wang arxiv

The rapid advancement of visual generation models has outpaced traditional evaluation approaches, necessitating the adoption of Vision-Language Models as surrogate judges. In this work, we systematically investigate the reliability of the prevailing absolute pointwise scoring standard, across a wide spectrum of visual generation tasks. Our analysis reveals that this paradigm is limited due to stochastic inconsistency and poor alignment with human perception. To resolve these limitations, we introduce GenArena, a unified evaluation framework that leverages a pairwise comparison paradigm to ensure stable and human-aligned evaluation. Crucially, our experiments uncover a transformative finding that simply adopting this pairwise protocol enables off-the-shelf open-source models to outperform top-tier proprietary models. Notably, our method boosts evaluation accuracy by over 20% and achieves a Spearman correlation of 0.86 with the authoritative LMArena leaderboard, drastically surpassing the 0.36 correlation of pointwise methods. Based on GenArena, we benchmark state-of-the-art visual generation models across diverse tasks, providing the community with a rigorous and automated evaluation standard for visual generation.

📄 PDF Abstract BibTeX arXiv:2602.06013

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RefReward-SR: LR-Conditioned Reward Modeling for Preference-Aligned Super-Resolution

2026-03-25 · Yushuai Song, Weize Quan, Weining Wang, Jiahui Sun 외 arxiv

Recent advances in generative super-resolution (SR) have greatly improved visual realism, yet existing evaluation and optimization frameworks remain misaligned with human perception. Full-Reference and No-Reference metri…

GEditBench v2: A Human-Aligned Benchmark for General Image Editing

2026-03-30 · Zhangqi Jiang, Zheng Sun, Xianfang Zeng, Yufeng Yang 외 arxiv

Recent advances in image editing have enabled models to handle complex instructions with impressive realism. However, existing evaluation frameworks lag behind: current benchmarks suffer from narrow task coverage, while …

Image Editing

FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding

2026-01-24 · João Pereira, Vasco Lopes, João Neves, David Semedo arxiv

Video Anomaly Understanding (VAU) is a novel task focused on describing unusual occurrences in videos. Despite growing interest, the evaluation of VAU remains an open challenge. Existing benchmarks rely on n-gram-based m…

DB-3DME: From Dataset to Benchmark for Human-aligned Automatic 3D Mesh Evaluation

2026-06-08 · Nanshan Jia, Zhenyu Zhao, Sui Huang, Jingshen Wang 외 arxiv

Recent advances in 3D generation have led to substantial improvements in realism, controllability, and efficiency, yet the evaluation of 3D assets remains underexplored. Existing evaluation paradigms, including human eva…

3D Generation

WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics

2026-03-11 · Yuhong Dai, Yanlin Lai, Mitt Huang, Hangyu Guo 외 arxiv

Existing web-generation benchmarks rely on text prompts or static screenshots as input. However, videos naturally convey richer signals such as interaction flow, transition timing, and motion continuity, which are essent…