paper-with-me

홈 › Papers

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation

2024-11-27 · CVPR 2025 1 · Pengfei Zhou, Xiaopeng Peng, Jiajun Song, Chuanhao Li, Zhaopan Xu, Yue Yang, Ziyao Guo, Hao Zhang, Yuqi Lin, Yefei He, Lirui Zhao, Shuo Liu, Tianhua Li, Yuxuan Xie, Xiaojun Chang, Yu Qiao, Wenqi Shao, Kaipeng Zhang

Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding and generation tasks. However, generating interleaved image-text content remains a challenge, which requires integrated multimodal understanding and generation abilities. While the progress in unified models offers new solutions, existing benchmarks are insufficient for evaluating these methods due to limitations in data size and diversity. To bridge this gap, we introduce OpenING, a comprehensive benchmark comprising 5,400 high-quality human-annotated instances across 56 real-world tasks. OpenING covers diverse daily scenarios such as travel guide, design, and brainstorming, offering a robust platform for challenging interleaved generation methods. In addition, we present IntJudge, a judge model for evaluating open-ended multimodal generation methods. Trained with a novel data pipeline, our IntJudge achieves an agreement rate of 82.42% with human judgments, outperforming GPT-based evaluators by 11.34%. Extensive experiments on OpenING reveal that current interleaved generation methods still have substantial room for improvement. Key findings on interleaved image-text generation are further presented to guide the development of next-generation models.

📄 PDF Abstract BibTeX arXiv:2411.18499

Code (1)

LanceZPF/OpenING pytorch

Tasks

Image Generationmultimodal generationText Generation

Methods 이 논문이 사용한 방법론

Travel 설명 없음

Similar Papers 제목 키워드 기반

SLMJury: Can Small Language Models Judge as Well as Large Ones?

2026-06-05 · Anish Laddha, Nitesh Pradhan, Gaurav Srivastava arxiv

Large language models (LLMs) are widely used as judges for evaluating model outputs, but their high cost, latency, and opacity limit scalability. We introduce SLMJury, a framework for evaluating small language models (SL…

Domain Generalization

SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

2026-05-29 · Wai-Chung Kwan, Aryo Pradipta Gema, Joshua Ong Jun Leang, Pasquale Minervini arxiv

Self-play can train language models without external supervision. However, existing methods require rule-checkable answers, leaving open-ended tasks dependent on curated prompts or frontier-model judges. We introduce SCO…

StoryER: Automatic Story Evaluation via Ranking, Rating and Reasoning

2022-10-16 · Hong Chen, Duc Minh Vo, Hiroya Takamura, Yusuke Miyao 외

Existing automatic story evaluation methods place a premium on story lexical level coherence, deviating from human preference. We go beyond this limitation by considering a novel \textbf{Story} \textbf{E}valuation method…

Comment GenerationDecoder

U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs

2024-12-04 · Konstantin Chernyshev, Vitaliy Polshkov, Ekaterina Artemova, Alex Myasnikov 외

The current evaluation of mathematical skills in LLMs is limited, as existing benchmarks are either relatively small, primarily focus on elementary and high-school problems, or lack diversity in topics. Additionally, the…

DiversityMath

ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks

2025-03-10 · Yan Yang, Dongxu Li, HaoNing Wu, Bei Chen 외

Solving expert-level multimodal tasks is a key milestone towards general intelligence. As the capabilities of multimodal large language models (MLLMs) continue to improve, evaluation of such advanced multimodal intellige…