paper-with-me

홈 › Papers

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

2026-06-04 · Yida Yin, Harish Krishnakumar, Chung Peng Lee, Boya Zeng, Wenhao Chai, Shengbang Tong, Wenhu Chen, Hu Xu, Xingyu Fu, Gabriel Sarch, Aleksandra Korolova, Zhuang Liu arxiv

In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types without capturing the visual diversity needed to handle open-ended visual inputs. We present WorldBench, a challenging and visually diverse reasoning benchmark to evaluate Multimodal Large Language Models (MLLMs). We build a taxonomy of thousands of visual concepts across multiple domains (e.g., living things). Guided by this taxonomy, we curate a broad collection of images from search engines and existing datasets to comprehensively represent the visual world. Through structured trial-and-error, we manually design challenging questions that frontier MLLMs fail to answer. On quantitative and human evaluations, WorldBench achieves higher visual diversity than any existing diverse benchmark. Evaluating 15 MLLMs on WorldBench reveals weaknesses in visual understanding: even the strongest model reaches only 64.0% accuracy, while some models perform marginally above chance-level. We hope our work highlights the importance of visual diversity in building multimodal benchmarks.

📄 PDF Abstract BibTeX arXiv:2606.06538

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation

2025-07-24 · Yubin Chen, Xuyang Guo, Zhenmei Shi, Zhao Song 외 arxiv

Text-to-video (T2V) models have shown remarkable performance in generating visually reasonable scenes, while their capability to leverage world knowledge for ensuring semantic consistency and factual accuracy remains lar…

Text-to-Video Generation

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning

2025-05-26 · Minheng Ni, Zhengyuan Yang, Linjie Li, Chung-Ching Lin 외

Recent advances in large language models have significantly improved textual reasoning through the effective use of Chain-of-Thought (CoT) and reinforcement learning. However, extending these successes to vision-language…

document understandingMultimodal ReasoningVisual Reasoning

WorldBench: Benchmarking Physical Understanding of World Models by Isolating Physics Concepts

2026-01-29 · Rishi Upadhyay, Howard Zhang, Jim Solomon, Ayush Agrawal 외 arxiv

Recent advances in generative foundational models, often termed "world models," have propelled interest in applying them to critical tasks like robotic planning and autonomous system training. For reliable deployment, th…

Video Generation

DeepVision-103K: A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning

2026-02-18 · Haoxiang Sun, Lizhen Xu, Bing Zhao, Wotao Yin 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has been shown effective in enhancing the visual reflection and reasoning capabilities of Large Multimodal Models (LMMs). However, existing datasets are predominantly…

Reinforcement LearningMultimodal Reasoning

Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models

2026-03-23 · Meiqi Wu, Zhixin Cai, Fufangchen Zhao, Xiaokun Feng 외 arxiv

Video--based world models have emerged along two dominant paradigms: video generation and 3D reconstruction. However, existing evaluation benchmarks either focus narrowly on visual fidelity and text--video alignment for …

3D ReconstructionVideo GenerationVideo Alignment