paper-with-me

홈 › Papers

HappyWorld-Bench

2026-09-21 · Zhiqi Bai, Junai Cai, Yixin Chen, Jingrun Du, Tao Feng, Wei Gong, Siyuan Huang, Xiao Lin, Jiaheng Liu, Jun Luo, Yongzhe Lyu, Liya Ma, Zenan Meng, Lin Qu, Wenbo Su, Jiaming Wang, Qinghe Wang, Shaofei Wang, Yanghai Wang, Zequn Wang, Ziming Wang, Hu Wei, Jiangtao Wu, Ruiqi Wu, Jiaxin Xie, Yuchi Xu, Ze Xu, Chengting Yu, Liangyu Yuan, Gang Zeng, Yawen Zeng, Xingyao Zhang, Zizheng Zhang, Bo Zheng, Jiancheng Zhu, Song-Chun Zhu hf

Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, instantiated across three independent evaluation tracks: video world models, spatial world models, and embodied world models. HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. Across all three tracks, we build and operate HappyWorld-Arena to organize human A/B comparisons and derive model-level Elo ratings, which complement newly designed automated metrics that capture behavioral correctness. We evaluate 14 video world models, 9 spatial systems, and 8 embodied candidates under this unified framework. Results reveal remaining reliability gaps across all three tracks: video models exhibit reduced consistency during extended rollouts and revisits, spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution, and embodied models struggle to preserve state across multi-step actions and respond precisely to altered action conditions and physical rules. These findings highlight the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.

📄 PDF Abstract BibTeX arXiv:2609.24308

Code (1)

BaiShuanghao/my_arXiv_daily ★ 213

Similar Papers 제목 키워드 기반

Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench

2024-07-18 · Yotam Perlitz, Ariel Gera, Ofir Arviv, Asaf Yehudai 외

Recent advancements in Language Models (LMs) have catalyzed the creation of multiple benchmarks, designed to assess these models' general capabilities. A crucial task, however, is assessing the validity of the benchmarks…

Language Modelling

BENCHIP: Benchmarking Intelligence Processors

2017-10-23 · Jinhua Tao, Zidong Du, Qi Guo, Huiying Lan 외

The increasing attention on deep learning has tremendously spurred the design of intelligence processing hardware. The variety of emerging intelligence processors requires standard benchmarks for fair comparison and syst…

BenchmarkingDiversity

Omnibenchmark (alpha) for continuous and open benchmarking in bioinformatics

2024-09-25 · Izaskun Mallona, Almut Luetge, Ben Carrillo, Daniel Incicau 외

Benchmarking in bioinformatics is a process of designing, running and disseminating rigorous performance evaluations of methods (software). Benchmarking systems facilitate the benchmarking process by providing an entrypo…

Benchmarking

Mapping global dynamics of benchmark creation and saturation in artificial intelligence

2022-03-09 · Simon Ott, Adriano Barbosa-Silva, Kathrin Blagec, Jan Brauner 외

Benchmarks are crucial to measuring and steering progress in artificial intelligence (AI). However, recent studies raised concerns over the state of AI benchmarking, reporting issues such as benchmark overfitting, benchm…

Benchmarking

EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models

2023-12-11 · Samuel J. Paech

We introduce EQ-Bench, a novel benchmark designed to evaluate aspects of emotional intelligence in Large Language Models (LLMs). We assess the ability of LLMs to understand complex emotions and social interactions by ask…

BenchmarkingEmotional IntelligenceMMLU