paper-with-me

홈 › Papers

LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition

2025-10-10 · Yushuo Zheng, Zicheng Zhang, Xiongkuo Min, Huiyu Duan, Guangtao Zhai arxiv

Existing benchmarks for large multimodal models (LMMs) often fail to capture their performance in real-time, adversarial environments. We introduce LM Fight Arena (Large Model Fight Arena), a novel framework that evaluates LMMs by pitting them against each other in the classic fighting game Mortal Kombat II, a task requiring rapid visual understanding and tactical, sequential decision-making. In a controlled tournament, we test six leading open- and closed-source models, where each agent operates controlling the same character to ensure a fair comparison. The models are prompted to interpret game frames and state data to select their next actions. Unlike static evaluations, LM Fight Arena provides a fully automated, reproducible, and objective assessment of an LMM's strategic reasoning capabilities in a dynamic setting. This work introduces a challenging and engaging benchmark that bridges the gap between AI evaluation and interactive entertainment.

📄 PDF Abstract BibTeX arXiv:2510.08928

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Skill-Based Differences in Spatio-Temporal Team Behavior in Defence of The Ancients 2

2016-03-24 · Anders Drachen, Matthew Yancey, John Maguire, Derrek Chu 외

Multiplayer Online Battle Arena (MOBA) games are among the most played digital games in the world. In these games, teams of players fight against each other in arena environments, and the gameplay is focused on tactical …

ClusteringDota 2Time SeriesTime Series Analysis+1

EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments

2025-06-09 · Zefang Liu, Yinzhu Quan

We introduce EconWebArena, a benchmark for evaluating autonomous agents on complex, multimodal economic tasks in realistic web environments. The benchmark comprises 360 curated tasks from 82 authoritative websites spanni…

BenchmarkingNavigateVisual Grounding

Introduction to Behavior Algorithms for Fighting Games

2020-07-06 · Ignacio Gajardo, Felipe Besoain, Nicolas A. Barriga

The quality of opponent Artificial Intelligence (AI) in fighting videogames is crucial. Some other game genres can rely on their story or visuals, but fighting games are all about the adversarial experience. In this pape…

Game Arena: Strategic LLM Evaluation in Competitive Environments

2026-09-25 · Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu, Hann Wang 외 hf

We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head mat…

GameArena: Evaluating LLM Reasoning through Live Computer Games

2024-12-09 · Lanxiang Hu, Qiyu Li, Anze Xie, Nan Jiang 외

Evaluating the reasoning abilities of large language models (LLMs) is challenging. Existing benchmarks often depend on static datasets, which are vulnerable to data contamination and may get saturated over time, or on bi…

Chatbot