paper-with-me

Papers

VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance

2025-05-21 · Mohammad Reza Taesiri, Abhijay Ghildyal, Saman Zadtootaghaj, Nabajeet Barman, Cor-Paul Bezemer

With video games now generating the highest revenues in the entertainment industry, optimizing game development workflows has become essential for the sector's sustained growth. Recent advancements in Vision-Language Models (VLMs) offer considerable potential to automate and enhance various aspects of game development, particularly Quality Assurance (QA), which remains one of the industry's most labor-intensive processes with limited automation options. To accurately evaluate the performance of VLMs in video game QA tasks and determine their effectiveness in handling real-world scenarios, there is a clear need for standardized benchmarks, as existing benchmarks are insufficient to address the specific requirements of this domain. To bridge this gap, we introduce VideoGameQA-Bench, a comprehensive benchmark that covers a wide array of game QA activities, including visual unit testing, visual regression testing, needle-in-a-haystack tasks, glitch detection, and bug report generation for both images and videos of various games. Code and data are available at: https://asgaardlab.github.io/videogameqa-bench/

📄 PDF Abstract BibTeX arXiv:2505.15952

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VideoGameBench: Can Vision-Language Models complete popular video games?

2025-05-23 · Alex L. Zhang, Thomas L. Griffiths, Karthik R. Narasimhan, Ofir Press

Vision-language models (VLMs) have achieved strong results on coding and math benchmarks that are challenging for humans, yet their ability to perform tasks that come naturally to humans--such as perception, spatial navi…

Math

Frame Sampling Strategies Matter: A Benchmark for small vision language models

2025-09-18 · Marija Brkic, Anas Filali Razzouki, Yannis Tevissen, Khalil Guetari 외 arxiv

Comparing vision language models on videos is particularly complex, as the performances is jointly determined by the model's visual representation capacity and the frame-sampling strategy used to construct the input. Cur…

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments

2025-02-10 · Sankalp Nagaonkar, Augustya Sharma, Ashish Choithani, Ashutosh Trivedi

This paper introduces an open-source benchmark for evaluating Vision-Language Models (VLMs) on Optical Character Recognition (OCR) tasks in dynamic video environments. We present a curated dataset containing 1,477 manual…

BenchmarkingOptical Character RecognitionOptical Character Recognition (OCR)

VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality

2026-08-26 · Longteng Jiang, DanDan Zheng, Qianqian Qiao, Heng Huang 외 arxiv

We introduce VGA-BenchV2, an extended human-aligned benchmark and optimization framework for jointly evaluating and improving video generation quality and aesthetic value. Built upon VGA-Bench, VGA-BenchV2 preserves the …

Reinforcement LearningVideo Generation

T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation

2025-07-24 · Yubin Chen, Xuyang Guo, Zhenmei Shi, Zhao Song 외 arxiv

Text-to-video (T2V) models have shown remarkable performance in generating visually reasonable scenes, while their capability to leverage world knowledge for ensuring semantic consistency and factual accuracy remains lar…

Text-to-Video Generation