paper-with-me

홈 › Papers

Do Vision Language Models Understand Human Engagement in Games?

2026-03-19 · Ziyi Wang, Qizan Guo, Rishitosh Singh, Xiyang Hu arxiv

Inferring human engagement from gameplay video is important for game design and player-experience research, yet it remains unclear whether vision--language models (VLMs) can infer such latent psychological states from visual cues alone. Using the GameVibe Few-Shot dataset across nine first-person shooter games, we evaluate three VLMs under six prompting strategies, including zero-shot prediction, theory-guided prompts grounded in Flow, GameFlow, Self-Determination Theory, and MDA, and retrieval-augmented prompting. We consider both pointwise engagement prediction and pairwise prediction of engagement change between consecutive windows. Results show that zero-shot VLM predictions are generally weak and often fail to outperform simple per-game majority-class baselines. Memory- or retrieval-augmented prompting improves pointwise prediction in some settings, whereas pairwise prediction remains consistently difficult across strategies. Theory-guided prompting alone does not reliably help and can instead reinforce surface-level shortcuts. These findings suggest a perception--understanding gap in current VLMs: although they can recognize visible gameplay cues, they still struggle to robustly infer human engagement across games.

📄 PDF Abstract BibTeX arXiv:2603.18480

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Understanding Virality: A Rubric based Vision-Language Model Framework for Short-Form Edutainment Evaluation

2025-12-24 · Arnav Gupta, Gurekas Singh Sahney, Hardik Rathi, Abhishek Chandwani 외 arxiv

Evaluating short-form video content requires moving beyond surface-level quality metrics toward human-aligned, multimodal reasoning. While existing frameworks like VideoScore-2 assess visual and semantic fidelity, they d…

Multimodal ReasoningFeature Importance

Improving Solvability for Procedurally Generated Challenges in Physical Solitaire Games Through Entangled Components

2018-10-03 · Mark Goadrich, James Droscha

Challenges for physical solitaire puzzle games are typically designed in advance by humans and limited in number. Alternatively, some games incorporate rules for stochastic setup, where the human solver randomly sets up …

Solitaire

Gameplay Highlights Generation

2025-05-12 · Vignesh Edithal, Le Zhang, Ilia Blank, Imran Junejo

In this work, we enable gamers to share their gaming experience on social media by automatically generating eye-catching highlight reels from their gameplay session Our automation will save time for gamers while increasi…

Event DetectionHighlight DetectionOptical Character Recognition (OCR)Prompt Engineering+3

Can Large Language Models Capture Video Game Engagement?

2025-02-05 · David Melhart, Matthew Barthet, Georgios N. Yannakakis

Can out-of-the-box pretrained Large Language Models (LLMs) detect human affect successfully when observing a video? To address this question, for the first time, we evaluate comprehensively the capacity of popular LLMs t…

Inducing Cooperative behaviour in Sequential-Social dilemmas through Multi-Agent Reinforcement Learning using Status-Quo Loss

2020-01-15 · Pinkesh Badjatiya, Mausoom Sarkar, Abhishek Sinha, Siddharth Singh 외

In social dilemma situations, individual rationality leads to sub-optimal group outcomes. Several human engagements can be modeled as a sequential (multi-step) social dilemmas. However, in contrast to humans, Deep Reinfo…

ClusteringDeep Reinforcement LearningMulti-agent Reinforcement LearningReinforcement Learning