paper-with-me

홈 › Papers

Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities

2025-10-22 · Nishant Balepur, Dang Nguyen, Dayeon Ki arxiv

Multi-modal large language models (MLMs) are often assessed on static, individual benchmarks -- which cannot jointly assess MLM capabilities in a single task -- or rely on human or model pairwise comparisons -- which is highly subjective, expensive, and allows models to exploit superficial shortcuts (e.g., verbosity) to inflate their win-rates. To overcome these issues, we propose game-based evaluations to holistically assess MLM capabilities. Games require multiple abilities for players to win, are inherently competitive, and are governed by fix, objective rules, and makes evaluation more engaging, providing a robust framework to address the aforementioned challenges. We manifest this evaluation specifically through Dixit, a fantasy card game where players must generate captions for a card that trick some, but not all players, into selecting the played card. Our quantitative experiments with five MLMs show Dixit win-rate rankings are perfectly correlated with those on popular MLM benchmarks, while games between human and MLM players in Dixit reveal several differences between agent strategies and areas of improvement for MLM reasoning.

📄 PDF Abstract BibTeX arXiv:2510.19892

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay

2025-10-11 · Yunxiang Mo, Tianshi Zheng, Qing Zong, Jiayu Liu 외 arxiv

Multimodal abductive reasoning--the generation and selection of explanatory hypotheses from partial observations--is a cornerstone of intelligence. Current evaluations of this ability in vision-language models (VLMs) are…

Dixit: Interactive Visual Storytelling via Term Manipulation

2019-03-06 · Chao-Chun Hsu, Yu-Hua Chen, Zi-Yuan Chen, Hsin-Yu Lin 외

In this paper, we introduce Dixit, an interactive visual storytelling system that the user interacts with iteratively to compose a short story for a photo sequence. The user initiates the process by uploading a sequence …

DecoderVisual Storytelling

Creative Captioning: An AI Grand Challenge Based on the Dixit Board Game

2020-09-30 · Maithilee Kunda, Irina Rabkina

We propose a new class of "grand challenge" AI problems that we call creative captioning---generating clever, interesting, or abstract captions for images, as well as understanding such captions. Creative captioning draw…

Common Sense Reasoning

Know your audience: specializing grounded language models with listener subtraction

2022-06-16 · Aaditya K. Singh, David Ding, Andrew Saxe, Felix Hill 외

Effective communication requires adapting to the idiosyncrasies of each communicative context--such as the common ground shared with each partner. Humans demonstrate this ability to specialize to their audience in many c…

Language ModellingLarge Language Model

Beneath the Surface: Investigating LLMs' Capabilities for Communicating with Subtext

2026-04-07 · Kabir Ahuja, Yuxuan Li, Andrew Kyle Lampinen arxiv

Human communication is fundamentally creative, and often makes use of subtext -- implied meaning that goes beyond the literal content of the text. Here, we systematically study whether language models can use subtext in …