paper-with-me

Papers

Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests

2025-02-20 · Filippo Momentè, Alessandro Suglia, Mario Giulianelli, Ambra Ferrari, Alexander Koller, Oliver Lemon, David Schlangen, Raquel Fernández, Raffaella Bernardi

We examine three evaluation paradigms: large question-answering benchmarks (e.g., MMLU and BBH), interactive games (e.g., Signalling Games or Taboo), and cognitive tests (e.g., for working memory or theory of mind). First, we investigate which of the former two-benchmarks or games-is most effective at discriminating LLMs of varying quality. Then, inspired by human cognitive assessments, we compile a suite of targeted tests that measure cognitive abilities deemed essential for effective language use, and we investigate their correlation with model performance in benchmarks and games. Our analyses reveal that interactive games are superior to standard benchmarks in discriminating models. Causal and logical reasoning correlate with both static and interactive tests, while differences emerge regarding core executive functions and social/emotional skills, which correlate more with games. We advocate the development of new interactive benchmarks and targeted cognitive tasks inspired by assessing human abilities but designed specifically for LLMs.

📄 PDF Abstract BibTeX arXiv:2502.14359

Code (0)

등록된 구현이 없습니다.

Tasks

Logical ReasoningMMLUQuestion Answering

Similar Papers 제목 키워드 기반

Play to Generalize: Learning to Reason Through Game Play

2025-06-09 · Yunfei Xie, Yinsong Ma, Shiyi Lan, Alan Yuille 외

Developing generalizable reasoning capabilities in multimodal large language models (MLLMs) remains challenging. Motivated by cognitive science literature suggesting that gameplay promotes transferable cognitive skills, …

Domain GeneralizationMathMultimodal ReasoningReinforcement Learning (RL)

Show, Don't Tell: Evaluating Large Language Models Beyond Textual Understanding with ChildPlay

2024-07-12 · Gonçalo Hora de Carvalho, Oscar Knap, Robert Pollice

We develop a systematic benchmark set to test the generalization of state-of-the-art large language models on broader problems beyond linguistic tasks and evaluate it on a systematic progression of GPT models (GPT-3.5, G…

Spatial Reasoning

WebGames: Challenging General-Purpose Web-Browsing AI Agents

2025-02-25 · George Thomas, Alex J. Chan, Jikun Kang, Wenqi Wu 외

We introduce WebGames, a comprehensive benchmark suite designed to evaluate general-purpose web-browsing AI agents through a collection of 50+ interactive challenges. These challenges are specifically crafted to be strai…

LETGAMES: An LLM-Powered Gamified Approach to Cognitive Training for Patients with Cognitive Impairment

2026-02-18 · Jingwei Shi, Shengyu Tao, Xinxiang Yin, Chen Huang 외 arxiv

The application of games as a therapeutic tool for cognitive training is beneficial for patients with cognitive impairments. However, effective game design for individual patient is resource-intensive. To this end, we pr…

Building Machines That Learn and Think Like People

2016-04-01 · Brenden M. Lake, Tomer D. Ullman, Joshua B. Tenenbaum, Samuel J. Gershman

Recent progress in artificial intelligence (AI) has renewed interest in building systems that learn and think like people. Many advances have come from using deep neural networks trained end-to-end in tasks such as objec…

Board GamesObject Recognition