paper-with-me

홈 › Papers

A little less conversation, a little more action, please: Investigating the physical common-sense of LLMs in a 3D embodied environment

2024-10-30 · Matteo G. Mecattaf, Ben Slater, Marko Tešić, Jonathan Prunty, Konstantinos Voudouris, Lucy G. Cheke

As general-purpose tools, Large Language Models (LLMs) must often reason about everyday physical environments. In a question-and-answer capacity, understanding the interactions of physical objects may be necessary to give appropriate responses. Moreover, LLMs are increasingly used as reasoning engines in agentic systems, designing and controlling their action sequences. The vast majority of research has tackled this issue using static benchmarks, comprised of text or image-based questions about the physical world. However, these benchmarks do not capture the complexity and nuance of real-life physical processes. Here we advocate for a second, relatively unexplored, approach: 'embodying' the LLMs by granting them control of an agent within a 3D environment. We present the first embodied and cognitively meaningful evaluation of physical common-sense reasoning in LLMs. Our framework allows direct comparison of LLMs with other embodied agents, such as those based on Deep Reinforcement Learning, and human and non-human animals. We employ the Animal-AI (AAI) environment, a simulated 3D virtual laboratory, to study physical common-sense reasoning in LLMs. For this, we use the AAI Testbed, a suite of experiments that replicate laboratory studies with non-human animals, to study physical reasoning capabilities including distance estimation, tracking out-of-sight objects, and tool use. We demonstrate that state-of-the-art multi-modal models with no finetuning can complete this style of task, allowing meaningful comparison to the entrants of the 2019 Animal-AI Olympics competition and to human children. Our results show that LLMs are currently outperformed by human children on these tasks. We argue that this approach allows the study of physical reasoning using ecologically valid experiments drawn directly from cognitive science, improving the predictability and reliability of LLMs.

📄 PDF Abstract BibTeX arXiv:2410.23242

Code (1)

kinds-of-intelligence-cfi/llm-aai 공식 구현

Tasks

Common Sense ReasoningDeep Reinforcement Learning

Similar Papers 제목 키워드 기반

Statistical laws and linguistics differ in naturalistic video and fictional conversations

2025-12-19 · Ashley M. A. Fehr, Calla G. Beauregard, Julia Witte Zimmerman, Katie Ekström 외 arxiv

Conversation is a cornerstone of social connection and is linked to well-being outcomes. Conversations vary widely in type with some portion generating complex, dynamic stories. One approach to studying how conversations…

GCDF1: A Goal- and Context- Driven F-Score for Evaluating User Models

2021-11-01 · EANCS 2021 11 · Alexandru Coca, Bo-Hsiang Tseng, Bill Byrne

The evaluation of dialogue systems in interaction with simulated users has been proposed to improve turn-level, corpus-based metrics which can only evaluate test cases encountered in a corpus and cannot measure system’s …

Dialogue EvaluationTask-Oriented Dialogue Systems

Embarrassed to observe: The effects of directive language in brand conversation

2025-08-18 · Andria Andriuzzi, Géraldine Michel arxiv

In social media, marketers attempt to influence consumers by using directive language, that is, expressions designed to get consumers to take action. While the literature has shown that directive messages in advertising …

Effortless Integration of Memory Management into Open-Domain Conversation Systems

2023-05-23 · Eunbi Choi, Kyoung-Woon On, Gunsoo Han, Sungwoong Kim 외

Open-domain conversation systems integrate multiple conversation skills into a single system through a modular approach. One of the limitations of the system, however, is the absence of management capability for external…

Management

“It seemed like an annoying woman”: On the Perception and Ethical Considerations of Affective Language in Text-Based Conversational Agents

2021-11-01 · CoNLL (EMNLP) 2021 11 · Lindsey Vanderlyn, Gianna Weber, Michael Neumann, Dirk Väth 외

Previous research has found that task-oriented conversational agents are perceived more positively by users when they provide information in an empathetic manner compared to a plain, emotionless information exchange. How…

Chatbot