paper-with-me

홈 › Papers

PokemonChat: Auditing ChatGPT for Pokémon Universe Knowledge

2023-06-05 · Laura Cabello, Jiaang Li, Ilias Chalkidis

The recently released ChatGPT model demonstrates unprecedented capabilities in zero-shot question-answering. In this work, we probe ChatGPT for its conversational understanding and introduce a conversational framework (protocol) that can be adopted in future studies. The Pok\'emon universe serves as an ideal testing ground for auditing ChatGPT's reasoning capabilities due to its closed world assumption. After bringing ChatGPT's background knowledge (on the Pok\'emon universe) to light, we test its reasoning process when using these concepts in battle scenarios. We then evaluate its ability to acquire new knowledge and include it in its reasoning process. Our ultimate goal is to assess ChatGPT's ability to generalize, combine features, and to acquire and reason over newly introduced knowledge from human feedback. We find that ChatGPT has prior knowledge of the Pokemon universe, which can reason upon in battle scenarios to a great extent, even when new information is introduced. The model performs better with collaborative feedback and if there is an initial phase of information retrieval, but also hallucinates occasionally and is susceptible to adversarial attacks.

📄 PDF Abstract BibTeX arXiv:2306.03024

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalQuestion AnsweringRetrieval

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Computational Natural Philosophy: A Thread from Presocratics through Turing to ChatGPT

2023-09-22 · Gordana Dodig-Crnkovic

Modern computational natural philosophy conceptualizes the universe in terms of information and computation, establishing a framework for the study of cognition and intelligence. Despite some critiques, this computationa…

Philosophy

ChatGPT-based Investment Portfolio Selection

2023-08-11 · Oleksandr Romanko, Akhilesh Narayan, Roy H. Kwon

In this paper, we explore potential uses of generative AI models, such as ChatGPT, for investment portfolio selection. Trusting investment advice from Generative Pre-Trained Transformer (GPT) models is a challenge due to…

Decision MakingPortfolio Optimization

An Explainable AI Approach to Large Language Model Assisted Causal Model Auditing and Development

2023-12-23 · YanMing Zhang, Brette Fitzgibbon, Dino Garofolo, Akshith Kota 외

Causal networks are widely used in many fields, including epidemiology, social science, medicine, and engineering, to model the complex relationships between variables. While it can be convenient to algorithmically infer…

Causal InferenceEpidemiologyLanguage ModelingLanguage Modelling+2

Challenges of Auditing: Variability in Outputs of Large Language Models for Health

2026-09-15 · Yuan Pu, Yewon Chang, Furong Jia, Xunjian Yin 외 arxiv

People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings. Here, we find systematic differences across access modes. Because …

KnowledgeBerg: Evaluating Systematic Knowledge Coverage and Compositional Reasoning in Large Language Models

2026-04-19 · Xiao Zhang, Qianru Meng, Yongjian Chen, Yumeng Wang 외 arxiv

Many real-world questions appear deceptively simple yet implicitly demand two capabilities: (i) systematic coverage of a bounded knowledge universe and (ii) compositional set-based reasoning over that universe, a phenome…