paper-with-me

홈 › Papers

A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks

2025-02-12 · Karahan Sarıtaş, Kıvanç Tezören, Yavuz Durmazkeser

In recent years, evaluating the Theory of Mind (ToM) capabilities of large language models (LLMs) has received significant attention within the research community. As the field rapidly evolves, navigating the diverse approaches and methodologies has become increasingly complex. This systematic review synthesizes current efforts to assess LLMs' ability to perform ToM tasks, an essential aspect of human cognition involving the attribution of mental states to oneself and others. Despite notable advancements, the proficiency of LLMs in ToM remains a contentious issue. By categorizing benchmarks and tasks through a taxonomy rooted in cognitive science, this review critically examines evaluation techniques, prompting strategies, and the inherent limitations of LLMs in replicating human-like mental state reasoning. A recurring theme in the literature reveals that while LLMs demonstrate emerging competence in ToM tasks, significant gaps persist in their emulation of human cognitive abilities.

📄 PDF Abstract BibTeX arXiv:2502.08796

Code (1)

mars-tin/awesome-theory-of-mind 공식 구현

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

A Unified Study of LoRA Variants: Taxonomy, Review, Codebase, and Empirical Evaluation

2026-01-30 · Haonan He, Jingqi Ye, Minglei Li, Zhengbo Wang 외 arxiv

Low-Rank Adaptation (LoRA) is a fundamental parameter-efficient fine-tuning method that balances efficiency and performance in large-scale neural networks. However, the proliferation of LoRA variants has led to fragmenta…

parameter-efficient fine-tuningNatural Language UnderstandingImage Classification

A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations

2024-07-04 · Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman 외

Large Language Models (LLMs) have recently gained significant attention due to their remarkable capabilities in performing diverse tasks across various domains. However, a thorough evaluation of these models is crucial b…

Zero-shot Generative Large Language Models for Systematic Review Screening Automation

2024-01-12 · Shuai Wang, Harrisen Scells, Shengyao Zhuang, Martin Potthast 외

Systematic reviews are crucial for evidence-based medicine as they comprehensively analyse published research findings on specific questions. Conducting such reviews is often resource- and time-intensive, especially in t…

Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review

2025-09-27 · Sydney Peters, Nan Zhang, Hong Jiao, Ming Li 외 arxiv

Item difficulty plays a crucial role in test performance, interpretability of scores, and equity for all test-takers, especially in large-scale assessments. Traditional approaches to item difficulty modeling rely on fiel…

Feature Engineering

Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi

2025-04-28 · Dandan Chen Kaptur, Yue Huang, Xuejun Ryan Ji, Yanhui Guo 외

This research delved into GPT-4 and Kimi, two Large Language Models (LLMs), for systematic reviews. We evaluated their performance by comparing LLM-generated codes with human-generated codes from a peer-reviewed systemat…