paper-with-me

Papers

CogDevelop2K: Reversed Cognitive Development in Multimodal Large Language Models

2024-10-06 · Yijiang Li, Qingying Gao, Haoran Sun, Haiyun Lyu, Dezhi Luo, Hokin Deng

Are Multi-modal Large Language Models (MLLMs) stochastic parrots? Do they genuinely understand? This paper aims to explore the core cognitive abilities that human intelligence builds upon to perceive, comprehend, and reason in MLLMs. To this end, we propose CogDevelop2K, a comprehensive benchmark that spans 12 sub-concepts from primitive knowledge like object permanence and boundary to more complex abilities like intentionality understanding, structured via the developmental trajectory of a human mind. We evaluate 46 MLLMs on our benchmarks. Surprisingly, we observe a reversed cognitive developmental trajectory compared to humans. Comprehensively, we further evaluate the influence of evaluation strategies and prompting techniques. Website with this $\href{https://growing-ai-like-a-child.github.io/}{link}$.

📄 PDF Abstract BibTeX arXiv:2410.10855

Code (1)

siquanhuang/Multi-metrics_against_backdoors_in_FL pytorch

Similar Papers 제목 키워드 기반

Failures in Perspective-taking of Multimodal AI Systems

2024-09-20 · Bridget Leonard, Kristin Woodard, Scott O. Murray

This study extends previous research on spatial representations in multimodal AI systems. Although current models demonstrate a rich understanding of spatial information from images, this information is rooted in proposi…

Blocks as Probes: Dissecting Categorization Ability of Large Multimodal Models

2024-09-03 · Bin Fu, Qiyang Wan, Jialin Li, Ruiping Wang 외

Categorization, a core cognitive ability in humans that organizes objects based on common features, is essential to cognitive science as well as computer vision. To evaluate the categorization ability of visual AI models…

Question AnsweringVisual Question Answering

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice

2026-05-02 · Rongyang Wang, Shuang Zhou, Jiashuo Wang, Wenya Xie 외 arxiv

Multimodal large language models (MLLMs) have emerged as a promising paradigm for dental image analysis. However, their ability to capture the multi-level cognitive processes required for radiographic analysis remains un…

Multimodal Grounding for Language Processing

2018-06-17 · COLING 2018 8 · Lisa Beinborn, Teresa Botschen, Iryna Gurevych

This survey discusses how recent developments in multimodal processing facilitate conceptual grounding of language. We categorize the information flow in multimodal processing with respect to cognitive models of human in…

Survey

ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning

2025-09-26 · Jasin Cekinmez, Omid Ghahroodi, Saad Fowad Chandle, Dhiman Gupta 외 arxiv

We introduce ADAM (A Diverse Archive of Mankind), a framework for evaluating and improving multimodal large language models (MLLMs) in biographical reasoning. To the best of our knowledge, this is the first work to syste…