paper-with-me

홈 › Papers

SCOP: Evaluating the Comprehension Process of Large Language Models from a Cognitive View

2025-06-05 · Yongjie Xiao, Hongru Liang, Peixin Qin, Yao Zhang, Wenqiang Lei

Despite the great potential of large language models(LLMs) in machine comprehension, it is still disturbing to fully count on them in real-world scenarios. This is probably because there is no rational explanation for whether the comprehension process of LLMs is aligned with that of experts. In this paper, we propose SCOP to carefully examine how LLMs perform during the comprehension process from a cognitive view. Specifically, it is equipped with a systematical definition of five requisite skills during the comprehension process, a strict framework to construct testing data for these skills, and a detailed analysis of advanced open-sourced and closed-sourced LLMs using the testing data. With SCOP, we find that it is still challenging for LLMs to perform an expert-level comprehension process. Even so, we notice that LLMs share some similarities with experts, e.g., performing better at comprehending local information than global information. Further analysis reveals that LLMs can be somewhat unreliable -- they might reach correct answers through flawed comprehension processes. Based on SCOP, we suggest that one direction for improving LLMs is to focus more on the comprehension process, ensuring all comprehension skills are thoroughly developed during training.

📄 PDF Abstract BibTeX arXiv:2506.05000

Code (0)

등록된 구현이 없습니다.

Tasks

Reading Comprehension

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding

2025-07-18 · Wan-Cyuan Fan, Yen-Chun Chen, Mengchen Liu, Alexander Jacobson 외 arxiv

Recent methods for customizing Large Vision Language Models (LVLMs) for domain-specific tasks have shown promising results in scientific chart comprehension. However, existing approaches face two major limitations: First…

MenatQA: A New Dataset for Testing the Temporal Comprehension and Reasoning Abilities of Large Language Models

2023-10-08 · Yifan Wei, Yisong Su, Huanhuan Ma, Xiaoyan Yu 외

Large language models (LLMs) have shown nearly saturated performance on many natural language processing (NLP) tasks. As a result, it is natural for people to believe that LLMs have also mastered abilities such as time u…

counterfactual

When Do Discourse Markers Affect Computational Sentence Understanding?

2023-09-01 · RuiQi Li, Liesbeth Allein, Damien Sileo, Marie-Francine Moens

The capabilities and use cases of automatic natural language processing (NLP) have grown significantly over the last few years. While much work has been devoted to understanding how humans deal with discourse connectives…

Sentence

Sentence Extraction-Based Machine Reading Comprehension for Vietnamese

2021-05-19 · Phong Nguyen-Thuan Do, Nhat Duy Nguyen, Tin Van Huynh, Kiet Van Nguyen 외

The development of natural language processing (NLP) in general and machine reading comprehension in particular has attracted the great attention of the research community. In recent years, there are a few datasets for m…

ArticlesMachine Reading ComprehensionQuestion AnsweringReading Comprehension+2

Attribution and Alignment: Effects of Local Context Repetition on Utterance Production and Comprehension in Dialogue

2023-11-21 · Aron Molnar, Jaap Jumelet, Mario Giulianelli, Arabella Sinclair

Language models are often used as the backbone of modern dialogue systems. These models are pre-trained on large amounts of written fluent language. Repetition is typically penalised when evaluating language model genera…

Dialogue GenerationLanguage ModelingLanguage Modelling