paper-with-me

홈 › Papers

Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds

2023-05-24 · Victoria Basmov, Yoav Goldberg, Reut Tsarfaty

We evaluate LLMs' language understanding capacities on simple inference tasks that most humans find trivial. Specifically, we target (i) grammatically-specified entailments, (ii) premises with evidential adverbs of uncertainty, and (iii) monotonicity entailments. We design evaluation sets for these tasks and conduct experiments in both zero-shot and chain-of-thought setups, and with multiple prompts and LLMs. The models exhibit moderate to low performance on these evaluation sets. Subsequent experiments show that embedding the premise in syntactic constructions that should preserve the entailment relations (presupposition triggers) or change them (non-factives), further confuses the models, causing them to either under-predict or over-predict certain entailment labels regardless of the true relation, and often disregarding the nature of the embedding context. Overall these results suggest that, despite LLMs' celebrated language understanding capacity, even the strongest models have blindspots with respect to certain types of entailments, and certain information-packaging structures act as ``blinds'' overshadowing the semantics of the embedded premise.

📄 PDF Abstract BibTeX arXiv:2305.14785

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Language and Thought: The View from LLMs

2025-05-19 · Daniel Rothschild

Daniel Dennett speculated in *Kinds of Minds* 1996: "Perhaps the kind of mind you get when you add language to it is so different from the kind of mind you can have without language that calling them both minds is a mist…

Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles

2025-06-16 · Antara Raaghavi Bhattacharya, Isabel Papadimitriou, Kathryn Davidson, David Alvarez-Melis

Across languages, numeral systems vary widely in how they construct and combine numbers. While humans consistently learn to navigate this diversity, large language models (LLMs) struggle with linguistic-mathematical puzz…

DiversityMathematical ReasoningNavigate

How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs

2025-07-18 · Karin de Langis, Jong Inn Park, Andreas Schramm, Bin Hu 외 arxiv

Large language models (LLMs) exhibit increasingly sophisticated linguistic capabilities, yet the extent to which these behaviors reflect human-like cognition versus advanced pattern recognition remains an open question. …

Human-Level Reasoning: A Comparative Study of Large Language Models on Logical and Abstract Reasoning

2025-10-28 · Benjamin Grando Moreira arxiv

Evaluating reasoning ability in Large Language Models (LLMs) is important for advancing artificial intelligence, as it transcends mere linguistic task performance. It involves understanding whether these models truly und…

Bridging the Gap between Recognition-level Pre-training and Commonsensical Vision-language Tasks

2022-05-01 · CSRR (ACL) 2022 5 · Yue Wan, Yueen Ma, Haoxuan You, Zhecan Wang 외

Large-scale visual-linguistic pre-training aims to capture the generic representations from multimodal features, which are essential for downstream vision-language tasks. Existing methods mostly focus on learning the sem…

DiversityInformativenessType predictionVisual Question Answering (VQA)