Assessing Step-by-Step Reasoning against Lexical Negation: A Case Study on Syllogism
Large language models (LLMs) take advantage of step-by-step reasoning instructions, e.g., chain-of-thought (CoT) prompting. Building on this, their ability to perform CoT-style reasoning robustly is of interest from a probing perspective. In this study, we inspect the step-by-step reasoning ability of LLMs with a focus on negation, which is a core linguistic phenomenon that is difficult to process. In particular, we introduce several controlled settings (e.g., reasoning in case of fictional entities) to evaluate the logical reasoning abilities of the models. We observed that dozens of modern LLMs were not robust against lexical negation (e.g., plausible ->implausible) when performing CoT-style reasoning, and the results highlight unique limitations in each LLM family.
Code (0)
등록된 구현이 없습니다.
Tasks
Logical ReasoningNegationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
First Heuristic Then Rational: Dynamic Use of Heuristics in Language Model Reasoning
Multi-step reasoning instruction, such as chain-of-thought prompting, is widely adopted to explore better language models (LMs) performance. We report on the systematic strategy that LMs employ in such a multi-step reaso…
Language ModelingLanguage ModellingEnhancing Table Reasoning with Deterministic Table-State Rewards
Large Language Models (LLMs) struggle with multi-step reasoning over structured tables. The primary reason is the lack of explicit supervision for intermediate reasoning states. Existing learned reward models or executor…
Text SummarizationEvaluating Step-by-step Reasoning Traces: A Survey
Step-by-step reasoning is widely used to enhance the reasoning ability of large language models (LLMs) in complex problems. Evaluating the quality of reasoning traces is crucial for understanding and improving LLM reason…
SurveyProbing Cross-Modal Representations in Multi-Step Relational Reasoning
We investigate the representations learned by vision and language models in tasks that require relational reasoning. Focusing on the problem of assessing the relative size of objects in abstract visual contexts, we analy…
DiagnosticRelational ReasoningGRACE: Gradient-aligned Reasoning Data Curation for Efficient Post-training
Existing reasoning data curation pipelines score whole samples, treating every intermediate step as equally valuable. In reality, steps within a trace contribute very unevenly, and selecting reasoning data well requires …