paper-with-me

홈 › Papers

BrainBench: Exposing the Commonsense Reasoning Gap in Large Language Models

2026-03-16 · Yuzhe Tang arxiv

Large language models (LLMs) achieve impressive scores on standard benchmarks yet routinely fail questions that any human would answer correctly in seconds. We introduce BrainBench, a benchmark of 100 brainteaser questions spanning 20 carefully designed categories, each targeting a specific commonsense reasoning failure mode in LLMs. Categories range from implicit physical constraints ("Should I walk or drive my rental car to the return lot?") to semantic scope tricks and default assumption hijacks. We evaluate eight frontier models -- four from the Claude family and four from the GPT family -- using a zero-shot protocol with 10 independent runs per question. The best model, Claude Opus 4.6 with extended thinking, achieves only 80.3% accuracy; the worst, GPT-4o, scores 39.7%. Even top-performing models exhibit a 6-16 percentage-point gap between accuracy and consistency, revealing stochastic reasoning. Cross-lingual evaluation in Chinese shows most models degrade by 2-8 percentage points, confirming that these failures reflect reasoning deficits rather than language-specific artifacts. BrainBench provides a fine-grained diagnostic tool for identifying where and why LLMs substitute surface heuristics for genuine commonsense reasoning.

📄 PDF Abstract BibTeX arXiv:2603.14761

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks

2025-11-02 · Zhihao Peng, Cheng Wang, Shengyuan Liu, Zhiying Liang 외 arxiv

Brain imaging analysis is crucial for diagnosing and treating brain disorders, and multimodal large language models (MLLMs) are increasingly supporting it. However, current brain imaging visual question-answering (VQA) b…

Evaluate Confidence Instead of Perplexity for Zero-shot Commonsense Reasoning

2022-08-23 · Letian Peng, Zuchao Li, Hai Zhao

Commonsense reasoning is an appealing topic in natural language processing (NLP) as it plays a fundamental role in supporting the human-like actions of NLP systems. With large-scale language models as the backbone, unsup…

Language ModelingLanguage ModellingQuestion AnsweringUnsupervised Pre-training

Improving Unsupervised Commonsense Reasoning Using Knowledge-Enabled Natural Language Inference

2021-11-01 · Findings (EMNLP) 2021 11 · Canming Huang, Weinan He, Yongmei Liu

Recent methods based on pre-trained language models have shown strong supervised performance on commonsense reasoning. However, they rely on expensive data annotation and time-consuming training. Thus, we focus on unsupe…

Natural Language InferenceTransfer LearningWinowhy

Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models

2023-12-29 · Yuqing Wang, Yun Zhao

The burgeoning interest in Multimodal Large Language Models (MLLMs), such as OpenAI's GPT-4V(ision), has significantly impacted both academic and industrial realms. These models enhance Large Language Models (LLMs) with …

HellaSwag

Generated Knowledge Prompting for Commonsense Reasoning

2021-11-16 · ACL ARR November 2021 11 · Anonymous

It remains an open question whether incorporating external knowledge benefits commonsense reasoning while maintaining the flexibility of pretrained sequence models. To investigate this question, we develop generated know…

Language ModelingLanguage ModellingOpen-Ended Question Answering