paper-with-me

홈 › Papers

Large Language Models Are Struggle to Cope with Unreasonability in Math Problems

2024-03-28 · Jingyuan Ma, Damai Dai, Zihang Yuan, Rui Li, Weilin Luo, Bin Wang, Qun Liu, Lei Sha, Zhifang Sui

Recent research have demonstrated LLMs' impressive performance in math and reasoning. However, the capacity of LLMs to address math problems under unconventional conditions, such as internal inconsistencies and flawed assumptions, remains largely unexplored. In this paper, we propose a novel benchmark Unreasonable Math Problem (UMP) designed to assess LLMs' ability to recognize and respond to unreasonability in math problem. The benchmark consists of a carefully curated collection of unreasonable math questions across diverse types. Based on extensive experiments covering 19 LLMs, we observe that even state-of-the-art models such as GPT-4o achieve only limited performance of 0.6 in UMP, while reasoning models such as DeepSeek-R1 are prone to overthinking and unstable. We further explore strategies for improving the recognition of unreasonable inputs, shedding light on both the possibility and limitations of LLMs in this challenging setting.

📄 PDF Abstract BibTeX arXiv:2403.19346

Code (0)

등록된 구현이 없습니다.

Tasks

Math

Similar Papers 제목 키워드 기반

Domain-Specific Knowledge Graphs in RAG-Enhanced Healthcare LLMs

2026-01-21 · Sydney Anuyah, Mehedi Mahmud Kaushik, Hao Dai, Rakesh Shiradkar 외 arxiv

Large Language Models (LLMs) generate fluent answers but can struggle with trustworthy, domain-specific reasoning. We evaluate whether domain knowledge graphs (KGs) improve Retrieval-Augmented Generation (RAG) for health…

Knowledge Graphs

NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models

2026-01-07 · Weiqi Liu, Yongliang Miao, Haiyan Zhao, Yanguang Liu 외 arxiv

Neuron-level interpretation in large language models (LLMs) is fundamentally challenged by widespread polysemanticity, where individual neurons respond to multiple distinct semantic concepts. Existing single-pass interpr…

SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese

2024-01-22 · Liang Xu, Hang Xue, Lei Zhu, Kangkang Zhao

We introduce SuperCLUE-Math6(SC-Math6), a new benchmark dataset to evaluate the mathematical reasoning abilities of Chinese language models. SC-Math6 is designed as an upgraded Chinese version of the GSM8K dataset with e…

DiversityGSM8KMathMathematical Reasoning

AnyEdit: Edit Any Knowledge Encoded in Language Models

2025-02-08 · Houcheng Jiang, Junfeng Fang, Ningyu Zhang, Guojun Ma 외

Large language models (LLMs) often produce incorrect or outdated information, necessitating efficient and precise knowledge updates. Current model editing methods, however, struggle with long-form knowledge in diverse fo…

FormImage Editingknowledge editingModel Editing

SCOPE: Sign Language Contextual Processing with Embedding from LLMs

2024-09-02 · Yuqi Liu, Wenqian Zhang, Sihan Ren, Chengyu Huang 외

Sign languages, used by around 70 million Deaf individuals globally, are visual languages that convey visual and contextual information. Current methods in vision-based sign language recognition (SLR) and translation (SL…

DiversityLanguage ModelingLanguage ModellingLarge Language Model+1