paper-with-me

홈 › Papers

Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities

2026-05-10 · Ryan Albright, Golam Md Muktadir, Zarif Ikram, S M Jubaer, Mehrab Hossain, Dianbo Liu arxiv

While extremely powerful and versatile at various tasks, the thinking capabilities of large language models (LLMs) are often put under scrutiny as they sometimes fail to solve problems that humans can systematically solve. However, recent literature focuses on breaking LLM reasoning with increasingly complex problems, and whether an LLM is robust in simple logical reasoning remains underexplored. This paper proposes Absurd World, a benchmarking framework, to test LLMs against altered realism, where scenarios are logically coherent, and humans can easily solve the tasks. Absurd World breaks a real-world model into symbols, actions, sequences, and events, which are automatically altered to create absurd worlds where the logic to solve the tasks remains the same. It evaluates a large collection of models with simple and advanced prompting techniques, and proves that it is an effective tool to determine LLMs' ability to think logically, ignoring the patterns learned from the real world. One can use this framework to extensively test an LLM against a real-world problem to verify whether the LLM's reasoning capability is robust against variations of the task.

📄 PDF Abstract BibTeX arXiv:2605.09678

Code (0)

등록된 구현이 없습니다.

Tasks

Logical Reasoning

Similar Papers 제목 키워드 기반

True or False: Does the Deep Learning Model Learn to Detect Rumors?

2021-12-01 · Shiwen Ni, Jiawen Li, Hung-Yu Kao

It is difficult for humans to distinguish the true and false of rumors, but current deep learning models can surpass humans and achieve excellent accuracy on many rumor datasets. In this paper, we investigate whether dee…

Common Sense Reasoning

The Quasi-Creature and the Uncanny Valley of Agency: A Synthesis of Theory and Evidence on User Interaction with Inconsistent Generative AI

2025-08-25 · Mauricio Manhaes, Christine Miller, Nicholas Schroeder arxiv

The user experience with large-scale generative AI is paradoxical: superhuman fluency meets absurd failures in common sense and consistency. This paper argues that the resulting potent frustration is an ontological probl…

The Idola Tribus of AI: Large Language Models tend to perceive order where none exists

2025-10-10 · Shin-nosuke Ishikawa, Masato Todo, Taiki Ogihara, Hirotsugu Ohba arxiv

We present a tendency of large language models (LLMs) to generate absurd patterns despite their clear inappropriateness in a simple task of identifying regularities in number series. Several approaches have been proposed…

Logical Reasoning

When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination

2026-09-09 · Karan Parekh, Sanjana Pendyala Ravinder, Sana Mhapsekar, Medina Maloku arxiv

Large language models are increasingly proposed as automated auditors of document quality, yet their reliability as detectors of planted errors is poorly characterised. We construct a contaminated corpus of 150 academic …

A Multi-Modal Method for Satire Detection using Textual and Visual Cues

2020-10-13 · NLP4IF (COLING) 2020 12 · Lily Li, Or Levi, Pedram Hosseini, David A. Broniatowski

Satire is a form of humorous critique, but it is sometimes misinterpreted by readers as legitimate news, which can lead to harmful consequences. We observe that the images used in satirical news articles often contain ab…

ArticlesImage ForensicsImage ManipulationSatire Detection