paper-with-me

홈 › Papers

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

2025-09-18 · Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Jiyuan Fang, Heyuan Guan, Xingfu Li, Naoya Shiraishi, Yo Nakawake, Le Minh Nguyen arxiv

Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems. We introduce the NazoNazo Benchmark, a renewable and extensible evaluation dataset derived from Japanese children's riddles that isolates a specific class of reasoning processes: insight-like representational restructuring and metacognitive evaluation. Rather than modeling reasoning in general, these tasks provide a focused test of failure modes that are difficult to detect in standard benchmarks. We curate 201 riddles and establish a human reference on a 120-item subset (n = 126; mean accuracy 52.9%). The benchmark is fully open, low-cost to refresh, and designed for continual evaluation under reduced contamination risk. We evaluate 38 frontier LLMs (2023-2025) under a strict retrieval-free, zero-shot protocol. On the human-comparison subset, non-reasoning models achieve 7.6% accuracy and reasoning-oriented models reach 17.6%, compared with a human mean of 52.9%, although performance varies substantially across models. Beyond accuracy, qualitative analysis of model-generated thought-logs identifies a distinctive failure mode, which we call verification failure: models generate a correct intermediate candidate but fail to endorse it as their final answer. This dissociation between candidate generation and endorsement reveals a metacognitive bottleneck: across the models with usable thought-logs, verification failures account for between 5% and 39% of a model's incorrect answers. By isolating the gap between generation and verification, this work provides a practical framework for diagnosing reasoning reliability and suggests concrete directions for improvement, including better calibration, structured verification, and stopping mechanisms.

📄 PDF Abstract BibTeX arXiv:2509.14704

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adaptive Originality Filtering: Rejection Based Prompting and RiddleScore for Culturally Grounded Multilingual Riddle Generation

2025-08-26 · Duy Le, Kent Ziti, Evan Girard-Sun, Bakr Bouhaya 외 arxiv

Language models are increasingly tested on multilingual creativity, demanding culturally grounded, abstract generations. Standard prompting methods often produce repetitive or shallow outputs. We introduce Adaptive Origi…

CC-Riddle: A Question Answering Dataset of Chinese Character Riddles

2022-06-28 · Fan Xu, Yunxiang Zhang, Xiaojun Wan

The Chinese character riddle is a unique form of cultural entertainment specific to the Chinese language. It typically comprises two parts: the riddle description and the solution. The solution to the riddle is a single …

General KnowledgeLanguage ModellingMultiple-choiceQuestion Answering

Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models

2024-07-28 · Nitzan Bitton-Guetta, Aviv Slobodkin, Aviya Maimon, Eliya Habba 외

Imagine observing someone scratching their arm; to understand why, additional context would be necessary. However, spotting a mosquito nearby would immediately offer a likely explanation for the person's discomfort, ther…

World Knowledge

Automatically generating Riddles aiding Concept Attainment

2023-10-27 · Niharika Sri Parasa, Chaitali Diwan, Srinath Srinivasa

One of the primary challenges in online learning environments, is to retain learner engagement. Several different instructional strategies are proposed both in online and offline environments to enhance learner engagemen…

Riddle Generation using Word Associations

2016-05-01 · LREC 2016 5 · Paloma Galv{\'a}n, Virginia Francisco, Raquel Herv{\'a}s, Gonzalo M{\'e}ndez

In knowledge bases where concepts have associated properties, there is a large amount of comparative information that is implicitly encoded in the values of the properties these concepts share. Although there have been p…