paper-with-me

홈 › Papers

ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning

2025-02-22 · Shulin Huang, Linyi Yang, Yan Song, Shuang Chen, Leyang Cui, Ziyu Wan, Qingcheng Zeng, Ying Wen, Kun Shao, Weinan Zhang, Jun Wang, Yue Zhang

Evaluating large language models (LLMs) poses significant challenges, particularly due to issues of data contamination and the leakage of correct answers. To address these challenges, we introduce ThinkBench, a novel evaluation framework designed to evaluate LLMs' reasoning capability robustly. ThinkBench proposes a dynamic data generation method for constructing out-of-distribution (OOD) datasets and offers an OOD dataset that contains 2,912 samples drawn from reasoning tasks. ThinkBench unifies the evaluation of reasoning models and non-reasoning models. We evaluate 16 LLMs and 4 PRMs under identical experimental conditions and show that most of the LLMs' performance are far from robust and they face a certain level of data leakage. By dynamically generating OOD datasets, ThinkBench effectively provides a reliable evaluation of LLMs and reduces the impact of data contamination.

📄 PDF Abstract BibTeX arXiv:2502.16268

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLMThinkBench: Towards Basic Math Reasoning and Overthinking in Large Language Models

2025-07-05 · Gaurav Srivastava, Aafiya Hussain, Sriram Srinivasan, Xuan Wang

Large Language Models (LLMs) have achieved remarkable performance on complex mathematical benchmarks, yet often struggle with simple arithmetic tasks and exhibit a tendency toward over-explaining or "overthinking" answer…

BenchmarkingGPUMath

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

2025-11-06 · Jingqi Tong, Yurong Mou, Hangcheng Li, Mingzhe Li 외 arxiv

The "Thinking with Text" and "Thinking with Images" paradigms significantly improve the reasoning abilities of large language models (LLMs) and Vision-Language Models (VLMs). However, these paradigms have inherent limita…

Multimodal ReasoningVideo Generation

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

2026-07-29 · Siyu Yan, Zhuoran Yan, Haiying Xu, Panhao Zhou 외 arxiv

Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are lim…

Visual Reasoning

Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge

2026-05-11 · Wenbo Zhang, Lijinghua Zhang, Liner Xiang, Hengrui Cai arxiv

Reasoning-capable large language models (LLMs) have recently been adopted as automated judges, but their benefits and costs in LLM-as-a-Judge settings remain unclear. Through controlled comparisons between reasoning and …

RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling

2025-06-10 · Yang Liu, Jiaqi Li, Zilong Zheng

Rule-based reasoning has been acknowledged as one of the fundamental problems in reasoning, while deviations in rule formats, types, and complexity in real-world applications pose severe challenges. Recent studies have s…

Computational EfficiencyReinforcement Learning (RL)