paper-with-me

Papers

Evaluating Step-by-step Reasoning Traces: A Survey

2025-02-17 · Jinu Lee, Julia Hockenmaier

Step-by-step reasoning is widely used to enhance the reasoning ability of large language models (LLMs) in complex problems. Evaluating the quality of reasoning traces is crucial for understanding and improving LLM reasoning. However, the evaluation criteria remain highly unstandardized, leading to fragmented efforts in developing metrics and meta-evaluation benchmarks. To address this gap, this survey provides a comprehensive overview of step-by-step reasoning evaluation, proposing a taxonomy of evaluation criteria with four top-level categories (groundedness, validity, coherence, and utility). We then categorize metrics based on their implementations, survey which metrics are used for assessing each criterion, and explore whether evaluator models can transfer across different criteria. Finally, we identify key directions for future research.

📄 PDF Abstract BibTeX arXiv:2502.12289

Code (0)

등록된 구현이 없습니다.

Tasks

Survey

Similar Papers 제목 키워드 기반

CrossTrace: A Cross-Domain Dataset of Grounded Scientific Reasoning Traces for Hypothesis Generation

2026-03-30 · Andrew Bouras, OMS-II Research Fellow arxiv

Scientific hypothesis generation is a critical bottleneck in accelerating research, yet existing datasets for training and evaluating hypothesis-generating models are limited to single domains and lack explicit reasoning…

A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages

2025-10-10 · Raoyuan Zhao, Yihong Liu, Hinrich Schütze, Michael A. Hedderich arxiv

Large reasoning models (LRMs) increasingly rely on step-by-step Chain-of-Thought (CoT) reasoning to improve task performance, particularly in high-resource languages such as English. While recent work has examined final-…

Cognitive Profiling of LRMs' Reasoning Traces Using Bloom's Taxonomy

2026-08-24 · Maria-Eleni Zoumpoulidi, Georgios Paraskevopoulos, Alexandros Potamianos arxiv

Large Reasoning Models (LRMs) have revolutionized reasoning in LLMs, and the increasing public availability of reasoning traces creates valuable opportunities to study model behavior not only at the surface level but als…

Unveiling Fine-Grained Visual Traces: Evaluating Multimodal Interleaved Reasoning Chains in Multimodal STEM Tasks

2026-04-21 · Jing Jin, Hao Liu, Yan Bai, Yihang Lou 외 arxiv

Multimodal large language models (MLLMs) have shown promising reasoning abilities, yet evaluating their performance in specialized domains remains challenging. STEM reasoning is a particularly valuable testbed because it…

Multimodal Reasoning

Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models

2026-07-24 · Junlin Fang, Do Nguyen-Thanh, Xiaogang Xu, Zhen Fang 외 hf

Large reasoning models (LRMs) generate long reasoning traces before producing final answers. While these traces may contain useful signals for hallucination detection, harnessing them is non-trivial because long trajecto…