paper-with-me

홈 › Papers

Reliable Chain-of-Thought via Prefix Consistency

2026-05-08 · Naoto Iwase, Yuki Ichihara, Mohammad Atif Quamar, Junpei Komiyama arxiv

Large Language Models often improve accuracy on reasoning tasks by sampling multiple Chain-of-Thought (CoT) traces and aggregating them with majority voting (MV), a test-time technique called self-consistency. When we truncate a CoT partway through and regenerate the remainder, we observe that traces with correct answers reproduce their original answer more often than traces with wrong answers. We use this difference as a reliability signal, prefix consistency, that weights each candidate answer by how often it reappears under regeneration. It requires no access to token log-probabilities or self-rating prompts. Across five reasoning models and four math and science benchmarks, prefix consistency is the best correctness predictor in most settings, and reweighting votes by it reaches Standard MV plateau accuracy at up to 21x fewer tokens (median 4.6x). Our code is available at https://github.com/naoto-iwase/prefix-consistency.

📄 PDF Abstract BibTeX arXiv:2605.07654

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning

2026-04-12 · Yuxi Sun, Aoqi Zuo, Haotian Xie, Wei Gao 외 arxiv

Chain-of-Thought (CoT) prompting has improved LLM reasoning, but models often generate explanations that appear coherent while containing unfaithful intermediate steps. Existing self-evaluation approaches are prone to in…

Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models

2026-05-27 · Xinyuan Cheng, Beiduo Chen, Philipp Mondorf, Barbara Plank arxiv

Large reasoning models (LRMs) often generate extensive chain-of-thought (CoT) traces before producing a final answer. As explicit textual artifacts, these traces can be passed to other models to solve the same task, enab…

Self-Consistency Improves Chain of Thought Reasoning in Language Models

2022-03-21 · Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le 외

Chain-of-thought prompting combined with pre-trained large language models has achieved encouraging results on complex reasoning tasks. In this paper, we propose a new decoding strategy, self-consistency, to replace the …

ARCArithmetic ReasoningGSM8KLanguage Modelling+2

Streaming Hallucination Detection in Long Chain-of-Thought Reasoning

2026-01-05 · Haolang Lu, Minghui Pan, Ripeng Li, Guoshun Nan 외 arxiv

Long chain-of-thought (CoT) reasoning improves the performance of large language models, yet hallucinations in such settings often emerge subtly and propagate across reasoning steps. We suggest that hallucination in long…

The Cost of Reasoning: Chain-of-Thought Induces Overconfidence in Vision-Language Models

2026-03-17 · Robert Welch, Emir Konuk, Kevin Smith arxiv

Vision-language models (VLMs) are increasingly deployed in high-stakes settings where reliable uncertainty quantification (UQ) is as important as predictive accuracy. Extended reasoning via chain-of-thought (CoT) prompti…