paper-with-me

홈 › Papers

A Verifiable Search Is Not a Learnable Chain-of-Thought

2026-06-20 · Harsh Patel arxiv

It is tempting to assume any task solvable by a short program can be taught to a model as its chain-of-thought: write the steps out, fine-tune, and the model follows. This paper shows the assumption fails for an identifiable class of procedures. The testbed is nine reasoning tasks, each from a deterministic generator; public and hidden splits share generators, so held-out data proxies test accuracy. I reverse-engineer the generators into Python solvers, render them as chain-of-thought, and distill into a rank-<= 32 LoRA over a 30B (3.5B-active) Nemotron model. Forward-computable tasks install readily: lookup/arithmetic and an 8-bit boolean task transfer (>= 0.99 and 0.68). Cryptarithm does not: distilling its backtracking search holds at 0.01-0.07 across eleven chain-of-thought designs, RL from verifiable rewards, and self-training, even though a search solver answers 71% of instances. This is not a capability gap. The model does the arithmetic on 97-100% of lines and ranks the correct cipher in its top eight on 71%; it cannot carry the search forward as a left-to-right derivation. Fine-tuning learns the shape of a verifiable elimination step while its verdicts become unconditional templates, correct only 16-57% of the time ("verdict-as-token"). The ceiling holds across backbones from 3B to 671B and across fine-tuning and prompting; a controlled intervention isolates the cause: revealing the cipher key, which turns the derivation forward, lifts the same instances from 0.03 to 0.57. When a procedure's only solution is search over information-free structure, no faithful forward chain-of-thought exists to imitate. The task becomes learnable only by removing the search, precomputing its combinatorial core into a catalog and reducing the trace to recall plus verification; the 1st-place solution reaches Private LB 0.92 this way. What distills is memorization and verification, not search.

📄 PDF Abstract BibTeX arXiv:2606.21884

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RusFinChain: A Russian Benchmark for Verifiable Chain-of-Thought Reasoning in Finance with Fuzzy-Aligned Evaluation

2026-07-01 · M. K. Arabov arxiv

Multi-step symbolic reasoning is essential for robust financial analysis, yet most benchmarks neglect intermediate reasoning steps. FINCHAIN introduced verifiable Chain-of-Thought (CoT) evaluation but is limited to Engli…

Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base

2025-10-30 · Yu Li, Yuan Huang, Tao Wang, Caiyu Fan 외 arxiv

Most scientific materials compress reasoning, presenting conclusions while omitting the derivational chains that justify them. This compression hinders verification by lacking explicit, step-wise justifications and inhib…

Forethought: Verifiable Reasoning from Neurosymbolic Primitive Programming

2026-07-05 · Vishvesh Bhat, Jay Vaghasiya, Emmanuel Anaya Gonzalez arxiv

Current agentic workflows usually involve decomposing user requests into sequences of tool calls with correctly resolved parameters, the results of which are processed through reasoning traces in the language model's con…

Reinforcement Learning

Eliciting Chain-of-Thought Reasoning for Time Series Analysis using Reinforcement Learning

2025-10-01 · Felix Parker, Nimeesha Chan, Chi Zhang, Kimia Ghobadi arxiv

Complex numerical time series analysis often demands multi-step reasoning capabilities beyond current models' reach. Tasks like medical diagnosis and weather forecasting require sequential reasoning processes - including…

Reinforcement LearningTime Series AnalysisWeather ForecastingMedical Diagnosis

Masked-and-Reordered Self-Supervision for Reinforcement Learning from Verifiable Rewards

2025-11-21 · Zhen Wang, Zhifeng Gao, Guolin Ke arxiv

Test-time scaling has been shown to substantially improve large language models' (LLMs) mathematical reasoning. However, for a large portion of mathematical corpora, especially theorem proving, RLVR's scalability is limi…

Reinforcement LearningMathematical Reasoning