paper-with-me

홈 › Papers

Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models

2026-09-22 · Xiaoyu Luo, Tao Ren, Wenrui Yu, Xiao Li, Qiongxiu Li, Johannes Bjerva hf

The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to externalize intermediate reasoning. Because these traces may reflect post-hoc rationalization rather than genuine reasoning, we first evaluate against native CoT on open-source models and extend to closed-source frontier models including GPT-6 Astra. We find that the extracted reasoning matches native reasoning performance and substantially outperforms no-reasoning baselines, across competition mathematics, science, and code generation. We then characterize how frontier models structure their intermediate reasoning. Across token efficiency, reasoning-step types, and induced reasoning trees, we identify systematic differences in how models externalize, compress, and organize reasoning. We find that Astra exhibits token-efficient directed reasoning, selecting a correct trajectory earlier, while resolving elementary steps internally and externalizing only crucial reasoning. These findings provide a behavioral lens on frontier-model reasoning beyond benchmark scores.

📄 PDF Abstract BibTeX arXiv:2609.26637

Code (0)

등록된 구현이 없습니다.

Tasks

Code Generation

Similar Papers 제목 키워드 기반

Factorized Asymptotic Bayesian Inference for Factorial Hidden Markov Models

2015-06-26 · Shaohua Li, Ryohei Fujimaki, Chunyan Miao

Factorial hidden Markov models (FHMMs) are powerful tools of modeling sequential data. Learning FHMMs yields a challenging simultaneous model selection issue, i.e., selecting the number of multiple Markov chains and the …

Bayesian InferenceModel Selection

Parsimonious HMMs for Offline Handwritten Chinese Text Recognition

2018-08-13 · Wenchao Wang, Jun Du, Zi-Rui Wang

Recently, hidden Markov models (HMMs) have achieved promising results for offline handwritten Chinese text recognition. However, due to the large vocabulary of Chinese characters with each modeled by a uniform and fixed …

Handwritten Chinese Text Recognition

Parsimonious Bayesian deep networks

2018-05-22 · NeurIPS 2018 12 · Mingyuan Zhou

Combining Bayesian nonparametrics and a forward model selection strategy, we construct parsimonious Bayesian deep networks (PBDNs) that infer capacity-regularized network architectures from the data and require neither c…

Model Selection

Geometric Signatures of Reasoning: A Spectral Perspective on Task Hardness

2026-07-02 · Aria Masoomi, Mahsa Bazzaz, Adel Javanmard, Vahab Mirrokni arxiv

Chain-of-thought (CoT) reasoning enables large language models (LLMs) to solve complex problems by generating intermediate reasoning steps. While much attention has been paid to the length and content of these reasoning …

Mathematical Reasoning

MedMobile: A mobile-sized language model with expert-level clinical capabilities

2024-10-11 · Krithik Vishwanath, Jaden Stryker, Anton Alaykin, Daniel Alexander Alber 외

Language models (LMs) have demonstrated expert-level reasoning and recall abilities in medicine. However, computational costs and privacy concerns are mounting barriers to wide-scale implementation. We introduce a parsim…

Language ModelingLanguage ModellingMedQAQuestion Answering+2