paper-with-me

홈 › Papers

Learning to chain-of-thought with Jensen's evidence lower bound

2025-03-25 · Yunhao Tang, Sid Wang, Rémi Munos

We propose a way to optimize chain-of-thought with reinforcement learning, but without external reward function. Our algorithm relies on viewing chain-of-thought as latent variable as part of a probabilistic inference problem. Contrary to the full evidence lower bound, we propose to apply a much simpler Jensen's lower bound, which derives tractable objectives with simple algorithmic components (e.g., without the need for parametric approximate posterior), making it more conducive to modern large-scale training. The lower bound approach naturally interpolates other methods such as supervised fine-tuning and online reinforcement learning, whose practical trade-offs we will illustrate. Finally, we show that on mathematical reasoning problems, optimizing with Jensen's lower bound is as effective as policy gradient with external reward. Taken together, our results showcase as a proof of concept to this new algorithmic paradigm's potential to more generic applications.

📄 PDF Abstract BibTeX arXiv:2503.19618

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoningreinforcement-learningReinforcement Learning

Similar Papers 제목 키워드 기반

Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers

2025-02-04 · Alireza Amiri, Xinting Huang, Mark Rofin, Michael Hahn

Chain-of-thought reasoning and scratchpads have emerged as critical tools for enhancing the computational capabilities of transformers. While theoretical results show that polynomial-length scratchpads can extend transfo…

Hard Attention

Connecting Jensen-Shannon and Kullback-Leibler Divergences: A New Bound for Representation Learning

2025-10-23 · Reuben Dorent, Polina Golland, William Wells arxiv

Mutual Information (MI) is a fundamental measure of statistical dependence widely used in representation learning. While direct optimization of MI via its definition as a Kullback-Leibler divergence (KLD) is often intrac…

Representation Learning

Quantifying the Necessity of Chain of Thought through Opaque Serial Depth

2026-03-10 · Jonah Brown-Cohen, David Lindner, Rohin Shah arxiv

Large language models (LLMs) tend to externalize their reasoning in their chain of thought, making the chain of thought a good target for monitoring. This is partially an inherent feature of the Transformer architecture:…

Tight Sample Complexity of Transformers

2026-06-08 · Chenxiao Yang, Nathan Srebro, Zhiyuan Li arxiv

We tightly characterize the VC dimension of depth-$L$ Transformers with a total of $W$ parameters, mapping an input sequence of length $T$ to a single output, establishing an upper bound of $O(L W \log (T W))$ and a near…

COTCAgent: Preventive Consultation via Probabilistic Chain-of-Thought Completion

2026-05-14 · Zihan Deng, Xiaozhen Zhong, Chuanzhi Xu arxiv

As large language models empower healthcare, intelligent clinical decision support has developed rapidly. Longitudinal electronic health records (EHR) provide essential temporal evidence for accurate clinical diagnosis a…