paper-with-me

홈 › Papers

Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO

2025-09-29 · Hongcheng Wang, Yinuo Huang, Sukai Wang, Guanghui Ren, Hao Dong arxiv

Group Relative Policy Optimization (GRPO) trains Chain-of-Thought reasoning with verifiable rewards, but estimating thought-level advantages without value functions often suffers from high variance. Although tree-style branching is used in practice to reduce variance, it lacks a theoretical explanation of why it works and whether it is important or potentially necessary. We study thought-level advantage estimation in GRPO from a variance perspective under a minimal tree-style setting where multiple continuations are sampled for each thought. Using the multivariate delta method, we reveal a sampling-dimension asymmetry. Increasing sampled thoughts ($K$) leaves a strictly positive estimation-variance floor, whereas increasing continuations per thought ($M$) drives the leading-order estimation variance to zero at rate $1/M$. This implies that, within the fixed-temperature GRPO-style estimator without value models studied here, accurate thought-level advantage estimation cannot be achieved by scaling thought sampling alone, making continuation-level branching a principled and potentially necessary mechanism rather than a heuristic. Experiments further provide empirical evidence for its effectiveness and potential necessity, demonstrating improved optimization stability, training efficiency, and final performance not only in math but also across vision domains and under different model architectures and sizes.

📄 PDF Abstract BibTeX arXiv:2509.24494

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Neural Architecture Search by Learning a Hierarchical Search Space

2025-03-27 · Mehraveh Javan Roshtkhari, Matthew Toews, Marco Pedersoli

Monte-Carlo Tree Search (MCTS) is a powerful tool for many non-differentiable search related problems such as adversarial games. However, the performance of such approach highly depends on the order of the nodes that are…

image-classificationImage ClassificationNeural Architecture Search

LEAF: Growing Trees Without Branching for Speech-Aware Large Language Model Post-Training

2026-05-29 · Argyrios Gerogiannis, Yekaterina Yegorova, Mark Hasegawa-Johnson, Venugopal V. Veeravalli arxiv

State-of-the-art GRPO-style methods for speech-aware large language model post-training suffer from coarse credit assignment, broadcasting the same terminal-reward advantage to every token in a response. This ignores use…

Question Answering

Reasoning Topology Matters: Network-of-Thought for Complex Reasoning Tasks

2026-03-21 · Fan Huang arxiv

Existing prompting paradigms structure LLM reasoning in limited topologies: Chain-of-Thought (CoT) produces linear traces, while Tree-of-Thought (ToT) performs branching search. Yet complex reasoning often requires mergi…

Logical Reasoning

ArborKV: Structure-Aware KV Cache Management for Scaling Tree-based LLM Reasoning

2026-05-21 · Yeqiu Chen, Ziyan Liu, Zhenxin Huang, Runquan Gui 외 arxiv

Recent progress in LLM reasoning has increasingly shifted from single-pass generation to explicit search over intermediate reasoning states. Tree-of-Thoughts (ToT) organizes inference to tree-structured search with branc…

Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model

2025-12-09 · Wenjiang Xu, Cindy Wang, Rui Fang, Mingkang Zhang 외 arxiv

World models have emerged as a pivotal component in robot manipulation planning, enabling agents to predict future environmental states and reason about the consequences of actions before execution. While video-generatio…

Robot Manipulation