paper-with-me

Papers

Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs

2026-08-28 · Vishvesh Bhat arxiv

Post training a language model to reason means updating its weights. Supervised finetuning and reinforcement learning both place the acquired capability inside the model where it cannot be inspected cannot be checked step by step and cannot be moved to another model. We argue that for tasks whose intermediate steps admit verification, reasoning is better placed outside the base models weights as an explicit program composed from deterministic and neural primitives. We introduce PLVR (Program Learning with Verifiable Rewards): a post training method that learns such programs directly from input-output examples. Its mechanism is symbolic backpropagation: each program layer carries a typed ontology a loss is computed at the output against ground truth and required input ontologies are propagated backward by type inference over primitive signatures: an analogue of the chain rule in which credit assignment is a derivation rather than an estimate. Where RLVR verifies a terminal outcome, PLVRs reward is a per step contract verdict dense over program structure. On LiveCodeBench v6 and Tau2Bench, 30B base models with PLVR outperform RL at matched budget by 27.8 points on average and frontier models an order of magnitude larger by 13.6 points. A single primitive library serves two benchmarks, so the marginal cost of a new task is 100 examples of program search and no new finetuning data. Replacing the loss guided search with uniform sampling over the same type admissible space at equal budget collapses the median program from 65.6 to 17.5, identifying the backward pass rather than the type system as the source of the advantage. We release the symbolic backpropagation library and a conformance checker so the method can be applied to primitive libraries other than our own.

📄 PDF Abstract BibTeX arXiv:2608.28421

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Forethought: Verifiable Reasoning from Neurosymbolic Primitive Programming

2026-07-05 · Vishvesh Bhat, Jay Vaghasiya, Emmanuel Anaya Gonzalez arxiv

Current agentic workflows usually involve decomposing user requests into sequences of tool calls with correctly resolved parameters, the results of which are processed through reasoning traces in the language model's con…

Reinforcement Learning

Incentivizing Vision Language Models to Search for Long Video Question Answering

2026-07-03 · Harsh Goel, S P Sharan, Sahil Shah, Minkyu Choi 외 arxiv

We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process. VSeek utilizes a natural language-driven sear…

Video Question AnsweringReinforcement Learning

Verifiable Process Rewards for Agentic Reasoning

2026-05-11 · Huining Yuan, Zelai Xu, Huaijie Wang, Xiangmin Yi 외 arxiv

Reinforcement learning from verifiable rewards (RLVR) has improved the reasoning abilities of large language models (LLMs), but most existing approaches rely on sparse outcome-level feedback. This sparsity creates a cred…

Reinforcement LearningLogical Reasoning

Symbolic Rule Extraction from Attention-Guided Sparse Representations in Vision Transformers

2025-05-10 · Parth Padalkar, Gopal Gupta

Recent neuro-symbolic approaches have successfully extracted symbolic rule-sets from CNN-based models to enhance interpretability. However, applying similar techniques to Vision Transformers (ViTs) remains challenging du…

DiscoverDCP: A Data-Driven Approach for Construction of Disciplined Convex Programs via Symbolic Regression

2025-12-03 · Sveinung Myhre arxiv

We propose DiscoverDCP, a data-driven framework that integrates symbolic regression with the rule sets of Disciplined Convex Programming (DCP) to perform system identification. By enforcing that all discovered candidate …