paper-with-me

Papers

Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair

2026-08-18 · Jesus Salas arxiv

Machine-verifiable workflows produce governance records linking a task contract, model attempt, verifier decision, accepted output, and target origin. We test whether these records can supervise bounded models, consolidating occasional or expensive capability into reliable one-shot execution. On fresh, structure-disjoint PlanBench replanning cases, Qwen3-14B thinking generated 24 plans admitted by the independently authored VAL verifier. Those plans trained the same checkpoint for non-thinking execution, without oracle targets or a stronger teacher. On 80 unopened cases, VAL-accepted plans increased from 1 to 57, with 56 paired gains and zero regressions; thinking reached 30. The adapter was schema-valid on all cases and used approximately 1/56 of thinking's mean latency. The separate paired interface-cure gate did not pass. A matched ablation fixed the source cases, 52-candidate pool, 24-target count, model, recipe, and seed while changing target selection. On 160 new cases, base, schema-selected, model-self-selected, and VAL-selected execution reached 1, 55, 69, and 102 accepted plans. VAL exceeded self-selection by paired net +33 (p=0.0000019647), with gains in both difficulty strata. Independent semantic selection is therefore load-bearing relative to matched alternatives within this band. A complementary Phi stronger-teacher arm raised base Phi-4 from 2 to 51 accepted plans and from 35 to 80 schema-valid outputs. Earlier synthetic experiments establish teachability, cumulative learning, construction robustness, and stopping boundaries. The results support verifier-selected supervision for bounded, machine-checkable capabilities, not arbitrary planning, enterprise validity, or unrestricted self-improvement.

📄 PDF Abstract BibTeX arXiv:2608.18324

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization

2026-04-17 · Yixu Huang, Xinglei Yu, Zhongyu Wei arxiv

Large Language Models (LLMs) excel at code generation but remain heavily reliant on large-scale annotated solutions and verification-based supervision, which constrains scalability and hinders sustained self-improvement.…

Code Generation

Rationale-Aware Answer Verification by Pairwise Self-Evaluation

2024-10-07 · Akira Kawabata, Saku Sugawara

Answer verification identifies correct solutions among candidates generated by large language models (LLMs). Current approaches typically train verifier models by labeling solutions as correct or incorrect based solely o…

ARCStrategyQAvalid

SAPO: Self-Adaptive Process Optimization Makes Small Reasoners Stronger

2026-01-28 · Kaiyuan Chen, Guangmin Zheng, Jin Wang, Xiaobing Zhou 외 arxiv

Existing self-evolution methods overlook the influence of fine-grained reasoning steps, which leads to the reasoner-verifier gap. The computational inefficiency of Monte Carlo (MC) process supervision further exacerbates…

CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-Evolution

2026-03-18 · Teng Pan, Yuchen Yan, Zixuan Wang, Ruiqing Zhang 외 arxiv

Label-free reinforcement learning enables large language models to improve reasoning capabilities without ground-truth supervision, typically by treating majority-voted answers as pseudo-labels. However, we identify a cr…

Reinforcement LearningMathematical Reasoning

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

2026-07-08 · Mingguang Chen, Licheng Wang, Bo Qu arxiv

AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, training on data they generate, and, increasingly, conducting AI research itself. This…