paper-with-me

홈 › Papers

BlueprintRepair: Typed Local Edits for Failed Lean Proof Blueprints

2026-07-30 · Ruslan Khrulev arxiv

LLM-based Lean proving systems increasingly organize a proof as a blueprint: a dependency graph of formal statements. We introduce BlueprintRepair, a repair interface that lets a model change this graph through ten schema-checked local operations. An operation names the node it edits, so the target theorem cannot be changed. Lean checks every applied change, and an accepted repair must declare every blueprint lemma its proof uses. We also construct BlueprintTrace, a benchmark of 142 controlled failures with complete accepted and rejected repair trajectories. We compare typed edits, exact source patches, and complete module rewrites under matched source, feedback, model, and budget, one episode per state and interface. With DeepSeek-V4-Flash, the three interfaces solve almost the same number of the benchmark's localized failures. Typed repair is the cheapest per solved state (patching is 1.30x as expensive, rewriting 2.06x), and within 10,000 completion tokens per task it reaches almost all of its final coverage, while both free-form interfaces are well behind. A second model, Qwen3.6-Flash, solves fewer states but keeps typed repair cheapest, puts it ahead on the proof-authoring states, and repeats the localized pattern.

📄 PDF Abstract BibTeX arXiv:2607.28110

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PROJECTMEM: A Local-First, Event-Sourced Memory and Judgment Layer for AI Coding Agents

2026-06-10 · Ripon Chandra Malo, Tong Qiu arxiv

AI coding assistants now support a growing share of software work, from quick scripts to production applications. Yet these agents remain largely stateless: each new session re-reads project files, re-derives prior decis…

Kintsugi: Learning Policies by Repairing Executable Knowledge Bases

2026-05-10 · Teng Cao, Yu Deng, Hikaru Shindo, Quentin Delfosse 외 arxiv

Modern embodied agents achieve impressive performance, but their task knowledge is often stored in neural weights, latent state, or prompt-bound memory, making individual policy knowledge difficult to inspect, validate, …

LPDP: Inference-Time Reward Control for Variable-Length DNA Generation with Edit Flows

2026-05-12 · Jeongchan Kim, Yunkyung Ko, Jong Chul Ye arxiv

We study the application of recent Edit Flows for inference-time reward control for DNA sequence generation. Unlike most reward-guided DNA generation frameworks, which operate on fixed-length sequence spaces, Edit Flows …

Learning Explicit Behavioral Models with Adaptive Questions and World-Model Probes

2026-06-05 · Hikaru Shindo, Yu Deng, Teng Cao, Quentin Delfosse 외 arxiv

Interactive agents trained only against task return can achieve high scores while failing to represent the mechanisms that make their actions succeed. This makes brittle behavior difficult to diagnose and limits adaptati…

Question AnsweringCode Repair

Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning

2026-05-27 · Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov 외 arxiv

Lean is increasingly used to judge natural-language mathematical answers, but its signal is partial: many answers never formalize, and a failed proof may reflect an ill-typed statement or a missing library fact, not a wr…

Mathematical Reasoning