paper-with-me

Papers

The Signal-Coverage Matrix: Stratifying Type and Semantic Errors in Statement Autoformalization

2026-06-26 · Chengxiao Dai, Zhaokun Yan, Zhanhui Lin arxiv

Headline type-correctness (TC\%) of LLM autoformalization has climbed from $\sim$53\% to $\sim$76\% in two years, yet this scalar conceals which errors each method resolves. We propose a signal-coverage matrix that crosses the Lean elaborator (pass/fail) with a semantic-equivalence judgment (equivalent/not), sorting every output into one of four cells: true success (TS), type-only (TO), semantic-only (SO), or both fail (BF). On ProofNet\# and MiniF2F-test with DeepSeek V4-Pro across Vanilla, Lean-Retry, Sample-Filter, and Stratified Autoformalization (SAF): (1) the +34 to +36 TS gain across the three elab-feedback methods is $\sim$64\% type-stratum recovery, with SO flat on net (87.5\% of original semantic errors rescued, 8 newly created). (2) The TO-to-TS rate is 23/61 for each method (Wilson 95\% CI [26.6\%, 50.3\%]), and this stratum-level recovery rate predicts $Δ$TS on held-out methods to within 2/186 and renders $Δ$TC linear in the Vanilla elab-fail rate across six (model, dataset) cells ($R^2=0.96$). (3) The two judges disagree by 26 to 37 pp on elab-feedback outputs (vs. 7 pp on Vanilla), with 30 to 56\% of symbolic-judge false negatives traceable to elaborator-forced rewrites. The persistent residual reduces to two gold-formalization errors. TC\% gains should be credited by which cell moved, not by the scalar alone.

📄 PDF Abstract BibTeX arXiv:2606.28013

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SemLoc: Structured Grounding of Free-Form LLM Reasoning for Fault Localization

2026-03-31 · Zhaorui Yang, Haichao Zhu, Qian Zhang, Rajiv Gupta 외 arxiv

Fault localization identifies program locations responsible for observed failures. Existing techniques rank suspicious code using syntactic spectra--signals derived from execution structure such as statement coverage, co…

Grids Often Outperform Implicit Neural Representations

2025-06-10 · Namhoon Kim, Sara Fridovich-Keil

Implicit Neural Representations (INRs) have recently shown impressive results, but their fundamental capacity, implicit biases, and scaling behavior remain poorly understood. We investigate the performance of diverse INR…

DenoisingSuper-Resolution

Deep Survival Analysis

2016-08-06 · Rajesh Ranganath, Adler Perotte, Noémie Elhadad, David Blei

The electronic health record (EHR) provides an unprecedented opportunity to build actionable tools to support physicians at the point of care. In this paper, we investigate survival analysis in the context of EHR data. W…

Survival Analysis

Weak Detection in the Spiked Wigner Model with General Rank

2020-01-16 · Ji Hyung Jung, Hye Won Chung, Ji Oon Lee

We study the statistical decision process of detecting the signal from a `signal+noise' type matrix model with an additive Wigner noise. We propose a hypothesis test based on the linear spectral statistics of the data ma…

Vocal Bursts Type Prediction

Stratifying Reinforcement Learning with Signal Temporal Logic

2026-04-06 · Justin Curry, Alberto Speranzon arxiv

In this paper, we develop a stratification-based semantics for Signal Temporal Logic (STL) in which each atomic predicate is interpreted as a membership test in a stratified space. This perspective reveals a novel corres…

Reinforcement Learning