paper-with-me

홈 › Papers

VERGE: Formal Refinement and Guidance Engine for Verifiable LLM Reasoning

2026-01-27 · Vikash Singh, Darion Cassel, Nathaniel Weir, Nick Feng, Sam Bayless arxiv

Despite the syntactic fluency of Large Language Models (LLMs), ensuring their logical correctness in high-stakes domains remains a fundamental challenge. We present a neurosymbolic framework that combines LLMs with SMT solvers to produce verification-guided answers through iterative refinement. Our approach decomposes LLM outputs into atomic claims, autoformalizes them into first-order logic, and verifies their logical consistency using automated theorem proving. We introduce three key innovations: (1) multi-model consensus via formal semantic equivalence checking to ensure logic-level alignment between candidates, eliminating the syntactic bias of surface-form metrics, (2) semantic routing that directs different claim types to appropriate verification strategies: symbolic solvers for logical claims and LLM ensembles for commonsense reasoning, and (3) precise logical error localization via Minimal Correction Subsets (MCS), which pinpoint the exact subset of claims to revise, transforming binary failure signals into actionable feedback. Our framework classifies claims by their logical status and aggregates multiple verification signals into a unified score with variance-based penalty. The system iteratively refines answers using structured feedback until acceptance criteria are met or convergence is achieved. This hybrid approach delivers formal guarantees where possible and consensus verification elsewhere, advancing trustworthy AI. With the GPT-OSS-120B model, VERGE demonstrates an average performance uplift of 18.7% at convergence across a set of reasoning benchmarks compared to single-pass approaches.

📄 PDF Abstract BibTeX arXiv:2601.20055

Code (0)

등록된 구현이 없습니다.

Tasks

Automated Theorem Proving

Similar Papers 제목 키워드 기반

Uncertainty Reasoning with Large Language Models for Explainable Disease Diagnosis

2026-05-25 · Xiaoyang Fan, Yufan Cai, Zhe Hou, Jin Song Dong arxiv

Clinical decision-making requires reasoning over incomplete, imprecise, and linguistically expressed patient narratives. While large language models (LLMs) excel at extracting latent information from natural language, th…

Medical DiagnosisFormal Logic

Milestones over Outcome: Unlocking Geometric Reasoning with Sub-Goal Verifiable Reward

2026-01-08 · Jianlong Chen, Daocheng Fu, Shengze Xu, Jiawei Chen 외 arxiv

Multimodal Large Language Models (MLLMs) struggle with complex geometric reasoning, largely because "black box" outcome-based supervision fails to distinguish between lucky guesses and rigorous deduction. To address this…

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement

2026-06-16 · You Li, Samuel Mandell, David Z. Pan arxiv

Large language models (LLMs) have achieved remarkable success in software development. However, they are susceptible to hallucinations, meaning that they can introduce subtle semantic and logical errors. Due to the high …

From Defects to Demands: A Unified, Iterative, and Heuristically Guided LLM-Based Framework for Automated Software Repair and Requirement Realization

2024-12-06 · Alex, Liu, Vivian, Chi

This manuscript signals a new era in the integration of artificial intelligence with software engineering, placing machines at the pinnacle of coding capability. We present a formalized, iterative methodology proving tha…

Ingenuitytest driven development

Statistical Proof of Execution (SPEX)

2025-03-24 · Michele Dallachiesa, Antonio Pitasi, David Pinger, Josh Goodbody 외

Many real-world applications are increasingly incorporating automated decision-making, driven by the widespread adoption of ML/AI inference for planning and guidance. This study examines the growing need for verifiable c…

Decision Making