paper-with-me

홈 › Papers

VerifiAgent: a Unified Verification Agent in Language Model Reasoning

2025-04-01 · Jiuzhou Han, Wray Buntine, Ehsan Shareghi

Large language models demonstrate remarkable reasoning capabilities but often produce unreliable or incorrect responses. Existing verification methods are typically model-specific or domain-restricted, requiring significant computational resources and lacking scalability across diverse reasoning tasks. To address these limitations, we propose VerifiAgent, a unified verification agent that integrates two levels of verification: meta-verification, which assesses completeness and consistency in model responses, and tool-based adaptive verification, where VerifiAgent autonomously selects appropriate verification tools based on the reasoning type, including mathematical, logical, or commonsense reasoning. This adaptive approach ensures both efficiency and robustness across different verification scenarios. Experimental results show that VerifiAgent outperforms baseline verification methods (e.g., deductive verifier, backward verifier) among all reasoning tasks. Additionally, it can further enhance reasoning accuracy by leveraging feedback from verification results. VerifiAgent can also be effectively applied to inference scaling, achieving better results with fewer generated samples and costs compared to existing process reward models in the mathematical reasoning domain. Code is available at https://github.com/Jiuzhouh/VerifiAgent

📄 PDF Abstract BibTeX arXiv:2504.00406

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingMathematical Reasoning

Similar Papers 제목 키워드 기반

SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning

2025-11-11 · Xuchen Li, Ruitao Wu, Xuanbo Liu, Xukai Wang 외 arxiv

Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified …

Code as Agent Harness

2026-05-18 · Xuying Ning, Katherine Tieu, Dongqi Fu, Tianxin Wei 외 arxiv

Recent large language models (LLMs) have demonstrated strong capabilities in understanding and generating code, from competitive programming to repository-level software engineering. In emerging agentic systems, code is …

From Solving to Verifying: A Unified Objective for Robust Reasoning in LLMs

2025-11-19 · Xiaoxuan Wang, Bo Liu, Song Jiang, Jingzhou Liu 외 arxiv

The reasoning capabilities of large language models (LLMs) have been significantly improved through reinforcement learning (RL). Nevertheless, LLMs still struggle to consistently verify their own reasoning traces. This r…

Reinforcement Learning

SiliconMind-V1: Multi-Agent Distillation and Debug-Reasoning Workflows for Verilog Code Generation

2026-02-10 · Mu-Chi Chen, Yu-Hung Kao, Po-Hsuan Huang, Shao-Chun Ho 외 arxiv

Large language models (LLMs) have recently emerged as a promising approach for automating Verilog code generation; however, existing methods primarily emphasize syntactic correctness and often rely on commercial models o…

Code Generation

FactGuard: Agentic Video Misinformation Detection via Reinforcement Learning

2026-02-26 · Zehao Li, Hongwei Yu, Hao Jiang, Qiang Sheng 외 arxiv

Multimodal large language models (MLLMs) have substantially advanced video misinformation detection through unified multimodal reasoning, but they often rely on fixed-depth inference and place excessive trust in internal…

Reinforcement LearningMultimodal ReasoningDecision Making