paper-with-me

홈 › Papers

AutoVerifier: Residual-Guided Non-Parametric Optimization for Reference-Based Answer Verification

2026-08-26 · Zebei Zhao, Zhihao Shi, Minqi Shi arxiv

Reference-based verifiers are important for evaluating reasoning models and providing accurate outcome rewards in reinforcement learning with verifiable rewards. To improve verification accuracy, prior work has explored rule-based, model-based, and tool-augmented verifiers for checking answer equivalence across diverse answer forms. However, the equivalence of answer forms such as $1+3.14$ and $1+π$ may depend on the question and scoring criterion. We frame such implicit assumptions as verifier inductive biases. To address this challenge, we propose AutoVerifier, a residual-guided non-parametric optimization method that learns these biases from recurring verifier errors. Specifically, AutoVerifier records these biases in rule cards and promotes them to code modules or prompt guidance only after replay validation detects no direct regressions, keeping accepted updates auditable, editable, and reusable. Experiments on four verifier benchmarks demonstrate that AutoVerifier outperforms state-of-the-art verifiers by a large margin.

📄 PDF Abstract BibTeX arXiv:2608.25637

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

AutoVerifier: An Agentic Automated Verification Framework Using Large Language Models

2026-04-03 · Yuntao Du, Minh Dinh, Kaiyuan Zhang, Ninghui Li arxiv

Scientific and Technical Intelligence (S&TI) analysis requires verifying complex technical claims across rapidly growing literature, where existing approaches fail to bridge the verification gap between surface-level acc…

Knowledge Graphs

A Statistical Framework for Alignment with Biased AI Feedback

2026-02-09 · Xintao Xia, Zhiqiu Xia, Linjun Zhang, Zhanrui Cai arxiv

Modern alignment pipelines are increasingly replacing expensive human preference labels with evaluations from large language models (LLM-as-Judge). However, AI labels can be systematically biased compared to high-quality…

Computational Efficiency

Follow the Mean: Reference-Guided Flow Matching

2026-05-11 · Pedro M. P. Curvo, Maksim Zhdanov, Floor Eijkelboom, Jan-Willem van de Meent arxiv

Existing approaches to controllable generation typically rely on fine-tuning, auxiliary networks, or test-time search. We show that flow matching admits a different control interface: adaptation through examples. For det…

LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection

2026-07-28 · Can Wang, Yuhao Wang, Yushe Cao, Canran Xiao 외 arxiv

Recent generative models can produce images with few obvious visual artifacts, weakening detectors and explanations that rely only on surface appearance. We present LaP-Forensics, a multimodal framework that augments RGB…

Multimodal ReasoningDeepFake Detection

ResAD: Normalized Residual Trajectory Modeling for End-to-End Autonomous Driving

2025-10-09 · Zhiyu Zheng, Shaoyu Chen, Haoran Yin, Xinbang Zhang 외 arxiv

End-to-end autonomous driving (E2EAD) systems, which learn to predict future trajectories directly from sensor data, are fundamentally challenged by the inherent spatio-temporal imbalance of trajectory data. This imbalan…

Trajectory ModelingAutonomous Driving