paper-with-me

홈 › Papers

KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling

2026-07-10 · Peng Kuang, Haibo Jin, Xiaoyu Han, Yanli Wang, Xiaopeng Yuan, Ye Yu, Kaidi Xu, Haohan Wang arxiv

Process Reward Models (PRMs) have been proven to be highly effective in guiding test-time scaling (TTS) methods, which significantly boost the capabilities of LLM-based multi-agent systems. However, existing PRMs are text-based: they re-encode the entire trajectory text from scratch. In long multi-agent rollouts, the scoring cost, growing quadratically with respect to sequence length L, creates a severe computational bottleneck, severely limiting PRMs' application in long-context scenarios. To resolve this, we introduce KV-PRM, a highly efficient process reward model that eliminates the heavy text re-encoding by directly reading the KV cache produced naturally during the LLM's generation phase. By processing a single "verify token" against the pre-existing KV cache, KV-PRM reduces the scoring cost from O(L^2) to O(L). We formally prove that the KV cache contains strictly greater information capacity than text, and is more efficient for downstream reward modeling. Empirically, across the MATH, GSM8K, and AIME benchmarks, KV-PRM matches or strictly outperforms text-PRMs under various TTS methods such as Beam Search, MCTS, and Weighted Voting, with up to a 5,000x reduction in scoring FLOPs, a 37x reduction in latency, and a 34x reduction in per-sequence memory footprint compared to text-based PRMs.

📄 PDF Abstract BibTeX arXiv:2607.09153

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CacheRL:Multi-Turn Tool-Calling Agents via Cached Rollouts and Hybrid Reward

2026-06-12 · Md Amirul Islam, Sumiran Thakur, Huancheng Chen, Su Min Park 외 arxiv

We present CacheRL, a system for training small agent foundation models that achieves 92 percent process accuracy on multi-step tool-calling tasks, approaching GPT-5's 94 percent while requiring 100 times less compute. O…

Reinforcement Learning

HiCache: A Plug-in Scaled-Hermite Upgrade for Taylor-Style Cache-then-Forecast Diffusion Acceleration

2025-08-23 · Liang Feng, Shikang Zheng, Jiacheng Liu, Yuqi Lin 외 arxiv

Diffusion models have achieved remarkable success in content generation but often incur prohibitive computational costs due to iterative sampling. Recent feature caching methods accelerate inference via temporal extrapol…

Video Generation

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

2025-05-26 · Kunjun Li, Zigeng Chen, Cheng-Yen Yang, Jenq-Neng Hwang

Visual Autoregressive (VAR) modeling has garnered significant attention for its innovative next-scale prediction approach, which yields substantial improvements in efficiency, scalability, and zero-shot generalization. N…

Zero-shot Generalization

Neurocache: Efficient Vector Retrieval for Long-range Language Modeling

2024-07-02 · Ali Safaya, Deniz Yuret

This paper introduces Neurocache, an approach to extend the effective context size of large language models (LLMs) using an external vector cache to store its past states. Like recent vector retrieval approaches, Neuroca…

Few-Shot LearningLanguage ModelingLanguage ModellingQuestion Answering+2

Cross-lingual Transfer of Reward Models in Multilingual Alignment

2024-10-23 · Jiwoo Hong, Noah Lee, Rodrigo Martínez-Castaño, César Rodríguez 외

Reinforcement learning with human feedback (RLHF) is shown to largely benefit from precise reward models (RMs). However, recent studies in reward modeling schemes are skewed towards English, limiting the applicability of…

Cross-Lingual TransferInstruction Following