paper-with-me

홈 › Papers

Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agents

2026-02-18 · Yun-Shiuan Chuang, Chaitanya Kulkarni, Alec Chiu, Avinash Thangali, Zijie Pan, Shivani Shekhar, Yirou Ge, Yixi Li, Uma Kona, Linsey Pang, Prakhar Mehrotra arxiv

Interactive large language model (LLM) agents operating via multi-turn dialogue and multi-step tool calling are increasingly used in production. Benchmarks for these agents must both reliably compare models and yield on-policy training data. Prior agentic benchmarks, such as tau-bench, tau^2-bench, and AppWorld, rely on fully deterministic backends, which are costly to build and iterate. We propose Proxy State-Based Evaluation, an LLM-driven simulation framework that preserves final state-based evaluation without a deterministic database. Specifically, a scenario specifies the user goal, user/system facts, expected final state, and expected agent behavior, and an LLM state tracker infers a structured proxy state from the full interaction trace. LLM judges then verify goal completion and detect tool/user hallucinations against scenario constraints. Empirically, our benchmark produces stable, model-differentiating rankings across model families and inference-time reasoning efforts, and its on-/off-policy rollouts provide supervision that transfers to unseen scenarios. Careful scenario specification yields near-zero simulator hallucination rates, as supported by ablation studies. The framework also supports sensitivity analyses over user personas. Human-LLM judge agreement exceeds 90%, indicating reliable automated evaluation. Overall, proxy state-based evaluation offers a practical, scalable alternative to deterministic agentic benchmarks for industrial LLM agents.

📄 PDF Abstract BibTeX arXiv:2602.16246

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

2026-07-26 · Qinsi Wang, Jing Shi, Huazheng Wang, Kun Wan 외 hf

Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited…

Self-Supervised LearningReinforcement LearningMathematical ReasoningText Summarization

What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards

2025-11-29 · Minh-Quan Le, Yuanzhi Zhu, Vicky Kalogeiton, Dimitris Samaras arxiv

Recent video diffusion models can synthesize visually compelling clips, yet often violate basic physical laws-objects float, accelerations drift, and collisions behave inconsistently-revealing a persistent gap between vi…

Video Generation

Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards

2025-09-29 · Chenhui Xu, Fuxun Yu, Michael J. Bianco, Jacob Kovarskiy 외 arxiv

Training robust reasoning vision-language models (VLMs) in rare domains (such as geospatial) is fundamentally constrained by supervision scarcity. While raw geospatial imagery is abundant, the amount of task-direct super…

Reinforcement Learning

Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations

2026-07-01 · Mehul Damani, Isha Puri, Idan Shenfeld, Jacob Andreas arxiv

RL with verifiable rewards (RLVR) has emerged as a powerful paradigm for training LMs on tasks with well-defined success metrics, such as code generation and mathematical reasoning. However, current RLVR methods optimize…

Mathematical ReasoningStory GenerationCode Generation

How to Evaluate Reward Models for RLHF

2024-10-18 · Evan Frick, Tianle Li, Connor Chen, Wei-Lin Chiang 외

We introduce a new benchmark for reward models that quantifies their ability to produce strong language models through RLHF (Reinforcement Learning from Human Feedback). The gold-standard approach is to run a full RLHF t…