paper-with-me

홈 › Papers

RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement

2026-05-10 · Will LeVine, Brendan Evers, Sam Saltwick, Abhay Venkatesh arxiv

Iterative self-refinement is a popular inference-time reliability technique, but its effectiveness in code-mode tool use depends heavily on the structure of the feedback signal: unstructured critique helps inconsistently across models, and even revision with real execution feedback improves only modestly. The dominant failures are inter-tool contract violations (wrong output shape, incorrect tool routing, broken argument provenance) that run to completion without raising errors, making runtime feedback insufficient. We introduce RubricRefine, a training-free method for pre-execution contract checking that generates task- and registry-specific rubrics, scores candidate code against explicit contract checks, and iteratively repairs failures before any execution occurs. RubricRefine reaches $0.86$, averaged across seven models, on M3ToolEval with zero execution attempts, improving over prior inference-time baselines at lower latency than rubric-guided reranking. Performance remains flat on the predominantly single-step API-Bank, consistent with the method's reliance on inter-tool contract structure. Results on AppWorld demonstrate that our method maintains an advantage in the multi-turn setting. Because the rubric is derived from the supplied tool documentation, the method's advantage survives incomplete documentation but reverses under incorrect documentation. A rubric-category ablation identifies which rules are load-bearing, and top-bin calibration enables early stopping even where aggregate calibration is poor.

📄 PDF Abstract BibTeX arXiv:2605.09730

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ASA: Backbone-Training-Free Representation Engineering for Tool-Calling Agents

2026-02-04 · Youjin Wang, Run Zhou, Yingjie Ma, Rong Fu 외 arxiv

Adapting LLM agents to domain-specific tool calling remains notably brittle under evolving interfaces. Prompt and schema engineering is easy to deploy but often fragile under distribution shift and strict parsers, while …

parameter-efficient fine-tuning

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines

2026-04-23 · Abel Yagubyan arxiv

Large language model (LLM) agents with tool-calling capabilities are increasingly deployed in production systems, yet a fundamental reliability question remains under-explored: does the same agent behave the same way twi…

PathoSage: Towards Multi-Source Evidence Adjudication in Pathology via Experience-Aware Agentic Workflow

2026-05-18 · Chengyang Zhang, Wenchuan Zhang, Bo Li, Mengran Li 외 arxiv

Recent advances in Multimodal Large Language Models (MLLMs) and agent workflows have shown strong promise for computational pathology, yet reliable patch-level reasoning remains challenging. End-to-end pathology MLLMs of…

Multimodal Reasoning

Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence

2026-08-13 · Haokai Zhang, Yuhang Ding, Yunshu Zhou, Xinze Du 외 arxiv

Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. One lin…

Reinforcement LearningSpatial Reasoning3D ReconstructionDepth Estimation

Judge Reliability Harness: Stress Testing the Reliability of LLM Judges

2026-03-05 · Sunishchal Dev, Andrew Sloan, Joshua Kavner, Nicholas Kong 외 arxiv

We present the Judge Reliability Harness, an open source library for constructing validation suites that test the reliability of LLM judges. As LLM based scoring is widely deployed in AI benchmarks, more tooling is neede…