paper-with-me

홈 › Papers

Coverage, Not Credit: Failure-Credit Routing of Zeroth-Order Perturbation Budgets Does Not Improve On-Pool Sample Efficiency for LLM Agents

2026-08-28 · Yuxu Ge arxiv

Trajectory-level credit assignment can localize which module of a tool-using LLM agent causes failures using only verifiable signals. We ask whether such failure credit should route a fixed zeroth-order/evolution-strategies (ZO/ES) perturbation budget. Across a synthetic environment and frozen Qwen2.5-1.5B/3B and SmolLM2-1.7B agents, three task families, six allocation schemes, a credit-noise sweep, paired seeds, and exact sign-flip tests, we find no statistically detectable improvement over uniform allocation in any on-pool comparison (no gain of at least 2 percentage points). The joint soft-plus-sigma scheme is equivalent to uniform within a +/- 0.02 AUC margin on 1.5B and 3B; concentrating the full budget on the credit argmax is marginally equivalent on 1.5B, where that module is the verified bottleneck, and significantly worse on 3B. Inverse-propensity debiasing does not rescue routing, and misrouting costs up to -0.074 AUC in-house and -0.118 end-to-end on the BFCL-derived family. Across six fixed-step schedules, loss is linear in bottleneck starvation rate (R^2 = 0.94, descriptive), and a preregistered credit-free coverage floor removes detected harm. Matched-budget burst and step-compensating catch-up schedules are consistent with harm arising from insufficient cumulative parameter movement rather than update frequency. Our primary estimand is optimization efficiency on a fixed task pool. On unseen BFCL functions, the study's one exception is that soft routing exceeds uniform on held-out endpoints (+0.047, p = 0.031, n = 6). A plausible but untested reading is that routing-favored caller improvements transfer while uniform's on-pool gains reflect a synthesizer behavior specific to our harness. We report this exception explicitly and document three failure modes that can silently invalidate ZO/ES experiments on frozen LLMs.

📄 PDF Abstract BibTeX arXiv:2608.28011

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SwarmHarness: Skill-Based Task Routing via Decentralized Incentive-Aligned AI Agent Networks

2026-05-27 · Edwin Jose arxiv

Vast quantities of compute (GPU cycles on personal workstations, idle inference servers, and edge devices between jobs) go unused because no incentive-aligned protocol exists for their owners to share them safely and pro…

Online Pseudo-Zeroth-Order Training of Neuromorphic Spiking Neural Networks

2024-07-17 · Mingqing Xiao, Qingyan Meng, Zongpeng Zhang, Di He 외

Brain-inspired neuromorphic computing with spiking neural networks (SNNs) is a promising energy-efficient computational approach. However, successfully training SNNs in a more biologically plausible and neuromorphic-hard…

Credit Risk Assessment Model for UAE Commercial Banks: A Machine Learning Approach

2024-07-02 · Aditya Saxena, Dr Parizad Dungore

Credit ratings are becoming one of the primary references for financial institutions of the country to assess credit risk in order to accurately predict the likelihood of business failure of an individual or an enterpris…

Dimensionality Reduction

Identifying Corporate Credit Risk Sentiments from Financial News

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Credit risk management is one major practice for financial institutions, that helps them measure and understand the inherent risk within their portfolios. Historically, they relied on the assessment of default probabilit…

Managementtext-classificationText Classification

Identifying Corporate Credit Risk Sentiments from Financial News

2022-07-01 · NAACL (ACL) 2022 7 · Noujoud Ahbali, Xinyuan Liu, Albert Nanda, Jamie Stark 외

Credit risk management is one central practice for financial institutions, and such practice helps them measure and understand the inherent risk within their portfolios. Historically, firms relied on the assessment of de…

Managementtext-classificationText Classification