paper-with-me

Papers

Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

2026-07-14 · Oleg Solozobov arxiv

Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may sit atop materially different evidence regimes. No vendor-neutral, runnable instrument scores reconstructability as an evaluation-validity metric: whether captured evidence can reconstruct the decision a claim depends on. This paper introduces a property-level reconstructability metric over eight decision-property classes and a cross-harness adapter emitting per-decision Evidence Sufficiency Cards backing a per-run monitor-coverage release check. It specifies a counterfactual-replay intervention protocol, implements its replayability-precondition probe, and defines a claim-evidence overclaim gap. On public and bundled traces, without new model runs, twelve-field sufficiency spans 0.458-0.833 across four inputs sharing a surface reading; replay preconditions are unmet in every scored trace. In a synthetic release-gate pair, the sufficiency gate blocks the raw variant (0.542) and passes the instrumented (0.667). Safety-evaluation claims should travel with their reconstructability vector; a reproducibility package regenerates every reported number.

📄 PDF Abstract BibTeX arXiv:2607.12469

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Electromagnetic-Thermal Analyses of Distributed Antennas Embedded into a Load Bearing Wall

2022-07-13 · Lauri Vähä-Savo, Katsuyuki Haneda, Clemens Icheln, Xiaoshu Lü

The importance of indoor mobile connectivity has increased during the last years, especially during the Covid-19 pandemic. In contrast, new energy-efficient buildings contain structures like low-emissive windows and mult…

Wire twisting stiffness modelling with application in wire race ball bearings. Derivation of analytical formula and Finite Element validation

2024-11-06 · Josu Aguirrebeitia, Inigo Martin, Iker Heras, Mikel Abasolo 외

Since Erich Franke produced the first wire race bearings in 1934, they have not been used profusely until these last years in applications such as computerized tomography, X-ray machines, wheels with direct drive... wher…

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

2026-05-21 · Piercosma Bisconti, Matteo Prandi, Federico Pierucci, Federico Sartore 외 arxiv

Background. Traditional safety benchmarks for language models evaluate generated text: whether a model outputs toxic language, reproduces bias, or follows harmful instructions. When models are deployed as agents, the saf…

Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents

2026-06-22 · Aman Mehta, Anupam Datta arxiv

Long-horizon agents depend on context management: systems compress, summarize, and evict old tokens so tasks can continue beyond finite windows. That is safe only when dropped information is no longer needed or has been …

From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents

2026-06-03 · Yiqi Wang, Jiaqi Zhang, Taotao Cai, Zirui Liu 외 arxiv

Large language model (LLM)-based agents are evolving from passive text generators into autonomous systems capable of planning, tool use, retrieval, memory access, environmental interaction, and multi-agent collaboration.…