paper-with-me

Papers

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents

2026-07-08 · Xing Zhang, Yanwei Cui, Guanghui Wang, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He arxiv

A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the structural constraint that keeps a growing library from drifting below the no-skill baseline, but its guarantee assumes an unbiased reward, which is false for the LLM judges that reference-free tasks require. We show that a biased judge does not merely add noise; it \emph{silently switches off the curator}. We make this precise with a corrupted-reward analysis, then a behavioral study on a reference-free report-writing testbed with a code-generation cross-check, injecting corruption on top of a deterministic reward to isolate the causal channel. Symmetric noise leaves retirement intact, but \emph{false-pass} bias (failures slipping through as passes) disables contribution-based retirement past a sharp threshold (here a false-pass rate of $0.45$) that no amount of data can cross. Separating genuine retirement from cap-eviction churn shows this \emph{mechanism} failure is universal, holding across domains and failure rates and sparing only near-zero-false-pass, verifier-like graders. The downstream \emph{outcome}, though, is regime-dependent: eval quality degrades only where the same corruption also starves skill synthesis, and otherwise holds steady, so the disabled curator is \emph{silent}, surfacing in no aggregate metric. The contribution is a behavioral safety result, not a performance one. A cheap defect-injection audit then tells an operator, before deployment, which side of the threshold their judge occupies.

📄 PDF Abstract BibTeX arXiv:2607.07436

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models

2026-05-29 · Iosif Tsangko, Andreas Triantafyllopoulos, George Margetis, Ioana Crihana 외 arxiv

Blind and low-vision (BLV) audiences remain underserved by visual art descriptions, particularly across languages and in museum settings where privacy and intellectual-property constraints may favour small on-premise vis…

A Judge Agent Closes the Reliability Gap in AI-Generated Scientific Simulation

2026-03-26 · Chengshuai Yang arxiv

Large language models can generate scientific simulation code, but the generated code silently fails on most non-textbook problems. We show that classical mathematical validation -- well-posedness, convergence, and error…

CuratorKIT : Data Curation and Synthetic Data Generation for LLM Post-Training

2026-06-19 · Soham Bhattacharjee, Karun Sharma, Vinay Kumar Sankarapu, Pratinav Seth arxiv

Data curation is a critical part of post-training pipelines for large language models, yet existing tools often treat ingestion, deduplication, synthetic generation, and quality filtering as separate stages. This fragmen…

Synthetic Data Generation

!Imperio, smolVLA: The Implications of Data Poisoning on Open Source Robotics

2026-07-05 · Stefan Bühler, Mark Schutera arxiv

This work establishes that trigger-word data poisoning of vision language action models is practical, while at the same time the open-source robotics ecosystem holds trust assumptions about community contributions. A few…

Memory-Induced Tool-Drift in LLM Agents

2026-05-24 · Mahavir Dabas, Jihyun Jeong, Ming Jin, Ruoxi Jia arxiv

Modern LLM agents combine long-term memory for personalization with tool-calling interfaces for taking actions in the world -- a combination underpinning contemporary production systems. We study a previously unexamined …