paper-with-me

홈 › Papers

Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents

2026-05-22 · Yuandao Cai, Yuzhang Zhu, Liyou Gao, Wensheng Tang, Shengchao Qin arxiv

Long-horizon language agents can make many plausible local tool calls yet fail to persist until a requested count is actually complete. We study this gap as Quantitative Goal Persistence (QGP): whether an agent keeps working until an external verifier confirms enough distinct valid items. PushBench turns this into a benchmark for repository-artifact collection and verifier-backed work units, so repeated work, duplicate submissions, false completion, and progress drift are measured directly rather than hidden behind a final success flag. In matched controller comparisons, a state-tracking retrieval controller reaches 69-78% success while eliminating duplicate submissions, and a backlog-tracking work-unit controller reaches 25-50% success in settings where standard and completion-gated controllers complete no task instances. Black-box frontier-agent evaluations with Claude Code (Sonnet 4.6) and Codex CLI (gpt-5.4) solve many 50-artifact tasks but drop to 3 out of 9 successes per condition at 100 artifacts. The results show that quantitative goals stress a different reliability requirement from local task competence: agents must maintain verified progress and stop only when the requested work is complete.

📄 PDF Abstract BibTeX arXiv:2605.23574

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Free and Fair Economy: A Game of Justice and Inclusion

2021-07-27 · Ghislain H. Demeze-Jouatsa, Roland Pongou, Jean-Baptiste Tondji

Frequent violations of fair principles in real-life settings raise the fundamental question of whether such principles can guarantee the existence of a self-enforcing equilibrium in a free economy. We show that elementar…

EthicsFairness

Measuring Competency of Machine Learning Systems and Enforcing Reliability

2022-12-02 · M. Planer, J. M. Sierchio, for BAE Systems

We explore the impact of environmental conditions on the competency of machine learning agents and how real-time competency assessments improve the reliability of ML agents. We learn a representation of conditions which …

PushupBench: Your VLM is not good at counting pushups

2026-04-25 · Shengzhi Li, Jiarun Chen, Karun Sharma, Jiaqi Su 외 arxiv

Large vision-language models (VLMs) can recognize \textit{what} happens in video but fail to count \textit{how many} times. We introduce \textbf{PushupBench}, 446 long-form clips (avg. 36.7s) for evaluating repetition co…

Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion Models

2023-12-21 · CVPR 2024 1 · Huan Ling, Seung Wook Kim, Antonio Torralba, Sanja Fidler 외

Text-guided diffusion models have revolutionized image and video generation and have also been successfully used for optimization-based 3D object synthesis. Here, we instead focus on the underexplored text-to-4D setting …

Synthetic Data GenerationVideo Generation

Mind Your Language: Learning Visually Grounded Dialog in a Multi-Agent Setting

2018-05-25 · Anonymous

The task of visually grounded dialog involves learning goal-oriented cooperative dialog between autonomous agents who exchange information about a scene through several rounds of questions and answers. We posit that requ…