paper-with-me

홈 › Papers

Post-Edit Re-Verification in Simulator-Backed Engineering Agents: A Controlled Comparison of Verification-Cadence Guidance

2026-08-28 · Qingchuan Zhu, Shuyue Tong, Pengju Ren arxiv

Engineering agents that interact with external simulators may need to coordinate design modification with reacquisition of engineering evidence for the modified state. We ask whether first post-edit re-verification changes when explicit verification-cadence guidance is retained versus omitted while verification-relevant state/facts are held constant. Cadence-Guided (CG) retained an instruction to request a new simulation after a substantive modification, whereas Cadence-Omitted (CO) removed that instruction; neither condition used a hard gate. The study therefore measures instruction-conditioned post-edit verification-policy adherence rather than spontaneous recognition that prior evidence has become stale. Using DWSIM as the simulator backend and continuous valve-pressure adjustment, five Alibaba/Qwen models were evaluated on eight synthetic cases; each model-case-condition combination was executed three times via live API calls, yielding 120 evaluation slots per condition. Re-verification was observed in 94/120 CG slots versus 32/120 CO slots; cadence violations occurred in 26/120 versus 87/120; and bounded final success was reached in 95/120 versus 35/120. qwen3.5-35b-a3b showed minimal re-verification (1/24 in CG and 0/24 in CO) and no final success in either condition. Within this bounded protocol, explicit post-edit verification-cadence guidance was associated with more re-verification, fewer cadence violations, and more frequent bounded final success, supporting the treatment of verification cadence as an explicit interaction-protocol component.

📄 PDF Abstract BibTeX arXiv:2608.28147

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Instrumented data for causal scientific machine learning

2026-06-05 · Daniel N. Wilke arxiv

Scientific machine learning is limited less by model size than by the data it is trained on. Observational data records what happened but not why; template synthetic data has a known generating process but only for the s…

Position Paper: Post-Solve Robustness in Decision Engines: Feasible Regions and Smoothness Under Perturbations

2026-03-25 · Yi-Xiang Hu arxiv

Mixed-Integer Linear Programming (MILP) decision engines routinely output nominally optimal plans for high-stakes industrial systems. Yet deployment rarely matches solve-time assumptions: small perturbations in costs, de…

Adversarial Robustness

Reinforcement Learning for Machine Learning Engineering Agents

2025-09-01 · Sherry Yang, Joy He-Yueya, Percy Liang arxiv

Existing agents for solving tasks such as ML engineering rely on prompting powerful language models. As a result, these agents do not improve with more experience. In this paper, we show that agents backed by weaker mode…

Reinforcement Learning

Verification-Guided Context Optimization for Tool Calling via Hierarchical LLMs-as-Editors

2025-12-15 · Henger Li, Shuangjie You, Flavio Di Palo, Yiyue Qian 외 arxiv

Tool calling enables large language models (LLMs) to interact with external environments through tool invocation, providing a practical way to overcome the limitations of pretraining. However, the effectiveness of tool u…

Prompt Engineering

Multilingual Reference Need Assessment System for Wikipedia

2026-03-17 · Aitolkyn Baigutanova, Francisco Navas, Pablo Aragon, Mykola Trokhymovych 외 arxiv

Wikipedia is a critical source of information for millions of users across the Web. It serves as a key resource for large language models, search engines, question-answering systems, and other Web-based applications. In …

Computational Efficiency