paper-with-me

홈 › Papers

Learning Correct Behavior from Examples: Validating Sequential Execution in Autonomous Agents

2026-05-04 · Reshabh K Sharma, Gaurav Mittal, Yu Hu arxiv

As autonomous agents become increasingly sophisticated, validating their sequential behavior presents a significant challenge. Traditional testing approaches require manual specification, exact sequence matching, or thousands of training examples. We present a novel algorithm that automatically learns correct behavior from just 2-10 passing execution traces and validates new executions against this learned model. Our approach combines dominator analysis from compiler theory with multimodal large language model-powered semantic understanding to identify essential states and handle non-deterministic behavior. The system constructs a generalized ground truth model using Prefix Tree Acceptors, merges traces through multi-tiered equivalence detection, and validates new executions via topological subsequence matching. In controlled experiments, our system achieved high accuracy in detecting product bugs and false successes using only 3 training traces. This approach provides explainable validation results with coverage metrics and works across diverse domains including UI testing, code generation, and robotic processes.

📄 PDF Abstract BibTeX arXiv:2605.03159

Code (0)

등록된 구현이 없습니다.

Tasks

Code Generation

Similar Papers 제목 키워드 기반

SolidCoder: Bridging the Mental-Reality Gap in LLM Code Generation through Concrete Execution

2026-04-20 · Woojin Lee, Jin-Xia Huang arxiv

State-of-the-art code generation frameworks rely on mental simulation, where LLMs internally trace execution to verify correctness. We expose a fundamental limitation: the Mental-Reality Gap -- where models hallucinate e…

Code Generation

AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration

2025-02-13 · Jizhou Chen, Samuel Lee Cong

The integration of tool use into large language models (LLMs) enables agentic systems with real-world impact. In the meantime, unlike standalone LLMs, compromised agents can execute malicious workflows with more conseque…

Replanning Human-Robot Collaborative Tasks with Vision-Language Models via Semantic and Physical Dual-Correction

2026-02-16 · Taichi Kato, Takuya Kiyokawa, Namiko Saito, Kensuke Harada arxiv

Human-Robot Collaboration (HRC) plays an important role in assembly tasks by enabling robots to plan and adjust their motions based on interactive, real-time human instructions. However, such instructions are often lingu…

DriveGPT: Scaling Autoregressive Behavior Models for Driving

2024-12-19 · Xin Huang, Eric M. Wolff, Paul Vernaza, Tung Phan-Minh 외

We present DriveGPT, a scalable behavior model for autonomous driving. We model driving as a sequential decision-making task, and learn a transformer model to predict future agent states as tokens in an autoregressive fa…

Autonomous DrivingDecision MakingSequential Decision Making

An Approach to Checking Correctness for Agentic Systems

2025-08-19 · Thomas J Sheffler arxiv

This paper presents a temporal expression language for monitoring AI agent behavior, enabling systematic error-detection of LLM-based agentic systems that exhibit variable outputs due to stochastic generation processes. …

Prompt Engineering