paper-with-me

홈 › Papers

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents

2026-05-17 · Sirui Hong, Zhijie Liu, Tengfei Li, Wei Tao, Yifan Wu, Chenglin Wu arxiv

Evaluating LLM-generated interactive software requires execution in addition to static analysis. The key difficulty is that correctness is a graph-level reachable property over latent UI state-transition graphs, whereas a GUI evaluator observes only a single execution trajectory. A failed rollout therefore rules out only one realized path, leaving failure attribution ambiguous between evaluator-side execution error and genuine software defect. We present DiagEval, a trajectory-conditioned diagnostic evaluation protocol for post-failure GUI-agent evaluation of interactive software. Rather than blindly retrying from scratch, DiagEval reuses the failed trajectory to choose targeted diagnostic probes and aggregates their outcomes into an internal attribution signal. The latent-graph view motivates the diagnostic problem; DiagEval does not reconstruct the graph or estimate calibrated posterior probabilities. We evaluate DiagEval on WebDevJudge-Unit and RealDevBench across multiple GUI-agent evaluators and LLM backbones. On false-negative cases, DiagEval recovers 45.6-62.1% of failures that were initially misattributed to software defects, outperforming retry-based baselines with 34.4-160.6% relative gains. On the full evaluation sets, this recovery improves accuracy from 69.9% to 78.3% on WebDevJudge-Unit and from 65.0% to 81.6% on RealDevBench. These results suggest that reliable GUI-agent evaluation requires not only stronger execution, but also active failure diagnosis to disambiguate evaluator-side errors from genuine software defects. Our code is available at https://github.com/scutGit/DiagEval.

📄 PDF Abstract BibTeX arXiv:2605.17439

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents

2025-10-29 · Jiayi Kuang, Yinghui Li, Xin Zhang, Yangning Li 외 arxiv

Large language model-based agents show promise for software engineering, but environment configuration remains a bottleneck due to heavy manual effort and scarce large-scale, high-quality datasets. Existing benchmarks as…

Reliable Trajectory Prediction and Uncertainty Quantification with Conditioned Diffusion Models

2024-05-23 · Marion Neumeier, Sebastian Dorn, Michael Botsch, Wolfgang Utschick

This work introduces the conditioned Vehicle Motion Diffusion (cVMD) model, a novel network architecture for highway trajectory prediction using diffusion models. The proposed model ensures the drivability of the predict…

PredictionTrajectory PredictionUncertainty Quantification

HumanDiffusion: A Vision-Based Diffusion Trajectory Planner with Human-Conditioned Goals for Search and Rescue UAV

2026-01-21 · Faryal Batool, Iana Zhura, Valerii Serpiva, Roohan Ahmed Khan 외 arxiv

Reliable human--robot collaboration in emergency scenarios requires autonomous systems that can detect humans, infer navigation goals, and operate safely in dynamic environments. This paper presents HumanDiffusion, a lig…

Uncertainty-Aware Longitudinal Forecasting of Alzheimer's Disease Progression Using Deep Learning

2026-06-23 · Arya Hariharan, Shreyank N Gowda, Anala M R arxiv

Longitudinal modelling of Alzheimer's disease progression is clinically useful only if it can describe not just the most likely next diagnosis, but how a patient may evolve over time and how reliable that forecast is. Mo…

Language-Conditioned Safe Trajectory Generation for Spacecraft Rendezvous

2025-12-09 · Yuji Takubo, Arpit Dwivedi, Sukeerth Ramkumar, Luis A. Pabon 외 arxiv

Reliable real-time trajectory generation is essential for future autonomous spacecraft. While recent progress in nonconvex guidance and control is paving the way for onboard autonomous trajectory optimization, these meth…