paper-with-me

홈 › Papers

Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents

2025-10-29 · Jiayi Kuang, Yinghui Li, Xin Zhang, Yangning Li, Di Yin, Xing Sun, Ying Shen, Philip S. Yu arxiv

Large language model-based agents show promise for software engineering, but environment configuration remains a bottleneck due to heavy manual effort and scarce large-scale, high-quality datasets. Existing benchmarks assess only end-to-end build/test success, obscuring where and why agents succeed or fail. We introduce the Environment Configuration Diagnosis Benchmark, Enconda-bench, which provides process-level trajectory assessment of fine-grained agent capabilities during environment setup-planning, perception-driven error diagnosis, feedback-driven repair, and action to execute final environment configuration. Our task instances are automatically constructed by injecting realistic README errors and are validated in Docker for scalable, high-quality evaluation. Enconda-bench combines process-level analysis with end-to-end executability to enable capability assessments beyond aggregate success rates. Evaluations across state-of-the-art LLMs and agent frameworks show that while agents can localize errors, they struggle to translate feedback into effective corrections, limiting end-to-end performance. To our knowledge, Enconda-bench is the first framework to provide process-level internal capability assessment for environment configuration, offering actionable insights for improving software engineering agents.

📄 PDF Abstract BibTeX arXiv:2510.25694

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Differentiable Environment-Trajectory Co-Optimization for Safe Multi-Agent Navigation

2026-04-08 · Zhan Gao, Gabriele Fadini, Stelian Coros, Amanda Prorok arxiv

The environment plays a critical role in multi-agent navigation by imposing spatial constraints, rules, and limitations that agents must navigate around. Traditional approaches treat the environment as fixed, without exp…

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design

2026-05-14 · Marilyn Zhang, Tianfeng Chen, Fabián Barzuna, Ankita Rathod 외 arxiv

LLMs are increasingly deployed in autonomous laboratories, under the assumption that their domain priors and reasoning over iterative feedback let them converge on good designs in fewer iterations than feedback-only base…

Reinforcement Learning

Online Vehicle Trajectory Prediction using Policy Anticipation Network and Optimization-based Context Reasoning

2019-03-03 · Wenchao Ding, Shaojie Shen

In this paper, we present an online two-level vehicle trajectory prediction framework for urban autonomous driving where there are complex contextual factors, such as lane geometries, road constructions, traffic regulati…

Autonomous DrivingTrajectory Prediction

C-Free-Uniform: A Map-Conditioned Trajectory Sampler for Model Predictive Path Integral Control

2025-10-19 · Yukang Cao, Rahul Moorthy, O. Goktug Poyrazoglu, Volkan Isler arxiv

Trajectory sampling is a key component of sampling-based control mechanisms. Trajectory samplers rely on control input samplers, which generate control inputs u from a distribution p(u | x) where x is the current state. …

Game Semantics and Linear Logic in the Cognition Process

2018-12-27 · Dmitry Maximov

A description of the environment cognition process by intelligent systems with a fixed set of system goals is suggested. Such a system is represented by the set of its goals only without any models of the system elements…