paper-with-me

홈 › Papers

Constraint Decay: The Fragility of LLM Agents in Backend Code Generation

2026-05-07 · Francesco Dente, Dario Satriani, Paolo Papotti arxiv

Large Language Model (LLM) agents demonstrate strong performance in autonomous code generation under loose specifications. However, production-grade software requires strict adherence to structural constraints, such as architectural patterns, databases, and object-relational mappings. Existing benchmarks often overlook these non-functional requirements, rewarding functionally correct but structurally arbitrary solutions. We present a systematic study evaluating how well agents handle structural constraints in multi-file backend generation. By fixing a unified API contract across 80 greenfield generation tasks and 20 feature-implementation tasks spanning eight web frameworks, we isolate the effect of structural complexity using a dual evaluation with end-to-end behavioral tests and static verifiers. Our findings reveal a phenomenon of constraint decay: as structural requirements accumulate, agent performance exhibits a substantial decline. Capable configurations lose 30 points on average in assertion pass rates from baseline to fully specified tasks, while some weaker configurations approach zero. Framework sensitivity analysis exposes significant performance disparities: agents succeed in minimal, explicit frameworks (e.g., Flask) but perform substantially worse on average in convention-heavy environments (e.g., FastAPI, Django). Finally, error analysis identifies data-layer defects (e.g., incorrect query composition and ORM runtime violations) as the leading root causes. This work highlights that jointly satisfying functional and structural requirements remains a key open challenge for coding agents.

📄 PDF Abstract BibTeX arXiv:2605.06445

Code (0)

등록된 구현이 없습니다.

Tasks

Code Generation

Similar Papers 제목 키워드 기반

One Goal, Many Commands: Characterizing Denylist Fragility in AI Agents

2026-06-14 · Chuyang Chen, Zhiqiang Lin arxiv

The adoption of AI agents is increasing rapidly. Terminal AI agents, i.e., AI agents that run in terminal environments, are a widely used type of AI agents. Terminal AI agents rely heavily on shell command execution to i…

OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents

2026-08-04 · Jingsheng Zheng, Xinyuan Fang, Jintian Zhang, Zhengke Gui 외 hf

LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goals and constraints ac…

AI-Driven Alpha Decay: Algorithmic Homogenization, Reflexive Signal Erosion, and the Paradox of Intelligent Markets

2026-03-23 · Shuchen Meng, Xupeng Chen arxiv

We show that AI-driven investment strategies are inherently self-defeating at scale. As AI adoption rises, three mutually reinforcing channels -- signal crowding, performative signal erosion, and Red Queen competition --…

Simplicial persistence of financial markets: filtering, generative processes and portfolio risk

2020-09-16 · Jeremy D. Turiel, Paolo Barucca, Tomaso Aste

We introduce simplicial persistence, a measure of time evolution of network motifs in subsequent temporal layers. We observe long memory in the evolution of structures from correlation filtering, with a two regime power …

Time SeriesTime Series Analysis

Omission Constraints Decay While Commission Constraints Persist in Long-Context LLM Agents

2026-04-22 · Yeran Gamage arxiv

LLM agents deployed in production operate under operator-defined behavioral policies (system-prompt instructions such as prohibitions on credential disclosure, data exfiltration, and unauthorized output) that safety eval…