paper-with-me

홈 › Papers

From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering

2025-12-29 · Tao Dong, Harini Sampath, Ja Young Lee, Sherry Y. Shi, Andrew Macvean arxiv

As Large Language Models (LLMs) evolve from code generators into collaborative partners for software engineers, our methods for evaluation are lagging. Current benchmarks, focused on code correctness, fail to capture the nuanced, interactive behaviors essential for successful human-AI partnership. To bridge this evaluation gap, this paper makes two core contributions. First, we present a foundational taxonomy of desirable agent behaviors for enterprise software engineering, derived from an analysis of 91 sets of user-defined agent rules. This taxonomy defines four key expectations of agent behavior: Adhere to Standards and Processes, Ensure Code Quality and Reliability, Solving Problems Effectively, and Collaborating with the User. Second, recognizing that these expectations are not static, we introduce the Context-Adaptive Behavior (CAB) Framework. This emerging framework reveals how behavioral expectations shift along two empirically-derived axes: the Time Horizon (from immediate needs to future ideals), established through interviews with 15 expert engineers, and the Type of Work (from enterprise production to rapid prototyping, for example), identified through a prompt analysis of a prototyping agent. Together, these contributions offer a human-centered foundation for designing and evaluating the next generation of AI agents, moving the field's focus from the correctness of generated code toward the dynamics of true collaborative intelligence.

📄 PDF Abstract BibTeX arXiv:2512.23844

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Human-Centered Human-AI Collaboration (HCHAC)

2025-05-28 · Qi Gao, Wei Xu, Hanxi Pan, Mowei Shen 외

In the intelligent era, the interaction between humans and intelligent systems fundamentally involves collaboration with autonomous intelligent agents. Human-AI Collaboration (HAC) represents a novel type of human-machin…

Autonomous Vehicles

Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures

2025-12-25 · Hua Shen, Tiffany Knearem, Divy Thakkar, Pat Pataranutaporn 외 arxiv

The rapid integration of generative AI into everyday life underscores the need to move beyond unidirectional alignment models that only adapt AI to human values. This workshop focuses on bidirectional human-AI alignment,…

From Accuracy to Readiness: Metrics and Benchmarks for Human-AI Decision-Making

2026-03-19 · Min Hun Lee arxiv

Artificial intelligence (AI) systems are deployed as collaborators in human decision-making. Yet, evaluation practices focus primarily on model accuracy rather than whether human-AI teams are prepared to collaborate safe…

From Prompt to Product: A Human-Centered Benchmark of Agentic App Generation Systems

2025-12-19 · Marcos Ortiz, Justin Hill, Collin Overbay, Ingrida Semenec 외 arxiv

Agentic AI systems capable of generating full-stack web applications from natural language prompts ("prompt- to-app") represent a significant shift in software development. However, evaluating these systems remains chall…

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

2026-08-04 · Prince Zizhuang Wang, Aojie Yuan, Haiyue Zhang, Xiyang Hu 외 arxiv

Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and c…