paper-with-me

홈 › Papers

CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks

2026-04-20 · Yijia Shao, Zora Zhiruo Wang, Neel Ahuja, Yicheng Wang, Bowen Liu, Diyi Yang arxiv

AI agents are reshaping the workspace, leading to drastic change of how humans work. Despite the considerable potential of human-agent collaboration both in preserving human agency and generating economic value, this paradigm remains largely absent from occupational task evaluation, hindered by the difficulty of gathering real human data and accounting for inter-human variability. We introduce CollabSkill, a framework for evaluating human-agent collaboration on real-world occupational tasks. CollabSkill pairs real human workers with AI agents on tasks matched to their occupational background, collecting data that capture the complexity of economically valuable tasks and the usage patterns of real workers. To account for inter-human variability, CollabSkill employs a Bayesian skill rating system to disentangle and quantify the skill contributions of both humans and AI agents. Drawing on over 1,500 prompts from 386 working sessions contributed by 93 human workers, our analysis yields insights on two fronts: on the agent side, rankings on CollabSkill diverge meaningfully from those of existing fully autonomous benchmarks where Codex leads, with Claude Code ranking first; on the human side, CollabSkill reveals that practical experience emerges as the primary driver of collaboration skill, with hands-on collaboration meaningfully shifting workers' AI literacy. Together, we hope CollabSkill enables the community to invest in systematic evaluation of human-agent collaboration and spurs development efforts aimed at building AI agents that genuinely augment human workers.

📄 PDF Abstract BibTeX arXiv:2606.09833

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Effective Human-in-the-Loop Assistive AI Agents

2025-07-24 · Filippos Bellos, Yayuan Li, Cary Shu, Ruey Day 외 arxiv

Effective human-AI collaboration for physical task completion has significant potential in both everyday activities and professional domains. AI agents equipped with informative guidance can enhance human performance, bu…

Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration

2024-12-20 · Yijia Shao, Vinay Samuel, Yucheng Jiang, John Yang 외

Recent advancements in language models (LMs) have sparked growing interest in developing LM agents. While fully autonomous agents could excel in many scenarios, numerous use cases inherently require them to collaborate w…

Human Agent Collaboration

Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration

2026-06-04 · Jiaju Chen, Yuxuan Lu, Jiayi Su, Chaoran Chen 외 arxiv

Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly position these agents as human collaborators. Effective collaboration, ho…

AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics

2026-08-28 · Tejas Srinivasan, Shikib Mehri, Nandita Shankar Naik, Anirban Das 외 arxiv

User preferences in user-agent collaboration are rarely static and fully-specified upfront: preferences are formed, revealed, adjusted, and relaxed during interaction. Existing benchmarks for evaluating user-agent collab…

AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?

2026-05-27 · Maharshi Gor, Yoo Yeon Sung, Yu Hou, Eve Fleisig 외 arxiv

AI systems are fallible, and humans can make mistakes in deciding whether to trust AI over their own judgment. Thus, improving human-AI collaboration requires understanding when, why, and how humans decide to rely on AI.…

Question Answering