paper-with-me

Papers

MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents

2026-06-15 · Lawrence Keunho Jang, Andrew Keunwoo Jang, Jing Yu Koh, Ruslan Salakhutdinov arxiv

Current benchmarks for computer-use agents evaluate models in impersonal environments. This leaves a gap between evaluation and deployment where personal assistants are expected to work across a user's whole digital life, including their context, historical data, and logged-in accounts. This gap is widest on web tasks, where live web evaluations cannot exercise sites that require logging in or personal information, the kind of site a real personal assistant has to drive. We introduce MyPCBench, which tests computer-use agents as personal assistants on a Linux desktop populated with 17 simulated real-world web applications and a full desktop stack, all seeded for one canonical persona, Michael Scott from The Office. We define 184 tasks in this environment, each inspired by a real request drawn from the OpenClaw community, and benchmark six closed and open-weight models with a uniform computer+bash tool surface. We find that the best model, Claude Opus 4.6, fully solves 55.4\% of the tasks, the only model above 50\%. Model failures cluster on tasks that span many applications and on long trajectories, where personalization stresses an assistant the most. We release the environment, task set, and agent harness at https://mypcbench.com.

📄 PDF Abstract BibTeX arXiv:2606.16748

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

iOSWorld: A Benchmark for Personally Intelligent Phone Agents

2026-06-08 · Lawrence Keunho Jang, Mareks Woodside, Geronimo Carom, Andrew Keunwoo Jang 외 arxiv

A useful phone agent needs to be personally intelligent. It should reason over a user's identity, history, and preferences as they exist on the device, not just follow isolated instructions in an impersonal sandbox. Exis…

WebPII: Benchmarking Visual PII Detection for Computer-Use Agents

2026-03-18 · Nathan Zhao arxiv

Computer use agents create new privacy risks: training data collected from real websites inevitably contains sensitive information, and cloud-hosted inference exposes user screenshots. Detecting personally identifiable i…

Fair Algorithms for Multi-Agent Multi-Armed Bandits

2020-07-13 · NeurIPS 2021 12 · Safwan Hossain, Evi Micha, Nisarg Shah

We propose a multi-agent variant of the classical multi-armed bandit problem, in which there are $N$ agents and $K$ arms, and pulling an arm generates a (possibly different) stochastic reward for each agent. Unlike the c…

FairnessMulti-Armed Bandits

Gapoera: Application Programming Interface for AI Environment of Indonesian Board Game

2021-10-22 · Rian Adam Rajagede, Galang Prihadi Mahardhika

Currently, the development of computer games has shown a tremendous surge. The ease and speed of internet access today have also influenced the development of computer games, especially computer games that are played onl…

Board Games

Artificially intelligent agents in the social and behavioral sciences: A history and outlook

2025-10-07 · Petter Holme, Milena Tsvetkova arxiv

We review the historical development and current trends of artificially intelligent agents (agentic AI) in the social and behavioral sciences: from the first programmable computers, and social simulations soon thereafter…