paper-with-me

홈 › Papers

StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents

2026-07-24 · Yan Yang, Xiangru Jian, Ziyang Luo, Zirui Zhao, Yutong Dai, Ziji Shi, Hanshu Yan, Jun Hao Liew, Silvio Savarese, Junnan Li hf

Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screenshot is only a lossy rendering of the underlying program state, e.g., the files, application backends, and DOM that hold the task data. Different states can produce the same pixels, while code can inspect and modify that state directly. StateAct is a code-first, multi-agent harness built around this distinction. Its main agent works directly with program state by using code, while a dedicated GUI subagent handles screenshot-and-click interaction on the few subgoals that need it, just 28 of 108 tasks and 1.1% of main-agent steps. The same direct access to program state also supports verification: an independent finish gate double-checks the saved result for structural failures, e.g., output that is missing, unsaved, or written to the wrong path. To stay on track over hundreds of steps, the main agent hands subgoals to fresh subagents, keeping its own context focused. On OSWorld 2.0, StateAct lifts Claude Opus 4.8 from 20.6% to 26.9% on binary success, and from 54.8% to 61.6% on partial success, at ~ 9x lower cost per task than the same model driven by screenshots alone; a code-only variant with no GUI subagent reaches only 45.9% partial, below that screenshot-based baseline's 54.8%. In general, grounding action, verification, and memory in state, what we call state-grounding, shifts the main bottleneck from perception toward reasoning: failures depend more on what the agent thinks than on what it sees.

📄 PDF Abstract BibTeX arXiv:2607.22798

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

StateAct: State Tracking and Reasoning for Acting and Planning with Large Language Models

2024-09-21 · Nikolai Rozanov, Marek Rei

Planning and acting to solve `real' tasks using large language models (LLMs) in interactive environments has become a new frontier for AI methods. While recent advances allowed LLMs to interact with online tools, solve r…

In-Context Learning

Maximum Cohesive Grid of Superpixels for Fast Object Localization

2013-06-01 · CVPR 2013 6 · Liang Li, Wei Feng, Liang Wan, Jiawan Zhang

This paper addresses a challenging problem of regularizing arbitrary superpixels into an optimal grid structure, which may significantly extend current low-level vision algorithms by allowing them to use superpixels (SPs…

ObjectObject LocalizationSuperpixels

Neural Computers

2026-04-07 · Mingchen Zhuge, Changsheng Zhao, Haozhe Liu, Zijian Zhou 외 arxiv

We propose a new frontier: Neural Computers (NCs) that unify computation, memory, and I/O of traditional computers in a learned runtime state. Our long-term goal is the Completely Neural Computer (CNC): the mature, gener…

Programming with Pixels: Computer-Use Meets Software Engineering

2025-02-24 · Pranjal Aggarwal, Sean Welleck

Recent advancements in software engineering (SWE) agents have largely followed a $\textit{tool-based paradigm}$, where agents interact with hand-engineered tool APIs to perform specific tasks. While effective for special…

Visual Grounding

GOGGLES: Automatic Image Labeling with Affinity Coding

2019-03-11 · Nilaksh Das, Sanya Chaba, Renzhi Wu, Sakshi Gandhi 외

Generating large labeled training data is becoming the biggest bottleneck in building and deploying supervised machine learning models. Recently, the data programming paradigm has been proposed to reduce the human cost i…

Few-Shot Learning