paper-with-me

Papers

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

2026-08-03 · Ziyu Ma, Hailang Huang, Shun Zou, Yong Wang, Shidong Yang, Yiming Hu, Fei Wei, XiangXiang Chu hf

Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses maintain task execution, task state, and completion assessment within a growing context, making the state difficult to track and allowing incorrect self-assessments to propagate into later decisions. We reformulate long-horizon execution as a task-state management problem and propose LongHorizon-Harness, which maintains the task state explicitly outside execution and updates it only with facts independently verified from the environment. Its Manage-Execute-Audit(MEA) loop uses a manager to maintain the task state and determine the next subtask, a fresh-context executor to perform it, and a read-only auditor to verify the resulting environment state before the next round. A lightweight AgentAdapter supports interchangeable model and harness backends without modifying their native agent loops. LongHorizon-Harness improves Qwen~3.7-Plus from 51.8% to 80.7% on WeaveBench, from 69.7% to 77.2% on Terminal-Bench~2.1, and from 2.8% to 8.3% on OSWorld~2.0. It also raises Claude Opus~4.7 from 20.0% to 34.3% on an OSWorld2.0 subset, demonstrating consistent gains across models, harnesses, and interaction domains.

📄 PDF Abstract BibTeX arXiv:2608.01964

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents

2026-02-15 · Zheng Chu, Xiao Wang, Jack Hong, Huiming Fan 외 arxiv

Large language models are transitioning from generalpurpose knowledge engines to realworld problem solvers, yet optimizing them for deep search tasks remains challenging. The central bottleneck lies in the extreme sparsi…

Reinforcement Learning

General Modular Harness for LLM Agents in Multi-Turn Gaming Environments

2025-07-15 · Yuxuan Zhang, Haoyang Yu, Lanxiang Hu, Haojian Jin 외 arxiv

We introduce a modular harness design for LLM agents that composes of perception, memory, and reasoning components, enabling a single LLM or VLM backbone to tackle a wide spectrum of multi turn gaming environments withou…

OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents

2026-08-04 · Jingsheng Zheng, Xinyuan Fang, Jintian Zhang, Zhengke Gui 외 hf

LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goals and constraints ac…

Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents

2026-04-21 · Vasundra Srininvasan arxiv

Long-horizon enterprise agents make high-stakes decisions (loan underwriting, claims adjudication, clinical review, prior authorization) under lossy memory, multi-step reasoning, and binding regulatory constraints. Curre…

Prime Agent: A Self-Improving RLM Harness

2026-08-24 · Seth Karten, Alex L. Zhang, Kevin Thomas, Sebastian Müller 외 arxiv

Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context. Prime Agent is an open-source harness for long-horizon evaluation …