paper-with-me

홈 › Papers

Evaluating Language-Model Agents on Realistic Autonomous Tasks

2023-12-18 · Megan Kinniment, Lucas Jun Koba Sato, Haoxing Du, Brian Goodrich, Max Hasin, Lawrence Chan, Luke Harold Miles, Tao R. Lin, Hjalmar Wijk, Joel Burget, Aaron Ho, Elizabeth Barnes, Paul Christiano

In this report, we explore the ability of language model agents to acquire resources, create copies of themselves, and adapt to novel challenges they encounter in the wild. We refer to this cluster of capabilities as "autonomous replication and adaptation" or ARA. We believe that systems capable of ARA could have wide-reaching and hard-to-anticipate consequences, and that measuring and forecasting ARA may be useful for informing measures around security, monitoring, and alignment. Additionally, once a system is capable of ARA, placing bounds on a system's capabilities may become significantly more difficult. We construct four simple example agents that combine language models with tools that allow them to take actions in the world. We then evaluate these agents on 12 tasks relevant to ARA. We find that these language model agents can only complete the easiest tasks from this list, although they make some progress on the more challenging tasks. Unfortunately, these evaluations are not adequate to rule out the possibility that near-future agents will be capable of ARA. In particular, we do not think that these evaluations provide good assurance that the ``next generation'' of language models (e.g. 100x effective compute scaleup on existing models) will not yield agents capable of ARA, unless intermediate evaluations are performed during pretraining. Relatedly, we expect that fine-tuning of the existing models could produce substantially more competent agents, even if the fine-tuning is not directly targeted at ARA.

📄 PDF Abstract BibTeX arXiv:2312.11671

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingmodel

Similar Papers 제목 키워드 기반

VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

2024-01-24 · Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur 외

Autonomous agents capable of planning, reasoning, and executing actions on the web offer a promising avenue for automating computer tasks. However, the majority of existing benchmarks primarily focus on text-based agents…

WebArena: A Realistic Web Environment for Building Autonomous Agents

2023-07-25 · Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 외

With advances in generative AI, there is now potential for autonomous agents to manage daily tasks via natural language commands. However, current agents are primarily created and tested in simplified synthetic environme…

EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments

2025-06-09 · Zefang Liu, Yinzhu Quan

We introduce EconWebArena, a benchmark for evaluating autonomous agents on complex, multimodal economic tasks in realistic web environments. The benchmark comprises 360 curated tasks from 82 authoritative websites spanni…

BenchmarkingNavigateVisual Grounding

ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows

2025-05-26 · Qiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding 외

Large Language Models (LLMs) have extended their impact beyond Natural Language Processing, substantially fostering the development of interdisciplinary research. Recently, various LLM-based agents have been developed to…

Astronomyscientific discovery

Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges

2026-04-21 · Ali Al-Kaswan, Maksim Plotnikov, Maxim Hájek, Roland Vízner 외 arxiv

Large Language Model (LLM) agents are increasingly proposed for autonomous cybersecurity tasks, but their capabilities in realistic offensive settings remain poorly understood. We present DeepRed, an open-source benchmar…