paper-with-me

Papers

AgentLens: Visual Analysis for Agent Behaviors in LLM-based Autonomous Systems

2024-02-14 · Jiaying Lu, Bo Pan, Jieyi Chen, Yingchaojie Feng, Jingyuan Hu, Yuchen Peng, Wei Chen

Recently, Large Language Model based Autonomous system(LLMAS) has gained great popularity for its potential to simulate complicated behaviors of human societies. One of its main challenges is to present and analyze the dynamic events evolution of LLMAS. In this work, we present a visualization approach to explore detailed statuses and agents' behavior within LLMAS. We propose a general pipeline that establishes a behavior structure from raw LLMAS execution events, leverages a behavior summarization algorithm to construct a hierarchical summary of the entire structure in terms of time sequence, and a cause trace method to mine the causal relationship between agent behaviors. We then develop AgentLens, a visual analysis system that leverages a hierarchical temporal visualization for illustrating the evolution of LLMAS, and supports users to interactively investigate details and causes of agents' behaviors. Two usage scenarios and a user study demonstrate the effectiveness and usability of our AgentLens.

📄 PDF Abstract BibTeX arXiv:2402.08995

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents

2026-04-22 · Jeonghyeon Kim, Byeongjun Joung, Junwon Lee, Joohyung Lee 외 arxiv

Mobile GUI agents can automate smartphone tasks by interacting directly with app interfaces, but how they should communicate with users during execution remains underexplored. Existing systems rely on two extremes: foreg…

AgentLens: Interpretable Safety Steering via Mechanistic Subspaces for Multi-Turn Coding Agent

2026-06-21 · Weidi Luo, Qiming Zhang, Yihao Quan, Mingyu Jin 외 arxiv

Coding agents based on large language models (LLMs) demonstrate remarkable autonomous capabilities, but they also introduce significant safety and misuse risks during multi-turn interactions with external environments. E…

AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation

2026-05-13 · Priyam Sahoo, Gaurav Mittal, Xiaomin Li, Shengjie Ma 외 arxiv

Evaluation of software engineering (SWE) agents is dominated by a binary signal: whether the final patch passes the tests. This outcome-only view treats a principled solution and a chaotic trial-and-error process as equi…

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

2026-07-07 · Andrey Podivilov, Vadim Lomshakov, Sergey Savin, Matvei Startsev 외 hf

We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience …

Data-driven Analysis for Understanding Team Sports Behaviors

2021-02-15 · Keisuke Fujii

Understanding the principles of real-world biological multi-agent behaviors is a current challenge in various scientific and engineering fields. The rules regarding the real-world biological multi-agent behaviors such as…

counterfactual