paper-with-me

홈 › Papers

NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration

2026-04-03 · Ahmed Twabi, Yepeng Ding, Tohru Kondo arxiv

As agentic network management gains popularity, there is a critical need for evaluation frameworks that transcend static, one-shot testing. To address this, we introduce NetAgentBench, a dynamic benchmark that evaluates agent interactions through a Finite State Machine (FSM) formalization guaranteeing determinism, correctness, and bounded execution. This provides the networking landscape with a rigorous foundation to measure complex, multi-turn operational behaviors. Our empirical evaluation of four state-of-the-art LLM agents through diverse network configuration tasks reveals stark deficiencies: while agents can solve basic tasks, they suffer severe exploration meltdowns and coherence collapse during expert-level configurations. Ultimately, NetAgentBench demonstrates that systematically evaluating multi-turn behavioral stability is an indispensable step toward realizing trustworthy, fully autonomous networks.

📄 PDF Abstract BibTeX arXiv:2604.09678

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

2025-05-30 · Tajamul Ashraf, Amal Saqib, Hanan Ghani, Muhra AlMahri 외

Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understanding. However, existing benchmarks typically evaluate agents with fully syntheti…

Autonomous DrivingMathMultimodal ReasoningVisual Reasoning

SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models

2025-11-07 · Jingxuan Xu, Ken Deng, Weihao Li, Songwei Yu 외 arxiv

Evaluating large language models (LLMs) for software engineering has been limited by narrow task coverage, language bias, and insufficient alignment with real-world developer workflows. Existing benchmarks often focus on…

MPCI-Bench: A Benchmark for Multimodal Pairwise Contextual Integrity Evaluation of Language Model Agents

2026-01-13 · Shouju Wang, Haopeng Zhang arxiv

As language-model agents evolve from passive chatbots into proactive assistants that handle personal data, evaluating their adherence to social norms becomes increasingly critical, often through the lens of Contextual In…

FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering

2025-08-07 · Chanyeol Choi, Jihoon Kwon, Alejandro Lopez-Lira, Chaewoon Kim 외 arxiv

Accurate information retrieval (IR) is critical in the financial domain, where investors must identify relevant information from large collections of documents. Traditional IR methods -- whether sparse or dense -- often …

Information RetrievalSemantic SimilarityQuestion Answering

AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications

2026-02-26 · Yujie Zhao, Boqin Yuan, Junbo Huang, Haocheng Yuan 외 arxiv

Large Language Models (LLMs) are increasingly used as autonomous agents in complex, long-horizon applications, where effective memory is critical for sustained performance. Yet existing memory benchmarks are largely dial…