paper-with-me

홈 › Papers

FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets

2026-08-27 · Kuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige, Kunjan Patel, Kosta Dakic, Suranga Seneviratne arxiv

LLM-based agents are increasingly proposed for network fault diagnosis, but existing benchmarks evaluate them only on accurate tickets and always assume a fault is present, conditions rarely met in practice. We present FaulT-Bench, a benchmark of 200 troubleshooting scenarios across eight network topologies, five reimplemented from public practitioner labs, spanning genuine faults, false fault reports, incorrect device attribution, and incorrect root-cause claims. To isolate how ticket wording affects diagnosis, we further rewrite 72 false-premise tickets into five reporter personas that vary reporter confidence and verifiable detail one factor at a time, holding the network state fixed. Our automated harness deploys each scenario in Kathará, lets agents interact through the NIKA tool interface, and scores free-text diagnoses with an LLM judge across outcome, fix, and reasoning quality. Evaluating SADE, ReAct, and Claude Code, we find all three are near-saturated on accurate tickets and robust to misdirection, yet degrade sharply when the network is healthy and the ticket is wrong, probing until a benign condition can be promoted to a root cause rather than concluding nothing is wrong. Persona rewrites show that how a ticket is written matters more than what it claims: a confidently wrong report is handled about as well as an accurate one, while a vague, underspecified report degrades performance sharply. The three agents also fail differently, from constant over-diagnosis to unanswered runs, at very different cost. These results position FaulT-Bench as a benchmark for developing agentic systems that can reason reliably over the noisy, unreliable tickets of real-world network troubleshooting.

📄 PDF Abstract BibTeX arXiv:2608.27021

Code (0)

등록된 구현이 없습니다.

Tasks

Fault Diagnosis

Similar Papers 제목 키워드 기반

Towards a Playground to Democratize Experimentation and Benchmarking of AI Agents for Network Troubleshooting

2025-07-01 · Zhihao Wang, Alessandro Cornacchia, Franco Galante, Carlo Centofanti 외 arxiv

Recent research has demonstrated the effectiveness of Artificial Intelligence (AI), and more specifically, Large Language Models (LLMs), in supporting network configuration synthesis and automating network diagnosis task…

SADE: Symptom-Aware Diagnostic Escalation for LLM-Based Network Troubleshooting

2026-05-06 · Kuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige, Kosta Dekic 외 arxiv

Large language model (LLM) agents are increasingly applied to network troubleshooting, but root-cause localization on public benchmarks remains well below practical deployment thresholds. We argue this is because existin…

TelcoAgent-Bench: A Multilingual Benchmark for Telecom AI Agents

2026-03-16 · Lina Bariah, Brahim Mefgouda, Farbod Tavakkoli, Enrique Molero 외 arxiv

The integration of large language model (LLM) agents into telecom networks introduces new challenges, related to intent recognition, tool execution, and resolution generation, while taking into consideration different op…

Intent Recognition

Decision-Theoretic Troubleshooting: A Framework for Repair and Experiment

2013-02-13 · John S. Breese, David Heckerman

We develop and extend existing decision-theoretic methods for troubleshooting a nonfunctioning device. Traditionally, diagnosis with Bayesian networks has focused on belief updating---determining the probabilities of var…

Leveraging Multi-Agent System (MAS) and Fine-Tuned Small Language Models (SLMs) for Automated Telecom Network Troubleshooting

2025-11-01 · Chenhua Shi, Bhavika Jalli, Gregor Macdonald, John Zou 외 arxiv

Telecom networks are rapidly growing in scale and complexity, making effective management, operation, and optimization increasingly challenging. Although Artificial Intelligence (AI) has been applied to many telecom task…