A Reproducible Semantic Benchmark for Multivendor DSM-to-CLI Translation
Translating high-level network intents into correct multivendor configurations remains a central challenge in network automation, as syntactically valid outputs may still violate the intended operational state. Despite recent advances in Large Language Models (LLMs), the field still lacks reproducible semantic benchmarks for rigorous cross-vendor evaluation. This paper presents a reproducible DSM-to-CLI semantic benchmark covering five cloud LLMs, three vendors, five representative use cases, and ten repeated runs per experimental cell under fixed judges and an explicit failure taxonomy. Our results show that semantic quality and operational reliability are orthogonal, vendor effects dominate use-case effects, and repeated-run dispersion strongly predicts vote instability, with Huawei VRP exposing failure modes hidden by aggregate metrics. These findings demonstrate that multivendor, repeated-execution semantic benchmarks are essential for scientifically rigorous comparison of LLM-based network configuration systems.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Mammographic density: Comparison of visual assessment with fully automatic calculation on a multivendor dataset
Objectives: To compare breast density (BD) assessment provided by an automated BD evaluator (ABDE) with that provided by a panel of experienced breast radiologists, on a multivendor dataset. Methods: Twenty-one radiologi…
Binary ClassificationClassificationDiagnosticGeneral ClassificationAfroMT: Pretraining Strategies and Reproducible Benchmarks for Translation of 8 African Languages
Reproducible benchmarks are crucial in driving progress of machine translation research. However, existing machine translation benchmarks have been mostly limited to high-resource or well-represented languages. Despite a…
Cross-Lingual TransferData AugmentationMachine TranslationTranslationHybrid Responsible AI-Stochastic Approach for SLA Compliance in Multivendor 6G Networks
The convergence of AI and 6G network automation introduces new challenges in maintaining transparency, fairness, and accountability across multivendor management systems. Although closed-loop AI orchestration improves ad…
Stochastic OptimizationRecovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets
The reliability of multilingual Large Language Model (LLM) evaluation is currently compromised by the inconsistent quality of translated benchmarks. Existing resources often suffer from semantic drift and context loss, w…
Toxicity-Aware Few-Shot Prompting for Low-Resource Singlish Translation
As online communication increasingly incorporates under-represented languages and colloquial dialects, standard translation systems often fail to preserve local slang, code-mixing, and culturally embedded markers of harm…
Semantic SimilarityPrompt Engineering