paper-with-me

Papers

Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain

2026-07-04 · Dmitrii Khizbullin, Zaid Alyafeai, Abdelrahman Eldesokey, Nourah AlSultan, Raghad Alshalan, Bernard Ghanem, David R. Pugh arxiv

We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Telco-GAIA comprises 100 human-verified question-answering tasks, in English and Arabic, that each demand multi-hop reasoning (4.2 hops on average) over three heterogeneous sources: a static website snapshot (HTML, images, and linked PDFs), a synthetic relational SQL database, and external web archives, spanning text, image, and tabular modalities. The benchmark is delivered as a sandboxed Docker environment and scored by normalized exact string matching, making evaluation objective, deterministic, and reproducible over time without any LLM-as-a-Judge. Evaluating a purpose-built reference agent across twelve commercial and open LLMs, we find Telco-GAIA challenging: even the strongest model solves only 71% of tasks; under a moderate cost budget, this falls to about 40%, and the visually grounded categories remain the weakest, where the average backend scores below 30%, leaving substantial headroom in document and image understanding. Telco-GAIA offers a rigorous, reproducible testbed for enterprise agents and a template for constructing closed-domain benchmarks.

📄 PDF Abstract BibTeX arXiv:2607.20510

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TelcoAgent-Bench: A Multilingual Benchmark for Telecom AI Agents

2026-03-16 · Lina Bariah, Brahim Mefgouda, Farbod Tavakkoli, Enrique Molero 외 arxiv

The integration of large language model (LLM) agents into telecom networks introduces new challenges, related to intent recognition, tool execution, and resolution generation, while taking into consideration different op…

Intent Recognition

Telco-oRAG: Optimizing Retrieval-augmented Generation for Telecom Queries via Hybrid Retrieval and Neural Routing

2025-05-17 · Andrei-Laurentiu Bornea, Fadhel Ayed, Antonio De Domenico, Nicola Piovesan 외

Artificial intelligence will be one of the key pillars of the next generation of mobile networks (6G), as it is expected to provide novel added-value services and improve network performance. In this context, large langu…

RAGRetrievalRetrieval-augmented Generation

Telco-RAG: Navigating the Challenges of Retrieval-Augmented Language Models for Telecommunications

2024-04-24 · Andrei-Laurentiu Bornea, Fadhel Ayed, Antonio De Domenico, Nicola Piovesan 외

The application of Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems in the telecommunication domain presents unique challenges, primarily due to the complex nature of telecom standard documen…

RAGRetrievalRetrieval-augmented Generation

TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain

2024-12-20 · Camille Barboule, Viet-Phi Huynh, Adrien Bufort, Yoan Chabot 외

Despite outstanding processes in many tasks, Large Language Models (LLMs) still lack accuracy when dealing with highly technical domains. Especially, telecommunications (telco) is a particularly challenging domain due th…

Benchmarking

MM-Telco: Benchmarks and Multimodal Large Language Models for Telecom Applications

2025-11-17 · Anshul Kumar, Gagan Raj Gupta, Manish Rai, Apu Chakraborty 외 arxiv

Large Language Models (LLMs) have emerged as powerful tools for automating complex reasoning and decision-making tasks. In telecommunications, they hold the potential to transform network optimization, automate troublesh…