paper-with-me

홈 › Papers

TeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation

2025-12-05 · Anas Ezzakri, Nicola Piovesan, Mohamed Sana, Antonio De Domenico, Fadhel Ayed, Haozhe Zhang arxiv

Large Language Models (LLMs) are increasingly applied to telecom engineering tasks, yet perform poorly on 3GPP specifications. These standards encode much of their technical information in complex tables, but LLM knowledge and interpretation of such tables remain largely unexplored. We introduce TeleTables, a benchmark comprising 2,220 tables from 13 3GPP specifications in four formats and 500 human-verified MCQs spanning direct retrieval to multi-step reasoning. Evaluating 20 open-weight LLMs across non reasoning, multimodal, reasoning, and table specialized architectures reveals two distinct performance bottlenecks. In the closed-book setting, domain knowledge is the primary constraint, with no general-purpose model exceeding 41% accuracy. When the table is provided as context, the best models exceed 90%, but performance degrades systematically with reasoning depth, evidence scope, and structural complexity, with a 32.2pp spread across reasoning skills. Table specialization on non-telecom data provides no consistent benefit, while strong reasoning capabilities remain essential for reliable interpretation of complex technical tables.

📄 PDF Abstract BibTeX arXiv:2601.04202

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TelecomGPT: A Framework to Build Telecom-Specfic Large Language Models

2024-07-12 · Hang Zou, Qiyang Zhao, Yu Tian, Lina Bariah 외

Large Language Models (LLMs) have the potential to revolutionize the Sixth Generation (6G) communication networks. However, current mainstream LLMs generally lack the specialized knowledge in telecom domain. In this pape…

Code GenerationMathOpen-Ended Question AnsweringQuestion Answering

TelcoAgent-Bench: A Multilingual Benchmark for Telecom AI Agents

2026-03-16 · Lina Bariah, Brahim Mefgouda, Farbod Tavakkoli, Enrique Molero 외 arxiv

The integration of large language model (LLM) agents into telecom networks introduces new challenges, related to intent recognition, tool execution, and resolution generation, while taking into consideration different op…

Intent Recognition

TeleQnA: A Benchmark Dataset to Assess Large Language Models Telecommunications Knowledge

2023-10-23 · Ali Maatouk, Fadhel Ayed, Nicola Piovesan, Antonio De Domenico 외

We introduce TeleQnA, the first benchmark dataset designed to evaluate the knowledge of Large Language Models (LLMs) in telecommunications. Comprising 10,000 questions and answers, this dataset draws from diverse sources…

ArticlesQuestion GenerationQuestion-Generation

MM-Telco: Benchmarks and Multimodal Large Language Models for Telecom Applications

2025-11-17 · Anshul Kumar, Gagan Raj Gupta, Manish Rai, Apu Chakraborty 외 arxiv

Large Language Models (LLMs) have emerged as powerful tools for automating complex reasoning and decision-making tasks. In telecommunications, they hold the potential to transform network optimization, automate troublesh…

TeleCom-Bench: How Far Are Large Language Models from Industrial Telecommunication Applications?

2026-05-18 · Jieting Xiao, Yun Lin, Huizhen Qiu, Rui Ma 외 arxiv

While Large Language Models have achieved remarkable integration in various vertical scenarios, their deployment in the telecommunications domain remains exploratory due to the lack of a standardized evaluation framework…

Intent Recognition