paper-with-me

Papers

WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications

2025-05-20 · Xin Li, Mengbing Liu, Li Wei, Jiancheng An, Mérouane Debbah, Chau Yuen

Large Language Models (LLMs) have achieved impressive results across a broad array of tasks, yet their capacity for complex, domain-specific mathematical reasoning-particularly in wireless communications-remains underexplored. In this work, we introduce WirelessMathBench, a novel benchmark specifically designed to evaluate LLMs on mathematical modeling challenges to wireless communications engineering. Our benchmark consists of 587 meticulously curated questions sourced from 40 state-of-the-art research papers, encompassing a diverse spectrum of tasks ranging from basic multiple-choice questions to complex equation completion tasks, including both partial and full completions, all of which rigorously adhere to physical and dimensional constraints. Through extensive experimentation with leading LLMs, we observe that while many models excel in basic recall tasks, their performance degrades significantly when reconstructing partially or fully obscured equations, exposing fundamental limitations in current LLMs. Even DeepSeek-R1, the best performer on our benchmark, achieves an average accuracy of only 38.05%, with a mere 7.83% success rate in full equation completion. By publicly releasing WirelessMathBench along with the evaluation toolkit, we aim to advance the development of more robust, domain-aware LLMs for wireless system analysis and broader engineering applications.

📄 PDF Abstract BibTeX arXiv:2505.14354

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical ReasoningMultiple-choice

Methods 이 논문이 사용한 방법론

FAVOR+ 설명 없음
Performer Performer is a Transformer architecture which can estimate regular…

Similar Papers 제목 키워드 기반

WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning

2025-09-27 · Xin Li, Mengbing Liu, Yiyang Zhu, Wenhe Zhang 외 arxiv

Large language models (LLMs) excel at general mathematical reasoning but fail catastrophically on specialized technical mathematics. In wireless communications, where problems require precise manipulation of information-…

Reinforcement LearningMathematical Reasoning

SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning

2026-08-18 · Guozheng Sun arxiv

Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved the mathematical reasoning capabilities of large language models (LLMs). However, t…

Reinforcement LearningMathematical Reasoning

Mamo: a Mathematical Modeling Benchmark with Solvers

2024-05-21 · Xuhan Huang, Qingning Shen, Yan Hu, Anningzhe Gao 외

Mathematical modeling involves representing real-world phenomena, systems, or problems using mathematical expressions and equations to analyze, understand, and predict their behavior. Given that this process typically re…

Large Multi-Modal Models (LMMs) as Universal Foundation Models for AI-Native Wireless Systems

2024-01-30 · Shengzhe Xu, Christo Kurisummoottil Thomas, Omar Hashash, Nikhil Muralidhar 외

Large language models (LLMs) and foundation models have been recently touted as a game-changer for 6G systems. However, recent efforts on LLMs for wireless networks are limited to a direct application of existing languag…

Mathematical ReasoningRAGRetrieval-augmented Generation

ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges

2025-05-21 · Cheng Qian, Hongyi Du, Hongru Wang, Xiusi Chen 외

Recent progress in large language models (LLMs) has enabled substantial advances in solving mathematical problems. However, existing benchmarks often fail to reflect the complexity of real-world problems, which demand op…

Mathvalid